Jinjian Wu

dblp:01/8056 · DBLP profile ↗
← Back
151ranked-venue papers
30as first author
97since 2021 · last 2026
0000-0001-7501-0009ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 106 · 24 first-author · 62 since 2021Artificial intelligence and machine learning · 37 · 32 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 10 since 2021Systems, architecture and hardware · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MCIB: Multi-Modal Complementary Information Bottleneck for Hyperspectral and LiDAR Classification
abstract
The effective fusion of multi-modal remote sensing images, particularly hyperspectral imagery (HSI) and light detection and ranging (LiDAR) data, is pivotal for accurate land use and land cover (LULC) classification. However, this process is hindered by two inherent challenges: pervasive data redundancy and the underutilization of cross-modal complementarity, largely due to the lack of a unifying theoretical framework. To address these limitations, we propose the multi-modal complementary information bottleneck (MCIB) framework, which extends the IB principle to learn compact, sufficient, and complementary representations for multi-modal scenes. From a theoretical perspective, we formalize the MCIB objective and introduce structured priors to derive tractable information-theoretic bounds, providing a principled and computationally feasible approach to reduce redundancy and enhance complementarity simultaneously. Building on the obtained theoretical insights, we design an end-to-end variational optimization strategy with a novel supervised conditional InfoNCE (SCInfoNCE). Efficiently reusing existing model components, this new supervised contrastive method optimizes the conditional mutual information terms crucial for synergy. Extensive experiments on benchmark HSI-LiDAR datasets demonstrate superior classification performance of MCIB. This work not only fills a theoretical gap in multi-modal representation learning, but offers a robust and principled solution for LULC classification using complex heterogeneous remote sensing images.
Hao Zhu 0009, Bo Yang 0047, Changzhe Jiao, Jie Feng 0003, Jinjian Wu
IEEE Trans. Image Process.6
2026 Exploring Cross-Modal Mutual Prompt Learning for Video Quality Assessment
abstract
Enhancing video quality assessment (VQA) through semantic information integration is a critical research focus. Recent research has employed the Contrastive Language-Image Pre-training (CLIP) model as a foundation to improve semantic perception. However, the image-text alignment inherent in these pre-trained Vision-Language (VL) models frequently results in suboptimal VQA performance. While prompt engineering has recently targeted the language component to address this alignment issue, the unique insights resided in visual analysis is still overlooked for further advancing VQA tasks. Additionally, seeking a trade-off between quality separability and domain invariance in VQA remains largely unresolved within the VL paradigm. In this paper, we introduce a novel cross-modal prompt-based approach to tackle these challenges. Specifically, we propose learnable prompts within the vision branch to foster synergy between visual and language modalities through a language-to-vision coupling function. The multi-view backbone is then carefully crafted with content enhancement and distortion-aware temporal modulation to ensure quality separability. The language prompts, derived from visual representations, are further supported by adaptive weighting mechanisms to optimize the balance between quality separability and domain invariance. Experimental results demonstrate the effectiveness of our proposed method over leading VQA models, showing significant improvements in generalization across diverse datasets. The source code for this work is publicly available athttps://github.com/cpf0079/CM2PL.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Jiebin Yan, Vinit Jakhetiya, Aladine Chetouani
IEEE Trans. Multim.3
2026 Geo-SelfSSC: Integrating Dense Geometric Priors for Enhanced Self-Supervised Semantic Scene Completion
abstract
Accurate 3D scene understanding is vital for applications like autonomous driving and robotics. However, existing voxel-based methods struggle with the reliance on large-scale labeled data, inherent voxel-pixel misalignments, and high computational costs. Recent self-supervised Semantic Scene Completion (SSC) methods using Neural Radiance Fields (NeRF) reduce 3D annotation needs but assume Lambertian surfaces and rely on photometric consistency, which fails in real-world scenes with non-Lambertian effects and sparse camera coverage. In this paper, we introduce Geo-SelfSSC, a self-supervised framework that leverages temporal information from consecutive frames and slight variations in camera poses for supervision, while integrating dense geometric cues to enhance the reconstruction quality and efficiency of neural implicit models. Specifically, Geo-SelfSSC leverages depth priors along with complementary semi-local and local geometric supervision to facilitate efficient sampling, while ensuring effective complementarity between photometric and geometric cues. Our method yields significant improvements in reconstructing reflective and under-observed regions where conventional photometric-based strategies struggle. Comprehensive experiments on challenging tasks demonstrate that Geo-SelfSSC not only achieves strong results in semantic scene completion but also establishes competitive performance in geometry-only 3D occupancy prediction and monocular depth estimation. Our code and models are available at: https://github.com/Xidian AIGroup190726/GeoSelfSSC.
Hao Zhu 0009, Pute Guo, Longsheng Qu, Jinjian Wu
IEEE Trans. Multim.5
2025 Asymmetric Hierarchical Difference-aware Interaction Network for Event-guided Motion Deblurring
abstract
Event cameras are bio-inspired sensors that are capable of capturing motion information with high temporal resolution, which show potential in aiding image motion deblurring recently. Most existing methods indiscriminately handle feature fusion of two modalities with symmetric unidirectional/bidirectional interactions at different-level layers in feature encoder, while ignoring the different dependencies between cross-modal hierarchical features. To tackle these limitations, we propose a novel Asymmetric Hierarchical Difference-aware Interaction Network (AHDINet) for event-based motion deblurring, which explores the complementarity of two modalities with differential dependence modeling of cross-modal hierarchical features. Thereby, an event-assisted edge complement module is designed to leverage event modality to enhance the edge details of the image features in low-level encoder stage, and an image-assisted semantic complement module is developed to transfer contextual semantics of image features to event branch in high-level encoder stage. Benefiting from the proposed differentiated interaction mode, the respective advantages of image and event modalities are fully exploited. Extensive experiments on both synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance.
Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
AAAI2
2025 Simultaneous Denoising and Compression for DVS with Partitioned Cache-Like Spatiotemporal Filter
abstract
Dynamic vision sensor (DVS) is a novel neuromorphic imaging device that asynchronously generates event data corresponding to changes in light intensity at each pixel. However, the differential imaging paradigm of DVS renders it highly sensitive to background noise. Additionally, the substantial volume of event data produced in a very short time presents significant challenges for data transmission and processing. In this work, we present a novel spatiotemporal filter design, named PCLF, to achieve simultaneous denoising and compression for the first time. The PCLF employs a hierarchical memory structure that utilizes symmetric multi-bank cache-like row and column memories to store event data from a partitioned pixel array, which exhibits low memory complexity of O(m + n) for an$\mathrm{m}\times \mathrm{n}$DVS. Furthermore, we propose a probability-based criterion to effectively control the compression ratio. We have implemented our design on an FPGA, demonstrating capabilities for real-time operation$(\leq 60\ \text{ns})$and low power consumption$(< 200\text{mW})$. Extensive experiments conducted on real-world DVS data across various tasks indicate that our design enables a reduction of event data by 30% to 68%, while maintaining or even enhancing the performance of the tasks.
Qinghang Zhao, Yixi Ji, Jinjian Wu, Guangming Shi
DATE4
2025 SNNPTrack: Spiking Neural Network Based Prompt for High-Accuracy RGBE Tracking
abstract
RGBE object tracking is an emerging field that integrates RGB frames and event data to achieve more robust tracking results, particularly in challenging scenarios. However, existing methodologies predominantly focus on transforming sparse event streams into event frames, thereby neglecting the potential of rich temporal information. To address this limitation, we introduce Spiking Neural Network-based Prompt Tracking (SNNPTrack), a hybrid framework designed for temporally adaptive RGBE tracking, aimed at achieving high-accuracy tracking performance. SNNPTrack includes a Leaky Integrate-and-Fire (LIF)-based Spiking Neural Network (SNN) module for temporal feature extraction, a Cross-Modality Fusion module for feature fusion across both domains, and a pre-trained RGB-based transformer model for dual-modal feature extraction and interaction. Extensive experimental evaluations demonstrate that, with only a modest increase in the number of parameters, our SNNPTrack framework surpasses state-of-the-art methods on the VisEvent, FE108, and COESOT datasets, highlighting its potential as a promising solution for real-world tracking applications.
Yixi Ji, Qinghang Zhao, Yuping Liang, Jinjian Wu
ICASSP4
2025 Towards Syn-to-Real IQA: A Novel Perspective on Reshaping Synthetic Data Distributions
abstract
Blind Image Quality Assessment (BIQA) has advanced significantly through deep learning, but the scarcity of large-scale labeled datasets remains a challenge. While synthetic data offers a promising solution, models trained on existing synthetic datasets often show limited generalization ability. In this work, we make a key observation that representations learned from synthetic datasets often exhibit a discrete and clustered pattern that hinders regression performance: features of high-quality images cluster around reference images, while those of low-quality images cluster based on distortion types. Our analysis reveals that this issue stems from the distribution of synthetic data rather than model architecture. Consequently, we introduce a novel framework SynDR-IQA, which reshapes synthetic data distribution to enhance BIQA generalization. Based on theoretical derivations of sample diversity and redundancy's impact on generalization error, SynDR-IQA employs two strategies: distribution-aware diverse content upsampling, which enhances visual diversity while preserving content distribution, and density-aware redundant cluster downsampling, which balances samples by reducing the density of densely clustered areas. Extensive experiments across three cross-dataset settings (synthetic-to-authentic, synthetic-to-algorithmic, and synthetic-to-synthetic) demonstrate the effectiveness of our method. The code is available at https://github.com/Li-aobo/SynDR-IQA.
Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong
NeurIPS2
2025 Attribute-guided feature fusion network with knowledge-inspired attention mechanism for multi-source remote sensing classification
Changzhe Jiao, Bo Yang 0047, Hao Zhu 0009, Jinjian Wu
Neural Networks5
2025 Fast Window-Based Event Denoising With Spatiotemporal Correlation Enhancement
abstract
Previous deep learning-based event denoising methods mostly suffer from poor interpretability and difficulty in real-time processing due to their complex architecture designs. In this paper, we propose window-based event denoising, which simultaneously deals with a stack of events while existing element-based denoising focuses on one event each time. Besides, we give the theoretical analysis based on probability distributions in both temporal and spatial domains to improve interpretability. In temporal domain, we use timestamp deviations between processing events and central event to judge the temporal correlation and filter out temporal-irrelevant events. In spatial domain, we choose maximum a posteriori (MAP) to discriminate real-world event and noise and use the learned convolutional sparse coding to optimize the objective function. Based on the theoretical analysis, we build Temporal Window (TW) module and Soft Spatial Feature Embedding (SSFE) module to process temporal and spatial information separately, and construct a novel multi-scale window-based event denoising network, named WedNet. The high denoising accuracy and fast running speed of our WedNet enables us to achieve real-time denoising in complex scenes. Extensive experimental results verify the effectiveness and robustness of our WedNet. Our algorithm can remove event noise effectively and efficiently and improve the performance of downstream tasks.
Huachen Fang, Jinjian Wu, Qibin Hou, Weisheng Dong, Guangming Shi
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Self-Supervised Learning of LiDAR 3D Point Clouds via 2D-3D Neural Calibration
abstract
This paper introduces a novel self-supervised learning framework for enhancing 3D perception in autonomous driving scenes. Specifically, our approach, namely NCLR, focuses on 2D-3D neural calibration, a novel pretext task that estimates the rigid pose aligning camera and LiDAR coordinate systems. First, we propose the learnable transformation alignment to bridge the domain gap between image and point cloud data, converting features into a unified representation space for effective comparison and matching. Second, we identify the overlapping area between the image and point cloud with the fused features. Third, we establish dense 2D-3D correspondences to estimate the rigid pose. The framework not only learns fine-grained matching from points to pixels but also achieves alignment of the image and point cloud at a holistic level, understanding the LiDAR-to-camera extrinsic parameters. We demonstrate the efficacy of NCLR by applying the pre-trained backbone to downstream tasks, such as LiDAR-based 3D semantic segmentation, object detection, and panoptic segmentation. Comprehensive experiments on various datasets illustrate the superiority of NCLR over existing self-supervised methods. The results confirm that joint learning from different modalities significantly enhances the network's understanding abilities and effectiveness of learned representation.
Yifan Zhang 0036, Junhui Hou, Jinjian Wu, Yixuan Yuan, Guangming Shi
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Source-free collaborative domain adaptation via multi-perspective feature enrichment for functional MRI analysis
Yuqi Fang, Jinjian Wu, Qianqian Wang 0004, Shijun Qiu, Andrea Bozoki, Mingxia Liu 0001
Pattern Recognit.2
2025 An event-based motion scene feature extraction framework
Zhaoxin Liu, Jinjian Wu, Guangming Shi, Wen Yang 0008, Jupo Ma
Pattern Recognit.2
2025 Brain anatomy prior modeling to forecast clinical progression of cognitive impairment with structural MRI
Jinjian Wu, Li Wang 0026, David C. Steffens, Shijun Qiu, Guy G. Potter, Mingxia Liu 0001
Pattern Recognit.2
2025 Towards Explainable Image Aesthetics Assessment With Attribute-Oriented Critiques Generation
abstract
Compared with the unimodal image aesthetics assessment (IAA), multimodal IAA has demonstrated superior performance. This indicates that the critiques could provide rich aesthetics-aware semantic information, which also enhance the explainability of IAA models. However, images are not always accompanied with critiques in real-world situation, rendering multimodal IAA inapplicable in most cases. Therefore, it would be interesting to investigate whether we can generate aesthetic critiques to facilitate image aesthetic representation learning and enhance model explainability. Motivated by these facts, this paper presents an attribute-oriented Critiques Generation framework for explainable IAA, dubbed CG-IAA, which consists of three major components, i.e., Vision-Language Aesthetic Pretraining (VLAP), Multi-Attribute Experts Learning (MAEL) and Multimodal Aesthetics Prediction (MAP). Specifically, the vanilla CLIP is first finetuned on a multimodal IAA database. Considering that the aesthetic critiques typically consist of multiple attributes, a new multimodal IAA database which contains over 1 million critiques with up to four aesthetic attributes is constructed with the language model-based knowledge transfer. Then, CLIP-based multi-attribute experts are trained based on this database. Finally, the pretrained experts are utilized to generate aesthetic critiques for assisting unimodal image aesthetics prediction. Extensive experiments have been done on four popular IAA databases, and the results demonstrate the advantage of CG-IAA over the state-of-the-arts. Furthermore, CG-IAA features better explainability and generalization with the assistance of generated critiques. The source code is available athttps://github.com/sxfly99/CG-IAA.
Leida Li, Xiangfei Sheng, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong
IEEE Trans. Circuits Syst. Video Technol.4
2025 Scene Prior Constrained Self-Paced Learning for Unsupervised Satellite Video Vehicle Detection
abstract
Recently, deep learning has significantly advanced the satellite object detection. However, the effectiveness of these methods heavily relies on abundant and accurate annotations, which are extremely labor-intensive for satellite videos. Meanwhile, the robustness of traditional difference-based methods is limited by the hand-craft feature from the satellite videos with low-resolution and frame misalignment. To address this problem, an unsupervised deep satellite video vehicle detection framework based on scene prior constrained and self-paced learning (S-SPL) is proposed in this paper. S-SPL obtains the initial pseudo label by the difference-based methods, and employs a deep learning-based detector and refiner to detect objects and update labels respectively. In the train phase, to alleviate the deviation of feature expression caused by noise samples, a novel cooperative self-paced learning scheme is designed to improve the label quality and model accuracy in an alternating optimization manner. Furthermore, considering the semantic relationship between the scene and the object distribution, multi-cue prior knowledge is introduced to provide scene-level constraints, with which samples in high-confidence scenes are emphasized to improve the self-paced learning process. The experimental results on Jilin-1 and SkySat satellite videos demonstrate the superiority of S-SPL.
Yuping Liang, Guangming Shi, Jinjian Wu
IEEE Trans. Circuits Syst. Video Technol.3
2025 Scene-Modulated High-Order Statistical Representation Learning for No-Reference Super-Resolution Image Quality Assessment
abstract
With the rapid development of single image super-resolution (SR) technology, there is an urgent need to develop a fair no reference Super-Resolution image Quality Assessment (SRQA) method. Existing no reference SRQA methods primarily concentrate on SR artifacts including structural distortion and texture distortion by extracting spatial features, but ignore the inductive bias of Deep Neural Network (DNN)-based SR models. As a result, they function effectively for interpolation-based and dictionary-based algorithms, but struggle to perform as effectively with DNN-based SR algorithms. We found that the visual content generated by DNN-based SR models under different inductive biases often carries a content-invariant model-specific style, which can be captured by the correlations between hierarchical representation channels. To that end, we propose a novel Scene-modulated High-order Statistical Representation network (SmHSR) built on a multi-scale over-complete transformation. We quantify the perceptual quality of SR images as the shift of high-order statistical properties in their multi-scale over-complete representation, where intra-channel statistics are used to capture spatial correlations and inter-channel statistics are used to capture the inductive bias of SR models. In addition, the scene information implicit in the deep over-complete representation is used to modulate the high-order statistical properties, which simulates the top-down regulation of cognition on perception. Under the modulation of scene information, SmHSR can learn more sophisticated scene-aware statistical representation. The MultiLayer Perceptron (MLP) is used to map the high-order statistical representation to an overall quality. We test our method on multiple SR image quality databases. Experimental results show that our method outperforms the state-of-the-art SRQA methods.
Yongwei Mao, Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong
IEEE Trans. Circuits Syst. Video Technol.2
2025 Defect Detection in Remote Sensing Satellite Images: A New Dataset and Algorithm
abstract
Satellite observation is an important way to understand the earth. However, due to the problems such as satellite aging, cloud obstruction, and other interferences during the imaging and transmission process, remote sensing images inevitably produce various defects. Hence, it is necessary to quickly detect defects to calibrate the imaging system and avoid the waste of satellite resource. Current researches on defect detection in remote sensing images are not comprehensive, which only focus on partial defect categories, such as cloud and stripe. To this end, we construct the first large-scale High-resolution Remote Sensing image Defect detection dataset (HRSD). The proposed dataset contains more than 1.2 million manually annotated patches from eight different satellites, covering various common defect categories and including multiple image modalities (i.e., panchromatic and multispectral). The dataset also has rich diversity which covers different landforms in multiple regions. Furthermore, to realize the detection of multiple defect categories simultaneously, we design a feature aggregation graph network (FAGN) based on the position correlation and semantic similarity among image patches, which fully utilizes the distribution characteristics of defects to achieve accurate defect detection. Extensive experiments on the HRSD dataset demonstrated the effectiveness of FAGN. We will release the HRSD dataset and FAGN model later.
Hengchao Hu, Jupo Ma, Qi Wang 0053, Yuanshi Zheng, Jinjian Wu
IEEE Trans. Geosci. Remote. Sens.6
2025 Oriented Vehicle Joint Detection and Tracking in Satellite Video via Identifier-Free Point Supervision
abstract
Oriented vehicle detection and tracking play a crucial role in various real-world applications. Yet, existing advanced models heavily rely on abundant and accurate oriented bounding box and tracking identifier annotations, which are extremely labor-intensive for satellite videos. In this paper, we endeavor to employ identifier-free point annotation to achieve the competitive performance while minimizing annotation costs. Specifically, each instance across video frames are labeled by single points, without providing its instance identifier. Building upon this setting, we introduce an oriented vehicle joint detection and tracking framework for satellite video, focusing on enhancing model performance by carefully-designed sample acquisition and robust learning processes. Firstly, we leverage temporal and visual information to generate sequence-aligned pseudo-labels and visually-aligned synthetic objects, which complement each other during training by providing both exact appearance and annotation information. Secondly, a novel spatio-temporal consistency metric is developed to assess sample quality, which is then incorporated into a curriculum learning schedule. This strategy facilitates a gradual learning progression from high-quality data to low-quality or noisy examples, thereyby minimizing interference from potentially misleading samples. Finally, an end-to-end oriented object joint detection and tracking network is constructed to enable effective oriented vehicle dynamic analysis. Extensive ablation and experimental results on two satellite video datasets demonstrate the superiority of our proposed method.
Yuping Liang, Jinjian Wu, Junpeng Zhang 0002, Yuxuan Chang, Jie Feng 0003, Guangming Shi
IEEE Trans. Geosci. Remote. Sens.2
2025 Proxy-Enhanced Prototype Memory Network for Weakly Supervised Hyperspectral Target Detection
abstract
Hyperspectral target detection (HTD) holds significant promise in numerous earth vision applications, yet it encounters challenges in acquiring high-quality prior target signatures, capturing target spectral variability, and dealing with sample imbalance. To address these issues, we propose a weakly supervised solution, the Proxy-Enhanced Prototype Memory Network (PE-PMN), for HTD tasks. It relies solely on region-level weakly labeled data, eliminating the need for strict prior target knowledge (e.g., handcrafted target signatures or pixel-level annotations). To fully describe target variations and background diversity, two memory prototype networks are introduced to extract, store, and retrieve prototypes of targets and backgrounds, providing comprehensive spectral information. Additionally, a proxy-based enhancement approach is incorporated to enrich the prototypes in the memory banks and boost the separation between target and background features. To mitigate sample imbalance in PE-PMN, we develop the Bag Mix-Up (BMU) strategy based on the Unconstrained Linear Mixture Model (ULMM) to construct a sufficient training dataset. Experimental results on three simulated datasets and three real datasets demonstrate that the proposed PE-PMN significantly outperforms other competitive weakly supervised HTD methods.
Bo Yang 0047, Jinjian Wu, Changzhe Jiao
IEEE Trans. Geosci. Remote. Sens.2
2025 HiCAL: Hierarchical Consistency-Based Active Learning for Drone-View Object Detection
abstract
The recent years have witnessed the great progress of drone-view object detection in both economic and military applications. Generally, the good performance of drone-view object detection requires a large amount of annotated data, which has imposed significant demands on human and material resources. To optimize the labelling expenses, previous work has introduced active learning to select the most valuable samples for annotation, and balances the annotation cost and model performance. However, existing active learning methods are primarily controlled by the “absolute” prediction of the model (e.g., the predicted categories for diversity, and the classification confidence for uncertainty). It would be highly misleading when the model outputs wrong prediction but with high confidence. This confident misleading is more severe in drone-view object detection as the targets are captured with varied viewpoints, illumination conditions, and possible occlusion. In this paper, we refresh the active learning with Perturbation Consistency Test (PCT), which transforms the absolute prediction into the relative error to address the situation where the absolute prediction is unreliable. The basic idea is to test the prediction consistency when the input samples are with/without perturbation, and regards the inconsistency as a measurement of the model’s resilience to guide the active selection. To this end, a Hierarchical Consistency-based Active Learning (HiCAL) is built, which constructs adversarially pair-wise inputs with hierarchical perturbation. The samples are perturbed with multi-granularity (i.e., pixel level, feature level, and object level) and afterwards, the entropy difference of the paired outputs before/after perturbation is calculated as the measurement. The samples with high difference are selected to follow a standard active learning loop. Experimental results show that HiCAL can achieve superior performance in different datasets and is easy to adapt to various types of object detectors. The code will be available on: https://github.com/zstar1003/HiCAL.
Yongxu Liu 0001, Qinghang Zhao, Jinjian Wu
IEEE Trans. Geosci. Remote. Sens.6
2025 CrossEI: Boosting Motion-Oriented Object Tracking With an Event Camera
abstract
With the differential sensitivity and high time resolution, event cameras can record detailed motion clues, which form a complementary advantage with frame-based cameras to enhance the object tracking, especially in challenging dynamic scenes. However, how to better match heterogeneous event-image data and exploit rich complementary cues from them still remains an open issue. In this paper, we align event-image modalities by proposing a motion adaptive event sampling method, and we revisit the cross-complementarities of event-image data to design a bidirectional-enhanced fusion framework. Specifically, this sampling strategy can adapt to different dynamic scenes and integrate aligned event-image pairs. Besides, we design an image-guided motion estimation unit for extracting explicit instance-level motions, aiming at refining the uncertain event clues to distinguish primary objects and background. Then, a semantic modulation module is devised to utilize the enhanced object motion to modify the image features. Coupled with these two modules, this framework learns both the high motion sensitivity of events and the full texture of images to achieve more accurate and robust tracking. The proposed method is easily embedded in existing tracking pipelines, and trained end-to-end. We evaluate it on four large benchmarks, i.e. FE108, VisEvent, FE240hz and CoeSot. Extensive experiments demonstrate our method achieves state-of-the-art performance, and large improvements are pointed as contributions by our sampling strategy and fusion concept.
Zhiwen Chen 0002, Jinjian Wu, Weisheng Dong, Leida Li, Guangming Shi
IEEE Trans. Image Process.2
2025 Modeling State Shifting via Local-Global Distillation for Event-Frame Gaze Tracking
abstract
This paper tackles the problem of passive gaze estimation using both event and frame (or 2D image) data. Considering the inherently different physiological structures, it is intractable to accurately estimate gaze purely based on a given state. Thus, we reformulate gaze estimation as the quantification of the state shifting from the current state to several prior registered anchor states. Specifically, we propose a two-stage learning-based gaze estimation framework that divides the whole gaze estimation process into a coarse-to-fine approach involving anchor state selection and final gaze location. Moreover, to improve the generalization ability, instead of learning a large gaze estimation network directly, we align a group of local experts with a student network, where a novel denoising distillation algorithm is introduced to utilize denoising diffusion techniques to iteratively remove inherent noise in event data. Extensive experiments demonstrate the effectiveness of the proposed method, which surpasses state-of-the-art methods by a large margin of 15$\%$. The code will be publicly available athttps://github.com/ZHU-Zhiyu/Event_Gaze_Tracking.
Jinhui Hou, Jiading Li, Jinjian Wu, Junhui Hou
IEEE Trans. Mob. Comput.4
2025 Progressive Semi-Decoupled Detector for Accurate Object Detection
abstract
Inconsistent accuracy between classification and localization tasks is a common challenge in modern object detection. Task decoupling, which employs distinct features or labeling strategies for each task, is a widely used approach to address this issue. Although it has led to noteworthy advancements, this approach is insufficient as it neglects task interdependence and lacks an explicit consistency constraint. To bridge this gap, this paper proposes the Progressive Semi-Decoupled Detector (ProSDD) to enhance both classification and localization accuracy. Specifically, a new detection head is designed that incorporates feature suppression and enhancement mechanism (FSEM) and bidirectional interaction module (BIM). Compared with the decoupled head, it not only filters out task-irrelevant information and enhances task-related information, but also avoids excessive decoupling at the feature level. Moreover, both FSEM and BIM are used multiple times, thus forming a progressive semi-decoupled head. Then, a novel consistency loss is proposed and integrated into the loss function of object detection, ensuring harmonic performance in classification and localization. Experimental results demonstrate that the proposed ProSDD effectively alleviates inconsistent accuracy and achieves high-quality object detection. Taking the pretrained ResNet-50 as the backbone, ProSDD achieves a remarkable 43.3 AP on the MS COCO dataset, surpassing contemporary state-of-the-art detectors by a substantial margin under the equivalent configurations. Code is available athttps://github.com/HB-X/ProSDD.
Bo Han 0004, Lihuo He, Junjie Ke, Jinjian Wu, Xinbo Gao 0001
IEEE Trans. Multim.4
2025 Variational Multiple-Instance Learning With Embedding Correlation Modeling for Hyperspectral Target Detection
abstract
The hyperspectral target detection is widely concerned in geoscience and remote sensing due to the abundant spectral information in hyperspectral imagery. However, the detection performance is highly dependent on the high-quality target signature or pixel-level supervised signals, which are extremely challenging and costly. In this article, we propose a variational multiple-instance neural network with embedding correlation modeling (VMIL-ECM) for weakly supervised hyperspectral target detection, which relaxes the rigid target prior (e.g., target signatures and/or pixel-level annotations), and only region-level labels are required. VMIL-ECM explicitly models the location of the targets within the region as a latent variable under the nonindependent and identically distributed (non-i.i.d.) assumption to estimate the underlying ground-truth target locations. The expectation-maximization (EM) algorithm is employed to iteratively optimize the posterior distribution of latent variables and learn discriminative spectral features for the target detection. To fully utilize the contextual information within the hyperspectral region, a permutation-invariant transformer-based structure is devised to explore the embedding correlation among instances. Moreover, a dynamic thresholding strategy is adopted to produce the reliable fine-grained supervised signals. Extensive experiments on three simulated datasets and two real-field datasets are conducted to verify the effectiveness of VMIL-ECM, and the state-of-the-art performance has been achieved over the existing comparison methods. The code for the VMIL-ECM is publicly available at: https://github.com/BoYangXDU/VMIL-ECM.
Bo Yang 0047, Changzhe Jiao, Jinjian Wu, Leida Li
IEEE Trans. Neural Networks Learn. Syst.3
2024 Scaling and Masking: A New Paradigm of Data Sampling for Image and Video Quality Assessment
abstract
Quality assessment of images and videos emphasizes both local details and global semantics, whereas general data sampling methods (e.g., resizing, cropping or grid-based fragment) fail to catch them simultaneously. To address the deficiency, current approaches have to adopt multi-branch models and take as input the multi-resolution data, which burdens the model complexity. In this work, instead of stacking up models, a more elegant data sampling method (named as SAMA, scaling and masking) is explored, which compacts both the local and global content in a regular input size. The basic idea is to scale the data into a pyramid first, and reduce the pyramid into a regular data dimension with a masking strategy. Benefiting from the spatial and temporal redundancy in images and videos, the processed data maintains the multi-scale characteristics with a regular input size, thus can be processed by a single-branch model. We verify the sampling method in image and video quality assessment. Experiments show that our sampling method can improve the performance of current single-branch models significantly, and achieves competitive performance to the multi-branch models without extra model complexity. The source code will be available at https://github.com/Sissuire/SAMA.
Yongxu Liu 0001, Yinghui Quan, Guoyao Xiao, Jinjian Wu
AAAI5
2024 Motion Deblurring via Spatial-Temporal Collaboration of Frames and Events
abstract
Motion deblurring can be advanced by exploiting informative features from supplementary sensors such as event cameras, which can capture rich motion information asynchronously with high temporal resolution. Existing event-based motion deblurring methods neither consider the modality redundancy in spatial fusion nor temporal cooperation between events and frames. To tackle these limitations, a novel spatial-temporal collaboration network (STCNet) is proposed for event-based motion deblurring. Firstly, we propose a differential-modality based cross-modal calibration strategy to suppress redundancy for complementarity enhancement, and then bimodal spatial fusion is achieved with an elaborate cross-modal co-attention mechanism to weight the contributions of them for importance balance. Besides, we present a frame-event mutual spatio-temporal attention scheme to alleviate the errors of relying only on frames to compute cross-temporal similarities when the motion blur is significant, and then the spatio-temporal features from both frames and events are aggregated with the custom cross-temporal coordinate attention. Extensive experiments on both synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/STCNet.
Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Guangming Shi
AAAI2
2024 Segment Any Event Streams via Weighted Adaptation of Pivotal Tokens
abstract
In this paper, we delve into the nuanced challenge of tailoring the Segment Anything Models (SAMs) for integration with event data, with the overarching objective of attaining robust and universal object segmentation within the event-centric domain. One pivotal issue at the heart of this endeavor is the precise alignment and calibration of embeddings derived from event-centric data such that they harmoniously coincide with those originating from RGB imagery. Capitalizing on the vast repositories of datasets with paired events and RGB images, our proposition is to harness and extrapolate the profound knowledge encapsulated within the pretrained SAM framework. As a cornerstone to achieving this, we introduce a multi-scale feature distillation methodology. This methodology rigorously optimizes the alignment of token embeddings originating from event data with their RGB image counterparts, thereby preserving and enhancing the robustness of the overall architecture. Considering the distinct significance that token embeddings from intermediate layers hold for higher-level embeddings, our strategy is centered on accurately calibrating the pivotal token embeddings. This targeted calibration is aimed at effectively managing the discrepancies in high-level embeddings originating from both the event and image domains. Extensive experiments on different datasets demonstrate the effectiveness of the proposed distillation method. Code in https://github.com/happychenpipi/EventSAM.
Zhiwen Chen 0002, Yifan Zhang 0036, Junhui Hou, Guangming Shi, Jinjian Wu
CVPR6
2024 Bridging the Synthetic-to-Authentic Gap: Distortion-Guided Unsupervised Domain Adaptation for Blind Image Quality Assessment
abstract
The annotation of blind image quality assessment (BIQA) is labor-intensive and time-consuming, especially for authentic images. Training on synthetic data is expected to be beneficial, but synthetically trained models often suf-fer from poor generalization in real domains due to domain gaps. In this work, we make a key observation that introducing more distortion types in the synthetic dataset may not improve or even be harmful to generalizing au-thentic image quality assessment. To solve this challenge, we propose distortion-guided unsupervised domain adaptationfor BIQA (DGQA), a novel framework that leverages adaptive multi-domain selection via prior knowledge from distortion to match the data distribution between the source domains and the target domain, thereby reducing negative transfer from the outlier source domains. Extensive experiments on two cross-domain settings (synthetic distortion to authentic distortion and synthetic distortion to algorith-mic distortion) have demonstrated the effectiveness of our proposed DGQA. Besides, DGQA is orthogonal to existing model-based BIQA methods, and can be used in combi-nation with such models to improve performance with less training data.
Jinjian Wu, Yongxu Liu 0001, Leida Li
CVPR2
2024 An O(m+n)-Space Spatiotemporal Denoising Filter with Cache-Like Memories for Dynamic Vision Sensors
abstract
Dynamic vision sensor (DVS) is novel neuromorphic imaging device that generates asynchronous events. Despite the high temporal resolution and high dynamic range features, DVS is faced with background noise problem. Spatiotemporal filter is an effective and hardware-friendly solution for DVS denoising but previous designs have large memory overhead or degraded performance issues. In this paper, we present a lightweight and real-time spatiotemporal denoising filter with set-associative cache-like memories, which has low space complexity of O(m+n) for DVS of m×n resolution. A two-stage pipeline for memory access with read cancellation feature is proposed to reduce power consumption. Further the bitwidth redundancy for event storage is exploited to minimize the memory footprint. We implemented our design on FPGA and experimental results show that it achieves state-of-the-art performance compared with previous spatiotemporal filters while maintaining low resource utilization and low power consumption of about 125mW to 210mW at 100MHz clock frequency.
Qinghang Zhao, Yixi Ji, Jinjian Wu, Guangming Shi
ICCAD4
2024 Learnable Prompts-Based Transformers for Domain Generalization of Hyperspectral Image Classification
abstract
Extensive pre-trained visual-language alignment models, such as Contrastive Language-Image Pre-training (CLIP), have demonstrated significant potential for learning representations transferable to domain generation tasks. In hyperspectral image (HSI) classification, a major challenge in deploying such models lies in prompt engineering, which requires particular expertise and substantial time investment. Moreover, existing methods ignore correlation information cross spectral bands. To address these issues, a novel method named learnable prompts-based Transformer (LPFormer) is proposed in this paper. In LPFormer, cross-band correlation information is extracted by self-attention of the transformer, which converted into positional embedding within the transformer framework to obtain the visual features. Subsequently, prompt words are modeled using learnable parameters that turn into efficient expertise. Finally, contrast learning method is used to align visual and textual features. Experimental results on two HSI datasets shows that the proposed LPFormer outperforms other domain adaptation methods.
Baofa He, Jie Feng 0003, Ronghua Shang, Jinjian Wu, Licheng Jiao
IGARSS5
2024 Semantics-Aware Image Aesthetics Assessment using Tag Matching and Contrastive Ranking
abstract
The perception of image aesthetics is built upon the understanding of semantic content. However, how to evaluate the aesthetic quality of images with diversified semantic backgrounds remains challenging in image aesthetics assessment (IAA). To address the dilemma, this paper presents a semantics-aware image aesthetics assessment approach, which first analyzes the semantic content of images and then models the aesthetic distinctions among images from two perspectives, i.e., aesthetic attribute and aesthetic level. Concretely, we propose two strategies, dubbed tag matching and contrastive ranking, to extract knowledge pertaining to image aesthetics. The tag matching identifies the semantic category and the dominant aesthetic attributes based on predefined tag libraries. The contrastive ranking is designed to uncover the comparative relationships among images with different aesthetic levels but similar semantic backgrounds. In the process of contrastive ranking, the impact of long-tailed distribution of aesthetic data is also considered by balanced sampling and traversal contrastive learning. Extensive experiments and comparisons on three benchmark IAA databases demonstrate the superior performance of the proposed model in terms of both prediction accuracy and alleviating long-tailed effect. The code will be public at https://github.com/yzc-ippl/TMCR **REMOVE 2nd URL**://github.com/yzc-ippl/TMCR.
Zhichao Yang 0013, Leida Li, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong
ACM Multimedia4
2024 E-Motion: Future Motion Simulation via Event Sequence Diffusion
abstract
Forecasting a typical object's future motion is a critical task for interpreting and interacting with dynamic environments in computer vision. Event-based sensors, which could capture changes in the scene with exceptional temporal granularity, may potentially offer a unique opportunity to predict future motion with a level of detail and precision previously unachievable. Inspired by that, we propose to integrate the strong learning capacity of the video diffusion model with the rich motion information of an event camera as a motion simulation framework. Specifically, we initially employ pre-trained stable video diffusion models to adapt the event sequence dataset. This process facilitates the transfer of extensive knowledge from RGB videos to an event-centric domain. Moreover, we introduce an alignment mechanism that utilizes reinforcement learning techniques to enhance the reverse generation trajectory of the diffusion model, ensuring improved performance and accuracy. Through extensive testing and validation, we demonstrate the effectiveness of our method in various complex scenarios, showcasing its potential to revolutionize motion flow prediction in computer vision applications such as autonomous vehicle guidance, robotic navigation, and interactive media. Our findings suggest a promising direction for future research in enhancing the interpretative power and predictive accuracy of computer vision systems. The source code is publicly available at https://github.com/p4r4mount/E-Motion.
Junhui Hou, Guangming Shi, Jinjian Wu
NeurIPS5
2024 Quality-aware blind image motion deblurring
Tianshu Song, Leida Li, Jinjian Wu, Weisheng Dong, Deqiang Cheng 0001
Pattern Recognit.3
2024 Learning real-world heterogeneous noise models with a benchmark dataset
Jie Lin 0008, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi
Pattern Recognit.5
2024 Emotion-aware hierarchical interaction network for multimodal image aesthetics assessment
Tong Zhu 0003, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001
Pattern Recognit.4
2024 Motion-Oriented Hybrid Spiking Neural Networks for Event-Based Motion Deblurring
abstract
Image deblurring based only on the blurry image is challenging as motion information is lost while imaging. Event cameras capture the texture of moving objects in high temporal resolution with asynchronous events. In this paper, we extract motion features from events and fuse them with background features from the image for event-based image deblurring. Spiking neural network (SNN), a widely recognized event feature extractor, is well suited for motion feature extraction due to its high temporal resolution. However, extracting motion information from events exclusively with SNN is challenging. We propose a novel Temporal-local-Spatio Spiking Transformer (TSST) to extract motion intensity and motion attention regions in the spatio-temporal domain. Motion intensity extracted from spiking features is represented as a high temporal resolution motion attention map to guide the fusion of the two networks. In the temporal domain, motion intensity maps spiking features to CNN features as motion features to avoid blurring. In the spatial domain, the motion intensity shows the motion regions and gives the weight of the motion feature during fusion. Moreover, a hybrid feature extraction encoder (HFEE) is introduced, which fully fuses the motion and background features for deblurring. The gradient is back-propagated from CNN to SNN, and the hybrid deblurring network is jointly optimized. We evaluated the performance of our model on the public dataset GoPro and a real event dataset we captured. Codes and pretrained models are available athttps://github.com/XDULzx/MotionSNN.
Zhaoxin Liu, Jinjian Wu, Guangming Shi, Wen Yang 0008, Weisheng Dong, Qinghang Zhao
IEEE Trans. Circuits Syst. Video Technol.2
2024 Active Learning-Based Sample Selection for Label-Efficient Blind Image Quality Assessment
abstract
Despite the considerable effort devoted to high-generalizable blind image quality assessment (BIQA), the generalization performance of the state-of-the-art metrics remains limited when facing new visual scenes. A straightforward way to address the dilemma is labeling a great number of images from the new scene and subsequently training a new model, which is quite labor-intensive and cost-expensive. Hence, there is an urgent need to mitigate the dependency on labeled samples by designing a data-efficient BIQA algorithm. Motivated by the above facts, this paper presents an Active Learning-based IQA (AL-IQA) framework, which reduces the requirement for training samples by selecting representative images from two perspectives, including distortion and content. Specifically, in terms of distortion, we design distortion prompts and adopt Contrastive Language-Image Pre-Training (CLIP) to predict image distortion in a zero-shot manner. Then, we employ curriculum learning-inspired strategy to select samples with gradually increasing difficulty (measured by prediction uncertainty of CLIP), in order to facilitate model training. Meantime, in terms of content, we adopt distribution matching-based dataset distillation to distill unlabeled images into several high-density informative synthetic images. Then, feature distances between unlabeled images and distilled images are compared to identify images with the most representative content. Finally, Borda count is adopted to capture a consensus of both distortion and content through weighted counting, and prompt tuning is utilized for adapting the model to the IQA task. Extensive experiments are conducted on five IQA datasets, and the results demonstrate that the proposed AL-IQA not only effectively reduces the number of training samples but also achieves state-of-the-art prediction accuracy and generalization performance. The source code is available athttps://github.com/esnthere/AL-IQA.
Tianshu Song, Leida Li, Deqiang Cheng 0001, Pengfei Chen 0003, Jinjian Wu
IEEE Trans. Circuits Syst. Video Technol.5
2024 Class-Aligned and Class-Balancing Generative Domain Adaptation for Hyperspectral Image Classification
abstract
The task of hyperspectral image (HSI) classification is fundamental and crucial in HSI processing. Currently, domain adaptive methods have become a research hotspot in HSI classification. However, most domain adaptive methods ignore the class alignment in different domains. Additionally, HSIs have the characteristics of category imbalance and complex spatial-spectral distribution, which restricts the adaptation performance in HSIs. To address these problems, a class-aligned and class-balancing generative domain adaptation (CCGDA) method is proposed for HSI classification. The architecture of CCGDA is designed by using the classifier, domain discriminator, sampler and two weight-sharing generators. In the classifier, split-level capsule network is constructed by extracting rich spatial information of shallow layer and spectral features of deep layer with equivariant characteristic. Then, the classifier provides the pseudo label of samples in the target domain. To prevent the generators from mode collapse caused by category imbalance, the sampler is designed. It samples and re-samples the samples of the target domain in an adaptive proportion according to the statistical calculation through confidence and distribution of pseudo labels. Finally, a novel class-aligned domain adversarial loss is defined to jointly optimize the generators and discriminator. It incorporates the class shift adjusting and adaptive sampling for the samples of the target domain to better adapt the discriminant boundary of the classifier to the target domain. Experiments on benchmark HSI datasets verify the superiority of the proposed method for domain adaptive classification.
Jie Feng 0003, Ziyu Zhou 0009, Ronghua Shang, Jinjian Wu, Xiangrong Zhang, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.4
2024 Learning Frame-Event Fusion for Motion Deblurring
abstract
Motion deblurring is a highly ill-posed problem due to the significant loss of motion information in the blurring process. Complementary informative features from auxiliary sensors such as event cameras can be explored for guiding motion deblurring. The event camera can capture rich motion information asynchronously with microsecond accuracy. In this paper, a novel frame-event fusion framework is proposed for event-driven motion deblurring (FEF-Deblur), which can sufficiently explore long-range cross-modal information interactions. Firstly, different modalities are usually complementary and also redundant. Cross-modal fusion is modeled as complementary-unique features separation-and-aggregation, avoiding the modality redundancy. Unique features and complementary features are first inferred with parallel intra-modal self-attention and inter-modal cross-attention respectively. After that, a correlation-based constraint is designed to act between unique and complementary features to facilitate their differentiation, which assists in cross-modal redundancy suppression. Additionally, spatio-temporal dependencies among neighboring inputs are crucial for motion deblurring. A recurrent cross attention is introduced to preserve inter-input attention information, in which the current spatial features and aggregated temporal features are attending to each other by establishing the long-range interaction between them. Extensive experiments on both synthetic and real-world motion deblurring datasets demonstrate our method outperforms state-of-the-art event-based and image/video-based methods. The code will be made publicly available.
Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Image Process.2
2024 Coarse-to-Fine Image Aesthetics Assessment With Dynamic Attribute Selection
abstract
Image aesthetics assessment (IAA) is an interesting but challenging task, owing to the ineffable nature of human sense of beauty. The study of IAA has evolved from simple binary classification to more complex score regression and distribution prediction. It is effortless for people to perform aesthetic binary classification,i.e., aesthetically pleasing or not. However, further judgment on the fine-level scalar aesthetic score is complex and typically determined by aesthetic attributes presented in the image, such as content, lighting and color. Motivated by the above facts, this paper presents a Coarse-to-fine image Aesthetics assessment model guided by Dynamic Attribute Selection, dubbed CADAS. The underlying idea is to simulate the process of human aesthetic perception by performing coarse-to-fine aesthetic reasoning. Specifically, a hierarchical AttributeNet is first pre-trained by imitating the staged mechanism of human aesthetic experience, producing the candidate aesthetic attributes. Then, an AestheticNet is introduced to perform the coarse-level binary classification, based on which a confidence-based attribute selection strategy is designed to dynamically pick out the dominant aesthetic attributes from the candidate ones. Finally, a self-attention-based FusionNet is designed to explore the interaction between dominant aesthetic attributes and aesthetic features, producing the fine-level aesthetic prediction. Extensive experiments demonstrate that the proposed model is superior to the state-of-the-arts. Furthermore, CADAS is also able to output the dominant aesthetic attributes in images, facilitating model explainability.
Yipo Huang, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001, Guangming Shi
IEEE Trans. Multim.4
2024 Blind Image Quality Assessment Based on Perceptual Comparison
abstract
Blind image quality assessment (BIQA) is a regression task with continuous label space, the feature space of which is expected to have a corresponding continuity in the target space. However, existing approaches typically learn quality score regression directly in an end-to-end fashion, which leaves networks susceptible to interference from task-agnostic information, and fails to capture the continuity of BIQA. In this work, by explicitly establishing inter-sample associations, a simple yet effective BIQA framework based on perceptual comparison is proposed to capture the continuity. To this end, besides the basic quality score regression, the relative quality scores between images are predicted to exploit the relative quality relationships between samples for optimizing the representation of image perceptual quality. In addition, based on the human perceptual characteristic, we derive a novel sample weighting strategy to dynamically adjust the weights for different samples in the network learning process for further improving the robustness of the model. The performances on both single-database and cross-database experiments achieve state-of-the-art, indicating the effectiveness of the proposed method. Besides, the proposed framework is model-agnostic, which can effectively improve the performance of the benchmark model with no extra inference cost.
Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Multim.2
2024 Quality Assessment for Stitched Panoramic Images via Patch Registration and Bidimensional Feature Aggregation
abstract
Quality assessment for stitched panoramic images (SPIQA) is of great significance for the stitching algorithm optimization. By contrast, this task is much more challenging and arduous than traditional IQA task due to the high resolution of stitched panoramic images and the particularity and complexity of stitching distortions. For this task, we propose an effective method based on patch registration and bidimensional feature aggregation (PRBFA). First, inspired by the attention mechanism of the human visual system and the limited range of human vision, a soft patch segmentation and selection method is presented to determine the key patches in panoramic images to participate in the following patch matching and feature alignment stages, achieving patch registration between the panoramic image and the corresponding constituent images. Further, to fully simulate the human visual perception process from local viewport to panorama, the feature exploration is successively performed from local to global, which is also adaptive to the complexity of the distortions in stitched panoramic images. For performance testification, extensive experiments are conducted on the publicly released SPIQA database, the results of which prove the performance superiority of the PRBFA method.
Yu Zhou 0009, Weikang Gong, Yanjing Sun, Leida Li, Ke Gu 0001, Jinjian Wu
IEEE Trans. Multim.6
2023 Self-supervised Non-uniform Kernel Estimation with Flow-based Motion Prior for Blind Image Deblurring
abstract
Many deep learning-based solutions to blind image deblurring estimate the blur representation and reconstruct the target image from its blurry observation. However, these methods suffer from severe performance degradation in real-world scenarios because they ignore important prior information about motion blur (e.g., real-world motion blur is diverse and spatially varying). Some methods have attempted to explicitly estimate non-uniform blur kernels by CNNs, but accurate estimation is still challenging due to the lack of ground truth about spatially varying blur kernels in real-world images. To address these issues, we propose to represent the field of motion blur kernels in a latent space by normalizing flows, and design CNNs to predict the latent codes instead of motion kernels. To further improve the accuracy and robustness of non-uniform kernel estimation, we introduce uncertainty learning into the process of estimating latent codes and propose a multi-scale kernel attention module to better integrate image features with estimated kernels. Extensive experimental results, especially on real-world blur datasets, demonstrate that our method achieves state-of-the-art results in terms of both subjective and objective quality as well as excellent generalization performance for non-uniform image deblurring. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/UFPNet.htm.
Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi
CVPR5
2023 Attribute-assisted Multimodal Network for Image Aesthetics Assessment
abstract
Image aesthetics assessment (IAA) is challenging due to its highly abstract nature. Nowadays, people tend to share images and comment them on social networks, which can provide rich information for judging image aesthetics. As a result, user comments of an image can be jointly utilized to learn better feature representations for IAA. Previous researches have shown that aesthetic attributes are crucial factors in determining image aesthetic quality and influencing people’s aesthetic perception. Accordingly, when commenting an image, people usually give descriptions from the perspective of aesthetic attributes. Inspired by this, this paper presents a new Attribute-Assisted Multimodal network (AAM-Net) for image aesthetics assessment. Specifically, we propose a cross-modal attribute interaction module to explore the related aesthetic attribute semantics shared by an image and the corresponding aesthetic comments. Then, a cross-modal gate unit is introduced to further refine significant attribute semantics interactively. Finally, informative aesthetic features can be obtained for predicting image aesthetic distributions. Experimental results on two public multimodal IAA databases demonstrate the superiority of the proposed model over the state-of-the-art methods.
Tong Zhu 0003, Leida Li, Pengfei Chen 0003, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo
ICME4
2023 Brain Anatomy-Guided MRI Analysis for Assessing Clinical Progression of Cognitive Impairment with Structural MRI
Jinjian Wu, Li Wang 0026, David C. Steffens, Shijun Qiu, Guy G. Potter, Mingxia Liu 0001
MICCAI (8)2
2023 AesCLIP: Multi-Attribute Contrastive Learning for Image Aesthetics Assessment
abstract
Image aesthetics assessment (IAA) aims at predicting the aesthetic quality of images. Recently, large pre-trained vision-language models, like CLIP, have shown impressive performances on various visual tasks. When it comes to IAA, a straightforward way is to finetune the CLIP image encoder using aesthetic images. However, this can only achieve limited success without considering the uniqueness of multimodal data in the aesthetics domain. People usually assess image aesthetics according to fine-grained visual attributes, e.g., color, light and composition. However, how to learn aesthetics-aware attributes from CLIP-based semantic space has not been addressed before. With this motivation, this paper presents a CLIP-based multi-attribute contrastive learning framework for IAA, dubbed AesCLIP. Specifically, AesCLIP consists of two major components, i.e., aesthetic attribute-based comment classification and attribute-aware learning. The former classifies the aesthetic comments into different attribute categories. Then the latter learns an aesthetic attribute-aware representation by contrastive learning, aiming to mitigate the domain shift from the general visual domain to the aesthetics domain. Extensive experiments have been done by using the pre-trained AesCLIP on four popular IAA databases, and the results demonstrate the advantage of AesCLIP over the state-of-the-arts. The source code will be public at https://github.com/OPPOMKLab/AesCLIP.
Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Jinjian Wu, Weisheng Dong, Yuzhe Yang 0001, Liwu Xu, Guangming Shi
ACM Multimedia4
2023 Event-based Motion Deblurring with Modality-Aware Decomposition and Recomposition
abstract
Event camera responds to the brightness changes at each pixel independently with microsecond accuracy. Event cameras offer attractive property that can record well high-speed scene but ignore static and non-moving areas, while conventional frame cameras are able to acquire the whole intensity information of the scene but suffer from motion blur. Therefore, it would be desirable to combine the best of two cameras for reconstructing high quality intensity frame with no motion blur. The human visual system presents a two-pathway procedure for non-action-based representation and objects motion perception, which corresponds well to the hybrid frame and event. In this paper, inspired by the two-pathway visual system, a novel dual-stream based framework is proposed for motion deblurring (DS-Deblur), which flexibly utilizes the respective advantages from frame and event. A complementary-unique information splitting based feature fusion module is firstly proposed to adaptively aggregate the frame and event progressively at multiple levels, which is well-grounded on the hierarchical process in twopathway visual system. Then, a recurrent spatio-temporal feature transformation module is designed to exploit relevant information between adjacent frames, in which features of both current and previous frames are transformed in a global-local manner. Extensive experiments on both synthetic and real motion blur datasets demonstrate our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/Motion-Deblurringwith-Hybrid-Frames-and-Events.
Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
ACM Multimedia2
2023 Technical Quality-Assisted Image Aesthetics Quality Assessment
Xiangfei Sheng, Leida Li, Pengfei Chen 0003, Jinjian Wu, Liwu Xu, Yuzhe Yang 0001
PRCV (11)4
2023 Memory Based Temporal Fusion Network for Video Deblurring
Chaohua Wang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi
Int. J. Comput. Vis.5
2023 Transfer learning for just noticeable difference estimation
Yongwei Mao, Jinjian Wu, Leida Li, Weisheng Dong
Inf. Sci.2
2023 Deep Gaussian Scale Mixture Prior for Image Reconstruction
abstract
Image reconstruction from partial observations has attracted increasing attention. Conventional image reconstruction methods with hand-crafted priors often fail to recover fine image details due to the poor representation capability of the hand-crafted priors. Deep learning methods attack this problem by directly learning mapping functions between the observations and the targeted images can achieve much better results. However, most powerful deep networks lack transparency and are nontrivial to design heuristically. This paper proposes a novel image reconstruction method based on the Maximum a Posterior (MAP) estimation framework using learned Gaussian Scale Mixture (GSM) prior. Unlike existing unfolding methods that only estimate the image means (i.e., the denoising prior) but neglected the variances, we propose characterizing images by the GSM models with learned means and variances through a deep network. Furthermore, to learn the long-range dependencies of images, we develop an enhanced variant based on the Swin Transformer for learning GSM models. All parameters of the MAP estimator and the deep network are jointly optimized through end-to-end training. Extensive simulation and real data experimental results on spectral compressive imaging and image super-resolution demonstrate that the proposed method outperforms existing state-of-the-art methods.
Xin Yuan 0002, Weisheng Dong, Jinjian Wu, Guangming Shi
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Adaptive Search-and-Training for Robust and Efficient Network Pruning
abstract
Both network pruning and neural architecture search (NAS) can be interpreted as techniques to automate the design and optimization of artificial neural networks. In this paper, we challenge the conventional wisdom of training before pruning by proposing a joint search-and-training approach to learn a compact network directly from scratch. Using pruning as a search strategy, we advocate three new insights for network engineering: 1) to formulate adaptive search as a cold start strategy to find a compact subnetwork on the coarse scale; and 2) to automatically learn the threshold for network pruning; 3) to offer flexibility to choose between efficiency and robustness. More specifically, we propose an adaptive search algorithm in the cold start by exploiting the randomness and flexibility of filter pruning. The weights associated with the network filters will be updated by ThreshNet, a flexible coarse-to-fine pruning method inspired by reinforcement learning. In addition, we introduce a robust pruning strategy leveraging the technique of knowledge distillation through a teacher-student network. Extensive experiments on ResNet and VGGNet have shown that our proposed method can achieve a better balance in terms of efficiency and accuracy and notable advantages over current state-of-the-art pruning methods in several popular datasets, including CIFAR10, CIFAR100, and ImageNet. The code associate with this paper is available at: https://see.xidian.edu.cn/faculty/wsdong/Projects/AST-NP.htm.
Xiaotong Lu, Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Dynamic Expert-Knowledge Ensemble for Generalizable Video Quality Assessment
abstract
Despite the impressive progress of supervised methods in quality assessment for in- the-wild videos, models trained on one domain often fail to generalize well to others due to the domain shifts caused by distortion diversity and content variation. Domain generalizable video quality assessment (VQA) methods that can work across domains remain an open research challenge. Although combining more data following the mixed-domain training strategy can improve the generalization performance to a certain extent, the specific knowledge from each source domain, which could potentially be useful for improving unseen domain generalization, is ignored in this principle. Motivated by this, we propose a domain generalizable VQA method named Dynamic Ensemble of Expert-Knowledge (DEEK), a novel framework that dynamically exploits the expert-knowledge from each source domain to achieve a generalizable ensemble prediction. Specifically, based on the multiple experts each trained to specialize in a particular source domain, we aim to exploit complementary information provided by the expert-knowledge. We effectively train an ensemble model by proposing a quality-sensitive InfoNCE loss to regularize the collaborative training of all experts in the contrastive learning formulation, aiming to exploit complementary information provided by the expert-knowledge when forming the ensemble. By dynamically integrating the experts according to their relevances to the target data, these expert-knowledge could be leveraged for better generalization. Experiments on five VQA datasets verify that our approach outperforms the state-of-the-arts by large margins.
Pengfei Chen 0003, Leida Li, Haoliang Li, Jinjian Wu, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.4
2023 ECSNet: Spatio-Temporal Feature Learning for Event Camera
abstract
The neuromorphic event cameras can efficiently sense the latent geometric structures and motion clues of a scene by generating asynchronous and sparse event signals. Due to the irregular layout of the event signals, how to leverage their plentiful spatio-temporal information for recognition tasks remains a significant challenge. Existing methods tend to treat events as dense image-like or point-serie representations. However, they either suffer from severe destruction on the sparsity of event data or fail to encode robust spatial cues. To fully exploit their inherent sparsity with reconciling the spatio-temporal information, we introduce a compact event representation, namely 2D-1T event cloud sequence (2D-1T ECS). We couple this representation with a novel light-weight spatio-temporal learning framework (ECSNet) that accommodates both object classification and action recognition tasks. The core of our framework is a hierarchical spatial relation module. Equipped with specially designed surface-event-based sampling unit and local event normalization unit to enhance the inter-event relation encoding, this module learns robust geometric features from the 2D event clouds. And we propose a motion attention module for efficiently capturing long-term temporal context evolving with the 1T cloud sequence. Empirically, the experiments show that our framework achieves par or even better state-of-the-art performance. Importantly, our approach cooperates well with the sparsity of event data without any sophisticated operations, hence leading to low computational costs and prominent inference speeds.
Zhiwen Chen 0002, Jinjian Wu, Junhui Hou, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.2
2023 Theme-Aware Visual Attribute Reasoning for Image Aesthetics Assessment
abstract
People usually assess image aesthetics according to visual attributes, e.g., interesting content, good lighting and vivid color, etc. Further, the perception of visual attributes depends on the image theme. Therefore, the inherent relationship between visual attributes and image theme is crucial for image aesthetics assessment (IAA), which has not been comprehensively investigated. With this motivation, this paper presents a new IAA model based on Theme-Aware Visual Attribute Reasoning (TAVAR). The underlying idea is to simulate the process of human perception in image aesthetics by performing bilevel reasoning. Specifically, a visual attribute analysis network and a theme understanding network are first pre-trained to extract aesthetic attribute features and theme features, respectively. Then, the first level Attribute-Theme Graph (ATG) is built to investigate the coupling relationship between visual attributes and image theme. Further, a flexible aesthetics network is introduced to extract general aesthetic features, based on which we built the second level Attribute-Aesthetics Graph (AAG) to mine the relationship between theme-aware visual attributes and aesthetic features, producing the final aesthetic prediction. Extensive experiments on four public IAA databases demonstrate the superiority of the proposed TAVAR model over the state-of-the-arts. Furthermore, TAVAR features better explainability due to the use of visual attributes.
Leida Li, Yipo Huang, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.3
2023 Quality Assessment of UGC Videos Based on Decomposition and Recomposition
abstract
The prevalence of short-video applications imposes more requirements for video quality assessment (VQA). User-generated content (UGC) videos are captured under an unprofessional environment, thus suffering from various dynamic degradations, such as camera shaking. To cover the dynamic degradations, existing recurrent neural network-based UGC-VQA methods can only provide implicit modeling, which is unclear and difficult to analyze. In this work, we consider explicit motion representation for dynamic degradations, and propose a motion-enhanced UGC-VQA method based on decomposition and recomposition. In the decomposition stage, a dual-stream decomposition module is built, and VQA task is decomposed into single frame-based quality assessment problem and cross frames-based motion understanding. The dual streams are well grounded on the two-pathway visual system during perception, and require no extra UGC data due to knowledge transfer. Hierarchical features from shallow to deep layers are gathered to narrow the gaps from tasks and domains. In the recomposition stage, a progressively residual aggregation module is built to recompose features from the dual streams. Representations with different layers and pathways are interacted and aggregated in a progressive and residual manner, which keeps a good trade-off between representation deficiency and redundancy. Extensive experiments on UGC-VQA databases verify that our method achieves the state-of-the-art performance and keeps a good capability of generalization. The source code will be available inhttps://github.com/Sissuire/DSD-PRO.
Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.2
2023 Depth Perception Assessment of 3D Videos Based on Stereoscopic and Spatial Orientation Structural Features
abstract
Depth quality of stereoscopic three-dimensional (S3D) videos is a significant factor which directly affects the quality of experience (QoE) associated with 3D video applications and services. Nevertheless, there remain limited reports on the investigation of depth perception and depth quality evaluation of S3D videos, which impedes further advancement and deployment of 3D video technology. This paper reports a series of subjective experiments which have been conducted to investigate the depth perception and its related properties of the human visual system (HVS) using S3D video compressed by the H.264/AVC standard. The experimental results reveal that the HVS response in depth perception varies at different frequencies and in varying orientations, and the distortions introduced by video coding can cause the loss of and/or variation in depth perception. By integration of binocular and monocular features (BM) extracted from left and right views of S3D video with respect to depth perception, a depth quality assessment model, herein referred to as BM-DQAM, is devised by training these stereoscopic and spatial orientation structural features with a support vector regression model. It is shown that the BM-DQAM provides a novel no-reference metric for evaluation of the depth quality of S3D videos. Based on two publicly available 3D video databases and the proposed depth perception assessment database, the experimental results show that the BM-DQAM has demonstrated better performance in assessing the depth quality in S3D video viewing than that of other metrics reported in the published literatures, correlating well with the HVS response in the depth perception assessment experiment.
Wenfei Wan, Dengjia Huang, Bin Shang, Shengyu Wei, Hong Ren Wu, Jinjian Wu, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.6
2023 Multiple-Instance Metric Learning Network for Hyperspectral Target Detection
abstract
Target detection becomes increasingly important in hyperspectral image analysis but is limited by difficulties in acquiring accurate pixel-level training labels. This paper proposes a multiple instance metric learning neural network (MIML-Net) for hyperspectral target detection tasks, which only requires region-level labels and greatly alleviates the laborious pixel-level annotation problems. Our method learns the embeddings of regions with weak labels under attention-based multiple instance learning framework. Based on which, we impose a novel metric-based regularizer to constrain target and background embeddings to two learnable compact clusters with distinct centroids, which further boosts the spectral feature representation ability. The proposed metric-based regularizer enforces a discriminative detector due to its capability to reduce the intra-class variations and encourage the inter-class separations simultaneously. Extensive experimental results from both simulated and real-field data sets demonstrate the effectiveness of the proposed MIML-Net in comparison with the state-of-the-art weakly supervised techniques.
Bo Yang 0047, Changzhe Jiao, Guozhen Wang, Lei Wang 0258, Jinjian Wu
IEEE Trans. Geosci. Remote. Sens.7
2023 Searching Efficient Model-Guided Deep Network for Image Denoising
abstract
Unlike the success of neural architecture search (NAS) in high-level vision tasks, it remains challenging to find computationally efficient and memory-efficient solutions to low-level vision problems such as image restoration through NAS. One of the fundamental barriers to differential NAS-based image restoration is the optimization gap between the super-network and the sub-architectures, causing instability during the searching process. In this paper, we present a novel approach to fill this gap in image denoising application by connecting model-guided design (MoD) with NAS (MoD-NAS). Specifically, we propose to construct a new search space under a model-guided framework and develop more stable and efficient differential search strategies. MoD-NAS employs a highly reusable width search strategy and a densely connected search block to automatically select the operations of each layer as well as network width and depth via gradient descent. During the search process, the proposed MoD-NAS remains stable because of the smoother search space designed under the model-guided framework. Experimental results on several popular datasets show that our MoD-NAS method has achieved at least comparable even better PSNR performance than current state-of-the-art methods with fewer parameters, fewer flops, and less testing time. "The code associate with this paper is available at: https://see.xidian.edu.cn/faculty/wsdong/Projects/Mod-NAS.htm".
Qian Ning, Weisheng Dong, Xin Li 0005, Jinjian Wu
IEEE Trans. Image Process.4
2023 Knowledge-Guided Blind Image Quality Assessment With Few Training Samples
abstract
Blind image quality assessment (BIQA) for in-the-wild images has achieved great progress by training advanced deep neural networks. However, the current BIQA models are suffering the generalization challenge, meaning that a well-trained BIQA model is still very limited in evaluating images with different distributions. Deep BIQA models are data-intensive, but the annotation of image quality labels is extremely expensive. To design a generalizable BIQA model with few training samples is highly desired. Motivated by the above fact, this paper presents a knowledge-guided BIQA (KG-IQA) framework by integrating domain knowledge from the human visual system (HVS) and natural scene statistics (NSS). Specifically, the quality-aware HVS and NSS features are first extracted as prior knowledge. Then, we embed the two types of knowledge into the conventional deep neural network by learning to predict the HVS and NSS features, producing the knowledge-enhanced quality features, based on which the final image quality score is obtained. We conduct extensive experiments and comparisons on five authentically distorted IQA datasets. The experimental results demonstrate that the introduction of knowledge greatly reduces the requirement on the amount of training images, and the proposed KG-IQA model achieves superior performance in terms of both prediction accuracy and generalization ability.
Tianshu Song, Leida Li, Jinjian Wu, Yuzhe Yang 0001, Yandong Guo, Guangming Shi
IEEE Trans. Multim.3
2023 Pyramid Feature Aggregation for Hierarchical Quality Prediction of Stitched Panoramic Images
abstract
Panoramic image quality assessment (PIQA) is crucial to the successful application of technologies that can provide immersive visual experience. Stitching distortions are one of the main types of distortions that result in panoramic image degradation. However, most existing PIQA methods are general-purpose ones, which ignore the special characteristics of the stitching distortions caused by imperfect stitching algorithms. This results in unsatisfactory performance. To this end, we propose an effective stitched PIQA method, which consists of an imaginary reference generation (IRG) module and a hierarchical quality prediction (HQP) module. Among them, the IRG module is proposed to mimic the capability of the human visual system in imagining the raw version in the face of a degraded image. For the IRG module learning, we construct a large-scale database. The HQP module is presented to adapt to the particularity and complexity of stitching distortions, which is achieved by the pyramid feature aggregation. Extensive experiments and comparisons have been performed on the stitched PIQA database and the experimental results demonstrate the superiority of the proposed method in evaluating the quality of stitched panoramic images.
Yu Zhou 0009, Weikang Gong, Yanjing Sun, Leida Li, Jinjian Wu, Xinbo Gao 0001
IEEE Trans. Multim.5
2022 Robust Depth Completion with Uncertainty-Driven Loss Functions
abstract
Recovering a dense depth image from sparse LiDAR scans is a challenging task. Despite the popularity of color-guided methods for sparse-to-dense depth completion, they treated pixels equally during optimization, ignoring the uneven distribution characteristics in the sparse depth map and the accumulated outliers in the synthesized ground truth. In this work, we introduce uncertainty-driven loss functions to improve the robustness of depth completion and handle the uncertainty in depth completion. Specifically, we propose an explicit uncertainty formulation for robust depth completion with Jeffrey's prior. A parametric uncertain-driven loss is introduced and translated to new loss functions that are robust to noisy or missing data. Meanwhile, we propose a multiscale joint prediction model that can simultaneously predict depth and uncertainty maps. The estimated uncertainty map is also used to perform adaptive prediction on the pixels with high uncertainty, leading to a residual map for refining the completion results. Our method has been tested on KITTI Depth Completion Benchmark and achieved the state-of-the-art robustness performance in terms of MAE, IMAE, and IRMSE metrics.
Yufan Zhu, Weisheng Dong, Leida Li, Jinjian Wu, Xin Li 0005, Guangming Shi
AAAI4
2022 Uncertainty Learning in Kernel Estimation for Multi-stage Blind Image Super-Resolution
Zhenxuan Fang, Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
ECCV (18)4
2022 Self-feature Distillation with Uncertainty Modeling for Degraded Image Recognition
Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
ECCV (24)4
2022 AEDNet: Asynchronous Event Denoising with Spatial-Temporal Correlation among Irregular Data
abstract
Dynamic Vision Sensor (DVS) is a compelling neuromorphic camera compared to conventional camera, but it suffers from fiercer noise. Due to the nature of irregular format and asynchronous readout, DVS data is always transformed into a regular tensor (e.g., 3D voxel or image) for deep learning method, which corrupts its own asynchronous properties. To maintain asynchronous, we establish an innovative asynchronous event denoise neural network, named AEDNet, which directly consumes the correlation of the irregular signal in spatial-temporal range without destroying its original structural property. Based on the property of continuation in temporal domain and discreteness in spatial domain, we decompose the DVS signal into two parts, i.e., temporal correlation and spatial affinity, and separately process these two parts. Our spatial feature embedding unit is a unique feature extraction module that extracts feature from event-level, which perfectly maintains its spatial-temporal correlation. To test effectiveness, we build a novel dataset named DVSCLEAN containing both simulated and real-world data. The experimental results of AEDNet achieve SOTA.
Huachen Fang, Jinjian Wu, Leida Li, Junhui Hou, Weisheng Dong, Guangming Shi
ACM Multimedia2
2022 Learning for Motion Deblurring with Hybrid Frames and Events
abstract
Event camera responds to the brightness changes at each pixel independently with microsecond accuracy. Event cameras offer attractive property that can record well high-speed scene but ignore static and non-moving areas, while conventional frame cameras are able to acquire the whole intensity information of the scene but suffer from motion blur. Therefore, it would be desirable to combine the best of two cameras for reconstructing high quality intensity frame with no motion blur. The human visual system presents a two-pathway procedure for non-action-based representation and objects motion perception, which corresponds well to the hybrid frame and event. In this paper, inspired by the two-pathway visual system, a novel dual-stream based framework is proposed for motion deblurring (DS-Deblur), which flexibly utilizes the respective advantages from frame and event. A complementary-unique information splitting based feature fusion module is firstly proposed to adaptively aggregate the frame and event progressively at multiple levels, which is well-grounded on the hierarchical process in twopathway visual system. Then, a recurrent spatio-temporal feature transformation module is designed to exploit relevant information between adjacent frames, in which features of both current and previous frames are transformed in a global-local manner. Extensive experiments on both synthetic and real motion blur datasets demonstrate our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/Motion-Deblurringwith-Hybrid-Frames-and-Events.
Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Weisheng Dong, Guangming Shi
ACM Multimedia2
2022 Semantic Attribute Guided Image Aesthetics Assessment
abstract
Image aesthetics assessment (IAA) measures the perceived beauty of images using a computational approach. People usually assess the aesthetics of an image according to semantic attributes, e.g., lighting, color, object emphasis, etc. However, the state-of-the-art IAA approaches usually follow the data-driven framework without considering the rich attributes contained in images. With this motivation, this paper presents a new semantic attribute guided IAA model, where the attention maps of semantic attributes are employed to enhance the representation ability of general aesthetic features for more effective aesthetics assessment. Specifically, we first design an attribute attention generation network to obtain the attention maps for different semantic attributes, which are utilized to weight the general aesthetic features, producing the semantic attribute-enhanced feature representations. Then, the Graph Convolutional Network (GCN) is employed to further investigate the inherent relationship among the enhanced aesthetic features, producing the final image aesthetics prediction. Extensive experiments and comparisons on three public IAA databases demonstrate the effectiveness of the proposed method.
Jiachen Duan, Pengfei Chen 0003, Leida Li, Jinjian Wu, Guangming Shi
VCIP4
2022 Robust Dynamic Background Modeling for Foreground Estimation
abstract
Separating the background and foreground components from video frames is important to many tasks in computer vision and multimedia. As of today, robust principal component analysis (RPCA) has shown highly promising performance with the assumption that the background is low-rank and the foreground is sparse. However, existing RPCA-based methods have overlooked the uncertainty that some parts of the background (e.g., moving leaves in a dynamic background) or even the whole background (e.g., camera jittering) can be moving, which violates the low-rank assumption. To address this issue, we propose a novel enhanced RPCA framework (called ERPCA) by robustly modeling the dynamic background. Different from traditional RPCA framework, the background is decomposed into a low-rank component and a sparse component in the proposed ERPCA framework. Specifically, the sparse parts including foreground and dynamic parts of the background are modeled by Gaussian scale mixture (GSM) model. Moreover, those sparse components are further constrained by temporal consistency using nonzeromeans Gaussian models; the correspondences between sparse pixels in adjacent frames are explored by optical flow. Experimental results on 40 real videos demonstrate the superiority of our proposed method, with better average results than current state-of-the-art foreground estimation methods.
Qian Ning, Weisheng Dong, Jinjian Wu, Guangming Shi, Xin Li 0005
VCIP4
2022 Correlation filters based on spatial-temporal Gaussion scale mixture modelling for visual tracking
Guangming Shi, Weisheng Dong, Tianzhu Zhang 0001, Jinjian Wu, Xuemei Xie, Xin Li 0005
Neurocomputing5
2022 Blind image quality assessment based on progressive multi-task learning
Jinjian Wu, Shiwei Tian, Leida Li, Weisheng Dong, Guangming Shi
Neurocomputing2
2022 SPIQ: A Self-Supervised Pre-Trained Model for Image Quality Assessment
abstract
Blind image quality assessment (BIQA) has witnessed a flourishing progress due to the rapid advances in deep learning technique. The vast majority of prior BIQA methods try to leverage models pre-trained on ImageNet to mitigate the data shortage problem. These well-trained models, however, can be sub-optimal when applied to BIQA task that varies considerably from the image classification domain. To address this issue, we make the first attempt to leverage the plentiful unlabeled data to conduct self-supervised pre-training for BIQA task. Based on the distorted images generated from the high-quality samples using the designed distortion augmentation strategy, the proposed pre-training is implemented by a feature representation prediction task. Specifically, patch-wise feature representations corresponding to a certain grid are integrated to make prediction for the representation of the patch below it. The prediction quality is then evaluated using a contrastive loss to capture quality-aware information for BIQA task. Experimental results conducted on KADID-10 k and KonIQ-10 k databases demonstrate that the learned pre-trained model can significantly benefit the existing learning based IQA models.
Pengfei Chen 0003, Leida Li, Qingbo Wu 0001, Jinjian Wu
IEEE Signal Process. Lett.4
2022 Bayesian Correlation Filter Learning With Gaussian Scale Mixture Model for Visual Tracking
abstract
Correlation filters (CF), a popular tool for visual tracking, suffer from unwanted boundary effects due to the periodic assumption needed for FFT implementation. To address this issue, spatially regularized discriminative correlation filters (SRDCF) have been proposed by introducing a weighting matrix to the regularization term. However, the existing design of spatial weighting matrix is often heuristic and non-adaptive. Inspired by recent advances in joint discrimination and reliability learning for correlation tracking, we propose a principled Bayesian correlation filter learning method using Gaussian scale mixture (GSM) model. The key idea is to decompose each CF coefficient into the product of a positive scalar multiplier and a Gaussian random variable. Treating positive multipliers as weighting coefficients, GSM-based modeling of CFs leads to a spatially adaptive regularization strategy with improved capability of handling various appearance-related uncertainty factors (e.g., scale variation, out-of-plane rotation, and motion blur). Moreover, by imposing a sparse prior over the multipliers, we can jointly learn multipliers and CFs under a unified Bayesian estimation framework. Structured GSM model allows us to better exploit the spatial correlations among CFs and further improve the tracking performance. Experimental results on OTB-2013, OTB-2015, Temple Color-128, VOT-2016, and VOT-2017 show that our tracking method performs favorably when compared with current state-of-the-art methods.
Guangming Shi, Tianzhu Zhang 0001, Weisheng Dong, Jinjian Wu, Xuemei Xie, Xin Li 0005
IEEE Trans. Circuits Syst. Video Technol.5
2022 Blind Image Quality Index for Authentic Distortions With Local and Global Deep Feature Aggregation
abstract
Blind image quality assessment (BIQA) for authentic distortions is still a great challenge, even in today’s deep learning era. It has been widely acknowledged that local and global features are both indispensable for IQA, which play complementary roles. While combining local and global features is straightforward in traditional handcrafted feature-based IQA metrics, it is not an easy task in the deep learning framework. This is mainly due to the fact that deep neural networks typically require input images with a fixed size. Current metrics either resize the image or use local patches as input, which are problematic in that they cannot integrate local and global aspects as well as their interactions to achieve comprehensive quality evaluation. Motivated by the above facts, this paper presents a new BIQA metric for authentic distortions by aggregating local and global deep features in a Vision-Transformer framework. In the proposed metric, selective local regions and global content are simultaneously input for complementary feature extraction, and the Vision-Transformer is employed to build the relationship between different local patches and image quality. Self-attention mechanism is further adopted to explore the interaction between local and global deep features, producing the final image quality score. Extensive experiments on five authentically distorted IQA databases demonstrate that the proposed metric outperforms the state-of-the-arts in terms of both prediction performance and generalization ability.
Leida Li, Tianshu Song, Jinjian Wu, Weisheng Dong, Jiansheng Qian, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.3
2022 Spatiotemporal Representation Learning for Blind Video Quality Assessment
abstract
Blind video quality assessment (BVQA) is of great importance for video-related applications, yet still challenging even in this deep learning era. The difficulty lies in the shortage of large-scale labeled data, thus making it hard to train a robust spatiotemporal encoder for BVQA. To relieve such difficulty, we first build a video dataset, which contains over 320K samples suffering from various compression and transmission artifacts. While manually annotating the dataset with subjective perception is much labor-intensive and time-consuming, we adopt reference-based VQA algorithms to weakly label the data automatically. We consider that single weak label is derived from single knowledge, which is deficient and incomplete for VQA. To alleviate the bias from single weak label (i.e., single knowledge) in the weakly labeled dataset, we propose HEterogeneous Knowledge Ensemble (HEKE) for spatiotemporal representation learning. Compared to learning from single knowledge, learning with HEKE is thought to achieve a lower infimum theoretically, and obtain richer representation. On the basis of the built dataset and the HEKE methodology, a feature encoder specific to BVQA is formed, and directly extract spatiotemporal representation from videos. Then, the video quality can be either acquired in a completely BVQA manner without ground truth, or via a finetuning-based regressor with labels. Extensive experiments on various VQA databases show that our BVQA model with the pretrained encoder achieves the state-of-the-art performance. More surprisingly, even trained on the synthetic data, our model still shows competitive performance on authentic databases. The data and source code will be available athttps://github.com/Sissuire/BVQA-HEKE.
Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.2
2022 Generalizable No-Reference Image Quality Assessment via Deep Meta-Learning
abstract
Recently, researchers have shown great interest in using convolutional neural networks (CNNs) for no-reference image quality assessment (NR-IQA). Due to the lack of big training data, the efforts of existing metrics in optimizing CNN-based NR-IQA models remain limited. Furthermore, the diversity of distortions in images result in the generalization problem of NR-IQA models when trained with known distortions and tested on unseen distortions, which is an easy task for human. Hence, we propose a NR-IQA metric via deep meta-learning, which is highly generalizable in the face of unseen distortions. The fundamental idea is to learn the meta-knowledge shared by human when evaluating the quality of images with diversified distortions. Specifically, we define NR-IQA of different distortions as a series of tasks and propose a task selection strategy to build two task sets, which are characterized by synthetic to synthetic and synthetic to authentic distortions, respectively. Based on these two task sets, an optimization-based meta-learning is proposed to learn the generalized NR-IQA model, which can be directly used to evaluate the quality of images with unseen distortions. Extensive experiments demonstrate that our NR-IQA metric outperforms the state-of-the-arts in terms of both evaluation performance and generalization ability.
Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi
IEEE Trans. Circuits Syst. Video Technol.3
2022 Personalized Image Aesthetics Assessment via Meta-Learning With Bilevel Gradient Optimization
abstract
Typical image aesthetics assessment (IAA) is modeled for the generic aesthetics perceived by an "average" user. However, such generic aesthetics models neglect the fact that users' aesthetic preferences vary significantly depending on their unique preferences. Therefore, it is essential to tackle the issue for personalized IAA (PIAA). Since PIAA is a typical small sample learning (SSL) problem, existing PIAA models are usually built by fine-tuning the well-established generic IAA (GIAA) models, which are regarded as prior knowledge. Nevertheless, this kind of prior knowledge based on "average aesthetics" fails to incarnate the aesthetic diversity of different people. In order to learn the shared prior knowledge when different people judge aesthetics, that is, learn how people judge image aesthetics, we propose a PIAA method based on meta-learning with bilevel gradient optimization (BLG-PIAA), which is trained using individual aesthetic data directly and generalizes to unknown users quickly. The proposed approach consists of two phases: 1) meta-training and 2) meta-testing. In meta-training, the aesthetics assessment of each user is regarded as a task, and the training set of each task is divided into two sets: 1) support set and 2) query set. Unlike traditional methods that train a GIAA model based on average aesthetics, we train an aesthetic meta-learner model by bilevel gradient updating from the support set to the query set using many users' PIAA tasks. In meta-testing, the aesthetic meta-learner model is fine-tuned using a small amount of aesthetic data of a target user to obtain the PIAA model. The experimental results show that the proposed method outperforms the state-of-the-art PIAA metrics, and the learned prior model of BLG-PIAA can be quickly adapted to unseen PIAA tasks.
Hancheng Zhu, Leida Li, Jinjian Wu, Sicheng Zhao, Guiguang Ding, Guangming Shi
IEEE Trans. Cybern.3
2022 Discriminative Multiple-Instance Hyperspectral Subpixel Target Characterization
abstract
Subpixel target detection in hyperspectral imagery is challenging since subpixel targets are smaller in size than the resolution of a single pixel and accurate pixel-level labels on subpixel targets are often unavailable. In particular, this article addresses the problem of learning a prime prototype target signature from imprecisely labeled highly mixed hyperspectral data. Two algorithms, multiple-instance subpixel adaptive cosine estimator (MI-SPACE) and multiple-instance subpixel spectral matched filter (MI-SPSMF), based on multiple-instance learning framework are presented. The proposed methods aim to learn a discriminative prime target signature by maximizing the posterior detection statistics of subpixel hyperspectral targets for the correspondingly proposed subpixel adaptive cosine estimator (SPACE) and subpixel spectral matched filter (SPSMF) detectors, which are also developed in this article. Experimental results demonstrate the effectiveness of the proposed methods on both simulated and real-field hyperspectral subpixel target detection tasks.
Changzhe Jiao, Bo Yang 0047, Qi Wang 0053, Guozhen Wang, Jinjian Wu
IEEE Trans. Geosci. Remote. Sens.5
2022 Contrastive Self-Supervised Pre-Training for Video Quality Assessment
abstract
Video quality assessment (VQA) task is an ongoing small sample learning problem due to the costly effort required for manual annotation. Since existing VQA datasets are of limited scale, prior research tries to leverage models pre-trained on ImageNet to mitigate this kind of shortage. Nonetheless, these well-trained models targeting on image classification task can be sub-optimal when applied on VQA data from a significantly different domain. In this paper, we make the first attempt to perform self-supervised pre-training for VQA task built upon contrastive learning method, targeting at exploiting the plentiful unlabeled video data to learn feature representation in a simple-yet-effective way. Specifically, we implement this idea by first generating distorted video samples with diverse distortion characteristics and visual contents based on the proposed distortion augmentation strategy. Afterwards, we conduct contrastive learning to capture quality-aware information by maximizing the agreement on feature representations of future frames and their corresponding predictions in the embedding space. In addition, we further introduce distortion prediction task as an additional learning objective to push the model towards discriminating different distortion categories of the input video. Solving these prediction tasks jointly with the contrastive learning not only provides stronger surrogate supervision signals, but also learns the shared knowledge among the prediction tasks. Extensive experiments demonstrate that our approach sets a new state-of-the-art in self-supervised learning for VQA task. Our results also underscore that the learned pre-trained model can significantly benefit the existing learning based VQA models. Source code is available at https://github.com/cpf0079/CSPT.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi
IEEE Trans. Image Process.3
2022 Deep Posterior Distribution-Based Embedding for Hyperspectral Image Super-Resolution
abstract
In this paper, we investigate the problem of hyperspectral (HS) image spatial super-resolution via deep learning. Particularly, we focus on how to embed the high-dimensional spatial-spectral information of HS images efficiently and effectively. Specifically, in contrast to existing methods adopting empirically-designed network modules, we formulate HS embedding as an approximation of the posterior distribution of a set of carefully-defined HS embedding events, including layer-wise spatial-spectral feature extraction and network-level feature aggregation. Then, we incorporate the proposed feature embedding scheme into a source-consistent super-resolution framework that is physically-interpretable, producing PDE-Net, in which high-resolution (HR) HS images are iteratively refined from the residuals between input low-resolution (LR) HS images and pseudo-LR-HS images degenerated from reconstructed HR-HS images via probability-inspired HS embedding. Extensive experiments over three common benchmark datasets demonstrate that PDE-Net achieves superior performance over state-of-the-art methods. Besides, the probabilistic characteristic of this kind of networks can provide the epistemic uncertainty of the network outputs, which may bring additional benefits when used for other HS image-based applications. The code will be publicly available at https://github.com/jinnh/PDE-Net.
Jinhui Hou, Junhui Hou, Huanqiang Zeng, Jinjian Wu, Jiantao Zhou 0001
IEEE Trans. Image Process.5
2022 Fine-Grained Image Quality Caption With Hierarchical Semantics Degradation
abstract
Blind image quality assessment (BIQA), which is capable of precisely and automatically estimating human perceived image quality with no pristine image for comparison, attracts extensive attention and is of wide applications. Recently, many existing BIQA methods commonly represent image quality with a quantitative value, which is inconsistent with human cognition. Generally, human beings are good at perceiving image quality in terms of semantic description rather than quantitative value. Moreover, cognition is a needs-oriented task where humans are able to extract image contents with local to global semantics as they need. The mediocre quality value represents coarse or holistic image quality and fails to reflect degradation on hierarchical semantics. In this paper, to comply with human cognition, a novel quality caption model is inventively proposed to measure fine-grained image quality with hierarchical semantics degradation. Research on human visual system indicates there are hierarchy and reverse hierarchy correlations between hierarchical semantics. Meanwhile, empirical evidence shows that there are also bi-directional degradation dependencies between them. Thus, a novel bi-directional relationship-based network (BDRNet) is proposed for semantics degradation description, through adaptively exploring those correlations and degradation dependencies in a bi-directional manner. Extensive experiments demonstrate that our method outperforms the state-of-the-arts in terms of both evaluation performance and generalization ability.
Wen Yang 0008, Jinjian Wu, Shiwei Tian, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Image Process.2
2022 Video Quality Assessment With Serial Dependence Modeling
abstract
Video quality assessment (VQA) is much more challenging than image quality assessment, due to the difficulty of modeling temporal influence among frames. Most of the existing VQA methods usually isolate each moment within the video (i.e., it neglects the sequential nature), leading to a large gap from the subjective perception. Recent research on neuroscience suggests a serially dependent perception (SDP) mechanism in the human visual system (HVS). Namely, the HVS tends to incorporate the recent past visual experience to predict the present perception. Inspired by the SDP, we suggest that the HVS prefers stable and continuous degradations in videos due to their predictability, and exhibits less tolerance to interrupted and unpredictable disturbances. Thus, we introduce a novel serial dependence modeling (SDM) framework for full-reference VQA in this paper. Firstly, the instantaneous degradation is measured on both the static appearance and motion information for each glimpse of scenes. Since motion plays an important role in videos, two types of structures are extracted for motion representation, namely, an explicit content-based 3D structure and an implicit feature-based 2D structure. Next, an assessment-directed long-short term memory (A-LSTM) is proposed to capture the serial dependence among instantaneous degradations. With the consideration of the perceptual effect from the previous moment on the current one, especially the effect from the perceptually worst moment, the serially dependent degradation is characterized. Finally, by mimicking the subjective rating for video-viewing, an attention-based quality decision procedure is presented to acquire the final video quality. Experimental results on publicly available VQA databases demonstrate that the proposed method maintains good consistency with the subjective perception.
Yongxu Liu 0001, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin
IEEE Trans. Multim.2
2021 Deep Gaussian Scale Mixture Prior for Spectral Compressive Imaging
abstract
In coded aperture snapshot spectral imaging (CASSI) system, the real-world hyperspectral image (HSI) can be reconstructed from the captured compressive image in a snapshot. Model-based HSI reconstruction methods employed hand-crafted priors to solve the reconstruction problem, but most of which achieved limited success due to the poor representation capability of these hand-crafted priors. Deep learning based methods learning the mappings between the compressive images and the HSIs directly achieved much better results. Yet, it is nontrivial to design a powerful deep network heuristically for achieving satisfied results. In this paper, we propose a novel HSI reconstruction method based on the Maximum a Posterior (MAP) estimation framework using learned Gaussian Scale Mixture (GSM) prior. Different from existing GSM models using hand-crafted scale priors (e.g., the Jeffrey’s prior), we propose to learn the scale prior through a deep convolutional neural network (DCNN). Furthermore, we also propose to estimate the local means of the GSM models by the DCNN. All the parameters of the MAP estimation algorithm and the DCNN parameters are jointly optimized through end-to-end training. Extensive experimental results on both synthetic and real datasets demonstrate that the proposed method outperforms existing state-of-the-art methods. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/DGSM-SCI.htm.
Weisheng Dong, Xin Yuan 0002, Jinjian Wu, Guangming Shi
CVPR4
2021 Unsupervised Curriculum Domain Adaptation for No-Reference Video Quality Assessment
abstract
During the last years, convolutional neural networks (C-NNs) have triumphed over video quality assessment (VQA) tasks. However, CNN-based approaches heavily rely on annotated data which are typically not available in VQA, leading to the difficulty of model generalization. Recent advances in domain adaptation technique makes it possible to adapt models trained on source data to unlabeled target data. However, due to the distortion diversity and content variation of the collected videos, the intrinsic subjectivity of VQA tasks hampers the adaptation performance. In this work, we propose a curriculum-style unsupervised domain adaptation to handle the cross-domain no-reference VQA problem. The proposed approach could be divided into two stages. In the first stage, we conduct an adaptation between source and target domains to predict the rating distribution for target samples, which can better reveal the subjective nature of VQA. From this adaptation, we split the data in target domain into confident and uncertain subdomains using the proposed uncertainty-based ranking function, through measuring their prediction confidences. In the second stage, by regarding samples in confident subdomain as the easy tasks in the curriculum, a fine-level adaptation is conducted between two subdomain-s to fine-tune the prediction model. Extensive experimental results on benchmark datasets highlight the superiority of the proposed method over the competing methods in both accuracy and speed. The source code is released at https://github.com/cpf0079/UCDA.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi
ICCV3
2021 Multiple Instance Constrained Energy Minimization for Discriminative Hyperspectral Target Characterization
abstract
In hyperspectral imagery, target detection is challenging since lots of pixels are a mixture of more than one distinct substance and precise pixel-wise labels are often infeasible to obtain. To address this problem, the Multiple Instance Constrained Energy Minimization (MI-CEM) for estimating a discriminative target signature from inaccurately labeled and mixed hyperspectral data is introduced in this paper. The proposed method maximizes the posterior detection statistics of the constrained energy minimization sub-pixel detector and estimates a discriminative target signature. The learned target signature can be applied to CEM for sub-pixel target detection. Experiments on both simulated and real-world data demonstrate that MI-CEM achieves competitive performance compared with the state-of-the-art algorithms.
Changzhe Jiao, Bo Yang 0047, Jinjian Wu
IGARSS3
2021 No-Reference Video Quality Assessment with Heterogeneous Knowledge Ensemble
abstract
Blind assessment of video quality is still challenging even in this deep learning era. The limited number of samples in existing databases is insufficient to learn a good feature extractor for video quality assessment (VQA), while manually labeling a larger database with subjective perception is very labor-intensive and time-consuming. To relieve such difficulty, we first collect 3589 high-quality video clips as the reference and build a large VQA dataset. The dataset contains more than 300K samples degraded by various distortion types due to compression and transmission error, and provides weak labels for each distorted sample with several full-reference VQA algorithms. To learn effective representation from the weakly labeled data, we alleviate the bias of single weak label (i.e., single knowledge) via learning from multiple heterogeneous knowledge. To this end, we propose a novel no-reference VQA (NR-VQA) method with HEterogeneous Knowledge Ensemble (HEKE). Comparing to learning from single knowledge, HEKE can theoretically reach a lower infimum, and learn richer representation due to the heterogeneity. Extensive experimental results show that the proposed HEKE outperforms existing NR-VQA methods, and achieves the state-of-the-art performance. The source code will be available at https://github.com/Sissuire/BVQA-HEKE.
Jinjian Wu, Yongxu Liu 0001, Leida Li, Weisheng Dong, Guangming Shi
ACM Multimedia1
2021 Image Quality Caption with Attentive and Recurrent Semantic Attractor Network
abstract
In this paper, a novel quality caption model is inventively developed to assess the image quality with hierarchical semantics. Existing image quality assessment (IQA) methods usually represent image quality with a quantitative value, resulting in inconsistency with human cognition. Generally, human beings are good at perceiving image quality in terms of semantic description rather than quantitative value. Moreover, cognition is a needs-oriented task where hierarchical semantics are extracted. The mediocre quality value fails to reflect degradations on hierarchical semantics. Therefore, a new IQA framework is proposed to describe the quality for needs-oriented cognition. A novel quality caption procedure is firstly introduced, in which the quality is represented as patterns of activation distributed across the diverse degradations on hierarchical semantics. Then, an attentive and recurrent semantic attractor network (ARSANet) is designed to activate the distributed patterns for image quality description. Experiments demonstrate that our method achieves superior performance and is highly compliant with human cognition.
Wen Yang 0008, Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi
ACM Multimedia2
2021 Uncertainty-Driven Loss for Single Image Super-Resolution
abstract
In low-level vision such as single image super-resolution (SISR), traditional MSE or L1 loss function treats every pixel equally with the assumption that the importance of all pixels is the same. However, it has been long recognized that texture and edge areas carry more important visual information than smooth areas in photographic images. How to achieve such spatial adaptation in a principled manner has been an open problem in both traditional model-based and modern learning-based approaches toward SISR. In this paper, we propose a new adaptive weighted loss for SISR to train deep networks focusing on challenging situations such as textured and edge pixels with high uncertainty. Specifically, we introduce variance estimation characterizing the uncertainty on a pixel-by-pixel basis into SISR solutions so the targeted pixels in a high-resolution image (mean) and their corresponding uncertainty (variance) can be learned simultaneously. Moreover, uncertainty estimation allows us to leverage conventional wisdom such as sparsity prior for regularizing SISR solutions. Ultimately, pixels with large certainty (e.g., texture and edge pixels) will be prioritized for SISR according to their importance to visual quality. For the first time, we demonstrate that such uncertainty-driven loss can achieve better results than MSE or L1 loss for a wide range of network architectures. Experimental results on three popular SISR networks show that our proposed uncertainty-driven loss has achieved better PSNR performance than traditional loss functions without any increased computation during testing. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/UDL-SR.htm
Qian Ning, Weisheng Dong, Xin Li 0005, Jinjian Wu, Guangming Shi
NeurIPS4
2021 Deep Maximum a Posterior Estimator for Video Denoising
Weisheng Dong, Xin Li 0005, Jinjian Wu, Leida Li, Guangming Shi
Int. J. Comput. Vis.4
2021 Blind image quality prediction with hierarchical feature aggregation
Jinjian Wu, Wen Yang 0008, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin
Inf. Sci.1
2021 Robust subspace clustering network with dual-domain regularization
Guangming Shi, Xin Li 0005, Weisheng Dong, Jinjian Wu
Pattern Recognit. Lett.6
2021 Predicting the Quality of View Synthesis With Color-Depth Image Fusion
abstract
With the increasing prevalence of free-viewpoint video applications, virtual view synthesis has attracted extensive attention. In view synthesis, a new viewpoint is generated from the input color and depth images with a depth-image-based rendering (DIBR) algorithm. Current quality evaluation models for view synthesis typically operate on the synthesized images, i.e. after the DIBR process, which is computationally expensive. So a natural question is that can we infer the quality of DIBR-based synthesized images using the input color and depth images directly without performing the intricate DIBR operation. With this motivation, this paper presents a no-reference image quality prediction model for view synthesis via COlor-Depth Image Fusion, dubbed CODIF, where the actual DIBR is not needed. First, object boundary regions are detected from the color image, and a Wavelet-based image fusion method is proposed to imitate the interaction between color and depth images during the DIBR process. Then statistical features of the interactional regions and natural regions are extracted from the fused color-depth image to portray the influences of distortions in color/depth images on the quality of synthesized views. Finally, all statistical features are utilized to learn the quality prediction model for view synthesis. Extensive experiments on public view synthesis databases demonstrate the advantages of the proposed metric in predicting the quality of view synthesis, and it even suppresses the state-of-the-art post-DIBR view synthesis quality metrics.
Leida Li, Yipo Huang, Jinjian Wu, Ke Gu 0001, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Temporal Reasoning Guided QoE Evaluation for Mobile Live Video Broadcasting
abstract
Quality of experience (QoE) that serves as a direct evaluation of viewing experience from the end users is of vital importance for network optimization, and should be constantly monitored. Unlike existing video-on-demand streaming services, real-time interactivity is critical to the mobile live broadcasting experience for both broadcasters and their audiences. While existing QoE metrics that are validated on limited video contents and synthetic stall patterns have shown effectiveness in their trained QoE benchmarks, a common caveat is that they often encounter challenges in practical live broadcasting scenarios, where one needs to accurately understand the activity in the video with fluctuating QoE and figure out what is going to happen to support the real-time feedback to the broadcaster. In this paper, we propose a temporal relational reasoning guided QoE evaluation approach for mobile live video broadcasting, namely TRR-QoE, which explicitly attends to the temporal relationships between consecutive frames to achieve a more comprehensive understanding of the distortion-aware variation. In our design, video frames are first processed by deep neural network (DNN) to extract quality-indicative features. Afterwards, besides explicitly integrating features of individual frames to account for the spatial distortion information, multi-scale temporal relational information corresponding to diverse temporal resolutions are made full use of to capture temporal-distortion-aware variation. As a result, the overall QoE prediction could be derived by combining both aspects. The results of experiments conducted on a number of benchmark databases demonstrate the superiority of TRR-QoE over the representative state-of-the-art metrics.
Pengfei Chen 0003, Leida Li, Jinjian Wu, Yabin Zhang 0002, Weisi Lin
IEEE Trans. Image Process.3
2021 Model-Guided Deep Hyperspectral Image Super-Resolution
abstract
The trade-off between spatial and spectral resolution is one of the fundamental issues in hyperspectral images (HSI). Given the challenges of directly acquiring high-resolution hyperspectral images (HR-HSI), a compromised solution is to fuse a pair of images: one has high-resolution (HR) in the spatial domain but low-resolution (LR) in spectral-domain and the other vice versa. Model-based image fusion methods including pan-sharpening aim at reconstructing HR-HSI by solving manually designed objective functions. However, such hand-crafted prior often leads to inevitable performance degradation due to a lack of end-to-end optimization. Although several deep learning-based methods have been proposed for hyperspectral pan-sharpening, HR-HSI related domain knowledge has not been fully exploited, leaving room for further improvement. In this paper, we propose an iterative Hyperspectral Image Super-Resolution (HSISR) algorithm based on a deep HSI denoiser to leverage both domain knowledge likelihood and deep image prior. By taking the observation matrix of HSI into account during the end-to-end optimization, we show how to unfold an iterative HSISR algorithm into a novel model-guided deep convolutional network (MoG-DCN). The representation of the observation matrix by subnetworks also allows the unfolded deep HSISR network to work with different HSI situations, which enhances the flexibility of MoG-DCN. Extensive experimental results are reported to demonstrate that the proposed MoG-DCN outperforms several leading HSISR methods in terms of both implementation cost and visual quality. The code is available at https://see.xidian.edu.cn/faculty/wsdong/Projects/MoG-DCN.htm.
Weisheng Dong, Chen Zhou 0005, Jinjian Wu, Guangming Shi, Xin Li 0005
IEEE Trans. Image Process.4
2021 Blind Image Quality Assessment With Active Inference
abstract
Blind image quality assessment (BIQA) is a useful but challenging task. It is a promising idea to design BIQA methods by mimicking the working mechanism of human visual system (HVS). The internal generative mechanism (IGM) indicates that the HVS actively infers the primary content (i.e., meaningful information) of an image for better understanding. Inspired by that, this paper presents a novel BIQA metric by mimicking the active inference process of IGM. Firstly, an active inference module based on the generative adversarial network (GAN) is established to predict the primary content, in which the semantic similarity and the structural dissimilarity (i.e., semantic consistency and structural completeness) are both considered during the optimization. Then, the image quality is measured on the basis of its primary content. Generally, the image quality is highly related to three aspects, i.e., the scene information (content-dependency), the distortion type (distortion-dependency), and the content degradation (degradation-dependency). According to the correlation between the distorted image and its primary content, the three aspects are analyzed and calculated respectively with a multi-stream convolutional neural network (CNN) based quality evaluator. As a result, with the help of the primary content obtained from the active inference and the comprehensive quality degradation measurement from the multi-stream CNN, our method achieves competitive performance on five popular IQA databases. Especially in cross-database evaluations, our method achieves significant improvements.
Jupo Ma, Jinjian Wu, Leida Li, Weisheng Dong, Xuemei Xie, Guangming Shi, Weisi Lin
IEEE Trans. Image Process.2
2021 Quality Evaluation for Image Retargeting With Instance Semantics
abstract
To meet the ever-increasing demand for devices with diversified displays, image retargeting has become a prevalent technique for adaptive image resizing. In practice, the retargeting operation inevitably causes impairments in the images; thus, image retargeting quality assessment (IRQA) is urgently needed and, can be used to guide algorithm optimization, selection and design. Unlike traditional image quality assessment, image retargeting introduces geometric distortions, which typically affect high-level image semantics. With this motivation, this paper presents a quality evaluation model for image retargeting based on INstance SEMantics (INSEM). Considering that the human visual system (HVS) perceives images highly dependent on apprehensible areas and that impairments in image retargeting mainly degrade the salient instances, an image instance is utilized as the basic semantic unit, and a top-down method is devised to extract instance-level semantic features for IRQA. In addition, taking into account the influence of semantic categories on the perception of retargeting quality, we further propose Semantic-based self-adaptive pooling (SSAP) to integrate instance-based semantic features. Finally, global features are incorporated to generate quality scores that are more consistent with people's perceptions. Extensive experiments and comparisons of three public databases, in terms of both intradatabase and cross-database settings, demonstrate the superiority of the proposed metric over state-of-the-art methods.
Leida Li, Jinjian Wu, Lin Ma 0002, Yuming Fang 0001
IEEE Trans. Multim.3
2021 Quality Index for View Synthesis by Measuring Instance Degradation and Global Appearance
abstract
Virtual view synthesis plays a vital role in the application of multi-view and free-viewpoint videos. Depth-image-based rendering (DIBR) is the most commonly used approach in view synthesis, and many DIBR algorithms have been proposed. However, how to evaluate the quality of DIBR-synthesized images and benchmark the DIBR algorithms are still very challenging, which may hinder the further development of the view synthesis technique. Hence, an effective quality metric for evaluating the distortions in view synthesis is urgently needed. With this motivation, this paper presents a quality index for view synthesis by simultaneously measuring local Instance DEgradation and global Appearance (IDEA). Due to the imperfection of rendering algorithms, local geometric distortions are easily introduced around instance contours, causing instance degradation, which is the dominant distortion in synthesized views. In this work, image instances are first detected and local instance degradation is measured based on discrete orthogonal moments. Meantime, we propose to measure the global appearance of synthesized images based on the superpixel representation. By integrating both local and global aspects of the distortions, a more accurate quality model is built for view synthesis. Extensive experiments and comparisons have demonstrated the superiority of the proposed method in evaluating the quality of DIBR-synthesized images and benchmarking the performance of view synthesis algorithms.
Leida Li, Yu Zhou 0009, Jinjian Wu, Fu Li 0002, Guangming Shi
IEEE Trans. Multim.3
2021 Probabilistic Undirected Graph Based Denoising Method for Dynamic Vision Sensor
abstract
Dynamic Vision Sensor (DVS) is a new type of neuromorphic event-based sensor, which has an innate advantage in capturing fast-moving objects. Due to the interference of DVS hardware itself and many external factors, noise is unavoidable in the output of DVS. Different from frame/image with structural data, the output of DVS is in the form of address-event representation (AER), which means that the traditional denoising methods cannot be used for the output (i.e., event stream) of the DVS. In this paper, we propose a novel event stream denoising method based on probabilistic undirected graph model (PUGM). The motion of objects always shows a certain regularity/trajectory in space and time, which reflects the spatio-temporal correlation between effective events in the stream. Meanwhile, the event stream of DVS is composed by the effective events and random noise. Thus, a probabilistic undirected graph model is constructed to describe such priori knowledge (i.e., spatio-temporal correlation). The undirected graph model is factorized into the product of the cliques energy function, and the energy function is defined to obtain the complete expression of the joint probability distribution. Better denoising effect means a higher probability (lower energy), which means the denoising problem can be transfered into energy optimization problem. Thus, the iterated conditional modes (ICM) algorithm is used to optimize the model to remove the noise. Experimental results on denoising show that the proposed algorithm can effectively remove noise events. Moreover, with the preprocessing of the proposed algorithm, the recognition accuracy on AER data can be remarkably promoted.
Jinjian Wu, Chuanwei Ma, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Multim.1
2020 Spatial-Temporal Gaussian Scale Mixture Modeling for Foreground Estimation
abstract
Subtracting the backgrounds from the video frames is an important step for many video analysis applications. Assuming that the backgrounds are low-rank and the foregrounds are sparse, the robust principle component analysis (RPCA)-based methods have shown promising results. However, the RPCA-based methods suffered from the scale issue, i.e., the ℓ1-sparsity regularizer fails to model the varying sparsity of the moving objects. While several efforts have been made to address this issue with advanced sparse models, previous methods cannot fully exploit the spatial-temporal correlations among the foregrounds. In this paper, we proposed a novel spatial-temporal Gaussian scale mixture (STGSM) model for foreground estimation. In the proposed STGSM model, a temporal consistent constraint is imposed over the estimated foregrounds through nonzero-means Gaussian models. Specifically, the estimates of the foregrounds obtained in the previous frame are used as the prior for these of the current frame, and nonzero means Gaussian scale mixture models (GSM) are developed. To better characterize the temporal correlations, the optical flow has been used to model the correspondences between foreground pixels in adjacent frames. The spatial correlations have also been exploited by considering that local correlated pixels should be characterized by the same STGSM model, leading to further performance improvements. Experimental results on real video datasets show that the proposed method performs comparably or even better than current state-of-the-art background subtraction methods.
Qian Ning, Weisheng Dong, Jinjian Wu, Jie Lin 0008, Guangming Shi
AAAI4
2020 MetaIQA: Deep Meta-Learning for No-Reference Image Quality Assessment
abstract
Recently, increasing interest has been drawn in exploiting deep convolutional neural networks (DCNNs) for no-reference image quality assessment (NR-IQA). Despite of the notable success achieved, there is a broad consensus that training DCNNs heavily relies on massive annotated data. Unfortunately, IQA is a typical small sample problem. Therefore, most of the existing DCNN-based IQA metrics operate based on pre-trained networks. However, these pre-trained networks are not designed for IQA task, leading to generalization problem when evaluating different types of distortions. With this motivation, this paper presents a no-reference IQA metric based on deep meta-learning. The underlying idea is to learn the meta-knowledge shared by human when evaluating the quality of images with various distortions, which can then be adapted to unknown distortions easily. Specifically, we first collect a number of NR-IQA tasks for different distortions. Then meta-learning is adopted to learn the prior knowledge shared by diversified distortions. Finally, the quality prior model is fine-tuned on a target NR-IQA task for quickly obtaining the quality model. Extensive experiments demonstrate that the proposed metric outperforms the state-of-the-arts by a large margin. Furthermore, the meta-model learned from synthetic distortions can also be easily generalized to authentic distortions, which is highly desired in real-world applications of IQA metrics.
Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi
CVPR3
2020 Denoising of Event-Based Sensors with Spatial-Temporal Correlation
abstract
As a novel asynchronous-driven cameras, event-based sensors are with high sensitivity, fast speed, low power consumption and low data volume, but with abundant noise. Since the output of event-based sensors is in the form of address-event-representation (AER), the traditional frame-based denoising method cannot be used. In this paper, we introduce a novel event stream denoising method for such sensors. Effective events tend to show temporal and spatial regularity, while noise events show a kind of randomness. Thus, we build a probabilistic undirected graph model to describe this difference, with which the denoising problem is converted to a probability maximization problem. Then, the model is decomposed into the product of the energy function on the maximum cliques, and the iterated condition model (ICM) is used for energy minimization to obtain the denoised event stream. Experiments show that our method can effectively remove noise events directly from the event stream and significantly improve event recognition rate.
Jinjian Wu, Chuanwei Ma, Xiaojie Yu, Guangming Shi
ICASSP1
2020 Active Inference of GAN for No-Reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) is a challenging task. It is a promising idea to design NR-IQA algorithms by mimicking how human visual system (HVS) works. The internal generative mechanism (IGM) indicates that HVS actively infers the primary content of an image for better understanding. Inspired by that, a novel NR-IQA method with active inference is proposed in this paper. First, a generative adversarial network (GAN) is proposed to predict the primary content of a distorted image, in which two IGM-inspired constraints are considered during the optimization. Next, based on the correlation between the distorted image and its primary content, different degradations (i.e., the content/distortion-/structure-dependency degradation) are measured simultaneously with a multi-stream convolutional neural network (CNN) for NR-IQA. Benefit from the primary content obtained from GAN and the multiple degradations measurement of CNN, our method achieves the state-of-the-art on five public IQA databases.
Jupo Ma, Jinjian Wu, Leida Li, Weisheng Dong, Xuemei Xie
ICME2
2020 RIRNet: Recurrent-In-Recurrent Network for Video Quality Assessment
abstract
Video quality assessment (VQA), which is capable of automatically predicting the perceptual quality of source videos especially when reference information is not available, has become a major concern for video service providers due to the growing demand for video quality of experience (QoE) by end users. While significant advances have been achieved from the recent deep learning techniques, they often lead to misleading results in VQA tasks given their limitations on describing 3D spatio-temporal regularities using only fixed temporal frequency. Partially inspired by psychophysical and vision science studies revealing the speed tuning property of neurons in visual cortex when performing motion perception (i.e., sensitive to different temporal frequencies), we propose a novel no-reference (NR) VQA framework named Recurrent-In-Recurrent Network (RIRNet) to incorporate this characteristic to prompt an accurate representation of motion perception in VQA task. By fusing motion information derived from different temporal frequencies in a more efficient way, the resulting temporal modeling scheme is formulated to quantify the temporal motion effect via a hierarchical distortion description. It is found that the proposed framework is in closer agreement with quality perception of the distorted videos since it integrates concepts from motion perception in human visual system (HVS), which is manifested in the designed network structure composed of low- and high- level processing. A holistic validation of our methods on four challenging video quality databases demonstrates the superior performances over the state-of-the-art methods.
Pengfei Chen 0003, Leida Li, Lei Ma 0003, Jinjian Wu, Guangming Shi
ACM Multimedia4
2020 No-reference quality index of depth images based on statistics of edge profiles for view synthesis
Leida Li, Jinjian Wu, Shiqi Wang 0001, Guangming Shi
Inf. Sci.3
2020 Subjective and objective quality assessment for image restoration: A critical survey
Bo Hu 0008, Leida Li, Jinjian Wu, Jiansheng Qian
Signal Process. Image Commun.3
2020 End-to-End Blind Image Quality Prediction With Cascaded Deep Neural Network
abstract
The deep convolutional neural network (CNN) has achieved great success in image recognition. Many image quality assessment (IQA) methods directly use recognition-oriented CNN for quality prediction. However, the properties of IQA task is different from image recognition task. Image recognition should be sensitive to visual content and robust to distortion, while IQA should be sensitive to both distortion and visual content. In this paper, an IQA-oriented CNN method is developed for blind IQA (BIQA), which can efficiently represent the quality degradation. CNN is large-data driven, while the sizes of existing IQA databases are too small for CNN optimization. Thus, a large IQA dataset is firstly established, which includes more than one million distorted images (each image is assigned with a quality score as its substitute of Mean Opinion Score (MOS), abbreviated as pseudo-MOS). Next, inspired by the hierarchical perception mechanism (from local structure to global semantics) in human visual system, a novel IQA-orientated CNN method is designed, in which the hierarchical degradation is considered. Finally, by jointly optimizing the multilevel feature extraction, hierarchical degradation concatenation (HDC) and quality prediction in an end-to-end framework, the Cascaded CNN with HDC (named as CaHDC) is introduced. Experiments on the benchmark IQA databases demonstrate the superiority of CaHDC compared with existing BIQA methods. Meanwhile, the CaHDC (with about 0.73M parameters) is lightweight comparing to other CNN-based BIQA models, which can be easily realized in the microprocessing system. The dataset and source code of the proposed method are available at https://web.xidian.edu.cn/wjj/paper.html.
Jinjian Wu, Jupo Ma, Fuhu Liang, Weisheng Dong, Guangming Shi, Weisi Lin
IEEE Trans. Image Process.1
2019 End-to-End Blind Image Quality Assessment with Cascaded Deep Features
abstract
The convolutional neural network (CNN) has achieved great success in many visual tasks. However, it has limited progress on image quality assessment (IQA) due to the lacking of IQA-oriented CNN framework which can efficiently represent the hierarchical quality degradation. In this paper, inspired by the hierarchical perception mechanism (from local structure to global semantics) in the human visual system, we design an end-to-end cascaded CNN framework for blind IQA (BIQA), in which multilevel features are extracted and concatenated to represent the hierarchical quality degradation. By jointly optimizing the feature extraction, hierarchical degradation integration, and quality prediction in an end-to-end manner, the novel cascaded CNN with hierarchical feature integration (CaHFI) for BIQA is designed. Experimental results on five benchmark IQA databases demonstrate that the proposed CaHFI achieves the state-of-the-art. And experiments on cross-database evaluation further prove the high generalization ability of the proposed CaHFI.
Jinjian Wu, Jupo Ma, Fuhu Liang, Weisheng Dong, Guangming Shi
ICME1
2019 A Real-Time Rock-Paper-Scissor Hand Gesture Recognition System Based on FlowNet and Event Camera
Xuemei Xie, Jinjian Wu, Guangming Shi
PRCV (1)3
2019 Survey of visual just noticeable difference estimation
Jinjian Wu, Guangming Shi, Weisi Lin
Frontiers Comput. Sci.1
2019 SISRSet: Single image super-resolution subjective evaluation test and objective quality assessment
Guangming Shi, Wenfei Wan, Jinjian Wu, Xuemei Xie, Weisheng Dong, Hong Ren Wu
Neurocomputing3
2019 No-reference image quality assessment with visual pattern degradation
Jinjian Wu, Man Zhang 0007, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin
Inf. Sci.1
2019 Blind image quality assessment with semantic information
Weiping Ji, Jinjian Wu, Guangming Shi, Wenfei Wan, Xuemei Xie
J. Vis. Commun. Image Represent.2
2019 Blind image quality assessment with hierarchy: Degradation from local structure to deep semantics
Jinjian Wu, Jichen Zeng, Weisheng Dong, Guangming Shi, Weisi Lin
J. Vis. Commun. Image Represent.1
2019 No-Reference Quality Assessment for View Synthesis Using DoG-Based Edge Statistics and Texture Naturalness
abstract
View synthesis is a key technique in free-viewpoint video, which renders virtual views based on texture and depth images. The distortions in synthesized views come from two stages, i.e., the stage of the acquisition and processing of texture and depth images, and the rendering stage using depth-image-based-rendering (DIBR) algorithms. The existing view synthesis quality metrics are designed for the distortions caused by a single stage, which cannot accurately evaluate the quality of the entire view synthesis process. With the considerations that the distortions introduced by two stages both cause edge degradation and texture unnaturalness, and the Difference-of-Gaussian (DoG) representation is powerful in capturing image edge and texture characteristics by simulating the center-surrounding receptive fields of retinal ganglion cells of human eyes, this paper presents a no-reference quality index for Synthesized views using DoG-based Edge statistics and Texture naturalness (SET). To mimic the multi-scale property of the Human Visual System (HVS), DoG images are first calculated at multiple scales. Then the orientation selective statistics features and the texture naturalness features are calculated on the DoG images and the coarsest scale image, producing two groups of quality-aware features. Finally, the quality model is learnt from these features using the random forest regression model. Experimental results on two view synthesis image databases demonstrate that the proposed metric is advantageous over the relevant state-of-the-arts in dealing with the distortions in the whole view synthesis process.
Yu Zhou 0009, Leida Li, Shiqi Wang 0001, Jinjian Wu, Yuming Fang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.4
2019 Quality Assessment for Video With Degradation Along Salient Trajectories
abstract
With the rapid growth of digital video through the Internet, a reliable objective video-quality assessment (VQA) algorithm is in great demand for video management. Motion information plays a dominant role for video perception, and the human visual system (HVS) is able to track moving objects effectively with eye movement. Moreover, the middle temporal area of the brain is selective for moving objects with particular velocities. In other words, visual contents that are along the motion trajectories will automatically attract our attention for dedicated processing. Inspired by the motion-related process in the HVS, we suggest analyzing the degradation along attended motion trajectories for VQA. The characteristic of motion velocity along each trajectory is analyzed for temporal quality measurement. Meanwhile, visual information along each trajectory is extracted for joint spatial-temporal quality measurement. Finally, considering the spatial-quality degradation from each frame, a novel full-reference assessor along salient trajectories (FAST) for VQA (which combines the spatial, temporal, and joint spatial-temporal quality degradations) is introduced. Experimental results on five publicly available VQA databases demonstrate that the proposed FAST VQA model performs consistently with the subjective perception. The source code of the proposed method is available at http://web.xidian.edu.cn/wjj/paper.html.
Jinjian Wu, Yongxu Liu 0001, Weisheng Dong, Guangming Shi, Weisi Lin
IEEE Trans. Multim.1
2018 Super-Resolution Quality Assessment: Subjective Evaluation Database and Quality Index Based on Perceptual Structure Measurement
abstract
With the outstanding performance of deep learning based single image super-resolution (SISR) methods, the traditional SISR evaluation metrics (e.g., PSNR and SSIM, which measure the per-pixel differences and simple structure similarities respectively) are facing great challenges. When assessing SISR algorithms, they generally are hardly consistent with the human visual system (HVS). According to the psychological studies, the HVS presents different sensitivities to the plain, edge and texture regions, which are difficult to be accurately identified and measured with the existing quality indexes, especially for SR images. To deal with this problem, we firstly build a SISR subjective assessment database including several major deep learning based SR methods. Then we propose a more accurate perception structure measurement and use their similarity comparisons to evaluate the SR algorithms. Experimental results on the databases demonstrate that the proposed method performs well consistent with the human visual perception.
Wenfei Wan, Jinjian Wu, Guangming Shi, Weisheng Dong
ICME2
2018 Motion Trajectory based Spatial-Temporal Degradation Measurement for Video Quality Assessment
abstract
With the rapid growth of digital video through the Internet, a reliable video quality assessment (VQA) technology is greatly demanded for video management. Motion information plays a dominant role for video perception, however it is too difficult to be accurately analyzed for VQA. The human visual system (HVS) is highly adaptive to track moving objects with pursuit eye movement. Inspired by the motion process in the HVS, we suggest to analyze the degradation along attended motion trajectories for VQA. As a convenient representation of motion, optical flow is calculated for motion trajectory searching. Next, the degradation on the motion velocity along each trajectory is analyzed with the optical flow for temporal quality measurement. Meanwhile, visual information along each trajectory is extracted for joint spatial-temporal quality measurement. Finally, considering the spatial quality degradation from each frame, a novel VQA model is introduced. Experimental results on the public available VQA databases demonstrate that the proposed VQA model performs highly consistency with the subjective perception.
Jinjian Wu, Yongxu Liu 0001, Guangming Shi
VCIP1
2018 No-reference quality assessment of DIBR-synthesized videos by measuring temporal flickering
Yu Zhou 0009, Leida Li, Shiqi Wang 0001, Jinjian Wu, Yun Zhang 0002
J. Vis. Commun. Image Represent.4
2018 No Reference Quality Assessment for Screen Content Images With Both Local and Global Feature Representation
abstract
In this paper, we propose a novel no reference quality assessment method by incorporating statistical luminance and texture features (NRLT) for screen content images (SCIs) with both local and global feature representation. The proposed method is designed inspired by the perceptual property of the human visual system (HVS) that the HVS is sensitive to luminance change and texture information for image perception. In the proposed method, we first calculate the luminance map through the local normalization, which is further used to extract the statistical luminance features in global scope. Second, inspired by existing studies from neuroscience that high-order derivatives can capture image texture, we adopt four filters with different directions to compute gradient maps from the luminance map. These gradient maps are then used to extract the second-order derivatives by local binary pattern. We further extract the texture feature by the histogram of high-order derivatives in global scope. Finally, support vector regression is applied to train the mapping function from quality-aware features to subjective ratings. Experimental results on the public large-scale SCI database show that the proposed NRLT can achieve better performance in predicting the visual quality of SCIs than relevant existing methods, even including some full reference visual quality assessment methods.
Yuming Fang 0001, Jiebin Yan, Leida Li, Jinjian Wu, Weisi Lin
IEEE Trans. Image Process.4
2018 Image Super-Resolution With Parametric Sparse Model Learning
abstract
Recovering a high-resolution (HR) image from its low-resolution (LR) version is an ill-posed inverse problem. Learning accurate prior of HR images is of great importance to solve this inverse problem. Existing super-resolution (SR) methods either learn a non-parametric image prior from training data (a large set of LR/HR patch pairs) or estimate a parametric prior from the LR image analytically. Both methods have their limitations: the former lacks flexibility when dealing with different SR settings; while the latter often fails to adapt to spatially varying image structures. In this paper, we propose to take a hybrid approach toward image SR by combining those two lines of ideas - that is, a parametric sparse prior of HR images is learned from the training set as well as the input LR image. By exploiting the strengths of both worlds, we can more accurately recover the sparse codes and therefore HR image patches than conventional sparse coding approaches. Experimental results show that the proposed hybrid SR method significantly outperforms existing model-based SR methods and is highly competitive to current state-of-the-art learning-based SR methods in terms of both subjective and objective image qualities.
Weisheng Dong, Xuemei Xie, Guangming Shi, Jinjian Wu, Xin Li 0005
IEEE Trans. Image Process.5
2018 Robust Foreground Estimation via Structured Gaussian Scale Mixture Modeling
abstract
Recovering the background and foreground parts from video frames has important applications in video surveillance. Under the assumption that the background parts are stationary and the foreground are sparse, most of existing methods are based on the framework of robust principal component analysis (RPCA), i.e., modeling the background and foreground parts as a low-rank and sparse matrices, respectively. However, in realistic complex scenarios, the conventional norm sparse regularizer often fails to well characterize the varying sparsity of the foreground components. How to select the sparsity regularizer parameters adaptively according to the local statistics is critical to the success of the RPCA framework for background subtraction task. In this paper, we propose to model the sparse component with a Gaussian scale mixture (GSM) model. Compared with the conventional norm, the GSM-based sparse model has the advantages of jointly estimating the variances of the sparse coefficients (and hence the regularization parameters) and the unknown sparse coefficients, leading to significant estimation accuracy improvements. Moreover, considering that the foreground parts are highly structured, a structured extension of the GSM model is further developed. Specifically, the input frame is divided into many homogeneous regions using superpixel segmentation. By characterizing the set of sparse coefficients in each homogeneous region with the same GSM prior, the local dependencies among the sparse coefficients can be effectively exploited, leading to further improvements for background subtraction. Experimental results on several challenging scenarios show that the proposed method performs much better than most of existing background subtraction methods in terms of both performance and speed.
Guangming Shi, Weisheng Dong, Jinjian Wu, Xuemei Xie
IEEE Trans. Image Process.4
2018 Blind Quality Index for Multiply Distorted Images Using Biorder Structure Degradation and Nonlocal Statistics
abstract
In the past decade, extensive image quality metrics have been proposed. The majority of them are tailored for the images that contain a specific type of distortion. However, in practice, the images are usually degraded by different types of distortions simultaneously. This poses great challenges to the existing quality metrics. Motivated by this, this paper proposes a no-reference quality index for the multiply distorted images using the biorder structure degradation and the nonlocal statistics. The design philosophy is inspired by the fact that the human visual system (HVS) is highly sensitive to the degradations of both the spatial contrast and the spatial distribution, which are prone to be changed by the joint effects of the multiple distortions. Specifically, the multiresolution representation of the image is first built by downsampling to simulate the hierarchical property of the HVS. Then, the structure degradation is calculated to measure the spatial contrast. Considering the fact that the human visual cortex has the separate mechanisms to perceive the first- and second-order structures, dubbed biorder structures, the degradations of biorder structures are calculated to account for the spatial contrast, producing the first group of the quality-aware features. Furthermore, the nonlocal self-similarity statistics is calculated to measure the spatial distribution, producing the second group of features. Finally, all the features are fed into the random forest regression model to learn the quality model for the multiply distorted images. Extensive experimental results conducted on the three public databases demonstrate the superiority of the proposed metric to the state-of-the-art metrics. Moreover, the proposed metric is also advantageous over the existing metrics in terms of the generalization ability.
Yu Zhou 0009, Leida Li, Jinjian Wu, Ke Gu 0001, Weisheng Dong, Guangming Shi
IEEE Trans. Multim.3
2017 No-reference image quality assessment with orientation selectivity mechanism
abstract
No-reference (NR) image quality assessment (IQA) technology is greatly required in quality-orientated visual signal processing systems. However, without the guidance of the reference information, it is still a great challenge for NR IQA to perform consistent with the subjective perception. Researches on cognitive neuroscience state that the human visual system (HVS) presents substantially orientation selectivity mechanism, within which the visual structures are extracted in the local receptive fields for scene understanding. Inspired by this mechanism, a set of orientation selectivity based visual patterns are designed. By analyzing the quality degradation on those patterns, a novel visual pattern degradation based NR IQA method is proposed. Experimental results on large databases demonstrate that the proposed method outperforms the existing NR IQA methods.
Jinjian Wu, Man Zhang 0007, Guangming Shi, Xuemei Xie, Weisi Lin
ICIP1
2017 Saliency change based reduced reference image quality assessment
abstract
The image quality assessment (IQA) technique, which aims to perform coherently with subjective perception, is useful in quality-orientated image processing systems. In this paper, we suggest to take the saliency change into account for reduced reference (RR) IQA model. Generally, a saliency region will attract more attention, and our human vision is more sensitive to quality degradation on such region. Inspired by this, saliency values are firstly used to highlight these sensitive regions, and a local saliency weighted histogram (LSWH) based on visual orientation pattern is generated for visual feature extraction. Next, strong distortion may change the saliency from the reference to the distorted images. Thus, the saliency of each visual orientation pattern is measured, and a global saliency based histogram (GSBH) is created. Finally, by combining the LSWH and GSBH, a novel IQA model for reduced reference is introduced. Experimental results on five publicly available databases demonstrate that the proposed model uses only several values (9 values) as reference information, and performs consistently with subjective perception.
Jinjian Wu, Yongxu Liu 0001, Guangming Shi, Weisi Lin
VCIP1
2017 Bag-of-words feature representation for blind image quality assessment with local quantized pattern
Xuemei Xie, Yazhong Zhang, Jinjian Wu, Guangming Shi, Weisheng Dong
Neurocomputing3
2017 No-reference quality assessment of compressive sensing image recovery
Bo Hu 0008, Leida Li, Jinjian Wu, Shiqi Wang 0001, Lu Tang 0001, Jiansheng Qian
Signal Process. Image Commun.3
2017 Enhanced Just Noticeable Difference Model for Images With Pattern Complexity
abstract
The just noticeable difference (JND) in an image, which reveals the visibility limitation of the human visual system (HVS), is widely used for visual redundancy estimation in signal processing. To determine the JND threshold with the current schemes, the spatial masking effect is estimated as the contrast masking, and this cannot accurately account for the complicated interaction among visual contents. Research on cognitive science indicates that the HVS is highly adapted to extract the repeated patterns for visual content representation. Inspired by this, we formulate the pattern complexity as another factor to determine the total masking effect: the interaction is relatively straightforward with a limited masking effect in a regular pattern, and is complicated with a strong masking effect in an irregular pattern. From the orientation selectivity mechanism in the primary visual cortex, the response of each local receptive field can be considered as a pattern; therefore, in this paper, the orientation that each pixel presents is regarded as the fundamental element of a pattern, and the pattern complexity is calculated as the diversity of the orientation in a local region. Finally, considering both pattern complexity and luminance contrast, a novel spatial masking estimation function is deduced, and an improved JND estimation model is built. Experimental results on comparing with the latest JND models demonstrate the effectiveness of the proposed model, which performs highly consistent with the human perception. The source code of the proposed model is publicly available at http://web.xidian.edu.cn/wjj/en/index.html.
Jinjian Wu, Leida Li, Weisheng Dong, Guangming Shi, Weisi Lin, C.-C. Jay Kuo
IEEE Trans. Image Process.1
2017 Color-Guided Depth Recovery via Joint Local Structural and Nonlocal Low-Rank Regularization
abstract
High-quality depth recovery from RGB-D data has received increasingly more attention in recent years due to their wide applications from depth-based image rendering to three-dimensional imaging and video. Sharp contrast between high-quality color images and low-quality depth maps presents severe challenges to the development of color-guided depth recovery techniques. Previous works have emphasized either locally varying characteristics of color-depth dependence or nonlocal similarities around the discontinuities of the scene geometry. Therefore, it is desirable to exploit both local and nonlocal structural constraints for optimizing the performance of color-guided depth recovery. In this work, we propose a unified variational approach via joint local and nonlocal regularization. The local regularization term consists of two complementary parts-one characterizing the color-depth dependence in the gradient domain and the other in the spatial domain; nonlocal regularization involves a low-rank constraint suitable for large-scale depth discontinuities. Extensive experimental results are reported to show that our approach outperforms several existing state-of-the-art depth recovery methods on both synthetic and real-world data sets.
Weisheng Dong, Guangming Shi, Xin Li 0005, Kefan Peng, Jinjian Wu, Zhenhua Guo 0001
IEEE Trans. Multim.5
2016 Enhanced just noticeable difference model with visual regularity consideration
abstract
Just noticeable difference (JND) reveals the visibility of our human visual system (HVS), below which changes cannot be perceived by the human. Though dozens of JND estimation models have been introduced during the past decade, how to accurately estimate the JND thresholds for different content regions (e.g., edge and texture region) is still an open problem. Research on cognitive science indicates that the HVS is adaptive to extract the visual regularities from an input scene for content perception and understanding. Thus, we analyze the effect of content regularity on visual sensitivity, and suggest that the visual regularity is another important factor that determines the JND threshold. According to the orientation distributions of local regions, the content regularities are firstly calculated. Then, by considering the effect from content regularity, luminance adaptation, and contrast masking, a novel JND model is proposed. Experimental results demonstrate that the proposed model can effectively estimate the JND thresholds of regions with different visual contents.
Jinjian Wu, Guangming Shi, Weisi Lin, C.-C. Jay Kuo
ICASSP1
2016 Visual information measurement with quality assessment
abstract
The quantity of visual information measurement is significant for many perception-oriented signal processing system. The classical Shannon theory, which based on the probability of signal, is useful to measure the quantity of channel information. However, it fails to accurately measure the quantity of visual information of a given image. Image quality refers to the subjective perception on the visual information that an image carried. An image with high quality carries more information than that of a low quality image. Thus, the quality of an image can effectively represent its quantity of visual information. In this paper, we propose a novel visual information measurement, and verify it with quality assessment. Firstly, a dictionary is learned from natural images, in which the content change of each atom is calculated to present its quantity of visual information. Then, a testing image is represented by the dictionary, and the sparse coefficients for each local block in the image are acquired. Finally, according to the sparse coefficients, the quantity of visual information is measured. The information measurement result is verified with the subjective quality score. Experimental results on a large amount of images demonstrate the accuracy of the proposed method for quantity of visual information measurement.
Jinjian Wu, Guangming Shi, Man Zhang 0007, Guanmi Chen
VCIP1
2016 High quality impulse noise removal via non-uniform sampling and autoregressive modelling based super-resolution
abstract
The challenge of image impulse noise removal is to restore spatial details from damaged pixels using remaining ones in random locations. Most existing methods use all uncontaminated pixels within a local window to estimate the centred noisy one via a statistic way. These kinds of methods have two defects. First, all noisy pixels are treated as independent individuals and estimated by their neighbours one by one, with the correlation between their true values ignored. Second, the image structure as a natural feature is usually ignored. This study proposes a new denoising framework, in which all noisy pixels are jointly restored via non‐uniform sampling and supervised piecewise autoregressive modelling based super‐resolution. In this method, the noisy pixels are jointly estimated in groups through solving a well‐designed optimisation problem, in which image structure feature is considered as an important constraint. Another contribution is that piecewise autoregressive model is not simply adopted but carefully designed so that all noise‐free pixels can be used to supervise the model training and optimisation problem solving for higher accuracy. The experimental results demonstrate that the proposed method exhibits good denoising performance in a large noise density range (10–90%).
Xiaotian Wang 0001, Guangming Shi, Jinjian Wu, Fu Li 0002, Yantao Wang
IET Image Process.4
2016 No-reference quality assessment of deblocked images
Leida Li, Yu Zhou 0009, Weisi Lin, Jinjian Wu, Xinfeng Zhang 0001, Beijing Chen
Neurocomputing4
2016 Orientation selectivity based visual pattern for reduced-reference image quality assessment
Jinjian Wu, Weisi Lin, Guangming Shi, Leida Li, Yuming Fang 0001
Inf. Sci.1
2016 Color image quality assessment based on sparse representation and reconstruction residual
Leida Li, Wenhan Xia, Yuming Fang 0001, Ke Gu 0001, Jinjian Wu, Weisi Lin, Jiansheng Qian
J. Vis. Commun. Image Represent.5
2016 Visual structural degradation based reduced-reference image quality assessment
Jinjian Wu, Weisi Lin, Yuming Fang 0001, Leida Li, Guangming Shi, S. Issac Niwas
Signal Process. Image Commun.1
2016 Hyperspectral Image Super-Resolution via Non-Negative Structured Sparse Representation
abstract
Hyperspectral imaging has many applications from agriculture and astronomy to surveillance and mineralogy. However, it is often challenging to obtain high-resolution (HR) hyperspectral images using existing hyperspectral imaging techniques due to various hardware limitations. In this paper, we propose a new hyperspectral image super-resolution method from a low-resolution (LR) image and a HR reference image of the same scene. The estimation of the HR hyperspectral image is formulated as a joint estimation of the hyperspectral dictionary and the sparse codes based on the prior knowledge of the spatial-spectral sparsity of the hyperspectral image. The hyperspectral dictionary representing prototype reflectance spectra vectors of the scene is first learned from the input LR image. Specifically, an efficient non-negative dictionary learning algorithm using the block-coordinate descent optimization technique is proposed. Then, the sparse codes of the desired HR hyperspectral image with respect to learned hyperspectral basis are estimated from the pair of LR and HR reference images. To improve the accuracy of non-negative sparse coding, a clustering-based structured sparse coding method is proposed to exploit the spatial correlation among the learned sparse codes. The experimental results on both public datasets and real LR hypspectral images suggest that the proposed method substantially outperforms several existing HR hyperspectral image recovery techniques in the literature in terms of both objective quality metrics and computational efficiency.
Weisheng Dong, Fazuo Fu, Guangming Shi, Xun Cao, Jinjian Wu, Xin Li 0005
IEEE Trans. Image Process.5
2016 Image Sharpness Assessment by Sparse Representation
abstract
Recent advances in sparse representation show that overcomplete dictionaries learned from natural images can capture high-level features for image analysis. Since atoms in the dictionaries are typically edge patterns and image blur is characterized by the spread of edges, an overcomplete dictionary can be used to measure the extent of blur. Motivated by this, this paper presents a no-reference sparse representation-based image sharpness index. An overcomplete dictionary is first learned using natural images. The blurred image is then represented using the dictionary in a block manner, and block energy is computed using the sparse coefficients. The sharpness score is defined as the variance-normalized energy over a set of selected high-variance blocks, which is achieved by normalizing the total block energy using the sum of block variances. The proposed method is not sensitive to training images, so a universal dictionary can be used to evaluate the sharpness of images. Experiments on six public image quality databases demonstrate the advantages of the proposed method.
Leida Li, Jinjian Wu, Haoliang Li, Weisi Lin, Alex Chichung Kot
IEEE Trans. Multim.3
2015 Reduced-reference image quality assessment based on entropy differences in DCT domain
abstract
Reduced-reference image quality assessment (RR-IQA) algorithm aims to automatically evaluate the image quality using only partial information about the reference image. In this paper, we propose a new RR-IQA metric by employing the entropy features of each frequency band in the DCT domain. It is well known that human eyes have different sensitivity to different bands, and distortions on each band result in individual quality degradations. Therefore, we suggest to separately compute the visual information degradations on different band for quality assessment. The degradations on each DCT band are firstly analyzed according to the entropy difference. And then, the quality score is obtained using the weighted sum of the entropy difference of each band from low frequency to high frequency. Experimental results on several public image databases show that the proposed method uses limited reference data (8 values) and performs highly consistent with human perception.
Yazhong Zhang, Jinjian Wu, Guangming Shi, Xuemei Xie
ISCAS2
2015 GridSAR: Grid strength and regularity for robust evaluation of blocking artifacts in JPEG images
Leida Li, Yu Zhou 0009, Jinjian Wu, Weisi Lin, Haoliang Li
J. Vis. Commun. Image Represent.3
2015 Visual Orientation Selectivity Based Structure Description
abstract
The human visual system is highly adaptive to extract structure information for scene perception, and structure character is widely used in perception-oriented image processing works. However, the existing structure descriptors mainly describe the luminance contrast of a local region, but cannot effectively represent the spatial correlation of structure. In this paper, we introduce a novel structure descriptor according to the orientation selectivity mechanism in the primary visual cortex. Research on cognitive neuroscience indicate that the arrangement of excitatory and inhibitory cortex cells arise orientation selectivity in a local receptive field, within which the primary visual cortex performs visual information extraction for scene understanding. Inspired by the orientation selectivity mechanism, we compute the correlations among pixels in a local region based on the similarities of their preferred orientation. By imitating the arrangement of the excitatory/inhibitory cells, the correlations between a central pixel and its local neighbors are binarized, and the spatial correlation is represented with a set of binary values, which is named the orientation selectivity-based pattern. Then, taking both the gradient magnitude and the orientation selectivity-based pattern into account, a rotation invariant structure descriptor is introduced. The proposed structure descriptor is applied in texture classification and reduced reference image quality assessment, as two different application domains to verify its generality and robustness. Experimental results demonstrate that the orientation selectivity-based structure descriptor is robust to disturbance, and can effectively represent the structure degradation caused by different types of distortion.
Jinjian Wu, Weisi Lin, Guangming Shi, Yazhong Zhang, Weisheng Dong, Zhibo Chen 0001
IEEE Trans. Image Process.1
2014 Reduced-reference image quality assessment with local binary structural pattern
abstract
Reduced-reference (RR) image quality assessment (IQA) aims to use less reference data and achieve higher quality prediction accuracy. Recent researches confirm that the human visual system (HVS) is adapted to extract structural information and is sensitive to structure degradation. Therefore, in this paper, we try to represent image contents with several structural patterns, and measure image quality according to the structural degradation on these patterns. The classic local binary patterns (LBPs) are firstly employed to extract image structures and create LBP based structural histogram. And then, the structural degradation is computed as the histogram distance between the reference and distorted images. Experimental results on three large databases demonstrate that the proposed RR IQA method greatly improved the quality prediction accuracy.
Jinjian Wu, Weisi Lin, Guangming Shi, Long Xu 0001
ISCAS1
2014 Correlation based universal image/video coding loss recovery
Jinjian Wu, Weisi Lin, Guangming Shi, Jimin Xiao
J. Vis. Commun. Image Represent.1
2014 Image Quality Assessment with Degradation on Spatial Structure
abstract
In this letter, we introduce an improved structural degradation based image quality assessment (IQA) method. Most of the existing structural similarity based IQA metrics mainly consider the spatial contrast degradation but have not fully considered the changes on the spatial distribution of structures. Since the human visual system (HVS) is sensitive to degradations on both spatial contrast and spatial distribution, both factors need to be considered for IQA. In order to measure the structural degradation on spatial distribution, the local binary patterns (LBPs) are first employed to extract structural information. And then, the LBP shift between the reference and distorted images is computed, because noise distorts structural patterns. Finally, the spatial contrast degradation on each pair of LBP shifts is calculated for quality assessment. Experimental results on three large benchmark databases confirm that the proposed IQA method is highly consistent with the subjective perception.
Jinjian Wu, Weisi Lin, Guangming Shi
IEEE Signal Process. Lett.1
2013 Visual masking estimation based on structural uncertainty
abstract
A model of visual masking, which reveals the visible threshold of human perception, is useful in perceptual based image/video processing. The existing visual masking formulation, which mainly considers luminance contrast, cannot accurately estimate the visible threshold. Recent researches indicate that human perception is highly adaptive to extract orderly structures and is insensitive to disorderly structures. Therefore, we suggest that the structural characteristic is another determining factor for visual masking, and deduce a novel visual masking function based on structural uncertainty. Experimental results demonstrate that the proposed model is more consistent with human perception than the existing visual masking model.
Jinjian Wu, Weisi Lin, Guangming Shi
ISCAS1
2013 Perceptual Quality Metric With Internal Generative Mechanism
abstract
Objective image quality assessment (IQA) aims to evaluate image quality consistently with human perception. Most of the existing perceptual IQA metrics cannot accurately represent the degradations from different types of distortion, e.g., existing structural similarity metrics perform well on content-dependent distortions while not as well as peak signal-to-noise ratio (PSNR) on content-independent distortions. In this paper, we integrate the merits of the existing IQA metrics with the guide of the recently revealed internal generative mechanism (IGM). The IGM indicates that the human visual system actively predicts sensory information and tries to avoid residual uncertainty for image perception and understanding. Inspired by the IGM theory, we adopt an autoregressive prediction algorithm to decompose an input scene into two portions, the predicted portion with the predicted visual content and the disorderly portion with the residual content. Distortions on the predicted portion degrade the primary visual information, and structural similarity procedures are employed to measure its degradation; distortions on the disorderly portion mainly change the uncertain information and the PNSR is employed for it. Finally, according to the noise energy deployment on the two portions, we combine the two evaluation results to acquire the overall quality score. Experimental results on six publicly available databases demonstrate that the proposed metric is comparable with the state-of-the-art quality metrics.
Jinjian Wu, Weisi Lin, Guangming Shi, Anmin Liu
IEEE Trans. Image Process.1
2013 Pattern Masking Estimation in Image With Structural Uncertainty
abstract
A model of visual masking, which reveals the visibility of stimuli in the human visual system (HVS), is useful in perceptual based image/video processing. The existing visual masking function mainly considers luminance contrast, which always overestimates the visibility threshold of the edge region and underestimates that of the texture region. Recent research on visual perception indicates that the HVS is sensitive to orderly regions that possess regular structures and insensitive to disorderly regions that possess uncertain structures. Therefore, structural uncertainty is another determining factor on visual masking. In this paper, we introduce a novel pattern masking function based on both luminance contrast and structural uncertainty. Through mimicking the internal generative mechanism of the HVS, a prediction model is firstly employed to separate out the unpredictable uncertainty from an input image. In addition, an improved local binary pattern is introduced to compute the structural uncertainty. Finally, combining luminance contrast with structural uncertainty, the pattern masking function is deduced. Experimental result demonstrates that the proposed pattern masking function outperforms the existing visual masking function. Furthermore, we extend the pattern masking function to just noticeable difference (JND) estimation and introduce a novel pixel domain JND model. Subjective viewing test confirms that the proposed JND model is more consistent with the HVS than the existing JND models.
Jinjian Wu, Weisi Lin, Guangming Shi, Xiaotian Wang 0001, Fu Li 0002
IEEE Trans. Image Process.1
2013 Reduced-Reference Image Quality Assessment With Visual Information Fidelity
abstract
Reduced-reference (RR) image quality assessment (IQA) aims to use less data about the reference image and achieve higher evaluation accuracy. Recent research on brain theory suggests that the human visual system (HVS) actively predicts the primary visual information and tries to avoid the residual uncertainty for image perception and understanding. Therefore, the perceptual quality relies to the information fidelities of the primary visual information and the residual uncertainty. In this paper, we propose a novel RR IQA index based on visual information fidelity. We advocate that distortions on the primary visual information mainly disturb image understanding, and distortions on the residual uncertainty mainly change the comfort of perception. We separately compute the quantities of the primary visual information and the residual uncertainty of an image. Then the fidelities of the two types of information are separately evaluated for quality assessment. Experimental results demonstrate that the proposed index uses few data (30 bits) and achieves high consistency with human perception.
Jinjian Wu, Weisi Lin, Guangming Shi, Anmin Liu
IEEE Trans. Multim.1
2013 Just Noticeable Difference Estimation for Images With Free-Energy Principle
abstract
In this paper, we introduce a novel just noticeable difference (JND) estimation model based on the unified brain theory, namely the free-energy principle. The existing pixel-based JND models mainly consider the orderly factors and always underestimate the JND threshold of the disorderly region. Recent research indicates that the human visual system (HVS) actively predicts the orderly information and avoids the residual disorderly uncertainty for image perception and understanding. Thus, we suggest that there exists disorderly concealment effect which results in high JND threshold of the disorderly region. Beginning with the Bayesian inference, we deduce an autoregressive model to imitate the active prediction of the HVS. Then, we estimate the disorderly concealment effect for the novel JND model. Experimental results confirm that the proposed JND model outperforms the relevant existing ones. Furthermore, we apply the proposed JND model in image compression, and around 15% of bit rate can be reduced without jeopardizing the perceptual quality.
Jinjian Wu, Guangming Shi, Weisi Lin, Anmin Liu, Fei Qi 0001
IEEE Trans. Multim.1
2012 Self-similarity based structural regularity for just noticeable difference estimation
Jinjian Wu, Fei Qi 0001, Guangming Shi
J. Vis. Commun. Image Represent.1
2012 Non-local spatial redundancy reduction for bottom-up saliency estimation
Jinjian Wu, Fei Qi 0001, Guangming Shi, Yongheng Lu
J. Vis. Commun. Image Represent.1
2010 An improved model of pixel adaptive just-noticeable difference estimation
abstract
A pixel-wise adaptive model for estimating the just-noticeable difference (JND) in spatial domain is proposed in this paper. As the human visual system (HVS) can be considered as a multichannel system, we assume that there exist two channels in the HVS, which deliver luminance adaption factor and texture masking factor, respectively. Both channels affect the JND threshold in a cooperative manner. The texture regions are with abundant redundancy and can tolerate much noise. The disorder degree and spatial masking of the texture are considered to estimate the texture masking effect, for deducing such JND threshold that coincides with the HVS. Finally, the luminance adaptation factor and texture masking factor are combined nonlinearly. Various experiments confirm the improved model has a better visual effect than models proposed before.
Jinjian Wu, Fei Qi 0001, Guangming Shi
ICASSP1
2009 Extracting regions of attention by imitating the human visual system
abstract
Detecting and segmenting out the regions of interest (ROIs) is one of the foundations in image processing and analysis. Because the final information sink of images is human, for segmenting out the ROIs effectively, we need to study human visual system (HVS) and imitate the behaviors when human viewing a scene. Researchers have found several factors which affect human attentions by studying eye movements when one views an image. In this paper, a method is proposed to detect the ROIs automatically based on HVS. In the proposed algorithm, the properties of pixels such as the contrast, location and edges are analyzed, and the pixels are enhanced according to the sensitivity of HVS. Then these factors are combined to a salient map, which classifies each pixel of the image in relation to its perceptual importance. Finally, the ROIs are segmented according to the salient map. This algorithm is easy to work, and can segment the objects from complex background efficiently.
Fei Qi 0001, Jinjian Wu, Guangming Shi
ICASSP2