VLDB 2026 Research / reviewers in the wild / expert
Lu Zhang 0037
dblp:82/10609-37
· DBLP profile ↗
56ranked-venue papers
2as first author
28since 2021 · last 2026
0000-0002-8859-5453ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 1 first-author · 26 since 2021Human-computer interaction and ubiquitous computing · 9 · 5 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LISA: A New Subjective Test Protocol and Tool for Local Image Quality Assessment
Ewen Démézet, Meriem Outtas, Séverine Baudry, Luce Morin, Lu Zhang 0037 |
QoMEX | 5 |
| 2026 | No-Reference Quality Assessment of 3D Models Represented in Neural Radiance Fields and 3D Gaussian SplattingsabstractThe continuous breakthroughs in 3D reconstruction and rendering technologies have enabled synthesized 3D models to achieve exceptional realism. In particular, Neural Radiance Fields (NeRF) and 3D Gaussian Splattings (3DGS) have gained significant attention due to their impressive ability to deliver high-quality 3D models. However, research on the quality assessment of NeRF/3DGS models remains an underexplored area, hindering the development of relevant generation, compression, and transmission algorithms. To fill this gap, in this paper, we propose the first no-reference quality assessment metric for NeRF/3DGS models rendered in Processed Video Sequences (PVS). Considering the uniqueness and diversity of distortions introduced in NeRF/3DGS models, the core idea behind the proposed metric is to extract universal quality-aware features that are generalizable across various distortions, rather than targeting a specific one. Specifically, inspired by the fact that spatial distortions typically alter the statistical distributions, we first measure the spatial fidelity of the rendered PVS by analyzing the spatial explicit statistics of textural variation and naturalness. Then, motivated by the ability of implicit energy composition changes to reflect temporal distortions, we propose to evaluate the temporal consistency of the rendered PVS through inter-frame discrepancy energy and multi-frame motion energy in the Singular Value Decomposition (SVD) domain. Finally, the spatial explicit statistics and temporal implicit energy are combined as perceptual features to evaluate the quality of NeRF/3DGS models via Support Vector Regression (SVR). Extensive experimental results on three representative databases demonstrate the superiority of the proposed metric in various aspects, such as predictive accuracy, performance stability, and cross-database generalizability. The source code will be publicly available at https://github.com/ZhengyuZhang96/PVS-3DMQA. Yuhang Zhang 0011, Tiantian Zeng, Shishun Tian, Lu Zhang 0037 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | MUVOD: A Novel Multi-View Video Object Segm entation Dataset and a Benchmark for 3D SegmentationabstractThe application of methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D GS) have steadily gained popularity in the field of 3D object segmentation in static scenes. These approaches demonstrate efficacy in a range of 3D scene understanding and editing tasks. Nevertheless, the 4D object segmentation of dynamic scenes remains an underexplored field due to the absence of a sufficiently extensive and accurately labelled multi-view video dataset. In this paper, we present MUVOD, a new multi-view video dataset for training and evaluating object segmentation in reconstructed real-world scenarios. The 17 selected scenes, describing various indoor or outdoor activities, are collected from different sources of datasets originating from various types of camera rigs. Each scene contains a minimum of 9 views and a maximum of 46 views. We provide 7830 RGB images (30 frames per video) with their corresponding segmentation mask in 4D motion, meaning that any object of interest in the scene could be tracked across temporal frames of a given view or across different views belonging to the same camera rig. This dataset, which contains 459 instances of 73 categories, is intended as a basic benchmark for the evaluation of multi-view video segmentation methods. We also present an evaluation metric and a baseline segmentation approach to encourage and evaluate progress in this evolving field. Additionally, we propose a new benchmark for 3D object segmentation task with a subset of annotated multi-view images selected from our MUVOD dataset. This subset contains 50 objects of different conditions in different scenarios, providing a more comprehensive analysis of state-of-the-art 3D object segmentation methods. Our proposed MUVOD dataset is available athttps://volumetric-repository.labs.b-com.com/#/muvod. Bangning Wei, Joshua Maraval, Meriem Outtas, Kidiyo Kpalma, Nicolas Ramin, Lu Zhang 0037 |
IEEE Trans. Multim. | 6 |
| 2025 | Semantic-Guided Residual Learning for the Quality Assessment of Enhanced ImagesabstractImage enhancement algorithms are essential for improving visual quality but often introduce new distortions, highlighting the need for reliable image quality assessment (IQA). However, existing IQA methods typically focus on semantic information or distortion-prone regions while ignoring their interactions, resulting in unsatisfactory performance. To address this issue, we propose to integrate semantic information with edge residual learning and design a semantic-guided residual learning IQA framework tailored for enhanced images across diverse scenarios. Specifically, the proposed framework utilizes a covariance-guided encoder to extract semantic information, which is then enhanced using a semantic refinement module. The refined semantic information is subsequently utilized to guide edge residual feature learning in the decoder. Extensive experiments on multiple tasks such as deraining, dehazing, and low-light enhancement demonstrate that our method outperforms state-of-the-art approaches. Shishun Tian, Zhiwei Lan, Ting Su 0004, Xia Li 0006, Lu Zhang 0037 |
ICME | 6 |
| 2025 | Objective quality assessment of medical images and videos: review and challengesabstractAbstract Quality assessment is a key element for the evaluation of hardware and software involved in image and video acquisition, processing, and visualization. In the medical field, user-based quality assessment is still considered more reliable than objective methods, which allow the implementation of automated and more efficient solutions. Regardless of increasing research on this topic in the last decade, defining quality standards for medical content remains a non-trivial task, as the focus should be on the diagnostic value assessed by expert viewers rather than the perceived quality from naïve viewers, and objective quality metrics should aim at estimating the first rather than the latter. In this paper, we present a survey of methodologies used for the objective quality assessment of medical images and videos, dividing them into visual quality-based and task-based approaches. Visual quality-based methods compute a quality index directly from visual attributes, while task-based methods, being increasingly explored, measure the impact of quality impairments on the performance of a specific task. A discussion on the limitations of state-of-the-art research on this topic is also provided, along with future challenges to be addressed. Rafael Rodrigues, Lucie Lévêque, Jesús Gutiérrez 0001, Houda Jebbari, Meriem Outtas, Lu Zhang 0037, Aladine Chetouani, Shaymaa Al-Juboori, Maria G. Martini, António M. G. Pinheiro |
Multim. Tools Appl. | 6 |
| 2025 | A New Benchmark Database and Objective Metric for Light Field Image Quality EvaluationabstractLight Field Image (LFI) records both angular and spatial information and provides immersive experiences for observers by rendering a scene from multiple perspectives. To cope with the resolution limitations of capture hardware, LFI angular reconstruction and spatial super-resolution are two widely-used methods, but they can also induce some special types of distortions, especially when two methods are adopted in combination. To this end, new challenges have been brought in assessing the quality of these distorted LFIs. In this paper, firstly, we conduct subjective experiments to evaluate the distorted LFI quality and present a novel perceptual quality assessment database with the associated subjective quality scores. Specifically, the proposed database focuses on the distortions introduced by deep learning-based LFI angular reconstruction and spatial super-resolution methods, individually and multiplely. Besides, in the case of multiple distortions, the adoption order of two distortions is taken into consideration. Further, our database presents three types of LFIs that suffer from distortions: real-world, dense synthesis, and sparse synthesis. As a result, the quality of distorted LFIs was subjectively assessed by 32 valid observers using the Pairwise Comparison (PC) protocol. Secondly, we develop a novel objective No-Reference (NR) metric for LFI quality evaluation, based on the features extracted from spatial gradients, angular-spatial statistics, and binocular disparity. Finally, a benchmark of the proposed metric and numerous state-of-the-art quality assessment metrics on the proposed database is presented. Experimental results demonstrate the superiority of the proposed metric over most existing metrics in various aspects. The proposed database and metric will be publicly available athttps://github.com/ZhengyuZhang96/IETR-LFI. Shishun Tian, Jinjia Zhou, Luce Morin, Lu Zhang 0037 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Prototypical Progressive Alignment and Reweighting for Generalizable Semantic SegmentationabstractGeneralizable semantic segmentation, aims to excel on unseen target domains, as a critical focus due to the widespread practical applications requiring high generalizability. Class-wise prototypes, which depict class-wise centroids, as a type of domain-invariant information are key to improving the model generalizability due to its stability and representativeness. However, this manner faces some challenges. First, the existing methods adopt a coarse prototypical alignment form, potentially compromising performance. Second, the naive prototype generally serves as the class centroid generated by an average operation from source data batches, risks source domain overfitting, and may be detrimentally impacted by unrelated source data. Third, from a broader perspective, rather than just from a prototypical alignment perspective, the existing methods treat all samples equally, which is against the conclusion that different source features have different adaptation difficulties. To tackle these issues, we propose a novel method for generalizable semantic segmentation called Prototypical Progressive Alignment and Reweighting (PPAR) depending on the strong generalized representation of the Contrastive Language-Image Pretraining (CLIP) model. In particular, we first define the Original Text Prototype (OTP) and Visual Text Prototype (VTP) generated by the CLIP model, laying the foundation for the subsequent effective alignment strategy. Then, we propose a prototypical progressive alignment strategy by an easy-to-difficult alignment form to reduce domain-variant information progressively instead of directly. Finally, we propose a prototypical reweighting learning strategy that estimates the importance of the source data and corrects its learning weight to alleviate the influence of unrelated source features, i.e. alleviate negative transfer. Moreover, we also offer a theoretical insight into our method and it shows that our method compiles well on the domain generalization theory. Extensive experiments on several popular datasets demonstrate that our PPAR method achieves superior performance, proving the effectiveness of our method. The source code will available at: https://github.com/Hectoor/PPAR Yuhang Zhang 0011, Muxin Liao, Shishun Tian, Wenbin Zou, Lu Zhang 0037, Chen Xu 0004 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | A Dual Rig Approach for Multi-View Video and Spatialized Audio Capture in Medical TrainingabstractWe present a multi-view camera and spatialized audio microphone capture system designed for computer vision applications in free navigation immersive experiences. We propose a dataset of two long and complex in-situ training situations in the medical field. The scenarios in the dataset feature precise gestures for the learner to reproduce during complex situations with multiple simultaneous visual and auditory cues important for training. 3D computer vision techniques are used to reconstruct a 4D scene model from a set of videos to render novel views from unseen viewpoints. However, the quality of the rendered objects is directly dependent on the density of coverage by reference views. To ensure maximum Quality of Experience, we propose a dual rig of cameras, a central rig that captures the details of the gesture zone of the training scenarios and a peripheral rig that captures the environment of the room and the interactions occurring around the gesture zone. The central rig provides dense coverage of the central content, facilitating high-quality reconstruction on novel views of the captured gestures. Recordings include audio interactions of multiple actors, captured by Ambisonic microphones spatially distributed around the scene. The captured scenes are real-world educational content for medical courses, so this dataset provides a rare opportunity to assess the Quality of Experience of volumetric video techniques on realistic content, and to compare their pedagogical capabilities with standard multi-view video content. Joshua Maraval, Bangning Wei, David Pesce, Yann Gayral, Meriem Outtas, Nicolas Ramin, Lu Zhang 0037 |
QoMEX | 7 |
| 2023 | The impact of the affinity on ASD people visual engagementabstractAutism spectrum disorders affect the way people perceive their environment and interact with it. Many autistic people have a passion for an object or a topic, such as a film, planes, or geography maps to name very few of them, which is called an affinity. This affinity is sometimes described as an obsession that prevents the ASD subjects to connect with the surrounding world, but it is also considered as a key to the autistic world and a way to make a connection. In this paper we investigate the specific role of affinity in the autistic person’s attention. We have conducted eye tracking experiments over 44 autistic subjects from 3 different institutions. We have shown them neutral images and images with their own affinity and recorded their gaze position. Results are not conclusive in all the 3 institutions, but in the 2 first ones we got significant differences between the 2 sets of images indicating a higher visual attention for the affinity. Julie Fournier, Elise Etchamendy, Myriam Cherel, Meriem Outtas, Lu Zhang 0037 |
CBMI | 5 |
| 2023 | Predicting personalized saliency map for people with autism spectrum disorderabstractPeople with autism spectrum disorder usually exhibit heterogeneous gaze patterns. Universal saliency prediction, which generates salient regions based on high fixations across all observers, is limited to analyzing the visual attention of autism spectrum disorders. To solve the problem, we propose a learning-based method named PSMANet to predict the personalized saliency map based on personal information. Collecting personal information and collecting large-scale datasets are challenging tasks for people with autism spectrum disorders, since they often suffer from deficits in social communication and interaction. The proposed approach introduces the image-similarity-measure based embedding to extract personal information and transfers the saliency distribution knowledge from universal saliency prediction to personalized saliency prediction. For evaluating our network, two popular metrics, Normalized Scanpath Salience (NSS) and Area Under Curve (AUC), are used. The experimental results show that it achieves good performance on the databases of people with autism spectrum disorder. Meriem Outtas, Julie Fournier, Elise Etchamendy, Myriam Cherel, Lu Zhang 0037 |
CBMI | 6 |
| 2023 | Temporal Down-sampling based Video Coding with Frame-Recurrent EnhancementabstractIn many digital systems, the transmission bandwidth, as well as storage capacity, are usually very limited. This introduces challenges for both video transmission and video storage. To seek lower bit rates and further obtain high-quality up-sampled videos, this paper proposes a temporal down-sampling based video coding system and a frame-recurrent enhancement based video upsampling strategy. The structure of our proposed method is shown in Fig. 1. Unlike the existing work [1], instead of downsampling all video frames, only the intermediate frames are downsampled and two frames remain with high quality on the video coding system. Then, these two high-quality frames are used to iteratively enhance the quality of the low-bitrate low-quality frames through a deep-learned enhancement network. Compared to the latest video coding standard Versatile Video Coding (VVC), our work can obtain a BD-rate reduction from $39.261 {\%} \sim 85. 455$ % in All-Intra and Low-Delay-P configurations on the downsampled frames. A temporal down-sampling based video coding framework (TDS) is proposed. It can be combined with all the existing coding standards including HEVC/H.265 and VVC/H.266. A method of super-resolution with frame recurrent image enhancement (SRFR) is applied to up-sampling the frames by the neighboring high resolution frame. The temporal information from high resolution frames can be fully used to improve the video quality through frame recurrent. Keren He, Chi Do-Kim Pham, Lu Zhang 0037, Jinjia Zhou |
DCC | 4 |
| 2023 | TRG-DQA: Texture Residual-Guided Dehazed Image Quality AssessmentabstractImage dehazing algorithms have emerged to solve the visual impairment caused by haze. It is important to establish dehazed image quality assessment (DQA) methods that can accurately evaluate the dehazed image quality and the performance of dehazing algorithms. However, classical image quality assessment (IQA) and most hand-crafted feature based DQA methods may not be able to adequately measure complex distortions of dehazed images. To address this issue, this paper proposes a Texture Residual-Guided Dehazed image Quality Assessment (TRG-DQA) method. Specifically, we first introduce a global and local feature extraction module employing a combination of the Transformer and convolutional neural networks (CNN) for extracting the comprehensive features. Considering that texture residual maps represent haze density and artifact distortion information, we propose a residual-guided module to guide the model for efficient learning. Additionally, to mitigate the information loss issue that occurs in deeper networks, a distortion-aware feature enhancement module is proposed. Extensive experiments on six DQA databases demonstrate the proposed TRG-DQA achieves superior performance among all the state-of-the-art methods. Tiantian Zeng, Lu Zhang 0037, Wenbin Zou, Xia Li 0006, Shishun Tian |
ICIP | 2 |
| 2023 | Blind Quality Assessment of Light Field Image Based on Spatio-Angular Textural VariationabstractLight Field Image Quality Assessment (LF-IQA) is vitally important to facilitate the development of immersive technologies. However, current state-of-the-art LF-IQA metrics still struggle to handle Light Field Image (LFI) with massive data in an efficient manner. To cope with this challenge, we propose a simple yet effective Blind LF-IQA metric based on Spatio-Angular Textural Variation, named SATV-BLiF. Given a distorted LFI, we first apply Local Binary Pattern (LBP) operator to measure the textural variation in the spatial and angular domains respectively. Then the generated spatial and angular textural matrices are merged and further transformed into statistical textural histogram features. Finally, Support Vector Regression (SVR) is employed to construct a nonlinear mapping function between the statistical textural histogram features and the perceptual quality score of the distorted LFI. Experimental results on three representative light field databases show that the proposed metric achieves state-of-the-art quality evaluation performance, while having much lower complexity than the existing No-Reference (NR) LF-IQA metrics. The code of the proposed SATV-BLiF metric is available at https://github.com/ZhengyuZhang96/SATV-BLiF. Shishun Tian, Wenbin Zou, Yuhang Zhang 0011, Luce Morin, Lu Zhang 0037 |
ICIP | 6 |
| 2023 | Towards Machine Perception Aware Image Quality AssessmentabstractOver the years, the objective of image and video compression has been to preserve perceived quality according to the Human Visual System (HVS) with minimal rate. Traditional encoders achieve this with the use of Rate-Distortion Optimization (RDO) techniques along with Image Quality Assessment (IQA) metrics that are correlated with human perception. Nowadays, a fast-growing number of applications fall within the realm of Video Coding for Machines (VCM), where the final recipient of compressed data is not a human but a machine performing a vision task. Recently, the lack of correlation between existing distortion measures and machine perception has been revealed, especially for RDO algorithms where distortion measures are computed on a local scale. In this paper, we propose a machine perception-aware metric designed to be incorporated into a standard-compliant Versatile Video Coding (VVC) encoder. Our proposed metric relies on a supervised training procedure as well as additional information available on the encoder side. In terms of correlation with machine perception, our metric significantly outperforms existing distortion measures in the literature. Alban Marie, Karol Desnos, Jinjia Zhou, Luce Morin, Lu Zhang 0037 |
MMSP | 6 |
| 2023 | Evaluation of Image Quality Assessment Metrics for Semantic Segmentation in a Machine-to-Machine Communication ScenarioabstractImage and video compression aims at finding an optimal trade-off between rate and distortion. This is done through Rate-Distortion Optimization (RDO) in traditional en-coders with the use of Image Quality Assessment (IQA) metrics. While it is known that most IQA metrics are designed to be correlated with human perception, there is no evidence that this observation can be generalized in a Video Coding for Machines (VCM) context, where the receiver is not a human anymore but a machine. In this paper, we propose an evaluation protocol to measure the correlation level between conventional Full-Reference (FR) IQA metrics and machine perception through the semantic segmentation vision task. Experiments showed a relatively low correlation between them when measured on the block-level. This observation implies the need of RDO algorithms that are better suited for Machine-to-Machine (M2M) communications. In order to facilitate the emergence of IQA metrics that better reflect machine perception, the code and dataset used to perform this study is made freely available at https://github.com/albmarie/iqa_m2m_segmentation. Alban Marie, Karol Desnos, Luce Morin, Lu Zhang 0037 |
QoMEX | 4 |
| 2023 | Denoised CT Images Quality Assessment Through COVID-19 Pneumonia Detection TaskabstractMedical images largely contribute to the diagnosis of lung diseases, especially pneumonia, an inflammation of lungs tissue. Since the emergence of COVID-19 in late 2019, medical imaging systems, notably computed tomography (CT) scans, have considerably helped in its diagnosis as well as revealing its infection severity. Serving as such an important role in clinical practice, the quality of medical images is therefore crucial for an accurate diagnosis. Denoising techniques, as a common image processing method, are being more and more used in medical imaging. However, how image denoising technique influences medical images' quality in terms of diagnostic performance still remains to be answered. In this paper, a primary study was carried out thanks to a detection task-based image quality assessment experiment, where we explored the performance of COVID-19 classifiers on both original and denoised chest CT scans. Two different denoising methods, i.e., anisotropic diffusion (AD) and total variation (TV) filters, were used. Results showed that the TV denoised model performed better than both baseline and AD denoised model, despite its less favorable mathematical image quality metrics. Lumi Xia, Houda Jebbari, Olivier Déforges, Lu Zhang 0037, Lucie Lévêque, Meriem Outtas |
QoMEX | 4 |
| 2023 | EDDMF: An Efficient Deep Discrepancy Measuring Framework for Full-Reference Light Field Image Quality AssessmentabstractThe increasing demand for immersive experience has greatly promoted the quality assessment research of Light Field Image (LFI). In this paper, we propose an efficient deep discrepancy measuring framework for full-reference light field image quality assessment. The main idea of the proposed framework is to efficiently evaluate the quality degradation of distorted LFIs by measuring the discrepancy between reference and distorted LFI patches. Firstly, a patch generation module is proposed to extract spatio-angular patches and sub-aperture patches from LFIs, which greatly reduces the computational cost. Then, we design a hierarchical discrepancy network based on convolutional neural networks to extract the hierarchical discrepancy features between reference and distorted spatio-angular patches. Besides, the local discrepancy features between reference and distorted sub-aperture patches are extracted as complementary features. After that, the angular-dominant hierarchical discrepancy features and the spatial-dominant local discrepancy features are combined to evaluate the patch quality. Finally, the quality of all patches is pooled to obtain the overall quality of distorted LFIs. To the best of our knowledge, the proposed framework is the first patch-based full-reference light field image quality assessment metric based on deep-learning technology. Experimental results on four representative LFI datasets show that our proposed framework achieves superior performance as well as lower computational complexity compared to other state-of-the-art metrics. Shishun Tian, Wenbin Zou, Luce Morin, Lu Zhang 0037 |
IEEE Trans. Image Process. | 5 |
| 2022 | Intra-Modal Constraint Loss for Image-Text RetrievalabstractCross-modal retrieval has drawn much attention in both computer vision and natural language processing domains. With the development of convolutional and recurrent neural networks, the bottleneck of retrieval across image-text modalities is no longer the extraction of image and text features but an efficient loss function learning in embedding space. Many loss functions try to closer pairwise features from heterogeneous modalities. This paper proposes a method for learning joint embedding of images and texts using an intra-modal constraint loss function to reduce the violation of negative pairs from the same homogeneous modality. Experimental results show that our approach outperforms state-of-the-art bi-directional image-text retrieval methods on Flickr30K and Microsoft COCO datasets. Our code is publicly available1 Jianan Chen 0002, Lu Zhang 0037, Cong Bai, Kidiyo Kpalma |
ICIP | 2 |
| 2022 | Deeblif: Deep Blind Light Field Image Quality Assessment by Extracting Angular and Spatial InformationabstractIn the era of immersive media, the high-dimensional Light Field Image (LFI) puts forward higher requirements for Light Field Image Quality Assessment (LF-IQA). However, currently most existing LF-IQA metrics still rely on sophisticated hand-crafted feature extraction, which fail to predict the quality of LFI accurately. In this paper, we propose a patch-based Deep Blind Light Field image quality assessment metric (abbreviated as DeeBLiF), by employing a two-stream Convolutional Neural Network (CNN) model specifically designed for extracting the angular and spatial information of LFI. Firstly, the spatio-angular patches are generated as input data, which effectively reflect the spatio-angular information of LFI. After that, a two-stream CNN model is exploited to extract the patch features and further predict the patch scores. Finally, all the patch scores are pooled into an overall quality score of LFI. Experimental results on the LFI dataset demonstrate that the proposed DeeBLiF outperforms the state-of-the-art LF-IQA metrics. The code will be publicly available at https://github.com/ZhengyuZhang96/DeeBLiF. Shishun Tian, Wenbin Zou, Luce Morin, Lu Zhang 0037 |
ICIP | 5 |
| 2022 | Video Coding for Machines: Large-Scale Evaluation of Deep Neural Networks Robustness to Compression Artifacts for Semantic SegmentationabstractIn the Video Coding for Machines (VCM) context where visual content is compressed before being transmitted to a vision task algorithm, appropriate trade-off between the compression level and the vision task performance must be chosen. In this paper, a Deep Neural Networks (DNN) based semantic segmentation algorithm robustness to compression artifacts is evaluated with a total of 1486 different coding configurations. Results indicate the importance of using an appropriate image resolution to overcome the block-partitioning limitations in existing compression algorithms, allowing 58.3%, 49.8%, 33.5% and 24.3% bitrate savings at equivalent prediction accuracy for JPEG, JM, x265 and VVenC, respectively. Surprisingly, JPEG can achieve 73.41% bitrate reduction with the inclusion of compressed images at training time over VVC Test Model (VTM) with a DNN trained on pristine data, which implies that DNN generalization ability must not be overlooked. Alban Marie, Karol Desnos, Luce Morin, Lu Zhang 0037 |
MMSP | 4 |
| 2021 | Learning Synergistic Attention for Light Field Salient Object Detection
Yi Zhang 0076, Geng Chen 0001, Yong Xia 0001, Olivier Déforges, Wassim Hamidouche, Lu Zhang 0037 |
BMVC | 8 |
| 2021 | A Multi-View Stereoscopic Video Database With Green Screen (MTF) For Video Transition Quality-of-Experience AssessmentabstractWe introduce a multi-view stereoscopic video database with a green screen, called MTF, for the usages in computer vision applications, in particular for free navigation, free-viewpoint television, and video transition quality-of-experience (QoE) assessment. The MTF contains full-HD videos of real storytelling made up of 3 scenes. One particularity of this dataset is that to understand its storytelling, users must change their point of view in the scene at a given time. To this end, we usually need to generate a transition to link two points of view in the same scene. Computer vision techniques that enable such transitions like view synthesis methods, rely on a set of images of the scene to render some new views from different viewpoints of this scene. However, these methods may have many failure cases that lead to artifacts in the final rendered video transition. In most view synthesis QoE tests, the contents are not designed to make the transition between two points of view useful or interesting for the viewers, e.g. they don't need to make a transition to capture more information to better understand the content. We thus, assume that participants will harshly judge artifacts and imperfections in the rendered transition. Thus, the MTF is expected to enable a better analysis of the visual impact of persistent artifacts in the final rendered transition. In our dataset, all the scenes are recorded in a green screen studio, which is often used to superimpose special effects and scenery during editing according to specific needs. Our dataset also presents a wide baseline camera-setup, a challenging constraint for view synthesis techniques. Finally, The MTF can also be used as a complementary dataset with others in literature in various computer vision applications, such as video compression, 3D video content, immersive virtual reality environment, optical flow estimation... Nour Hobloss, Lu Zhang 0037, Marco Cagnazzo |
QoMEX | 2 |
| 2021 | Quality assessment of DIBR-synthesized views: An overview
Shishun Tian, Lu Zhang 0037, Wenbin Zou, Xia Li 0006, Ting Su 0004, Luce Morin, Olivier Déforges |
Neurocomputing | 2 |
| 2021 | Hybrid dual stream blender for wide baseline view synthesis
Nour Hobloss, Lu Zhang 0037, Stéphane Lathuilière, Marco Cagnazzo, Attilio Fiandrotti |
Signal Process. Image Commun. | 2 |
| 2021 | A Lightweight Depth Estimation Network for Wide-Baseline Light FieldsabstractExisting traditional and ConvNet-based methods for light field depth estimation mainly work on the narrow-baseline scenario. This paper explores the feasibility and capability of ConvNets to estimate depth in another promising scenario: wide-baseline light fields. Due to the deficiency of training samples, a large-scale and diverse synthetic wide-baseline dataset with labelled data is introduced for depth prediction tasks. Considering the practical goal for real-world applications, we design an end-to-end trained lightweight convolutional network to infer depths from light fields, called LLF-Net. The proposed LLF-Net is built by incorporating a cost volume which allows variable angular light field inputs and an attention module that enables to recover details at occlusion areas. Evaluations are made on the synthetic and real-world wide-baseline light fields, and experimental results show that the proposed network achieves the best performance when compared to recent state-of-the-art methods. We also evaluate our LLF-Net on narrow-baseline datasets, and it consequently improves the performance of previous methods. Yan Li 0083, Lu Zhang 0037, Gauthier Lafruit |
IEEE Trans. Image Process. | 3 |
| 2021 | SC-RPN: A Strong Correlation Learning Framework for Region ProposalabstractCurrent state-of-the-art two-stage detectors heavily rely on region proposals to guide the accurate detection for objects. In previous region proposal approaches, the interaction between different functional modules is correlated weakly, which limits or decreases the performance of region proposal approaches. In this paper, we propose a novel two-stage strong correlation learning framework, abbreviated as SC-RPN, which aims to set up stronger relationship among different modules in the region proposal task. Firstly, we propose a Light-weight IoU-Mask branch to predict intersection-over-union (IoU) mask and refine region classification scores as well, it is used to prevent high-quality region proposals from being filtered. Furthermore, a sampling strategy named Size-Aware Dynamic Sampling (SADS) is proposed to ensure sampling consistency between different stages. In addition, point-based representation is exploited to generate region proposals with stronger fitting ability. Without bells and whistles, SC-RPN achieves AR100014.5% higher than that of Region Proposal Network (RPN), surpassing all the existing region proposal approaches. We also integrate SC-RPN into Fast R-CNN and Faster R-CNN to test its effectiveness on object detection task, the experimental results achieve a gain of 3.2% and 3.8% in terms of mAP compared to the original ones. Wenbin Zou, Yingqing Peng, Canqun Xiang, Shishun Tian, Lu Zhang 0037 |
IEEE Trans. Image Process. | 6 |
| 2021 | Unsupervised Adversarial Instance-Level Image RetrievalabstractWith the wide use of visual sensors in the Internet of Things (IoT) in the past decades, huge amounts of images are captured in people's daily lives, which poses challenges to traditional deep-learning-based image retrieval frameworks. Most such frameworks need a large amount of annotated training data, which are expensive. Moreover, machines still lack human intelligence, as illustrated by the fact that they pay less attention to the interesting regions that humans generally focus on when searching for images. Hence, this paper proposes a novel unsupervised framework that focuses on the instance object in the image and integrates human intelligence into the deep-learning-based image retrieval. This framework is called adversarial instance-level image retrieval (AILIR). We incorporate adversarial training and an attention mechanism into this framework that considers human intelligence with artificial intelligence. The generator and discriminator are redesigned to guarantee that the generator retrieves similar images while the discriminator selects unmatched images and creates an adversarial reward for the generator. A minimax game is conducted by the adversarial reward retrieval mechanism until the discriminator is unable to judge whether the image sequence retrieved matches the query. Comparison and ablation experiments on four benchmark datasets prove that the proposed adversarial training framework indeed improves instance retrieval and outperforms the state-of-the-art methods focused on instance retrieval. Cong Bai, Jinglin Zhang 0003, Ling Huang 0003, Lu Zhang 0037 |
IEEE Trans. Multim. | 5 |
| 2021 | A Multi-FoV Viewport-Based Visual Saliency Model Using Adaptive Weighting Losses for 360$^\circ$ Imagesabstract360$^\circ$media allows observers to explore the scene in all directions. The consequence is that the human visual attention is guided by not only the perceived area in the viewport but also the overall content in 360$^\circ$. In this paper, we propose a method to estimate the 360$^\circ$saliency map which extracts salient features from the entire 360$^\circ$image in each viewport in three different Field of Views (FoVs). Our model is first pretrained with a large-scale 2D image dataset to enable the interpretation of semantic contents, then fine-tuned with a relative small 360$^\circ$image dataset. A novel weighting loss function attached with stretch weighted maps is introduced to adaptively weight the losses of three evaluation metrics and attenuate the impact of stretched regions in equirectangular projection during training process. Experimental results demonstrate that our model achieves better performance with the integration of three FoVs and its diverse viewport images. Results also show that the adaptive weighting losses and stretch weighted maps effectively enhance the evaluation scores compared to the fixed weighting losses solutions. Comparing to other state of the art models, our method surpasses them on three different datasets and ranks the top using 5 performance evaluation metrics on the Salient360! benchmark set. The code is available athttps://github.com/FannyChao/MV-SalGAN360. Fang-Yi Chao, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges |
IEEE Trans. Multim. | 2 |
| 2020 | Binary Probability Model for Learning Based Image CompressionabstractIn this paper, we propose to enhance learned image compression systems with a richer probability model for the latent variables. Previous works model the latents with a Gaussian or a Laplace distribution. Inspired by binary arithmetic coding, we propose to signal the latents with three binary values and one integer, with different probability models.A relaxation method is designed to perform gradient-based training. The richer probability model results in a better entropy coding leading to lower rate. Experiments under the Challenge on Learned Image Compression (CLIC) test conditions demonstrate that this method achieves 18 % rate saving compared to Gaussian or Laplace models. Théo Ladune, Pierrick Philippe, Wassim Hamidouche, Lu Zhang 0037, Olivier Déforges |
ICASSP | 4 |
| 2020 | Manet: Multi-Scale Aggregated Network For Light Field Depth EstimationabstractWe present a novel end-to-end network, MANet, for light field depth estimation. MANet is a parameter-effective and effi-cient multi-scale aggregated network, which is about 3 times smaller and 3 times faster than the current top-performing method Epinet. The MANet architecture is performed for estimating depth from light field plenoptic cameras, and experimental results show that the proposed MANet outperforms state-of-the-art methods on HCI, CVIA-HCI and EPFL Lytro light field datasets. Yan Li 0083, Lu Zhang 0037, Gauthier Lafruit |
ICASSP | 2 |
| 2020 | A Fixation-Based 360° Benchmark Dataset For Salient Object DetectionabstractFixation prediction (FP) in panoramic contents has been widely investigated along with the booming trend of virtual reality (VR) applications. However, another issue within the field of visual saliency, salient object detection (SOD), has been seldom explored in 360° or omnidirectional) images due to the lack of datasets representative of real scenes with pixel-level annotations. Toward this end, we collect 107 equirectangular panoramas with challenging scenes and multiple object classes. Based on the consistency between FP and explicit saliency judgements, we further manually annotate 1,165 salient objects over the collected images with precise masks under the guidance of real human eye fixation maps. Six state-of-the-art SOD models are then benchmarked on the proposed fixation-based 360° image dataset (F-360iSOD), by applying a multiple cubic projection-based fine-tuning method. Experimental results show a limitation of the current methods when used for SOD in panoramic images, which indicates the proposed dataset is challenging. Key issues for 360° SOD is also discussed. The proposed dataset is available at https://github.com/PanoAsh/F-360iSOD. Yi Zhang 0076, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges |
ICIP | 2 |
| 2020 | Optical Flow and Mode Selection for Learning-based Video CodingabstractThis paper introduces a new method for inter-frame coding based on two complementary autoencoders: MOFNet and CodecNet. MOFNet aims at computing and conveying the Optical Flow and a pixel-wise coding Mode selection. The optical flow is used to perform a prediction of the frame to code. The coding mode selection enables competition between direct copy of the prediction or transmission through CodecNet.The proposed coding scheme is assessed under the Challenge on Learned Image Compression 2020 (CLIC20) P-frame coding conditions, where it is shown to perform on par with the state-of-the-art video codec ITU/MPEG HEVC. Moreover, the possibility of copying the prediction enables to learn the optical flow in an end-to-end fashion i.e. without relying on pre-training and/or a dedicated loss term. Théo Ladune, Pierrick Philippe, Wassim Hamidouche, Lu Zhang 0037, Olivier Déforges |
MMSP | 4 |
| 2020 | Towards Audio-Visual Saliency Prediction for Omnidirectional Video with Spatial AudioabstractOmnidirectional videos (ODVs) with spatial audio enable viewers to perceive 360° directions of audio and visual signals during the consumption of ODVs with head-mounted displays (HMDs). By predicting salient audio-visual regions, ODV systems can be optimized to provide an immersive sensation of audio-visual stimuli with high-quality. Despite the intense recent effort for ODV saliency prediction, the current literature still does not consider the impact of auditory information in ODVs. In this work, we propose an audio-visual saliency (AVS360) model that incorporates 360° spatial-temporal visual representation and spatial auditory information in ODVs. The proposed AVS360 model is composed of two 3D residual networks (ResNets) to encode visual and audio cues. The first one is embedded with a spherical representation technique to extract 360° visual features, and the second one extracts the features of audio using the log mel-spectrogram. We emphasize sound source locations by integrating audio energy map (AEM) generated from spatial audio description (i.e., ambisonics) and equator viewing behavior with equator center bias (ECB). The audio and visual features are combined and fused with AEM and ECB via attention mechanism. Our experimental results show that the AVS360 model has significant superiority over five state-of-the-art saliency models. To the best of our knowledge, it is the first w ork that develops the audio-visual saliency model in ODVs. The code will be publicly available to foster future research on audio-visual saliency in ODVs. Fang-Yi Chao, Cagri Ozcinar, Lu Zhang 0037, Wassim Hamidouche, Olivier Déforges, Aljoscha Smolic |
VCIP | 3 |
| 2020 | Discriminative feature representation for Noisy image quality assessment
Yunbo Gu, Tianling Lv, Yang Chen 0008, Lu Zhang 0037, Jian Yang 0009, Huazhong Shu, Limin Luo 0001, Gouenou Coatrieux |
Multim. Tools Appl. | 6 |
| 2020 | Overview of deep-learning based methods for salient object detection in videos
Lu Zhang 0037, Yan Li 0083, Kidiyo Kpalma |
Pattern Recognit. | 2 |
| 2020 | Salient video object detection using a virtual border and guided filter
Lu Zhang 0037, Wenbin Zou, Kidiyo Kpalma |
Pattern Recognit. | 2 |
| 2019 | How Well Current Saliency Prediction Models Perform on UAVs Videos?
Anne-Flore Perrin, Lu Zhang 0037, Olivier Le Meur |
CAIP (1) | 2 |
| 2019 | Laboratory and Crowdsourcing Studies of Lip Sync Effect on the Audio-Video Quality Assessment for Videoconferencing ApplicationabstractLip sync is one of the factors that impacts a lot the quality of the videoconferencing experience. In this paper we study the end-user perception of the asynchrony and we try to determine the annoyance threshold of lip-synch error. We are also interested in investigating the mutual interaction between the asynchrony annoyance and changes in the video coding bit rate, spatial video resolution, video IP packet loss, and audio IP packet loss. We conducted two subjective tests in two different environments: laboratory and crowdsourcing. The experimental results show that the audio-video desynchronization annoyance is not an independent factor, but influenced by the video and audio quality. Furthermore, by comparing the results of the two tested methodologies, we show that the crowdsourcing methodology for the quality assessment might be used for audiovisual, video and asynchrony perception assessment, but there is some challenges to consider for the audio quality test. Inès Saidi, Lu Zhang 0037, Vincent Barriac, Olivier Déforges |
ICIP | 2 |
| 2019 | An effective general-purpose NR-IQA model using natural scene statistics (NSS) of the luminance relative order
Tonghan Wang 0002, Lu Zhang 0037, Huizhen Jia |
Signal Process. Image Commun. | 2 |
| 2019 | A Benchmark of DIBR Synthesized View Quality Assessment Metrics on a New Database for Immersive Media ApplicationsabstractDepth-image-based rendering (DIBR) is a fundamental technology in several 3-D-related applications, such as free viewpoint video, virtual reality, and augmented reality. However, new challenges have also been brought in assessing the quality of DIBR-synthesized views since this process induces some new types of distortions, which are inherently different from the distortion caused by video coding. In this paper, we present a new DIBR-synthesized image database with the associated subjective scores. We also test the performances of the state-of-the-art objective quality metrics on this database. This paper focuses on the distortions only induced by different DIBR synthesis methods. Seven state-of-the-art DIBR algorithms, including inter-view synthesis and single-view-based synthesis methods, are considered in this database. The quality of synthesized views was assessed subjectively by 41 observers and objectively using 14 state-of-the-art objective metrics. Subjective test results show that the interview synthesis methods, having more input information, significantly outperform the single-view-based ones. Correlation results between the tested objective metrics and the subjective scores on this database reveal that further studies are still needed for a better objective quality metric dedicated to the DIBR-synthesized views. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
IEEE Trans. Multim. | 2 |
| 2018 | Low Complexity Joint RDO of Prediction Units Couples for HEVC Intra CodingabstractHEVC is the latest block-based video compression standard, outperforming H.264/AVC by 50% bitrate savings for the same perceptual quality. An HEVC encoder provides Rate-Distortion optimization coding tools for block-wise compression. Because of complexity limitations, Rate-Distortion Optimization (RDO) is usually performed independently for each block, assuming coding efficiency losses to be negligible. In this paper, we propose an acceleration solution for the Intra coding scheme named Dual-JRDO, which takes advantage of Inter-Block dependencies related to both predictive coding and CABAC. The Dual-JRDO improves Intra coding efficiency at the expense of higher computational complexity. The acceleration of the Dual-JRDO scheme includes adaptive use of the Dual-JRDO model based on source analysis, short-listing and early decisions strategies. The proposed Fast Dual-JRDO reduces the original model complexity by 89.54%, while providing tractable computation for average R-D gains of -0.45% (up to -0.82%) in the HM16.12 reference software model. Maxime Bichon, Julien Le Tanou, Michaël Ropert, Wassim Hamidouche, Luce Morin, Lu Zhang 0037 |
ICASSP | 6 |
| 2018 | Temporal Adaptive Quantization using Accurate Estimations of Inter and Skip ProbabilitiesabstractHybrid video coding systems use spatial and temporal predictions in order to remove redundancies within the video source signal. These predictions create coding-scheme-related dependencies, often neglected for sake of simplicity. The R-D Spatio-Temporal Adaptive Quantization (RDSTQ) solution uses such dependencies to achieve better coding efficiency. It models the temporal distortion propagation by estimating the probability of a Coding Unit (CU) to be Inter coded. Uased on this probability, each CU is given a weight depending on its relative importance compared to other CUs. However, the initial approach roughly estimates the Inter probability and does not take into account the Skip mode characteristics in the propagation. It induces important Target uitrate Deviation (TBD) compared to the reference target rate. This paper provides undeniable improvements of the original RDSTQ model in using a more accurate estimation of the Inter probability. Then a new analytical solution for local quantizers is obtained by introducing the Skip probability of a CU into the temporal distortion propagation model. The proposed solution brings -2.05% BD-BR gain in average over the RDSTQ at low rate, which corresponds to -13.54% BD-BR gain in average against no local quantization. Moreover, the TBD is reduced from 38% to 14%. Maxime Bichon, Julien Le Tanou, Michaël Ropert, Wassim Hamidouche, Luce Morin, Lu Zhang 0037 |
PCS | 6 |
| 2018 | On the Subjective Assessment of the Perceived Quality of Medical Images and VideosabstractMedical professionals are viewing an increasing number of images and videos in their clinical routine. However, various types of distortions can affect medical imaging data, and therefore impact the viewers' experienced quality and their clinical practice. Thus it is necessary to quantify this impact and understand how the viewers, i.e., medical experts, perceive the quality of (distorted) images and videos. In this paper, we present an up-to-date review of the methodologies used in the literature for the subjective quality assessment of medical images and videos and discuss their merits and drawbacks depending on the use case. Lucie Lévêque, Hantao Liu, Sabina Barakovic, Jasmina Barakovic, Maria G. Martini, Meriem Outtas, Lu Zhang 0037, Asli Kumcu, Ljiljana Platisa, Rafael Rodrigues, António M. G. Pinheiro, Athanassios N. Skodras |
QoMEX | 7 |
| 2018 | Evaluation of No-reference quality metrics for Ultrasound liver imagesabstractAlthough assessing post-processed medical images is still done by radiologists (rather than computers), numerous algorithms dedicated to medical image processing are developed without taking into consideration the expert's perceived quality scores. In order to evaluate these algorithms, we study in this paper four No-Reference(NR) quality assessment metrics in terms of correlation with perceived scores of experts. These scores were obtained through subjective tests conducted on ultrasound (US) livers images. Results show that one NR metric among the four evaluated performs the best for assessing the quality of US images. However, further study is needed for the development of more suitable NR metrics. Meriem Outtas, Lu Zhang 0037, Olivier Déforges, Wassim Hamidouche, Amina Serir |
QoMEX | 2 |
| 2018 | SC-IQA: Shift compensation based image quality assessment for DIBR-synthesized viewsabstractDepth-image-based-rendering (DIBR) has been used to generate the virtual views for Multi-view videos and Free-viewpoint videos. However, the quality assessment of DIBR-synthesized views is very challenging owing to the new types of distortions induced by inaccurate depth maps, dis-occlusions and image inpainting methods. There exist a large number of object shifts and geometric distortions in the synthesized view which the traditional 2D quality metrics may fail to assess. In this paper, we propose a shift compensation based image quality assessment metric (SC-IQA) for DIBR-synthesized views. Firstly, the global geometric shift is compensated roughly by an SURF + RANSAC homography approach. Then, a multi-resolution block matching method, which performs a more accurate matching, is used to precisely compensate the shift and penalize the local geometric distortion as well. In addition, a visual saliency map is also used as a weighting function. To calculate the final overall quality scores, only the worst blocks are utilized since the biggest distortions have the most effects on the overall perceptual quality. The results show that the proposed metric significantly outperforms the state-of-the-art synthesized view dedicated metrics and the conventional 2D IQA metrics. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
VCIP | 2 |
| 2018 | MS-CapsNet: A Novel Multi-Scale Capsule NetworkabstractCapsule network is a novel architecture to encode the properties and spatial relationships of the feature in an image, which shows encouraging results on image classification. However, the original capsule network is not suitable for some classification tasks, where the target objects are complex internal representations. Hence, we propose a multi-scale capsule network that is more robust and efficient for feature representation in image classification. The proposed multi-scale capsule network consists of two stages. In the first stage, structural and semantic information are obtained by multi-scale feature extraction. In the second stage, the hierarchy of features is encoded to multi-dimensional primary capsules. Moreover, we propose an improved dropout to enhance the robustness of the capsule network. Experimental results show that our method has a competitive performance on FashionMNIST and CIFAR10 datasets. Canqun Xiang, Lu Zhang 0037, Yi Tang 0008, Wenbin Zou, Chen Xu 0004 |
IEEE Signal Process. Lett. | 2 |
| 2018 | NIQSV+: A No-Reference Synthesized View Quality Assessment MetricabstractBenefiting from multi-view video plus depth and depth-image-based-rendering technologies, only limited views of a real 3-D scene need to be captured, compressed, and transmitted. However, the quality assessment of synthesized views is very challenging, since some new types of distortions, which are inherently different from the texture coding errors, are inevitably produced by view synthesis and depth map compression, and the corresponding original views (reference views) are usually not available. Thus the full-reference quality metrics cannot be used for synthesized views. In this paper, we propose a novel no-reference image quality assessment method for 3-D synthesized views (called NIQSV+). This blind metric can evaluate the quality of synthesized views by measuring the typical synthesis distortions: blurry regions, black holes, and stretching, with access to neither the reference image nor the depth map. To evaluate the performance of the proposed method, we compare it with four full-reference 3-D (synthesized view dedicated) metrics, five full-reference 2-D metrics, and three no-reference 2-D metrics. In terms of their correlations with subjective scores, our experimental results show that the proposed no-reference metric approaches the best of the state-of-the-art full reference and no-reference 3-D metrics; and outperforms the widely used no-reference and full-reference 2-D metrics significantly. In terms of its approximation of human ranking, the proposed metric achieves the best performance in the experimental test. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
IEEE Trans. Image Process. | 2 |
| 2017 | Inter-block dependencies consideration for intra coding in H.264/AVC and HEVC standardsabstractRecent MPEG video compression standards are still block-based: blocks of pixels are sequentially coded using spatial or temporal prediction schemes. For each block, a vector of coding parameters has to be selected. In order to limit the complexity of this decision, independence between blocks is assumed, and coding parameters are locally optimized to maximize the coding efficiency. Few studies have investigated the benefits of inter-block dependencies consideration using Joint Rate-Distortion Optimization (JRDO), especially in Intra coding. To the best of our knowledge, maximum achievable gains of such approaches have never been exhibited. In this paper, we propose two JRDO models performing joint optimization of multiple blocks applied to intra prediction mode decision. The proposed models have been evaluated in both H.264/AVC and HEVC standards. These two models enables a bitrate saving with respect to the classical RDO model up to -3.10% and -2.31% in H.264/AVC and HEVC, respectively. Maxime Bichon, Julien Le Tanou, Michaël Ropert, Wassim Hamidouche, Luce Morin, Lu Zhang 0037 |
ICASSP | 6 |
| 2017 | NIQSV: A no reference image quality assessment metric for 3D synthesized viewsabstractThe popularity of 3D applications, such as Free View-point TV (FTV) and Multi-view Video plus Depth (MVD), induces a heavy requirement of synthesized views. However, the quality assessment of synthesized views is very challenging because the corresponding original views (reference views) are usually not available at both encoder and decoder sides. In this paper, we propose a new no-reference quality assessment model to evaluate the quality of 3D synthesized views, called NIQSV (No-reference Image Quality assessment of Synthesized Views). This metric is based on the hypothesis that a good quality image is composed of flat areas (objects) separated by sharp edges, and the quality estimation involves only a set of simple morphological operators. NIQSV integrates the distortions of all the components, and then uses an edge image to weight the final distortions since the distortions of synthesized views mainly happen around object edges. The experimental results show that the proposed metric outperforms traditional 2D metrics and ranks among the best of dedicated 3D synthesized and full reference metrics. Shishun Tian, Lu Zhang 0037, Luce Morin, Olivier Déforges |
ICASSP | 2 |
| 2017 | Multi-output speckle reduction filter for ultrasound medical images based on multiplicative multiresolution decompositionabstractUltrasonographic examination, either as visual inspection or quantitative analysis, is less effective than other medical imaging systems due to speckle noise. The state-of-the-art speckle reduction methods often offers an effective speckle reduction but generally they suffer from oversmoothig, blurring effect and man-made/artificial appearance. In this paper, a new Multi-Output Filter based on a Multiplicative Multiresolution Decomposition (MOF-MMD) is proposed. This multiscale based method, particularly efficient in the case of multiplicative noise, enhances distinctively three outputs: edges, texture and the global image. The multi-output filter aims at offering an enhanced images according to the features desired by radiologists. The different structures, textures and edges are filtered according to the contour image obtained by morphological operators. Finally, we compare the MOF-MMD method with two state-of-the-art speckle reduction methods in terms of speckle reduction capacity and image quality improvement. The results show that the proposed method offers an effective speckle reduction with an improvement of the image quality without blurry and over-smoothing effect. Meriem Outtas, Lu Zhang 0037, Olivier Déforges, Amina Serir, Wassim Hamidouche |
ICIP | 2 |
| 2017 | Evaluation of single-artifact based video quality metrics in video communication contextabstractFor an accurate assessment of media quality, it is essential not only to compute an overall quality measure, but also to identify the type of occurring distortions. In this paper, we focus on a set of no-reference single-artifact based metrics developed by the MOAVI project. We carried out a correlation analysis in order to evaluate the performance of these metrics on three databases with a large sample of distortion types. This study will be used for setting up a video quality monitoring tool box. Inès Saidi, Lu Zhang 0037, Vincent Barriac, Olivier Déforges |
QoMEX | 2 |
| 2016 | Audiovisual quality study for videoconferencing on IP networksabstractIn this paper, an audiovisual quality assessment experiment was conducted on audiovisual clips collected using a PC-based videoconferencing application connected via a local IP network. The analyses of experimental results provided a better understanding of the influence of network impairments (packet loss, jitter, delay) on perceived audio and video qualities, as well as their interaction effect on the overall audiovisual quality in videoconferencing applications. We updated the human perception acceptability limits of audio-video synchronization for video conferencing. Further, we investigated the contribution of this synchronization to the audiovisual quality independently and accompanied with network impairments. Finally, we proposed an integration model to estimate the audiovisual quality in the studied context. Inès Saidi, Lu Zhang 0037, Vincent Barriac, Olivier Déforges |
MMSP | 2 |
| 2016 | Interactive vs. non-interactive subjective evaluation of IP network impairments on audiovisual quality in videoconferencing contextabstractIn this paper, we present a subjective audiovisual quality assessment experiment realized using a PC based video-conferencing application connected via a local IP network. The experiment was conducted under two different scenarios: a non-interactive and an interactive conversational one. We present the effects of network impairments (packet loss, delay) on perceived audiovisual, audio and video quality. We evaluate the impact of scene complexity on the quality perception in case of video calls. We establish a comparison between the perception of multimedia quality in interactive and non-interactive context. The results presented in this paper show a dependency between the perceived quality and the scene complexity: the perceived quality of a high spatial and temporal complex scene is inferior to that of the less complex one. In addition, our findings show that the audio-video synchronization acceptability do not differ between the interactive and the non-interactive experimental context. Inès Saidi, Lu Zhang 0037, Vincent Barriac, Olivier Déforges |
QoMEX | 2 |
| 2016 | Multiscale contrast similarity deviation: An effective and efficient index for perceptual image quality assessment
Tonghan Wang 0002, Lu Zhang 0037, Huizhen Jia, Huazhong Shu |
Signal Process. Image Commun. | 2 |
| 2015 | A multi-slice model observer for medical image quality assessmentabstractModel observers (MOs) have been developed for the medical image quality assessment. Nowadays, numerous modern medical instruments are capable of producing 3D images, while few researchers have conducted MO studies on 3D data. In this paper, we propose a multi-slice MO when considering a relatively more realistic diagnostic task: the detection-localization of simulated multiple-sclerosis (MS) lesions on 3D magnetic resonance (MR) images. The jackknife free-response receiver operating characteristic (JAFROC) method was used to quantitatively analyse its performances and compare them with those of human observers. Our preliminary results showed that the proposed framework has the potential to approach human detection-localization task performance. Lu Zhang 0037, Christine Cavaro-Ménard, Patrick Le Callet, Di Ge |
ICASSP | 1 |
| 2012 | A Perceptually Relevant Channelized Joint Observer (PCJO) for the Detection-Localization of Parametric SignalsabstractMany numerical observers have been proposed in the framework of task-based approach for medical image quality assessment. However, the existing numerical observers are still limited in diagnostic tasks: the detection task has been largely studied, while the localization task concerning one signal has been little studied and the localization of multiple signals has not been studied yet. In addition, most existing numerical observers need a priori knowledge about all the parameters of the underdetection signals, while only a few of them need at least two signal parameters. In this paper, we propose a novel numerical observer called the perceptually relevant channelized joint observer (PCJO), which cannot only detect but also localize multiple signals with unknown amplitude, orientation, size and location. We validated the PCJO for predicting human observer task performance by conducting a clinically relevant free-response subjective experiment in which six radiologists (including two experts) had to detect and localize multiple sclerosis (MS) lesions on magnetic resonance (MR) images. By using the jackknife alternative free-response operating characteristic (JAFROC) as the figure of merit (FOM), the detection-localization task performance of the PCJO was evaluated and then compared to that of the radiologists and two other numerical observers--channelized hotelling observer (CHO) and Goossenss CHO for detecting asymmetrical signals with random orientations. Overall, the results show that the PCJO performance was closer to that of the experts than to that of the other radiologists. The JAFROC1 FOMs of the PCJO (around 0.75) are not significantly different from those of the two experts (0.7672 and 0.7110), while the JAFROC1 FOMs of the numerical observers mentioned above (always over 0.84) outperform those of the experts. This indicates that the PCJO is a promising method for predicting radiologists' performance in the joint detection-localization task. Lu Zhang 0037, Christine Cavaro-Ménard, Patrick Le Callet, Jean-Yves Tanguy |
IEEE Trans. Medical Imaging | 1 |