Fernando Pereira 0001

dblp:49/3721-1 · also Fernando Manuel Bernardo Pereira · DBLP profile ↗
← Back
180ranked-venue papers
24as first author
26since 2021 · last 2026
0000-0001-6100-947XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 174 · 22 first-author · 25 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 JPEG AI Compressed Domain Face Detection: A Multi-Scale Bridging Perspective
abstract
Learning-based image coding is showing improved compression efficiency, while also offering a novel advantage in enabling computer vision tasks directly within the compressed domain. The latent representation created by deep learning methods inherently contains all visual features, without a computationally expensive synthesis process at the decoder. This paper is an invited extension of a previous solution for JPEG AI compressed domain face detection that adapts a RetinaFace-based detector to operate directly on the latent tensor. In addition to a former single-scale bridging solution, this work provides a novel multi-scale bridging architecture to enable a more effective multi-scale compressed domain face detection. The results show a significant performance gain, improving accuracy up to 20% for detection of tiny faces on the WIDER FACE dataset compared to single-scale bridging, and further narrowing the gap when compared to detection on uncompressed or JPEG AI decoded images. Furthermore, since the computationally expensive decoding step is bypassed and since the bridges consist of lower-complexity networks, the overall processing cost is significantly reduced. Single and multi-scale bridging, respectively, have about 10% and 32% the complexity of applying pixel domain face detection on decoded images. The proposed architecture is expected to be extended to other multiscale sensitive vision tasks, as JPEG AI is not specifically designed for any single downstream application.
Ayman Alkhateeb, Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, João Ascenso, Fernando Pereira 0001
IEEE Trans. Multim.6
2026 Scalable Graph-Guided Transformer for Point Cloud Geometry Coding
abstract
Attention models, particularly Transformers, have significantly advanced deep learning in fields like natural language processing and computer vision by capturing contextual relationships in both sequential and spatial data. This ability is valuable for Point Clouds (PC), which are unstructured sets of points in 3D space. Transformers can effectively identify correlations between distant points, allowing them to focus on the most critical regions of the data. To demonstrate this capability, this paper proposes a novel, scalable Graph-Guided Transformer model, labeled 2GFormer, for static PC geometry. This model is built using a scalable architecture that leverages Graph Convolutions to enhance a Relational Neighborhood SelfAttention (RNSA) base layer model. Both models are integrated into the JPEG Pleno Learning-based Point Cloud Coding (JPEG PCC) standard, resulting in the creation of two attention-enabled codecs for static PC coding: JPEG RNSA and JPEG 2GFormer. While JPEG RNSA codec delivers significant compression improvements for solid and dense PCs compared to the baseline JPEG PCC standard, JPEG 2GFormer extends these gains to solid, dense, and sparse PCs with only a marginal increase in model parameters. Additionally, JPEG 2GFormer outperforms both conventional and learning-based state-of-the-art PC codecs. These results position JPEG 2GFormer as a highly efficient solution for versatile PC coding.
Mohammadreza Ghafari, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001
IEEE Trans. Multim.4
2025 Deep Learning-Based Point Cloud Coding and Super-Resolution: A Joint Geometry and Color Approach
abstract
In this golden age of multimedia, realistic content is in high demand with users seeking more immersive and interactive experiences. As a result, new image modalities for 3D representations have emerged in recent years, among which point clouds have deserved especial attention. Naturally, with this increase in demand, efficient storage and transmission became a must, with standardization groups such as MPEG and JPEG entering the scene, as it happened before with other types of visual media. In a surprising development, JPEG issued a Call for Proposals on point cloud coding targeting exclusively learning-based solutions, in parallel to a similar call for image coding. This is a natural consequence of the growing popularity of deep learning, which due to its excellent performances is currently dominant in the multimedia processing field, including coding. This paper presents the coding solution selected by JPEG as the best-performing response to the Call for Proposals and adopted as the first version of the JPEG Pleno Point Cloud Coding Verification Model, in practice the first step for developing a standard. The proposed solution offers a novel joint geometry and color approach for point cloud coding, in which a single deep learning model processes both geometry and color simultaneously. To maximize the RD performance for a large range of point clouds, the proposed solution uses down-sampling and learning-based super-resolution as pre- and post-processing steps. Compared to the MPEG point cloud coding standards, the proposed coding solution comfortably outperforms G-PCC, for both geometry, color, and joint quality metrics.
André F. R. Guarda, Manuel Ruivo, Luís Coelho, Abdelrahman Seleem, Nuno M. M. Rodrigues, Fernando Pereira 0001
IEEE Trans. Multim.6
2024 XAIface: A Framework and Toolkit for Explainable Face Recognition
abstract
Artificial intelligence-based face recognition solutions are becoming increasingly popular. Therefore, it is crucial to fully understand and explain how these technologies work in order to make them more effective and acceptable to society. This is the goal of the CHIST-ERA project XAIface, the final results of which are reported in this article: a framework and toolkit for improving AI decision explainability, in the context of automated face recognition, through several novel methods are presented. These methods are integrated into an end-to-end face recognition demonstrator system, which facilitates studying the impact of various influencing factors and system processes on recognition performance. By doing so, we can visually explain the decisions made by the face verification pipeline for specific instances in our test set using heatmaps and locally interpretable features. Furthermore, we offer a comprehensive explanation of the end-to-end model by examining the relationship between verification failures and misclassifications of soft biometric facial traits.
Nélida Mirabet-Herranz, Naima Bousnina, Jonas Pfister, Chiara Galdi, Jean-Luc Dugelay, Werner Bailer, Touradj Ebrahimi, Paulo Lobato Correia, Fernando Pereira 0001, Felix Schmautzer, Erich Schweighofer
CBMI11
2024 An International Standard For Assessing Trustworthiness In Media
abstract
The proliferation of synthetic media generation technologies, such as generative AI, has led to a surge of media content generation and consumption. While this progress opens new opportunities, especially in creative industries, it also causes challenges, including piracy, fake media distribution, and concerns about trust and privacy. In the creative sector, media modifications are often part of the production pipelines and in many application domains, creators need or want to declare the type of modifications that were performed on the media asset. The cryptographically signed association of provenance information with the media asset itself provides a trust link between the owner or editor of a media asset and its consumers. The absence of such assertions may reveal the lack of trustworthiness in media assets or worse, the intention to hide the existence of manipulations. This paper describes the JPEG Trust framework (ISO/IEC 21617) that aims to establish trust in digital media creation, modification, annotation, distribution and consumption. The framework provides standardized protocols to extract indicators to assess trustworthiness, means to annotate media provenance, and securely link the assets and associated annotations together.
Deepayan Bhowmik, Sabrina B. Caldwell, Jaime Delgado, Touradj Ebrahimi, Nikolaos Fotos, Xiaojun Gu, Ziyuan Hu, Xin Kang 0001, Fernando Pereira 0001, Leonard Rosenthol, Frederik Temmermans
ICIP9
2024 Learning-Based Point Cloud Decoding with Independent and Scalable Reduced Complexity
abstract
Point Clouds (PCs) have gained significant attention due to their usage in diverse application domains, notably virtual and augmented reality. While PCs excel in providing detailed 3D visualization, this typically requires millions of points which must be efficiently coded for real-world deployment, notably storage and streaming. Recently, learning-based coding solutions have been adopted, notably in the JPEG Pleno Point Coding (PCC) standard, which uses a coding model with millions of model parameters. This requires the use of high-performance computing devices, which may not be available, notably at the decoder side. In this context, this paper proposes two reduced complexity decoding solutions, based on the adoption of scalability principles, to decode the same JPEG PCC compliant bitstreams. The design of these solutions is based on two innovative model pruning strategies which reduce the decoding complexity. The experimental results demonstrate the effective capability to significantly reduce the number of decoding model parameters with an acceptable penalty on Rate-Distortion (RD) performance compared to the full complexity model.
Mohammadreza Ghafari, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001
ICIP4
2024 Point Cloud Geometry Scalable Coding with a Quality-Conditioned Latents Probability Estimator
abstract
The widespread usage of point clouds (PC) for immersive visual applications has resulted in the use of very heterogeneous receiving conditions and devices, notably in terms of network, hardware, and display capabilities. In this scenario, quality scalability, i.e., the ability to reconstruct a signal at different qualities by progressively decoding a single bitstream, is a major requirement that has yet to be conveniently addressed, notably in most learning-based PC coding solutions. This paper proposes a quality scalability scheme, named Scalable Quality Hyperprior (SQH), adaptable to learning-based static point cloud geometry codecs, which uses a Quality-conditioned Latents Probability Estimator (QuLPE) to decode a high-quality version of a PC learning-based representation, based on an available lower quality base layer. SQH is integrated in the future JPEG PC coding standard, allowing to create a layered bitstream that can be used to progressively decode the PC geometry with increasing quality and fidelity. Experimental results show that SQH offers the quality scalability feature with very limited or no compression performance penalty at all when compared with the corresponding non-scalable solution, thus preserving the significant compression gains over other state-of-the-art PC codecs.
Daniele Mari, André F. R. Guarda, Nuno M. M. Rodrigues, Simone Milani, Fernando Pereira 0001
ICIP5
2024 JPEG AI Compressed Domain Face Detection
abstract
Learning-based image coding has achieved competitive performance in terms of compression efficiency, while also gaining a key advantage in the ability to carry out computer vision tasks directly in the compressed domain. In fact, the latent representation which is generated using deep learning techniques may natively encapsulate all visual features needed for processing tasks, thereby eliminating the need to perform the expensive synthesis transform process at the decoder side. In this paper, it is proposed to perform face detection using the latent code present in the JPEG AI architecture. First, some experiments show how decoded images can be efficiently processed for face detection without retraining, albeit with some performance degradation. Then, for the first time a compressed domain RetinaFace-based detector applied to JPEG AI latent representations is competitively proposed. The performance achieved is comparable to the performance of the original RetinaFace applied to the reconstructed JPEG AI images, while reducing computational complexity since it bypasses the image decoding process. It is expected that this approach might be extended to other vision tasks since the JPEG AI representation format is not tailored specifically for any computer vision task.
Ayman Alkhateeb, Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, João Ascenso, Fernando Pereira 0001
MMSP6
2024 Point Cloud Geometry Coding with Relational Neighborhood Self-Attention
abstract
In the ever-evolving landscape of deep learning, attention models have contributed to boost the performance in diverse fields such as computer vision and natural language processing. Following this trend, this paper proposes a novel Relational Neighborhood Self-Attention (RNSA) model, specifically designed for Point Cloud (PC) geometry coding to be integrated in the emerging learning-based JPEG PCC standard. The RNSA model proposes three new methods: first, to effectively learn correlations between the points by capturing the relational features and positions of neighboring points; second, to address the inefficiencies of conventional dot product attention, a novel Relational Scoring method to generate an attention map able to capture both linear and non-linear relationships between points and their neighbors is adopted; third, the created attention maps are normalized by Sparsemax instead of Softmax to generate sparse probabilities and assigns higher scores to the most important neighbors while marginalizing the less significant ones. Experimental results show that the proposed attention model achieves around 8% gains in both BD-Rate PSNR Dl and PSNR D2 compared to the baseline codec, i.e., JPEG PCC, while adding a small number of model parameters to JPEG PCC.
Mohammadreza Ghafari, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001
MMSP4
2023 Learning-based Point Cloud Geometry Coding Rate Control
abstract
Multimedia applications have been evolving towards providing users with more immersive and realistic experiences. A common way to model the light available for the users’ eyes is the so-called plenoptic function – a powerful 7D representation of light. There are three main types of 3D representation models for the plenoptic function, capable of expressing the light information needed to offer 6-Degrees of Freedom (DoF) experiences, namely light fields, meshes, and Point Clouds (PCs). This paper focuses on PCs since they allow representing and processing objects directly in the 3D space, facilitating user interaction and navigation in a multitude of application domains. Since the illusion of real surfaces is provided by high-density point sets, a good quality of experience requires a rather large set of points to represent a single PC, thus originating huge amounts of data to be stored and/or transmitted. Consequently, PC Coding (PCC) with significant compression levels is a must to reduce the PC data to more manageable sizes and bring PC-based applications to practical deployment. The promising results for image coding led the Joint Photographic Experts Group (JPEG) to launch a standardization project especially targeting Deep Learning (DL)-based PCC, with a final Call for Proposals in January 2022. The best performing response to this call [1] became the JPEG Pleno Learning-based PCC Verification Model (VM), which is the seed codec for the final standard. In this codec, the rate may be controlled through a set of coding parameters, largely depending on the specific PC to code, notably its sparsity and homogeneity.
Manuel Ruivo, André F. R. Guarda, Fernando Pereira 0001
DCC3
2023 Point Cloud Geometry and Color Coding in a Learning-Based Ecosystem for JPEG Coding Standards
abstract
Despite its novelty, learning-based coding for images and point clouds is already outperforming some of the best long-standing conventional codecs. In addition to its rising compression performance, learning-based coding has opened new opportunities, notably the use of a single compressed domain representation to provide both high fidelity reconstructions for human visualization as well as effective performance for computer vision tasks, effectively unifying the visual language for man and machine. This paper proposes a new double learning-based static point cloud geometry and color coding solution, which targets point cloud component scalability, rate control flexibility at coding time, and a unified compressed domain representation. The proposed solution exploits the synergies between learning-based coding for images and point clouds, through the current JPEG PCC standard for geometry coding and the JPEG AI standard for image/color coding, establishing a learning-based ecosystem for JPEG coding standards. The proposed solution is able to overcome some design limitations of the current JPEG PCC Verification Model, and significantly improve its RD performance, becoming competitive with MPEG PCC standards.
André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001
ICIP3
2023 Learning-Based Rate Control for Learning-Based Point Cloud Geometry Coding
abstract
Point Clouds represent one of the most versatile 3D visual representation models as they can provide the user the six degrees of freedom required for a truly immersive experience. In the last decade, several point cloud coding solutions have been proposed using distinct approaches, notably two MPEG standards, addressing static and dynamic point cloud coding. More recently, learning-based coding approaches started to be considered also for point cloud coding. The performance of these solutions has been so competitive that JPEG already decided to develop a point cloud coding standard adopting this novel approach. This paper proposes the first learning-based rate control mechanism to minimize the complexity associated to the selection of appropriate coding parameters for the learning-based point cloud geometry codec adopted as the initial Verification Model for the development of the JPEG Pleno Learning-based Point Cloud Coding standard.
Manuel Ruivo, André F. R. Guarda, Fernando Pereira 0001
ICIP3
2023 Deep Learning-Based Compressed Domain Point Cloud Classification
abstract
Deep learning (DL) based tools have recently reached performance levels similar to state-of-the-art hand-crafted methods for Point Cloud (PC) coding and classification. In 2022, JPEG issued a Call for Proposals for a Learning-based PC Coding (PCC) standard that envisions a unified representation, targeting both human visualization and computer vision tasks. This paper proposes the first DL-based Compressed Domain PC CLassifier (CD-PCCL), built on the PointGrid classifier, for geometry-only PCs coded with the current DL-based JPEG Pleno PCC Verification Model. The performance of compressed domain PC classification is studied against using voxel domain classification, notably for original, voxelized, and decompressed PCs. Experimental results with the ModelNet40 PC dataset show the proposed CD-PCCL can achieve significant PC classification gains regarding decompressed domain classification, while reducing the PC classifier complexity.
Abdelrahman Seleem, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001
ICIP4
2023 Deep Learning-based Point Cloud Geometry Coding with Attention Models
abstract
Recent advancements in Deep Learning (DL)-based architectures have demonstrated that integrating attention models can substantially enhance the performance across various tasks, including computer vision and visual coding. In accordance with this trend, this paper proposes a framework for incorporating attention models into DL-based Point Cloud (PC) geometry coding, namely JPEG Pleno PC Coding (JPEG PCC) and a definition of its key architectural design options. Experimental results show that the integration of attention models in JPEG PCC can provide a trade-off between compression and complexity, notably compression gains at the cost of an increase in model complexity and number of parameters. If the priority is on compression gains, the use of attention models can lead to a rate reduction up to 5.7%, for both the PSNR DI and PSNR D2 geometry quality metrics, at the cost of a 47% increase on the number of parameters and a 15.1% increase on the Floating-Point Operations (FLOPs) complexity. This obtained trade-off depends on the attention model integration configuration which can be defined depending on the application requirements.
Mohammadreza Ghafari, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001
ISM4
2022 Impact of Conventional and Deep Learning-based Point Cloud Geometry Coding on Deep Learning-based Classification Performance
abstract
Deep learning (DL)-based point cloud (PC) classification is a key computer vision task for many applications, notably autonomous driving, surveillance, and cultural heritage. In many application scenarios, PCs must be coded to reach practical rates for storage and transmission purposes, and thus they suffer from more or less intense compression artifacts. After the specification of two MPEG PC coding standards, DL-based PC coding has gained momentum, reaching competitive compression performance, especially for dense PCs. Since using decoded PCs, which may suffer from compression artifacts, may impact the final classification performance, the main goal of this paper is to study the impact of static PC geometry coding on DL-based classification. This study is performed on the ModelNet40 test dataset using the conventional G-PCC coding standard and the DL-based PC geometry codec which was the top performing solution responding to the recent JPEG Pleno PC Coding Call for Proposals. Two highly performing DL-based classifiers are used, considering the original PC geometry before and after voxelization, as well as the decoded PC geometry for different rates and qualities. As expected, coding has an impact on the classification performance, especially for the lower rates/qualities. For very sparse PCs, conventional coding still has advantage, contrarily to dense PCs, but this should change in the future with DL-based tools becoming the most natural solutions for both PC geometry coding and classification.
Abdelrahman Seleem, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001
ISM4
2022 Deep Learning-based Point Cloud Coding for Immersive Experiences
abstract
The recent advances in visual data acquisition and consumption have led to the emergence of the so-called plenoptic visual models, where Point Clouds (PCs) are playing an increasingly important role. Point clouds are a 3D visual model where the visual scene is represented through a set of points and associated attributes, notably color. To offer realistic and immersive experiences, point clouds need to have millions, or even billions, of points, thus asking for efficient representation and coding solutions. This is critical for emerging applications and services, notably virtual and augmented reality, personal communications and meetings, education and medical applications and virtual museum tours. The point cloud coding field has received many contributions in recent years, notably adopting deep learning-based approaches, and it is critical for the future of immersive media experiences. In this context, the key objective of this tutorial is to review the most relevant point cloud coding solutions available in the literature with a special focus on deep learning-based solutions and its specific novel features. Special attention will be dedicated to the ongoing standardization projects in this domain, notably in JPEG and MPEG.
Fernando Pereira 0001
ACM Multimedia1
2021 Multi-Perspective LSTM for Joint Visual Representation Learning
abstract
We present a novel LSTM cell architecture capable of learning both intra- and inter-perspective relationships available in visual sequences captured from multiple perspectives. Our architecture adopts a novel recurrent joint learning strategy that uses additional gates and memories at the cell level. We demonstrate that by using the pro-posed cell to create a network, more effective and richer visual representations are learned for recognition tasks. We validate the performance of our proposed architecture in the context of two multi-perspective visual recognition tasks namely lip reading and face recognition. Three relevant datasets are considered and the results are compared against fusion strategies, other existing multi-input LSTM architectures, and alternative recognition solutions. The experiments show the superior performance of our solution over the considered benchmarks, both in terms of recognition accuracy and complexity. We make our code publicly available at https://github.com/arsm/MPLSTM.
Alireza Sepas-Moghaddam, Fernando Pereira 0001, Paulo Lobato Correia, Ali Etemad
CVPR2
2021 Image Coding with Neural Network-Based Colorization
abstract
Automatic colorization is a process with the objective of inferring the color of grayscale images. This process is frequently used for artistic purposes and to restore the color in old or damaged images. Motivated by the excellent results obtained with deep learning-based solutions in the area of automatic colorization, this paper proposes an image coding solution integrating a deep learning-based colorization process to estimate the chrominance components based on the decoded luminance which is regularly encoded with a conventional image coding standard. In this case, the chrominance components are not coded and transmitted as usual, notably after some subsampling, as only some color hints, i.e. chrominance values for specific pixel locations, may be sent to the decoder to help it creating more accurate colorizations. To boost the colorization and final compression performance, intelligent ways to select the color hints are proposed. Experimental results show performance improvements with the increased level of intelligence in the color hints extraction process and a good subjective quality of the final decoded (and colorized) images.
Diogo Lopes, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICASSP4
2021 A Point-to-Distribution Joint Geometry and Color Metric for Point Cloud Quality Assessment
abstract
Point clouds (PCs) are a powerful 3D visual representation paradigm for many emerging application domains, especially virtual and augmented reality, and autonomous vehicles. However, the large amount of PC data required for highly immersive and realistic experiences requires the availability of efficient, lossy PC coding solutions are critical. Recently, two MPEG PC coding standards have been developed to address the relevant application requirements and further developments are expected in the future. In this context, the assessment of PC quality, notably for decoded PCs, is critical and asks for the design of efficient objective PC quality metrics. In this paper, a novel point-to-distribution metric is proposed for PC quality assessment considering both the geometry and texture. This new quality metric exploits the scale-invariance property of the Mahalanobis distance to assess first the geometry and color point-to-distribution distortions, which are after fused to obtain a joint geometry and color quality metric. The proposed quality metric significantly outperforms the best PC quality assessment metrics in the literature.
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
MMSP3
2021 Performance Evaluation of Objective Image Quality Metrics on Conventional and Learning-Based Compression Artifacts
abstract
Lossy image compression is a popular, simple and effective solution to reduce the amount of data representing digital pictures. In most lossy compression methods, the reduced volume of data in bits is achieved at the expense of introducing visual artifacts in the picture. The perceptual quality impact of such artifacts can be assessed with expensive and time-consuming subjective image quality experiments or through objective image quality metrics. However, the faster and less resource demanding objective quality metrics are not always able to reliably predict the quality as perceived by human observers. In this paper, the performance of 14 objective image quality metrics is benchmarked against a dataset of compressed images labeled with their subjective quality scores. Moreover, the performance of the above objective quality metrics in predicting the subjective quality of images distorted by both conventional and learning-based lossy compression artifacts is assessed and conclusions are drawn.
Michela Testolina, Evgeniy Upenik, João Ascenso, Fernando Pereira 0001, Touradj Ebrahimi
QoMEX4
2021 Large-Scale Crowdsourcing Subjective Quality Evaluation of Learning-Based Image Coding
abstract
Learning-based image codecs produce different compression artifacts, when compared to the blocking and blurring degradation introduced by conventional image codecs, such as JPEG, JPEG 2000 and HEIC. In this paper, a crowdsourcing based subjective quality evaluation procedure was used to benchmark a representative set of end-to-end deep learning-based image codecs submitted to the MMSP'2020 Grand Challenge on Learning-Based Image Coding and the JPEG AI Call for Evidence. For the first time, a double stimulus methodology with a continuous quality scale was applied to evaluate this type of image codecs. The subjective experiment is one of the largest ever reported including more than 240 pair-comparisons evaluated by 118 naïve subjects. The results of the benchmarking of learning-based image coding solutions against conventional codecs are organized in a dataset of differential mean opinion scores along with the stimuli and made publicly available.
Evgeniy Upenik, Michela Testolina, João Ascenso, Fernando Pereira 0001, Touradj Ebrahimi
VCIP4
2021 Lenslet Light Field Image Coding: Classifying, Reviewing and Evaluating
abstract
In recent years, visual sensors have been quickly improving, notably targeting richer acquisitions of the light present in a visual scene. In this context, the so-called lenslet light field (LLF) cameras are able to go beyond the conventional 2D visual acquisition models, by enriching the visual representation with directional light measures for each pixel position. LLF imaging is associated to large amounts of data, thus critically demanding efficient coding solutions in order applications involving transmission and storage may be deployed. For this reason, considerable research efforts have been invested in recent years in developing increasingly efficient LLF imaging coding (LLFIC) solutions. In this context, the main objective of this paper is to review and evaluate some of the most relevant LLFIC solutions in the literature, guided by a novel classification taxonomy, which allows better organizing this field. In this way, more solid conclusions can be drawn about the current LLFIC status quo, thus allowing to better drive future research and standardization developments in this technical area.
Catarina Brites, João Ascenso, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Long Short-Term Memory With Gate and State Level Fusion for Light Field-Based Face Recognition
abstract
Long Short-Term Memory (LSTM) is a prominent recurrent neural network for extracting dependencies from sequential data such as time-series and multi-view data, having achieved impressive results for different visual recognition tasks. A conventional LSTM network, hereafter referred only as LSTM network, can learn a model to posteriorly extract information from one input sequence. However, if two or more dependent sequences of data are simultaneously acquired, the LSTM networks may only process those sequences consecutively, not taking benefit of the information carried out by their mutual dependencies. In this context, this paper proposes two novel LSTM cell architectures that are able to jointly learn from multiple sequences simultaneously acquired, targeting to create richer and more effective models for recognition tasks. The efficacy of the novel LSTM cell architectures is assessed by integrating them into deep learning-based methods for face recognition with multi-view, light field images. The new cell architectures jointly learn the scene horizontal and vertical parallaxes available in a light field image, to capture richer spatio-angular information from both directions. A comprehensive evaluation, with the IST-EURECOM LFFD dataset using three challenging evaluation protocols, shows the advantage of using the novel LSTM cell architectures for face recognition over the state-of-the-art light field-based methods. These results highlight the added value of the novel cell architectures when learning from correlated input sequences.
Alireza Sepas-Moghaddam, Ali Etemad, Fernando Pereira 0001, Paulo Lobato Correia
IEEE Trans. Inf. Forensics Secur.3
2021 CapsField: Light Field-Based Face and Expression Recognition in the Wild Using Capsule Routing
abstract
Light field (LF) cameras provide rich spatio-angular visual representations by sensing the visual scene from multiple perspectives and have recently emerged as a promising technology to boost the performance of human-machine systems such as biometrics and affective computing. Despite the significant success of LF representation for constrained facial image analysis, this technology has never been used for face and expression recognition in the wild. In this context, this paper proposes a new deep face and expression recognition solution, called CapsField, based on a convolutional neural network and an additional capsule network that utilizes dynamic routing to learn hierarchical relations between capsules. CapsField extracts the spatial features from facial images and learns the angular part-whole relations for a selected set of 2D sub-aperture images rendered from each LF image. To analyze the performance of the proposed solution in the wild, the first in the wild LF face dataset, along with a new complementary constrained face dataset captured from the same subjects recorded earlier have been captured and are made available. A subset of the in the wild dataset contains facial images with different expressions, annotated for usage in the context of face expression recognition tests. An extensive performance assessment study using the new datasets has been conducted for the proposed and relevant prior solutions, showing that the CapsField proposed solution achieves superior performance for both face and expression recognition tasks when compared to the state-of-the-art.
Alireza Sepas-Moghaddam, Ali Etemad, Fernando Pereira 0001, Paulo Lobato Correia
IEEE Trans. Image Process.3
2021 Constant Size Point Cloud Clustering: A Compact, Non-Overlapping Solution
abstract
Point clouds have recently become a popular 3D representation model for many application domains, notably virtual and augmented reality. Since point cloud data is often very large, processing a point cloud may require that it be segmented into smaller clusters. For example, the input to deep learning-based methods like auto-encoders should be constant size point cloud clusters, which are ideally compact and non-overlapping. However, given the unorganized nature of point clouds, defining the specific data segments to code is not always trivial. This paper proposes a point cloud clustering algorithm which targets five main goals: i) clusters with a constant number of points; ii) compact clusters, i.e., with low dispersion; iii) non-overlapping clusters, i.e., not intersecting each other; iv) ability to scale with the number of points; and v) low complexity. After appropriate initialization, the proposed algorithm transfers points between neighboring clusters as a propagation wave, filling or emptying clusters until they achieve the same size. The proposed algorithm is unique since there is no other point cloud clustering method available in the literature offering the same clustering features for large point clouds at such low complexity.
André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001
IEEE Trans. Multim.3
2021 Point Cloud Rendering After Coding: Impacts on Subjective and Objective Quality
abstract
Recently, point clouds have shown to be a promising way to represent 3D visual data for a wide range of immersive applications, from augmented reality to autonomous cars. Emerging imaging sensors have made easier to perform richer and denser point cloud acquisition, notably with millions of points, thus raising the need for efficient point cloud coding solutions. In such scenario, it is important to evaluate the impact and performance of several processing steps in a point cloud communication system, notably the degradations associated to point cloud coding solutions. Moreover, since point clouds are not directly visualized but rather processed with a rendering algorithm before shown on any display, the perceived quality of point cloud data highly depends on the rendering solution. In this context, the main objective of this paper is to study the impact of several coding and rendering solutions on the perceived user quality and in the performance of available objective assessment metrics. Another contribution regards the assessment of recent MPEG point cloud coding solutions for several popular rendering methods, which was never presented before. The conclusions regard the visibility of three types of coding artifacts for the three considered rendering approaches as well as the strengths and weaknesses of objective metrics when point clouds are rendered after coding.
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
IEEE Trans. Multim.3
2020 Facial Emotion Recognition Using Light Field Images with Deep Attention-Based Bidirectional LSTM
abstract
Light field cameras are able to capture the intensity of light rays coming from multiple directions, thus representing the visual scene from multiple viewpoints. This paper exploits the rich spatio-angular information available in light field images for facial emotion recognition. In this context, a new deep network is proposed that first extracts spatial features using a VGG16 convolutional neural network. Then, a Bidirectional Long Short-Term Memory (Bi-LSTM) recurrent neural network is used to learn spatio-angular features from viewpoint feature sequences, exploring both forward and backward angular relationships. Additionally, an attention mechanism allows our model to selectively focus on the most important spatio-angular features, thus enabling a more effective learning outcome. Finally, a fusion scheme is adopted to obtain the emotion recognition classification results. Comprehensive experiments have been conducted on the IST-EURECOM Light Field Face database using two challenging evaluation protocols, showing the superiority of our method over the state-of-the-art.
Alireza Sepas-Moghaddam, Ali Etemad, Fernando Pereira 0001, Paulo Lobato Correia
ICASSP3
2020 Point Cloud Geometry Scalable Coding With a Single End-to-End Deep Learning Model
abstract
Point clouds are gaining importance as the format to represent complex 3D objects and scenes, offering high user immersion and interaction, although at the cost of requiring massive data. Scalable coding is an important feature for point cloud coding, especially for real-time applications, where the fast and bitrate efficient access to a decoded point cloud is important; however, this issue is still rather unexplored in the literature. With the rise of deep learning methods as a promising solution for efficient coding, this paper proposes the first deep learning-based point cloud geometry scalable coding solution. Experimental results show that the proposed scalable coding solution consistently outperforms the MPEG standard for static point cloud geometry coding. In this way, a new research path is open for point cloud scalable coding technology.
André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001
ICIP3
2020 Improving Psnr-Based Quality Metrics Performance For Point Cloud Geometry
abstract
An increased interest in immersive applications has drawn attention to emerging 3D imaging representation formats, notably light fields and point clouds (PCs). Nowadays, PCs are one of the most popular 3D media formats, due to recent developments in PC acquisition, namely with new depth sensors and signal processing algorithms. To obtain high fidelity 3D representations of visual scenes a huge amount of PC data is typically acquired, which demands efficient compression solutions. As in 2D media formats, the final perceived PC quality plays an importance role in the overall user experience and, thus, objective metrics capable to measure the PC quality in a reliable way are essential. In this context, this paper proposes and evaluates a set of objective quality metrics for the geometry component of PC data, which plays a very important role on the final perceived quality. Based on the popular PSNR PC geometry quality metric, novel improved PSNR-based metrics are proposed by exploiting the intrinsic PC characteristics and the rendering process that must occur before visualization. The experimental results show the superiority of the best proposed metrics over state-of-the-art, obtaining an improvement up to 32% in the Pearson correlation coefficient.
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
ICIP3
2020 Deep Learning-based Point Cloud Geometry Coding with Resolution Scalability
abstract
Point clouds are a 3D visual representation format that has recently become fundamentally important for immersive and interactive multimedia applications. Considering the high number of points of practically relevant point clouds, and their increasing market demand, efficient point cloud coding has become a vital research topic. In addition, scalability is an important feature for point cloud coding, especially for real-time applications, where the fast and rate efficient access to a decoded point cloud is important; however, this issue is still rather unexplored in the literature. In this context, this paper proposes a novel deep learning-based point cloud geometry coding solution with resolution scalability via interlaced sub-sampling. As additional layers are decoded, the number of points in the reconstructed point cloud increases as well as the overall quality. Experimental results show that the proposed scalable point cloud geometry coding solution outperforms the recent MPEG Geometry-based Point Cloud Compression standard which is much less scalable.
André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001
MMSP3
2020 A Generalized Hausdorff Distance Based Quality Metric for Point Cloud Geometry
abstract
Reliable quality assessment of decoded point cloud geometry is essential to evaluate the compression performance of emerging point cloud coding solutions and guarantee some target quality of experience. This paper proposes a novel point cloud geometry quality assessment metric based on a generalization of the Hausdorff distance. To achieve this goal, the so-called generalized Hausdorff distance for multiple rankings is exploited to identify the best performing quality metric in terms of correlation with the MOS scores obtained from a subjective test campaign. The experimental results show that the quality metric derived from the classical Hausdorff distance leads to low objective-subjective correlation and, thus, fails to accurately evaluate the quality of decoded point clouds for emerging codecs. However, the quality metric derived from the generalized Hausdorff distance with an appropriately selected ranking, outperforms the MPEG adopted geometry quality metrics when decoded point clouds with different types of coding distortions are considered.
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
QoMEX3
2020 Visual monitoring of High-Sea fishing activities using deep learning-based image processing
Pedro Perdigão, Pedro Lousã, João Ascenso, Fernando Pereira 0001
Multim. Tools Appl.4
2020 Point cloud coding: A privileged view driven by a classification taxonomy
Fernando Pereira 0001, Antoine Dricot, João Ascenso, Catarina Brites
Signal Process. Image Commun.1
2020 Versatile Video Coding Based Quality Scalability With Joint Layer Reference
abstract
Scalability is an essential coding feature for adaptive video streaming applications, notably considering the growing heterogeneity of the transmission, display and consumption environments. Versatile video coding (VVC) is the emerging video coding standard, targeting offering higher compression efficiency regarding previous standards to further facilitate already available and novel video applications, notably at higher spatial resolutions. In this context, this letter proposes the first VVC-based quality scalability extension, targeting to offer higher compression efficiency than the native VVC quality scalability solution. The proposed Quality Scalable Versatile Video Coding (QS-VVC) solution is designed based on a layered coding approach with one base layer (BL) and one or more enhancement layers (EL). To achieve higher compression performance, a novel joint layer referencing approach is proposed where the base and enhancement layers decoded information are jointly exploited to create a new EL coding reference. Experimental results shown that the proposed QS-VVC codec outperforms the most relevant benchmarks, notably VVC-based simulcasting, native VVC quality scalability, and the previous Scalable High Efficiency Video Coding (SHVC) standard.
Xiem HoangVan, Sang NguyenQuang, Fernando Pereira 0001
IEEE Signal Process. Lett.3
2020 Mahalanobis Based Point to Distribution Metric for Point Cloud Geometry Quality Evaluation
abstract
Nowadays, point clouds (PCs) are a promising representation format for immersive content and target several emerging applications, notably in virtual and augmented reality. However, efficient coding solutions are critically needed due to the large amount of PC data required for high quality user experiences. To address these needs, several PC coding standards were developed and thus, objective PC quality metrics able to accurately account for the subjective impact of coding artifacts are needed. In this paper, a scale-invariant PC geometry quality assessment metric is proposed based on a new type of correspondence, namely between a point and a distribution of points. This metric is able to reliably measure the geometry quality for PCs with different intrinsic characteristics and degraded by several coding solutions. Experimental results show the superiority of the proposed PC quality metric over relevant state-of-the-art.
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
IEEE Signal Process. Lett.3
2020 A Double-Deep Spatio-Angular Learning Framework for Light Field-Based Face Recognition
abstract
Face recognition has attracted increasing attention due to its wide range of applications, but it is still challenging when facing large variations in the biometric data characteristics. Lenslet light field cameras have recently come into prominence to capture rich spatio-angular information, thus offering new possibilities for advanced biometric recognition systems. This paper proposes a double-deep spatio-angular learning framework for light field-based face recognition, which is able to model both the intra-view/spatial and inter-view/angular information using two deep networks in sequence. This is a novel recognition framework that has never been proposed in the literature for face recognition or any other visual recognition task. The proposed double-deep learning framework includes a long short-term memory (LSTM) recurrent network, whose inputs are VGG-Face descriptions, computed using a VGG-16 convolutional neural network (CNN). The VGG-Face spatial descriptions are extracted from a selected set of 2D sub-aperture (SA) images rendered from the light field image, corresponding to different observation angles. A sequence of the VGG-Face spatial descriptions is then analyzed by the LSTM network. A comprehensive set of experiments has been conducted using the IST-EURECOM light field face database, addressing varied and challenging recognition tasks. The results show that the proposed framework achieves superior face recognition performance when compared to the state of the art.
Alireza Sepas-Moghaddam, Mohammad A. Haque, Paulo Lobato Correia, Kamal Nasrollahi, Thomas B. Moeslund, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.6
2019 A Deep Framework for Facial Emotion Recognition using Light Field Images
abstract
Light field cameras capture the intensity of light rays coming from multiple directions, thus allowing a set of 2D images, named sub-aperture (SA) images, to be rendered. These images correspond to observations of the scene from slightly different angles. The rich spatio-angular information obtained using these cameras is exploited in this paper, for the first time, in the context of facial emotion recognition. A deep learning spatio-angular fusion framework is adopted which is able to model both the intra-view/spatial and inter-view/angular information, using a VGG-16 convolutional neural network and a long short-term memory (LSTM) recurrent network. The proposed solution, based on the adopted deep spatio-angular fusion framework, creates two view sequences, horizontal and vertical, with selected SA images, for which VGG-Face descriptions are extracted. The resulting descriptions are fed to two LSTM networks, with the aim of independently learning horizontal and vertical classification models. The softmax classifier scores obtained for the horizontal and vertical descriptors are then fused to obtain the final emotion recognition labels. A comprehensive set of experiments has been conducted on the IST-EURECOM light field face database using two assessment protocols. The adopted framework achieves superior emotion recognition performance when compared with state-of-the-art benchmarking methods.
Alireza Sepas-Moghaddam, Ali Etemad, Paulo Lobato Correia, Fernando Pereira 0001
ACII4
2019 Adaptive Plane Projection for Video-Based Point Cloud Coding
abstract
One of the most promising emerging 3D representation paradigms is the point cloud (PC) model, notably due to the new set of applications that it enables, from immersive telepresence to 3D geographic information systems. Recognizing the potential of this representation model, MPEG has launched a standardization project to specify efficient PC coding solutions. This project has led to the so-called Video-based Point Cloud Coding (V-PCC) standard, which is based on the idea of projecting the dynamic PC geometry and texture into a sequence of frames to be coded with the highly efficient HEVC video coding standard. In V-PCC, the projection of points and color attributes is always performed using the same, rigid set of projection planes, independently of the PC characteristics. This paper proposes a more flexible coding solution, which improves the V-PCC Intra coding mode by adopting a content dependent set of projection planes, thus, more adapted to the characteristics of the PCs to be coded, targeting a better compression performance. The experimental results show an average total bitrate saving of around 15% regarding V-PCC with clear benefits in terms of subjective assessment.
Eurico Lopes, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICME4
2019 Improved Patch Packing for the MPEG V-PCC Standard
abstract
Point cloud representation is an emerging visual data technology, targeting immersive 3D experiences in the context of multiple applications scenarios, notably entertainment, geographical information systems, medicine, architecture, and robotics. Since a point cloud may easily involve millions of points, and thus an enormous amount of data, its effective storage and transmission critically asks for efficient coding solutions. With this purpose in mind, several point cloud coding (PCC) solutions have been proposed in the literature; special emphasis is due to the recent MPEG standards, which target interoperability in this domain, notably the MPEG V-PCC standard. The objective of this paper is to improve the V-PCC standard compression efficiency by proposing novel solutions for the V-PCC packing module without compromising in any way the V-PCC stream (syntax and semantics) and decoder compliance. In this context, several patch packing solutions are proposed, including new packing algorithms and associated sorting and positioning metrics; for the metrics, both absolute and relative approaches are proposed. The RD performance results show BD-Rate savings up to 0.8% for the best packing solution regarding the V-PCC benchmark. Moreover, the packing map size reductions can go up to 12%, on average.
Afonso Costa, Antoine Dricot, Catarina Brites, João Ascenso, Fernando Pereira 0001
MMSP5
2019 Point Cloud Coding: Adopting a Deep Learning-based Approach
abstract
Point clouds have recently become an important visual representation format, especially for virtual and augmented reality applications, thus making point cloud coding a very hot research topic. Deep learning-based coding methods have recently emerged in the field of image coding with increasing success. These coding solutions take advantage of the ability of convolutional neural networks to extract adaptive features from the images to create a latent representation that can be efficiently coded. In this context, this paper extends the deep-learning coding approach to point cloud coding using an autoencoder network design. Performance results are very promising, showing improvements over the Point Cloud Library codec often taken as benchmark, thus suggesting a significant margin of evolution for this new point cloud coding paradigm.
André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001
PCS3
2019 Graph-Based Static 3D Point Clouds Geometry Coding
abstract
Recently, 3D visual representation models such as light fields and point clouds are becoming popular due to their capability to represent the real world in a more complete and immersive way, paving the road for new and more advanced visual experiences. The point cloud representation model is able to efficiently represent the surface of objects/scenes by means of a set of 3D points and associated attributes and is increasingly being used from autonomous cars to augmented reality. Emerging imaging sensors have made it easier to perform richer and denser point cloud acquisitions, notably with millions of points, making it impossible to store and transmit these very high amounts of data without appropriate coding. This bottleneck has raised the need for efficient point cloud coding solutions in order to offer more immersive visual experiences and better quality of experience to the users. In this context, this paper proposes an efficient lossy coding solution for the geometry of static point clouds. The proposed coding solution uses an octree-based approach for a base layer and a graph-based transform approach for the enhancement layer where an Inter-layer residual is coded. The performance assessment shows very significant compression gains regarding the state-of-the-art, especially for the most relevant lower and medium rates.
Paulo de Oliveira Rente, Catarina Brites, João Ascenso, Fernando Pereira 0001
IEEE Trans. Multim.4
2018 A Study on the 4D Sparsity of JPEG Pleno Light Fields Using the Discrete Cosine Transform
abstract
In this work we study the 4D sparsity of light fields using as main tool the 4D- Discrete Cosine Transform. We analyze the two JPEG Pleno light field datasets, namely the lenslet-based and the High-Density Camera Array (HDCA) datasets. The results suggest that the lenslets datasets exhibit a high 4D redundancy, with a larger inter-view sparsity than the intra-view one. For the HDCA datasets, there is also 4D redundancy worthy to be exploited, yet in a smaller degree. Unlike the lenslets case, the intra-view redundancy is much larger than the inter-view one. The results and conclusions of this first study for light field imaging may have a strong impact on the current and future design of efficient codecs for this type of emerging data, notably in the context of the JPEG Pleno standard.
Gustavo Alves 0001, Marcio P. Pereira, Murilo B. de Carvalho, Fernando Pereira 0001, Carla L. Pagliari, Vanessa Testoni, Eduardo A. B. da Silva
ICIP4
2018 A 4D DCT-Based Lenslet Light Field Codec
abstract
Light fields aim to represent visual information in 3D space. They are 4D structures that contain the images of a given scene from a sampled 2D range of viewpoints. When acquired using a lenslet camera, in addition to the ordinary intra-view redundancy, these views have a great deal of inter-view redundancy. In this work we propose a light field codec that fully exploits the 4D redundancy of light fields by using a 4D transform and hexadeca-trees. It initially divides the light field into 4D blocks and computes a 4D Discrete Cosine Transform of each one. Then the transform coefficients of the 4D block are grouped using hexadeca-trees on a bitplane-by-bitplane basis, and the generated stream is encoded using an adaptive arithmetic coder. The proposed codec has been employed to encode the JPEG Pleno lenslet light fields. The rate-distortion results have been assessed using test conditions comparable to the ones presented at the ICIP 2017 Light Field Coding Grand Challenge. The proposed codec, despite being conceptually simple, achieves competitive rate-distortion performance.
Murilo B. de Carvalho, Marcio P. Pereira, Gustavo Alves 0001, Eduardo A. B. da Silva, Carla L. Pagliari, Fernando Pereira 0001, Vanessa Testoni
ICIP6
2018 Rate-Distortion Driven Adaptive Partitioning for Octree-Based Point Cloud Geometry Coding
abstract
Point clouds are widely recognized as a promising 3D visual representation model for future 3D applications, e.g. Augmented/Virtual Reality and autonomous mobile navigation. To enable immersive and high quality experiences, a large amount of data needs to be transmitted, and thus, efficient compression techniques are a key factor for its success. Nowadays, many point cloud coding techniques rely on layered octree data structures able to provide multiple benefits, notably level of detail scalability. However, the octree creation process still lacks flexibility compared to the optimal partitioning methods popular in 2D encoders, which are mainly driven by rate-distortion trade-offs. The goal of this paper is to propose a novel octree creation mechanism which is able to exploit the intrinsic point cloud resolution by iteratively partitioning the octree voxels by means of split decisions based on a Rate-Distortion Optimization process. In addition, the arithmetic coding step is adapted to the density at each octree depth. This novel approach allows reaching a significant bitrate reduction, notably 5% in average, and up to 9.8%, over the adopted coding anchor using state-of-the-art octree based coding.
Antoine Dricot, Fernando Pereira 0001, João Ascenso
ICIP2
2018 Optimal Lagrange multipliers for dependent rate allocation in video coding
Ana De Abreu, Gene Cheung, Pascal Frossard, Fernando Pereira 0001
Signal Process. Image Commun.4
2018 Light Field-Based Face Presentation Attack Detection: Reviewing, Benchmarking and One Step Further
abstract
Vulnerability of face recognition systems to presentation attacks has attracted increasing attention from the biometrics and forensics communities. Moreover, the recent availability of light field cameras is opening new possibilities for designing improved face presentation attack detection solutions. In this context, this paper provides the first review and benchmarking study in the literature on light field-based face presentation attack detection solutions. State-of-the-art solutions are assessed in terms of accuracy, generalization and complexity, using a common, representative evaluation framework. This paper also proposes a novel face presentation attack detection solution, based on a histogram of oriented gradients descriptor, which exploits the disparity information available in light field imaging. The evaluation of the proposed face presentation attack detection solution for different presentation attack types shows a very effective and stable performance, notably performing better than the state-of-the-art alternatives.
Alireza Sepas-Moghaddam, Fernando Pereira 0001, Paulo Lobato Correia
IEEE Trans. Inf. Forensics Secur.2
2018 Holographic Data Coding: Benchmarking and Extending HEVC With Adapted Transforms
abstract
Holography is an emerging technology to represent and display visual information with high expectations in terms of user experience. A hologram is a reproduction of a light field represented through the interference pattern between two wavefields the reference and the object wavefields. Whatever their creation process holograms may have a digital representation using some appropriate format. Moreover considering the huge amounts of data involved digital holographic data have to be compressed using appropriate coding solutions for example available image coding standard solutions or efficient extensions of them. In this context this paper contributes to advance the state-of-the-art on holographic data coding by: 1) benchmarking the most relevant available image coding standard solutions when using the most relevant holographic data representation formats; 2) proposing a novel mode depend directional transform-based HEVC coding solution trained with holographic data. Experimental results obtained under meaningful test conditions show that the proposed coding solution outperforms the state-of-the-art HEVC coding standard for specific formats and conditions. Altogether these two contributions are critical to understand the current status quo and advance the state-of-the-art on holographic data coding.
Jose Peixeiro, Catarina Brites, João Ascenso, Fernando Pereira 0001
IEEE Trans. Multim.4
2017 Light field local binary patterns description for face recognition
abstract
Light field cameras are emerging as powerful sensor devices to capture the full spatio-angular visual information in a viewing range. As more information should allow better analysis performance, this paper proposes a simple, yet effective descriptor, named Light Field Local Binary Patterns (LFLBP), able to exploit the richer information available in light field images for face recognition. The LFLBP descriptor combines two main components, the spatial, local LBP and the angular LBP, to capture not only the usual spatial information but also the light field angular information associated to the set of sub-aperture images, corresponding to different viewpoints. Experiments were conducted with the novel IST-EURECOM light field face database. When compared with competing methods, the proposed descriptor has shown superior face recognition performance under varied and challenging acquisition conditions. Moreover, the proposed light field angular LBP descriptor can be flexibly combined with any available spatial descriptor to derive combined descriptors for enhanced light field based face recognition performance.
Alireza Sepas-Moghaddam, Paulo Lobato Correia, Fernando Pereira 0001
ICIP3
2017 Improving point cloud to surface reconstruction with generalized Tikhonov regularization
abstract
Point cloud rendering has a vital role in the user Quality of Experience for applications adopting point cloud based representations. While this is not a new area, it has recently become more relevant with the recent interest on point cloud coding by major standardization groups, notably JPEG and MPEG. The screened Poisson surface reconstruction is a state-of-the-art technique for generating a watertight surface mesh from the point cloud samples. While its screening component allows the surface to better fit the cloud points, this fitting may lead to undesired artifacts in the surface, notably when the point cloud is noisy. This paper proposes to improve this reconstruction method by making it more robust to noise by adopting a generalized Tikhonov regularization term. The proposed regularization approach smooths regions that should be flat while keeping the important details in the edges, thus creating more pleasant surface reconstructions.
André F. R. Guarda, José M. Bioucas-Dias, Nuno M. M. Rodrigues, Fernando Pereira 0001
MMSP4
2017 Subjective and objective quality evaluation of compressed point clouds
abstract
The increasing availability of point cloud data in recent years is demanding high performance compression solutions. Naturally, methods to perform objective quality assessment of compressed point clouds are also very much needed, namely metrics to measure the geometry distortion of point clouds when positioning errors are present. This is a rather challenging problem since this 3D representation format is unstructured and it is typically not directly visualized. In this context, the objective of this paper is to perform subjective and objective quality assessment of point clouds degraded by compression artifacts and to evaluate the correlation of the most popular objective quality metrics with human perception. In this work, subjective experiments conducted at Instituto Superior Técnico (IST) are described with point clouds compressed with two different but yet promising solutions, one based on the octree representation of the 3D space and another based on the rather popular graph transform. As far as the authors know, this is the first study of this type made available and should have a key role on the future development and evaluation of point cloud coding solutions.1
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
MMSP3
2017 A geometric space-view redundancy descriptor for light fields: Predicting the compression potential of the JPEG Pleno light field datasets
abstract
The representation of data in terms of its statistical properties is valuable in many applications. This work uses statistics obtained from 4D scene geometry to characterize, in terms of redundancy, the content produced by lenslet-based light field cameras and by high-density arrays of cameras for the JPEG Pleno Call for Proposals on Light Field Coding. This paper proposes a novel so-called geometric space-view redundancy (GSVR) descriptor, which is able to characterize the amount of redundancy in light fields thus bringing information about the trade-offs involved in effectively exploring redundancy for efficient coding. The redundancy is here measured by the probability, for each block size and range of views, that the image of a given 3D point belongs to the block in all views. For a given probability, the GSVR descriptor models the spaceview correlation, i.e. the correlation between the intra-view block dimensions and the number of views. Therefore, it is a descriptor that has application on dataset selection and encoder control and optimization. The JPEG Pleno datasets are analyzed in terms of the GSVR descriptor in all views.
Marcio P. Pereira, Gustavo Alves 0001, Carla L. Pagliari, Murilo B. de Carvalho, Eduardo A. B. da Silva, Fernando Pereira 0001
MMSP6
2017 Epipolar based light field key-location detector
abstract
Nowadays, visual features play a key role, as they can provide a concise representation of visual data that is efficient for multiple tasks, notably content retrieval and object recognition. In parallel, visual sensors have been improving, targeting richer acquisitions of the light in a visual scene. In this context, the so-called light field cameras, which have recently emerged, are able to go beyond the standard acquisition models, by enriching the visual representation with directional light measures for each pixel position, e.g. by using a so-called lenslet light field camera. At this stage, not much research has been made in the field of feature detection and description for the emerging lenslet light field format. In this context, this paper proposes a feature detector suitable for lenslet light field images based on the exploitation of an alternative visual parametrization of the light field, called the Epipolar Planar Image (EPI). The proposed detector is heavily based on line detection in the EPI representation, since 3D points in the visual scene are mapped to line segments in an EPI, and the detector output is referred as key-locations. The proposed light field key-location detector is assessed with a solid evaluation framework using a large light field dataset. In comparison to the 2D SIFT detector, up to 10% improvements were achieved for the widely used repeatability metric.
José Abecasis Teixeira, Catarina Brites, Fernando Pereira 0001, João Ascenso
MMSP3
2017 Saliency-driven omnidirectional imaging adaptive coding: Modeling and assessment
abstract
Omnidirectional imaging, also known as 360° and spherical imaging, records all 360° of a scene from a specific spatial position, thus offering the user the capability to enjoy three rotational degrees of freedom (3-DoF). To offer a good quality of experience, omnidirectional imaging requires very high bitrates as high spatial resolution are a must and, ideally, also high frame rates. Due to the lack of video coding solutions specifically designed for omnidirectional imaging, this type of content is typically coded with the available image and video coding standards, such as JPEG, H.264/AVC and HEVC, after applying a 2D rectangular projection. In this context, this paper proposes an omnidirectional imaging coding solution allowing to reach improved coding performance by using an adaptive coding solution where the most visually salient image/video regions are coded with higher quality in a process appropriately controlled by the quantization parameter. To determine the saliency of the various omnidirectional imaging regions, a machine-learning based saliency detection model is proposed. The proposed coding solution achieves compression gains as measured by a novel objective quality metric also driven by saliency. This novel objective quality metric is validated by formal subjective testing where very high correlations with the subjective tests scores are achieved.
Guilherme Luz Tortorella, João Ascenso, Catarina Brites, Fernando Pereira 0001
MMSP4
2017 Adaptive Scalable Video Coding: An HEVC-Based Framework Combining the Predictive and Distributed Paradigms
abstract
The emerging scalable High Efficiency Video Coding (SHVC) video coding standard provides an efficient solution for transmission of video over heterogeneous and time dynamic networks, terminals, and usage environments. The encoding complexity and the error sensitivity associated with the efficient HEVC coding tools adopted in SHVC make this scalable codec less attractive to some emerging applications such as video surveillance, visual sensor network, and remote space transmission where these requirements are critical. To address the requirements of these application scenarios including scalability, this paper proposes a novel HEVC-based framework offering quality scalability on top of an HEVC compliant base layer while appropriately combining the predictive and distributed coding paradigms. To achieve the best enhancement layer compression efficiency, two novel coding tools are proposed, notably a machine learning-based side information creation mechanism and an adaptive correlation modeling process. The experimental results reveal that the rate-distortion performance of the proposed distributed scalable video coding-HEVC solution outperforms the relevant alternative coding solutions, notably by up to 52.9% and 23.7% BD-rate gains regarding the HEVC-Simulcast and SHVC standard solutions, respectively, for an equivalent prediction configuration, while achieving a lower encoding complexity.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.3
2016 Improving SHVC performance with a joint layer coding mode
abstract
The growing need for a powerful scalable video coding engine targeting the heterogeneous landscape of network, devices, and consumption environments has led to the development of the Scalable High Efficiency Video Coding (SHVC) standard, an extension of the High Efficiency Video Coding (HEVC) standard. To improve the SHVC compression efficiency, this paper proposes a novel joint layer coding mode to be integrated in the SHVC codec. In the proposed coding mode, the base layer (BL) and enhancement layer (EL) decoded information are linearly combined at the pixel level to create an additional coding mode. To fuse the BL and EL driven predictions, a weighting term is defined to indicate the contributions of each of them for the final joint layer prediction. To reach high adaptability, these weights are computed at pixel level in the prediction unit. Moreover, to achieve the highest compression efficiency, the proposed joint layer coding mode is adaptively selected using a rate distortion optimization (RDO) mechanism. Experiments conducted for a rich set of test conditions have shown that significant compression efficiency gains can be achieved with the proposed joint layer coding mode, notably up to 4.3 % in BD-Rate savings regarding the standard SHVC quality scalable codec.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
ICASSP3
2016 Multi-view distributed source coding of binary features for visual sensor networks
abstract
Visual analysis algorithms have been mostly developed for a centralized scenario where all visual data is acquired and processed at a central location. However, in visual sensor networks (VSN), several constraints in computational power, energy and bandwidth require a radically different approach, notably a paradigm shift from centralized to distributed visual processing. In the new paradigm, visual data is acquired and features are extracted at the sensing nodes locations to be after transmitted to enable further analysis at some central location. In such scenario, one of the key challenges is to design suitable feature coding schemes that are able to exploit the correlation among the features corresponding to (partially) overlapped views of the same visual scene. To achieve efficient coding, it is proposed to employ the distributed source coding paradigm as it does not require any communication between the sensing nodes (rather expensive in VSN) and it is parsimonious in terms of computational resources. Experimental results show that significant accuracy and compression gains (up to 37.36%) can be achieved when coding features extracted from multiple views.
Nuno Monteiro, Catarina Brites, Fernando Pereira 0001, João Ascenso
ICASSP3
2016 Multi-view distributed coding and selection of local binary features
abstract
Recently, the latest advances in compact feature representation and feature learning have provided an efficient framework for several visual analysis tasks, such as object recognition. However, when multiple cameras with overlapping fields-of-view are employed, other visual analysis tasks such as depth estimation can be supported and object recognition accuracy can be improved. In this paper the problem of distributed visual analysis from multiple views of a scene is addressed, considering that computational power and bandwidth, at each camera sensor, are rather limited. More specifically, an efficient coding technique for local binary features is proposed which exploits the correlation at the decoder side between each descriptor and its quantized representation. Moreover, considering that descriptors representing the same visual feature across different views are well correlated, a technique to avoid the transmission of redundant descriptors from multiple views is proposed. At the decoder, the joint statistics of all descriptors from all views is used to drive the selection of the best descriptors to be transmitted by each sensing node. The proposed multi-view feature coding and selection techniques allow obtaining bitrate reductions up to 80%, with respect to the uncompressed descriptor rate, for a certain task accuracy.
Nuno Monteiro, Catarina Brites, Fernando Pereira 0001, João Ascenso
ICME3
2016 Digital holography: Benchmarking coding standards and representation formats
abstract
Holography is an emerging technology to represent and display visual information and the associated experience expectations on the layman imaginary are huge. A hologram is a reproduction of a light field represented through the interference pattern between two wavefields, the reference and object wavefields. Holograms may be optically generated from real objects or computationally generated from synthetic objects leading to the so-called computer generated holograms. Whatever the creation process, holograms may have a digital representation using some appropriate representation format. Moreover, considering the huge amounts of data involved, digital holographic data has to be compressed using appropriate coding solutions, possibly available standard coding solutions. Finally, the decoded holographic data has to lead to reconstructed images using some appropriate reconstruction method. Holographic data coding is an emerging field of research where literature is still very scarce. In this context, before starting designing coding solutions considering the specific characteristics of holographic data, it is essential to assess the current compression performance associated to the most relevant available standard coding solutions and alternative representations formats. This paper has the main objective to benchmark the most relevant available standard coding solutions and the main alternative representation formats under meaningful test conditions and for the test material currently available. This assessment is critical to understand the current status quo and launch future research.
Jose Peixeiro, Catarina Brites, João Ascenso, Fernando Pereira 0001
ICME4
2016 Efficient plenoptic imaging representation: Why do we need it?
abstract
The 3D representation of the world visual information has been a challenge for a long time both in the analogue and digital domains. At least in the past decade, 3D stereo-based solutions have become very common. However, several constraints and limitations ended up causing a negative impact on its user popularity and market deployment. Recent developments in terms of acquisition and display devices have shown that it is possible to offer more immersive and powerful 3D experiences by adopting higher dimensional representations. In this context, the so-called plenoptic function offers an excellent framework to analyze and discuss the recent and future developments towards improved 3D imaging representations, functionalities and applications. Since they are associated to huge amounts of data, the new imaging modalities such as light fields and point clouds critically ask for appropriate efficient coding solutions. In this context, the main objective of this paper is to present, organize and discuss the recent trends and future developments on 3D visual data representation in a plenoptic function framework. This is critical to effectively plan the next research and standardization steps on 3D imaging representation and coding.
Fernando Pereira 0001, Eduardo A. B. da Silva
ICME1
2016 Feature-based video coding: Designing an RD efficient and search friendly framework
abstract
To provide more powerful video enabled applications, e.g. in video surveillance environments, it is increasingly more critical not only to have access to the decoded video but also to, e.g. efficiently search for similar videos. In this context, this paper proposes a feature-based video coding solution adopting a hybrid approach where both pixels and local visual features are exploited for coding. In this novel solution, part of the frames are coded using a set of key point matches, thus allowing not only to decode the usual frames for visualization but also valuable key point information extracted from uncompressed frames which is instrumental for searching. Experimental results for video surveillance like sequences and conditions show bitrate savings regarding the state-of-the-art HEVC standard while additionally facilitating more accurate searching.
Renam C. da Silva, Fernando Pereira 0001, Eduardo A. B. da Silva
PCS2
2016 Objective and subjective evaluation of light field image compression algorithms
abstract
This paper reports results of subjective and objective quality assessments of responses to a grand challenge on light field image compression. The goal of the challenge was to collect and evaluate new compression algorithms for light field images. In total seven proposals were received, out of which five were accepted for further evaluations. For objective evaluations, conventional metrics were used, whereas the double stimulus continuous quality scale method was selected to perform subjective assessments. Results show competitive performance among submitted proposals. However, in low bitrates, one proposal outperforms the others.
Irene Viola 0001, Martin Rerábek, Tim Bruylants, Peter Schelkens, Fernando Pereira 0001, Touradj Ebrahimi
PCS5
2015 Stereo Based Tracking-by-Detection for Visual Sensor Networks
abstract
Visual binary descriptors have successfully been employed in several applications such as visual search, object recognition and visual tracking. In particular, binary descriptors are suitable for scenarios where computational, storage and energy resources are constrained and have been previously exploited to track an object along a video sequence. In this paper, binary descriptors are used to perform visual tracking in a stereo-based system, i.e. when two cameras with overlapping views are employed in a cooperative way. The proposed stereo-based visual tracker follows the tracking-by-detection approach where features extracted from different cameras are used to characterize the object appearance with a suitable model. Moreover visual tracking is performed at a central controller by just using the features transmitted from the two camera nodes. To achieve this target, efficient coding techniques are proposed to reduce the amount of feature data that is transmitted through the network. The performance of the proposed stereo-based visual tracker is evaluated in terms of rate-accuracy, i.e. using quantitative metrics to assess the accuracy of the visual tracker as a function of the coding bitrate.
Berner Panti, Fernando Pereira 0001, João Ascenso
ISM2
2015 Epipolar plane image based rendering for 3D video coding
abstract
In current 3D video coding solutions, such as the 3D-HEVC standard, depth data is instrumental to have a continuum of views synthesized at the decoder based on a limited set of coded views. In order view synthesis may be performed at the decoder, depth data is currently directly acquired or estimated at the encoder based on very few neighboring views and transmitted to the decoder after appropriate compression. At the decoder, further views then those decoded are synthesized using again very few neighboring decoded views, thus using a local synthesis approach. A promising alternative synthesis approach may consider not a few but rather all the views available at the decoder, thus offering a scene global approach to synthesis. One way to implement this approach involves cutting the views cube along the viewpoint direction, creating the so-called epipolar plane images (EPI) which provide a rather compact representation of the scene. In this context, this paper proposes an EPI based view rendering framework for 3D video coding solution and identifies the major benefits of such framework, notably in comparison with the traditional local synthesis approach.
Catarina Brites, João Ascenso, Fernando Pereira 0001
MMSP3
2015 Report on the evaluation of current and future image compression technologies
abstract
This document reports the conclusions, comments and recommendations resulting from the final panel discussion which happened at the Session on Evaluation of Current and Future Image Compression Technologies held at the Picture Coding Symposium, Cairns, Australia, in June 2015.
Fernando Pereira 0001, Ralf Schaefer, Touradj Ebrahimi, Jörn Ostermann, Edward J. Delp
PCS1
2015 Improving enhancement layer merge mode for HEVC scalable extension
abstract
In the HEVC scalable extension (SHVC), the merge mode prediction plays an important role due to its high selection probability and its low associated bitrate. In the SHVC enhancement layer (EL) merge mode prediction, the motion information is selected from the merge candidates, in this case the spatial, temporal, and inter-layer candidates. Therefore, the merge mode prediction is usually inefficient when the motion vector field (MVF) correlation between the merge candidates and the current EL block is low, especially in video sequences containing high motion activity. To address this problem, this paper proposes an improved EL merge mode prediction solution that adaptively refines the motion vector candidates by using both base layer (BL) and EL decoded information and after linearly combines the motion compensated samples with the BL reconstructed samples to achieve a better merge mode prediction quality. Experiments conducted for a rich set of test conditions have shown that significant compression efficiency gains can be achieved with the proposed improved scalable coding solution, notably up to 4.57% in BD-rate savings regarding the standard SHVC quality scalable codec.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
PCS3
2015 Optimal layered representation for adaptive interactive multiview video streaming
Ana De Abreu, Laura Toni, Nikolaos Thomos, Thomas Maugey, Fernando Pereira 0001, Pascal Frossard
J. Vis. Commun. Image Represent.5
2015 Multiview side information creation for efficient Wyner-Ziv video coding: Classifying and reviewing
Catarina Brites, Fernando Pereira 0001
Signal Process. Image Commun.2
2015 Distributed video coding: Assessing the HEVC upgrade
Catarina Brites, Fernando Pereira 0001
Signal Process. Image Commun.2
2015 HEVC backward compatible scalability: A low encoding complexity distributed video coding based approach
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
Signal Process. Image Commun.3
2014 Correlation noise modeling for multiview transform domain Wyner-Ziv video coding
abstract
Multiview Wyner-Ziv (MV-WZ) video coding rate-distortion (RD) performance is highly influenced by the adopted correlation noise model (CNM). In the related literature, the statistics of the correlation noise between the original frame and the side information (SI), typically resulting from the fusion of temporally and inter-view created SIs, is modelled by a Laplacian distribution. In most cases, the Laplacian CNM parameter is estimated using an offline approach, assuming that either the SI is available at the encoder or the originals are available at the decoder which is not realistic. In this context, this paper proposes the first practical, online CNM solution for a multiview transform domain WZ (MV-TDWZ) video codec. The online estimation of the Laplacian CNM parameter is performed at the decoder based on metrics exploring both the temporal and inter-view correlations with two levels of granularity, notably transform band and transform coefficient. The results obtained show that better RD performance is achieved for the finest granularity level since the inter-view, temporal and spatial correlations are exploited with the highest adaptation.
Catarina Brites, Fernando Pereira 0001
ICIP2
2014 Local feature selection for efficient binary descriptor coding
abstract
In a visual sensor network, a large number of camera nodes are able to acquire and process image data locally, collaborate with other camera nodes and provide a description about the captured events. Typically, camera nodes have severe constraints in terms of energy, bandwidth resources and processing capabilities. Considering these unique characteristics, coding and transmission of the pixel-level representation of the visual scene must be avoided, due to the energy resources required. A promising approach is to extract at the camera nodes, compact visual features that are coded to meet the bandwidth and power requirements of the underlying network and devices. Since the total number of features extracted from an image may be rather significant, this paper proposes a novel method to select the most relevant features before the actual coding process. The solution relies on a score that estimates the accuracy of each local feature. Then, local features are ranked and only the most relevant features are coded and transmitted. The selected features must maximize the efficiency of the image analysis task but also minimize the required computational and transmission resources. Experimental results show that higher efficiency is achieved when compared to the previous state-of-the-art.
Pedro Monteiro, João Ascenso, Fernando Pereira 0001
ICIP3
2014 Correlation modeling for a distributed scalable video codec based on the HEVC standard
abstract
The growing heterogeneity of networks, devices and consumption conditions asks for flexible and adaptive video coding solutions. The compression power of the HEVC standard and the benefits of the distributed video coding paradigm allow designing novel scalable coding solutions with improved error robustness and low encoding complexity while still achieving competitive compression efficiency. In this context, this paper proposes a novel scalable video coding scheme using a HEVC Intra compliant base layer and a distributed coding approach in the enhancement layers (EL). This design inherits the HEVC compression efficiency while providing low encoding complexity at the enhancement layers. The temporal correlation is exploited at the decoder to create the EL side information (SI) residue, an estimation of the original residue. The EL encoder sends only the data that cannot be inferred at the decoder, thus exploiting the correlation between the original and SI residues; however, this correlation must be characterized with an accurate correlation model to obtain coding efficiency improvements. Therefore, this paper proposes a correlation modeling solution to be used at both encoder and decoder, without requiring a feedback channel. Experiments results confirm that the proposed scalable coding scheme has lower encoding complexity and provides BD-Rate savings up to 3.43% in comparison with the HEVC Intra scalable extension under development.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
MMSP3
2014 H.264/AVC backward compatible bit-depth scalable video coding
abstract
As high dynamic range video is gaining popularity, video coding solutions able to efficiently provide both low and high dynamic range video, notably with a single bitstream, are increasingly important. While simulcasting can provide both dynamic range videos at the cost of some compression efficiency penalty, bit-depth scalable video coding can provide a better trade-off between compression efficiency, adaptation flexibility and computational complexity. Considering the widespread use of H.264/AVC video, this paper proposes a H.264/AVC backward compatible bit-depth scalable video coding solution offering a low dynamic range base layer and two high dynamic range enhancement layers with different qualities, at low complexity. Experimental results show that the proposed solution has an acceptable rate-distortion performance penalty regarding the HDR H.264/AVC single-layer coding solution.
Vasco Nascimento, João Ascenso, Fernando Pereira 0001
MMSP3
2014 Multiview video representations for quality-scalable navigation
abstract
Interactive multiview video (IMV) applications offer to users the freedom of selecting their preferred viewpoint. Usually, in these systems texture and depth maps of captured views are available at the user side, as they permit the rendering of intermediate virtual views. However, the virtual views' quality depends on the distance to the available views used as references and on their quality, which is generally constrained by the heterogeneous capabilities of the users. In this context, this work proposes an IMV scalable system, where views are optimally organized in layers, each one offering an incremental improvement in the interactive navigation quality. We propose a distortion model for the rendered virtual views and an algorithm that selects the optimal views' subset per layer. Simulation results show the efficiency of the proposed distortion model, and that the careful choice of reference cameras permits to have a graceful quality degradation for clients with limited capabilities.
Ana De Abreu, Laura Toni, Thomas Maugey, Nikolaos Thomos, Pascal Frossard, Fernando Pereira 0001
VCIP6
2014 Statistical reconstruction for predictive video coding
abstract
Substantial rate-distortion (RD) gains have been achieved in video coding standards by increasing the encoder complexity while maintaining the decoder complexity the lowest possible. On the other hand, the alternative distributed video coding (DVC) approach proposes to exploit the video redundancy mostly at the decoder side, keeping the encoder as simple as possible. One of the most characteristic DVC tools is the statistical reconstruction of the DCT coefficients, which plays a similar role to the inverse scalar quantization (ISQ) in predictive codecs. The main objective of this paper is to propose a statistical reconstruction approach for predictive coding (notably the H.264/AVC standard) as a substitute to ISQ, thus creating a coding architecture with a mix of predictive and distributed coding tools. Experimental results show that the proposed statistical reconstruction solution allows achieving Bjontegaard bitrate savings up to 2.4% regarding the ISQ based H.264/AVC High profile codec.
Catarina Brites, Vitor Gomes 0003, João Ascenso, Fernando Pereira 0001
VCIP4
2014 Optimal reconstruction for a HEVC backward compatible distributed scalable video codec
abstract
In a landscape of heterogeneous networks, terminals and usage environments, a low encoding complexity scalable video coding engine is required for many emerging applications such as wireless video surveillance, visual sensor networks and remote space transmission. To fulfil this need, a distributed scalable video coding (DSVC) framework has been developed combining the predictive and distributed coding paradigms. The DSVC framework provides High Efficiency Video Coding (HEVC) backward compatibility at the base layer (BL) and adopts a distributed coding approach for the enhancement layers (ELs). In DSVC, the decoder reconstruction plays a key role as it strongly impacts the EL decoded frame quality and, thus, the final DSVC compression efficiency. In this context, this paper proposes a statistically inspired decoder reconstruction solution for the EL coded information based on the so-called side information residue and its correlation with the encoder EL residue. Experimental results confirm that significant RD performance gains can be achieved with the proposed statistical reconstruction technique, notably up to 11.28% BD-Rate savings regarding the most relevant benchmark.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
VCIP3
2014 Perceptually driven video error protection using a distributed source coding approach
André Seixas Dias, Catarina Brites, João Ascenso, Fernando Pereira 0001
Signal Process. Image Commun.4
2014 Epipolar Geometry-Based Side Information Creation for Multiview Wyner-Ziv Video Coding
abstract
The side information (SI) quality significantly influences the rate-distortion (RD) performance of both monoview and multiview Wyner-Ziv (WZ) video coding. Efficient multiview WZ video coding schemes typically exploit both temporal and inter-view correlation during the SI creation process targeting to achieve high SI quality and, thus, high RD performance. In this context, the overall SI creation process typically involves fusing a temporally created SI with an inter-view created SI. Consequently, the final SI quality does not only depend on the temporal and inter-view SI creation techniques but also on the fusion process. Thus, the main objective of this paper is to propose an efficient SI creation solution for multiview WZ video coding by simultaneously tackling the two main issues aforementioned with the following technical novelty: 1) a disparity-based view synthesis technique accounting for the scene geometry to create high quality inter-view SI and 2) a time-view driven fusion technique efficiently selecting, on a pixel basis, the temporal or inter-view SI depending on which one is estimated to be closer the original WZ data. Experimental results show considerable RD performance improvements, notably gains up to 2 dB or more, regarding the best performing state-of-the-art multiview WZ video coding solutions available.
Catarina Brites, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.2
2013 Fast MVC prediction structure selection for interactive multiview video streaming
abstract
Multiview Video Coding (MVC) has been developed to efficiently compress a set of camera views by exploiting the spatial, temporal and interview correlations among images of the same scene. However, the resulting compressed data has a lot of prediction coding dependencies, which may not suit interactive multiview video streaming (IMVS) systems, where only one view is requested at a time by the end-user. This paper proposes a fast selection mechanism for effective interview prediction structure (PS) in IMVS while minimizing the point-to-point transmission rate, given some storage and visual distortion constraints, and a user interactive behavior model. Simulation results show that our novel fast MVC PS selection algorithm has high efficiency with low computational complexity that is reduced by more than 40% in comparison to the exhaustive searching benchmark.
Ana De Abreu, Pascal Frossard, Fernando Pereira 0001
PCS3
2013 Improving scalable video coding performance with decoder side information
abstract
In a heterogeneous landscape of networks, devices and consumption environments, scalability is one of the most important video coding features. To achieve higher scalable video compression efficiency, this paper proposes a novel scalable video coding framework based on predictive video coding but also exploiting some additional decoder side information. The side information is estimated at both encoder and decoder using a motion compensated temporal interpolation technique, commonly used in distributed video coding solutions. To improve the B-slices compression efficiency, the side information independently created at each coding layer, notably the base and enhancement layers, is inserted in the corresponding layer decoded picture buffer to be exploited as an additional reference frame in the scalable predictive coding process. Experimental results have shown significant compression efficiency gains, notably up to around 3.5% in bitrate savings regarding the state-of-the-art SVC standard.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
PCS3
2013 Side information creation for efficient Wyner-Ziv video coding: Classifying and reviewing
Catarina Brites, João Ascenso, Fernando Pereira 0001
Signal Process. Image Commun.3
2013 Optimized MVC Prediction Structures for Interactive Multiview Video Streaming
abstract
The Multiview Video Coding (MVC) standard efficiently compresses multiview video by considering spatial, temporal and interview correlations. This letter studies the impact of the MVC interview prediction structure on both the transmission and the overall coding rates for an interactive multiview video streaming system, considering both unicast and multicast scenarios, with the user interactive behavior represented by some view-popularity model. We propose a method to identify the optimal prediction structure minimizing the visual distortion, given some storage and link capacities constraints. Simulation results confirm that the optimal prediction structure results from a non-trivial tradeoff between the system constraints, the transmission model and the views' popularity.
Ana De Abreu, Pascal Frossard, Fernando Pereira 0001
IEEE Signal Process. Lett.3
2012 Improving predictive video coding performance with decoder side information
abstract
Although major achievements have been reached in terms of video compression efficiency, additional gains are still needed to satisfy current and emerging applications needs. This trend justifies the continuous efforts to go beyond the compression capabilities of the state-of-the-art H.264/AVC standard. This paper proposes to further bridge two video coding approaches, the popular predictive and the emerging distributed coding paradigms, by proposing a novel bidirectional (B) coding process to be integrated in the H.264/AVC codec, inspired by the side information creation module, central in distributed video coding. The novel B-slice coding process builds on a new reference frame, the side information frame, which is both created at the encoder and decoder, as the SI creation method does not need the original to be available. The experimental results for the novel coding solution show average bitrate savings of 5.89% both for the low and high bitrate regions.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
ICIP3
2012 Learning based decoding approach for improved Wyner-Ziv video coding
abstract
Wyner-Ziv (WZ) video coding compression efficiency depends critically both on the side information (SI) quality and the correlation noise model (CNM) accuracy. In this context, this paper proposes a learning based decoding approach for transform domain WZ video coding, notably in the context of the following techniques: i) fractional-pixel motion field learning to define the relevance of the SI block candidates, and ii) CNM parameter learning. Experimental results show the proposed learning approach brings consistent RD performance improvements, with coding gains up to 3.9 dB regarding the state-of-the-art DISCOVER WZ video codec for a GOP size of 8.
Catarina Brites, João Ascenso, Fernando Pereira 0001
PCS3
2012 Fast rate distortion optimization for the emerging HEVC standard
abstract
The under development High Efficiency Video Coding (HEVC) standard employs several powerful coding tools to obtain improved compression efficiency regarding the state-of-the-art H.264/AVC standard. To efficiently exploit the temporal and spatial redundancies, HEVC adopts a very flexible quadtree coding structure, allowing the encoder to use a block partition better matching the image features. The best combination of HEVC block partitioning and coding modes is found by means of a rate distortion optimization process. During this minimization, the encoder tests all the possible coding modes and block partitions and keeps those providing the smallest RD cost. Due to the large number of available modes and partitions, the minimization involves extremely high complexity, which may not be suitable for real-time application scenarios. To reduce the complexity, two novel fast RDO techniques are proposed in this paper: Top Skip and Early Termination. The experimental evaluation reveals that the proposed techniques significantly reduce the encoding time, notably up to 45% regarding the high complexity HEVC reference software codec, with a negligible quality loss never larger than 0.1 dB.
Michele Belotti Cassa, Matteo Naccari, Fernando Pereira 0001
PCS3
2012 Adaptive bilateral filter for improved in-loop filtering in the emerging high efficiency video coding standard
abstract
To face the still growing video compression needs, ITU and MPEG have jointly started a standardization project called High Efficiency Video Coding (HEVC) which aims to improve the compression efficiency of the state-of-the-art H.264/AVC standard for high and ultra high definition video. The HEVC codec still relies on the usual motion compensated, predictive, block based transform coding architecture with block sizes higher than 8×8 being now considered. Furthermore, the codec is also equipped with a Wiener in-loop filter which, together with the H.264/AVC deblocking filter, further reduces the distortion between the original and decoded frames introduced by lossy coding. While this filter allows indeed reducing the sum of square differences between the original and reconstructed frames, it is not specifically designed to reduce the ringing artifacts resulting from the use of large transform block sizes. In this context, this paper proposes to combine an adaptive bilateral filter together with the HEVC Wiener filter. The proposed combined filter allows an average bitrate reduction of about 7% regarding the HEVC codec without the Wiener filter and a 1.5% reduction regarding the HEVC codec with only the Wiener filter, always for the same quality. Moreover, the combined filter also reduces the ringing artifacts according to an objective metric specifically designed to quantify this type of coding artifact.
Matteo Naccari, Fernando Pereira 0001
PCS2
2012 Quadratic modeling rate control in the emerging HEVC standard
abstract
The High Efficiency Video Coding (HEVC) standardization project mainly aims at improving the compression efficiency beyond the state-of-the-art H.264/AVC standard for high and ultra high definition video contents. The HEVC codec reference software, still under development, encodes video content using a fixed quantization step, thus leading to a variable bitrate stream which may not be suitable for the many multimedia applications where a constant bandwidth is required. Therefore, a rate control algorithm must be integrated in the HEVC codec taking into account that the novel coding tools further reduce the bitrate associated to texture data while increase the bitrate associated to the auxiliary information, e.g. motion vectors and coding modes. In this context, this paper studies the integration of a quadratic modeling rate control algorithm in the HEVC codec. Experimental results reveal that the rate control enabled HEVC codec provides a lower rate-distortion performance regarding a fixed quantization step HEVC codec, notably spending up to 20.12% more bitrate for the same objective quality. However, the integrated rate control algorithm offers lower bitrate fluctuations along the video contents than the HEVC codec with fixed quantization step.
Matteo Naccari, Fernando Pereira 0001
PCS2
2011 Integrating a spatial just noticeable distortion model in the under development HEVC codec
abstract
Although great achievements have been obtained in the past, recent developments towards ultra high definition video have been asking for further compression efficiency gains regarding the H.264/AVC state-of-the-art. As an answer to these needs, ITU-T and MPEG started a new standardization project called High Efficiency Video Coding. This paper extends and integrates a perceptual visual model in the High Efficiency Video Coding experimental codec to improve its RD performance. The designed extensions are related to the new Integer Discrete Cosine Transform block sizes as well as the new Mode Dependent Directional Transform. Moreover, an estimation mechanism for the perceptual model at the decoder side is also integrated to avoid the rate burden of sending to the decoder the model parameters. The RD performance is assessed using a methodology taking into account the variability of subjective scores which is reflected into a nonlinear sensitivity of the subjective quality metric. The RD performance obtained for the perceptually driven codec shows a bitrate reduction of up to 43% regarding the non-perceptual video codec.
Matteo Naccari, Fernando Pereira 0001
ICASSP2
2011 A denoising approach for iterative side information creation in distributed video coding
abstract
In distributed video coding, motion estimation is typically performed at the decoder to generate the side information, increasing the decoder complexity while providing low complexity encoding in comparison with predictive video coding. Motion estimation can be performed once to create the side information or several times to refine the side information quality along the decoding process. In this paper, motion estimation is performed at the decoder side to generate multiple side information hypotheses which are adaptively and dynamically combined, whenever additional decoded information is available. The proposed iterative side information creation algorithm is inspired in video denoising filters and requires some statistics of the virtual channel between each side information hypothesis and the original data. With the proposed denoising algorithm for side information creation, a RD performance gain up to 1.2 dB is obtained for the same bitrate.
João Ascenso, Catarina Brites, Fernando Pereira 0001
ICIP3
2011 A new fast motion estimation and mode decision algorithm for H.264 depth maps encoding in free viewpoint TV
abstract
In this paper, we consider a scenario where 3D scenes are modeled through a View+Depth representation. This representation is to be used at the rendering side to generate synthetic views for free viewpoint video. The encoding of both type of data (view and depth) is carried out using two H.264/AVC encoders. In this scenario we address the reduction of the encoding complexity of depth data. Firstly, an analysis of the Mode Decision and Motion Estimation processes has been conducted for both view and depth sequences, in order to capture the correlation between them. Taking advantage of this correlation, we propose a fast mode decision and motion estimation algorithm for the depth encoding. Results show that the proposed algorithm reduces the computational burden with a negligible loss in terms of quality of the rendered synthetic views. Quality measurements have been conducted using the Video Quality Metric.
Gianluca Cernigliaro, Matteo Naccari, Fernando Jaureguizar, Julián Cabrera, Fernando Pereira 0001, Narciso García
ICIP5
2011 Low complexity deblocking filter perceptual optimization for the HEVC codec
abstract
The compression efficiency of the state-of-art H.264/AVC video coding standard must be improved to accommodate the compression needs of high definition videos. To this end, ITU and MPEG started a new standardization project called High Efficiency Video Coding. The video codec under development still relies on trans form domain quantization and includes the same in-loop deblocking filter adopted in the H.264/AVC standard to reduce quantization blocking artifacts. This deblocking filter provides two offsets to vary the amount of filtering for each image area. This paper proposes a perceptual optimization of these offsets based on a quality metric able to quantify the blocking artifacts impact on the perceived video quality. The proposed optimization involves low computational complexity and provides quality improvements with respect to a non-perceptually optimized H.264/AVC deblocking filter. Moreover, the proposed optimization allows up to 92% of complexity reduction regarding a brute force perceptual optimization which exhaustively tests all the possible offsets values.
Matteo Naccari, Catarina Brites, João Ascenso, Fernando Pereira 0001
ICIP4
2011 Binary tree decomposition depth coding for 3D video applications
abstract
Recent advances in three dimensional display technologies and the growing efforts put in the production of three dimensional videos are intensely demanding for new, more efficient three dimensional content representation formats. One promising for mat is the so-called multiview plus depth format where multiple views of the observed scene are represented together with a per pixel depth map. Naturally, this format requires to efficiently code not only the video views but also the depth data. In this context, this paper proposes a novel depth map codec encoding the depth data by means of a binary tree triangular decomposition and reconstructing the depth map values by means of a tri angle based planar approximation. For depth maps related to typical three dimensional video contents, the proposed depth map codec outperforms both the H.264/AVC standard with all intra coding modes enabled and the JPEG standard. In particular, for the same objective reconstruction quality, the proposed codec allows an average bitrate reduction of 60% and 90% regarding H.264/AVC Intra and JPEG coding, respectively.
Gonçalo Carmo, Matteo Naccari, Fernando Pereira 0001
ICME3
2011 Augmented LDPC graph for distributed video coding with multiple side information
abstract
The advances made in channel-capacity codes, such as turbo codes and low-density parity-check (LDPC) codes, have played a major role in the emerging distributed source coding paradigm. LDPC codes can be easily adapted to new source coding strategies due to their natural representation as bipartite graphs and the use of quasi-optimal decoding algorithms, such as belief propagation. This paper tackles a relevant scenario in distributed video coding: lossy source coding when multiple side information (SI) hypotheses are available at the decoder, each one correlated with the source according to different correlation noise channels. Thus, it is proposed to exploit multiple SI hypotheses through an efficient joint decoding technique with multiple LDPC syndrome decoders that exchange information to obtain coding efficiency improvements. At the decoder side, the multiple SI hypotheses are created with motion compensated frame interpolation and fused together in a novel iterative LDPC based Slepian-Wolf decoding algorithm. With the creation of multiple SI hypotheses and the proposed decoding algorithm, bitrate savings up to 8.0% are obtained for similar decoded quality.
João Ascenso, Catarina Brites, Fernando Pereira 0001
MMSP3
2011 Musical slideshow: boosting user experience in photo presentation
Bruno Tomás, Fernando Pereira 0001
Multim. Tools Appl.2
2011 Low delay distributed video coding with refined side information
António Tomé, Fernando Pereira 0001
Signal Process. Image Commun.2
2011 An Efficient Encoder Rate Control Solution for Transform Domain Wyner-Ziv Video Coding
abstract
Most Wyner-Ziv (WZ) video coding solutions in the literature use a feedback channel (FC) based decoder rate control (DRC) strategy to adjust the bitrate to correct the side information (SI) errors. More recently, some encoder rate control (ERC) strategies have been proposed to address application scenarios where a FC is not available. The ERC based WZ video coding RD performance depends not only on the (encoder) parity rate estimator (PRE) accuracy but also on the decoder “intelligence” in dealing with the residual errors due to parity rate underestimation. In this context, the main objective of this paper is to propose a more efficient and powerful ERC solution for transform domain WZ (TDWZ) video coding by simultaneously tackling the two issues aforementioned with the following technical novelty: 1) integration in an ERC context of Gray mapping for the quantized DCT coefficients to enhance the correlation between WZ and SI data; 2) more accurate PRE to better estimate the needed parity rate to avoid undesired parity rate underestimations and overestimations; 3) novel soft reconstruction function to reduce the impact of the residual bitplane errors in the decoded WZ frame quality; and 4) weighted overlapped block motion compensation technique to refine the SI used in an iterative WZ decoding framework with the correlation noise model parameters dynamically updated. Experimental results show a considerable RD performance improvement with a reduction of up to about 2 dB of the gap between the ERC and DRC based approaches in TDWZ video coding solutions, thus making this ERC based WZ codec the most efficient available and competitive regarding DRC based WZ video coding solutions.
Catarina Brites, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.2
2011 Advanced H.264/AVC-Based Perceptual Video Coding: Architecture, Tools, and Assessment
abstract
The characteristics of the human visual system may be further exploited in lossy video coding to improve the video compression efficiency beyond the state-of-the-art H.264/AVC standard. Although the literature is rich in solutions to model the human visual system characteristics, the performance and real benefits brought by these models have not been fully integrated and assessed yet. Moreover, the rate-distortion (RD) performance is usually measured by means of methodologies that do not account for the implicit variability of the observers when rating the video quality. In this context, the novelty brought by this paper is threefold: first, it proposes novel perceptual video coding tools, notably decoder side just noticeable distortion (JND) model estimation to perceptually allocate the available rate with the finest level of granularity while avoiding the extra rate associated to coding the varying quantization steps. Second, it proposes an integrated, powerful H.264/AVC-based perceptual video coding architecture embedding a state-of-the-art JND model based on spatio-temporal human visual system masking mechanisms; this model is exploited for both the aforementioned rate allocation as well as to perceptually weight the distortion used in the motion estimation and RD optimization. Finally, it proposes a relative assessment methodology to measure the RD performance of a perceptual video codec (PVC) with respect to another codec taken as reference. The methodology considers the implicit observers variability when rating video quality which leads to a nonlinear sensitivity of the objective metrics used for quality assessment. The obtained RD performance, measured according to this methodology, shows an average bitrate reduction of up to 30% when the proposed PVC is compared with the H.264/AVC High profile at the same objective quality level. Moreover, the proposed perceptual codec outperforms an alternative perceptual codec recently published in the literature.
Matteo Naccari, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.2
2010 Probability updating for decoder and encoder rate control turbo based Wyner-Ziv video coding
abstract
In Wyner-Ziv video coding (WZVC), powerful error correcting codes must be used to achieve high compression efficiency; turbo codes are the most commonly used error correcting codes in WZVC. To improve the turbo coding performance in the context of WZVC, this paper proposes a probability updating technique (PUT) acting as an outer loop of the common turbo decoding operation. Whenever a turbo decoded bitplane is not error-free, the proposed technique attempts to correct bitplane errors by updating the correlation noise probabilities for the most likely in error bits, followed by turbo redecoding. The new tool is evaluated both in the context of encoder rate control (ERC) and decoder rate control (DRC) turbo based WZVC scenarios with average overall PSNR gains up to about 0.5 dB in ERC and average WZ rate savings up to about 6% in DRC.
Catarina Brites, Fernando Pereira 0001
ICIP2
2010 Decoder side just noticeable distortion model estimation for efficient H.264/AVC based perceptual video coding
abstract
This paper proposes a decoder side estimation for a just noticeable distortion perceptual model integrated in the H.264/AVC codec to improve its RD performance. The proposed estimation avoids having to send the perceptual weights used during the quantization of each frequency coefficient. Experiments with high definition videos show that the novel perceptual video codec is able to reduce the bitrate up to about 34% at the cost of, at most, 1.2% objective multi scale structural similarity quality loss. Moreover, the RD performance with the estimation is comparable to the performance achieved if the original perceptual model is used at the decoder without considering the associated rate.
Matteo Naccari, Fernando Pereira 0001
ICIP2
2010 A flexible side information generation framework for distributed video coding
João Ascenso, Catarina Brites, Fernando Pereira 0001
Multim. Tools Appl.3
2010 Automatic MPEG-4 sprite coding - Comparison of integrated object segmentation algorithms
Alexander Glantz, Andreas Krutz, Thomas Sikora, Paulo J. L. Nunes, Fernando Pereira 0001
Multim. Tools Appl.5
2009 Automatic and adaptive network-aware macroblock intra refresh for error-resilient H.264/AVC video coding
abstract
In this paper, an automatic and adaptive network-aware macroblock Intra coding refresh method is proposed. It adaptively selects the amount of gracefully forced Intra macroblocks and the amount of cyclic Intra refresh (CIR) macroblocks based on the actual network error conditions, in terms of packet loss rate, an the target encoding bit rate. With the proposed method, the error robustness of H.264/AVC bitstreams can be significantly increased by efficiently taking into account the actual rate-distortion impact of Intra coding macroblock mode decisions, while simultaneously guaranteeing that errors do not propagate endlessly by selecting an adequate amount of CIR macroblocks per frame according to the network packet loss rate and the encoding target bit rate.
Paulo J. L. Nunes, Luís Ducla Soares, Fernando Pereira 0001
ICIP3
2009 Low complexity intra mode selection for efficient distributed video coding
abstract
Motion compensated frame interpolation (MCFI) is one of the most efficient solutions to generate side information (SI) in the context of distributed video coding. However, it creates SI with rather significant motion compensated errors for some frame regions while rather small for some other regions depending on the video content. In this paper, a low complexity intra mode selection algorithm is proposed to select the most dasiacriticalpsila blocks in the WZ frame and help the decoder with some reliable data for those blocks. For each block, the novel coding mode selection algorithm estimates the encoding rate for the intra based and WZ coding modes and determines the best coding mode while maintaining a low encoder complexity. The proposed solution is evaluated in terms of rate-distortion performance with improvements up to 1.2 dB regarding a WZ coding mode only solution.
João Ascenso, Fernando Pereira 0001
ICME2
2009 Video compression: Discussing the next steps
abstract
Video compression has been intensively evolving for more than 20 years. Until recently, around 50% compression gains every 5 years were obtained, resulting in the current set of video coding standards. More recently, after the development of the very successful H.264/AVC standard, compression gains have been short and more difficult to reach than usual. In this context, this talk will discuss the future of video compression considering the emerging industry needs, notably in terms of promising technological novelties, and recent standardization initiatives targeting 3D video and further video compression efficiency.
Fernando Pereira 0001
ICME1
2009 Distributed video coding: Basics, main solutions and trends
abstract
After the great success of the predictive video coding approach, which led to a number of largely deployed MPEG and ITU-T standards, the video coding research community has been working on a new video coding paradigm, so-called distributed video coding (DVC), which is based on some information theory results from the 70s: the Slepian-Wolf and the Wyner-Ziv theorems. The first practical solutions have emerged around 2002 at the Stanford University and the University of California, Berkeley. This talk will address the basics, main solutions and trends on distributed video coding with especial emphasis on the Stanford DVC codec which has deserved a larger research investment. The rate-distortion (RD) performance of a state-of-the-art Stanford based DVC codec will be presented and benchmarked by the relevant alternative standard based video coding solutions. Finally, some trends on the DVC research will be discussed.
Fernando Pereira 0001
ICME1
2009 Distributed Video Coding with multiple side information
abstract
Distributed Video Coding (DVC) is a new video coding paradigm which mainly exploits the source statistics at the decoder based on the availability of some decoder side information. The quality of the side information has a major impact on the DVC rate-distortion (RD) performance in the same way the quality of the predictions had a major impact in predictive video coding. In this paper, a DVC solution exploiting multiple side information is proposed; the multiple side information is generated by frame interpolation and frame extrapolation targeting to improve the side information of a single estimation mode. Compared with the best available single side information solutions, the proposed DVC solution with multiple side information robustly improves the RD performance for the set of test sequences.
Xin Huang 0004, Catarina Brites, João Ascenso, Fernando Pereira 0001, Søren Forchhammer
PCS4
2009 Adaptive deblocking filter for transform domain Wyner-Ziv video coding
abstract
Wyner–Ziv (WZ) video coding is a particular case of distributed video coding, the recent video coding paradigm based on the Slepian–Wolf and Wyner–Ziv theorems that exploits the source correlation at the decoder and not at the encoder as in predictive video coding. Although many improvements have been done over the last years, the performance of the state-of-the-art WZ video codecs still did not reach the performance of state-of-the-art predictive video codecs, especially for high and complex motion video content. This is also true in terms of subjective image quality mainly because of a considerable amount of blocking artefacts present in the decoded WZ video frames. This paper proposes an adaptive deblocking filter to improve both the subjective and objective qualities of the WZ frames in a transform domain WZ video codec. The proposed filter is an adaptation of the advanced deblocking filter defined in the H.264/AVC (advanced video coding) standard to a WZ video codec. The results obtained confirm the subjective quality improvement and objective quality gains that can go up to 0.63 dB in the overall for sequences with high motion content when large group of pictures are used.
Catarina Brites, João Ascenso, Fernando Pereira 0001
IET Image Process.4
2009 Refining Side Information for Improved Transform Domain Wyner-Ziv Video Coding
abstract
Wyner-Ziv (WZ) video coding is a particular case of distributed video coding, which is a recent video coding paradigm based on the Slepian-Wolf and WZ theorems. Contrary to available prediction-based standard video codecs, WZ video coding exploits the source statistics at the decoder, allowing the development of simpler encoders. Until now, WZ video coding did not reach the compression efficiency performance of conventional video coding solutions, mainly due to the poor quality of the side information, which is an estimate of the original frame created at the decoder in the most popular WZ video codecs. In this context, this paper proposes a novel side information refinement (SIR) algorithm for a transform domain WZ video codec based on a learning approach where the side information is successively improved as the decoding proceeds. The results show significant and consistent performance improvements regarding state-of-the-art WZ and standard video codecs, especially under critical conditions such as high motion content and long group of pictures sizes.
Catarina Brites, João Ascenso, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.4
2009 Joint Rate Control Algorithm for Low-Delay MPEG-4 Object-Based Video Encoding
abstract
This paper proposes an improved rate control algorithm for jointly encoding multiple arbitrarily shaped video objects in the context of low-delay MPEG-4 compliant video coding. The algorithm provides adequate mechanisms for dealing with deviations between the ideal and the actual behavior of video scene encoders, notably: 1) compensation mechanisms (e.g., rate control decisions) that are able to track these deviations and compensate them to allow a stable and efficient operation of the encoder, and 2) adaptation mechanisms (e.g., estimation of model parameters) that are able to instantaneously represent the actual behavior of the encoder and its rate controller. The proposed solution efficiently allocates the available resources, i.e., target bit rate and bitstream buffer space, aiming at maximizing the average scene quality and minimizing quality fluctuations along time and among the various video objects. The results show that this solution outperforms the usual reference solutions, notably those specified in the rate control informative annex of the MPEG-4 visual standard.
Paulo J. L. Nunes, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.2
2008 Design and performance of a novel low-density parity-check code for distributed video coding
abstract
Low-density parity-check (LDPC) codes are nowadays one of the hottest topics in coding theory, notably due to their advantages in terms of bit error rate performance and low complexity. In order to exploit the potential of the Wyner-Ziv coding paradigm, practical distributed video coding (DVC) schemes should use powerful error correcting codes with near-capacity performance. In this paper, new ways to design LDPC codes for the DVC paradigm are proposed and studied. The new LDPC solutions rely on merging parity-check nodes, which corresponds to reduce the number of rows in the parity-check matrix. This allows to change gracefully the compression ratio of the source (DCT coefficient bitplane) according to the correlation between the original and the side information. The proposed LDPC codes reach a good performance for a wide range of source correlations and achieve a better RD performance when compared to the popular turbo codes.
João Ascenso, Catarina Brites, Fernando Pereira 0001
ICIP3
2008 No-reference modeling of the channel induced distortion at the decoder for H.264/AVC video coding
abstract
This paper proposes a model for estimating, at the decoder side, the distortion induced by the transmission over an error-prone channel, when error-free reconstructed frames are not available as a reference. The proposed estimation model considers explicitly the temporal error concealment algorithm adopted at the decoder. In the evaluation of the induced distortion, we model the effects of the absence of motion vectors and prediction residuals in the decoding process. In addition, we take into account the error propagation along successive frames. Experimental results conducted over real video sequences coded with the state-of-art H.264/AVC video coding standard validate the proposed model. In fact, the distortion estimated when no reference is available is strongly correlated both at the frame and group of pictures level with the actual distortion. This technique represents an effective no-reference video quality monitoring tool that can be embedded in any H.264/AVC compliant decoder.
Matteo Naccari, Marco Tagliasacchi, Fernando Pereira 0001, Stefano Tubaro
ICIP3
2008 Error resilient macroblock rate control for H.264/AVC video coding
abstract
In this paper, an error resilient rate control scheme for the H.264/AVC standard is proposed. This scheme differs from traditional rate control schemes in that macroblock mode decisions are not made only to minimize their rate-distortion cost, but also take into account that the bitstream will have to be transmitted through an error-prone network. Since channel errors will probably occur, error propagation due to predictive coding should be mitigated by adequate Intra coding refreshes. The proposed scheme works by comparing the rate-distortion cost of coding a macroblock in Intra and Inter modes: if the cost of Intra coding is only slightly larger than the cost of Inter coding, the coding mode is changed to Intra, thus reducing error propagation. Additionally, cyclic Intra refresh is also applied to guarantee that all macroblocks are eventually refreshed. The proposed scheme outperforms the H.264/AVC reference software, for typical test sequences, for error-free transmission and several packet loss rates.
Paulo J. L. Nunes, Luís Ducla Soares, Fernando Pereira 0001
ICIP3
2008 Wyner-Ziv video coding: A review of the early architectures and further developments
abstract
In 2002, the video coding community faced the emergence of a new video coding paradigm, the so-called Wyner-Ziv video coding, which was represented by two early solutions designed by the Stanford University and the University of California, Berkeley research teams. This paper intends to briefly review, and compare these two early Wyner-Ziv video coding solutions, notably from the functional point of view. Moreover, this paper reviews some important developments of the Stanford Wyner-Ziv coding architecture, which has become the most popular in the literature.
Fernando Pereira 0001, Catarina Brites, João Ascenso, Marco Tagliasacchi
ICME1
2008 Advanced side information creation techniques and framework for Wyner-Ziv video coding
João Ascenso, Fernando Pereira 0001
J. Vis. Commun. Image Represent.2
2008 Multimedia Retrieval and Delivery: Essential Metadata Challenges and Standards
abstract
Multimedia information retrieval (MIR) and delivery plays an important role in many application domains due to the increasing need to identify, filter, and manage growing amounts of data, notably multimedia information. To efficiently manage and exchange multimedia information, interoperability between coded data and metadata is required and standardization is central to achieving the necessary level of interoperability. In the context of this paper, the term retrieval refers to the process by which a user, human or machine, identifies the content it needs, and the term delivery refers to the adaptive transport and consumption of the identified content in a particular context or usage environment. Both the retrieval and delivery processes may require content and context metadata. This paper will argue that maximum quality of experience depends not only on the content itself (and thus content metadata) but also on the consumption conditions (thus context metadata). Additionally, the rights and protection conditions have become critically important in recent years, especially with the explosion of electronic music commerce and different ldquoshoppingrdquo conditions. This paper will review existing multimedia standards related to information retrieval and adaptive delivery of multimedia content, emphasizing the need for such standards, and will show how these standards can help the development, dissemination, and valorization of MIR research results. Moreover, it will also discuss limitations of the current standards and anticipate what future standardization activities are relevant and needed. Due to space limitations, the paper will mainly concentrate on MPEG standards although many other relevant standards are also reviewed and discussed.
Fernando Pereira 0001, Anthony Vetro, Thomas Sikora
Proc. IEEE1
2008 Evaluating a feedback channel based transform domain Wyner-Ziv video codec
Catarina Brites, João Ascenso, José Quintas Pedro, Fernando Pereira 0001
Signal Process. Image Commun.4
2008 Special issue on distributed video coding
Christine Guillemot, Fernando Pereira 0001
Signal Process. Image Commun.2
2008 Automatic creation and evaluation of MPEG-7 compliant summary descriptions for generic audiovisual content
Nuno Matos, Fernando Pereira 0001
Signal Process. Image Commun.2
2008 Distributed Video Coding: Selecting the most promising application scenarios
Fernando Pereira 0001, Christine Guillemot, Touradj Ebrahimi, Riccardo Leonardi, Sven Klomp
Signal Process. Image Commun.1
2008 Special Issue on Video Surveillance
abstract
The 14 regular papers and two brief papers in this special issue capture some of the state-of-the-art research on video surveillance issues, provide comprehensive overview of existing techniques, and propose novel solutions for important research problems.
Ishfaq Ahmad 0001, Zhihai He, Hong-Yuan Mark Liao, Fernando Pereira 0001, Ming-Ting Sun
IEEE Trans. Circuits Syst. Video Technol.4
2008 Correlation Noise Modeling for Efficient Pixel and Transform Domain Wyner-Ziv Video Coding
abstract
In recent years, practical Wyner-Ziv (WZ) video coding solutions have been proposed with promising results. Most of the solutions available in the literature model the correlation noise (CN) between the original frame and its estimation made at the decoder, which is the so-called side information (SI), by a given distribution whose relevant parameters are estimated using an offline process, assuming that the SI is available at the encoder or the originals are available at the decoder. The major goal of this paper is to propose a more realistic WZ video coding approach by performing online estimation of the CN model parameters at the decoder, for pixel and transform domain WZ video codecs. In this context, several new techniques are proposed based on metrics which explore the temporal correlation between frames with different levels of granularity. For pixel-domain WZ (PDWZ) video coding, three levels of granularity are proposed: frame, block, and pixel levels. For transform-domain WZ (TDWZ) video coding, DCT bands and coefficients are the two granularity levels proposed. The higher the estimation granularity is, the better the rate-distortion performance is since the deeper the adaptation of the decoding process is to the video statistical characteristics, which means that the pixel and coefficient levels are the best performing for PDWZ and TDWZ solutions, respectively.
Catarina Brites, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.2
2007 Adaptive Hash-Based Side Information Exploitation for Efficient Wyner-Ziv Video Coding
abstract
Wyner-Ziv video coding is a lossy source coding paradigm where the video statistics are exploited, partially or totally at the decoder. The side information represents a noisy version of the original frame and is generated at the decoder with time consuming motion estimation and compensation tools. This paper proposes a novel bidirectional hash motion estimation framework which enables the decoder to choose between past and/or future reference frames for frame interpolation. New features include the coding of DCT hash with zero-motion, combination of trajectory-based motion interpolation with hash-based motion estimation and adaptive selection of the DCT bands which are sent to the decoder in order to guide the motion estimation procedure. Gains up to 1.2 dB compared to previous motion interpolation approaches may be reached.
João Ascenso, Fernando Pereira 0001
ICIP (3)2
2007 Encoder Rate Control for Transform Domain Wyner-Ziv Video Coding
abstract
Wyner-Ziv (WZ) video coding -a particular case of distributed video coding (DVC) -is a new video coding paradigm based on two major Information Theory results: the Slepian-Wolf and Wyner-Ziv theorems. Many of the practical WZ video coding solutions available in the literature make use of a feedback channel (FC) to perform rate control at the decoder which implies there must be a FC available in the application scenario addressed. The FC-based DVC solutions also have implications in terms of delay and decoder complexity since several iterative decoding operations may be needed to decode the data to the target quality level. In this context, this paper proposes an encoder rate control (ERC) solution for the transform domain WZ coding architecture previously using a FC driven rate control. Although this is the first solution in the literature, promising results are achieved with the proposed ERC solution without significantly increase the encoder complexity.
Catarina Brites, Fernando Pereira 0001
ICIP (2)2
2007 Improved Feedback Compensation Mechanisms for Multiple Video Object Encoding Rate Control
abstract
This paper proposes new buffer and video object distortion feedback compensation mechanisms for efficiently dealing with deviations between the ideal and the actual behavior of video scene encoders when jointly encoding multiple arbitrarily shaped video objects in the context of compliant low-delay object-based MPEG-4 video coding. The proposed solution computes target buffer occupancies for each encoding time instant based on the amount and complexity of the video data to encode, and the bit allocation for each encoding time instant is feedback adjusted according to deviations relatively to this ideal behavior. Additionally, each video object bit allocation is also feedback adjusted based on the relative distortion of the various video objects in the scene. The proposed solution outperforms the non-normative MPEG-4 reference rate control algorithm for a wide range of bit rates and spatio-temporal resolutions, for typical test sequences.
Paulo J. L. Nunes, Fernando Pereira 0001
ICIP (3)2
2007 MPEG multimedia standards: evolution and future developments
abstract
Multimedia communications play a growing role in the every day's life of modern societies. Until recently, and except for broadcast television and radio, voice was still the sole communication mechanism. However, the diffusion of digital processing algorithms and hardware has brought images, music, and video into everyday life. The availability of open standards (such as JPEG, MPEG-X Audio and Video, H.26X) has had a major impact on this progression, notably due to the easy interoperability. Such standards have made the creation, and communication of (digital) data aimed at our most important senses, sight and hearing, simple, inexpensive and commonplace. With time, multimedia standards have addressed a growing set of fields from coding and metadata to rights management and content adaptation, following the increasing (functional and technical) complexity of multimedia applications.
Fernando Pereira 0001
ACM Multimedia1
2007 Wyner-Ziv Stereo Video Coding using a Side Information Fusion Approach
abstract
Wyner-Ziv coding, also known as distributed video coding, is currently a very hot research topic in video coding due to the new opportunities it opens. This paper applies the distributed video coding principles to stereo video coding, to propose a practical solution for Wyner-Ziv stereo coding based on mask-based fusion of temporal and spatial side informations. The architecture includes a low-complexity encoder and avoids any communication between the cameras/encoders. While the rate-distortion (RD) performance strongly depends on the motion-based frame interpolation (MBFI) and disparity-based frame estimation (DBFE) solutions, first results show that the proposed approach is promising and there are still issues to address.
José Diogo Areia, João Ascenso, Catarina Brites, Fernando Pereira 0001
MMSP4
2007 Studying the GOP Size Impact on the Performance of a Feedback Channel-Based Wyner-Ziv Video Codec
Fernando Pereira 0001, João Ascenso, Catarina Brites
PSIVT1
2006 Improving Transform Domain Wyner-Ziv Video Coding Performance
abstract
Distributed video coding (DVC) is a new video coding paradigm based on two key information theory results: the Slepian-Wolf and Wyner-Ziv theorems. A particular case of DVC, the so-called Wyner-Ziv coding, deals with lossy source coding with side information at the decoder and enables a flexible allocation of complexity between the encoder and the decoder. This paper proposes an improved transform domain Wyner-Ziv video codec including: 1) the integer block-based transform defined in the H.264/MPEG-4 AVC standard, 2) a quantizer with a symmetrical interval around zero for AC coefficients, and a quantization step size adjusted to the transform coefficient bands dynamic range, and 3) advanced frame interpolation for side information generation. The combination of these tools brings significant rate-distortion (RD) gains regarding the state-of-the-art results available in the literature
Catarina Brites, João Ascenso, Fernando Pereira 0001
ICASSP (2)3
2006 Improving Turbo Codec Integration in Pixel-Domain Distributed Video Coding
abstract
The field of distributed video coding (DVC) theory has received a lot of attention in recent years and effective encoding techniques have been proposed. In the present work the framework of pixel domain Wyner-Ziv coding of video frames is considered, following the scheme proposed in A. Aaron et al. (2002). Some key frames are supposed to be available at the decoder while other frames are Wyner-Ziv encoded using turbo codes; at the decoder motion compensated interpolation between the key frames is performed in order to construct the side information for the Wyner-Ziv frame decoding. In this paper an improved model for the correlation noise between the side information frame and the original one is proposed. It is shown that modeling the nonstationary nature of the noise leads to substantial gain in the rate-distortion performance. Furthermore, by considering the memory of the noise, we show that some further gain can be obtained by placing an interleaver before the turbo codec so as to spread the correlation noise all over the frame
Marco Dalai, Riccardo Leonardi, Fernando Pereira 0001
ICASSP (2)3
2006 Intra Mode Decision Based on Spatio-Temporal Cues in Pixel Domain Wyner-ZIV Video Coding
abstract
Distributed source coding principles have been recently applied to video coding in order to achieve a flexible distribution of the complexity burden between the encoder and the decoder. In this paper we elaborate on a pixel based Wyner-Ziv video codec that shifts all the complexity of the motion estimation phase to the decoder, thus achieving light encoding. We observe that the correlation noise statistics describing the relationship between the frame to be encoded and the side information available at the decoder is not spatially stationary. For this reason we introduce a mode decision scheme either at the encoder or at the decoder in such a way that when the estimated correlation is weak we opt for intra coding on a block-by-block basis. Both spatial and temporal criteria are used to determine whether a block is better intra coded or not
Marco Tagliasacchi, Alan Trapanese, Stefano Tubaro, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICASSP (2)6
2006 Content Adaptive Wyner-ZIV Video Coding Driven by Motion Activity
abstract
In distributed video coding (DVC), the video statistics are exploited, partially or totally at the decoder. A particular case of DVC, Wyner-Ziv video coding deals with lossy source coding with side information at the decoder and allows moving part or the entire motion estimation task to the decoder. In this context, it is the decoder responsibility to obtain the side information, a guess of the encoded Wyner-Ziv frame and the encoder only sends parity bits to improve its quality. In this paper, a technique targeting the improvement of the quality of the side information, and thus of the rate-distortion performance of the Wyner-Ziv codec is proposed. This is achieved by adaptively adjusting the size of the motion interpolation structure (or GOP length) according to the motion activity along the sequence. Experimentally, this allows to achieve gains up to 0.8 dB without performing any motion estimation or complex mode decision at the encoder.
João Ascenso, Catarina Brites, Fernando Pereira 0001
ICIP3
2006 Studying Temporal Correlation Noise Modeling for Pixel Based Wyner-Ziv Video Coding
abstract
Wyner-Ziv (WZ) video coding-a particular case of distributed video coding (DVC)-is a new video coding paradigm based on two major information theory results: the Slepian-Wolf and Wyner-Ziv theorems. Recently, practical WZ video coding solutions were proposed with promising results. Most of the solutions available in the literature, model the correlation noise between the original frame and the so-called side information by a given distribution whose relevant parameters are estimated in an offline process, at the encoder. In this paper, three algorithms are proposed towards a more realistic WZ coding approach by performing online estimation of the error distribution at the decoder. Both algorithms explore temporal correlation between frames however with different levels of granularity: frame, block and pixel levels; better rate-distortion (RD) performance is achieved for lower granularity (pixel) level.
Catarina Brites, João Ascenso, Fernando Pereira 0001
ICIP3
2006 Spatio-Temporal Scene Level Error Concealment for Shape and Texture Data in Segmented Video Content
abstract
In this paper, a novel shape and texture error concealment technique for segmented object-based video scenes is proposed. This technique is different from existing concealment techniques because it considers not only the corrupted video objects to be concealed, but also the context/scene in which they are inserted. In the proposed technique, concealment is done by using information from the current time instant as well as from the past. The obtained results suggest that the use of this technique significantly improves the subjective visual impact of scenes on the end-user, when compared to independent concealment of video objects.
Luís Ducla Soares, Fernando Pereira 0001
ICIP2
2006 Exploiting Spatial Redundancy in Pixel Domain Wyner-Ziv Video Coding
abstract
Distributed video coding is a recent paradigm that enables a flexible distribution of the computational complexity between the encoder and the decoder building on top of distributed source coding principles. In this paper we focus on the scenario where most of the complexity is shifted to the decoder, thus achieving light encoding. We elaborate on a well known pixel based Wyner-Ziv architecture and we improve its coding efficiency by exploiting both spatial and temporal correlation at the decoder side, without the need of performing any transform at the encoder. In order to generate the side information, the decoder adaptively chooses spatial or temporal information, based on the local estimate of the correlation noise. Simulations on test sequences demonstrate that a coding gain of up to +1.8 dB can be obtained with respect to the case that generates the side information by motion interpolation only.
Marco Tagliasacchi, Alan Trapanese, Stefano Tubaro, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICIP6
2006 Multimedia Representation in MPEG Standards: Achievements and Challenges
Fernando Pereira 0001
SECRYPT1
2006 Temporal shape error concealment by global motion compensation with local refinement
abstract
This paper presents an original temporal shape error concealment technique based on a combination of global and local motion compensation. For this technique, which is especially useful for object-based video applications in error-prone environments (e.g., mobile networks), it is assumed that the shape of the corrupted object at hand is in the form of a binary alpha plane and some of the shape data is missing due to channel errors. To conceal the corrupted shape, the decoder first assumes that a global motion model can describe the shape changes in consecutive time instants. This way, based on locally estimated global motion parameters, the decoder attempts to conceal the corrupted alpha plane by global motion compensating the shape data from the previous time instant. Afterwards, since a global motion model cannot perfectly describe all alpha plane changes, a local motion refinement is applied to improve the concealment in areas of the object with significant local motion.
Luís Ducla Soares, Fernando Pereira 0001
IEEE Trans. Image Process.2
2005 Motion compensated refinement for low complexity pixel based distributed video coding
abstract
Distributed video coding (DVC) is a new coding paradigm that enables to exploit video statistics, partially or totally at the decoder. A particular case of DVC, Wyner-Ziv coding, deals with lossy source coding with side information at the decoder and allows a shift of complexity from the encoder to the decoder, theoretically without any penalty in the coding efficiency. The Wyner-Ziv solution here described encodes each video frame independently (intraframe coding), but decodes the same frame conditionally (interframe decoding). At the decoder, and compensation tools are responsible to obtain an accurate interpolation of the original frame using previously decoded (temporally adjacent) frames. This paper proposes a novel approach to improve the performance of pixel domain Wyner-Ziv video coding by using a motion compensated refinement of the decoded frame and use it as improved side information. More precisely, upon partial decoding of each frame, the decoder refines its motion trajectories in order to achieve a better reconstruction of the decoded frame.
João Ascenso, Catarina Brites, Fernando Pereira 0001
AVSS3
2005 A Triple User Characterization Model for Video Adaptation and Quality of Experience Evaluation
abstract
This paper proposes the triple sensation-perception-emotion user characterization model for content adaptation and discusses adaptation tools and quality of experience metrics in light of this model. In the context of this paper, the utility would be associated to the sensorial, perceptual and emotional dimensions of the quality of experience. The adequate usage by the adaptation mechanism of available user preferences data will imply higher quality perceptual and emotional experiences. The major purpose of this paper is to launch some new ideas in the field of video adaptation and quality of experience evaluation. This model is hierarchical in the sense that typically emotions build on perceptions and perceptions build on sensations, setting a hierarchy
Fernando Pereira 0001
MMSP1
2005 Special issue on European projects on visual representation systems and services
Fernando Pereira 0001, Eric Badiqué
Signal Process. Image Commun.1
2004 Drift reduction for a H.264/AVC fine grain scalability with motion compensation architecture
abstract
The recent advances in nonscalable video encoding brought by the H.264-AVC standard offered significant improvements in terms of rate-distortion performance. This paper proposes a H.264-AVC based fine grain scalable video encoder which also exploits the motion compensation tools of the H.264-AVC standard to explore the temporal redundancy in the enhancement layer. The enhancement layer is predicted from a high quality reference obtained from past information of the enhancement and base layers. One of the drawbacks of this architecture is the drift effect, which occurs when part of the enhancement layer used for prediction is not received by the decoder. The drift reduction approaches here proposed simultaneously allow improvements in the coding efficiency and a reduction of the drift effect. The experimental results show improvements up to 2 dB in coding efficiency in comparison to Intra coding (like used by the MPEG-4 FGS standard) using the MPEG-4 testing conditions.
João Ascenso, Fernando Pereira 0001
ICIP2
2004 Motion-based shape error concealment for object-based video
abstract
In this paper, an original motion-based shape error concealment technique, especially useful for object-based video applications in error-prone environments such as mobile networks, is proposed. It is assumed that the shape of the corrupted object at hand is in the form of a binary alpha plane and some of the shape data is missing due to channel errors. To conceal the corrupted shape, the decoder starts by assuming that the alpha plane changes in consecutive time instants can be described by a global motion model. This way, based on locally estimated global motion parameters, the decoder tries to conceal the corrupted alpha plane by global motion compensating the shape data from the previous time instant. Then, since not all alpha plane changes can be perfectly described by global motion, an additional local motion refinement is applied to deal with areas of the object that have significant motion.
Luís Ducla Soares, Fernando Pereira 0001
ICIP2
2004 Automatic video summarization based on MPEG-7 descriptions
Pedro Miguel Fonseca, Fernando Pereira 0001
Signal Process. Image Commun.2
2004 Using MPEG standards for multimedia customization
João Magalhães, Fernando Pereira 0001
Signal Process. Image Commun.2
2004 Classification of video segmentation application scenarios
abstract
Video analysis can be used in the context of a wide variety of applications and therefore a multiplicity of techniques has been proposed in the literature. Each of those techniques is usually devoted to solving a specific part of the complete analysis problem, unless the problem is rather simple. Typically, to be able to propose meaningful analysis solutions, the analysis problem must first be appropriately constrained, taking into account the relevant application environment. Then, complementary types of analysis techniques may have to be used in combination to achieve the desired results. This paper proposes a classification of segmentation applications into a set of scenarios, according to the different application constraints and goals. This allows an easier selection of the appropriate video segmentation solution for each specific application. Examples of segmentation solutions for the most relevant scenarios identified are presented.
Paulo Lobato Correia, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.2
2004 Spatial shape error concealment for object-based image and video coding
abstract
In this paper, an original spatial shape error-concealment technique, to be used in the context of object-based image and video coding schemes, is proposed. In this technique, it is assumed that the shape of the corrupted object at hand is in the form of a binary alpha plane, in which some of the shape data is missing due to channel errors. From this alpha plane, a contour corresponding to the border of the object can be extracted. However, due to errors, some parts of the contour will be missing and, therefore, the contour will be broken. The proposed technique relies on the interpolation of the missing contours with Bézier curves, which is done based on the available surrounding contours. After all the missing parts of the contour have been interpolated, the concealed alpha plane can be easily reconstructed from the fully recovered contour and used instead of the erroneous one improving the final subjective impact.
Luís Ducla Soares, Fernando Pereira 0001
IEEE Trans. Image Process.2
2004 Adaptive shape and texture intra refreshment schemes for improved error resilience in object-based video coding
abstract
Video encoders may use several techniques to improve error resilience. In particular, for video encoders that rely on predictive (inter) coding to remove temporal redundancy, intra coding refreshment is especially useful to stop temporal error propagation when errors occur in the transmission or storage of the coded streams, since these errors may cause the decoded quality to decay very rapidly. In the context of object-based video coding, intra coding refreshment can be applied to both the shape and texture data. In this paper, novel shape and texture intra refreshment schemes are proposed which can be used by object-based video encoders, such as MPEG-4 video encoders, independently or combined. These schemes allow to adaptively determine when the shape and texture of the various video objects in a scene should be refreshed in order to maximize the decoded video quality for a certain total bit rate.
Luís Ducla Soares, Fernando Pereira 0001
IEEE Trans. Image Process.2
2003 Methodologies for objective evaluation of video segmentation quality
Paulo Lobato Correia, Fernando Pereira 0001
VCIP2
2003 Special issue on multimedia adaptation
Fernando Pereira 0001, Ian S. Burnett, Shih-Fu Chang
Signal Process. Image Commun.1
2003 Objective evaluation of video segmentation quality
abstract
Video segmentation assumes a major role in the context of object-based coding and description applications. Evaluating the adequacy of a segmentation result for a given application is a requisite both to allow the appropriate selection of segmentation algorithms as well as to adjust their parameters for optimal performance. Subjective testing, the current practice for the evaluation of video segmentation quality, is an expensive and time-consuming process. Objective segmentation quality evaluation techniques can alternatively be used; however, it is recognized that, so far, much less research effort has been devoted to this subject than to the development of segmentation algorithms. This paper discusses the problem of video segmentation quality evaluation, proposing evaluation methodologies and objective segmentation quality metrics for individual objects as well as for complete segmentation partitions. Both stand alone and relative evaluation metrics are developed to cover the cases for which a reference segmentation is missing or available for comparison.
Paulo Lobato Correia, Fernando Pereira 0001
IEEE Trans. Image Process.2
2003 Refreshment need metrics for improved shape and texture object-based resilient video coding
abstract
Video encoders may use several techniques to improve error resilience. In particular, for video encoders that rely on predictive (inter) coding to remove temporal redundancy, intra coding refreshment is especially useful to stop error propagation when errors occur in the transmission or storage of the coded streams, which can cause the decoded quality to decay very rapidly. In the context of object-based video coding, the video encoder can apply intra coding refreshment to both the shape and the texture data. In this paper, shape refreshment need and texture refreshment need metrics are proposed which can be used by object-based video encoders, notably MPEG-4 video encoders, to determine when the shape and the texture of the various video objects in the scene should be refreshed in order to improve the decoded video quality, e.g., for a given bitrate.
Luís Ducla Soares, Fernando Pereira 0001
IEEE Trans. Image Process.2
2002 Shape refreshment need metric for object-based resilient video coding
abstract
Although there are several techniques that video encoders may use to improve error resilience, it is largely recognized that intra coding refreshment plays a major role. This technique is especially useful for video encoders that rely on predictive (inter) coding to remove temporal redundancy because, in these conditions, the decoded quality can decay very rapidly due to error propagation if errors occur in the transmission or storage of the coded streams. Therefore, in order to avoid error propagation for too long a time, the encoder can use a coding refreshment scheme to refresh the decoding process and stop (spatial and temporal) error propagation. In the context of object-based video coding, the video encoder can apply intra coding refreshment to both the shape and the texture data. A shape refreshment need metric is proposed which can be used by object-based video encoders, notably MPEG-4 video encoders, to determine when the shape of a given video object should be refreshed in order to improve the decoded video quality.
Luís Ducla Soares, Fernando Pereira 0001
ICIP (1)2
2002 Texture refreshment need metric for resilient object-based video coding
abstract
Video encoders may use several techniques to improve error resilience. In particular, for video encoders that rely on predictive (inter) coding to remove temporal redundancy, intra coding refreshment is especially useful to stop error propagation when errors occur in the transmission or storage of the coded streams, which can cause the decoded quality to decay very rapidly. In object-based video coders, intra coding refreshment can be applied to both shape and texture data. A texture refreshment need metric is proposed which can be used by object-based video encoders, notably MPEG-4 video encoders, to determine when the texture of the various video objects in a scene should be refreshed in order to improve the decoded video quality, e.g. for a certain amount of bitrate resources.
Luís Ducla Soares, Fernando Pereira 0001
ICME (2)2
2002 Special issue on multimedia adaptation
Fernando Pereira 0001, Ian S. Burnett, Shih-Fu Chang
Signal Process. Image Commun.1
2002 Evaluating MPEG-4 video decoding complexity for an alternative video complexity verifier model
abstract
MPEG-4 is the first object-based audiovisual coding standard. To control the minimum decoding complexity resources required at the decoder, the MPEG-4 Visual standard defines the so-called video buffering verifier mechanism, which includes three virtual buffer models, among them the video complexity verifier (VCV). This paper proposes an alternative VCV model, based on a set of macroblock (MB) relative decoding complexity weights assigned to the various MB coding types used in MPEG-4 video coding. The new VCV model allows a more efficient use of the available decoding resources by preventing the overevaluation of the decoding complexity of certain MB types and thus making it possible to encode scenes (for the same profile@level decoding resources) which otherwise would be considered too demanding.
João Valentim, Paulo J. L. Nunes, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.3
2001 An alternative complexity model for the MPEG-4 video verifier mechanism
abstract
MPEG-4 is the first object-based audiovisual coding standard. To control the minimum decoding complexity resources required at the decoder, the MPEG-4 visual standard defines the so-called video complexity verifier (VCV). This paper proposes an alternative VCV model, based on a set of relative macroblock (MB) complexity weights assigned to the various MB coding types used in MPEG-4 video coding. The new VCV model allows a more efficient use of the available decoding resources by preventing the over-evaluation of the decoding complexity of certain MB types and thus making possible to encode scenes (for the same profile@level decoding resources) which otherwise would be considered too demanding.
João Valentim, Paulo J. L. Nunes, Fernando Pereira 0001
ICIP (1)3
2001 Scene level rate control algorithm for MPEG-4 video coding
Paulo J. L. Nunes, Fernando Pereira 0001
VCIP2
2000 Objective Evaluation of Relative Segmentation Quality
abstract
When working in image and video segmentation, the major objective is to design an algorithm producing the appropriate segmentation results for the particular goals of the application addressed. Therefore, the assessment of the segmentation quality assumes a crucial importance to evaluate the likeliness that the application targets are met. Since no well-established methods for objective segmentation quality evaluation are currently available, this paper's major goal is to propose objective metrics for the evaluation of relative segmentation quality for both individual objects and the overall segmentation partition. The paper presents a methodology for performing objective evaluation of relative segmentation quality, identifies the relevant features to be compared against those of the reference segmentation, and proposes appropriate objective quality metrics. These metrics build on the existing knowledge of segmentation quality evaluation and also on some relevant aspects from the video quality evaluation field.
Paulo Lobato Correia, Fernando Pereira 0001
ICIP2
2000 Influence of Encoder Parameters on the Decoded Video Quality for MPEG-4 over W-CDMA Mobile Networks
abstract
The MPEG-4 standard provides error resilience tools that can significantly increase the decoded video quality when using error prone media to transmit or store the encoded data. However, the use of these tools introduces extra redundancy and overhead in the bitstreams, which means that the decoded video quality can be severely affected if no careful configuration of the relevant parameters is done while taking into account the error characteristics of the channel being used. In this paper, a videotelephony system over a W-CDMA mobile network is used to study the influence of the encoding parameters on the decoded video quality, notably its behavior and optimization.
Luís Ducla Soares, Satoru Adachi, Fernando Pereira 0001
ICIP3
2000 MPEG-7: A standardised description of audiovisual content
Rob Koenen, Fernando Pereira 0001
Signal Process. Image Commun.2
2000 A contour-based approach to binary shape coding using a multiple grid chain code
Paulo J. L. Nunes, Ferran Marqués, Fernando Pereira 0001, Antoni Gasull
Signal Process. Image Commun.3
2000 MPEG-4: Why, what, how and when?
Fernando Pereira 0001
Signal Process. Image Commun.1
2000 Editorial
Fernando Pereira 0001, Philippe Salembier
Signal Process. Image Commun.1
1999 Hierarchical Visual Description Schemes for Still Images and Video Sequences
abstract
This paper proposes two description schemes (DSs) to describe the visual information of an audio-visual (AV) document. The first one, is devoted to still images. It describes the image visual appearance and its structure with regions as well as its semantic content in terms of objects. The second DS is devoted to video sequences. It describes the sequence structure as well as its semantic content in terms of events. Features such as motion, camera activity, etc. are included in this DS. Moreover, it involves static visual representations such as key-frames, background mosaics and key-regions. These elements are considered as still images and are described by the first DS.
Philippe Salembier, Noel E. O'Connor, Paulo Lobato Correia, Fernando Pereira 0001
ICIP (2)4
1999 Error resilience and concealment performance for MPEG-4 frame-based video coding
Luís Ducla Soares, Fernando Pereira 0001
Signal Process. Image Commun.2
1999 MPEG-4 facial animation technology: survey, implementation, and results
abstract
The emerging MPEG-4 standard specifies an object-based audiovisual representation framework, integrating both natural and synthetic content. Tools supporting three-dimensional facial animation will be standardized for the first time. To support facial animation decoders with different degrees of complexity, MPEG-4 uses a profiling strategy, which foresees the specification of object types, profiles, and levels adequate to the various relevant application classes. This paper first gives an overview of the MPEG-4 facial animation technology. Subsequently, the paper describes the Institute Superior Tecnico implementation of an MPEG-4 facial animation system, then evaluates the performance of the various tools standardized, using the MPEG-4 test material.
Gabriel Antunes Abrantes, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.2
1999 Introduction to the special issue on object-based video coding and description
Fernando Pereira 0001, Shih-Fu Chang, Rob Koenen, Atul Puri, Olivier Avaro
IEEE Trans. Circuits Syst. Video Technol.1
1998 Proposal for an Integrated Video Analysis Framework
abstract
The analysis of video data targeting the identification of relevant objects and the extraction of associated descriptive characteristics will be the enabling factor for a number of multimedia applications. This process has intrinsic difficulties, and since semantic criteria are difficult to express, usually only a part of the desired analysis results can be automatically achieved. For many applications, the automatic tools can be complemented with user guidance to improve performance. This paper proposes an integrated framework for video analysis, addressing the video segmentation and feature extraction problems. The framework includes a set of modules that can be combined following specific application needs. It includes both automatic (more objective) and user interaction (more semantic) analysis modules. The paper also proposes a specific segmentation solution to one of the most relevant application scenarios considered-off-line applications requiring precise segmentation.
Paulo Lobato Correia, Fernando Pereira 0001
ICIP (2)2
1998 An Alternative to the MPEG-4 Object-based Error Resilient Video Syntax
Luís Ducla Soares, Fernando Pereira 0001
ICIP (3)2
1998 MPEG-4: a flexible coding standard for the emerging mobile multimedia applications
abstract
This paper analyses the relevance and performance of the emerging MPEG-4 audiovisual coding standard for emerging mobile multimedia applications. Some results are presented for one of the MPEG-4 profiles targeting mobile scenarios.
Luís Ducla Soares, Fernando Pereira 0001
PIMRC2
1998 The role of analysis in content-based video coding and indexing
Paulo Lobato Correia, Fernando Pereira 0001
Signal Process.2
1997 Multi-grid chain coding of binary shapes
abstract
This paper presents a chain code based approach to efficiently code binary shape information of video objects, in the context of object-based video coding. The proposed method tries to meet some of the requirements of the MPEG-4 standard, currently under development, notably efficient coding, and low delay. This approach allows several modes of operation depending on the application requirements, notably lossless, near-lossless, and lossy coding modes. For the lossless case a pure differential chain code method is proposed while for the near-lossless case a multi-grid chain code (MGCC) technique is adopted. Also both INTRA and INTER prediction modes can be used. For the INTER mode, motion compensation is applied without coding the residues. The MGCC is a near-lossless contour coding technique using a contour description based on edges, which combines both contour prediction and contour simplification.
Paulo J. L. Nunes, Fernando Pereira 0001, Ferran Marqués
ICIP (3)2
1997 Subjective evaluation of MPEG-4 video codec proposals: Methodological approach and test procedures
abstract
A new audio-visual coding standard, MPEG-4, is currently under development. MPEG-4 will address not only compression, but also completely new audio-video coding functionalities related to content-based interactivity and universal access. As part of the MPEG-4 standardization process, in November, 1995 assessments were performed on technologies proposed for incorporation in the standard. These assessments included formal subjective tests, as well as expert panel evaluations. This paper describes the MPEG-4 video formal subjective tests. Since MPEG-4 addresses new coding functionalities, and also operates at bit-rates lower than ever subjectively tested before on a large scale, standard ITU test methods were not directly applicable. These methods had to be adapted, and even new test methods devised, for the MPEG-4 video subjective tests. We describe here the test methods used in the MPEG-4 video subjective tests, how the tests were carried out, and how the test results were interpreted. We also evaluate the successes and shortcomings of the MPEG-4 video subjective tests, and suggest possible improvements for future tests. The MPEG-4 video subjective tests were successful, providing the MPEG community with critical information to guide in the selection of technologies for inclusion in the video part of the MPEG-4 standard.
Thierry Alpert, Vittorio Baroncini, D. Choi, Laura Contin, Rob Koenen, Fernando Pereira 0001, H. Peterson
Signal Process. Image Commun.6
1997 MPEG-4: Context and objectives
Rob Koenen, Fernando Pereira 0001, Leonardo Chiariglione
Signal Process. Image Commun.2
1997 MPEG-4, Part 2: Submitted papers
Fernando Pereira 0001, Kevin O'Connell, Rob Koenen, Minoru Etoh
Signal Process. Image Commun.1
1997 MPEG-4, part 1: Invited papers
Fernando Pereira 0001, Kevin O'Connell, Rob Koenen, Minoru Etoh
Signal Process. Image Commun.1
1997 MPEG-4 video subjective test procedures and results
abstract
In the recent years, the technical developments in the area of audio-visual communications, notably in video coding, encouraged the emergence of new services which are already changing our everyday life. The convergence of the telecommunications, computer, and TV/film technologies is leading to the intermixture of elements formerly characteristic of each one of these fields, creating new needs and new requirements. Among the most important trends is the need to increase the interaction capabilities between the user and the audio-visual information, notably by considering the scene as a composition of objects-the content-according to a script that describes their spatial and temporal behavior and not just a set of pixels. MPEG-4 is a new audio-visual standard aiming to establish a universal, efficient coding of different forms of audio-visual data, called audio-visual objects. To reach this target, MPEG-4 has called for proposals on techniques that may be instrumental to efficiently represent visual information, allowing simultaneously high degrees of content-based interactivity and error resilience. This paper addresses the conditions under which the proposals to the MPEG-4 first round of video subjective tests have been evaluated. Moreover, the most significative results of these tests are also presented.
Fernando Pereira 0001, Thierry Alpert
IEEE Trans. Circuits Syst. Video Technol.1
1996 Very low bit-rate audio-visual applications
Fernando Pereira 0001, Rob Koenen
Signal Process. Image Commun.1
1995 Image segmentation towards new image representation methods
Diogo Cortez, Paulo J. L. Nunes, Manuel Menezes de Sequeira, Fernando Pereira 0001
Signal Process. Image Commun.4
1993 Knowledge-based videotelephone sequence segmentation
abstract
This paper presents a robust knowledge-based segmentation algorithm for videotelephony sequences ranging from studio based to mobile. It is able to divide each image in a sequence in non-overlapping head, body, and background areas. Its robustness stems from its ability to cope with the peculiarities of mobile sequences, having very detailed, moving backgrounds as well as strong camera movements (originating from vibration in car videotelephones or from small hand movements in hand-held videotelephones). The proposed algorithm uses edge and changed areas (due to speaker's motion) detection, as well as the redundancy associated to the speaker's position, as the basis for the segmentation. Geometrical knowledge-based techniques are then used to define the complete regions. The algorithm includes a quality estimation and control procedure, which enables it to decide whether to accept or reject the current segmentation, and which can be input to the videotelephone coder.
Manuel Menezes de Sequeira, Fernando Pereira 0001
VCIP2
1990 Two-layers constant-quality video coding for ATM environments
abstract
The near future Broadband Integrated Services Digital Network (B-ISDN) represents a new mark in Telecommunications allowing service integration and simplifying the global communications structure. Since one of the most important services in the B-ISDN will be video communications, it is essential to identify the characteristics of the video coding schemes which exploit, with the best performance, the new network capabilities. This paper presents some ideas and results related with one interesting coding scheme, suited to the B-ISDN - the two-layers coding scheme.
Fernando Pereira 0001, Lorenzo Masera
VCIP1