João Ascenso

dblp:50/6689 · DBLP profile ↗
← Back
92ranked-venue papers
10as first author
27since 2021 · last 2026
0000-0001-9902-5926ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 92 · 10 first-author · 27 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 An Overview of the JPEG AI Learning-Based Image Coding Standard
abstract
JPEG AI is an emerging learning-based image coding standard developed by Joint Photographic Experts Group (JPEG). The scope of the JPEG AI is the creation of a practical learning-based image coding standard offering a single-stream, compact compressed domain representation, targeting both human visualization and machine consumption. Scheduled for completion in early 2025, the first version of JPEG AI focuses on human vision tasks, demonstrating significant BD-rate reductions compared to existing standards, in terms of MS-SSIM, FSIM, VIF, VMAF, PSNR-HVS, IW-SSIM and NLPD quality metrics. Designed to ensure broad interoperability, JPEG AI incorporates various design features to support deployment across diverse devices and applications. This paper provides an overview of the technical features and characteristics of the JPEG AI standard.
Semih Esenlik, Yaojun Wu 0001, Zhaobin Zhang, Ye-Kui Wang, Kai Zhang 0007, Li Zhang 0006, João Ascenso, Shan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.7
2026 JPEG AI Compressed Domain Face Detection: A Multi-Scale Bridging Perspective
abstract
Learning-based image coding is showing improved compression efficiency, while also offering a novel advantage in enabling computer vision tasks directly within the compressed domain. The latent representation created by deep learning methods inherently contains all visual features, without a computationally expensive synthesis process at the decoder. This paper is an invited extension of a previous solution for JPEG AI compressed domain face detection that adapts a RetinaFace-based detector to operate directly on the latent tensor. In addition to a former single-scale bridging solution, this work provides a novel multi-scale bridging architecture to enable a more effective multi-scale compressed domain face detection. The results show a significant performance gain, improving accuracy up to 20% for detection of tiny faces on the WIDER FACE dataset compared to single-scale bridging, and further narrowing the gap when compared to detection on uncompressed or JPEG AI decoded images. Furthermore, since the computationally expensive decoding step is bypassed and since the bridges consist of lower-complexity networks, the overall processing cost is significantly reduced. Single and multi-scale bridging, respectively, have about 10% and 32% the complexity of applying pixel domain face detection on decoded images. The proposed architecture is expected to be extended to other multiscale sensitive vision tasks, as JPEG AI is not specifically designed for any single downstream application.
Ayman Alkhateeb, Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, João Ascenso, Fernando Pereira 0001
IEEE Trans. Multim.5
2025 Fine-Grained Subjective Visual Quality Assessment for High-Fidelity Compressed Images
abstract
Advances in image compression, storage, and display technologies have made high-quality images and videos widely accessible. At this level of quality, distinguishing between compressed and original content becomes difficult, highlighting the need for assessment methodologies that are sensitive to even the smallest visual quality differences. Conventional subjective visual quality assessments often use absolute category rating scales, ranging from “excellent” to “bad”. While suitable for evaluating more pronounced distortions, these scales are inadequate for detecting subtle visual differences. The JPEG standardization project AIC is currently developing a subjective image quality assessment methodology for high-fidelity images. This paper presents the proposed assessment methods, a dataset of high-quality compressed images, and their corresponding crowdsourced visual quality ratings. It also outlines a data analysis approach that reconstructs quality scale values in just noticeable difference (JND) units. The assessment method uses boosting techniques on visual stimuli to help observers detect compression artifacts more clearly. This is followed by a rescaling process that adjusts the boosted quality values back to the original perceptual scale. This reconstruction yields a fine-grained, high-precision quality scale in JND units, providing more informative results for practical applications. The dataset and code to reproduce the results will be available at https://github.com/jpeg-aic/dataset-BTC-PTC-24.
Michela Testolina, Mohsen Jenadeleh, Shima Mohammadi, Shaolin Su, João Ascenso, Touradj Ebrahimi, Jon Sneyers, Dietmar Saupe
DCC5
2025 Attention-Enhanced Multi-Branch Spiking Neural Network for Event Stream Super-Resolution
abstract
Traditional visual sensors capture images by sampling light at fixed intervals, producing a sequence of frames. In contrast, event vision sensors detect changes in light intensity asynchronously at the pixel level and generate discrete events with precise timing information, allowing to capture challenging scenes with high-speed object motion and extreme lighting conditions accurately. This paper introduces an attention-enhanced, polarity-aware multi-branch Spiking Neural Network (SNN) that directly super-resolves low-resolution event streams while preserving their temporal accuracy. The proposed architecture uses two parallel branches with novel spike-based spatial and temporal attention modules that give more importance to salient spatio-temporal structures, yet remaining fully asynchronous and hardware-friendly. Experimental results demonstrate that the proposed spike-based attention with multi-branch SNN significantly outperforms existing state-of-the-art methods across all evaluation metrics, while effectively maintaining the underlying spatio-temporal characteristics of the event stream.
Ahmadreza Sezavar, Catarina Brites, João Ascenso
ISM3
2025 Subjective Visual Quality Assessment for High-Fidelity Learning-Based Image Compression
abstract
Learning-based image compression methods offer a promising alternative to traditional codecs by improving rate-distortion performance. JPEG AI is the first standard in this domain and leverages deep neural networks to achieve high-fidelity image reconstruction. In this work, we present a comprehensive subjective visual quality assessment of JPEG AI-compressed images using the JPEG AIC-3 methodology, which quantifies perceptual differences using Just Noticeable Difference (JND) units. We created a dataset of 50 compressed images with fine-grained distortion levels from five diverse source images and conducted a large-scale crowdsourced experiment that collected 96,200 triplet responses from 459 participants. We reconstructed JND-based quality scales using a unified model based on both boosted and plain triplet comparisons. We also evaluated how well objective image quality metrics align with human perception in the high-fidelity range. The CVVDP metric achieved the highest overall performance, however, most metrics, including CVVDP, were overly optimistic in estimating image quality, emphasizing the need for rigorous subjective evaluation. We introduced the Meng–Rosenthal–Rubin test into Quality of Experience research to compare metric correlations with a shared ground truth. The dataset is publicly available1.
Mohsen Jenadeleh, Jon Sneyers, Panqi Jia, Shima Mohammadi, João Ascenso, Dietmar Saupe
QoMEX5
2025 Fine-Grained HDR Image Quality Assessment From Noticeably Distorted to Very High Fidelity
abstract
High dynamic range (HDR) and wide color gamut (WCG) technologies significantly improve color reproduction compared to standard dynamic range (SDR) and standard color gamuts, resulting in more accurate, richer, and more immersive images. However, HDR increases data demands, posing challenges for bandwidth efficiency and compression techniques. Advances in compression and display technologies require more precise image quality assessment, particularly in the high-fidelity range where perceptual differences are subtle. To address this gap, we introduce AIC-HDR2025, the first such HDR dataset, comprising 100 test images generated from five HDR sources, each compressed using four codecs at five compression levels. It covers the high-fidelity range, from visible distortions to compression levels below the visually lossless threshold. A subjective study was conducted using the JPEG AIC-3 test methodology, combining plain and boosted triplet comparisons. In total, 34,560 ratings were collected from 151 participants across four fully controlled labs. The results confirm that AIC-3 enables precise HDR quality estimation, with 95% confidence intervals averaging a width of 0.27 at 1 JND. In addition, several recently proposed objective metrics were evaluated based on their correlation with subjective ratings. The dataset is publicly available1.
Mohsen Jenadeleh, Jon Sneyers, Davi Lazzarotto, Shima Mohammadi, Dominik Keller, Atanas Boev, Rakesh Rao Ramachandra Rao, António M. G. Pinheiro, Thomas Richter 0005, Alexander Raake, Touradj Ebrahimi, João Ascenso, Dietmar Saupe
QoMEX12
2025 GS-QA: Comprehensive Quality Assessment Benchmark for Gaussian Splatting View Synthesis
abstract
Gaussian Splatting (GS) offers a promising alternative to Neural Radiance Fields (NeRF) for real-time 3D scene rendering. Using a set of 3D Gaussians to represent complex geometry and appearance, GS achieves faster rendering times and reduced memory consumption compared to the neural network approach used in NeRF. However, quality assessment of GS-generated static content is not yet explored in-depth. This paper describes a subjective quality assessment study that aims to evaluate synthesized videos obtained with several static GS state-of-the-art methods. The methods were applied to diverse visual scenes, covering both 360° and forward-facing (FF) camera trajectories. Moreover, the performance of 18 objective quality metrics was analyzed using the scores resulting from the subjective study, providing insights into their strengths, limitations, and alignment with human perception. All videos and scores are made available providing a comprehensive database that can be used as benchmark on GS view synthesis and objective quality metrics.
Pedro Martin, António Rodrigues, João Ascenso, Tiago Rosa Maria Paula Queluz
QoMEX3
2025 Uncertainty-Driven Sampling for Efficient Pairwise Comparison Subjective Assessment
abstract
Assessing image quality is crucial in image processing tasks such as compression, super-resolution, and denoising. While subjective assessments involving human evaluators provide the most accurate quality scores, they are impractical for large-scale or continuous evaluations due to their high cost and time requirements. Pairwise comparison subjective assessment tests, which rank image pairs instead of assigning scores, offer more reliability and accuracy but require numerous comparisons, leading to high costs. Although objective quality metrics are more efficient, they lack the precision of subjective tests, which are essential for benchmarking and training learning-based quality metrics. This paper proposes an uncertainty-based sampling method to optimize the pairwise comparison subjective assessment process. By utilizing deep learning models to estimate human preferences and identify pairs that need human labeling, the approach reduces the number of required comparisons while maintaining high accuracy. The key contributions include modeling uncertainty for accurate preference predictions and for pairwise sampling. The experimental results demonstrate superior performance of the proposed approach compared to traditional active sampling methods. An implementation of the pairwise sampling method is publicly available athttps://github.com/shimamohammadi/LBPS-EIC
Shima Mohammadi, João Ascenso
IEEE Trans. Multim.2
2024 Evaluation of strategies for efficient rate-distortion NeRF streaming
abstract
Neural Radiance Fields (NeRF) have revolutionized the field of 3D visual representation by enabling highly realistic and detailed scene reconstructions from a sparse set of images. NeRF uses a volumetric functional representation that maps 3D points to their corresponding colors and opacities, enabling photorealistic view synthesis from arbitrary viewpoints. Despite its advancements, the efficient streaming of NeRF content remains a significant challenge due to the large amount of data involved. This paper investigates the rate-distortion performance of two NeRF streaming strategies: pixel-based and neural network (NN) parameter-based streaming. While in the former, images are coded and then transmitted throughout the network, in the latter, the respective NeRF model parameters are coded and transmitted instead. This work also highlights the trade-offs in complexity and performance, demonstrating that the NN parameter-based strategy generally offers superior efficiency, making it suitable for one-to-many streaming scenarios.
Pedro Martin, António Rodrigues, João Ascenso, Tiago Rosa Maria Paula Queluz
ISM3
2024 Low Complexity Learning-based Lossless Event-based Compression
abstract
Event cameras are a cutting-edge type of visual sensors that capture data by detecting brightness changes at the pixel level asynchronously. These cameras offer numerous benefits over conventional cameras, including high temporal resolution, wide dynamic range, low latency, and lower power consumption. However, the substantial data rates they produce require efficient compression techniques, while also fulfilling other typical application requirements, such as the ability to respond to visual changes in real-time or near real-time. Additionally, many event-based applications demand high accuracy, making lossless coding desirable, as it retains the full detail of the sensor data. Learningbased methods show great potential due to their ability to model the unique characteristics of event data thus allowing to achieve high compression rates. This paper proposes a low-complexity lossless coding solution based on the quadtree representation that outperforms traditional compression algorithms in efficiency and speed, ensuring low computational complexity and minimal delay for real-time applications. Experimental results show that the proposed method delivers better compression ratios, i.e., with fewer bits per event, and lower computational complexity compared to current lossless data compression methods.
Ahmadreza Sezavar, Catarina Brites, João Ascenso
ISM3
2024 JPEG AI Compressed Domain Face Detection
abstract
Learning-based image coding has achieved competitive performance in terms of compression efficiency, while also gaining a key advantage in the ability to carry out computer vision tasks directly in the compressed domain. In fact, the latent representation which is generated using deep learning techniques may natively encapsulate all visual features needed for processing tasks, thereby eliminating the need to perform the expensive synthesis transform process at the decoder side. In this paper, it is proposed to perform face detection using the latent code present in the JPEG AI architecture. First, some experiments show how decoded images can be efficiently processed for face detection without retraining, albeit with some performance degradation. Then, for the first time a compressed domain RetinaFace-based detector applied to JPEG AI latent representations is competitively proposed. The performance achieved is comparable to the performance of the original RetinaFace applied to the reconstructed JPEG AI images, while reducing computational complexity since it bypasses the image decoding process. It is expected that this approach might be extended to other vision tasks since the JPEG AI representation format is not tailored specifically for any computer vision task.
Ayman Alkhateeb, Alessandro Gnutti, Fabrizio Guerrini, Riccardo Leonardi, João Ascenso, Fernando Pereira 0001
MMSP5
2024 Learning-based Lossless Event Data Compression
abstract
Emerging event cameras acquire visual information by detecting time domain brightness changes asynchronously at the pixel level and, unlike conventional cameras, are able to provide high temporal resolution, very high dynamic range, low latency, and low power consumption. Considering the huge amount of data involved, efficient compression solutions are very much needed. In this context, this paper presents a novel deep-learning-based lossless event data compression scheme based on octree partitioning and a learned hyperprior model. The proposed method arranges the event stream as a 3D volume and employs an octree structure for adaptive partitioning. A deep neural network-based entropy model, using a hyperprior, is then applied. Experimental results demonstrate that the proposed method outperforms traditional lossless data compression techniques in terms of compression ratio and bits per event.
Ahmadreza Sezavar, Catarina Brites, João Ascenso
VCIP3
2024 Globally and locally optimized Pannini projection for high FoV rendering of 360° images
abstract
To render a spherical (360° or omnidirectional) image on planar displays, a 2D image - called as viewport - must be obtained by projecting a sphere region on a plane, according to the user's viewing direction and a predefined field of view (FoV). However, any sphere to plan projection introduces geometric distortions, such as object stretching and/or bending of straight lines, which intensity increases with the considered FoV. In this paper, a fully automatic content-aware projection is proposed, aiming to reduce the geometric distortions when high FoVs are used. This new projection is based on the Pannini projection, whose parameters are firstly globally optimized according to the image content, followed by a local conformality improvement of relevant viewport objects. A crowdsourcing subjective test showed that the proposed projection is the most preferred solution among the considered state-of-the-art sphere to plan projections, producing viewports with a more pleasant visual quality.
Falah Jabar, João Ascenso, Tiago Rosa Maria Paula Queluz
Signal Process. Image Commun.2
2024 Guest Editorial Special Section on Recent Standardization Efforts for Learning-Based Visual Data Coding
abstract
Visual data coding is an enabling technology for various applications and is now ubiquitously adopted in modern image processing, communications, and computer vision systems. To enable interoperability between devices manufactured and services provided by different enterprises, a series of standards targeting visual data coding have been crafted in the past three decades. Several standardization organizations, such as ISO/IEC JTC 1/SC 29 consisting of Joint Picture Experts Group (JPEG) and Moving Picture Experts Group (MPEG),1ITU-T SG 16 Video Coding Experts Group (VCEG),2IEEE Data Compression Standards Committee Audio Video Coding Working Group (1857 WG),3MPAI Community,4have been creating these standards from many contributions of academia and industry. While most of these visual coding standards have been successfully deployed in many applications, there are more challenges nowadays, especially to accommodate the large volume of visual data in limited storage and limited bandwidth transmission links. Compression efficiency improvements are still needed, especially considering emerging data representation formats ranging from 8K/HDR image/video to rich plenoptic data.
Dong Liu 0002, Shan Liu 0001, João Ascenso, Dong Tian, Lu Yu 0003
IEEE Trans. Circuits Syst. Video Technol.3
2023 Complexity Scalable Learning-Based Image Decoding
abstract
Recently, learning-based image compression has attracted a lot of attention, leading to the development of a new JPEG AI standard based on neural networks. Typically, this type of coding solution has much lower encoding complexity compared to conventional coding standards such as HEVC and VVC (Intra mode) but has much higher decoding complexity. Therefore, to promote the wide adoption of learning-based image compression, especially to resource-constrained (such as mobile) devices, it is important to achieve lower decoding complexity even if at the cost of some coding efficiency. This paper proposes a complexity scalable decoder that can control the decoding complexity by proposing a novel procedure to learn the filters of the convolutional layers at the decoder by varying the number of channels at each layer, effectively having simple to more complex decoding networks. A regularization loss is employed with pruning after training to obtain a set of scalable layers, which may use more or fewer channels depending on the complexity budget. Experimental results show that complexity can be significantly reduced while still allowing a competitive rate-distortion performance.
Md Tahsir Ahmed Munna, João Ascenso
ICIP2
2023 Predictive Sampling for Efficient Pairwise Subjective Image Quality Assessment
abstract
Subjective image quality assessment studies are used in many scenarios, such as the evaluation of compression, super-resolution, and denoising solutions. Among the available subjective test methodologies, pair comparison is attracting popularity due to its simplicity, reliability, and robustness to changes in the test conditions, e.g. display resolutions. The main problem that impairs its wide acceptance is that the number of pairs to compare by subjects grows quadratically with the number of stimuli that must be considered. Usually, the paired comparison data obtained is fed into an aggregation model to obtain a final score for each degraded image and thus, not every comparison contributes equally to the final quality score. In the past years, several solutions that sample pairs (from all possible combinations) have been proposed, from random sampling to active sampling based on the past subjects' decisions. This paper introduces a novel sampling solution called Predictive Sampling for Pairwise Comparison (PS-PC) which exploits the characteristics of the input data to make a prediction of which pairs should be evaluated by subjects. The proposed solution exploits popular machine learning techniques to select the most informative pairs for subjects to evaluate, while for the other remaining pairs, it predicts the subjects' preferences. The experimental results show that PS-PC is the best choice among the available sampling algorithms with higher performance for the same number of pairs. Moreover, since the choice of the pairs is done a priori before the subjective test starts, the algorithm is not required to run during the test and thus much more simple to deploy in online crowdsourcing subjective tests.
Shima Mohammadi, João Ascenso
ACM Multimedia2
2023 On the Performance of Subjective Visual Quality Assessment Protocols for Nearly Visually Lossless Image Compression
abstract
The past decades have witnessed rapid growth in imaging as a major form of communication between individuals. Due to recent advances in capture, storage, delivery and display technologies, consumers demand improved perceptual quality while requiring reduced storage. In this context, research and innovation in lossy image compression have steered towards methods capable of achieving high compression ratios without compromising the perceived visual quality of images, and in some cases even enhancing the latter. Subjective visual quality assessment of images plays a fundamental role in defining quality as perceived by human observers. Although the field of image compression is constantly evolving towards efficient solutions for higher visual qualities, standardized subjective visual quality assessment protocols are still limited to those proposed in ITU-R Recommendation BT.500 and JPEG AIC standards. The number of comprehensive and in-depth studies where different protocols are compared is still insufficient. Moreover, previous works have not investigated the effectiveness of these methods on higher quality ranges, using recent image compression methods. In this paper, subjective visual scores collected from three subjective image quality assessment protocols, namely the Double Stimulus Continuous Quality Scale (DSCQS) and two test methods described in the JPEG AIC Part 2 standard, are compared between different laboratories under similar controlled conditions. The analysis of the experimental results has revealed that the DSCQS protocol is highly influenced by the quality of the reference images and experience of the subjects, while the JPEG AIC Part 2 specifications produce more stable results but are expensive and only suitable for a limited range of qualities. These emphasize the need for new robust subjective image quality assessment methodologies able to discriminate in the range of qualities generally demanded by consumers, i.e. from high to nearly visually lossless.
Michela Testolina, Davi Lazzarotto, Rafael Rodrigues, Shima Mohammadi, João Ascenso, António M. G. Pinheiro, Touradj Ebrahimi
ACM Multimedia5
2023 NeRF-QA: Neural Radiance Fields Quality Assessment Database
abstract
This short paper proposes a new database - NeRF-QA - containing 48 videos synthesized with seven NeRF based methods, together with their perceived quality scores, resulting from subjective assessment tests; for the videos selection, both real and synthetic, 360 degrees scenes were considered. This database will allow to evaluate the suitability, to NeRF based synthesized views, of existing objective quality metrics, and also the development of new quality metrics.
Pedro Martin, António Rodrigues, João Ascenso, Tiago Rosa Maria Paula Queluz
QoMEX3
2023 Fidelity-preserving Learning-Based Image Compression: Loss Function and Subjective Evaluation Methodology
abstract
Learning-based image compression methods have emerged as state-of-the-art, showcasing higher performance compared to conventional compression solutions. These data-driven approaches aim to learn the parameters of a neural network model through iterative training on large amounts of data. The optimization process typically involves minimizing the distortion between the decoded and the original ground truth images. This paper focuses on perceptual optimization of learning-based image compression solutions and proposes: i) novel loss function to be used during training and ii) novel subjective test methodology that aims to evaluate the decoded image fidelity. According to experimental results from the subjective test taken with the new methodology, the optimization procedure can enhance image quality for low-rates while offering no advantage for high-rates.
Shima Mohammadi, Yaojun Wu 0001, João Ascenso
VCIP3
2022 Evaluation of Sampling Algorithms for a Pairwise Subjective Assessment Methodology
abstract
Subjective assessment tests are often employed to evaluate image processing systems, notably image and video compression, super-resolution among others and have been used as an indisputable way to provide evidence of the performance of an algorithm or system. While several methodologies can be used in a subjective quality assessment test, pairwise comparison tests are nowadays attracting a lot of attention due to their accuracy and simplicity. However, the number of comparisons in a pairwise comparison test increases quadratically with the number of stimuli and thus often leads to very long tests, which is impractical for many cases. However, not all the pairs contribute equally to the final score and thus, it is possible to reduce the number of comparisons without degrading the final accuracy. To do so, pairwise sampling methods are often used to select the pairs which provide more information about the quality of each stimuli. In this paper, a reliable and much-needed evaluation procedure is proposed and used for already available methods in the literature, especially considering the case of subjective evaluation of image and video codecs. The results indicate that an appropriate selection of the pairs allows to achieve very reliable scores while requiring the comparison of a much lower number of pairs.
Shima Mohammadi, João Ascenso
ISM2
2022 Perceptual impact of the loss function on deep-learning image coding performance
abstract
Nowadays, deep-learning image coding solutions have shown similar or better compression efficiency than conventional solutions based on hand-crafted transforms and spatial prediction techniques. These deep-learning codecs require a large training set of images and a training methodology to obtain a suitable model (set of parameters) for efficient compression. The training is performed with an optimization algorithm which provides a way to minimize the loss function. Therefore, the loss function plays a key role in the overall performance and includes a differentiable quality metric that attempts to mimic human perception. The main objective of this paper is to study the perceptual impact of several image quality metrics that can be used in the loss function of the training process, through a crowdsourcing subjective image quality assessment study. From this study, it is possible to conclude that the choice of the quality metric is critical for the perceptual performance of the deep-learning codec and that can vary depending on the image content.
Shima Mohammadi, João Ascenso
PCS2
2021 Image Coding with Neural Network-Based Colorization
abstract
Automatic colorization is a process with the objective of inferring the color of grayscale images. This process is frequently used for artistic purposes and to restore the color in old or damaged images. Motivated by the excellent results obtained with deep learning-based solutions in the area of automatic colorization, this paper proposes an image coding solution integrating a deep learning-based colorization process to estimate the chrominance components based on the decoded luminance which is regularly encoded with a conventional image coding standard. In this case, the chrominance components are not coded and transmitted as usual, notably after some subsampling, as only some color hints, i.e. chrominance values for specific pixel locations, may be sent to the decoder to help it creating more accurate colorizations. To boost the colorization and final compression performance, intelligent ways to select the color hints are proposed. Experimental results show performance improvements with the increased level of intelligence in the color hints extraction process and a good subjective quality of the final decoded (and colorized) images.
Diogo Lopes, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICASSP2
2021 A Point-to-Distribution Joint Geometry and Color Metric for Point Cloud Quality Assessment
abstract
Point clouds (PCs) are a powerful 3D visual representation paradigm for many emerging application domains, especially virtual and augmented reality, and autonomous vehicles. However, the large amount of PC data required for highly immersive and realistic experiences requires the availability of efficient, lossy PC coding solutions are critical. Recently, two MPEG PC coding standards have been developed to address the relevant application requirements and further developments are expected in the future. In this context, the assessment of PC quality, notably for decoded PCs, is critical and asks for the design of efficient objective PC quality metrics. In this paper, a novel point-to-distribution metric is proposed for PC quality assessment considering both the geometry and texture. This new quality metric exploits the scale-invariance property of the Mahalanobis distance to assess first the geometry and color point-to-distribution distortions, which are after fused to obtain a joint geometry and color quality metric. The proposed quality metric significantly outperforms the best PC quality assessment metrics in the literature.
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
MMSP4
2021 Performance Evaluation of Objective Image Quality Metrics on Conventional and Learning-Based Compression Artifacts
abstract
Lossy image compression is a popular, simple and effective solution to reduce the amount of data representing digital pictures. In most lossy compression methods, the reduced volume of data in bits is achieved at the expense of introducing visual artifacts in the picture. The perceptual quality impact of such artifacts can be assessed with expensive and time-consuming subjective image quality experiments or through objective image quality metrics. However, the faster and less resource demanding objective quality metrics are not always able to reliably predict the quality as perceived by human observers. In this paper, the performance of 14 objective image quality metrics is benchmarked against a dataset of compressed images labeled with their subjective quality scores. Moreover, the performance of the above objective quality metrics in predicting the subjective quality of images distorted by both conventional and learning-based lossy compression artifacts is assessed and conclusions are drawn.
Michela Testolina, Evgeniy Upenik, João Ascenso, Fernando Pereira 0001, Touradj Ebrahimi
QoMEX3
2021 Large-Scale Crowdsourcing Subjective Quality Evaluation of Learning-Based Image Coding
abstract
Learning-based image codecs produce different compression artifacts, when compared to the blocking and blurring degradation introduced by conventional image codecs, such as JPEG, JPEG 2000 and HEIC. In this paper, a crowdsourcing based subjective quality evaluation procedure was used to benchmark a representative set of end-to-end deep learning-based image codecs submitted to the MMSP'2020 Grand Challenge on Learning-Based Image Coding and the JPEG AI Call for Evidence. For the first time, a double stimulus methodology with a continuous quality scale was applied to evaluate this type of image codecs. The subjective experiment is one of the largest ever reported including more than 240 pair-comparisons evaluated by 118 naïve subjects. The results of the benchmarking of learning-based image coding solutions against conventional codecs are organized in a dataset of differential mean opinion scores along with the stimuli and made publicly available.
Evgeniy Upenik, Michela Testolina, João Ascenso, Fernando Pereira 0001, Touradj Ebrahimi
VCIP3
2021 Lenslet Light Field Image Coding: Classifying, Reviewing and Evaluating
abstract
In recent years, visual sensors have been quickly improving, notably targeting richer acquisitions of the light present in a visual scene. In this context, the so-called lenslet light field (LLF) cameras are able to go beyond the conventional 2D visual acquisition models, by enriching the visual representation with directional light measures for each pixel position. LLF imaging is associated to large amounts of data, thus critically demanding efficient coding solutions in order applications involving transmission and storage may be deployed. For this reason, considerable research efforts have been invested in recent years in developing increasingly efficient LLF imaging coding (LLFIC) solutions. In this context, the main objective of this paper is to review and evaluate some of the most relevant LLFIC solutions in the literature, guided by a novel classification taxonomy, which allows better organizing this field. In this way, more solid conclusions can be drawn about the current LLFIC status quo, thus allowing to better drive future research and standardization developments in this technical area.
Catarina Brites, João Ascenso, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Point Cloud Rendering After Coding: Impacts on Subjective and Objective Quality
abstract
Recently, point clouds have shown to be a promising way to represent 3D visual data for a wide range of immersive applications, from augmented reality to autonomous cars. Emerging imaging sensors have made easier to perform richer and denser point cloud acquisition, notably with millions of points, thus raising the need for efficient point cloud coding solutions. In such scenario, it is important to evaluate the impact and performance of several processing steps in a point cloud communication system, notably the degradations associated to point cloud coding solutions. Moreover, since point clouds are not directly visualized but rather processed with a rendering algorithm before shown on any display, the perceived quality of point cloud data highly depends on the rendering solution. In this context, the main objective of this paper is to study the impact of several coding and rendering solutions on the perceived user quality and in the performance of available objective assessment metrics. Another contribution regards the assessment of recent MPEG point cloud coding solutions for several popular rendering methods, which was never presented before. The conclusions regard the visibility of three types of coding artifacts for the three considered rendering approaches as well as the strengths and weaknesses of objective metrics when point clouds are rendered after coding.
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
IEEE Trans. Multim.4
2020 Improving Psnr-Based Quality Metrics Performance For Point Cloud Geometry
abstract
An increased interest in immersive applications has drawn attention to emerging 3D imaging representation formats, notably light fields and point clouds (PCs). Nowadays, PCs are one of the most popular 3D media formats, due to recent developments in PC acquisition, namely with new depth sensors and signal processing algorithms. To obtain high fidelity 3D representations of visual scenes a huge amount of PC data is typically acquired, which demands efficient compression solutions. As in 2D media formats, the final perceived PC quality plays an importance role in the overall user experience and, thus, objective metrics capable to measure the PC quality in a reliable way are essential. In this context, this paper proposes and evaluates a set of objective quality metrics for the geometry component of PC data, which plays a very important role on the final perceived quality. Based on the popular PSNR PC geometry quality metric, novel improved PSNR-based metrics are proposed by exploiting the intrinsic PC characteristics and the rendering process that must occur before visualization. The experimental results show the superiority of the best proposed metrics over state-of-the-art, obtaining an improvement up to 32% in the Pearson correlation coefficient.
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
ICIP4
2020 A Generalized Hausdorff Distance Based Quality Metric for Point Cloud Geometry
abstract
Reliable quality assessment of decoded point cloud geometry is essential to evaluate the compression performance of emerging point cloud coding solutions and guarantee some target quality of experience. This paper proposes a novel point cloud geometry quality assessment metric based on a generalization of the Hausdorff distance. To achieve this goal, the so-called generalized Hausdorff distance for multiple rankings is exploited to identify the best performing quality metric in terms of correlation with the MOS scores obtained from a subjective test campaign. The experimental results show that the quality metric derived from the classical Hausdorff distance leads to low objective-subjective correlation and, thus, fails to accurately evaluate the quality of decoded point clouds for emerging codecs. However, the quality metric derived from the generalized Hausdorff distance with an appropriately selected ranking, outperforms the MPEG adopted geometry quality metrics when decoded point clouds with different types of coding distortions are considered.
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
QoMEX4
2020 Visual monitoring of High-Sea fishing activities using deep learning-based image processing
Pedro Perdigão, Pedro Lousã, João Ascenso, Fernando Pereira 0001
Multim. Tools Appl.3
2020 Point cloud coding: A privileged view driven by a classification taxonomy
Fernando Pereira 0001, Antoine Dricot, João Ascenso, Catarina Brites
Signal Process. Image Commun.3
2020 Mahalanobis Based Point to Distribution Metric for Point Cloud Geometry Quality Evaluation
abstract
Nowadays, point clouds (PCs) are a promising representation format for immersive content and target several emerging applications, notably in virtual and augmented reality. However, efficient coding solutions are critically needed due to the large amount of PC data required for high quality user experiences. To address these needs, several PC coding standards were developed and thus, objective PC quality metrics able to accurately account for the subjective impact of coding artifacts are needed. In this paper, a scale-invariant PC geometry quality assessment metric is proposed based on a new type of correspondence, namely between a point and a distribution of points. This metric is able to reliably measure the geometry quality for PCs with different intrinsic characteristics and degraded by several coding solutions. Experimental results show the superiority of the proposed PC quality metric over relevant state-of-the-art.
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
IEEE Signal Process. Lett.4
2019 Hybrid Point Cloud Geometry Coding Using Planes and Octree Representation Models
abstract
Nowadays, point clouds are considered as a promising representation model for future 3D applications, such as augmented reality. However, large amounts of data are required to offer high quality immersive experiences, and therefore, compression efficiency is an important factor. Many of the current coding techniques rely on octree data structures, that offer several advantages such as scalability with layers offering multiple levels of detail. Although octrees are capable to efficiently represent a point cloud, geometric representations may also play an important role. In this context, the objective of this paper is to propose a hybrid point cloud compression solution based on plane estimation and coding that enhances the octree structure, thus exploiting two representation models for point clouds. In the proposed solution, the octree partitioning is adaptive and includes this novel plane coding mode for leaf nodes at different layers of the octree. A large increase in coding efficiency against the pure octree coding case is observed with average BD-rate gains of 35%, up to 68% reported. Moreover, in several cases it can also significantly outperform state-of-the-art solutions that are nowadays considered for standardization.
Antoine Dricot, João Ascenso
DCC2
2019 Content-Aware Perspective Projection Optimization for Viewport Rendering of 360° Images
abstract
To visualize the spherical (360°) content on 2D planar displays, a selected part of the sphere must be projected into a 2D plane. The resulting 2D area is called viewport and is created according some fixed and limited field of view, usually much less than 360°. To generate the viewport, two geometric projections are usually employed: rectilinear and stereographic. For both cases, this process leads to geometric distortions in the viewport image which, depending on the content, may be visually unpleasant to users. Rectilinear projection stretches the image borders while the stereographic projection bends straight lines. In this paper, the viewport rendering problem is tackled with the formulation of a generalized perspective projection (which includes many other popular projection, besides perspective and stereographic) and a perceptual optimization of its key parameter: the projection center. The optimization process is performed according to the viewport characteristics, namely with perceptually important features that characterize the amount of line distortion and stretching of salient regions. The aim is to find an optimal, content dependent, projection center that minimizes the content distortions and thus provides better quality of experience (QoE) in the visualization of spherical images. Experimental results on several 360° images show that the proposed optimized projection achieves better results than the benchmark projections both quantitatively and qualitatively.
Falah Jabar, João Ascenso, Tiago Rosa Maria Paula Queluz
ICME2
2019 Adaptive Plane Projection for Video-Based Point Cloud Coding
abstract
One of the most promising emerging 3D representation paradigms is the point cloud (PC) model, notably due to the new set of applications that it enables, from immersive telepresence to 3D geographic information systems. Recognizing the potential of this representation model, MPEG has launched a standardization project to specify efficient PC coding solutions. This project has led to the so-called Video-based Point Cloud Coding (V-PCC) standard, which is based on the idea of projecting the dynamic PC geometry and texture into a sequence of frames to be coded with the highly efficient HEVC video coding standard. In V-PCC, the projection of points and color attributes is always performed using the same, rigid set of projection planes, independently of the PC characteristics. This paper proposes a more flexible coding solution, which improves the V-PCC Intra coding mode by adopting a content dependent set of projection planes, thus, more adapted to the characteristics of the PCs to be coded, targeting a better compression performance. The experimental results show an average total bitrate saving of around 15% regarding V-PCC with clear benefits in terms of subjective assessment.
Eurico Lopes, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICME2
2019 Improved Patch Packing for the MPEG V-PCC Standard
abstract
Point cloud representation is an emerging visual data technology, targeting immersive 3D experiences in the context of multiple applications scenarios, notably entertainment, geographical information systems, medicine, architecture, and robotics. Since a point cloud may easily involve millions of points, and thus an enormous amount of data, its effective storage and transmission critically asks for efficient coding solutions. With this purpose in mind, several point cloud coding (PCC) solutions have been proposed in the literature; special emphasis is due to the recent MPEG standards, which target interoperability in this domain, notably the MPEG V-PCC standard. The objective of this paper is to improve the V-PCC standard compression efficiency by proposing novel solutions for the V-PCC packing module without compromising in any way the V-PCC stream (syntax and semantics) and decoder compliance. In this context, several patch packing solutions are proposed, including new packing algorithms and associated sorting and positioning metrics; for the metrics, both absolute and relative approaches are proposed. The RD performance results show BD-Rate savings up to 0.8% for the best packing solution regarding the V-PCC benchmark. Moreover, the packing map size reductions can go up to 12%, on average.
Afonso Costa, Antoine Dricot, Catarina Brites, João Ascenso, Fernando Pereira 0001
MMSP4
2019 Adaptive Multi-level Triangle Soup for Geometry-based Point Cloud Coding
abstract
Nowadays, point clouds are considered as a promising representation for future immersive 3D applications. However, to recreate a 3D object or scene with high fidelity, a large number of points is required, often with high coordinates precision. Thus, efficient compression schemes are much needed for transmission and storage of point cloud data. Several types of coding techniques (e.g. 2D mapping, graph transforms, etc.) have been proposed in the literature to solve this problem. Octree-based solutions provide compression efficiency and level-of-detail scalability, especially for large static point clouds. The upcoming MPEG Geometry based Point Cloud Coding (G-PCC) solution provides high coding performance and relies on a static pruned octree, combined with a trisoup (for triangle soup) surface reconstruction. In this paper, the geometry coding process of G-PCC is enhanced by introducing mode decisions, thus enabling the trisoup in leaf nodes at multiple levels of the octree, i.e. allowing a more adaptive octree partitioning. Experimental results report average BD-rate gains of 5.3% (up to 7%) for the geometry component (point-to-plane error) on dense point clouds, with no significant impact on the color coding performance.
Antoine Dricot, João Ascenso
MMSP2
2019 Graph-Based Static 3D Point Clouds Geometry Coding
abstract
Recently, 3D visual representation models such as light fields and point clouds are becoming popular due to their capability to represent the real world in a more complete and immersive way, paving the road for new and more advanced visual experiences. The point cloud representation model is able to efficiently represent the surface of objects/scenes by means of a set of 3D points and associated attributes and is increasingly being used from autonomous cars to augmented reality. Emerging imaging sensors have made it easier to perform richer and denser point cloud acquisitions, notably with millions of points, making it impossible to store and transmit these very high amounts of data without appropriate coding. This bottleneck has raised the need for efficient point cloud coding solutions in order to offer more immersive visual experiences and better quality of experience to the users. In this context, this paper proposes an efficient lossy coding solution for the geometry of static point clouds. The proposed coding solution uses an octree-based approach for a base layer and a graph-based transform approach for the enhancement layer where an Inter-layer residual is coded. The performance assessment shows very significant compression gains regarding the state-of-the-art, especially for the most relevant lower and medium rates.
Paulo de Oliveira Rente, Catarina Brites, João Ascenso, Fernando Pereira 0001
IEEE Trans. Multim.3
2019 Blind Quality Assessment of 3-D Synthesized Views Based on Hybrid Feature Classes
abstract
In this paper, a novel quality metric to evaluate depth-based synthesized views is proposed. This metric relies on a hybrid approach that uses features extracted in different phases of the image synthesis procedure, namely from the bitstream, from intermediate data produced during the synthesis process, and from the final synthesized view; these features are then combined using support vector regression. A new data set of synthesized images, with compression and rendering artifacts, was built and used to develop and assess the proposed metric. The metric performance is compared with conventional full-reference two-dimensional image quality assessment metrics and with quality assessment metrics developed specifically for synthesized images. The experimental results showed that the proposed solution outperforms the considered benchmark metrics, being able to predict the subjective quality scores of the synthesized images with a Pearson correlation coefficient close to 0.9.
Fábio Rodrigues, João Ascenso, António Rodrigues, Tiago Rosa Maria Paula Queluz
IEEE Trans. Multim.2
2018 Rate-Distortion Driven Adaptive Partitioning for Octree-Based Point Cloud Geometry Coding
abstract
Point clouds are widely recognized as a promising 3D visual representation model for future 3D applications, e.g. Augmented/Virtual Reality and autonomous mobile navigation. To enable immersive and high quality experiences, a large amount of data needs to be transmitted, and thus, efficient compression techniques are a key factor for its success. Nowadays, many point cloud coding techniques rely on layered octree data structures able to provide multiple benefits, notably level of detail scalability. However, the octree creation process still lacks flexibility compared to the optimal partitioning methods popular in 2D encoders, which are mainly driven by rate-distortion trade-offs. The goal of this paper is to propose a novel octree creation mechanism which is able to exploit the intrinsic point cloud resolution by iteratively partitioning the octree voxels by means of split decisions based on a Rate-Distortion Optimization process. In addition, the arithmetic coding step is adapted to the density at each octree depth. This novel approach allows reaching a significant bitrate reduction, notably 5% in average, and up to 9.8%, over the adopted coding anchor using state-of-the-art octree based coding.
Antoine Dricot, Fernando Pereira 0001, João Ascenso
ICIP3
2018 Holographic Data Coding: Benchmarking and Extending HEVC With Adapted Transforms
abstract
Holography is an emerging technology to represent and display visual information with high expectations in terms of user experience. A hologram is a reproduction of a light field represented through the interference pattern between two wavefields the reference and the object wavefields. Whatever their creation process holograms may have a digital representation using some appropriate format. Moreover considering the huge amounts of data involved digital holographic data have to be compressed using appropriate coding solutions for example available image coding standard solutions or efficient extensions of them. In this context this paper contributes to advance the state-of-the-art on holographic data coding by: 1) benchmarking the most relevant available image coding standard solutions when using the most relevant holographic data representation formats; 2) proposing a novel mode depend directional transform-based HEVC coding solution trained with holographic data. Experimental results obtained under meaningful test conditions show that the proposed coding solution outperforms the state-of-the-art HEVC coding standard for specific formats and conditions. Altogether these two contributions are critical to understand the current status quo and advance the state-of-the-art on holographic data coding.
Jose Peixeiro, Catarina Brites, João Ascenso, Fernando Pereira 0001
IEEE Trans. Multim.3
2017 Rate-Accuracy Optimization of Deep Convolutional Neural Network Models
abstract
Recently, deep learning has enjoyed a great deal of success for computer vision problems due to its capability to model highly complex tasks, such as image classification, object detection, face recognition, among many others. Although these neural networks are nowadays very powerful, there is a huge amount of parameters (i.e. the model) that need to be learned and require considerable storage space and bandwidth during transmission. This paper addresses the problems of storage and transmission of large deep learning models by proposing a compression solution that is independent of the model being trained as well as the data used for training. An efficient compression framework for the parameters of a neural network, more precisely the weights that interconnect the different neurons, which consume a significant amount of resources (memory, storage and bandwidth) is proposed. Several quantization strategies are considered as well as a statistical models for the different layers of a neural network, which are exploited by an arithmetic coding engine. Experimental results show that up to 92% bitrate savings can be obtained with minimal impact in terms of image classification accuracy.
Alessandro Filini, João Ascenso, Riccardo Leonardi
ISM2
2017 Perceptual Analysis of Perspective Projection for Viewport Rendering in 360° Images
abstract
Omnidirectional, also referred to as 360o, visual content provides an immersive experience since it allows users to view a visual scene from different directions. The overall content typically covers a full sphere, and omnidirectional videos or images are processed to obtain a projection on a 2D plane of a fraction of the sphere (aka viewport), which is shown to the user. Therefore, users can look around freely and only a restricted field of view (FoV) of the entire content is shown, according to the selected viewing direction. To maximize the immersion and engagement that omnidirectional content can provide, viewports must be generated and displayed with a large FoV. However, when large FoVs are used, geometric distortions (e.g., shape and scale deformations) can be present. In this paper, the perceptual impact of the perspective projection for viewport rendering of 360oimages is evaluated. This study is based on subjective experiments conducted at Instituto Superior Técnico (IST), aiming to assess the perceptual impact of the rendering projection parameters, such as FoV and centre of projection, as well as the influence of the image content.
Falah Jabar, João Ascenso, Tiago Rosa Maria Paula Queluz
ISM2
2017 Subjective and objective quality evaluation of compressed point clouds
abstract
The increasing availability of point cloud data in recent years is demanding high performance compression solutions. Naturally, methods to perform objective quality assessment of compressed point clouds are also very much needed, namely metrics to measure the geometry distortion of point clouds when positioning errors are present. This is a rather challenging problem since this 3D representation format is unstructured and it is typically not directly visualized. In this context, the objective of this paper is to perform subjective and objective quality assessment of point clouds degraded by compression artifacts and to evaluate the correlation of the most popular objective quality metrics with human perception. In this work, subjective experiments conducted at Instituto Superior Técnico (IST) are described with point clouds compressed with two different but yet promising solutions, one based on the octree representation of the 3D space and another based on the rather popular graph transform. As far as the authors know, this is the first study of this type made available and should have a key role on the future development and evaluation of point cloud coding solutions.1
Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso
MMSP4
2017 Epipolar based light field key-location detector
abstract
Nowadays, visual features play a key role, as they can provide a concise representation of visual data that is efficient for multiple tasks, notably content retrieval and object recognition. In parallel, visual sensors have been improving, targeting richer acquisitions of the light in a visual scene. In this context, the so-called light field cameras, which have recently emerged, are able to go beyond the standard acquisition models, by enriching the visual representation with directional light measures for each pixel position, e.g. by using a so-called lenslet light field camera. At this stage, not much research has been made in the field of feature detection and description for the emerging lenslet light field format. In this context, this paper proposes a feature detector suitable for lenslet light field images based on the exploitation of an alternative visual parametrization of the light field, called the Epipolar Planar Image (EPI). The proposed detector is heavily based on line detection in the EPI representation, since 3D points in the visual scene are mapped to line segments in an EPI, and the detector output is referred as key-locations. The proposed light field key-location detector is assessed with a solid evaluation framework using a large light field dataset. In comparison to the 2D SIFT detector, up to 10% improvements were achieved for the widely used repeatability metric.
José Abecasis Teixeira, Catarina Brites, Fernando Pereira 0001, João Ascenso
MMSP4
2017 Saliency-driven omnidirectional imaging adaptive coding: Modeling and assessment
abstract
Omnidirectional imaging, also known as 360° and spherical imaging, records all 360° of a scene from a specific spatial position, thus offering the user the capability to enjoy three rotational degrees of freedom (3-DoF). To offer a good quality of experience, omnidirectional imaging requires very high bitrates as high spatial resolution are a must and, ideally, also high frame rates. Due to the lack of video coding solutions specifically designed for omnidirectional imaging, this type of content is typically coded with the available image and video coding standards, such as JPEG, H.264/AVC and HEVC, after applying a 2D rectangular projection. In this context, this paper proposes an omnidirectional imaging coding solution allowing to reach improved coding performance by using an adaptive coding solution where the most visually salient image/video regions are coded with higher quality in a process appropriately controlled by the quantization parameter. To determine the saliency of the various omnidirectional imaging regions, a machine-learning based saliency detection model is proposed. The proposed coding solution achieves compression gains as measured by a novel objective quality metric also driven by saliency. This novel objective quality metric is validated by formal subjective testing where very high correlations with the subjective tests scores are achieved.
Guilherme Luz Tortorella, João Ascenso, Catarina Brites, Fernando Pereira 0001
MMSP2
2017 Adaptive Scalable Video Coding: An HEVC-Based Framework Combining the Predictive and Distributed Paradigms
abstract
The emerging scalable High Efficiency Video Coding (SHVC) video coding standard provides an efficient solution for transmission of video over heterogeneous and time dynamic networks, terminals, and usage environments. The encoding complexity and the error sensitivity associated with the efficient HEVC coding tools adopted in SHVC make this scalable codec less attractive to some emerging applications such as video surveillance, visual sensor network, and remote space transmission where these requirements are critical. To address the requirements of these application scenarios including scalability, this paper proposes a novel HEVC-based framework offering quality scalability on top of an HEVC compliant base layer while appropriately combining the predictive and distributed coding paradigms. To achieve the best enhancement layer compression efficiency, two novel coding tools are proposed, notably a machine learning-based side information creation mechanism and an adaptive correlation modeling process. The experimental results reveal that the rate-distortion performance of the proposed distributed scalable video coding-HEVC solution outperforms the relevant alternative coding solutions, notably by up to 52.9% and 23.7% BD-rate gains regarding the HEVC-Simulcast and SHVC standard solutions, respectively, for an equivalent prediction configuration, while achieving a lower encoding complexity.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.2
2016 Improving SHVC performance with a joint layer coding mode
abstract
The growing need for a powerful scalable video coding engine targeting the heterogeneous landscape of network, devices, and consumption environments has led to the development of the Scalable High Efficiency Video Coding (SHVC) standard, an extension of the High Efficiency Video Coding (HEVC) standard. To improve the SHVC compression efficiency, this paper proposes a novel joint layer coding mode to be integrated in the SHVC codec. In the proposed coding mode, the base layer (BL) and enhancement layer (EL) decoded information are linearly combined at the pixel level to create an additional coding mode. To fuse the BL and EL driven predictions, a weighting term is defined to indicate the contributions of each of them for the final joint layer prediction. To reach high adaptability, these weights are computed at pixel level in the prediction unit. Moreover, to achieve the highest compression efficiency, the proposed joint layer coding mode is adaptively selected using a rate distortion optimization (RDO) mechanism. Experiments conducted for a rich set of test conditions have shown that significant compression efficiency gains can be achieved with the proposed joint layer coding mode, notably up to 4.3 % in BD-Rate savings regarding the standard SHVC quality scalable codec.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
ICASSP2
2016 Multi-view distributed source coding of binary features for visual sensor networks
abstract
Visual analysis algorithms have been mostly developed for a centralized scenario where all visual data is acquired and processed at a central location. However, in visual sensor networks (VSN), several constraints in computational power, energy and bandwidth require a radically different approach, notably a paradigm shift from centralized to distributed visual processing. In the new paradigm, visual data is acquired and features are extracted at the sensing nodes locations to be after transmitted to enable further analysis at some central location. In such scenario, one of the key challenges is to design suitable feature coding schemes that are able to exploit the correlation among the features corresponding to (partially) overlapped views of the same visual scene. To achieve efficient coding, it is proposed to employ the distributed source coding paradigm as it does not require any communication between the sensing nodes (rather expensive in VSN) and it is parsimonious in terms of computational resources. Experimental results show that significant accuracy and compression gains (up to 37.36%) can be achieved when coding features extracted from multiple views.
Nuno Monteiro, Catarina Brites, Fernando Pereira 0001, João Ascenso
ICASSP4
2016 Multi-view distributed coding and selection of local binary features
abstract
Recently, the latest advances in compact feature representation and feature learning have provided an efficient framework for several visual analysis tasks, such as object recognition. However, when multiple cameras with overlapping fields-of-view are employed, other visual analysis tasks such as depth estimation can be supported and object recognition accuracy can be improved. In this paper the problem of distributed visual analysis from multiple views of a scene is addressed, considering that computational power and bandwidth, at each camera sensor, are rather limited. More specifically, an efficient coding technique for local binary features is proposed which exploits the correlation at the decoder side between each descriptor and its quantized representation. Moreover, considering that descriptors representing the same visual feature across different views are well correlated, a technique to avoid the transmission of redundant descriptors from multiple views is proposed. At the decoder, the joint statistics of all descriptors from all views is used to drive the selection of the best descriptors to be transmitted by each sensing node. The proposed multi-view feature coding and selection techniques allow obtaining bitrate reductions up to 80%, with respect to the uncompressed descriptor rate, for a certain task accuracy.
Nuno Monteiro, Catarina Brites, Fernando Pereira 0001, João Ascenso
ICME4
2016 Digital holography: Benchmarking coding standards and representation formats
abstract
Holography is an emerging technology to represent and display visual information and the associated experience expectations on the layman imaginary are huge. A hologram is a reproduction of a light field represented through the interference pattern between two wavefields, the reference and object wavefields. Holograms may be optically generated from real objects or computationally generated from synthetic objects leading to the so-called computer generated holograms. Whatever the creation process, holograms may have a digital representation using some appropriate representation format. Moreover, considering the huge amounts of data involved, digital holographic data has to be compressed using appropriate coding solutions, possibly available standard coding solutions. Finally, the decoded holographic data has to lead to reconstructed images using some appropriate reconstruction method. Holographic data coding is an emerging field of research where literature is still very scarce. In this context, before starting designing coding solutions considering the specific characteristics of holographic data, it is essential to assess the current compression performance associated to the most relevant available standard coding solutions and alternative representations formats. This paper has the main objective to benchmark the most relevant available standard coding solutions and the main alternative representation formats under meaningful test conditions and for the test material currently available. This assessment is critical to understand the current status quo and launch future research.
Jose Peixeiro, Catarina Brites, João Ascenso, Fernando Pereira 0001
ICME3
2015 Stereo Based Tracking-by-Detection for Visual Sensor Networks
abstract
Visual binary descriptors have successfully been employed in several applications such as visual search, object recognition and visual tracking. In particular, binary descriptors are suitable for scenarios where computational, storage and energy resources are constrained and have been previously exploited to track an object along a video sequence. In this paper, binary descriptors are used to perform visual tracking in a stereo-based system, i.e. when two cameras with overlapping views are employed in a cooperative way. The proposed stereo-based visual tracker follows the tracking-by-detection approach where features extracted from different cameras are used to characterize the object appearance with a suitable model. Moreover visual tracking is performed at a central controller by just using the features transmitted from the two camera nodes. To achieve this target, efficient coding techniques are proposed to reduce the amount of feature data that is transmitted through the network. The performance of the proposed stereo-based visual tracker is evaluated in terms of rate-accuracy, i.e. using quantitative metrics to assess the accuracy of the visual tracker as a function of the coding bitrate.
Berner Panti, Fernando Pereira 0001, João Ascenso
ISM3
2015 Epipolar plane image based rendering for 3D video coding
abstract
In current 3D video coding solutions, such as the 3D-HEVC standard, depth data is instrumental to have a continuum of views synthesized at the decoder based on a limited set of coded views. In order view synthesis may be performed at the decoder, depth data is currently directly acquired or estimated at the encoder based on very few neighboring views and transmitted to the decoder after appropriate compression. At the decoder, further views then those decoded are synthesized using again very few neighboring decoded views, thus using a local synthesis approach. A promising alternative synthesis approach may consider not a few but rather all the views available at the decoder, thus offering a scene global approach to synthesis. One way to implement this approach involves cutting the views cube along the viewpoint direction, creating the so-called epipolar plane images (EPI) which provide a rather compact representation of the scene. In this context, this paper proposes an EPI based view rendering framework for 3D video coding solution and identifies the major benefits of such framework, notably in comparison with the traditional local synthesis approach.
Catarina Brites, João Ascenso, Fernando Pereira 0001
MMSP2
2015 Improving enhancement layer merge mode for HEVC scalable extension
abstract
In the HEVC scalable extension (SHVC), the merge mode prediction plays an important role due to its high selection probability and its low associated bitrate. In the SHVC enhancement layer (EL) merge mode prediction, the motion information is selected from the merge candidates, in this case the spatial, temporal, and inter-layer candidates. Therefore, the merge mode prediction is usually inefficient when the motion vector field (MVF) correlation between the merge candidates and the current EL block is low, especially in video sequences containing high motion activity. To address this problem, this paper proposes an improved EL merge mode prediction solution that adaptively refines the motion vector candidates by using both base layer (BL) and EL decoded information and after linearly combines the motion compensated samples with the BL reconstructed samples to achieve a better merge mode prediction quality. Experiments conducted for a rich set of test conditions have shown that significant compression efficiency gains can be achieved with the proposed improved scalable coding solution, notably up to 4.57% in BD-rate savings regarding the standard SHVC quality scalable codec.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
PCS2
2015 HEVC backward compatible scalability: A low encoding complexity distributed video coding based approach
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
Signal Process. Image Commun.2
2014 Coding binary local features extracted from video sequences
abstract
Local features represent a powerful tool which is exploited in several applications such as visual search, object recognition and tracking, etc. In this context, binary descriptors provide an efficient alternative to real-valued descriptors, due to low computational complexity, limited memory footprint and fast matching algorithms. The descriptor consists of a binary vector, in which each bit is the result of a pairwise comparison between smoothed pixel intensities. In several cases, visual features need to be transmitted over a bandwidth-limited network. To this end, it is useful to compress the descriptor to reduce the required rate, while attaining a target accuracy for the task at hand. The past literature thoroughly addressed the problem of coding visual features extracted from still images and, only very recently, the problem of coding real-valued features (e.g., SIFT, SURF) extracted from video sequences. In this paper we propose a coding architecture specifically designed for binary local features extracted from video content. We exploit both spatial and temporal redundancy by means of intra-frame and inter-frame coding modes, showing that significant coding gains can be attained for a target level of accuracy of the visual analysis task.
Luca Baroffio, João Ascenso, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi
ICIP2
2014 Enabling visual analysis in wireless sensor networks
abstract
This demo showcases some of the results obtained by the GreenEyes project, whose main objective is to enable visual analysis on resource-constrained multimedia sensor networks. The demo features a multi-hop visual sensor network operated by BeagleBones Linux computers with IEEE 802.15.4 communication capabilities, and capable of recognizing and tracking objects according to two different visual paradigms. In the traditional compress-then-analyze (CTA) paradigm, JPEG compressed images are transmitted through the network from a camera node to a central controller, where the analysis takes place. In the alternative analyze-then-compress (ATC) paradigm, the camera node extracts and compresses local binary visual features from the acquired images (either locally or in a distributed fashion) and transmits them to the central controller, where they are used to perform object recognition/tracking. We show that, in a bandwidth constrained scenario, the latter paradigm allows to reach better results in terms of application frame rates, still ensuring excellent analysis performance.
Luca Baroffio, Antonio Canclini, Matteo Cesana, Alessandro Redondi, Marco Tagliasacchi, György Dán, Emil Eriksson, Viktoria Fodor, João Ascenso, Pedro Monteiro
ICIP9
2014 Local feature selection for efficient binary descriptor coding
abstract
In a visual sensor network, a large number of camera nodes are able to acquire and process image data locally, collaborate with other camera nodes and provide a description about the captured events. Typically, camera nodes have severe constraints in terms of energy, bandwidth resources and processing capabilities. Considering these unique characteristics, coding and transmission of the pixel-level representation of the visual scene must be avoided, due to the energy resources required. A promising approach is to extract at the camera nodes, compact visual features that are coded to meet the bandwidth and power requirements of the underlying network and devices. Since the total number of features extracted from an image may be rather significant, this paper proposes a novel method to select the most relevant features before the actual coding process. The solution relies on a score that estimates the accuracy of each local feature. Then, local features are ranked and only the most relevant features are coded and transmitted. The selected features must maximize the efficiency of the image analysis task but also minimize the required computational and transmission resources. Experimental results show that higher efficiency is achieved when compared to the previous state-of-the-art.
Pedro Monteiro, João Ascenso, Fernando Pereira 0001
ICIP2
2014 Correlation modeling for a distributed scalable video codec based on the HEVC standard
abstract
The growing heterogeneity of networks, devices and consumption conditions asks for flexible and adaptive video coding solutions. The compression power of the HEVC standard and the benefits of the distributed video coding paradigm allow designing novel scalable coding solutions with improved error robustness and low encoding complexity while still achieving competitive compression efficiency. In this context, this paper proposes a novel scalable video coding scheme using a HEVC Intra compliant base layer and a distributed coding approach in the enhancement layers (EL). This design inherits the HEVC compression efficiency while providing low encoding complexity at the enhancement layers. The temporal correlation is exploited at the decoder to create the EL side information (SI) residue, an estimation of the original residue. The EL encoder sends only the data that cannot be inferred at the decoder, thus exploiting the correlation between the original and SI residues; however, this correlation must be characterized with an accurate correlation model to obtain coding efficiency improvements. Therefore, this paper proposes a correlation modeling solution to be used at both encoder and decoder, without requiring a feedback channel. Experiments results confirm that the proposed scalable coding scheme has lower encoding complexity and provides BD-Rate savings up to 3.43% in comparison with the HEVC Intra scalable extension under development.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
MMSP2
2014 H.264/AVC backward compatible bit-depth scalable video coding
abstract
As high dynamic range video is gaining popularity, video coding solutions able to efficiently provide both low and high dynamic range video, notably with a single bitstream, are increasingly important. While simulcasting can provide both dynamic range videos at the cost of some compression efficiency penalty, bit-depth scalable video coding can provide a better trade-off between compression efficiency, adaptation flexibility and computational complexity. Considering the widespread use of H.264/AVC video, this paper proposes a H.264/AVC backward compatible bit-depth scalable video coding solution offering a low dynamic range base layer and two high dynamic range enhancement layers with different qualities, at low complexity. Experimental results show that the proposed solution has an acceptable rate-distortion performance penalty regarding the HDR H.264/AVC single-layer coding solution.
Vasco Nascimento, João Ascenso, Fernando Pereira 0001
MMSP2
2014 Statistical reconstruction for predictive video coding
abstract
Substantial rate-distortion (RD) gains have been achieved in video coding standards by increasing the encoder complexity while maintaining the decoder complexity the lowest possible. On the other hand, the alternative distributed video coding (DVC) approach proposes to exploit the video redundancy mostly at the decoder side, keeping the encoder as simple as possible. One of the most characteristic DVC tools is the statistical reconstruction of the DCT coefficients, which plays a similar role to the inverse scalar quantization (ISQ) in predictive codecs. The main objective of this paper is to propose a statistical reconstruction approach for predictive coding (notably the H.264/AVC standard) as a substitute to ISQ, thus creating a coding architecture with a mix of predictive and distributed coding tools. Experimental results show that the proposed statistical reconstruction solution allows achieving Bjontegaard bitrate savings up to 2.4% regarding the ISQ based H.264/AVC High profile codec.
Catarina Brites, Vitor Gomes 0003, João Ascenso, Fernando Pereira 0001
VCIP3
2014 Optimal reconstruction for a HEVC backward compatible distributed scalable video codec
abstract
In a landscape of heterogeneous networks, terminals and usage environments, a low encoding complexity scalable video coding engine is required for many emerging applications such as wireless video surveillance, visual sensor networks and remote space transmission. To fulfil this need, a distributed scalable video coding (DSVC) framework has been developed combining the predictive and distributed coding paradigms. The DSVC framework provides High Efficiency Video Coding (HEVC) backward compatibility at the base layer (BL) and adopts a distributed coding approach for the enhancement layers (ELs). In DSVC, the decoder reconstruction plays a key role as it strongly impacts the EL decoded frame quality and, thus, the final DSVC compression efficiency. In this context, this paper proposes a statistically inspired decoder reconstruction solution for the EL coded information based on the so-called side information residue and its correlation with the encoder EL residue. Experimental results confirm that significant RD performance gains can be achieved with the proposed statistical reconstruction technique, notably up to 11.28% BD-Rate savings regarding the most relevant benchmark.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
VCIP2
2014 Perceptually driven video error protection using a distributed source coding approach
André Seixas Dias, Catarina Brites, João Ascenso, Fernando Pereira 0001
Signal Process. Image Commun.3
2013 Rate-accuracy optimization of binary descriptors
abstract
Binary descriptors have recently emerged as low-complexity alternatives to state-of-the-art descriptors such as SIFT. The descriptor is represented by means of a binary string, in which each bit is the result of the pair-wise comparison of smoothed pixel values properly selected in a patch around each keypoint. Previous works have focused on the construction of the descriptor neglecting the opportunity of performing lossless compression. In this paper, we propose two contributions. First, design an entropy coding scheme that seeks the internal ordering of the descriptor that minimizes the number of bits necessary to represent it. Second, we compare different selection strategies that can be adopted to identify which pair-wise comparisons to use when building the descriptor. Unlike previous works, we evaluate the discriminative power of descriptors as a function of rate, in order to investigate the trade-offs in a bandwidth constrained scenario.
Alessandro Redondi, Luca Baroffio, João Ascenso, Matteo Cesana, Marco Tagliasacchi
ICIP3
2013 Improving scalable video coding performance with decoder side information
abstract
In a heterogeneous landscape of networks, devices and consumption environments, scalability is one of the most important video coding features. To achieve higher scalable video compression efficiency, this paper proposes a novel scalable video coding framework based on predictive video coding but also exploiting some additional decoder side information. The side information is estimated at both encoder and decoder using a motion compensated temporal interpolation technique, commonly used in distributed video coding solutions. To improve the B-slices compression efficiency, the side information independently created at each coding layer, notably the base and enhancement layers, is inserted in the corresponding layer decoded picture buffer to be exploited as an additional reference frame in the scalable predictive coding process. Experimental results have shown significant compression efficiency gains, notably up to around 3.5% in bitrate savings regarding the state-of-the-art SVC standard.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
PCS2
2013 Clustering based binary descriptor coding for efficient transmission in visual sensor networks
abstract
Nowadays, local feature descriptors have emerged as one of the most promising and powerful visual representation solutions. In fact, with a minimal amount of computational effort, the detection and extraction of visual features can provide reliable and a compact image representation that enables a rich set of image analysis tasks. In this paper, a visual sensor network scenario is considered with energy and bandwidth constraints at each sensor node. In this scenario, low-level binary features can be computed with low complexity and efficient coding schemes can reduce the data rate needed to transmit the features. This paper proposes two binary descriptor coding techniques that exploit the correlation between descriptors of the same image by clustering the extracted descriptors with two solutions: divisive-clustering and agglomerative-clustering. While the former starts with one cluster containing all descriptors which is recursively divided, the latter starts with as many clusters as descriptors which are recursively grouped. The disjoint sets of descriptors are differentially coded and the prediction residual is entropy coded. The experimental results show bitrate savings up to 34% without any impact in the accuracy of the final image retrieval task.
Pedro Monteiro, João Ascenso
PCS2
2013 Side information creation for efficient Wyner-Ziv video coding: Classifying and reviewing
Catarina Brites, João Ascenso, Fernando Pereira 0001
Signal Process. Image Commun.2
2012 Improving predictive video coding performance with decoder side information
abstract
Although major achievements have been reached in terms of video compression efficiency, additional gains are still needed to satisfy current and emerging applications needs. This trend justifies the continuous efforts to go beyond the compression capabilities of the state-of-the-art H.264/AVC standard. This paper proposes to further bridge two video coding approaches, the popular predictive and the emerging distributed coding paradigms, by proposing a novel bidirectional (B) coding process to be integrated in the H.264/AVC codec, inspired by the side information creation module, central in distributed video coding. The novel B-slice coding process builds on a new reference frame, the side information frame, which is both created at the encoder and decoder, as the SI creation method does not need the original to be available. The experimental results for the novel coding solution show average bitrate savings of 5.89% both for the low and high bitrate regions.
Xiem HoangVan, João Ascenso, Fernando Pereira 0001
ICIP2
2012 Learning based decoding approach for improved Wyner-Ziv video coding
abstract
Wyner-Ziv (WZ) video coding compression efficiency depends critically both on the side information (SI) quality and the correlation noise model (CNM) accuracy. In this context, this paper proposes a learning based decoding approach for transform domain WZ video coding, notably in the context of the following techniques: i) fractional-pixel motion field learning to define the relevance of the SI block candidates, and ii) CNM parameter learning. Experimental results show the proposed learning approach brings consistent RD performance improvements, with coding gains up to 3.9 dB regarding the state-of-the-art DISCOVER WZ video codec for a GOP size of 8.
Catarina Brites, João Ascenso, Fernando Pereira 0001
PCS2
2011 A denoising approach for iterative side information creation in distributed video coding
abstract
In distributed video coding, motion estimation is typically performed at the decoder to generate the side information, increasing the decoder complexity while providing low complexity encoding in comparison with predictive video coding. Motion estimation can be performed once to create the side information or several times to refine the side information quality along the decoding process. In this paper, motion estimation is performed at the decoder side to generate multiple side information hypotheses which are adaptively and dynamically combined, whenever additional decoded information is available. The proposed iterative side information creation algorithm is inspired in video denoising filters and requires some statistics of the virtual channel between each side information hypothesis and the original data. With the proposed denoising algorithm for side information creation, a RD performance gain up to 1.2 dB is obtained for the same bitrate.
João Ascenso, Catarina Brites, Fernando Pereira 0001
ICIP1
2011 Low complexity deblocking filter perceptual optimization for the HEVC codec
abstract
The compression efficiency of the state-of-art H.264/AVC video coding standard must be improved to accommodate the compression needs of high definition videos. To this end, ITU and MPEG started a new standardization project called High Efficiency Video Coding. The video codec under development still relies on trans form domain quantization and includes the same in-loop deblocking filter adopted in the H.264/AVC standard to reduce quantization blocking artifacts. This deblocking filter provides two offsets to vary the amount of filtering for each image area. This paper proposes a perceptual optimization of these offsets based on a quality metric able to quantify the blocking artifacts impact on the perceived video quality. The proposed optimization involves low computational complexity and provides quality improvements with respect to a non-perceptually optimized H.264/AVC deblocking filter. Moreover, the proposed optimization allows up to 92% of complexity reduction regarding a brute force perceptual optimization which exhaustively tests all the possible offsets values.
Matteo Naccari, Catarina Brites, João Ascenso, Fernando Pereira 0001
ICIP3
2011 Augmented LDPC graph for distributed video coding with multiple side information
abstract
The advances made in channel-capacity codes, such as turbo codes and low-density parity-check (LDPC) codes, have played a major role in the emerging distributed source coding paradigm. LDPC codes can be easily adapted to new source coding strategies due to their natural representation as bipartite graphs and the use of quasi-optimal decoding algorithms, such as belief propagation. This paper tackles a relevant scenario in distributed video coding: lossy source coding when multiple side information (SI) hypotheses are available at the decoder, each one correlated with the source according to different correlation noise channels. Thus, it is proposed to exploit multiple SI hypotheses through an efficient joint decoding technique with multiple LDPC syndrome decoders that exchange information to obtain coding efficiency improvements. At the decoder side, the multiple SI hypotheses are created with motion compensated frame interpolation and fused together in a novel iterative LDPC based Slepian-Wolf decoding algorithm. With the creation of multiple SI hypotheses and the proposed decoding algorithm, bitrate savings up to 8.0% are obtained for similar decoded quality.
João Ascenso, Catarina Brites, Fernando Pereira 0001
MMSP1
2010 A flexible side information generation framework for distributed video coding
João Ascenso, Catarina Brites, Fernando Pereira 0001
Multim. Tools Appl.1
2009 Low complexity intra mode selection for efficient distributed video coding
abstract
Motion compensated frame interpolation (MCFI) is one of the most efficient solutions to generate side information (SI) in the context of distributed video coding. However, it creates SI with rather significant motion compensated errors for some frame regions while rather small for some other regions depending on the video content. In this paper, a low complexity intra mode selection algorithm is proposed to select the most dasiacriticalpsila blocks in the WZ frame and help the decoder with some reliable data for those blocks. For each block, the novel coding mode selection algorithm estimates the encoding rate for the intra based and WZ coding modes and determines the best coding mode while maintaining a low encoder complexity. The proposed solution is evaluated in terms of rate-distortion performance with improvements up to 1.2 dB regarding a WZ coding mode only solution.
João Ascenso, Fernando Pereira 0001
ICME1
2009 Distributed Video Coding with multiple side information
abstract
Distributed Video Coding (DVC) is a new video coding paradigm which mainly exploits the source statistics at the decoder based on the availability of some decoder side information. The quality of the side information has a major impact on the DVC rate-distortion (RD) performance in the same way the quality of the predictions had a major impact in predictive video coding. In this paper, a DVC solution exploiting multiple side information is proposed; the multiple side information is generated by frame interpolation and frame extrapolation targeting to improve the side information of a single estimation mode. Compared with the best available single side information solutions, the proposed DVC solution with multiple side information robustly improves the RD performance for the set of test sequences.
Xin Huang 0004, Catarina Brites, João Ascenso, Fernando Pereira 0001, Søren Forchhammer
PCS3
2009 Adaptive deblocking filter for transform domain Wyner-Ziv video coding
abstract
Wyner–Ziv (WZ) video coding is a particular case of distributed video coding, the recent video coding paradigm based on the Slepian–Wolf and Wyner–Ziv theorems that exploits the source correlation at the decoder and not at the encoder as in predictive video coding. Although many improvements have been done over the last years, the performance of the state-of-the-art WZ video codecs still did not reach the performance of state-of-the-art predictive video codecs, especially for high and complex motion video content. This is also true in terms of subjective image quality mainly because of a considerable amount of blocking artefacts present in the decoded WZ video frames. This paper proposes an adaptive deblocking filter to improve both the subjective and objective qualities of the WZ frames in a transform domain WZ video codec. The proposed filter is an adaptation of the advanced deblocking filter defined in the H.264/AVC (advanced video coding) standard to a WZ video codec. The results obtained confirm the subjective quality improvement and objective quality gains that can go up to 0.63 dB in the overall for sequences with high motion content when large group of pictures are used.
Catarina Brites, João Ascenso, Fernando Pereira 0001
IET Image Process.3
2009 Refining Side Information for Improved Transform Domain Wyner-Ziv Video Coding
abstract
Wyner-Ziv (WZ) video coding is a particular case of distributed video coding, which is a recent video coding paradigm based on the Slepian-Wolf and WZ theorems. Contrary to available prediction-based standard video codecs, WZ video coding exploits the source statistics at the decoder, allowing the development of simpler encoders. Until now, WZ video coding did not reach the compression efficiency performance of conventional video coding solutions, mainly due to the poor quality of the side information, which is an estimate of the original frame created at the decoder in the most popular WZ video codecs. In this context, this paper proposes a novel side information refinement (SIR) algorithm for a transform domain WZ video codec based on a learning approach where the side information is successively improved as the decoding proceeds. The results show significant and consistent performance improvements regarding state-of-the-art WZ and standard video codecs, especially under critical conditions such as high motion content and long group of pictures sizes.
Catarina Brites, João Ascenso, Fernando Pereira 0001
IEEE Trans. Circuits Syst. Video Technol.3
2008 Design and performance of a novel low-density parity-check code for distributed video coding
abstract
Low-density parity-check (LDPC) codes are nowadays one of the hottest topics in coding theory, notably due to their advantages in terms of bit error rate performance and low complexity. In order to exploit the potential of the Wyner-Ziv coding paradigm, practical distributed video coding (DVC) schemes should use powerful error correcting codes with near-capacity performance. In this paper, new ways to design LDPC codes for the DVC paradigm are proposed and studied. The new LDPC solutions rely on merging parity-check nodes, which corresponds to reduce the number of rows in the parity-check matrix. This allows to change gracefully the compression ratio of the source (DCT coefficient bitplane) according to the correlation between the original and the side information. The proposed LDPC codes reach a good performance for a wide range of source correlations and achieve a better RD performance when compared to the popular turbo codes.
João Ascenso, Catarina Brites, Fernando Pereira 0001
ICIP1
2008 Wyner-Ziv video coding: A review of the early architectures and further developments
abstract
In 2002, the video coding community faced the emergence of a new video coding paradigm, the so-called Wyner-Ziv video coding, which was represented by two early solutions designed by the Stanford University and the University of California, Berkeley research teams. This paper intends to briefly review, and compare these two early Wyner-Ziv video coding solutions, notably from the functional point of view. Moreover, this paper reviews some important developments of the Stanford Wyner-Ziv coding architecture, which has become the most popular in the literature.
Fernando Pereira 0001, Catarina Brites, João Ascenso, Marco Tagliasacchi
ICME3
2008 Advanced side information creation techniques and framework for Wyner-Ziv video coding
João Ascenso, Fernando Pereira 0001
J. Vis. Commun. Image Represent.1
2008 Evaluating a feedback channel based transform domain Wyner-Ziv video codec
Catarina Brites, João Ascenso, José Quintas Pedro, Fernando Pereira 0001
Signal Process. Image Commun.2
2007 Adaptive Hash-Based Side Information Exploitation for Efficient Wyner-Ziv Video Coding
abstract
Wyner-Ziv video coding is a lossy source coding paradigm where the video statistics are exploited, partially or totally at the decoder. The side information represents a noisy version of the original frame and is generated at the decoder with time consuming motion estimation and compensation tools. This paper proposes a novel bidirectional hash motion estimation framework which enables the decoder to choose between past and/or future reference frames for frame interpolation. New features include the coding of DCT hash with zero-motion, combination of trajectory-based motion interpolation with hash-based motion estimation and adaptive selection of the DCT bands which are sent to the decoder in order to guide the motion estimation procedure. Gains up to 1.2 dB compared to previous motion interpolation approaches may be reached.
João Ascenso, Fernando Pereira 0001
ICIP (3)1
2007 Wyner-Ziv Stereo Video Coding using a Side Information Fusion Approach
abstract
Wyner-Ziv coding, also known as distributed video coding, is currently a very hot research topic in video coding due to the new opportunities it opens. This paper applies the distributed video coding principles to stereo video coding, to propose a practical solution for Wyner-Ziv stereo coding based on mask-based fusion of temporal and spatial side informations. The architecture includes a low-complexity encoder and avoids any communication between the cameras/encoders. While the rate-distortion (RD) performance strongly depends on the motion-based frame interpolation (MBFI) and disparity-based frame estimation (DBFE) solutions, first results show that the proposed approach is promising and there are still issues to address.
José Diogo Areia, João Ascenso, Catarina Brites, Fernando Pereira 0001
MMSP2
2007 Studying the GOP Size Impact on the Performance of a Feedback Channel-Based Wyner-Ziv Video Codec
Fernando Pereira 0001, João Ascenso, Catarina Brites
PSIVT2
2006 Improving Transform Domain Wyner-Ziv Video Coding Performance
abstract
Distributed video coding (DVC) is a new video coding paradigm based on two key information theory results: the Slepian-Wolf and Wyner-Ziv theorems. A particular case of DVC, the so-called Wyner-Ziv coding, deals with lossy source coding with side information at the decoder and enables a flexible allocation of complexity between the encoder and the decoder. This paper proposes an improved transform domain Wyner-Ziv video codec including: 1) the integer block-based transform defined in the H.264/MPEG-4 AVC standard, 2) a quantizer with a symmetrical interval around zero for AC coefficients, and a quantization step size adjusted to the transform coefficient bands dynamic range, and 3) advanced frame interpolation for side information generation. The combination of these tools brings significant rate-distortion (RD) gains regarding the state-of-the-art results available in the literature
Catarina Brites, João Ascenso, Fernando Pereira 0001
ICASSP (2)2
2006 Intra Mode Decision Based on Spatio-Temporal Cues in Pixel Domain Wyner-ZIV Video Coding
abstract
Distributed source coding principles have been recently applied to video coding in order to achieve a flexible distribution of the complexity burden between the encoder and the decoder. In this paper we elaborate on a pixel based Wyner-Ziv video codec that shifts all the complexity of the motion estimation phase to the decoder, thus achieving light encoding. We observe that the correlation noise statistics describing the relationship between the frame to be encoded and the side information available at the decoder is not spatially stationary. For this reason we introduce a mode decision scheme either at the encoder or at the decoder in such a way that when the estimated correlation is weak we opt for intra coding on a block-by-block basis. Both spatial and temporal criteria are used to determine whether a block is better intra coded or not
Marco Tagliasacchi, Alan Trapanese, Stefano Tubaro, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICASSP (2)4
2006 Content Adaptive Wyner-ZIV Video Coding Driven by Motion Activity
abstract
In distributed video coding (DVC), the video statistics are exploited, partially or totally at the decoder. A particular case of DVC, Wyner-Ziv video coding deals with lossy source coding with side information at the decoder and allows moving part or the entire motion estimation task to the decoder. In this context, it is the decoder responsibility to obtain the side information, a guess of the encoded Wyner-Ziv frame and the encoder only sends parity bits to improve its quality. In this paper, a technique targeting the improvement of the quality of the side information, and thus of the rate-distortion performance of the Wyner-Ziv codec is proposed. This is achieved by adaptively adjusting the size of the motion interpolation structure (or GOP length) according to the motion activity along the sequence. Experimentally, this allows to achieve gains up to 0.8 dB without performing any motion estimation or complex mode decision at the encoder.
João Ascenso, Catarina Brites, Fernando Pereira 0001
ICIP1
2006 Studying Temporal Correlation Noise Modeling for Pixel Based Wyner-Ziv Video Coding
abstract
Wyner-Ziv (WZ) video coding-a particular case of distributed video coding (DVC)-is a new video coding paradigm based on two major information theory results: the Slepian-Wolf and Wyner-Ziv theorems. Recently, practical WZ video coding solutions were proposed with promising results. Most of the solutions available in the literature, model the correlation noise between the original frame and the so-called side information by a given distribution whose relevant parameters are estimated in an offline process, at the encoder. In this paper, three algorithms are proposed towards a more realistic WZ coding approach by performing online estimation of the error distribution at the decoder. Both algorithms explore temporal correlation between frames however with different levels of granularity: frame, block and pixel levels; better rate-distortion (RD) performance is achieved for lower granularity (pixel) level.
Catarina Brites, João Ascenso, Fernando Pereira 0001
ICIP2
2006 Exploiting Spatial Redundancy in Pixel Domain Wyner-Ziv Video Coding
abstract
Distributed video coding is a recent paradigm that enables a flexible distribution of the computational complexity between the encoder and the decoder building on top of distributed source coding principles. In this paper we focus on the scenario where most of the complexity is shifted to the decoder, thus achieving light encoding. We elaborate on a well known pixel based Wyner-Ziv architecture and we improve its coding efficiency by exploiting both spatial and temporal correlation at the decoder side, without the need of performing any transform at the encoder. In order to generate the side information, the decoder adaptively chooses spatial or temporal information, based on the local estimate of the correlation noise. Simulations on test sequences demonstrate that a coding gain of up to +1.8 dB can be obtained with respect to the case that generates the side information by motion interpolation only.
Marco Tagliasacchi, Alan Trapanese, Stefano Tubaro, João Ascenso, Catarina Brites, Fernando Pereira 0001
ICIP4
2006 Hybrid Distributed Video Coding Using SCA Codes
abstract
We describe the architecture for our distributed video coding (DVC) system. Some key differences between our work and previous systems include a new method of enabling decoder motion compensation, and the use of serially concatenated accumulate syndrome codes for distributed source coding. To evaluate performance, we compare our system to the H.263+ and H.264/AVC video codecs. Experiments show that our system is comparable to DVC systems from Stanford and Berkeley in the sense that our system performs better than H.263+Intra, but worse than H.263+Inter and H.264/AVC
Emin Martinian, Anthony Vetro, Jonathan S. Yedidia, João Ascenso, Ashish Khisti, Dmitry Malioutov
MMSP4
2005 Motion compensated refinement for low complexity pixel based distributed video coding
abstract
Distributed video coding (DVC) is a new coding paradigm that enables to exploit video statistics, partially or totally at the decoder. A particular case of DVC, Wyner-Ziv coding, deals with lossy source coding with side information at the decoder and allows a shift of complexity from the encoder to the decoder, theoretically without any penalty in the coding efficiency. The Wyner-Ziv solution here described encodes each video frame independently (intraframe coding), but decodes the same frame conditionally (interframe decoding). At the decoder, and compensation tools are responsible to obtain an accurate interpolation of the original frame using previously decoded (temporally adjacent) frames. This paper proposes a novel approach to improve the performance of pixel domain Wyner-Ziv video coding by using a motion compensated refinement of the decoded frame and use it as improved side information. More precisely, upon partial decoding of each frame, the decoder refines its motion trajectories in order to achieve a better reconstruction of the decoded frame.
João Ascenso, Catarina Brites, Fernando Pereira 0001
AVSS1
2004 Drift reduction for a H.264/AVC fine grain scalability with motion compensation architecture
abstract
The recent advances in nonscalable video encoding brought by the H.264-AVC standard offered significant improvements in terms of rate-distortion performance. This paper proposes a H.264-AVC based fine grain scalable video encoder which also exploits the motion compensation tools of the H.264-AVC standard to explore the temporal redundancy in the enhancement layer. The enhancement layer is predicted from a high quality reference obtained from past information of the enhancement and base layers. One of the drawbacks of this architecture is the drift effect, which occurs when part of the enhancement layer used for prediction is not received by the decoder. The drift reduction approaches here proposed simultaneously allow improvements in the coding efficiency and a reduction of the drift effect. The experimental results show improvements up to 2 dB in coding efficiency in comparison to Intra coding (like used by the MPEG-4 FGS standard) using the MPEG-4 testing conditions.
João Ascenso, Fernando Pereira 0001
ICIP1