VLDB 2026 Research / reviewers in the wild / expert
Catarina Brites
dblp:15/6844
· DBLP profile ↗
52ranked-venue papers
16as first author
7since 2021 · last 2025
0000-0002-6011-4574ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 16 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Attention-Enhanced Multi-Branch Spiking Neural Network for Event Stream Super-ResolutionabstractTraditional visual sensors capture images by sampling light at fixed intervals, producing a sequence of frames. In contrast, event vision sensors detect changes in light intensity asynchronously at the pixel level and generate discrete events with precise timing information, allowing to capture challenging scenes with high-speed object motion and extreme lighting conditions accurately. This paper introduces an attention-enhanced, polarity-aware multi-branch Spiking Neural Network (SNN) that directly super-resolves low-resolution event streams while preserving their temporal accuracy. The proposed architecture uses two parallel branches with novel spike-based spatial and temporal attention modules that give more importance to salient spatio-temporal structures, yet remaining fully asynchronous and hardware-friendly. Experimental results demonstrate that the proposed spike-based attention with multi-branch SNN significantly outperforms existing state-of-the-art methods across all evaluation metrics, while effectively maintaining the underlying spatio-temporal characteristics of the event stream. Ahmadreza Sezavar, Catarina Brites, João Ascenso |
ISM | 2 |
| 2024 | Low Complexity Learning-based Lossless Event-based CompressionabstractEvent cameras are a cutting-edge type of visual sensors that capture data by detecting brightness changes at the pixel level asynchronously. These cameras offer numerous benefits over conventional cameras, including high temporal resolution, wide dynamic range, low latency, and lower power consumption. However, the substantial data rates they produce require efficient compression techniques, while also fulfilling other typical application requirements, such as the ability to respond to visual changes in real-time or near real-time. Additionally, many event-based applications demand high accuracy, making lossless coding desirable, as it retains the full detail of the sensor data. Learningbased methods show great potential due to their ability to model the unique characteristics of event data thus allowing to achieve high compression rates. This paper proposes a low-complexity lossless coding solution based on the quadtree representation that outperforms traditional compression algorithms in efficiency and speed, ensuring low computational complexity and minimal delay for real-time applications. Experimental results show that the proposed method delivers better compression ratios, i.e., with fewer bits per event, and lower computational complexity compared to current lossless data compression methods. Ahmadreza Sezavar, Catarina Brites, João Ascenso |
ISM | 2 |
| 2024 | Learning-based Lossless Event Data CompressionabstractEmerging event cameras acquire visual information by detecting time domain brightness changes asynchronously at the pixel level and, unlike conventional cameras, are able to provide high temporal resolution, very high dynamic range, low latency, and low power consumption. Considering the huge amount of data involved, efficient compression solutions are very much needed. In this context, this paper presents a novel deep-learning-based lossless event data compression scheme based on octree partitioning and a learned hyperprior model. The proposed method arranges the event stream as a 3D volume and employs an octree structure for adaptive partitioning. A deep neural network-based entropy model, using a hyperprior, is then applied. Experimental results demonstrate that the proposed method outperforms traditional lossless data compression techniques in terms of compression ratio and bits per event. Ahmadreza Sezavar, Catarina Brites, João Ascenso |
VCIP | 2 |
| 2021 | Image Coding with Neural Network-Based ColorizationabstractAutomatic colorization is a process with the objective of inferring the color of grayscale images. This process is frequently used for artistic purposes and to restore the color in old or damaged images. Motivated by the excellent results obtained with deep learning-based solutions in the area of automatic colorization, this paper proposes an image coding solution integrating a deep learning-based colorization process to estimate the chrominance components based on the decoded luminance which is regularly encoded with a conventional image coding standard. In this case, the chrominance components are not coded and transmitted as usual, notably after some subsampling, as only some color hints, i.e. chrominance values for specific pixel locations, may be sent to the decoder to help it creating more accurate colorizations. To boost the colorization and final compression performance, intelligent ways to select the color hints are proposed. Experimental results show performance improvements with the increased level of intelligence in the color hints extraction process and a good subjective quality of the final decoded (and colorized) images. Diogo Lopes, João Ascenso, Catarina Brites, Fernando Pereira 0001 |
ICASSP | 3 |
| 2021 | A Point-to-Distribution Joint Geometry and Color Metric for Point Cloud Quality AssessmentabstractPoint clouds (PCs) are a powerful 3D visual representation paradigm for many emerging application domains, especially virtual and augmented reality, and autonomous vehicles. However, the large amount of PC data required for highly immersive and realistic experiences requires the availability of efficient, lossy PC coding solutions are critical. Recently, two MPEG PC coding standards have been developed to address the relevant application requirements and further developments are expected in the future. In this context, the assessment of PC quality, notably for decoded PCs, is critical and asks for the design of efficient objective PC quality metrics. In this paper, a novel point-to-distribution metric is proposed for PC quality assessment considering both the geometry and texture. This new quality metric exploits the scale-invariance property of the Mahalanobis distance to assess first the geometry and color point-to-distribution distortions, which are after fused to obtain a joint geometry and color quality metric. The proposed quality metric significantly outperforms the best PC quality assessment metrics in the literature. Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso |
MMSP | 2 |
| 2021 | Lenslet Light Field Image Coding: Classifying, Reviewing and EvaluatingabstractIn recent years, visual sensors have been quickly improving, notably targeting richer acquisitions of the light present in a visual scene. In this context, the so-called lenslet light field (LLF) cameras are able to go beyond the conventional 2D visual acquisition models, by enriching the visual representation with directional light measures for each pixel position. LLF imaging is associated to large amounts of data, thus critically demanding efficient coding solutions in order applications involving transmission and storage may be deployed. For this reason, considerable research efforts have been invested in recent years in developing increasingly efficient LLF imaging coding (LLFIC) solutions. In this context, the main objective of this paper is to review and evaluate some of the most relevant LLFIC solutions in the literature, guided by a novel classification taxonomy, which allows better organizing this field. In this way, more solid conclusions can be drawn about the current LLFIC status quo, thus allowing to better drive future research and standardization developments in this technical area. Catarina Brites, João Ascenso, Fernando Pereira 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Point Cloud Rendering After Coding: Impacts on Subjective and Objective QualityabstractRecently, point clouds have shown to be a promising way to represent 3D visual data for a wide range of immersive applications, from augmented reality to autonomous cars. Emerging imaging sensors have made easier to perform richer and denser point cloud acquisition, notably with millions of points, thus raising the need for efficient point cloud coding solutions. In such scenario, it is important to evaluate the impact and performance of several processing steps in a point cloud communication system, notably the degradations associated to point cloud coding solutions. Moreover, since point clouds are not directly visualized but rather processed with a rendering algorithm before shown on any display, the perceived quality of point cloud data highly depends on the rendering solution. In this context, the main objective of this paper is to study the impact of several coding and rendering solutions on the perceived user quality and in the performance of available objective assessment metrics. Another contribution regards the assessment of recent MPEG point cloud coding solutions for several popular rendering methods, which was never presented before. The conclusions regard the visibility of three types of coding artifacts for the three considered rendering approaches as well as the strengths and weaknesses of objective metrics when point clouds are rendered after coding. Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso |
IEEE Trans. Multim. | 2 |
| 2020 | Improving Psnr-Based Quality Metrics Performance For Point Cloud GeometryabstractAn increased interest in immersive applications has drawn attention to emerging 3D imaging representation formats, notably light fields and point clouds (PCs). Nowadays, PCs are one of the most popular 3D media formats, due to recent developments in PC acquisition, namely with new depth sensors and signal processing algorithms. To obtain high fidelity 3D representations of visual scenes a huge amount of PC data is typically acquired, which demands efficient compression solutions. As in 2D media formats, the final perceived PC quality plays an importance role in the overall user experience and, thus, objective metrics capable to measure the PC quality in a reliable way are essential. In this context, this paper proposes and evaluates a set of objective quality metrics for the geometry component of PC data, which plays a very important role on the final perceived quality. Based on the popular PSNR PC geometry quality metric, novel improved PSNR-based metrics are proposed by exploiting the intrinsic PC characteristics and the rendering process that must occur before visualization. The experimental results show the superiority of the best proposed metrics over state-of-the-art, obtaining an improvement up to 32% in the Pearson correlation coefficient. Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso |
ICIP | 2 |
| 2020 | A Generalized Hausdorff Distance Based Quality Metric for Point Cloud GeometryabstractReliable quality assessment of decoded point cloud geometry is essential to evaluate the compression performance of emerging point cloud coding solutions and guarantee some target quality of experience. This paper proposes a novel point cloud geometry quality assessment metric based on a generalization of the Hausdorff distance. To achieve this goal, the so-called generalized Hausdorff distance for multiple rankings is exploited to identify the best performing quality metric in terms of correlation with the MOS scores obtained from a subjective test campaign. The experimental results show that the quality metric derived from the classical Hausdorff distance leads to low objective-subjective correlation and, thus, fails to accurately evaluate the quality of decoded point clouds for emerging codecs. However, the quality metric derived from the generalized Hausdorff distance with an appropriately selected ranking, outperforms the MPEG adopted geometry quality metrics when decoded point clouds with different types of coding distortions are considered. Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso |
QoMEX | 2 |
| 2020 | Point cloud coding: A privileged view driven by a classification taxonomy
Fernando Pereira 0001, Antoine Dricot, João Ascenso, Catarina Brites |
Signal Process. Image Commun. | 4 |
| 2020 | Mahalanobis Based Point to Distribution Metric for Point Cloud Geometry Quality EvaluationabstractNowadays, point clouds (PCs) are a promising representation format for immersive content and target several emerging applications, notably in virtual and augmented reality. However, efficient coding solutions are critically needed due to the large amount of PC data required for high quality user experiences. To address these needs, several PC coding standards were developed and thus, objective PC quality metrics able to accurately account for the subjective impact of coding artifacts are needed. In this paper, a scale-invariant PC geometry quality assessment metric is proposed based on a new type of correspondence, namely between a point and a distribution of points. This metric is able to reliably measure the geometry quality for PCs with different intrinsic characteristics and degraded by several coding solutions. Experimental results show the superiority of the proposed PC quality metric over relevant state-of-the-art. Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso |
IEEE Signal Process. Lett. | 2 |
| 2019 | Adaptive Plane Projection for Video-Based Point Cloud CodingabstractOne of the most promising emerging 3D representation paradigms is the point cloud (PC) model, notably due to the new set of applications that it enables, from immersive telepresence to 3D geographic information systems. Recognizing the potential of this representation model, MPEG has launched a standardization project to specify efficient PC coding solutions. This project has led to the so-called Video-based Point Cloud Coding (V-PCC) standard, which is based on the idea of projecting the dynamic PC geometry and texture into a sequence of frames to be coded with the highly efficient HEVC video coding standard. In V-PCC, the projection of points and color attributes is always performed using the same, rigid set of projection planes, independently of the PC characteristics. This paper proposes a more flexible coding solution, which improves the V-PCC Intra coding mode by adopting a content dependent set of projection planes, thus, more adapted to the characteristics of the PCs to be coded, targeting a better compression performance. The experimental results show an average total bitrate saving of around 15% regarding V-PCC with clear benefits in terms of subjective assessment. Eurico Lopes, João Ascenso, Catarina Brites, Fernando Pereira 0001 |
ICME | 3 |
| 2019 | Improved Patch Packing for the MPEG V-PCC StandardabstractPoint cloud representation is an emerging visual data technology, targeting immersive 3D experiences in the context of multiple applications scenarios, notably entertainment, geographical information systems, medicine, architecture, and robotics. Since a point cloud may easily involve millions of points, and thus an enormous amount of data, its effective storage and transmission critically asks for efficient coding solutions. With this purpose in mind, several point cloud coding (PCC) solutions have been proposed in the literature; special emphasis is due to the recent MPEG standards, which target interoperability in this domain, notably the MPEG V-PCC standard. The objective of this paper is to improve the V-PCC standard compression efficiency by proposing novel solutions for the V-PCC packing module without compromising in any way the V-PCC stream (syntax and semantics) and decoder compliance. In this context, several patch packing solutions are proposed, including new packing algorithms and associated sorting and positioning metrics; for the metrics, both absolute and relative approaches are proposed. The RD performance results show BD-Rate savings up to 0.8% for the best packing solution regarding the V-PCC benchmark. Moreover, the packing map size reductions can go up to 12%, on average. Afonso Costa, Antoine Dricot, Catarina Brites, João Ascenso, Fernando Pereira 0001 |
MMSP | 3 |
| 2019 | Graph-Based Static 3D Point Clouds Geometry CodingabstractRecently, 3D visual representation models such as light fields and point clouds are becoming popular due to their capability to represent the real world in a more complete and immersive way, paving the road for new and more advanced visual experiences. The point cloud representation model is able to efficiently represent the surface of objects/scenes by means of a set of 3D points and associated attributes and is increasingly being used from autonomous cars to augmented reality. Emerging imaging sensors have made it easier to perform richer and denser point cloud acquisitions, notably with millions of points, making it impossible to store and transmit these very high amounts of data without appropriate coding. This bottleneck has raised the need for efficient point cloud coding solutions in order to offer more immersive visual experiences and better quality of experience to the users. In this context, this paper proposes an efficient lossy coding solution for the geometry of static point clouds. The proposed coding solution uses an octree-based approach for a base layer and a graph-based transform approach for the enhancement layer where an Inter-layer residual is coded. The performance assessment shows very significant compression gains regarding the state-of-the-art, especially for the most relevant lower and medium rates. Paulo de Oliveira Rente, Catarina Brites, João Ascenso, Fernando Pereira 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Holographic Data Coding: Benchmarking and Extending HEVC With Adapted TransformsabstractHolography is an emerging technology to represent and display visual information with high expectations in terms of user experience. A hologram is a reproduction of a light field represented through the interference pattern between two wavefields the reference and the object wavefields. Whatever their creation process holograms may have a digital representation using some appropriate format. Moreover considering the huge amounts of data involved digital holographic data have to be compressed using appropriate coding solutions for example available image coding standard solutions or efficient extensions of them. In this context this paper contributes to advance the state-of-the-art on holographic data coding by: 1) benchmarking the most relevant available image coding standard solutions when using the most relevant holographic data representation formats; 2) proposing a novel mode depend directional transform-based HEVC coding solution trained with holographic data. Experimental results obtained under meaningful test conditions show that the proposed coding solution outperforms the state-of-the-art HEVC coding standard for specific formats and conditions. Altogether these two contributions are critical to understand the current status quo and advance the state-of-the-art on holographic data coding. Jose Peixeiro, Catarina Brites, João Ascenso, Fernando Pereira 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Subjective and objective quality evaluation of compressed point cloudsabstractThe increasing availability of point cloud data in recent years is demanding high performance compression solutions. Naturally, methods to perform objective quality assessment of compressed point clouds are also very much needed, namely metrics to measure the geometry distortion of point clouds when positioning errors are present. This is a rather challenging problem since this 3D representation format is unstructured and it is typically not directly visualized. In this context, the objective of this paper is to perform subjective and objective quality assessment of point clouds degraded by compression artifacts and to evaluate the correlation of the most popular objective quality metrics with human perception. In this work, subjective experiments conducted at Instituto Superior Técnico (IST) are described with point clouds compressed with two different but yet promising solutions, one based on the octree representation of the 3D space and another based on the rather popular graph transform. As far as the authors know, this is the first study of this type made available and should have a key role on the future development and evaluation of point cloud coding solutions.1 Alireza Javaheri, Catarina Brites, Fernando Pereira 0001, João Ascenso |
MMSP | 2 |
| 2017 | Epipolar based light field key-location detectorabstractNowadays, visual features play a key role, as they can provide a concise representation of visual data that is efficient for multiple tasks, notably content retrieval and object recognition. In parallel, visual sensors have been improving, targeting richer acquisitions of the light in a visual scene. In this context, the so-called light field cameras, which have recently emerged, are able to go beyond the standard acquisition models, by enriching the visual representation with directional light measures for each pixel position, e.g. by using a so-called lenslet light field camera. At this stage, not much research has been made in the field of feature detection and description for the emerging lenslet light field format. In this context, this paper proposes a feature detector suitable for lenslet light field images based on the exploitation of an alternative visual parametrization of the light field, called the Epipolar Planar Image (EPI). The proposed detector is heavily based on line detection in the EPI representation, since 3D points in the visual scene are mapped to line segments in an EPI, and the detector output is referred as key-locations. The proposed light field key-location detector is assessed with a solid evaluation framework using a large light field dataset. In comparison to the 2D SIFT detector, up to 10% improvements were achieved for the widely used repeatability metric. José Abecasis Teixeira, Catarina Brites, Fernando Pereira 0001, João Ascenso |
MMSP | 2 |
| 2017 | Saliency-driven omnidirectional imaging adaptive coding: Modeling and assessmentabstractOmnidirectional imaging, also known as 360° and spherical imaging, records all 360° of a scene from a specific spatial position, thus offering the user the capability to enjoy three rotational degrees of freedom (3-DoF). To offer a good quality of experience, omnidirectional imaging requires very high bitrates as high spatial resolution are a must and, ideally, also high frame rates. Due to the lack of video coding solutions specifically designed for omnidirectional imaging, this type of content is typically coded with the available image and video coding standards, such as JPEG, H.264/AVC and HEVC, after applying a 2D rectangular projection. In this context, this paper proposes an omnidirectional imaging coding solution allowing to reach improved coding performance by using an adaptive coding solution where the most visually salient image/video regions are coded with higher quality in a process appropriately controlled by the quantization parameter. To determine the saliency of the various omnidirectional imaging regions, a machine-learning based saliency detection model is proposed. The proposed coding solution achieves compression gains as measured by a novel objective quality metric also driven by saliency. This novel objective quality metric is validated by formal subjective testing where very high correlations with the subjective tests scores are achieved. Guilherme Luz Tortorella, João Ascenso, Catarina Brites, Fernando Pereira 0001 |
MMSP | 3 |
| 2016 | Multi-view distributed source coding of binary features for visual sensor networksabstractVisual analysis algorithms have been mostly developed for a centralized scenario where all visual data is acquired and processed at a central location. However, in visual sensor networks (VSN), several constraints in computational power, energy and bandwidth require a radically different approach, notably a paradigm shift from centralized to distributed visual processing. In the new paradigm, visual data is acquired and features are extracted at the sensing nodes locations to be after transmitted to enable further analysis at some central location. In such scenario, one of the key challenges is to design suitable feature coding schemes that are able to exploit the correlation among the features corresponding to (partially) overlapped views of the same visual scene. To achieve efficient coding, it is proposed to employ the distributed source coding paradigm as it does not require any communication between the sensing nodes (rather expensive in VSN) and it is parsimonious in terms of computational resources. Experimental results show that significant accuracy and compression gains (up to 37.36%) can be achieved when coding features extracted from multiple views. Nuno Monteiro, Catarina Brites, Fernando Pereira 0001, João Ascenso |
ICASSP | 2 |
| 2016 | Multi-view distributed coding and selection of local binary featuresabstractRecently, the latest advances in compact feature representation and feature learning have provided an efficient framework for several visual analysis tasks, such as object recognition. However, when multiple cameras with overlapping fields-of-view are employed, other visual analysis tasks such as depth estimation can be supported and object recognition accuracy can be improved. In this paper the problem of distributed visual analysis from multiple views of a scene is addressed, considering that computational power and bandwidth, at each camera sensor, are rather limited. More specifically, an efficient coding technique for local binary features is proposed which exploits the correlation at the decoder side between each descriptor and its quantized representation. Moreover, considering that descriptors representing the same visual feature across different views are well correlated, a technique to avoid the transmission of redundant descriptors from multiple views is proposed. At the decoder, the joint statistics of all descriptors from all views is used to drive the selection of the best descriptors to be transmitted by each sensing node. The proposed multi-view feature coding and selection techniques allow obtaining bitrate reductions up to 80%, with respect to the uncompressed descriptor rate, for a certain task accuracy. Nuno Monteiro, Catarina Brites, Fernando Pereira 0001, João Ascenso |
ICME | 2 |
| 2016 | Digital holography: Benchmarking coding standards and representation formatsabstractHolography is an emerging technology to represent and display visual information and the associated experience expectations on the layman imaginary are huge. A hologram is a reproduction of a light field represented through the interference pattern between two wavefields, the reference and object wavefields. Holograms may be optically generated from real objects or computationally generated from synthetic objects leading to the so-called computer generated holograms. Whatever the creation process, holograms may have a digital representation using some appropriate representation format. Moreover, considering the huge amounts of data involved, digital holographic data has to be compressed using appropriate coding solutions, possibly available standard coding solutions. Finally, the decoded holographic data has to lead to reconstructed images using some appropriate reconstruction method. Holographic data coding is an emerging field of research where literature is still very scarce. In this context, before starting designing coding solutions considering the specific characteristics of holographic data, it is essential to assess the current compression performance associated to the most relevant available standard coding solutions and alternative representations formats. This paper has the main objective to benchmark the most relevant available standard coding solutions and the main alternative representation formats under meaningful test conditions and for the test material currently available. This assessment is critical to understand the current status quo and launch future research. Jose Peixeiro, Catarina Brites, João Ascenso, Fernando Pereira 0001 |
ICME | 2 |
| 2015 | Epipolar plane image based rendering for 3D video codingabstractIn current 3D video coding solutions, such as the 3D-HEVC standard, depth data is instrumental to have a continuum of views synthesized at the decoder based on a limited set of coded views. In order view synthesis may be performed at the decoder, depth data is currently directly acquired or estimated at the encoder based on very few neighboring views and transmitted to the decoder after appropriate compression. At the decoder, further views then those decoded are synthesized using again very few neighboring decoded views, thus using a local synthesis approach. A promising alternative synthesis approach may consider not a few but rather all the views available at the decoder, thus offering a scene global approach to synthesis. One way to implement this approach involves cutting the views cube along the viewpoint direction, creating the so-called epipolar plane images (EPI) which provide a rather compact representation of the scene. In this context, this paper proposes an EPI based view rendering framework for 3D video coding solution and identifies the major benefits of such framework, notably in comparison with the traditional local synthesis approach. Catarina Brites, João Ascenso, Fernando Pereira 0001 |
MMSP | 1 |
| 2015 | Multiview side information creation for efficient Wyner-Ziv video coding: Classifying and reviewing
Catarina Brites, Fernando Pereira 0001 |
Signal Process. Image Commun. | 1 |
| 2015 | Distributed video coding: Assessing the HEVC upgrade
Catarina Brites, Fernando Pereira 0001 |
Signal Process. Image Commun. | 1 |
| 2014 | Correlation noise modeling for multiview transform domain Wyner-Ziv video codingabstractMultiview Wyner-Ziv (MV-WZ) video coding rate-distortion (RD) performance is highly influenced by the adopted correlation noise model (CNM). In the related literature, the statistics of the correlation noise between the original frame and the side information (SI), typically resulting from the fusion of temporally and inter-view created SIs, is modelled by a Laplacian distribution. In most cases, the Laplacian CNM parameter is estimated using an offline approach, assuming that either the SI is available at the encoder or the originals are available at the decoder which is not realistic. In this context, this paper proposes the first practical, online CNM solution for a multiview transform domain WZ (MV-TDWZ) video codec. The online estimation of the Laplacian CNM parameter is performed at the decoder based on metrics exploring both the temporal and inter-view correlations with two levels of granularity, notably transform band and transform coefficient. The results obtained show that better RD performance is achieved for the finest granularity level since the inter-view, temporal and spatial correlations are exploited with the highest adaptation. Catarina Brites, Fernando Pereira 0001 |
ICIP | 1 |
| 2014 | Statistical reconstruction for predictive video codingabstractSubstantial rate-distortion (RD) gains have been achieved in video coding standards by increasing the encoder complexity while maintaining the decoder complexity the lowest possible. On the other hand, the alternative distributed video coding (DVC) approach proposes to exploit the video redundancy mostly at the decoder side, keeping the encoder as simple as possible. One of the most characteristic DVC tools is the statistical reconstruction of the DCT coefficients, which plays a similar role to the inverse scalar quantization (ISQ) in predictive codecs. The main objective of this paper is to propose a statistical reconstruction approach for predictive coding (notably the H.264/AVC standard) as a substitute to ISQ, thus creating a coding architecture with a mix of predictive and distributed coding tools. Experimental results show that the proposed statistical reconstruction solution allows achieving Bjontegaard bitrate savings up to 2.4% regarding the ISQ based H.264/AVC High profile codec. Catarina Brites, Vitor Gomes 0003, João Ascenso, Fernando Pereira 0001 |
VCIP | 1 |
| 2014 | Perceptually driven video error protection using a distributed source coding approach
André Seixas Dias, Catarina Brites, João Ascenso, Fernando Pereira 0001 |
Signal Process. Image Commun. | 2 |
| 2014 | Epipolar Geometry-Based Side Information Creation for Multiview Wyner-Ziv Video CodingabstractThe side information (SI) quality significantly influences the rate-distortion (RD) performance of both monoview and multiview Wyner-Ziv (WZ) video coding. Efficient multiview WZ video coding schemes typically exploit both temporal and inter-view correlation during the SI creation process targeting to achieve high SI quality and, thus, high RD performance. In this context, the overall SI creation process typically involves fusing a temporally created SI with an inter-view created SI. Consequently, the final SI quality does not only depend on the temporal and inter-view SI creation techniques but also on the fusion process. Thus, the main objective of this paper is to propose an efficient SI creation solution for multiview WZ video coding by simultaneously tackling the two main issues aforementioned with the following technical novelty: 1) a disparity-based view synthesis technique accounting for the scene geometry to create high quality inter-view SI and 2) a time-view driven fusion technique efficiently selecting, on a pixel basis, the temporal or inter-view SI depending on which one is estimated to be closer the original WZ data. Experimental results show considerable RD performance improvements, notably gains up to 2 dB or more, regarding the best performing state-of-the-art multiview WZ video coding solutions available. Catarina Brites, Fernando Pereira 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Side information creation for efficient Wyner-Ziv video coding: Classifying and reviewing
Catarina Brites, João Ascenso, Fernando Pereira 0001 |
Signal Process. Image Commun. | 1 |
| 2012 | Learning based decoding approach for improved Wyner-Ziv video codingabstractWyner-Ziv (WZ) video coding compression efficiency depends critically both on the side information (SI) quality and the correlation noise model (CNM) accuracy. In this context, this paper proposes a learning based decoding approach for transform domain WZ video coding, notably in the context of the following techniques: i) fractional-pixel motion field learning to define the relevance of the SI block candidates, and ii) CNM parameter learning. Experimental results show the proposed learning approach brings consistent RD performance improvements, with coding gains up to 3.9 dB regarding the state-of-the-art DISCOVER WZ video codec for a GOP size of 8. Catarina Brites, João Ascenso, Fernando Pereira 0001 |
PCS | 1 |
| 2011 | A denoising approach for iterative side information creation in distributed video codingabstractIn distributed video coding, motion estimation is typically performed at the decoder to generate the side information, increasing the decoder complexity while providing low complexity encoding in comparison with predictive video coding. Motion estimation can be performed once to create the side information or several times to refine the side information quality along the decoding process. In this paper, motion estimation is performed at the decoder side to generate multiple side information hypotheses which are adaptively and dynamically combined, whenever additional decoded information is available. The proposed iterative side information creation algorithm is inspired in video denoising filters and requires some statistics of the virtual channel between each side information hypothesis and the original data. With the proposed denoising algorithm for side information creation, a RD performance gain up to 1.2 dB is obtained for the same bitrate. João Ascenso, Catarina Brites, Fernando Pereira 0001 |
ICIP | 2 |
| 2011 | Low complexity deblocking filter perceptual optimization for the HEVC codecabstractThe compression efficiency of the state-of-art H.264/AVC video coding standard must be improved to accommodate the compression needs of high definition videos. To this end, ITU and MPEG started a new standardization project called High Efficiency Video Coding. The video codec under development still relies on trans form domain quantization and includes the same in-loop deblocking filter adopted in the H.264/AVC standard to reduce quantization blocking artifacts. This deblocking filter provides two offsets to vary the amount of filtering for each image area. This paper proposes a perceptual optimization of these offsets based on a quality metric able to quantify the blocking artifacts impact on the perceived video quality. The proposed optimization involves low computational complexity and provides quality improvements with respect to a non-perceptually optimized H.264/AVC deblocking filter. Moreover, the proposed optimization allows up to 92% of complexity reduction regarding a brute force perceptual optimization which exhaustively tests all the possible offsets values. Matteo Naccari, Catarina Brites, João Ascenso, Fernando Pereira 0001 |
ICIP | 2 |
| 2011 | Augmented LDPC graph for distributed video coding with multiple side informationabstractThe advances made in channel-capacity codes, such as turbo codes and low-density parity-check (LDPC) codes, have played a major role in the emerging distributed source coding paradigm. LDPC codes can be easily adapted to new source coding strategies due to their natural representation as bipartite graphs and the use of quasi-optimal decoding algorithms, such as belief propagation. This paper tackles a relevant scenario in distributed video coding: lossy source coding when multiple side information (SI) hypotheses are available at the decoder, each one correlated with the source according to different correlation noise channels. Thus, it is proposed to exploit multiple SI hypotheses through an efficient joint decoding technique with multiple LDPC syndrome decoders that exchange information to obtain coding efficiency improvements. At the decoder side, the multiple SI hypotheses are created with motion compensated frame interpolation and fused together in a novel iterative LDPC based Slepian-Wolf decoding algorithm. With the creation of multiple SI hypotheses and the proposed decoding algorithm, bitrate savings up to 8.0% are obtained for similar decoded quality. João Ascenso, Catarina Brites, Fernando Pereira 0001 |
MMSP | 2 |
| 2011 | An Efficient Encoder Rate Control Solution for Transform Domain Wyner-Ziv Video CodingabstractMost Wyner-Ziv (WZ) video coding solutions in the literature use a feedback channel (FC) based decoder rate control (DRC) strategy to adjust the bitrate to correct the side information (SI) errors. More recently, some encoder rate control (ERC) strategies have been proposed to address application scenarios where a FC is not available. The ERC based WZ video coding RD performance depends not only on the (encoder) parity rate estimator (PRE) accuracy but also on the decoder “intelligence” in dealing with the residual errors due to parity rate underestimation. In this context, the main objective of this paper is to propose a more efficient and powerful ERC solution for transform domain WZ (TDWZ) video coding by simultaneously tackling the two issues aforementioned with the following technical novelty: 1) integration in an ERC context of Gray mapping for the quantized DCT coefficients to enhance the correlation between WZ and SI data; 2) more accurate PRE to better estimate the needed parity rate to avoid undesired parity rate underestimations and overestimations; 3) novel soft reconstruction function to reduce the impact of the residual bitplane errors in the decoded WZ frame quality; and 4) weighted overlapped block motion compensation technique to refine the SI used in an iterative WZ decoding framework with the correlation noise model parameters dynamically updated. Experimental results show a considerable RD performance improvement with a reduction of up to about 2 dB of the gap between the ERC and DRC based approaches in TDWZ video coding solutions, thus making this ERC based WZ codec the most efficient available and competitive regarding DRC based WZ video coding solutions. Catarina Brites, Fernando Pereira 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Probability updating for decoder and encoder rate control turbo based Wyner-Ziv video codingabstractIn Wyner-Ziv video coding (WZVC), powerful error correcting codes must be used to achieve high compression efficiency; turbo codes are the most commonly used error correcting codes in WZVC. To improve the turbo coding performance in the context of WZVC, this paper proposes a probability updating technique (PUT) acting as an outer loop of the common turbo decoding operation. Whenever a turbo decoded bitplane is not error-free, the proposed technique attempts to correct bitplane errors by updating the correlation noise probabilities for the most likely in error bits, followed by turbo redecoding. The new tool is evaluated both in the context of encoder rate control (ERC) and decoder rate control (DRC) turbo based WZVC scenarios with average overall PSNR gains up to about 0.5 dB in ERC and average WZ rate savings up to about 6% in DRC. Catarina Brites, Fernando Pereira 0001 |
ICIP | 1 |
| 2010 | A flexible side information generation framework for distributed video coding
João Ascenso, Catarina Brites, Fernando Pereira 0001 |
Multim. Tools Appl. | 2 |
| 2009 | Distributed Video Coding with multiple side informationabstractDistributed Video Coding (DVC) is a new video coding paradigm which mainly exploits the source statistics at the decoder based on the availability of some decoder side information. The quality of the side information has a major impact on the DVC rate-distortion (RD) performance in the same way the quality of the predictions had a major impact in predictive video coding. In this paper, a DVC solution exploiting multiple side information is proposed; the multiple side information is generated by frame interpolation and frame extrapolation targeting to improve the side information of a single estimation mode. Compared with the best available single side information solutions, the proposed DVC solution with multiple side information robustly improves the RD performance for the set of test sequences. Xin Huang 0004, Catarina Brites, João Ascenso, Fernando Pereira 0001, Søren Forchhammer |
PCS | 2 |
| 2009 | Adaptive deblocking filter for transform domain Wyner-Ziv video codingabstractWyner–Ziv (WZ) video coding is a particular case of distributed video coding, the recent video coding paradigm based on the Slepian–Wolf and Wyner–Ziv theorems that exploits the source correlation at the decoder and not at the encoder as in predictive video coding. Although many improvements have been done over the last years, the performance of the state-of-the-art WZ video codecs still did not reach the performance of state-of-the-art predictive video codecs, especially for high and complex motion video content. This is also true in terms of subjective image quality mainly because of a considerable amount of blocking artefacts present in the decoded WZ video frames. This paper proposes an adaptive deblocking filter to improve both the subjective and objective qualities of the WZ frames in a transform domain WZ video codec. The proposed filter is an adaptation of the advanced deblocking filter defined in the H.264/AVC (advanced video coding) standard to a WZ video codec. The results obtained confirm the subjective quality improvement and objective quality gains that can go up to 0.63 dB in the overall for sequences with high motion content when large group of pictures are used. Catarina Brites, João Ascenso, Fernando Pereira 0001 |
IET Image Process. | 2 |
| 2009 | Refining Side Information for Improved Transform Domain Wyner-Ziv Video CodingabstractWyner-Ziv (WZ) video coding is a particular case of distributed video coding, which is a recent video coding paradigm based on the Slepian-Wolf and WZ theorems. Contrary to available prediction-based standard video codecs, WZ video coding exploits the source statistics at the decoder, allowing the development of simpler encoders. Until now, WZ video coding did not reach the compression efficiency performance of conventional video coding solutions, mainly due to the poor quality of the side information, which is an estimate of the original frame created at the decoder in the most popular WZ video codecs. In this context, this paper proposes a novel side information refinement (SIR) algorithm for a transform domain WZ video codec based on a learning approach where the side information is successively improved as the decoding proceeds. The results show significant and consistent performance improvements regarding state-of-the-art WZ and standard video codecs, especially under critical conditions such as high motion content and long group of pictures sizes. Catarina Brites, João Ascenso, Fernando Pereira 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Design and performance of a novel low-density parity-check code for distributed video codingabstractLow-density parity-check (LDPC) codes are nowadays one of the hottest topics in coding theory, notably due to their advantages in terms of bit error rate performance and low complexity. In order to exploit the potential of the Wyner-Ziv coding paradigm, practical distributed video coding (DVC) schemes should use powerful error correcting codes with near-capacity performance. In this paper, new ways to design LDPC codes for the DVC paradigm are proposed and studied. The new LDPC solutions rely on merging parity-check nodes, which corresponds to reduce the number of rows in the parity-check matrix. This allows to change gracefully the compression ratio of the source (DCT coefficient bitplane) according to the correlation between the original and the side information. The proposed LDPC codes reach a good performance for a wide range of source correlations and achieve a better RD performance when compared to the popular turbo codes. João Ascenso, Catarina Brites, Fernando Pereira 0001 |
ICIP | 2 |
| 2008 | Wyner-Ziv video coding: A review of the early architectures and further developmentsabstractIn 2002, the video coding community faced the emergence of a new video coding paradigm, the so-called Wyner-Ziv video coding, which was represented by two early solutions designed by the Stanford University and the University of California, Berkeley research teams. This paper intends to briefly review, and compare these two early Wyner-Ziv video coding solutions, notably from the functional point of view. Moreover, this paper reviews some important developments of the Stanford Wyner-Ziv coding architecture, which has become the most popular in the literature. Fernando Pereira 0001, Catarina Brites, João Ascenso, Marco Tagliasacchi |
ICME | 2 |
| 2008 | Evaluating a feedback channel based transform domain Wyner-Ziv video codec
Catarina Brites, João Ascenso, José Quintas Pedro, Fernando Pereira 0001 |
Signal Process. Image Commun. | 1 |
| 2008 | Correlation Noise Modeling for Efficient Pixel and Transform Domain Wyner-Ziv Video CodingabstractIn recent years, practical Wyner-Ziv (WZ) video coding solutions have been proposed with promising results. Most of the solutions available in the literature model the correlation noise (CN) between the original frame and its estimation made at the decoder, which is the so-called side information (SI), by a given distribution whose relevant parameters are estimated using an offline process, assuming that the SI is available at the encoder or the originals are available at the decoder. The major goal of this paper is to propose a more realistic WZ video coding approach by performing online estimation of the CN model parameters at the decoder, for pixel and transform domain WZ video codecs. In this context, several new techniques are proposed based on metrics which explore the temporal correlation between frames with different levels of granularity. For pixel-domain WZ (PDWZ) video coding, three levels of granularity are proposed: frame, block, and pixel levels. For transform-domain WZ (TDWZ) video coding, DCT bands and coefficients are the two granularity levels proposed. The higher the estimation granularity is, the better the rate-distortion performance is since the deeper the adaptation of the decoding process is to the video statistical characteristics, which means that the pixel and coefficient levels are the best performing for PDWZ and TDWZ solutions, respectively. Catarina Brites, Fernando Pereira 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Encoder Rate Control for Transform Domain Wyner-Ziv Video CodingabstractWyner-Ziv (WZ) video coding -a particular case of distributed video coding (DVC) -is a new video coding paradigm based on two major Information Theory results: the Slepian-Wolf and Wyner-Ziv theorems. Many of the practical WZ video coding solutions available in the literature make use of a feedback channel (FC) to perform rate control at the decoder which implies there must be a FC available in the application scenario addressed. The FC-based DVC solutions also have implications in terms of delay and decoder complexity since several iterative decoding operations may be needed to decode the data to the target quality level. In this context, this paper proposes an encoder rate control (ERC) solution for the transform domain WZ coding architecture previously using a FC driven rate control. Although this is the first solution in the literature, promising results are achieved with the proposed ERC solution without significantly increase the encoder complexity. Catarina Brites, Fernando Pereira 0001 |
ICIP (2) | 1 |
| 2007 | Wyner-Ziv Stereo Video Coding using a Side Information Fusion ApproachabstractWyner-Ziv coding, also known as distributed video coding, is currently a very hot research topic in video coding due to the new opportunities it opens. This paper applies the distributed video coding principles to stereo video coding, to propose a practical solution for Wyner-Ziv stereo coding based on mask-based fusion of temporal and spatial side informations. The architecture includes a low-complexity encoder and avoids any communication between the cameras/encoders. While the rate-distortion (RD) performance strongly depends on the motion-based frame interpolation (MBFI) and disparity-based frame estimation (DBFE) solutions, first results show that the proposed approach is promising and there are still issues to address. José Diogo Areia, João Ascenso, Catarina Brites, Fernando Pereira 0001 |
MMSP | 3 |
| 2007 | Studying the GOP Size Impact on the Performance of a Feedback Channel-Based Wyner-Ziv Video Codec
Fernando Pereira 0001, João Ascenso, Catarina Brites |
PSIVT | 3 |
| 2006 | Improving Transform Domain Wyner-Ziv Video Coding PerformanceabstractDistributed video coding (DVC) is a new video coding paradigm based on two key information theory results: the Slepian-Wolf and Wyner-Ziv theorems. A particular case of DVC, the so-called Wyner-Ziv coding, deals with lossy source coding with side information at the decoder and enables a flexible allocation of complexity between the encoder and the decoder. This paper proposes an improved transform domain Wyner-Ziv video codec including: 1) the integer block-based transform defined in the H.264/MPEG-4 AVC standard, 2) a quantizer with a symmetrical interval around zero for AC coefficients, and a quantization step size adjusted to the transform coefficient bands dynamic range, and 3) advanced frame interpolation for side information generation. The combination of these tools brings significant rate-distortion (RD) gains regarding the state-of-the-art results available in the literature Catarina Brites, João Ascenso, Fernando Pereira 0001 |
ICASSP (2) | 1 |
| 2006 | Intra Mode Decision Based on Spatio-Temporal Cues in Pixel Domain Wyner-ZIV Video CodingabstractDistributed source coding principles have been recently applied to video coding in order to achieve a flexible distribution of the complexity burden between the encoder and the decoder. In this paper we elaborate on a pixel based Wyner-Ziv video codec that shifts all the complexity of the motion estimation phase to the decoder, thus achieving light encoding. We observe that the correlation noise statistics describing the relationship between the frame to be encoded and the side information available at the decoder is not spatially stationary. For this reason we introduce a mode decision scheme either at the encoder or at the decoder in such a way that when the estimated correlation is weak we opt for intra coding on a block-by-block basis. Both spatial and temporal criteria are used to determine whether a block is better intra coded or not Marco Tagliasacchi, Alan Trapanese, Stefano Tubaro, João Ascenso, Catarina Brites, Fernando Pereira 0001 |
ICASSP (2) | 5 |
| 2006 | Content Adaptive Wyner-ZIV Video Coding Driven by Motion ActivityabstractIn distributed video coding (DVC), the video statistics are exploited, partially or totally at the decoder. A particular case of DVC, Wyner-Ziv video coding deals with lossy source coding with side information at the decoder and allows moving part or the entire motion estimation task to the decoder. In this context, it is the decoder responsibility to obtain the side information, a guess of the encoded Wyner-Ziv frame and the encoder only sends parity bits to improve its quality. In this paper, a technique targeting the improvement of the quality of the side information, and thus of the rate-distortion performance of the Wyner-Ziv codec is proposed. This is achieved by adaptively adjusting the size of the motion interpolation structure (or GOP length) according to the motion activity along the sequence. Experimentally, this allows to achieve gains up to 0.8 dB without performing any motion estimation or complex mode decision at the encoder. João Ascenso, Catarina Brites, Fernando Pereira 0001 |
ICIP | 2 |
| 2006 | Studying Temporal Correlation Noise Modeling for Pixel Based Wyner-Ziv Video CodingabstractWyner-Ziv (WZ) video coding-a particular case of distributed video coding (DVC)-is a new video coding paradigm based on two major information theory results: the Slepian-Wolf and Wyner-Ziv theorems. Recently, practical WZ video coding solutions were proposed with promising results. Most of the solutions available in the literature, model the correlation noise between the original frame and the so-called side information by a given distribution whose relevant parameters are estimated in an offline process, at the encoder. In this paper, three algorithms are proposed towards a more realistic WZ coding approach by performing online estimation of the error distribution at the decoder. Both algorithms explore temporal correlation between frames however with different levels of granularity: frame, block and pixel levels; better rate-distortion (RD) performance is achieved for lower granularity (pixel) level. Catarina Brites, João Ascenso, Fernando Pereira 0001 |
ICIP | 1 |
| 2006 | Exploiting Spatial Redundancy in Pixel Domain Wyner-Ziv Video CodingabstractDistributed video coding is a recent paradigm that enables a flexible distribution of the computational complexity between the encoder and the decoder building on top of distributed source coding principles. In this paper we focus on the scenario where most of the complexity is shifted to the decoder, thus achieving light encoding. We elaborate on a well known pixel based Wyner-Ziv architecture and we improve its coding efficiency by exploiting both spatial and temporal correlation at the decoder side, without the need of performing any transform at the encoder. In order to generate the side information, the decoder adaptively chooses spatial or temporal information, based on the local estimate of the correlation noise. Simulations on test sequences demonstrate that a coding gain of up to +1.8 dB can be obtained with respect to the case that generates the side information by motion interpolation only. Marco Tagliasacchi, Alan Trapanese, Stefano Tubaro, João Ascenso, Catarina Brites, Fernando Pereira 0001 |
ICIP | 5 |
| 2005 | Motion compensated refinement for low complexity pixel based distributed video codingabstractDistributed video coding (DVC) is a new coding paradigm that enables to exploit video statistics, partially or totally at the decoder. A particular case of DVC, Wyner-Ziv coding, deals with lossy source coding with side information at the decoder and allows a shift of complexity from the encoder to the decoder, theoretically without any penalty in the coding efficiency. The Wyner-Ziv solution here described encodes each video frame independently (intraframe coding), but decodes the same frame conditionally (interframe decoding). At the decoder, and compensation tools are responsible to obtain an accurate interpolation of the original frame using previously decoded (temporally adjacent) frames. This paper proposes a novel approach to improve the performance of pixel domain Wyner-Ziv video coding by using a motion compensated refinement of the decoded frame and use it as improved side information. More precisely, upon partial decoding of each frame, the decoder refines its motion trajectories in order to achieve a better reconstruction of the decoded frame. João Ascenso, Catarina Brites, Fernando Pereira 0001 |
AVSS | 2 |