EDBT 2026 Demo / reviewers in the wild / expert
Nuno M. M. Rodrigues
dblp:48/1737
· DBLP profile ↗
43ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-9536-1017ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 4 first-author · 11 since 2021Systems, architecture and hardware · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scalable Graph-Guided Transformer for Point Cloud Geometry CodingabstractAttention models, particularly Transformers, have significantly advanced deep learning in fields like natural language processing and computer vision by capturing contextual relationships in both sequential and spatial data. This ability is valuable for Point Clouds (PC), which are unstructured sets of points in 3D space. Transformers can effectively identify correlations between distant points, allowing them to focus on the most critical regions of the data. To demonstrate this capability, this paper proposes a novel, scalable Graph-Guided Transformer model, labeled 2GFormer, for static PC geometry. This model is built using a scalable architecture that leverages Graph Convolutions to enhance a Relational Neighborhood SelfAttention (RNSA) base layer model. Both models are integrated into the JPEG Pleno Learning-based Point Cloud Coding (JPEG PCC) standard, resulting in the creation of two attention-enabled codecs for static PC coding: JPEG RNSA and JPEG 2GFormer. While JPEG RNSA codec delivers significant compression improvements for solid and dense PCs compared to the baseline JPEG PCC standard, JPEG 2GFormer extends these gains to solid, dense, and sparse PCs with only a marginal increase in model parameters. Additionally, JPEG 2GFormer outperforms both conventional and learning-based state-of-the-art PC codecs. These results position JPEG 2GFormer as a highly efficient solution for versatile PC coding. Mohammadreza Ghafari, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Deep Learning-Based Point Cloud Coding and Super-Resolution: A Joint Geometry and Color ApproachabstractIn this golden age of multimedia, realistic content is in high demand with users seeking more immersive and interactive experiences. As a result, new image modalities for 3D representations have emerged in recent years, among which point clouds have deserved especial attention. Naturally, with this increase in demand, efficient storage and transmission became a must, with standardization groups such as MPEG and JPEG entering the scene, as it happened before with other types of visual media. In a surprising development, JPEG issued a Call for Proposals on point cloud coding targeting exclusively learning-based solutions, in parallel to a similar call for image coding. This is a natural consequence of the growing popularity of deep learning, which due to its excellent performances is currently dominant in the multimedia processing field, including coding. This paper presents the coding solution selected by JPEG as the best-performing response to the Call for Proposals and adopted as the first version of the JPEG Pleno Point Cloud Coding Verification Model, in practice the first step for developing a standard. The proposed solution offers a novel joint geometry and color approach for point cloud coding, in which a single deep learning model processes both geometry and color simultaneously. To maximize the RD performance for a large range of point clouds, the proposed solution uses down-sampling and learning-based super-resolution as pre- and post-processing steps. Compared to the MPEG point cloud coding standards, the proposed coding solution comfortably outperforms G-PCC, for both geometry, color, and joint quality metrics. André F. R. Guarda, Manuel Ruivo, Luís Coelho, Abdelrahman Seleem, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Learning-Based Point Cloud Decoding with Independent and Scalable Reduced ComplexityabstractPoint Clouds (PCs) have gained significant attention due to their usage in diverse application domains, notably virtual and augmented reality. While PCs excel in providing detailed 3D visualization, this typically requires millions of points which must be efficiently coded for real-world deployment, notably storage and streaming. Recently, learning-based coding solutions have been adopted, notably in the JPEG Pleno Point Coding (PCC) standard, which uses a coding model with millions of model parameters. This requires the use of high-performance computing devices, which may not be available, notably at the decoder side. In this context, this paper proposes two reduced complexity decoding solutions, based on the adoption of scalability principles, to decode the same JPEG PCC compliant bitstreams. The design of these solutions is based on two innovative model pruning strategies which reduce the decoding complexity. The experimental results demonstrate the effective capability to significantly reduce the number of decoding model parameters with an acceptable penalty on Rate-Distortion (RD) performance compared to the full complexity model. Mohammadreza Ghafari, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
ICIP | 3 |
| 2024 | Point Cloud Geometry Scalable Coding with a Quality-Conditioned Latents Probability EstimatorabstractThe widespread usage of point clouds (PC) for immersive visual applications has resulted in the use of very heterogeneous receiving conditions and devices, notably in terms of network, hardware, and display capabilities. In this scenario, quality scalability, i.e., the ability to reconstruct a signal at different qualities by progressively decoding a single bitstream, is a major requirement that has yet to be conveniently addressed, notably in most learning-based PC coding solutions. This paper proposes a quality scalability scheme, named Scalable Quality Hyperprior (SQH), adaptable to learning-based static point cloud geometry codecs, which uses a Quality-conditioned Latents Probability Estimator (QuLPE) to decode a high-quality version of a PC learning-based representation, based on an available lower quality base layer. SQH is integrated in the future JPEG PC coding standard, allowing to create a layered bitstream that can be used to progressively decode the PC geometry with increasing quality and fidelity. Experimental results show that SQH offers the quality scalability feature with very limited or no compression performance penalty at all when compared with the corresponding non-scalable solution, thus preserving the significant compression gains over other state-of-the-art PC codecs. Daniele Mari, André F. R. Guarda, Nuno M. M. Rodrigues, Simone Milani, Fernando Pereira 0001 |
ICIP | 3 |
| 2024 | Point Cloud Geometry Coding with Relational Neighborhood Self-AttentionabstractIn the ever-evolving landscape of deep learning, attention models have contributed to boost the performance in diverse fields such as computer vision and natural language processing. Following this trend, this paper proposes a novel Relational Neighborhood Self-Attention (RNSA) model, specifically designed for Point Cloud (PC) geometry coding to be integrated in the emerging learning-based JPEG PCC standard. The RNSA model proposes three new methods: first, to effectively learn correlations between the points by capturing the relational features and positions of neighboring points; second, to address the inefficiencies of conventional dot product attention, a novel Relational Scoring method to generate an attention map able to capture both linear and non-linear relationships between points and their neighbors is adopted; third, the created attention maps are normalized by Sparsemax instead of Softmax to generate sparse probabilities and assigns higher scores to the most important neighbors while marginalizing the less significant ones. Experimental results show that the proposed attention model achieves around 8% gains in both BD-Rate PSNR Dl and PSNR D2 compared to the baseline codec, i.e., JPEG PCC, while adding a small number of model parameters to JPEG PCC. Mohammadreza Ghafari, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
MMSP | 3 |
| 2023 | Point Cloud Geometry and Color Coding in a Learning-Based Ecosystem for JPEG Coding StandardsabstractDespite its novelty, learning-based coding for images and point clouds is already outperforming some of the best long-standing conventional codecs. In addition to its rising compression performance, learning-based coding has opened new opportunities, notably the use of a single compressed domain representation to provide both high fidelity reconstructions for human visualization as well as effective performance for computer vision tasks, effectively unifying the visual language for man and machine. This paper proposes a new double learning-based static point cloud geometry and color coding solution, which targets point cloud component scalability, rate control flexibility at coding time, and a unified compressed domain representation. The proposed solution exploits the synergies between learning-based coding for images and point clouds, through the current JPEG PCC standard for geometry coding and the JPEG AI standard for image/color coding, establishing a learning-based ecosystem for JPEG coding standards. The proposed solution is able to overcome some design limitations of the current JPEG PCC Verification Model, and significantly improve its RD performance, becoming competitive with MPEG PCC standards. André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
ICIP | 2 |
| 2023 | Deep Learning-Based Compressed Domain Point Cloud ClassificationabstractDeep learning (DL) based tools have recently reached performance levels similar to state-of-the-art hand-crafted methods for Point Cloud (PC) coding and classification. In 2022, JPEG issued a Call for Proposals for a Learning-based PC Coding (PCC) standard that envisions a unified representation, targeting both human visualization and computer vision tasks. This paper proposes the first DL-based Compressed Domain PC CLassifier (CD-PCCL), built on the PointGrid classifier, for geometry-only PCs coded with the current DL-based JPEG Pleno PCC Verification Model. The performance of compressed domain PC classification is studied against using voxel domain classification, notably for original, voxelized, and decompressed PCs. Experimental results with the ModelNet40 PC dataset show the proposed CD-PCCL can achieve significant PC classification gains regarding decompressed domain classification, while reducing the PC classifier complexity. Abdelrahman Seleem, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
ICIP | 3 |
| 2023 | Deep Learning-based Point Cloud Geometry Coding with Attention ModelsabstractRecent advancements in Deep Learning (DL)-based architectures have demonstrated that integrating attention models can substantially enhance the performance across various tasks, including computer vision and visual coding. In accordance with this trend, this paper proposes a framework for incorporating attention models into DL-based Point Cloud (PC) geometry coding, namely JPEG Pleno PC Coding (JPEG PCC) and a definition of its key architectural design options. Experimental results show that the integration of attention models in JPEG PCC can provide a trade-off between compression and complexity, notably compression gains at the cost of an increase in model complexity and number of parameters. If the priority is on compression gains, the use of attention models can lead to a rate reduction up to 5.7%, for both the PSNR DI and PSNR D2 geometry quality metrics, at the cost of a 47% increase on the number of parameters and a 15.1% increase on the Floating-Point Operations (FLOPs) complexity. This obtained trade-off depends on the attention model integration configuration which can be defined depending on the application requirements. Mohammadreza Ghafari, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
ISM | 3 |
| 2022 | Impact of Conventional and Deep Learning-based Point Cloud Geometry Coding on Deep Learning-based Classification PerformanceabstractDeep learning (DL)-based point cloud (PC) classification is a key computer vision task for many applications, notably autonomous driving, surveillance, and cultural heritage. In many application scenarios, PCs must be coded to reach practical rates for storage and transmission purposes, and thus they suffer from more or less intense compression artifacts. After the specification of two MPEG PC coding standards, DL-based PC coding has gained momentum, reaching competitive compression performance, especially for dense PCs. Since using decoded PCs, which may suffer from compression artifacts, may impact the final classification performance, the main goal of this paper is to study the impact of static PC geometry coding on DL-based classification. This study is performed on the ModelNet40 test dataset using the conventional G-PCC coding standard and the DL-based PC geometry codec which was the top performing solution responding to the recent JPEG Pleno PC Coding Call for Proposals. Two highly performing DL-based classifiers are used, considering the original PC geometry before and after voxelization, as well as the decoded PC geometry for different rates and qualities. As expected, coding has an impact on the classification performance, especially for the lower rates/qualities. For very sparse PCs, conventional coding still has advantage, contrarily to dense PCs, but this should change in the future with DL-based tools becoming the most natural solutions for both PC geometry coding and classification. Abdelrahman Seleem, André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
ISM | 3 |
| 2021 | Light field image coding with flexible viewpoint scalability and random accessabstractThis paper proposes a novel light field image compression approach with viewpoint scalability and random access functionalities. Although current state-of-the-art image coding algorithms for light fields already achieve high compression ratios, there is a lack of support for such functionalities, which are important for ensuring compatibility with different displays/capturing devices, enhanced user interaction and low decoding delay. The proposed solution enables various encoding profiles with different flexible viewpoint scalability and random access capabilities, depending on the application scenario. When compared to other state-of-the-art methods, the proposed approach consistently presents higher bitrate savings (44% on average), namely when compared to pseudo-video sequence coding approach based on HEVC. Moreover, the proposed scalable codec also outperforms MuLE and WaSP verification models, achieving average bitrate saving gains of 37% and 47%, respectively. The various flexible encoding profiles proposed add fine control to the image prediction dependencies, which allow to exploit the tradeoff between coding efficiency and the viewpoint random access, consequently, decreasing the maximum random access penalties that range from 0.60 to 0.15, for lenslet and HDCA light fields. Ricardo J. S. Monteiro, Nuno M. M. Rodrigues, Sérgio M. M. de Faria, Paulo J. L. Nunes |
Signal Process. Image Commun. | 2 |
| 2021 | Constant Size Point Cloud Clustering: A Compact, Non-Overlapping SolutionabstractPoint clouds have recently become a popular 3D representation model for many application domains, notably virtual and augmented reality. Since point cloud data is often very large, processing a point cloud may require that it be segmented into smaller clusters. For example, the input to deep learning-based methods like auto-encoders should be constant size point cloud clusters, which are ideally compact and non-overlapping. However, given the unorganized nature of point clouds, defining the specific data segments to code is not always trivial. This paper proposes a point cloud clustering algorithm which targets five main goals: i) clusters with a constant number of points; ii) compact clusters, i.e., with low dispersion; iii) non-overlapping clusters, i.e., not intersecting each other; iv) ability to scale with the number of points; and v) low complexity. After appropriate initialization, the proposed algorithm transfers points between neighboring clusters as a propagation wave, filling or emptying clusters until they achieve the same size. The proposed algorithm is unique since there is no other point cloud clustering method available in the literature offering the same clustering features for large point clouds at such low complexity. André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Point Cloud Geometry Scalable Coding With a Single End-to-End Deep Learning ModelabstractPoint clouds are gaining importance as the format to represent complex 3D objects and scenes, offering high user immersion and interaction, although at the cost of requiring massive data. Scalable coding is an important feature for point cloud coding, especially for real-time applications, where the fast and bitrate efficient access to a decoded point cloud is important; however, this issue is still rather unexplored in the literature. With the rise of deep learning methods as a promising solution for efficient coding, this paper proposes the first deep learning-based point cloud geometry scalable coding solution. Experimental results show that the proposed scalable coding solution consistently outperforms the MPEG standard for static point cloud geometry coding. In this way, a new research path is open for point cloud scalable coding technology. André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
ICIP | 2 |
| 2020 | Deep Learning-based Point Cloud Geometry Coding with Resolution ScalabilityabstractPoint clouds are a 3D visual representation format that has recently become fundamentally important for immersive and interactive multimedia applications. Considering the high number of points of practically relevant point clouds, and their increasing market demand, efficient point cloud coding has become a vital research topic. In addition, scalability is an important feature for point cloud coding, especially for real-time applications, where the fast and rate efficient access to a decoded point cloud is important; however, this issue is still rather unexplored in the literature. In this context, this paper proposes a novel deep learning-based point cloud geometry coding solution with resolution scalability via interlaced sub-sampling. As additional layers are decoded, the number of points in the reconstructed point cloud increases as well as the overall quality. Experimental results show that the proposed scalable point cloud geometry coding solution outperforms the recent MPEG Geometry-based Point Cloud Compression standard which is much less scalable. André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
MMSP | 2 |
| 2019 | Point Cloud Coding: Adopting a Deep Learning-based ApproachabstractPoint clouds have recently become an important visual representation format, especially for virtual and augmented reality applications, thus making point cloud coding a very hot research topic. Deep learning-based coding methods have recently emerged in the field of image coding with increasing success. These coding solutions take advantage of the ability of convolutional neural networks to extract adaptive features from the images to create a latent representation that can be efficiently coded. In this context, this paper extends the deep-learning coding approach to point cloud coding using an autoencoder network design. Performance results are very promising, showing improvements over the Point Cloud Library codec often taken as benchmark, thus suggesting a significant margin of evolution for this new point cloud coding paradigm. André F. R. Guarda, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
PCS | 2 |
| 2017 | Improving point cloud to surface reconstruction with generalized Tikhonov regularizationabstractPoint cloud rendering has a vital role in the user Quality of Experience for applications adopting point cloud based representations. While this is not a new area, it has recently become more relevant with the recent interest on point cloud coding by major standardization groups, notably JPEG and MPEG. The screened Poisson surface reconstruction is a state-of-the-art technique for generating a watertight surface mesh from the point cloud samples. While its screening component allows the surface to better fit the cloud points, this fitting may lead to undesired artifacts in the surface, notably when the point cloud is noisy. This paper proposes to improve this reconstruction method by making it more robust to noise by adopting a generalized Tikhonov regularization term. The proposed regularization approach smooths regions that should be flat while keeping the important details in the edges, thus creating more pleasant surface reconstructions. André F. R. Guarda, José M. Bioucas-Dias, Nuno M. M. Rodrigues, Fernando Pereira 0001 |
MMSP | 3 |
| 2017 | Energy-Efficient and Portable Least Squares Prediction for Image Coding on a Mobile GPUabstractLeast squares prediction is a technique used to foresee pixel values during image coding by finding the minimum square error of neighbouring pixels. It has shown considerable quality gains especially for complex images with high variations in pixel intensities. The drawback of this technique consists of high computational complexity, consuming the most significant part of processing time and resources available, which makes it difficult to implement in fast, lossy image coders. One challenge is therefore to reduce the computational time of this predictor, namely through the use of new parallel programming techniques, making it more attractive for state-of-the-art coder-decoders. Also, new algorithmic propositions are made, trying to reduce the time spent in exchange for rate-distortion performance. These propositions are senseful since this predictor is used not only in lossless image coding, but also in lossy as well. Another aim of this article is to analyze energy efficiency among different types of platforms for this signal processing algorithm. Comparisons are provided on parallel computing processors ranging from very powerful Graphics Computing Units (GPUs) to mobile General-Purpose GPUs. Pedro Cordeiro, Gabriel Falcão Paiva Fernandes, Patrício Domingues, Nuno M. M. Rodrigues, Sérgio M. M. de Faria |
PDP | 4 |
| 2017 | A method to improve HEVC lossless coding of volumetric medical images
André F. R. Guarda, João M. Santos 0002, Luís Alberto da Silva Cruz, Pedro A. Amado Assunção, Nuno M. M. Rodrigues, Sérgio M. M. de Faria |
Signal Process. Image Commun. | 5 |
| 2017 | Lossless Compression of Medical Images Using 3-D PredictorsabstractThis paper describes a highly efficient method for lossless compression of volumetric sets of medical images, such as CTs or MRIs. The proposed method, referred to as 3-D-MRP, is based on the principle of minimum rate predictors (MRPs), which is one of the state-of-the-art lossless compression technologies presented in the data compression literature. The main features of the proposed method include the use of 3-D predictors, 3-D-block octree partitioning and classification, volume-based optimization, and support for 16-b-depth images. Experimental results demonstrate the efficiency of the 3-D-MRP algorithm for the compression of volumetric sets of medical images, achieving gains above 15% and 12% for 8- and 16-bit-depth contents, respectively, when compared with JPEG-LS, JPEG2000, CALIC, and HEVC, as well as other proposals based on the MRP algorithm. Luis F. R. Lucas, Nuno M. M. Rodrigues, Luís Alberto da Silva Cruz, Sérgio M. M. de Faria |
IEEE Trans. Medical Imaging | 2 |
| 2016 | Optimizing GPU Code for CPU Execution Using OpenCL and Vectorization: A Case Study on Image Coding
Pedro M. M. Pereira, Patrício Domingues, Nuno M. M. Rodrigues, Gabriel Falcão Paiva Fernandes, Sérgio M. M. de Faria |
ICA3PP | 3 |
| 2016 | Compression of medical images using MRP with bi-directional prediction and histogram packingabstractMedical imaging technology has become essential for the improvement of medical practice. This led to advances in the technology, namely in image sampling resolutions, pixel bit-depth and inter slice resolution. Additionally, common use of medical images, the life expectancy of patients and legal restrictions led to increasing storage costs. Therefore, efficient compression of medical image data is in high demand, for archiving and transmission. In this work we propose to improve the compression efficiency of the Minimum Rate Predictors lossless encoder, by adding bi-directional prediction support and a histogram packing technique. The results show that the proposed method presents a higher compression efficiency than state-of-the-art HEVC encoder. The compression efficiency is improved by 20%, on average, when compared to HEVC and by 46.1% when compared with the original MRP algorithm. João M. Santos 0002, André F. R. Guarda, Luís Alberto da Silva Cruz, Nuno M. M. Rodrigues, Sérgio M. M. de Faria |
PCS | 4 |
| 2016 | Image Coding Using Generalized Predictors Based on Sparsity and Geometric TransformationsabstractDirectional intra prediction plays an important role in current state-of-the-art video coding standards. In directional prediction, neighbouring samples are projected along a specific direction to predict a block of samples. Ultimately, each prediction mode can be regarded as a set of very simple linear predictors, a different one for each pixel of a block. Therefore, a natural question that arises is whether one could use the theory of linear prediction in order to generate intra prediction modes that provide increased coding efficiency. However, such an interpretation of each directional mode as a set of linear predictors is too poor to provide useful insights for their design. In this paper, we introduce an interpretation of directional prediction as a particular case of linear prediction, which uses the first-order linear filters and a set of geometric transformations. This interpretation motivated the proposal of a generalized intra prediction framework, whereby the first-order linear filters are replaced by adaptive linear filters with sparsity constraints. In this context, we investigate the use of efficient sparse linear models, adaptively estimated for each block through the use of different algorithms, such as matching pursuit, least angle regression, least absolute shrinkage and selection operator, or elastic net. The proposed intra prediction framework was implemented and evaluated within the state-of-the-art high efficiency video coding standard. Experiments demonstrated the advantage of this predictive solution, mainly in the presence of images with complex features and textured areas, achieving higher average bitrate savings than other related sparse representation methods proposed in the literature. Luis F. R. Lucas, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Carla L. Pagliari, Sérgio M. M. de Faria |
IEEE Trans. Image Process. | 2 |
| 2015 | Sparse least-squares prediction for intra image codingabstractThis paper presents a new intra prediction method for efficient image coding, based on linear prediction and sparse representation concepts, denominated sparse least-squares prediction (SLSP). The proposed method uses a low order linear approximation model which may be built inside a predefined large causal region. The high flexibility of the SLSP filter context allows the inclusion of more significant image features into the model for better prediction results. Experiments using an implementation of the proposed method in the state-of-the-art H.265/HEVC algorithm have shown that SLSP is able to improve the coding performance, specially in the presence of complex textures, achieving higher coding gains than other existing intra linear prediction methods. Luis F. R. Lucas, Nuno M. M. Rodrigues, Carla L. Pagliari, Eduardo A. B. da Silva, Sérgio M. M. de Faria |
ICIP | 2 |
| 2015 | Contributions to lossless coding of medical images using minimum rate predictorsabstractMedical imaging compression is experiencing a growth in terms of usage and image resolution, namely in diagnostics systems that require a large set of images, like MRI or CT. Furthermore, legal and diagnosis restrictions impose the use of lossless compression and data archival for several years. These facts create a demand for more efficient compression tools, used for archiving and communication. In this work, we first evaluate the performance of traditional medical image compression algorithms against that of recent state of the art lossless image encoders. We then propose a method to improve the Minimum Rate Predictors lossless encoder, by exploiting inter picture redundancy in volumetric anatomical images. Results show that the proposed method is more efficient than state of the art encoders, such as HEVC, by about 28.8%, and achieves a gain of up to 57.8% in compression ratio when compared with traditional methods. João M. Santos 0002, André F. R. Guarda, Nuno M. M. Rodrigues, Sérgio M. M. de Faria |
ICIP | 3 |
| 2015 | Intra Predictive Depth Map Coding Using Flexible Block PartitioningabstractA complete encoding solution for efficient intra-based depth map compression is proposed in this paper. The algorithm, denominated predictive depth coding (PDC), was specifically developed to efficiently represent the characteristics of depth maps, mostly composed by smooth areas delimited by sharp edges. At its core, PDC involves a directional intra prediction framework and a straightforward residue coding method, combined with an optimized flexible block partitioning scheme. In order to improve the algorithm in the presence of depth edges that cannot be efficiently predicted by the directional modes, a constrained depth modeling mode, based on explicit edge representation, was developed. For residue coding, a simple and low complexity approach was investigated, using constant and linear residue modeling, depending on the prediction mode. The performance of the proposed intra depth map coding approach was evaluated based on the quality of the synthesized views using the encoded depth maps and original texture views. The experimental tests based on all intra configuration demonstrated the superior rate-distortion performance of PDC, with average bitrate savings of 6%, when compared with the current state-of-the-art intra depth map coding solution present in the 3D extension of a high-efficiency video coding (3D-HEVC) standard. By using view synthesis optimization in both PDC and 3D-HEVC encoders, the average bitrate savings increase to 14.3%. This suggests that the proposed method, without using transform-based residue coding, is an efficient alternative to the current 3D-HEVC algorithm for intra depth map coding. Luis F. R. Lucas, Krzysztof Wegner, Nuno M. M. Rodrigues, Carla L. Pagliari, Eduardo A. B. da Silva, Sérgio M. M. de Faria |
IEEE Trans. Image Process. | 3 |
| 2014 | Optimizing Memory Usage and Accesses on CUDA-Based Recurrent Pattern Matching Image Compression
Patrício Domingues, João Silva 0001, Nuno M. M. Rodrigues, Murilo B. de Carvalho, Sérgio M. M. de Faria |
ICCSA (4) | 4 |
| 2013 | Predictive depth map coding for efficient virtual view synthesisabstractThis paper presents a novel approach to compress depth maps envisioned for virtual view synthesis. This proposal uses a sophisticated prediction model, combining the HEVC intra prediction modes with a flexible partitioning scheme. It exhaustively evaluates the prediction modes for a large amount of block sizes, in order to find the minimum coding cost for each depth map block. Unlike HEVC, no transform is used, the residue being trivially encoded through the transmission of just its mean value. The experimental results show that, when the encoding evaluation metric is the quality of the view synthesized using the encoded depth map against the map encoding rate, the proposed algorithm generates reconstructed depth maps that provide, for most bitrates, some of the best performances among state-of-the-art depth maps encoders. In addition, it runs approximately as fast as the HEVC HM. Luis F. R. Lucas, Nuno M. M. Rodrigues, Carla L. Pagliari, Eduardo A. B. da Silva, Sérgio M. M. de Faria |
ICIP | 2 |
| 2013 | Video compression using 3D multiscale recurrent patternsabstractIn this paper, we propose a new 3D pattern matching based video compression algorithm. Spatiotemporal prediction tools are used to exploit both the temporal and the spatial redundancies, and the resulting residue is encoded using a 3D data coding extension of the Multidimensional Multiscale Parser (MMP) algorithm. MMP was originally proposed as a generic lossy data compression algorithm. A high degree of adaptivity and versatility allowed it to be competitive with state-of-the-art transform-based compression methods for a wide range of applications and data sets. The performance of the proposed algorithm is compared with that of H.264/AVC, achieving close results for some data sets, which indicates the potential of the pattern matching paradigm as an alternative to the traditional hybrid video codecs. Nelson C. Francisco, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria |
ISCAS | 2 |
| 2013 | Lossy and lossless image encoding using multi-scale recurrent pattern matchingabstractIn this study, the authors investigate the use of multi‐scale recurrent pattern matching paradigm for lossless image compression. The multi‐scale multidimensional parser (MMP) algorithm is a successful implementation of this paradigm for lossy image compression, and can naturally perform lossless compression since it was first derived from a Lempel–Ziv lossless scheme. However, neither its recently adopted coding tools had been adapted for lossless coding nor a thorough analysis of its performance had been carried out. In this work, the authors evaluate MMP's lossless compression capability, proposing modifications for some of its predictions modes, as well as the inclusion of an adaptive prediction mode based on least squares. The residual information is also coded with well‐known techniques used in lossless compression. Experimental results for MMP show that the algorithm achieves a good performance for images such as computed generated graphics and scanned documents, whereas keeping a competitive performance for natural images. Since the algorithm's structure is exactly the same for lossless and lossy compression, the obtained results suggest that MMP is able to achieve a high compression performance for a wide range of images and rates, from lossy to lossless, without any prior analysis of the image to be coded. Danilo B. Graziosi, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria |
IET Image Process. | 2 |
| 2012 | Efficient depth map coding using linear residue approximation and a flexible prediction frameworkabstractThe importance to develop more efficient 3D and multiview data representation algorithms results from the recent market growth for 3D video equipments and associated services. One of the most investigated formats is video+depth which uses depth image based rendering (DIBR) to combine the information of texture and depth, in order to create an arbitrary number of views in the decoder. Such approach requires that depth information must be accurately encoded. However, methods usually employed to encode texture do not seem to be suitable for depth map coding. Luis F. R. Lucas, Nuno M. M. Rodrigues, Carla L. Pagliari, Eduardo A. B. da Silva, Sérgio M. M. de Faria |
ICIP | 2 |
| 2012 | A generic post-deblocking filter for block based image compression algorithms
Nelson C. Francisco, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Sérgio M. M. de Faria |
Signal Process. Image Commun. | 2 |
| 2012 | Efficient Recurrent Pattern Matching Video CodingabstractIn this paper, we propose a pattern-matching-based algorithm for video compression. This algorithm, named multidimensional multiscale parser (MMP)-Video, is based on the H.264/AVC video encoder, but uses a pattern-matching paradigm instead of the state-of-the-art transform-quantization-entropy encoding approach. The proposed method adopts the use of multiscale recurrent patterns to compress both spatial and temporal prediction residues, totally replacing the use of transforms and quantization. Experimental results show that the coding performance of MMP-Video is better than the one of H.264/AVC high profile, especially for medium to high bit-rates. The gains range up to 0.7 dB, showing that, in spite of its larger computational complexity, the use of multiscale recurrent pattern matching paradigm deserves being investigated as an alternative for video compression. Nelson C. Francisco, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Adaptive least squares prediction for stereo image codingabstractState-of-the art approaches towards stereo image coding exploit inter-view redundancy by employing block-matching methods for disparity estimation and compensation. However, the efficiency of these methods is affected by mismatched areas, due to occlusions, brightness variations, or perspective distortion between objects of the two views. In this paper we present a new prediction scheme for stereo image coding, that combines an implicit disparity estimation method, with an adaptive least squares (LS)-based filtering. The Multidimensional Multiscale Parser image coding algorithm was used to evaluate the efficiency of the proposed scheme. Experimental results demonstrate the advantage of LS prediction in stereo image coding. Furthermore, the rate-distortion performance of the MMP based stereo encoder is well above that of the state-of-the-art H.264/AVC Stereo Profile, especially at medium and high bit rates. Luis F. R. Lucas, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Sérgio M. M. de Faria |
ICIP | 2 |
| 2010 | Intra-prediction for color image coding using YUV correlationabstractIn this paper we present a new algorithm for chroma prediction in YUV images, based on inter component correlation. Despite the YUV color space transformation for inter component decorrelation, some dependency still exists between the Y, U and V chroma components. This dependency has been previously used to predict the chrominance data from the reconstructed luminance. In this paper we show that a chrominance component can be more efficiently predicted by using the reconstructed data from both the luminance and the remaining chrominance signal. The proposed chroma prediction is implemented and tested using the Multidimensional Multiscale Parser (MMP) image encoding algorithm. It is shown that the new color prediction mode outperforms the originally proposed prediction methods. Furthermore, by using the new color prediction scheme, MMP is consistently better than the state-of-the-art H.264/AVC for coding both for the luminance and the chrominance image components. Luis F. R. Lucas, Nuno M. M. Rodrigues, Sérgio M. M. de Faria, Eduardo A. B. da Silva, Murilo B. de Carvalho, Vítor Silva 0001 |
ICIP | 2 |
| 2010 | Subjective assessment of frame loss concealment methods in 3D videoabstractThis paper investigates the subjective impact resulting from different concealment methods for coping with lost frames in 3D video communication systems. It is assumed that a high priority channel is assigned to the main view and only the auxiliary view is subject to either transmission errors or packet loss, leading to missing frames at the decoder output. Three methods are used for frame concealment under different loss ratios. The results show that depth is well perceived by users and the subjective impact of frame loss not only depends on the concealment method but also exhibits high correlation with the disparity of the original sequence. It is also shown that under heavy loss conditions it is better to switch from 3D to 2D rather than presenting concealed 3D video to users. João Carreira 0003, Luís Pinto 0003, Nuno M. M. Rodrigues, Sérgio M. M. de Faria, Pedro A. Amado Assunção |
PCS | 3 |
| 2010 | Multiscale recurrent pattern matching approach for depth map codingabstractIn this article we propose to compress depth maps using a coding scheme based on multiscale recurrent pattern matching and evaluate its impact on depth image based rendering (DIBR). Depth maps are usually converted into gray scale images and compressed like a conventional luminance signal. However, using traditional transform-based encoders to compress depth maps may result in undesired artifacts at sharp edges due to the quantization of high frequency coefficients. The Multidimensional Multiscale Parser (MMP) is a pattern matching-based encoder, that is able to preserve and efficiently encode high frequency patterns, such as edge information. This ability is critical for encoding depth map images. Experimental results for encoding depth maps show that MMP is much more efficient in a rate-distortion sense than standard image compression techniques such as JPEG2000 or H.264/AVC. In addition, the depth maps compressed with MMP generate reconstructed views with a higher quality than all other tested compression algorithms. Danilo B. Graziosi, Nuno M. M. Rodrigues, Carla L. Pagliari, Eduardo A. B. da Silva, Sérgio M. M. de Faria, Marcelo M. Perez, Murilo B. de Carvalho |
PCS | 2 |
| 2010 | Scanned Compound Document Encoding Using Multiscale Recurrent PatternsabstractIn this paper, we propose a new encoder for scanned compound documents, based upon a recently introduced coding paradigm called multidimensional multiscale parser (MMP). MMP uses approximate pattern matching, with adaptive multiscale dictionaries that contain concatenations of scaled versions of previously encoded image blocks. These features give MMP the ability to adjust to the input image's characteristics, resulting in high coding efficiencies for a wide range of image types. This versatility makes MMP a good candidate for compound digital document encoding. The proposed algorithm first classifies the image blocks as smooth (texture) and nonsmooth (text and graphics). Smooth and nonsmooth blocks are then compressed using different MMP-based encoders, adapted for encoding either type of blocks. The adaptive use of these two types of encoders resulted in performance gains over the original MMP algorithm, further increasing the performance advantage over the current state-of-the-art image encoders for scanned compound images, without compromising the performance for other image types. Nelson C. Francisco, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001 |
IEEE Trans. Image Process. | 2 |
| 2009 | Improving multiscale recurrent pattern image coding with least-squares prediction modeabstractThe Multidimensional Multiscale Parser-based (MMP) image coding algorithm, when combined with flexible partitioning and predictive coding techniques (MMP-FP), provides state-of-the-art performance. In this paper we investigate the use of adaptive least-squares prediction in MMP. The linear prediction coefficients implicitly embed the local texture characteristics, and are computed based on a block's causal neighborhood (composed of already reconstructed data). Thus, the intra prediction mode is adaptively adjusted according to the local context and no extra overhead is needed for signaling the coefficients. We add this new context-adaptive linear prediction mode to the other MMP prediction modes, that are based on the ones used in H.264/AVC; the best mode is chosen through rate-distortion optimization. Simulation results show that least-squares prediction is able to significantly increase MMP-FPs rate-distortion performance for smooth images, leading to better results than the ones of state-of-theart, transform-based methods. Yet with the addition of least-squares prediction MMP-FP presents no performance loss when used for encoding non-smooth images, such as text and graphics. Danilo B. Graziosi, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Sérgio M. M. de Faria, Murilo B. de Carvalho |
ICIP | 2 |
| 2008 | Multiscale recurrent pattern image coding with a flexible partition schemeabstractIn this paper we present a new segmentation method for the multidimensional multiscale parser (MMP) algorithm. In previous works we have shown that, for text and compound images, MMP has better compression efficiency than state-of-the-art transform-based encoders like JPEG2000 and H.264/AVC; however, it is still inferior to them for smooth images. In this paper we improve the performance of MMP for smooth images by employing a more flexible block segmentation scheme than the one defined in the original algorithm. The new partition scheme allows MMP to exploit the image's structure in a much more adaptive and effective way. Experimental tests have shown consistent performance gains, mainly for smooth images. When employing the new block segmentation scheme, MMP outperforms the state-of-the-art JPEG2000 and H.264/AVC Intra-frame image coding algorithms for both smooth and non-smooth images, at low to medium compression ratios. Nelson C. Francisco, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001, Manuel J. C. S. Reis |
ICIP | 2 |
| 2008 | On Dictionary Adaptation for Recurrent Pattern Image CodingabstractIn this paper, we exploit a recently introduced coding algorithm called multidimensional multiscale parser (MMP) as an alternative to the traditional transform quantization-based methods. MMP uses approximate pattern matching with adaptive multiscale dictionaries that contain concatenations of scaled versions of previously encoded image blocks. We propose the use of predictive coding schemes that modify the source's probability distribution, in order to favour the efficiency of MMP's dictionary adaptation. Statistical conditioning is also used, allowing for an increased coding efficiency of the dictionaries' symbols. New dictionary design methods, that allow for an effective compromise between the introduction of new dictionary elements and the reduction of codebook redundancy, are also proposed. Experimental results validate the proposed techniques by showing consistent improvements in PSNR performance over the original MMP algorithm. When compared with state-of-the-art methods, like JPEG2000 and H.264/AVC, the proposed algorithm achieves relevant gains (up to 6 dB) for nonsmooth images and very competitive results for smooth images. These results strongly suggest that the new paradigm posed by MMP can be regarded as an alternative to the one traditionally used in image coding, for a wide range of image types. Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001 |
IEEE Trans. Image Process. | 1 |
| 2006 | Improving H.264/AVC Inter Compression with Multiscale Recurrent PatternsabstractIn this paper we describe the ongoing work on a new paradigm for compressing the motion predicted error in a video coder, referred to as MMP-Video. This new coding algorithm uses the multidimensional multiscale parser image coding algorithm to encode the residue error, in a H.264/AVC based video coder. MMP has shown to perform very well as a universal still image coding method, particularly when it is combined with intra prediction schemes. In addition, previously published preliminary results have also presented MMP as a promising video coding method. In this paper, we propose new dictionary updating techniques for MMP-video. Along with other functional optimizations, these techniques allow for a significant improvement in the encoder performance. Thus, we were able to achieve considerable gains over H.264/AVC for B slices, specially for medium and high bit-rates, while maintaining equivalent performance for the P slices. Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001 |
ICIP | 1 |
| 2006 | Efficient dictionary design for multiscale recurrent pattern image codingabstractMMP-Intra was recently proposed as a recurrent patterns based image encoder that combines the multidimensional multiscale parser (MMP) algorithm with intra prediction techniques. Our results show that this method is able to achieve considerable gains over state-of-the-art transform-based image encoders for a wide variety of types of images, like text, composed (text and graphics) and texture images, while having performance close to the one of traditional algorithms for smooth images. Because of this universal character, MMP-Intra can be regarded as a viable alternative to transform-based image coding. MMP-Intra uses a multiscale adaptive dictionary to approximate the original data blocks. It is composed of dilations, contractions and concatenations of previously encoded patterns. In this work we present a new method for controlling the dictionary adaptability, in which the dictionary is only updated if a certain distortion criterion is met in the block being encoded. Experimental results show that this scheme is able to consistently outperform the original method, while achieving relevant reductions in its computational complexity Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001, Frederico S. Pinagé |
ISCAS | 1 |
| 2005 | Universal image coding using multiscale recurrent patterns and predictionabstractIn this paper we present a new method for image coding that is able to achieve good results over a wide range of image types. This work is based on the multidimensional multiscale parser (MMP) algorithm (M. de Carvalho et al., 2002), allied with an intra frame image predictive coding scheme. MMP has been shown to have, for a large class of image data, including texts, graphics, mixed images and textures, a compression efficiency comparable (and, in several cases, well above) to the one of state-of-the-art encoders. However, for smooth grayscale images, its performance lags behind the one of wavelet-based encoders, as JPEG2000. In this paper we propose a novel encoder using MMP with intra predictive coding, similar to the one used in the H.264/AVC video coding standard. Experimental results show that this method closes the performance gap to JPEG-2000 for smooth images, with PSNR gains of up to 1.5 dB. Yet, it maintains the excellent performance level of the MMP for other types of image data, as text, graphics and compound images, lending it a useful universal character. Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001 |
ICIP (2) | 1 |
| 2001 | Hierarchical motion compensation with spatial and luminance transformationsabstractWe present a new method for motion compensation, which combines transformations in the spatial and luminance domains. In most standardised video compression algorithms, motion information is assumed to be translation only, as used by the traditional block matching algorithm (BMA). Such an assumption is not always correct, as in most scenes, namely head and shoulders, object motions are rotations and zooms, among other effects, which cannot be represented as translations. In order to compensate these complex motion features more accurately, we applied geometric transformations (BMGT) for motion compensation. Nevertheless, spatial transformations alone are unable to compensate situations like uncovered background, masking between objects or changes in lighting conditions. In this sense, we introduced an additional transformation in the luminance domain (BMGTI), which has proven to be appropriate to overcome such problems. Experimental results have shown that BMGTI exceeds the performance of BMGT by about 2 dB, and achieves better results than the global brightness compensation (GBC) technique. This method is implemented using a spatial hierarchical block structure, which allows video encoding at low bit rates and reduction of the computational complexity. Nuno M. M. Rodrigues, Vítor Silva 0001, Sérgio M. M. de Faria |
ICIP (3) | 1 |