Sérgio M. M. de Faria

dblp:29/1127 · also Sérgio Faria, Sérgio M. M. Faria, Sérgio M. Maciel de Faria · DBLP profile ↗
← Back
60ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-0993-9124ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 51 · 7 since 2021Systems, architecture and hardware · 5Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Fast adaptive QTMT partitioning for intra 360°video coding based on gradient boosted trees
Jose N. Filipe, Luis Tavora, Sérgio M. M. de Faria, Antonio Navarro 0002, Pedro A. Amado Assunção
J. Vis. Commun. Image Represent.3
2026 Iterative Occlusion-Aware Light Field Depth Estimation Using 4-D Geometrical Cues
abstract
Light field cameras and multi-camera arrays have emerged as promising solutions for accurately estimating depth by passively capturing light information. This is possible because the 3D information of a scene is embedded in the 4-D light field geometry. Commonly, depth estimation methods extract this information relying on gradient information, heuristic-based optimisation models, or learning-based approaches. This paper focuses mainly on explicitly understanding and exploiting 4-D geometrical cues for light field depth estimation. Thus, a novel method is proposed, based on a non-learning-based optimisation approach for depth estimation that explicitly considers surface normal accuracy and occlusion regions by utilising a fully explainable 4-D geometric model of the light field. The 4-D model performs depth/disparity estimation by determining the orientations and analysing the intersections of key 2D planes in 4-D space, which are the images of 3D-space points in the 4-D light field. Experimental results show that the proposed method outperforms both learning-based and non-learning-based state-of-the-art methods in terms of surface normal angle accuracy, achieving a Median Angle Error on planar surfaces, on average, 26.3% lower than the state-of-the-art, and still being competitive with state-of-the-art methods in terms of MSE $\boldsymbol {\times } 100$ and Badpix 0.07.
Rui Lourenço, Lucas A. Thomaz, Eduardo A. B. da Silva, Sérgio M. M. de Faria
IEEE Trans. Image Process.4
2024 Non-Separablewavelet Transform Using Learnable Convolutional Lifting Steps
abstract
Wavelet transforms have been a relevant topic in signal processing for many years. One of the most common strategies when designing wavelet transforms is the use of lifting schemes, known for their perfect reconstruction properties and flexible design. This paper introduces a novel 2D non-separable lifting design methodology based on deep learning architectures. The proposed method is assessed within the context of end-to-end lossless image compression. The attained results highlight a significant 99% reduction of the number of learnable coefficients, with a direct impact on the computational complexity, while still achieving higher compression performance when compared to other end-to-end state-of-the-art wavelet-based frameworks. The code for the models used to obtain the reported results are available at https://github.com/joaoparracho/2D-NSWT-LCLS.
Joao O. Parracho, Eduardo A. B. da Silva, Lucas A. Thomaz, Luis Tavora, Sérgio M. M. de Faria
ICIP5
2022 Melanoma classification using light-Fields with morlet scattering transform and CNN: Surface depth as a valuable tool to increase detection rate
Pedro M. M. Pereira, Lucas A. Thomaz, Luis Tavora, Pedro A. Amado Assunção, Rui Fonseca-Pinto, Rui Pedro Paiva, Sérgio M. M. de Faria
Medical Image Anal.7
2022 Hierarchical lossless coding of light fields with improved random access
João M. Santos 0002, Lucas A. Thomaz, Pedro A. Amado Assunção, Luís Alberto da Silva Cruz, Luis Tavora, Sérgio M. M. de Faria
Signal Process. Image Commun.6
2022 Lossless Coding of Light Fields Based on 4D Minimum Rate Predictors
abstract
Common representations of light fields use four-dimensional data structures, where a given pixel is closely related not only to its spatial neighbours within the same view, but also to its angular neighbours, co-located in adjacent views. Such structure presents increased redundancy between pixels, when compared with regular single-view images. Then, these redundancies are exploited to obtain compressed representations of the light field, using prediction algorithms specifically tailored to estimate pixel values based on both spatial and angular references. This paper proposes new encoding schemes which take advantage of the four-dimensional light field data structures to improve the coding performance of Minimum Rate Predictors. The proposed methods expand previous research on lossless coding beyond the current state-of-the-art. The experimental results, obtained using both traditional datasets and others more challenging, show bit-rate savings no smaller than 10%, when compared with existing methods for lossless light field compression.
João M. Santos 0002, Lucas A. Thomaz, Pedro A. Amado Assunção, Luís Alberto da Silva Cruz, Luis Tavora, Sérgio M. M. de Faria
IEEE Trans. Image Process.6
2021 Attention-driven tile splitting method for improved efficiency of omnidirectional versatile video coding
abstract
A common approach used in omnidirectional video coding is based on frame splitting into tiles, allowing partial delivery of only the subset of tiles that is necessary to render the user’s current viewing region, defined as a specific viewport or Field-of-View (FoV). Since tiles can be independently encoded, such mechanism provides a flexible solution for encoding planar representations with ultra-high definition (UHD), such as the Equirectangular Projection (ERP), using Versatile Video Coding (VVC). By only selecting and transmitting the coded data that is required to render the necessary FoV, rather than the full 360°, a great deal of bandwidth can be saved. While current solutions are based on splitting the omnidirectional video frames into tiles of equal size, this paper proposes a new approach based on adaptive tile size, driven by visual attention. Those regions where the visual attention is higher are partitioned in smaller tiles to obtain higher bit rate granularity, allowing to decode the most frequent FoVs with minimum out-of-FoV pixels and reduced bandwidth. Optimal tile boundaries are found by solving a lagrangian minimisation problem with a cost function that achieves the best tradeoff between the standard deviation and the average attention-weighted bit rate per tile. The experimental results show that an average of 7.17% and 17.73% of bit rate savings is obtained in comparison with conventional tilling methods for the commonly used FoVs of $90^{\circ} \times 90^{\circ}$ and $45^{\circ} \times 45^{\circ}$, respectively.
João Carreira 0003, Sérgio M. M. de Faria, Luis Tavora, Antonio Navarro 0002, Pedro A. Amado Assunção
ICIP2
2021 Light field image coding with flexible viewpoint scalability and random access
abstract
This paper proposes a novel light field image compression approach with viewpoint scalability and random access functionalities. Although current state-of-the-art image coding algorithms for light fields already achieve high compression ratios, there is a lack of support for such functionalities, which are important for ensuring compatibility with different displays/capturing devices, enhanced user interaction and low decoding delay. The proposed solution enables various encoding profiles with different flexible viewpoint scalability and random access capabilities, depending on the application scenario. When compared to other state-of-the-art methods, the proposed approach consistently presents higher bitrate savings (44% on average), namely when compared to pseudo-video sequence coding approach based on HEVC. Moreover, the proposed scalable codec also outperforms MuLE and WaSP verification models, achieving average bitrate saving gains of 37% and 47%, respectively. The various flexible encoding profiles proposed add fine control to the image prediction dependencies, which allow to exploit the tradeoff between coding efficiency and the viewpoint random access, consequently, decreasing the maximum random access penalties that range from 0.60 to 0.15, for lenslet and HDCA light fields.
Ricardo J. S. Monteiro, Nuno M. M. Rodrigues, Sérgio M. M. de Faria, Paulo J. L. Nunes
Signal Process. Image Commun.3
2020 Versatile Video Coding Of 360° Video Using Adaptive Resolution Change
abstract
Encoding 360° video with ultra high spatial resolution requires high bitrates to guarantee acceptable QoE in video delivery services. However, since in general the full Field-of-View (FoV), i.e., 360°, is not required at once, a great deal of bandwidth can be saved if only a limited FoV is delivered, according to content relevancy or/and user demand. This work addresses this problem using the concept of Adaptive Resolution Change (ARC) defined in the forthcoming Versatile Video Coding (VVC) standard, by dynamically mapping the full FoV into multiple video frames with different spatial resolutions. Those FoVs attracting more visual attention are encoded with higher resolution while the others are encoded with lower resolution, thus without compromising the visual quality and resolution of the most relevant regions. The simulation results show that the proposed adaptive coding scheme is able to deliver high quality video for the most relevant FoV at any time instant, achieving a maximum bitrate reduction of 37.2%.
João Carreira 0003, Sérgio M. M. de Faria, Luis Tavora, Antonio Navarro 0002, Pedro A. Amado Assunção
ICIP2
2020 Robust Depth Estimation From Multi-Focus Plenoptic Images
abstract
This paper describes a robust depth estimation algorithm for multi-focus plenoptic images. The main feature of the proposed method consists of a hybrid template matching scheme built-upon intensity and local phase information, which adapts to the blurriness of neighbouring lenslet microimages. By reducing the impact of defocusblur on the template matching accuracy, the proposed method efficiently handles the varying triangulation baseline over the depth-of-field, thus discarding the need for scene-related information such as the expected range of disparities. Experimental results demonstrate the robustness of the proposed method over the most used commercially available depth estimation algorithm, achieving a reduction of 73% on the depth estimation error.
Francisco Cunha, Lucas A. Thomaz, Luis Tavora, Pedro A. Amado Assunção, Rui Fonseca-Pinto, Sérgio M. M. de Faria
ICIP6
2020 Disparity compensation of light fields for improved efficiency in 4D transform-based encoders
abstract
Efficient light field encoders take advantage of the inherent 4D data structures to achieve high compression performance. This is accomplished by exploiting the redundancy of co-located pixels in different sub-aperture images (SAIs) through prediction and/or transform schemes to find a m ore compact representation of the signal. However, in image regions with higher disparity between SAIs, such scheme's performance tends to decrease, thus reducing the compression efficiency. This paper introduces a reversible pre-processing algorithm for disparity compensation that operates on the SAI domain of light field data. The proposed method contributes to improve the transform efficiency of the encoder, since the disparity-compensated data presents higher correlation between co-located image blocks. The experimental results show significant improvements in the compression performance of 4D light fields, achieving Bjontegaard delta rate gains of about 44% on average for MuLE codec using the 4D discrete cosine transform, when encoding High Density Camera Arrays (HDCA) light field images.
João M. Santos 0002, Lucas A. Thomaz, Pedro A. Amado Assunção, Luís Alberto da Silva Cruz, Luis Tavora, Sérgio M. M. de Faria
VCIP6
2020 Integer DCT Approximation With Arbitrary Size and Adjustable Precision
abstract
This letter proposes a method to obtain integer reversible discrete cosine transforms for generic transform-based coding schemes. The novelty of the proposed method, which is based on decomposition of the DCT-II matrix into two triangular and one diagonal matrices, is twofold: (i) the new matrices can be of arbitrary size, i.e., any square N × N dimension, thus suitable for applications where non power-of-2 dimensions are required; (ii) they can be designed with adjustable precision in a trade-off with the number of representation bits. Furthermore, improvements are also proposed over the base scheme to avoid numerical issues when working with large matrices and to obtain more reliable approximations. The performance evaluation demonstrate the effectiveness of the proposed transforms to approximate the coding gain capabilities of the original DCT-II.
Lucas A. Thomaz, Pedro A. Amado Assunção, Luis Tavora, Sérgio M. M. de Faria
IEEE Signal Process. Lett.4
2019 Lossless Compression of Light Fields Using Multi-reference Minimum Rate Predictors
abstract
This paper presents a method to improve the lossless compression efficiency of light field encoding based on Minimum Rate Predictors (MRP). The proposed method relies on the use of multiple references, either micro-images or sub-aperture images, to provide a richer set of correlated pixels for prediction. The results show better compression ratios than conventional versions of MRP, for both representation formats (micro-image and sub-aperture image arrays), achieving gains ranging from 16.9% to 21.2%. Furthermore, it is also shown that the proposed method consistently outperforms the state-of-the-art lossless encoders HEVC and JPEG-LS.
João M. Santos 0002, Pedro A. Amado Assunção, Luís Alberto da Silva Cruz, Luis Tavora, Rui Fonseca-Pinto, Sérgio M. M. de Faria
DCC6
2019 Versatile Video Coding of 360-Degree Video using Frame-Based FoV and Visual Attention
abstract
High quality omnidirectional video requires ultra high resolution formats encoded with very high bit rates to guarantee acceptable QoE in video delivery services. Since in general the full FoV, i.e., 360°, is not required at once by users, rather than agnostic encoding of the whole 360° video, this work proposes a flexible coding approach where the full FoV is mapped into video frames and efficiently encoded using intra-FoV prediction. By avoiding inter-FoV prediction, the proposed approach enables independent decoding of one or more FoVs extracted from a single compressed stream containing the full FoV video. To achieve improved quality in those FoVs which attract more visual attention, non-uniform coding is proposed for the new Versatile Video Coding standard (VVC), using perceptually-driven quantisation for each FoV. This strategy, makes use of visual attention maps to decrease the overall bit rate without compromising the quality of the most relevant regions. The simulation results show that the proposed coding mechanism achieves consistent quality gains in the relevant FoV without significant losses in the remaining ones. In comparison with the reference VVC, the proposed method is able to achieve average quality gains up to 1.56 dB and to efficiently adapt the coding parameters to the visual attention information.
João Carreira 0003, Sérgio M. M. de Faria, Luis Tavora, Antonio Navarro 0002, Pedro A. Amado Assunção
ISM2
2018 Silhouette Enhancement in Light Field Disparity Estimation Using the Structure Tensor
abstract
This paper presents a method to improve disparity maps computed from light field images using structure tensor methods, which tend to expand the borders of the occluding objects, enlarging their silhouette. The proposed method relies on the fact that such regions of the silhouette are defined by a mismatch between the disparity map edges, computed by the structure tensor, and those of the corresponding epipolar plane images (EPI) representing the light field. The proposed silhouette improvement method determines a correspondence between EPI edges and the disparity map edges, identifying the erroneous silhouette regions. The disparity map is corrected by using neighbouring values computed with high structure tensor reliability. The achieved results show that the disparity map is improved both around object edges and overall, reducing the MSE by 47.9% in comparison with other methods also based on the structure tensor.
Rui Lourenço, Pedro A. Amado Assunção, Luis Tavora, Rui Fonseca-Pinto, Sérgio M. M. de Faria
ICIP5
2018 Evaluation of Focus Metrics in Extended Depth-of-field Reconstruction
abstract
The performance of focus metrics in the evaluation of refocused images from light fields is investigated in this paper. To this aim, the paper presents a comprehensive study on the performance of a large set of different focus metrics (34 in total) in the evaluation of patch-based extended depth-of-field reconstructed images (e.g., all-in-focus). The new findings of this work demonstrate that optimal reconstruction of extended depth-of-field images, from light fields captured with focused plenoptic cameras, is not consistent with the focus level figures computed by available focus metrics. The results show that a higher focus level, as given by such computational methods, does not actually correspond to better all-in-focus images, thus indicating that currently available focus metrics are not suitable for evaluating the quality of extended depth-of-field reconstruction. The results obtained through subjective evaluation confirm that objective focus metrics fail to indicate the best focused images. Finally, the paper presents recommendations to adapt existing metrics to extended depth-of-field reconstruction in plenoptic imaging.
Jose N. Filipe, Luis Tavora, Pedro A. Amado Assunção, Rui Fonseca-Pinto, Sérgio M. M. de Faria
QoMEX5
2018 Lossless coding of light field images based on minimum-rate predictors
João M. Santos 0002, Pedro A. Amado Assunção, Luís Alberto da Silva Cruz, Luis Tavora, Rui Fonseca-Pinto, Sérgio M. M. de Faria
J. Vis. Commun. Image Represent.6
2018 A Two-Stage Approach for Robust HEVC Coding and Streaming
abstract
The increased compression ratios achieved by the High Efficiency Video Coding (HEVC) standard lead to reduced robustness of coded streams, with increased susceptibility to network errors and consequent video quality degradation. This paper proposes a method based on a two-stage approach to improve the error robustness of HEVC streaming, by reducing temporal error propagation in the case of frame loss. The prediction mismatch that occurs at the decoder after frame loss is reduced through the following two stages. First, at the encoding stage, the reference pictures are dynamically selected based on constraining conditions and Lagrangian optimization, which distributes the use of reference pictures, by reducing the number of prediction units that depend on a single reference. Second, at the streaming stage, a motion vector (MV) prioritization algorithm, based on spatial dependencies, selects an optimal subset of MVs to be transmitted, redundantly, as side information to reduce mismatched MV predictions at the decoder. The simulation results show that the proposed method significantly reduces the effect of temporal error propagation. Compared with the reference HEVC, the proposed reference picture selection method is able to improve the video quality at low-packet-loss rates (e.g., 1%) using the same bitrate, achieving quality gains up to 2.3 dB for 10% of packet loss ratio. It is shown, for instance, that the redundant MVs are able to boost the performance achieving quality gains of 3 dB when compared with the reference HEVC, at the cost using 4% increase in total bitrate.
João Carreira 0003, Pedro A. Amado Assunção, Sérgio M. M. de Faria, Erhan Ekmekcioglu, Ahmet M. Kondoz
IEEE Trans. Circuits Syst. Video Technol.3
2017 A robust video encoding scheme to enhance error concealment of intra frames
abstract
In this paper a robust encoding scheme is proposed to improve the visual quality of HEVC decoded video when intra frames are lost along the streaming path. For this purpose, the encoding process includes frame loss simulation and subsequent error concealment, to find the most efficient method that should be used by a decoder to recover lost intra frames. In this novel scheme, each image is divided into partitions, which are associated with the error concealment method that achieves the lowest distortion. Then this information is signalled to the decoder through SEI messages in the coded stream. In order to efficiently use the signalling overhead, rate-distortion optimisation is used to achieve the best trade-off between the number of transmitted symbols and distortion of reconstructed frames. Experimental results show the effectiveness of the proposed method to enhance the quality of reconstructed intra frames under different packet loss ratios (PLR). For PLR=10%, the robust coding scheme is able to improve the average PSNR of all frames affected by errors, up to 1.50 dB and 3.44 dB in Low-Delay and Random-Access configurations respectively, at a maximum overhead cost of 0.24%.
João Carreira 0003, Pedro A. Amado Assunção, Sérgio M. M. de Faria, Erhan Ekmekcioglu, Ahmet M. Kondoz
ISCAS3
2017 Energy-Efficient and Portable Least Squares Prediction for Image Coding on a Mobile GPU
abstract
Least squares prediction is a technique used to foresee pixel values during image coding by finding the minimum square error of neighbouring pixels. It has shown considerable quality gains especially for complex images with high variations in pixel intensities. The drawback of this technique consists of high computational complexity, consuming the most significant part of processing time and resources available, which makes it difficult to implement in fast, lossy image coders. One challenge is therefore to reduce the computational time of this predictor, namely through the use of new parallel programming techniques, making it more attractive for state-of-the-art coder-decoders. Also, new algorithmic propositions are made, trying to reduce the time spent in exchange for rate-distortion performance. These propositions are senseful since this predictor is used not only in lossless image coding, but also in lossy as well. Another aim of this article is to analyze energy efficiency among different types of platforms for this signal processing algorithm. Comparisons are provided on parallel computing processors ranging from very powerful Graphics Computing Units (GPUs) to mobile General-Purpose GPUs.
Pedro Cordeiro, Gabriel Falcão Paiva Fernandes, Patrício Domingues, Nuno M. M. Rodrigues, Sérgio M. M. de Faria
PDP5
2017 Spatial error concealment for intra-coded depth maps in multiview video-plus-depth
Pedro A. Amado Assunção, Sylvain Marcelino, Salviano F. S. P. Soares, Sérgio M. M. de Faria
Multim. Tools Appl.4
2017 A method to improve HEVC lossless coding of volumetric medical images
André F. R. Guarda, João M. Santos 0002, Luís Alberto da Silva Cruz, Pedro A. Amado Assunção, Nuno M. M. Rodrigues, Sérgio M. M. de Faria
Signal Process. Image Commun.6
2017 Lossless Compression of Medical Images Using 3-D Predictors
abstract
This paper describes a highly efficient method for lossless compression of volumetric sets of medical images, such as CTs or MRIs. The proposed method, referred to as 3-D-MRP, is based on the principle of minimum rate predictors (MRPs), which is one of the state-of-the-art lossless compression technologies presented in the data compression literature. The main features of the proposed method include the use of 3-D predictors, 3-D-block octree partitioning and classification, volume-based optimization, and support for 16-b-depth images. Experimental results demonstrate the efficiency of the 3-D-MRP algorithm for the compression of volumetric sets of medical images, achieving gains above 15% and 12% for 8- and 16-bit-depth contents, respectively, when compared with JPEG-LS, JPEG2000, CALIC, and HEVC, as well as other proposals based on the MRP algorithm.
Luis F. R. Lucas, Nuno M. M. Rodrigues, Luís Alberto da Silva Cruz, Sérgio M. M. de Faria
IEEE Trans. Medical Imaging4
2016 Optimizing GPU Code for CPU Execution Using OpenCL and Vectorization: A Case Study on Image Coding
Pedro M. M. Pereira, Patrício Domingues, Nuno M. M. Rodrigues, Gabriel Falcão Paiva Fernandes, Sérgio M. M. de Faria
ICA3PP5
2016 Compression of medical images using MRP with bi-directional prediction and histogram packing
abstract
Medical imaging technology has become essential for the improvement of medical practice. This led to advances in the technology, namely in image sampling resolutions, pixel bit-depth and inter slice resolution. Additionally, common use of medical images, the life expectancy of patients and legal restrictions led to increasing storage costs. Therefore, efficient compression of medical image data is in high demand, for archiving and transmission. In this work we propose to improve the compression efficiency of the Minimum Rate Predictors lossless encoder, by adding bi-directional prediction support and a histogram packing technique. The results show that the proposed method presents a higher compression efficiency than state-of-the-art HEVC encoder. The compression efficiency is improved by 20%, on average, when compared to HEVC and by 46.1% when compared with the original MRP algorithm.
João M. Santos 0002, André F. R. Guarda, Luís Alberto da Silva Cruz, Nuno M. M. Rodrigues, Sérgio M. M. de Faria
PCS5
2016 Reconstruction of lost depth data in multiview video-plus-depth communications using geometric transforms
Sylvain Marcelino, Salviano F. S. P. Soares, Sérgio M. M. de Faria, Pedro A. Amado Assunção
J. Vis. Commun. Image Represent.3
2016 Image Coding Using Generalized Predictors Based on Sparsity and Geometric Transformations
abstract
Directional intra prediction plays an important role in current state-of-the-art video coding standards. In directional prediction, neighbouring samples are projected along a specific direction to predict a block of samples. Ultimately, each prediction mode can be regarded as a set of very simple linear predictors, a different one for each pixel of a block. Therefore, a natural question that arises is whether one could use the theory of linear prediction in order to generate intra prediction modes that provide increased coding efficiency. However, such an interpretation of each directional mode as a set of linear predictors is too poor to provide useful insights for their design. In this paper, we introduce an interpretation of directional prediction as a particular case of linear prediction, which uses the first-order linear filters and a set of geometric transformations. This interpretation motivated the proposal of a generalized intra prediction framework, whereby the first-order linear filters are replaced by adaptive linear filters with sparsity constraints. In this context, we investigate the use of efficient sparse linear models, adaptively estimated for each block through the use of different algorithms, such as matching pursuit, least angle regression, least absolute shrinkage and selection operator, or elastic net. The proposed intra prediction framework was implemented and evaluated within the state-of-the-art high efficiency video coding standard. Experiments demonstrated the advantage of this predictive solution, mainly in the presence of images with complex features and textured areas, achieving higher average bitrate savings than other related sparse representation methods proposed in the literature.
Luis F. R. Lucas, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Carla L. Pagliari, Sérgio M. M. de Faria
IEEE Trans. Image Process.5
2015 Sparse least-squares prediction for intra image coding
abstract
This paper presents a new intra prediction method for efficient image coding, based on linear prediction and sparse representation concepts, denominated sparse least-squares prediction (SLSP). The proposed method uses a low order linear approximation model which may be built inside a predefined large causal region. The high flexibility of the SLSP filter context allows the inclusion of more significant image features into the model for better prediction results. Experiments using an implementation of the proposed method in the state-of-the-art H.265/HEVC algorithm have shown that SLSP is able to improve the coding performance, specially in the presence of complex textures, achieving higher coding gains than other existing intra linear prediction methods.
Luis F. R. Lucas, Nuno M. M. Rodrigues, Carla L. Pagliari, Eduardo A. B. da Silva, Sérgio M. M. de Faria
ICIP5
2015 Contributions to lossless coding of medical images using minimum rate predictors
abstract
Medical imaging compression is experiencing a growth in terms of usage and image resolution, namely in diagnostics systems that require a large set of images, like MRI or CT. Furthermore, legal and diagnosis restrictions impose the use of lossless compression and data archival for several years. These facts create a demand for more efficient compression tools, used for archiving and communication. In this work, we first evaluate the performance of traditional medical image compression algorithms against that of recent state of the art lossless image encoders. We then propose a method to improve the Minimum Rate Predictors lossless encoder, by exploiting inter picture redundancy in volumetric anatomical images. Results show that the proposed method is more efficient than state of the art encoders, such as HEVC, by about 28.8%, and achieves a gain of up to 57.8% in compression ratio when compared with traditional methods.
João M. Santos 0002, André F. R. Guarda, Nuno M. M. Rodrigues, Sérgio M. M. de Faria
ICIP4
2015 Reference picture selection using checkerboard pattern for resilient video coding
abstract
The improved compression efficiency achieved by the High Efficiency Video Coding (HEVC) standard has the counter-effect of decreasing error resilience in transmission over error-prone channels. To increase the error resilience of HEVC streams, this paper proposes a checkerboard reference picture selection method in order to reduce the prediction mismatch at the decoder in case of frame losses. The proposed approach not only allows to reduce the error propagation at the decoder, but also enhances the quality of reconstructed frames by selectively constraining the choice of reference pictures used for temporal prediction. The underlying approach is to increase the amount of accurate temporal information at the decoder when transmission errors occur, to improve the video quality by using an efficient combination of diverse motion fields. The proposed method compensates for the small loss of coding efficiency at frame loss rates as low as 3%. For a single frame-loss event the proposed method can achieve up to 2 dB of gain in the affected frames and an average quality gain of 0.84 dB for different error prone conditions.
João Carreira 0003, Pedro A. Amado Assunção, Sérgio M. M. de Faria, Erhan Ekmekcioglu, Ahmet M. Kondoz, H. Lim
VCIP3
2015 Intra Predictive Depth Map Coding Using Flexible Block Partitioning
abstract
A complete encoding solution for efficient intra-based depth map compression is proposed in this paper. The algorithm, denominated predictive depth coding (PDC), was specifically developed to efficiently represent the characteristics of depth maps, mostly composed by smooth areas delimited by sharp edges. At its core, PDC involves a directional intra prediction framework and a straightforward residue coding method, combined with an optimized flexible block partitioning scheme. In order to improve the algorithm in the presence of depth edges that cannot be efficiently predicted by the directional modes, a constrained depth modeling mode, based on explicit edge representation, was developed. For residue coding, a simple and low complexity approach was investigated, using constant and linear residue modeling, depending on the prediction mode. The performance of the proposed intra depth map coding approach was evaluated based on the quality of the synthesized views using the encoded depth maps and original texture views. The experimental tests based on all intra configuration demonstrated the superior rate-distortion performance of PDC, with average bitrate savings of 6%, when compared with the current state-of-the-art intra depth map coding solution present in the 3D extension of a high-efficiency video coding (3D-HEVC) standard. By using view synthesis optimization in both PDC and 3D-HEVC encoders, the average bitrate savings increase to 14.3%. This suggests that the proposed method, without using transform-based residue coding, is an efficient alternative to the current 3D-HEVC algorithm for intra depth map coding.
Luis F. R. Lucas, Krzysztof Wegner, Nuno M. M. Rodrigues, Carla L. Pagliari, Eduardo A. B. da Silva, Sérgio M. M. de Faria
IEEE Trans. Image Process.6
2014 Optimizing Memory Usage and Accesses on CUDA-Based Recurrent Pattern Matching Image Compression
Patrício Domingues, João Silva 0001, Nuno M. M. Rodrigues, Murilo B. de Carvalho, Sérgio M. M. de Faria
ICCSA (4)6
2014 Selective motion vector redundancies for improved error resilience in HEVC
abstract
This paper addresses the problem caused by motion vector coding dependencies on the error resilience performance of the emergent High Efficiency Video Coding (HEVC) standard. We propose a method based on the prediction dependency of motion vectors (MV) to select the most relevant ones for redundant coding with reduced overhead. The spatial dependencies are analysed in the encoder to prioritise the MVs that should be selected for redundancy, based on the number of subsequent dependent coding units. Then, a subset of prioritised MVs is transmitted as redundancy (referred to as side information in the paper), to reduce the use and propagation of mismatched MV predictions in case of transmission errors or data loss. The simulation results show that the proposed MV selection method can effectively identify the most relevant motion field, achieving improved error robustness with a reduced redundancy overhead. Exploiting only 30% of the generated MVs for redundancy, average quality gains of up to 1 dB are achieved compared to a uniform MV selection scheme, and up to 2 dB compared to the original HEVC standard with no redundant encoded information.
João Carreira 0003, Erhan Ekmekcioglu, Ahmet M. Kondoz, Pedro A. Amado Assunção, Sérgio M. M. de Faria, Varuna De Silva
ICIP5
2013 Predictive depth map coding for efficient virtual view synthesis
abstract
This paper presents a novel approach to compress depth maps envisioned for virtual view synthesis. This proposal uses a sophisticated prediction model, combining the HEVC intra prediction modes with a flexible partitioning scheme. It exhaustively evaluates the prediction modes for a large amount of block sizes, in order to find the minimum coding cost for each depth map block. Unlike HEVC, no transform is used, the residue being trivially encoded through the transmission of just its mean value. The experimental results show that, when the encoding evaluation metric is the quality of the view synthesized using the encoded depth map against the map encoding rate, the proposed algorithm generates reconstructed depth maps that provide, for most bitrates, some of the best performances among state-of-the-art depth maps encoders. In addition, it runs approximately as fast as the HEVC HM.
Luis F. R. Lucas, Nuno M. M. Rodrigues, Carla L. Pagliari, Eduardo A. B. da Silva, Sérgio M. M. de Faria
ICIP5
2013 Depth map concealment using interview warping vectors from geometric transforms
abstract
This paper deals with reconstruction of corrupted depth maps received by multiview video-plus-depth (MVD) decoders from error prone channels. An interview-based method is proposed using warping vectors obtained through a block matching approach with geometric transforms (BMGT) between two colour views. It is shown that BMGT is able to find efficient warping vectors for reconstruction of lost regions in the depth maps associated with the colour views. The proposed concealment method uses an additional contour reconstruction technique, applied to arbitrary shapes within the lost regions, which is used for weighted interpolation. In comparison to a reference method based on simple weighted interpolation, the proposed method is able to achieve PSNR gains in synthesised views up to 5.61 dB, at data loss ratios up to 40%.
Sylvain Marcelino, Pedro A. Amado Assunção, Sérgio M. M. de Faria, Salviano F. S. P. Soares
ICIP3
2013 Video compression using 3D multiscale recurrent patterns
abstract
In this paper, we propose a new 3D pattern matching based video compression algorithm. Spatiotemporal prediction tools are used to exploit both the temporal and the spatial redundancies, and the resulting residue is encoded using a 3D data coding extension of the Multidimensional Multiscale Parser (MMP) algorithm. MMP was originally proposed as a generic lossy data compression algorithm. A high degree of adaptivity and versatility allowed it to be competitive with state-of-the-art transform-based compression methods for a wide range of applications and data sets. The performance of the proposed algorithm is compared with that of H.264/AVC, achieving close results for some data sets, which indicates the potential of the pattern matching paradigm as an alternative to the traditional hybrid video codecs.
Nelson C. Francisco, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria
ISCAS5
2013 Lossy and lossless image encoding using multi-scale recurrent pattern matching
abstract
In this study, the authors investigate the use of multi‐scale recurrent pattern matching paradigm for lossless image compression. The multi‐scale multidimensional parser (MMP) algorithm is a successful implementation of this paradigm for lossy image compression, and can naturally perform lossless compression since it was first derived from a Lempel–Ziv lossless scheme. However, neither its recently adopted coding tools had been adapted for lossless coding nor a thorough analysis of its performance had been carried out. In this work, the authors evaluate MMP's lossless compression capability, proposing modifications for some of its predictions modes, as well as the inclusion of an adaptive prediction mode based on least squares. The residual information is also coded with well‐known techniques used in lossless compression. Experimental results for MMP show that the algorithm achieves a good performance for images such as computed generated graphics and scanned documents, whereas keeping a competitive performance for natural images. Since the algorithm's structure is exactly the same for lossless and lossy compression, the obtained results suggest that MMP is able to achieve a high compression performance for a wide range of images and rates, from lossy to lossless, without any prior analysis of the image to be coded.
Danilo B. Graziosi, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria
IET Image Process.5
2012 Efficient depth map coding using linear residue approximation and a flexible prediction framework
abstract
The importance to develop more efficient 3D and multiview data representation algorithms results from the recent market growth for 3D video equipments and associated services. One of the most investigated formats is video+depth which uses depth image based rendering (DIBR) to combine the information of texture and depth, in order to create an arbitrary number of views in the decoder. Such approach requires that depth information must be accurately encoded. However, methods usually employed to encode texture do not seem to be suitable for depth map coding.
Luis F. R. Lucas, Nuno M. M. Rodrigues, Carla L. Pagliari, Eduardo A. B. da Silva, Sérgio M. M. de Faria
ICIP5
2012 Lost block reconstruction in depth maps using color image contours
abstract
This paper presents a method to recover lost blocks in depth maps affected by data loss in 3D image/video communications over error prone networks. The proposed method relies on the color image for accurate reconstruction of the lost contour segments in the corresponding depth map areas. Such reconstructed depth map contours are then used as boundaries at different depth planes to recover the missing depth values through weighted interpolation. The method performance is evaluated by the objective quality (PSNR) of the synthesised views. Images decoded with the reconstructed depth maps are compared with those of the reference method. The proposed method exhibits PSNR gains up to 1.49dB higher than the reference one, and a better performance is consistently achieved for different 3D image content.
Sylvain Marcelino, Pedro A. Amado Assunção, Sérgio M. M. de Faria, Salviano F. S. P. Soares
PCS3
2012 A generic post-deblocking filter for block based image compression algorithms
Nelson C. Francisco, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Sérgio M. M. de Faria
Signal Process. Image Commun.4
2012 Efficient Recurrent Pattern Matching Video Coding
abstract
In this paper, we propose a pattern-matching-based algorithm for video compression. This algorithm, named multidimensional multiscale parser (MMP)-Video, is based on the H.264/AVC video encoder, but uses a pattern-matching paradigm instead of the state-of-the-art transform-quantization-entropy encoding approach. The proposed method adopts the use of multiscale recurrent patterns to compress both spatial and temporal prediction residues, totally replacing the use of transforms and quantization. Experimental results show that the coding performance of MMP-Video is better than the one of H.264/AVC high profile, especially for medium to high bit-rates. The gains range up to 0.7 dB, showing that, in spite of its larger computational complexity, the use of multiscale recurrent pattern matching paradigm deserves being investigated as an alternative for video compression.
Nelson C. Francisco, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria
IEEE Trans. Circuits Syst. Video Technol.5
2011 Adaptive least squares prediction for stereo image coding
abstract
State-of-the art approaches towards stereo image coding exploit inter-view redundancy by employing block-matching methods for disparity estimation and compensation. However, the efficiency of these methods is affected by mismatched areas, due to occlusions, brightness variations, or perspective distortion between objects of the two views. In this paper we present a new prediction scheme for stereo image coding, that combines an implicit disparity estimation method, with an adaptive least squares (LS)-based filtering. The Multidimensional Multiscale Parser image coding algorithm was used to evaluate the efficiency of the proposed scheme. Experimental results demonstrate the advantage of LS prediction in stereo image coding. Furthermore, the rate-distortion performance of the MMP based stereo encoder is well above that of the state-of-the-art H.264/AVC Stereo Profile, especially at medium and high bit rates.
Luis F. R. Lucas, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Sérgio M. M. de Faria
ICIP4
2011 Error recovery of image-based depth maps using Bézier curve fitting
abstract
This paper proposes a method to recover lost regions in image-based depth maps used in video plus depth 3D format. This method performs depth maps reconstruction taking into account depth contours within the lost regions. This is achieved by extracting the contours and recovering their lost segments based on Bézier curve fitting, followed by spatial interpolation. The proposed method maintains contour smoothness and uses them as the boundary limits of homogeneous depth regions, which are then filled through weighted pixel interpolation. The experimental results show that the proposed method yields better synthesized images than classic spatial concealment methods, uniquely based on pixel interpolation techniques. The method presented in this paper is able to outperform the reference method, in terms of PSNR by up to 1.91dB. The subjective quality is also shown as being significantly better.
Sylvain Marcelino, Pedro A. Amado Assunção, Sérgio M. M. de Faria, Salviano F. S. P. Soares
ICIP3
2010 Intra-prediction for color image coding using YUV correlation
abstract
In this paper we present a new algorithm for chroma prediction in YUV images, based on inter component correlation. Despite the YUV color space transformation for inter component decorrelation, some dependency still exists between the Y, U and V chroma components. This dependency has been previously used to predict the chrominance data from the reconstructed luminance. In this paper we show that a chrominance component can be more efficiently predicted by using the reconstructed data from both the luminance and the remaining chrominance signal. The proposed chroma prediction is implemented and tested using the Multidimensional Multiscale Parser (MMP) image encoding algorithm. It is shown that the new color prediction mode outperforms the originally proposed prediction methods. Furthermore, by using the new color prediction scheme, MMP is consistently better than the state-of-the-art H.264/AVC for coding both for the luminance and the chrominance image components.
Luis F. R. Lucas, Nuno M. M. Rodrigues, Sérgio M. M. de Faria, Eduardo A. B. da Silva, Murilo B. de Carvalho, Vítor Silva 0001
ICIP3
2010 Efficient MV prediction for zonal search in video transcoding
abstract
This paper proposes a method to efficiently find motion vector predictions for zonal search motion re-estimation in fast video transcoders. The motion information extracted from the incoming video stream is processed to generate accurate motion vector predictions for transcoding with reduced complexity. Our results demonstrate that motion vector predictions computed by the proposed method outperform those generated by the highly efficient EPZS (Enhanced Predictive Zonal Search) algorithm in H.264/AVC transcoders. The computational complexity is reduced up to 59.6% at negligible cost in R-D performance. The proposed method can be useful in multimedia systems and applications using any type of transcoder, such as transrating and/or spatial resolution downsizing.
Sylvain Marcelino, Sérgio M. M. de Faria, Pedro A. Amado Assunção, Sandro Moiron, Mohammed Ghanbari 0001
MMSP2
2010 Subjective assessment of frame loss concealment methods in 3D video
abstract
This paper investigates the subjective impact resulting from different concealment methods for coping with lost frames in 3D video communication systems. It is assumed that a high priority channel is assigned to the main view and only the auxiliary view is subject to either transmission errors or packet loss, leading to missing frames at the decoder output. Three methods are used for frame concealment under different loss ratios. The results show that depth is well perceived by users and the subjective impact of frame loss not only depends on the concealment method but also exhibits high correlation with the disparity of the original sequence. It is also shown that under heavy loss conditions it is better to switch from 3D to 2D rather than presenting concealed 3D video to users.
João Carreira 0003, Luís Pinto 0003, Nuno M. M. Rodrigues, Sérgio M. M. de Faria, Pedro A. Amado Assunção
PCS4
2010 Multiscale recurrent pattern matching approach for depth map coding
abstract
In this article we propose to compress depth maps using a coding scheme based on multiscale recurrent pattern matching and evaluate its impact on depth image based rendering (DIBR). Depth maps are usually converted into gray scale images and compressed like a conventional luminance signal. However, using traditional transform-based encoders to compress depth maps may result in undesired artifacts at sharp edges due to the quantization of high frequency coefficients. The Multidimensional Multiscale Parser (MMP) is a pattern matching-based encoder, that is able to preserve and efficiently encode high frequency patterns, such as edge information. This ability is critical for encoding depth map images. Experimental results for encoding depth maps show that MMP is much more efficient in a rate-distortion sense than standard image compression techniques such as JPEG2000 or H.264/AVC. In addition, the depth maps compressed with MMP generate reconstructed views with a higher quality than all other tested compression algorithms.
Danilo B. Graziosi, Nuno M. M. Rodrigues, Carla L. Pagliari, Eduardo A. B. da Silva, Sérgio M. M. de Faria, Marcelo M. Perez, Murilo B. de Carvalho
PCS5
2010 Scanned Compound Document Encoding Using Multiscale Recurrent Patterns
abstract
In this paper, we propose a new encoder for scanned compound documents, based upon a recently introduced coding paradigm called multidimensional multiscale parser (MMP). MMP uses approximate pattern matching, with adaptive multiscale dictionaries that contain concatenations of scaled versions of previously encoded image blocks. These features give MMP the ability to adjust to the input image's characteristics, resulting in high coding efficiencies for a wide range of image types. This versatility makes MMP a good candidate for compound digital document encoding. The proposed algorithm first classifies the image blocks as smooth (texture) and nonsmooth (text and graphics). Smooth and nonsmooth blocks are then compressed using different MMP-based encoders, adapted for encoding either type of blocks. The adaptive use of these two types of encoders resulted in performance gains over the original MMP algorithm, further increasing the performance advantage over the current state-of-the-art image encoders for scanned compound images, without compromising the performance for other image types.
Nelson C. Francisco, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001
IEEE Trans. Image Process.5
2009 Improving multiscale recurrent pattern image coding with least-squares prediction mode
abstract
The Multidimensional Multiscale Parser-based (MMP) image coding algorithm, when combined with flexible partitioning and predictive coding techniques (MMP-FP), provides state-of-the-art performance. In this paper we investigate the use of adaptive least-squares prediction in MMP. The linear prediction coefficients implicitly embed the local texture characteristics, and are computed based on a block's causal neighborhood (composed of already reconstructed data). Thus, the intra prediction mode is adaptively adjusted according to the local context and no extra overhead is needed for signaling the coefficients. We add this new context-adaptive linear prediction mode to the other MMP prediction modes, that are based on the ones used in H.264/AVC; the best mode is chosen through rate-distortion optimization. Simulation results show that least-squares prediction is able to significantly increase MMP-FPs rate-distortion performance for smooth images, leading to better results than the ones of state-of-theart, transform-based methods. Yet with the addition of least-squares prediction MMP-FP presents no performance loss when used for encoding non-smooth images, such as text and graphics.
Danilo B. Graziosi, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Sérgio M. M. de Faria, Murilo B. de Carvalho
ICIP4
2009 Video transcoding from H.264/AVC to MPEG-2 with reduced computational complexity
Sandro Moiron, Sérgio M. M. de Faria, Antonio Navarro 0002, Vítor Silva 0001, Pedro A. Amado Assunção
Signal Process. Image Commun.2
2008 Multiscale recurrent pattern image coding with a flexible partition scheme
abstract
In this paper we present a new segmentation method for the multidimensional multiscale parser (MMP) algorithm. In previous works we have shown that, for text and compound images, MMP has better compression efficiency than state-of-the-art transform-based encoders like JPEG2000 and H.264/AVC; however, it is still inferior to them for smooth images. In this paper we improve the performance of MMP for smooth images by employing a more flexible block segmentation scheme than the one defined in the original algorithm. The new partition scheme allows MMP to exploit the image's structure in a much more adaptive and effective way. Experimental tests have shown consistent performance gains, mainly for smooth images. When employing the new block segmentation scheme, MMP outperforms the state-of-the-art JPEG2000 and H.264/AVC Intra-frame image coding algorithms for both smooth and non-smooth images, at low to medium compression ratios.
Nelson C. Francisco, Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001, Manuel J. C. S. Reis
ICIP5
2008 On Dictionary Adaptation for Recurrent Pattern Image Coding
abstract
In this paper, we exploit a recently introduced coding algorithm called multidimensional multiscale parser (MMP) as an alternative to the traditional transform quantization-based methods. MMP uses approximate pattern matching with adaptive multiscale dictionaries that contain concatenations of scaled versions of previously encoded image blocks. We propose the use of predictive coding schemes that modify the source's probability distribution, in order to favour the efficiency of MMP's dictionary adaptation. Statistical conditioning is also used, allowing for an increased coding efficiency of the dictionaries' symbols. New dictionary design methods, that allow for an effective compromise between the introduction of new dictionary elements and the reduction of codebook redundancy, are also proposed. Experimental results validate the proposed techniques by showing consistent improvements in PSNR performance over the original MMP algorithm. When compared with state-of-the-art methods, like JPEG2000 and H.264/AVC, the proposed algorithm achieves relevant gains (up to 6 dB) for nonsmooth images and very competitive results for smooth images. These results strongly suggest that the new paradigm posed by MMP can be regarded as an alternative to the one traditionally used in image coding, for a wide range of image types.
Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001
IEEE Trans. Image Process.4
2007 Fast Interframe Transcoding from H.264 to MPEG-2
abstract
This paper deals with conversion from H.264/AVC (advanced video coding) coded video into the MPEG-2 format. The proposed approach exploits similarities between the coding techniques used in both standards in order to achieve a computationally efficient method for transcoding interframe coded slices. The conversion process is based on adaptation of both coding mode and motion information embedded in the H.264/AVC video stream, such that subsequent MPEG-2 encoding takes full advantage of the higher computational effort spent on the first encoding step. The proposed transcoding scheme significantly reduces the computational complexity needed for MPEG-2 interframe coding by reusing relevant information from the H.264 bitstream. The simulation results show that computational complexity savings up to 60%, with a marginal objective quality cost, can be achieved in comparison with a cascaded decoder-encoder.
Sandro Moiron, Sérgio M. M. de Faria, Pedro A. Amado Assunção, Vítor Silva 0001, Antonio Navarro 0002
ICIP (4)2
2006 Improving H.264/AVC Inter Compression with Multiscale Recurrent Patterns
abstract
In this paper we describe the ongoing work on a new paradigm for compressing the motion predicted error in a video coder, referred to as MMP-Video. This new coding algorithm uses the multidimensional multiscale parser image coding algorithm to encode the residue error, in a H.264/AVC based video coder. MMP has shown to perform very well as a universal still image coding method, particularly when it is combined with intra prediction schemes. In addition, previously published preliminary results have also presented MMP as a promising video coding method. In this paper, we propose new dictionary updating techniques for MMP-video. Along with other functional optimizations, these techniques allow for a significant improvement in the encoder performance. Thus, we were able to achieve considerable gains over H.264/AVC for B slices, specially for medium and high bit-rates, while maintaining equivalent performance for the P slices.
Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001
ICIP4
2006 Efficient dictionary design for multiscale recurrent pattern image coding
abstract
MMP-Intra was recently proposed as a recurrent patterns based image encoder that combines the multidimensional multiscale parser (MMP) algorithm with intra prediction techniques. Our results show that this method is able to achieve considerable gains over state-of-the-art transform-based image encoders for a wide variety of types of images, like text, composed (text and graphics) and texture images, while having performance close to the one of traditional algorithms for smooth images. Because of this universal character, MMP-Intra can be regarded as a viable alternative to transform-based image coding. MMP-Intra uses a multiscale adaptive dictionary to approximate the original data blocks. It is composed of dilations, contractions and concatenations of previously encoded patterns. In this work we present a new method for controlling the dictionary adaptability, in which the dictionary is only updated if a certain distortion criterion is met in the block being encoded. Experimental results show that this scheme is able to consistently outperform the original method, while achieving relevant reductions in its computational complexity
Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001, Frederico S. Pinagé
ISCAS4
2005 Universal image coding using multiscale recurrent patterns and prediction
abstract
In this paper we present a new method for image coding that is able to achieve good results over a wide range of image types. This work is based on the multidimensional multiscale parser (MMP) algorithm (M. de Carvalho et al., 2002), allied with an intra frame image predictive coding scheme. MMP has been shown to have, for a large class of image data, including texts, graphics, mixed images and textures, a compression efficiency comparable (and, in several cases, well above) to the one of state-of-the-art encoders. However, for smooth grayscale images, its performance lags behind the one of wavelet-based encoders, as JPEG2000. In this paper we propose a novel encoder using MMP with intra predictive coding, similar to the one used in the H.264/AVC video coding standard. Experimental results show that this method closes the performance gap to JPEG-2000 for smooth images, with PSNR gains of up to 1.5 dB. Yet, it maintains the excellent performance level of the MMP for other types of image data, as text, graphics and compound images, lending it a useful universal character.
Nuno M. M. Rodrigues, Eduardo A. B. da Silva, Murilo B. de Carvalho, Sérgio M. M. de Faria, Vítor Silva 0001
ICIP (2)4
2001 Hierarchical motion compensation with spatial and luminance transformations
abstract
We present a new method for motion compensation, which combines transformations in the spatial and luminance domains. In most standardised video compression algorithms, motion information is assumed to be translation only, as used by the traditional block matching algorithm (BMA). Such an assumption is not always correct, as in most scenes, namely head and shoulders, object motions are rotations and zooms, among other effects, which cannot be represented as translations. In order to compensate these complex motion features more accurately, we applied geometric transformations (BMGT) for motion compensation. Nevertheless, spatial transformations alone are unable to compensate situations like uncovered background, masking between objects or changes in lighting conditions. In this sense, we introduced an additional transformation in the luminance domain (BMGTI), which has proven to be appropriate to overcome such problems. Experimental results have shown that BMGTI exceeds the performance of BMGT by about 2 dB, and achieves better results than the global brightness compensation (GBC) technique. This method is implemented using a spatial hierarchical block structure, which allows video encoding at low bit rates and reduction of the computational complexity.
Nuno M. M. Rodrigues, Vítor Silva 0001, Sérgio M. M. de Faria
ICIP (3)3
1995 Motion compensation for very low bit-rate video
Mohammed Ghanbari 0001, Sérgio M. M. de Faria, I. N. Goh, K. T. Tan
Signal Process. Image Commun.2
1990 Parallel architecture for real-time video communications
abstract
A video codec based on several parallel digital signal processors is described. The digital signal processors (DSPs) can be easily programmed to implement the H.261 algorithm and are organized as a single instruction multiple data (SIMD) computing architecture. Both the encoder and the decoder divide a picture in regions of horizontal strips and use one local processor per region. These local processors code (decode) one horizontal strip of data which, using the terminology of the H.261 standard, corresponds to two group of blocks (GOBs). They also communicate to a central processor which multiplexes (demultiplexes) the coded data from (for) the processors in the encoder (decoder). In the case of the encoder the central processor also controls a data buffer for bit-rate adaptation. Lateral communication between adjacent processors is implemented to allow comparisons between blocks situated in neighbouring regions, as required by most motion estimation algorithms.
Luís Sá, Vítor Silva 0001, Fernando Perdigão, Sérgio M. M. de Faria, Pedro A. Amado Assunção
VCIP4
1990 A parallel architecture for real-time video coding
Luís Vieira de de Sá, Vítor Silva 0001, Fernando Perdigão, Sérgio M. M. de Faria, Pedro A. Amado Assunção
Microprocessing and Microprogramming4