Ricardo L. de Queiroz

dblp:66/1740 · DBLP profile ↗
← Back
136ranked-venue papers
41as first author
15since 2021 · last 2026
0000-0002-3911-1838ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 126 · 37 first-author · 14 since 2021Systems, architecture and hardware · 7 · 3 first-author · 1 since 2021Security and privacy · 2Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2026 A Rate-Distortion-Complexity Analysis of Neural Video CODECs
Ricardo L. de Queiroz, Diogo C. Garcia, Yi-Hsin Chen, Ruhan Conceição, Wen-Hsiao Peng, Luciano V. Agostini
ISCAS1
2025 Zerotree Coding of Subdivision Wavelet Coefficients in Dynamic Time-Varying Meshes
abstract
We propose a complete system to enable progressive coding with quality scalability of the mesh geometry, in MPEG's state-of-the-art Video-based Dynamic Mesh Coding (V-DMC) framework. In particular, we propose an alternative method for encoding the subdivision wavelet coefficients in V-DMC, using a zerotree coding approach that works directly in the native 3D mesh space. This allows us to identify parent-child relationships amongst the wavelet coefficients across different subdivision levels, which can be used to achieve an efficient and versatile coding mechanism. We demonstrate that, given a starting base mesh, a target subdivision surface and a desired maximum number of zerotree passes, our system produces an elegant and visually attractive lossy-to-lossless mesh geometry reconstruction with no further user intervention. Moreover, lossless coefficient encoding with our approach requires nearly the same bitrate as the default displacement coding methods in V-DMC. Yet, our approach provides several quality resolution levels embedded in the same bitstream, while the current V-DMC solutions encode a single quality level only. To the best of our knowledge, this is the first time that a zerotree-based method has been proposed and demonstrated to work for the compression of dynamic time-varying meshes, and the first time that an embedded quality-scalable approach has been used in the V-DMC framework.
Maja Krivokuca, Tomas M. Borges, Ricardo L. de Queiroz
IEEE Trans. Image Process.3
2024 Rate-Complexity Optimization in Lossless Neural-Based Image Compression
abstract
Neural networks are now widely used in image compression. Network architecture and hyperparameter choices impact both compression performance and complexity, but (as we show) there are many examples where higher complexity does not entail better compression. Thus, it is desirable to perform rate-complexity optimization over the space of hyperparameters. In the context of neural-based lossless image compression, we propose an algorithm that traces hyperparameter choices of points on or near the lower convex hull of the cloud of rate-complexity points produced by all combinations of hyperparameters, without having to know in advance the rate-complexity performance of each combination. This reduces the training/evaluation load of the rate-complexity optimization by over 50% in our experiments, for each of three measures of complexity: multiply/add operations per pixel, Joules per pixel, and encoded network size.
Lucas S. Lopes, Ricardo L. de Queiroz, Philip A. Chou
ICIP2
2024 Embedded Bit-Stream Region-of-Interest Coding of Point Cloud Attributes
abstract
This paper presents a novel approach to embedded Region-of-Interest (ROI) attribute coding for point clouds. The proposed method is an embedded point cloud attribute coder incorporating regions of interest. The method chosen was to interleave bit-streams for ROI and non-ROI regions, which brings us some advantages. Firstly, it is independent of any specific ROI detection algorithm. Secondly, no increased complexity is imposed on the decoder. The results demonstrate an improvement in the quality of reconstructed voxels within the ROI, accompanied by some degradation of reconstructed voxels outside the ROI.
Victor F. Figueiredo, Ricardo L. de Queiroz
MMSP2
2024 A Comparative Assessment of Implicit and Explicit Plenoptic Scene Representations
abstract
3D scene representation has been a central theme of study for a wide range of applications, and the representation of light behavior is one of the relevant topics when producing realistic models. In this work, we create a framework to assess the representation of non-Lambertian scenes by generating a pipeline to create plenoptic point clouds (PPCs) systematically and evaluating them against implicit solutions, such as Neural Radiance Fields (NeRF)-like models. We compare such approaches according to rendering quality and compression efficiency. On the compression side, we propose an encoding scheme for PPC, leveraging the occlusion masks of the points and the Moving Picture Expert Group's (MPEG) Geometry-Based Solid Content Test Model (GeS-TM). Rendering results over the training views show that the uncompressed PPC outperforms 3D Gaussian Splatting (3DGS) by 1.51 dB, on average, for the 8 scenes of the NeRF Synthetic 360 dataset. In compression efficiency, 3DGS outperforms the compressed PPCs by 0.7 dB in BD-PSNR on average. Our occlusion-aware encoding scheme reduces the size of uncompressed PPCs up to 800 times, outperforming current encoding schemes for PPC by 1.9 dB in BD-PSNR.
Davi Rabbouni Freitas, Ricardo L. de Queiroz, Ioan Tabus, Christine Guillemot
MMSP2
2024 Embedded Coding of Point Cloud Attributes
abstract
Point cloud compression (PCC) has been rapidly evolving in the context of international standards. Despite the inherent scalability of octree-based geometry descriptions, current attribute compression techniques prevent full scalability of compressed point clouds. We propose an improvement on an embedded attribute encoding method for point clouds based on set partitioning in hierarchical trees (SPIHT). We propose to use a multi-layer perceptron (MLP) to model contexts in order to further compress the final bit-stream. The encoder is used along with the region-adaptive hierarchical transform, which has been a popular transform for point cloud coding and is included in the standard geometry-based point cloud coder (G-PCC). The result is an encoder that is efficient, scalable, and, best of all, embedded. That is, higher compression is achieved by further trimming the single bit-stream. Experimental results show approximately 13% BD-Rate reduction using MLP-based context modeling.
Victor F. Figueiredo, Ricardo L. de Queiroz, Philip A. Chou, Lucas S. Lopes
IEEE Signal Process. Lett.2
2023 Motion-Compensated Predictive RAHT for Dynamic Point Clouds
abstract
We study the use of predictive approaches alongside the region-adaptive hierarchical transform (RAHT) in attribute compression of dynamic point clouds. The use of intra-frame prediction with RAHT was shown to improve attribute compression performance over pure RAHT and represents the state-of-the-art in attribute compression of point clouds, being part of MPEG's geometry-based test model. We studied a combination of inter-frame and intra-frame prediction for RAHT for the compression of dynamic point clouds. An adaptive zero-motion-vector (ZMV) scheme and an adaptive motion-compensated scheme are developed. The simple adaptive ZMV approach is able to achieve sizable gains over pure RAHT and over the intra-frame predictive RAHT (I-RAHT) for point clouds with little or no motion while ensuring similar compression performance to I-RAHT for point clouds with intense motion. The motion-compensated approach, more complex and more powerful, is able to achieve large gains across all of the tested dynamic point clouds.
André L. Souto, Ricardo L. de Queiroz, Camilo C. Dorea
IEEE Trans. Image Process.2
2022 On Quantization of Image Classification Neural Networks for Compression Without Retraining
abstract
We studied the quantization of neural networks for their compression and representation without retraining. The goal is to facilitate neural network representation and deployment in standard formats so that general networks may have their weights quantized and entropy coded within the deployment format. We relate weight entropy and model accuracy and try to evaluate distribution of weights against known distributions. Many scalar quantization strategies were tested. We have found that weights are typically approximated by a Laplacian distribution for which optimal quantizers are approximated by entropy-coded uniform quantizers with dead-zones. Results indicate that it is possible to reduce 8-fold the size of the popular image classification networks with accuracy losses near 1%.
Marcos Tonin, Ricardo L. de Queiroz
ICIP2
2022 Geometry-Based Compression of Plenoptic Point Clouds
abstract
Plenoptic point clouds (PPC) are novel data structures that represent the light from different viewing directions in order to provide a higher degree of realism to regular point clouds. This is achieved by associating each point to multiple colors instead of a single one. Here, we present a method to efficiently compress the attributes of a PPC, consisting of a Karhunen-Loève transform over the color attributes followed by multiple attribute coders with intra prediction capability. This compression scheme can be incorporated within the MPEG's geometry-based PCC (G-PCC) standard, using any of G-PCC's existing solutions for attribute coding. Compression performance assessment using PPCs of different spatial resolutions reveals competitive results in comparison to existing methods, such as RAHT-based or video-based PCC solutions. We believe our coder to be the new state of the art.
Davi Rabbouni Freitas, Gustavo L. Sandri, Ricardo L. de Queiroz
MMSP3
2022 Adaptive Context Modeling for Arithmetic Coding Using Perceptrons
abstract
Arithmetic coding is used in most media compression methods. Context modeling is usually done through frequency counting and look-up tables (LUTs). For long-memory signals, probability modeling with large context sizes is often infeasible. Recently, neural networks have been used to model probabilities of large contexts in order to drive arithmetic coders. These neural networks have been trained offline. We introduce anonlinemethod for training a perceptron-based context-adaptive arithmetic coder on-the-fly, calledadaptive perceptron coding, which continuously learns the context probabilities and quickly converges to the signal statistics. We test adaptive perceptron coding over a binary image database, with results always exceeding the performance of LUT-based methods for large context sizes and of recurrent neural networks. We also compare the method to a version requiring offline training, which leads to equally satisfactory results.
Lucas S. Lopes, Philip A. Chou, Ricardo L. de Queiroz
IEEE Signal Process. Lett.3
2022 Fractional Super-Resolution of Voxelized Point Clouds
abstract
We present a method to super-resolve voxelized point clouds downsampled by a fractional factor, using lookup-tables (LUT) constructed from self-similarities from their own downsampled neighborhoods. The proposed method was developed to densify and to increase the precision of voxelized point clouds, and can be used, for example, as improve compression and rendering. We super-resolve the geometry, but we also interpolate texture by averaging colors from adjacent neighbors, for completeness. Our technique, as we understand, is the first specifically developed for intra-frame super-resolution of voxelized point clouds, for arbitrary resampling scale factors. We present extensive test results over different point clouds, showing the effectiveness of the proposed approach against baseline methods.
Tomas M. Borges, Diogo C. Garcia, Ricardo L. de Queiroz
IEEE Trans. Image Process.3
2022 Differential Transform for Video-Based Plenoptic Point Cloud Coding
abstract
Point cloud compression has been studied in standard bodies and we are here concerned with the Moving Picture Experts Group video-based point cloud compression (V-PCC) solution. Plenoptic point clouds (PPC) is a novel volumetric data representation wherein points are associated with colors in all viewing directions to improve realism. It is sampled as a number ($N_{c}$) of attribute colors per point. We propose a new method for the efficient video-based compression of PPC that is backwards compatible with the existing single-color V-PCC decoder. V-PCC generates three image atlases which are encoded using an image/video encoder. We assume there may be a reference color which is to be encoded as the main payload. We generate$N_{c}+3$atlases and we produce$N_{c}$differential images against the reference color image. Those difference images are pixel-wise transformed using an$N_{c}$-point discrete cosine transform, generating$N_{c}$transformed atlases which are encoded, forming the secondary payload. Such secondary information is the plenoptic enhancement to the point cloud. If there is no reference attribute, we skip the differences and use the lowest frequency of the transformed atlases as the main payload. Results are presented that show an unrivaled performance of the proposed method.
Diogo C. Garcia, Camilo C. Dorea, Renan U. Ferreira, Davi Rabbouni Freitas, Ricardo L. de Queiroz, Rogério Higa, Ismael Seidel, Vanessa Testoni
IEEE Trans. Image Process.5
2021 Multi-Resolution Intra-Predictive Coding Of 3d Point Cloud Attributes
abstract
We propose an intra frame predictive strategy for compression of 3D point cloud attributes. Our approach is integrated with the region adaptive graph Fourier transform (RAGFT), a multi-resolution transform formed by a composition of localized block transforms, which produces a set of low pass (approximation) and high pass (detail) coefficients at multiple resolutions. Since the transform operations are spatially localized, RAGFT coefficients at a given resolution may still be correlated. To exploit this phenomenon, we propose an intraprediction strategy, in which decoded approximation coefficients are used to predict uncoded detail coefficients. The prediction residuals are then quantized and entropy coded. For the 8i dataset, we obtain gains up to 0. 5db as compared to intra predicted point cloud compresion based on the region adaptive Haar transform (RAHT).
Eduardo Pavez, André L. Souto, Ricardo L. de Queiroz, Antonio Ortega
ICIP3
2021 Memory-Friendly Segmentation Refinement for Video-Based Point Cloud Compression
abstract
Recently finalized, the MPEG Video-based Point Cloud Compression (V-PCC) standard leverages existing video codecs to compress point clouds. This approach relies on existing video coding hardware accelerators to enable its fast adoption. However, the processing steps to project 3D point clouds into 2D frames still have a considerable complexity. Thus, we propose two modifications to TMC2, V-PCC’s test model, considering the memory access pattern: one reducing unnecessary memory allocation and the other increasing data locality through a set of pre-processing steps. Our approach achieved up to 46.31% encoding self-time reduction without changing the resulting bitstream. This paper also provides an analysis of our implementation that may help future V-PCC codecs achieve real-time encoding.
Ismael Seidel, Davi Rabbouni Freitas, Camilo C. Dorea, Diogo C. Garcia, Renan U. Ferreira, Rogério Higa, Ricardo L. de Queiroz, Vanessa Testoni
ICIP7
2021 Set Partitioning in Hierarchical Trees for Point Cloud Attribute Compression
abstract
We propose an embedded attribute encoding method for point clouds based on set partitioning in hierarchical trees (SPIHT). The encoder is used with the region-adaptive hierarchical transform which has been a popular transform for point cloud coding, even included in the standard geometry-based point cloud coder (G-PCC). The result is an encoder that is efficient, scalable, and embedded. That is, higher compression is achieved by trimming the full bit-stream. G-PCCs RAHT coefficient prediction prevents the straightforward incorporation of SPIHT into G-PCC. However, our results over other RAHT-based coders are promising, improving over the original, nonpredictive RAHT encoder, while providing the key functionality of being embedded.
André L. Souto, Victor F. Figueiredo, Philip A. Chou, Ricardo L. de Queiroz
IEEE Signal Process. Lett.4
2020 Adaptive Block Partitioning of Point Clouds for Video-Based Color Compression
abstract
Point cloud compression presents novel and challenging demands. Video-based solutions can harness existing video codec technology for efficient compression, however, the cloud's 3D geometry and attributes must be presented within a compatible, regular 2D grid. We propose a method for generating equally-sized, square-shaped video blocks containing point cloud color information. Our video blocks may be readily tiled into an image for efficient color compression through a traditional video codec. The focus is on the development of a computationally simple, voxel-to-image projection methodology. Results provide evidence of competitive performance, with average PSNR gains of 0.68dB and attribute (color) bitrate savings of -15.10%, with respect to MPEG's G-PCC (TMC13 v5.1).
Camilo C. Dorea, Ricardo L. de Queiroz
ICIP2
2020 Lossy Point Cloud Geometry Compression Via Dyadic Decomposition
abstract
This paper proposes a lossy intra-frame coder of the geometry information of voxelized point clouds. Using an alternative approach to the widespread octree representation, this method represents the point cloud as an array of binary images. This algorithm works recursively using a dyadic decomposition that splits an interval of slices in two smaller intervals, depicting a binary tree traversal, and transmitting the occupancy information of each interval. The sequence of bi-level images are encoded in a lossless fashion until a fixed point in the tree, from where the algorithm “skips” the dyadic slicing and transmits all the k remaining slices as leaves of the tree, which are then encoded in a lossy fashion. The performance assessment shows that the proposed method outperforms state-of-the-art intra coders of lossy geometry for medium to higher bitrates on the public point cloud datasets tested.
Davi Rabbouni Freitas, Eduardo Peixoto, Ricardo L. de Queiroz, José Edil G. de Medeiros
ICIP3
2020 On Predictive RAHT For Dynamic Point Cloud Coding
abstract
We studied predictive coding applied to the region-adaptive hierarchical transform (RAHT) which is used for point cloud compression (PCC). RAHT is part of MPEG’s geometry-based PCC test model and an intra-frame prediction scheme for RAHT (URAHT), wherein the prediction residual is encoded rather than the voxel attributes themselves, has been shown to deliver large gains. We extend the scheme to inter-frame prediction and show that a combination of simple zero-motion-vector (ZMV) inter-frame and intra-frame predictions can provide sizeable gains over pure RAHT or over intra-frame-only prediction when compressing dynamic point clouds. An adaptive method is used such that sections where ZMV does not yield good prediction switch to intra-frame prediction, assuring the performance to be at least that of the intra-frame case. Be the gains large (in steady parts) or very small (where there is rapid motion) results show consistent positive gains coming from a simple inter- and intra-frame prediction combination.
André L. Souto, Ricardo L. de Queiroz
ICIP2
2020 Saliency Maps for Point Clouds
abstract
Algorithms for creating saliency maps are well established for images, even though there is no literature on such methods for point clouds. We use orthographic projections in 2D planes which are subject to well established saliency detection algorithms to create a 3D saliency map. The results of each saliency map are projected to the 3D voxels and the results of the many projections are used to generate a 3D saliency map. Simple compression tests were carried using soft region-of-interest maps. Results have shown an increase in the quality of the voxels inside the selective regions of increased levels of interest.
Victor F. Figueiredo, Gustavo L. Sandri, Ricardo L. de Queiroz, Philip A. Chou
MMSP3
2020 Geometry Coding for Dynamic Voxelized Point Clouds Using Octrees and Multiple Contexts
abstract
We present a method to compress geometry information of point clouds that explores redundancies across consecutive frames of a sequence. It uses octrees and works by progressively increasing resolution of the octree. At each branch of the tree, we generate an approximation of the child nodes by a number of methods which are used as contexts to drive an arithmetic coder. The best approximation, i.e. the context that yields the least amount of encoding bits, is selected and the chosen method is indicated as side information for replication at the decoder. The core of our method is a context-based arithmetic coder in which a reference octree is used as reference to encode the current octree, thus providing 255 contexts for each output octet. The 255×255 frequency histogram is viewed as a discrete 3D surface and is conveyed to the decoder using another octree. We present two methods to generate the predictions (contexts) which use adjacent frames in the sequence (inter-frame) and one method that works purely intra-frame. The encoder continuously switches the best mode among the three and conveys such information to the decoder. Since an intra-frame prediction is present, our coder can also work in purely intra-frame mode, as well. Extensive results are presented to show the method's potential against many compression alternatives for the geometry information in dynamic voxelized point clouds.
Diogo C. Garcia, Tiago A. da Fonseca, Renan U. Ferreira, Ricardo L. de Queiroz
IEEE Trans. Image Process.4
2019 Local Texture and Geometry Descriptors for Fast Block-Based Motion Estimation of Dynamic Voxelized Point Clouds
abstract
Motion estimation in dynamic point cloud analysis or compression is a computationally intensive procedure generally involving a large search space and often complex voxel matching functions. We present an extension and improvement on prior work to speed up block-based motion estimation between temporally adjacent point clouds. We introduce local, or block-based, texture descriptors as a complement to voxel geometry description. Descriptors are organized in an occupancy map which may be efficiently computed and stored. By consulting the map, a point cloud motion estimator may significantly reduce its search space while maintaining prediction distortion at similar quality levels. The proposed texture-based occupancy maps provide significant speedup, an average of 26.9% for the tested data set, with respect to prior work.
Camilo C. Dorea, Edson M. Hung, Ricardo L. de Queiroz
ICIP3
2019 Point Cloud Compression Incorporating Region of Interest Coding
abstract
We introduce Region-of-Interest (ROI) coding for point cloud attributes, using an input-weighted distortion measure where the weights are determined by the ROI. In terms of coding, we use the Region Adaptive Hierarchical Transform (RAHT), which relies on a set of weights. We use a measure-theoretic interpretation of RAHT to determine that the weights of the transform should be set to the weights of the distortion measure. The ROI is chosen as the 3D region of the face, which is detected from a set of 2D projections using the well-known Viola-Jones algorithm. Experimental results show subjectively meaningful improvements (7-8 dB PSNR) in a face ROI with subjectively insignificant degradations (under 1 dB PSNR) in the non-ROI.
Gustavo L. Sandri, Victor F. Figueiredo, Philip A. Chou, Ricardo L. de Queiroz
ICIP4
2019 Integer Alternative for the Region-Adaptive Hierarchical Transform
abstract
A recently-introduced coder based on region-adaptive hierarchical transform (RAHT) is being considered as a standard for the compression of point cloud attributes at moving picture experts group. The RAHT coefficients can be encoded in many ways and the transform is based on a series of orthogonal 2 × 2 transform matrices with geometry-dependent floating-point entries. In order to remove computation ambiguity and facilitate deployment, fixed-point operations are often preferred. In this letter, we present an alternative RAHT description that allows for fixed-point implementation of its transform steps. It is based on matrix decompositions akin to lifting steps and a scaling of the quantization steps. Results are presented to show that the new fixed-point transform is, in practical terms, equivalent to the floating-point RAHT. For that we use a reasonable number of precision bits for the integer operations, e.g. 8 b or more.
Gustavo L. Sandri, Philip A. Chou, Maja Krivokuca, Ricardo L. de Queiroz
IEEE Signal Process. Lett.4
2019 Compression of Plenoptic Point Clouds
abstract
Point clouds have been recently used in applications involving real-time capture and rendering of 3D objects. In a point cloud, for practical reasons, each point or voxel is usually associated with one single color along with other attributes. The region-adaptive hierarchical transform (RAHT) coder has been proposed for single-color point clouds. The cloud is usually captured by many cameras and the colors are averaged in some fashion to yield the point color. This approach may not be very realistic since, in real world objects, the reflected light may significantly change with the viewing angle, especially if specular surfaces are present. For that, we are interested in a more complete representation, the plenoptic point cloud, wherein every point has associated colors in all directions. Here, we propose a compression method for such a representation. Instead of encoding a continuous function, since there is only a finite number of cameras, it makes sense to compress as many colors per voxel as cameras, and to leave any intermediary color rendering interpolation to the decoder. Hence, each voxel is associated with a vector of color values, for each color component. We have here developed and evaluated four methods to expand the RAHT coder to encompass the multiple colors case. Experiments with synthetic data helped us to correlate specularity with the compression, since object specularity, at a given point in space, directly affects color disparity among the cameras, impacting the coder performance. Simulations were carried out using natural (captured) data and results are presented as rate-distortion curves that show that a combination of Kahunen-Loève transform and RAHT achieves the best performance.
Gustavo L. Sandri, Ricardo L. de Queiroz, Philip A. Chou
IEEE Trans. Image Process.2
2018 Block-Based Motion Estimation Speedup for Dynamic Voxelized Point Clouds
abstract
Motion estimation is a key component in dynamic point cloud analysis and compression. We present a method for reducing motion estimation computation when processing block-based partitions of temporally adjacent point clouds. We propose the use of an occupancy map containing information regarding size or other higher-order local statistics of the partitions. By consulting the map, the estimator may significantly reduce its search space, avoiding expensive block-matching evaluations. To form the maps we use 3D moment descriptors efficiently computed with one-pass update formulas and stored as scalar-values for multiple, subsequent references. Results show that a speedup of 2 produces a maximum distortion dropoff of less than 2% for the adopted PSNR-based metrics, relative to distortion of predictions attained from full search. Speedups of 5 and 10 are achievable with small average distortion dropoffs, less than 3% and 5%, respectively, for the tested data set.
Camilo C. Dorea, Ricardo L. de Queiroz
ICIP2
2018 Example-Based Super-Resolution for Point-Cloud Video
abstract
We propose a mixed-resolution point-cloud representation and an example-based super-resolution framework, from which several processing tools can be derived, such as compression, denoising and error concealment. By inferring the high-frequency content of low-resolution frames based on the similarities between adjacent full-resolution frames, the proposed framework achieves an average 1.18 dB gain over lowpass versions of the point-cloud, for a projection-based distortion metric [1], [2].
Diogo C. Garcia, Tiago A. da Fonseca, Ricardo L. de Queiroz
ICIP3
2018 Intra-Frame Context-Based Octree Coding for Point-Cloud Geometry
abstract
3D and free-viewpoint video has been slowly adopting a solid representation, such as using meshes and point-clouds. Among other characteristics, meshes provide direct surface representation, while point-cloud processing requires less computation. Points in the cloud are minimally represented by their geometry (3D position) and color. A common point-cloud geometry compression method is the octree representation, which acts on individual frames and can be further compressed by entropy encoding. This paper presents a lossless intra-frame compression method for point-cloud geometry, which uses the octree structure to provide better contexts for entropy coding. Results show that the proposed solution offers state-of-the-art performance, with an average rate reduction of 29% compared to the octree representation.
Diogo C. Garcia, Ricardo L. de Queiroz
ICIP2
2018 Compression of Plenoptic Point Clouds Using the Region-Adaptive Hierarchical Transform
abstract
Point clouds have recently gained interest for the represention of 3D scenes in augmented and virtual reality. In real-time applications point clouds typically assume one color per point. While this approach is suited to represent diffuse objects, it is less realistic with specular surfaces. We consider the compression of plenoptic point clouds, wherein each voxel is associated to colors as seen by different angles. We propose an efficiently compressible representation to incorporate the plenoptic information of each voxel. We have proposed three compression methods, one based on a cylindrical projection and two others based on the intersection of the line of view with the voxel's face, one using flat boundaries and the other using a spherical boundary. Extensive tests have shown that the last two have the best performance, which are much superior than independently encoding the color attributes from each of the cameras point of views.
Gustavo L. Sandri, Ricardo L. de Queiroz, Philip A. Chou
ICIP2
2018 Foreword to the Special Section on SIBGRAPI 2018
Arun Ross, Eduardo S. L. Gastal, Joaquim Jorge 0001, Ricardo L. de Queiroz
Comput. Graph.4
2018 Distance-Based Probability Model for Octree Coding
abstract
We present a context-driven method to encode nodes of an octree, which is typically used to encode the point-cloud geometry. Instead of using one bit per node of the tree, the context allows for deriving probabilities for that node based on distances of the actual voxel to the voxels in a reference point cloud. Accurate probabilities of the node state allow for the use of an arithmetic coder to reduce the bit rate. Results point to potentially large reductions in rate if there is a good model from which to derive the context, i.e., one can get large reductions if the reference-cloud geometry is close enough to the one being encoded.
Ricardo L. de Queiroz, Diogo C. Garcia, Philip A. Chou, Dinei A. F. Florêncio
IEEE Signal Process. Lett.1
2017 Context-based octree coding for point-cloud video
abstract
3D and free-viewpoint video has been moving towards solid-scene representation using meshes and point clouds. Point-cloud processing requires much less computation and the points in the cloud are minimally represented by their geometry (3D position) and color. A common point-cloud geometry compression method is the octree representation, which acts on individual frames. We present a lossless inter-frame compression method for point-cloud geometry, by reordering each octree based on previous frames prior to entropy coding. Results show that compared to the octree representation, the proposed solution offers an average rate reduction of 30%, while entropy-coding of the octree yields an average rate reduction of 24%.
Diogo C. Garcia, Ricardo L. de Queiroz
ICIP2
2017 Motion-compensated compression of point cloud video
abstract
3D immersive communications are trending as real-time point clouds capture and display of point cloud video become feasible. This paper presents a novel motion-compensated approach to encoding dynamic voxelized point clouds (VPC) at low bit rates. A simple coder breaks the VPC into blocks which are intra-frame coded or replaced by a motion-compensated version of a block in the previous frame. The decision is optimized in a rate-distortion sense, encoding with distortion both geometry and the color, at reduced bit-rates. In-loop filtering is employed to minimize compression artifacts caused by distortion in the geometry information. Simulations reveal that this simple motion compensated coder can efficiently extend the compression range of dynamic voxelized point clouds to rates below what intra-frame coding alone can accommodate, trading rate for geometry accuracy.
Ricardo L. de Queiroz, Philip A. Chou
ICIP1
2017 Motion estimation with multiple matching criteria
abstract
Block-based motion estimation is the method of choice in most video codecs to exploit temporal redundancy for compression. Since true rate-distortion evaluation for every candidate block is usually impractical, simple estimates are used instead as a matching criterion, e.g., the Sum of Absolute Differences (SAD) between the target and the candidate blocks weighted by its respective motion vector cost in bits. We show that different matching criteria may differ not only in the quality of the resulting motion estimation but may actually offer diverse motion estimates that can be locally selected to better suit the local data characteristics. As a proof of concept, we propose the Double Matching Criteria Algorithm (DMCA), a two-pass algorithm which performs two independent motion estimations for each block with different matching criteria and then locally selects the better one in a true rate-distortion sense. The algorithm is easily parallelizable and out-of-the-box compliant to any coding standard with support for block-based motion compensation. We also propose the Total Absolute Deviation from the Mean (TADM) as a matching criterion to be used along with the SAD in the DMCA framework. Unlike the SAD, which measures the size of the residue in a sense, the TADM is a measure of its dispersion. For demonstration purposes, we implemented the DMCA with the SAD and the TADM as matching criteria in a modified HM reference software encoder for the HEVC standard. We observed significant BD-rate gains with a fully compliant HEVC stream.
Gabriel L. de Oliveira, Eduardo Peixoto, Ricardo L. de Queiroz
MMSP3
2017 Convexity characterization of virtual view reconstruction error in multi-view imaging
abstract
Virtual view synthesis is a key component of multi-view imaging systems that enable visual immersion environments for emerging applications, e.g., virtual reality and 360-degree video. Using a small collection of captured reference view-points, this technique reconstructs any view of a remote scene of interest navigated by a user, to enhance the perceived immersion experience. We carry out a convexity characterization analysis of the virtual view reconstruction error that is caused by compression of the captured multi-view content. This error is expressed as a function of the virtual viewpoint coordinate relative to the captured reference viewpoints. We derive fundamental insights about the nature of this dependency and formulate a prediction framework that is able to accurately predict the specific dependency shape, convex or concave, for given reference views, multi-view content and compression settings. We are able to integrate our analysis into a proof-of-concept coding framework and demonstrate considerable benefits over a baseline approach.
Vladan Velisavljevic, Camilo C. Dorea, Jacob Chakareski, Ricardo L. de Queiroz
MMSP4
2017 A Quality-of-Content-Based Joint Source and Channel Coding for Human Detections in a Mobile Surveillance Cloud
abstract
More than 70% of consumer mobile Internet traffic will be mobile video transmissions by 2019. The development of wireless video transmission technologies has been boosted by the rapidly increasing demand of video streaming applications. Although more and more videos are delivered for video analysis (e.g., object detection/tracking and action recognition), most existing wireless video transmission schemes are developed to optimize human perception quality and are suboptimal for video analysis. In mobile surveillance networks, a cloud server collects videos from multiple moving cameras and detects suspicious persons in all camera views. Camera mobility in smartphones or dash cameras implies that video is to be uploaded through bandwidth-limited and error-prone wireless networks, which may cause quality degradation of the decoded videos and jeopardize the performance of video analyses. In this paper, we propose an effective rate-allocation scheme for multiple moving cameras in order to improve human detection (content) performance. Therefore, the optimization criterion of the proposed rate-allocation scheme is driven by quality of content (QoC). Both video source coding and application layer forward error correction coding rates are jointly optimized. Moreover, the proposed rate-allocation problem is formulated as a convex optimization problem and can be efficiently solved by standard solvers. Many simulations using High Efficiency Video Coding standard compression of video sequences and the deformable part model object detector are carried, and results demonstrate the effectiveness and favorable performance of our proposed QoC-driven scheme under different pedestrian densities and wireless conditions.
Xiang Chen 0003, Jenq-Neng Hwang, De Meng, Kuan-Hui Lee, Ricardo L. de Queiroz, Fu-Ming Yeh
IEEE Trans. Circuits Syst. Video Technol.5
2017 Attention-Weighted Texture and Depth Bit-Allocation in General-Geometry Free-Viewpoint Television
abstract
In a free-viewpoint television network, each viewer chooses its point of view from which to watch a scene. We use the concept of total observed distortion, wherein we aim to minimize the distortion of the view observed by the viewers as opposed to the distortion of each camera, to develop an optimized bit-rate allocation for each camera. Our attention-weighted approach effectively gives more bits to the cameras that are more watched. The more concentrated the viewer distribution, the larger the bit-rate savings, for a given total observed distortion, compared with the uniform rate allocation. We analyze and model the distortion of a synthesized view as a function of the distortions (both in texture and/or depth) of the nearby cameras. Based on such models, we develop optimal rate-allocation methods for texture images, considering a uniform bit allocation for depth, and for both texture and depth simultaneously. Simulation results are shown, demonstrating not only the correctness of the optimized solution, but also measuring its improvement against uniform rate allocation for a few viewer distributions.
Camilo C. Dorea, Ricardo L. de Queiroz
IEEE Trans. Circuits Syst. Video Technol.2
2017 Transform Coding for Point Clouds Using a Gaussian Process Model
abstract
We propose using stationary Gaussian processes (GPs) to model the statistics of the signal on points in a point cloud, which can be considered samples of a GP at the positions of the points. Furthermore, we propose using Gaussian process transforms (GPTs), which are Karhunen-Loève transforms of the GP, as the basis of transform coding of the signal. Focusing on colored 3D point clouds, we propose a transform coder that breaks the point cloud into blocks, transforms the blocks using GPTs, and entropy codes the quantized coefficients. The GPT for each block is derived from both the covariance function of the GP and the locations of the points in the block, which are separately encoded. The covariance function of the GP is parameterized, and its parameters are sent as side information. The quantized coefficients are sorted by the eigenvalues of the GPTs, binned, and encoded using an arithmetic coder with bin-dependent Laplacian models, whose parameters are also sent as side information. Results indicate that transform coding of 3D point cloud colors using the proposed GPT and entropy coding achieves superior compression performance on most of our data sets.
Ricardo L. de Queiroz, Philip A. Chou
IEEE Trans. Image Process.1
2017 Motion-Compensated Compression of Dynamic Voxelized Point Clouds
abstract
Dynamic point clouds are a potential new frontier in visual communication systems. A few articles have addressed the compression of point clouds, but very few references exist on exploring temporal redundancies. This paper presents a novel motion-compensated approach to encoding dynamic voxelized point clouds at low bit rates. A simple coder breaks the voxelized point cloud at each frame into blocks of voxels. Each block is either encoded in intra-frame mode or is replaced by a motion-compensated version of a block in the previous frame. The decision is optimized in a rate-distortion sense. In this way, both the geometry and the color are encoded with distortion, allowing for reduced bit-rates. In-loop filtering is employed to minimize compression artifacts caused by distortion in the geometry information. Simulations reveal that this simple motion-compensated coder can efficiently extend the compression range of dynamic voxelized point clouds to rates below what intra-frame coding alone can accommodate, trading rate for geometry accuracy.
Ricardo L. de Queiroz, Philip A. Chou
IEEE Trans. Image Process.1
2016 Gaussian process transforms
abstract
We introduce the Gaussian Process Transform (GPT), an orthogonal transform for signals defined on a finite but otherwise arbitrary set of points in a Euclidean domain. The GPT is obtained as the Karhunen-Loéve Transform (KLT) of the marginalization of a Gaussian Process defined on the domain. Compared to the Graph Transform (GT), which is the KLT of a Gauss Markov Random Field over the same set of points whose neighborhood structure is inherited from the Euclidean domain, the GPT has up to 6 dB higher coding gain.
Philip A. Chou, Ricardo L. de Queiroz
ICIP2
2016 Attention-weighted depth map rate-allocation in free-viewpoint television
abstract
We propose optimal rate-allocation, using viewer attention information among viewpoints, for depth map cameras within a free-viewpoint television broadcast system. An attention-weighted rate-allocation framework enables bit-rate, or quality, to be distributed across the multiple cameras in accordance with viewer interest, minimizing total observed distortions perceived among all viewers. Prior work has considered attention-weighted rate-allocation for texture cameras only in such systems. Depth maps, nonetheless, are a common requirement for view rendering in many systems. They may constitute a significant part of overall transmission bit-rate and they present their own distortion characteristics, different from those of texture. We model the effects of depth camera distortions upon virtual or synthesized views and propose an optimal attention-weighted rate-allocation among depth cameras. Results show significant gains in average PSNR of synthesized views and bit-rate savings of our proposal relative to the balanced rate-allocation alternative.
Camilo C. Dorea, Ricardo L. de Queiroz
ICIP2
2016 Compression of 3D Point Clouds Using a Region-Adaptive Hierarchical Transform
abstract
In free-viewpoint video, there is a recent trend to represent scene objects as solids rather than using multiple depth maps. Point clouds have been used in computer graphics for a long time, and with the recent possibility of real-time capturing and rendering, point clouds have been favored over meshes in order to save computation. Each point in the cloud is associated with its 3D position and its color. We devise a method to compress the colors in point clouds, which is based on a hierarchical transform and arithmetic coding. The transform is a hierarchical sub-band transform that resembles an adaptive variation of a Haar wavelet. The arithmetic encoding of the coefficients assumes Laplace distributions, one per sub-band. The Laplace parameter for each distribution is transmitted to the decoder using a custom method. The geometry of the point cloud is encoded using the well-established octtree scanning. Results show that the proposed solution performs comparably with the current state-of-the-art, while being much more computationally efficient. We believe this paper represents the state of the art in intra-frame compression of point clouds for real-time 3D video.
Ricardo L. de Queiroz, Philip A. Chou
IEEE Trans. Image Process.1
2015 Evaluating the effects of image compression in Moiré-pattern-based face-spoofing detection
abstract
Face-recognition biometric systems have been shown unreliable under the presence of face-spoofing images, creating the need for automatic spoofing detection. In this paper, the effect of image compression degradation on face-spoofing detection is evaluated, based on an algorithm that searches for Moiré patterns due to the overlap of the digital grids, through peak detection in the frequency domain. The proposed spoofing-detection algorithm performance is evaluated subject to H.264/AVC and JPEG image compression for an image database of facial shots under several conditions.
Diogo C. Garcia, Ricardo L. de Queiroz
ICIP2
2015 Clustering of matched features and gradient matching for mixed-resolution video super-resolution
abstract
This work presents a novel technique for image reconstruction applied to mixed-resolution video super-resolution. We segment an image into patches defined by the clustering of a vector flow generated from matching SIFT features. We reconstruct the segmented image by applying image projective transformation to a reference image. By varying the number of clusters, we composed a sequence of reconstructed images, which are then used to compose a codebook, through gradient matching. This idea is extended to use low and high-resolution image pairs for super-resolution. Our results indicate a 1.4dB gain, on average, over the use of overlapped-block motion-compensation (OBMC).
Renan U. Ferreira, Edson M. Hung, Ricardo L. de Queiroz
ISCAS3
2015 Quality-of-content (QoC)-driven rate allocation for video analysis in mobile surveillance networks
abstract
Nowadays, more and more videos are transmitted for video analytics purposes rather than human perceptions. In mobile surveillance networks, a cloud server collects videos delivered from multiple moving cameras and detects suspicious people in all the camera views. However, all the videos recorded by moving cameras such as phone or dash cameras are uploaded through bandwidth-limited wireless networks. Therefore, videos are required to be encoded with high compression ratio to satisfy the total data rate constraint, which may affect the video analyses (e.g., human detection/tracking and action recognition, etc.) performance due to the degraded video decoding qualities at the server side. In this paper, we propose an effective content-driven video source coding rate allocation scheme, which can improve the human detection success rate in mobile surveillance networks under a total data rate constraint. The proposed scheme allocates appropriate amount of data rate to each moving camera based on the corresponding content information (i.e., human detection results). A model of human detection accuracy based on object area and video quality is provided. The rate allocation problem is formulated as a convex optimization problem and can be solved by standard solvers. Simulations with real video sequences demonstrate the effectiveness of our proposed scheme.
Xiang Chen 0003, Jenq-Neng Hwang, Kuan-Hui Lee, Ricardo L. de Queiroz
MMSP4
2015 Face-Spoofing 2D-Detection Based on Moiré-Pattern Analysis
abstract
Biometric systems based on face recognition have been shown unreliable under the presence of face-spoofing images. Hence, automatic solutions for spoofing detection became necessary. In this paper, face-spoofing detection is proposed by searching for Moiré patterns due to the overlap of the digital grids. The conditions under which these patterns arise are first described, and their detection is proposed which is based on peak detection in the frequency domain. Experimental results for the algorithm are presented for an image database of facial shots under several conditions.
Diogo C. Garcia, Ricardo L. de Queiroz
IEEE Trans. Inf. Forensics Secur.2
2014 General rate-allocation in free-viewpoint television
abstract
We propose a framework for optimal rate-allocation in free-viewpoint television (FVTV) for a general camera arrangement based on the attention the viewers are paying to each camera. In a recent letter [1], the authors proposed a FVTV broadcast architecture and an optimal bit-allocation approach, assuming a uniformly-spaced one-dimensional arrangement of cameras. Quality (or bit-rate) at each camera was determined by viewer attention in order to minimize total observed distortion. Here, we extend the optimized bit-allocation scheme to allow for a more generic camera arrangement in FVTV. We present results on data sets from 1D and 2D array camera setups which show significant overall PSNR gains and bit-rate savings with respect to equally-balanced rate-allocation across cameras.
Camilo C. Dorea, Ricardo L. de Queiroz
ICIP2
2014 A fast HEVC transcoder based on content modeling and early termination
abstract
In this paper, a fast transcoding solution from H.264/AVC to HEVC bitstreams is presented. This solution is based on two main modules: a coding unit (CU) classification module that relies on a machine learning technique in order to map H.264/AVC macroblocks into HEVC CUs; and an early termination technique that is based on statistical modeling of the HEVC rate-distortion (RD) cost in order to further speed-up the transcoding. The transcoder is built around an established two-stage transcoding. In the first stage, called the training stage, full re-encoding is performed while the H.264/AVC and the HEVC information are gathered. This information is then used to build both the CU classification model and the early termination sieves, that are used in the second stage (called the transcoding stage). The solution is tested with well-known video sequences and evaluated in terms of RD and complexity. The proposed method is 3.83 times faster, on average, than the trivial transcoder, and 1.8 times faster than a previous transcoding solution, while yielding a RD loss of 4% compared to this solution.
Eduardo Peixoto, Bruno Macchiavello, Edson M. Hung, Ricardo L. de Queiroz
ICIP4
2014 Example-based Enhancement of Degraded Video
abstract
We present an example-based approach to general enhancement of degraded video frames. The method relies on building a dictionary with non-degraded parts of the video and to use such a dictionary to enhance the degraded parts. The image degradation has to originate from a “repeatable” process, so that the dictionary image patches (blocks) are equally degraded, thus originating a dictionary with degraded blocks and their residues (differences in between degraded and original blocks). Once a match is found between a degraded block in the video and a degraded block in the dictionary, the associated residue of the latter is soft-added to the block of the former. The method is a generalization of the method for example-based super-resolution. Results are presented to demonstrate the applicability of the method to many scenarios.
Edson M. Hung, Diogo C. Garcia, Ricardo L. de Queiroz
IEEE Signal Process. Lett.3
2013 Energy-constrained real-time H.264/AVC video coding
abstract
Energy consumption has become a leading design constraint for computing devices in order to defray electric bills for individuals and businesses. Over the past years, digital video communication technologies have demanded higher computing power availability and, therefore, higher energy expenditure. In order to meet the challenge to provide software-based video encoding solutions, we adopted an open source software implementation of an H.264 video encoder, the x264 encoder, and optimized its prediction stage in the energy sense (E). Thus, besides looking for the coding options which lead to the best coded representation in terms of rate and distortion (RD), we constrain the process to fit within a certain energy budget. i.e., an RDE optimization. We considered energy as the time integration of the real demanded electric power for a given system. We present an RDE-optimized framework which allows for software-based real-time video compression, meeting the desired targets of electrical consumption, hence, controlling carbon emissions. We show results of energy-constrained compression wherein one can save as much as 35% of the energy with small impact on RD performance.
Tiago A. da Fonseca, Ricardo L. de Queiroz
ICASSP2
2013 Depth-map super-resolution for asymmetric stereo images
abstract
We propose a mixed-resolution coding architecture for stereo color-plus-depth images, where encoding is performed at a low resolution, except for one of the color images. Super-resolution methods are proposed for the depth maps and for the low-resolution color image at the decoder side. Experiments are carried out for several real and synthetic images and reveal a reduction in complexity at the encoder associated with average PSNR gains of the super-resolved low-resolution color image with respect to regular interpolation.
Diogo C. Garcia, Camilo C. Dorea, Ricardo L. de Queiroz
ICASSP3
2013 Attention-Weighted Rate Allocation in Free-Viewpoint Television
abstract
An architecture for free-viewpoint broadcast television transmission is proposed where all the views are transmitted at potentially different qualities and watched by a large number of viewers. The quality (or bit-rate) of each view is controlled by the distribution of viewpoints chosen by the viewers. For example, if most viewers are watching synthetic views in between viewsnandn+1, those views are allocated more transmission bits than views that are scarcely watched. We developed an attention-weighted bit-rate-allocation method that is optimal in the total observer distortion sense. The optimality of the method relies on knowing the viewpoint probability distribution at every moment. Simulation results show that overall transmission rate can be reduced for the same total observed distortion.
Thacio Scandarolli, Ricardo L. de Queiroz, Dinei A. F. Florêncio
IEEE Signal Process. Lett.2
2013 Scanned Document Compression Using Block-Based Hybrid Video Codec
abstract
This paper proposes a hybrid pattern matching/transform-based compression method for scanned documents. The idea is to use regular video interframe prediction as a pattern matching algorithm that can be applied to document coding. We show that this interpretation may generate residual data that can be efficiently compressed by a transform-based encoder. The efficiency of this approach is demonstrated using H.264/advanced video coding (AVC) as a high-quality single and multipage document compressor. The proposed method, called advanced document coding (ADC), uses segments of the originally independent scanned pages of a document to create a video sequence, which is then encoded through regular H.264/AVC. The encoding performance is unrivaled. Results show that ADC outperforms AVC-I (H.264/AVC operating in pure intramode) and JPEG2000 by up to 2.7 and 6.2 dB, respectively. Superior subjective quality is also achieved.
Alexandre Zaghetto, Ricardo L. de Queiroz
IEEE Trans. Image Process.2
2012 Video super-resolution based on local invariant features matching
abstract
This paper presents an algorithm for video super-resolution based on scale-invariant feature transform (SIFT) matching. SIFT features are known to be a robust method for locating keypoints. The matching of these keypoints from different frames in a video allows us to infer high-frequency information in order to perform example-based super-resolution. We first apply a block constrained keypoint detection for a more precise superposition of features. Later, we extract high-frequency information with a gradient-based matching scheme. Our results indicate gains over interpolation and previous example-based super-resolution approaches.
Renan U. Ferreira, Edson M. Hung, Ricardo L. de Queiroz
ICIP3
2012 HEVC-based scanned document compression
abstract
This paper proposes a hybrid pattern matching/transform-based compression engine for scanned compound documents. The novelty of this approach is demonstrated by using a modified version of the HEVC (High Efficiency Video Coding) Test Model as a compound document compressor, here conveniently referred to as HEDC (High Efficiency Document Coder). The proposed method uses segments of a document to create a video sequence, which is then encoded by HEDC. The idea is to explore interframe prediction as a pattern matching algorithm for coding units pre-classified as text; and intraframe prediction for coding units pre-classified as image. Results show that HEDC outperforms AVC-I, HEVC-I (H.264/AVC and HEVC operating in pure intra mode), H.264/AVC and JPEG2000 by up to 3.3, 2.5, 1.7 and 5 dB, respectively. Furthermore, for most documents the proposed method yields practically the same rate-distortion performance as regular HEVC, but is approximately 5% to 20% faster due to a pre-classification algorithm that prevents it of performing all possible inter/intra prediction tests for each prediction unit.
Alexandre Zaghetto, Bruno Macchiavello, Ricardo L. de Queiroz
ICIP3
2012 Super Resolution for Multiview Images Using Depth Information
abstract
In stereoscopic and multiview video, binocular suppression theory states that the visual subjective quality of 3-D experience is not much affected by asymmetrical blurring of the individual views. Based on these studies, mixed-resolution frameworks applied for multiview systems offer great data-size reduction without incurring in significant quality degradation in 3-D video applications. However, it is interesting to recover high-frequency content of the blurred views, to reduce visual strain due to long-term exposure and to make the system suitable for free-viewpoint television. In this paper, we present a novel super-resolution technique, in which low-resolution views are enhanced with the aid of high-frequency content from neighboring full-resolution views, and the corresponding depth information for all views. Occlusions are handled by checking the consistency between views. Tests for synthetic and real image data in stereo and multiview cases are presented, and results show that significant objective quality gains can be achieved without any extra side information.
Diogo C. Garcia, Camilo C. Dorea, Ricardo L. de Queiroz
IEEE Trans. Circuits Syst. Video Technol.3
2012 Video Super-Resolution Using Codebooks Derived From Key-Frames
abstract
Example-based super-resolution (SR) is an attractive option to Bayesian approaches to enhance image resolution. We use a multiresolution approach to example-based SR and discuss codebook construction for video sequences. We match a block to be super-resolved to a low-resolution version of the reference high-resolution image blocks. Once the match is found, we carefully apply the high-frequency contents of the chosen reference block to the one to be super-resolved. In essence, the method relies on “betting” that if the low-frequency contents of two blocks are very similar, their high-frequency contents also might match. In particular, we are interested in scenarios where examples can be picked up from readily available high-resolution images that are strongly related to the frame to be super-resolved. Hence, they constitute an excellent source of material to construct a dynamic codebook. Here, we propose a method to super-resolve a video using multiple overlapped variable-block-size codebooks. We implemented a mixed-resolution video coding scenario, where some frames are encoded at a higher resolution and can be used to enhance the other lower-resolution ones. In another scenario, we consider the framework where the camera captures a video at a lower resolution and also takes periodic snapshots at a higher resolution. Results indicate substantial gains over interpolation and fixed-codebook SR, and significant gains over previous works as well.
Edson M. Hung, Ricardo L. de Queiroz, Fernanda Brandi, Karen França de Oliveira, Debargha Mukherjee
IEEE Trans. Circuits Syst. Video Technol.2
2011 Depth map reconstruction using color-based region merging
abstract
This paper presents a novel depth map enhancement method which takes as inputs a single view and an initial depth estimate. A region-based framework is introduced wherein a color-based partition of the image is created and depth uncertainty areas are identified according to the alignment of detected depth discontinuities and region borders. A color-based homogeneity criterion is used to guide a constrained region merging process and reconstruct depth estimates within the uncertainty areas. Experimental results on publicly available test sequences illustrate the potential of the algorithm in significantly improving low quality depth estimates.
Camilo C. Dorea, Ricardo L. de Queiroz
ICIP2
2011 Video compression complexity reduction with adaptive down-sampling
abstract
Encoding video sequences is a computation-demanding task in high-performance codecs. Optimizing this stage may result in a substantial encoding speed-up. In this paper, we propose a faster approach to encode sequences with the H.264/AVC codec, using mixed-resolution. By having some of the frames down-sampled, the overall computation is reduced, without greatly affecting the rate-distortion performance, which remains evaluated using the sequence's original resolution. In addition to this approach, a method which regards only the most frequent prediction modes of a frame is incorporated, so that the combined complexity reduction can be significant with negligible performance losses.
Diogo C. Garcia, Tiago A. da Fonseca, Ricardo L. de Queiroz
ICIP3
2011 Transform domain semi-super resolution
abstract
This paper presents a transform-based approach to semi-super resolution. The idea of semi-super resolution is to enhance low-resolution frames from a video sequence encoded with different resolutions among frames. The proposed framework uses a DCT-based down-sampling method at the encoder process. At the decoder, DCT-based up-sampling plus high-frequency information from adjacent frames in full resolution are used to enhance interpolated low-resolution frames. The results show an increase of the overall quality without using any extra information from the encoder. The proposed technique can also outperfom previous semi-super resolution methods operating in pixel domain.
Edson M. Hung, Diogo C. Garcia, Ricardo L. de Queiroz
ICIP3
2011 Complexity-constrained rate-distortion optimization for h.264/avc video coding
abstract
In order to enable real-time software-based video encoding, in this work we optimized the prediction stage of an H.264 video encoder, in the complexity sense. Thus, besides looking for the coding options which lead to the best coded representation in terms of rate and distortion (RD), we constrain to a complexity (C) budget. We present a complexity optimized framework (RDC-optimized) which allows for real-time video compression and that does not make use of frame-skipping to comply to the desired encoding speed. We developed our framework around an open source software implementation of the H.264/AVC, the the ×264 encoder. Results show that tight complexity control is attainable in practice, with very little loss in RD performance.
Tiago A. da Fonseca, Ricardo L. de Queiroz
ISCAS2
2010 Complexity-scalable H.264/AVC in an IPP-based video encoder
abstract
Real-time high-definition video encoding is a computation-hungry task that challenges software-based solutions. For that, in this work we adopted an Intel software implementation of an H.264 video encoder and optimized its prediction stage in the complexity sense (C). Thus, besides looking for the coding options which lead to the best coded representation in terms of rate and distortion, we constrain the process to fit within a certain time budget. We present an RDC-optimized framework which allows for real-time HD video compression.
Tiago A. da Fonseca, Ricardo L. de Queiroz, Debargha Mukherjee
ICIP2
2010 Super-resolution for multiview images using depth information
abstract
The joint usage of low- and full-resolution images in multiview systems provides an attractive opportunity for data size reduction while maintaining good quality in 3D applications. In this paper we present a novel application of a super-resolution method for usage within a mixed resolution multiview setup. The technique borrows high-frequency content from neighboring full resolution images to enhance particular low-resolution views. Occlusions are handled through matching of low-resolution images. Both the stereo and the more general multiview cases are considered using the multiview video-plus-depth format. Results demonstrate significant gains in PSNR and in visual quality for test sequences.
Diogo C. Garcia, Camilo C. Dorea, Ricardo L. de Queiroz
ICIP3
2010 Mixed-resolution distributed video codec without motion estimation at the encoder
abstract
Inspired by recent results showing that Wyner-Ziv coding using a combination of source and channel coding may be more efficient than pure channel coding, we have applied coset codes for the source coding part in the transform domain for Wyner-Ziv coding of video. The framework is a mixed-resolution approach where reduced encoding complexity is achieved by low resolution encoding of non-reference frames and regular encoding of the reference frames. Different from our previous works, no motion estimation is carried at the encoder, neither for the low resolution frames nor the reference frames. The entropy coders of the H.264/AVC codec were tuned to improve their performance for encoding the cosets. An encoding mode with lowest encoding complexity than H.264/AVC intra mode is achieved. Experimental results show a competitive rate-distortion performance especially at low bit rates.
Bruno Macchiavello, Edson M. Hung, Ricardo L. de Queiroz, Debargha Mukherjee
ICIP3
2010 High quality scanned book compression using pattern matching
abstract
This paper proposes a hybrid approximate pattern matching/transform-based compression engine. The idea is to use regular video interframe prediction as a pattern matching algorithm that can be applied to document coding. We show that this interpretation may generate residual data that can be efficiently compressed by a transform-based encoder. The novelty of this approach is demonstrated by using H.264/AVC, the newest video compression standard, as a high quality book compressor. The proposed method uses segments of the originally independent scanned pages of a book to create a video sequence, which is encoded through regular H.264/AVC. Results show that the proposed method outperforms AVC-I (H.264/AVC operating in pure intra mode) and JPEG2000 by up to 4 dB and 7 dB, respectively. Superior subjective quality is also achieved.
Alexandre Zaghetto, Ricardo L. de Queiroz
ICIP2
2010 Reversible color-to-gray mapping using subband domain texturization
Ricardo L. de Queiroz
Pattern Recognit. Lett.1
2010 Least-Squares Directional Intra Prediction in H.264/AVC
abstract
A new intra-prediction mode for the H.264/AVC standard is proposed. Each pixel within a block is predicted by a weighted sum of its neighbours, according to anNth order Markov linear model. The weights are obtained through a least-squares estimate from reconstructed data in the neighbouring blocks, so that no overhead is necessary to convey the weights to the decoder. Results show significant improvements in H.264/AVC compression for images that are rich in directional structures, and moderate improvements for the other images and sequences tested.
Diogo C. Garcia, Ricardo L. de Queiroz
IEEE Signal Process. Lett.2
2010 A Wyner-Ziv Video Transcoder
abstract
Wyner-Ziv (WZ) coding of video utilizes simple encoders and highly complex decoders. A transcoder from a WZ codec to a traditional codec can potentially increase the range of applications for WZ codecs. We present a transcoder scheme from the most popular WZ codec architecture to a differential pulse code modulation/discrete cosine transform codec. As a proof of concept, we implemented this transcoder using a simple pixel-domain WZ codec and the standard H.263+. The transcoder design aims at reducing complexity as a large amount of computation is saved by reusing the motion estimation, calculated at the side information generation process, and the l-frame streams. New approaches are used to generate side information and to map motion vectors for the transcoder. Results are presented to demonstrate the transcoder performance.
Eduardo Peixoto, Ricardo L. de Queiroz, Debargha Mukherjee
IEEE Trans. Circuits Syst. Video Technol.2
2009 Efficiency improvements for a geometric-partition-based video coder
abstract
H.264/AVC has brought an important increase in coding efficiency in comparison to previous video coding standards. One of its features is the use of macroblock partitioning in a tree-based structure. The use of macroblock partitions based in arbitrary line segments, like wedge partitions, has been reported to increase coding gains. The main problem of these non-standard new partitions schemes is the increase in computational complexity. Thus, our work proposes improvements in this extension to the H.264/AVC standard. First, we present a motion vector prediction scheme based on directional partitions. Second, we present a method for complexity reduction based on the most frequent partitions. The results show that it is possible to still produce good coding gains with lower complexity than previous approaches.
Renan U. Ferreira, Edson M. Hung, Ricardo L. de Queiroz, Debargha Mukherjee
ICIP3
2009 Macroblock sampling and mode ranking for complexity scalability in mobile H.264 video coding
abstract
We propose a framework for complexity scalability in H.264. The prediction is constrained so that only a subset of prediction modes are tested. The test subset is found by ranking the most ¿popular¿ modes (those the are most often picked as best) and selecting the modes that maximize their expected occurrence frequency given a complexity constraint. Ranking is performed by selecting a small set of macroblocks (the sampling population), for which all modes are tested. The remaining macroblocks only test the available dominant modes. The modes in the sampling population of a frame are used to process the next frame. Results are shown to verify the performance of the proposed method, which reveals sizeable complexity savings at small penalties.
Tiago A. da Fonseca, Ricardo L. de Queiroz
ICIP2
2009 Mapping motion vectors for Awyner-Ziv video transcoder
abstract
Wyner-Ziv (WZ) coding of video utilizes simple encoders and highly complex decoders. A transcoder from a WZ codec to a traditional codec can potentially increase the range of applications for WZ codecs. We present a transcoder scheme from the most popular WZ codec architecture to a DPCM/DCT codec. As a proof of concept, we implemented this transcoder using a simple pixel domain WZ codec and the standard H.263+. The transcoder design aims at reducing complexity, since the transcoder has to perform both WZ decoding and DPCM/DCT encoding, including motion estimation. New approaches are used to map motion vectors for such a transcoder. Results are presented to demonstrate the transcoder performance.
Eduardo Peixoto, Ricardo L. de Queiroz, Debargha Mukherjee
ICIP2
2009 Improved layer processing for MRC compression of scanned documents
abstract
The Mixed Raster Content (MRC) document compression is a well documented standard. Its efficiency for representing sharp text and graphics over a background has been extensively presented. Scanned documents, however, are difficult to be dealt with because of soft transitions. In one of our previous works we presented a pre/post-processing algorithm for MRC compression of scanned images that sharpens the edges, before compression, and softens it again, after reconstruction. The present paper uses the same basic idea but describes three new features that improve, in a rate-distortion sense, the compression performance. First, the real bitrate is used to determine the pre/post-processing parameters. Second, the mask layer and the edge sharpening/softening map have their quality improved and, third, layer resolution change has been explored. Experimental results show that the method not only provides higher subjective quality, but it can yield up to 2.5 dB gains in PSNR against single coder approaches, in the compression ranges of interest.
Alexandre Zaghetto, Ricardo L. de Queiroz
ICIP2
2009 Complexity-constrained H.264 HD video coding through mode ranking
abstract
Video compression is a computation-intensive task and real-time implementation of state-of-the-art video standards is a challenge. In this paper we propose a complexity scalable framework to H.264/AVC prediction module since the “intra” and “inter” prediction steps are responsible for almost all computation complexity. The idea is to employ a subset of prediction modes instead of testing all modes recommended by H.264/AVC standard. Only the most “popular” prediction modes are ranked and tested until reaching a complexity budget. The ranking is done by sampling macroblocks of previous frames. Results show a negligible the quality loss while achieving large computational savings.
Tiago A. da Fonseca, Ricardo L. de Queiroz
PCS2
2009 Iterative Side-Information Generation in a Mixed Resolution Wyner-Ziv Framework
abstract
We propose a mixed resolution framework based on full resolution key frames and spatial-reduction-based Wyner-Ziv coding of intermediate nonreference frames. Improved rate-distortion performance is achieved by enabling better side-information generation at the decoder side and better rate-allocation at the encoder side. The framework enables reduced encoding complexity by low resolution encoding of the nonreference frames, followed by Wyner-Ziv coding of the Laplacian residue. The quantized transform coefficients of the residual frame are mapped to cosets without the use of a feedback channel. A study to select optimal coding parameters in the creation of the memoryless cosets is made. Furthermore, a correlation estimation mechanism that guides the parameter choice process is proposed. The decoder first decodes the low resolution base layer and then generates a super-resolved side-information frame at full resolution using past and future key frames. Coset decoding is carried using side-information to obtain a higher quality version of the decoded frame. Implementation results are presented for the H.264/AVC codec.
Bruno Macchiavello, Debargha Mukherjee, Ricardo L. de Queiroz
IEEE Trans. Circuits Syst. Video Technol.3
2008 Super-resolution of video using key frames and motion estimation
abstract
Many scalable video coding systems use frame down- sampling in order to reduce complexity and to enable enhancement layers. Super-resolution (SR) can be used to help the up-sampling and recovering processes of those frames. We are interested in reversed-complexity (distributed) coding methods, wherein few key frames are encoded at normal resolution, while the rest are down- sampled and encoded at reduced resolution along with the enhancement layers. We are only interested on the decoder side, wherein we carry motion estimation of the down- sampled frames using the key frames as references. When a match is made, the high-frequency components of the key frames (KF) are used to super-resolve the non-key frames (NKF). The motion estimation process is performed using blocks of band-pass versions of the frames, rather than low- pass ones. Results indicate the improved performance of the proposed super-resolution algorithm.
Fernanda Brandi, Ricardo L. de Queiroz, Debargha Mukherjee
ICIP2
2008 Parameter estimation for an H.264-based distributed video coder
abstract
In this paper we present a statistical model used to select coding parameters for a mixed resolution Wyner-Ziv framework implemented using the H.264/AVC standard. This paper extends the results of a previous work for the H.263+ case to the H.264/AVC coder, since the parameters need to be recalculated for the H.264 case. The proposed correlation estimation mechanism guides the parameter choice process, and also yields the statistical model used for decoding. This mechanism is proposed based on extracting edge information and residual error rate in co-located blocks from the low resolution base layer that is available at both ends.
Bruno Macchiavello, Ricardo L. de Queiroz, Debargha Mukherjee
ICIP2
2008 Intra prediction versuswavelets and lapped transforms in an H.264/AVC coder
abstract
H.264/AVC is the latest video coding standard and, among other things, it uses a DCT-like transform and intra prediction modes. We are studying the possibility of replacing the modified DCT stages by lapped and wavelet transforms. Since those transforms have overlap, intra-frame prediction is not feasible, because of its block-recursive nature. Hence, intra-frame prediction is turned off. In essence, this paper contains a comparison among lapped (wavelet) transforms and linear prediction schemes, within the AVC framework. Results indicate that lapped transforms can outperform the intra prediction scheme, specially for high definition images.
Rafael Galvão de Oliveira, Ricardo L. de Queiroz
ICIP2
2008 Iterative pre- and post-processing for MRC layers of scanned documents
abstract
The mixed raster content (MRC) document compression standard (ITU T.44) specifies a multi-layer multi-resolution representation of a compound document. The model is very efficient for representing sharp text and graphics onto a background. However, since the mask layer is binary, it is difficult to deal with scanned data and soft edges. The edge transitions do not fully belong to the foreground neither to the background, and cause some "halo" to the object edges using the MRC model. This paper presents an algorithm that builds an edge sharpening map and iteratively parameterizes the original edge "softness" at the encoder. The generated map and the "softness" parameters are, then, used to reconstruct the original soft edges at the decoder. Experimental results are presented, showing that the method can yield 1.5 dB gains in PSNR, in the compression ranges of interest.
Alexandre Zaghetto, Ricardo L. de Queiroz
ICIP2
2008 Super resolution of video using key frames
abstract
In many video compression systems, the frames are down-sampled before transmission. Also, in many scalable systems, the residual after down- and up-sampling is encoded and transmitted. Sometimes, a few frames are encoded at normal resolution (key frames) while the other frames are encoded at reduced resolution. Super resolution can be used to enhance the up-sampling process, using motion information to improve traditional interpolation. In this paper, we propose to use a super resolution method to up-sample the non-key frames using the key frames as reference. We build dictionaries on-the- fly using the key frames instead of the traditional off-line training images. The high-frequency data of matching blocks are added to the low-resolution blocks. Since the key frames are very similar to the non-key frames, the method is robust enough to allow successful super resolution of highly compressed (severely degraded) sequences. Results are presented for many predefined block sizes, key frames frequencies, and compression parameters.
Fernanda Brandi, Ricardo L. de Queiroz, Debargha Mukherjee
ISCAS2
2007 Motion-Based Side-Information Generation for a Scalable Wyner-Ziv Video Coder
abstract
A motion-based side-information generation scheme with semi super-resolution for a scalable Wyner-Ziv coder framework is introduced. It is known that the performance of any Wyner-Ziv coder is heavily dependent on the efficiency of the side-information generation. We propose an iterative block based scheme to generate a semi super-resolution frame using the past and future reference frames which should be coded at full-resolution. To enable this side-information generation the framework should allow for low encoding complexity, reducing the spatial resolution only in the non-reference frames. The enhancement layer is produced using a residual frame of the reduced resolution encoded frame. The decoder first decodes the low resolution base layer and iteratively generates the side-information, along with channel decoding, to obtain a higher quality version of the decoded frame. Results of the implementation of the framework using the motion-based side-information in the H.263+ and H.264 standards are presented.
Bruno Macchiavello, Ricardo L. de Queiroz, Debargha Mukherjee
ICIP (6)2
2007 Using H.264/AVC-Intra for Segmentation-Driven Compound Document Coding
abstract
In this paper we explore H.264/AVC operating in intraframe mode to compress a mixed image, i.e. composed of text, graphics and pictures. Even though mixed contents (compound) documents usually require the use of multiple compressors, we apply a single compressor for both text and pictures. For that, distortion is taken into account differently between text and picture regions. Our approach is to use a segmentation-driven adaptation strategy to change the H.264/ AVC quantization parameter on a macroblock by macroblock basis, i.e. we deviate bits from pictorial regions to text in order to keep text edges sharp. We show results of a segmentation driven quantizer adaptation method applied to compress documents. Our reconstructed images have better text sharpness compared to straight unadapted coding, at negligible visual losses on pictorial regions. Our results also highlight the fact that H.264/AVC-INTRA outperforms coders such as JPEG-2000 as a single coder for compound images.
Alexandre Zaghetto, Ricardo L. de Queiroz
ICIP (2)2
2007 A simple reversed-complexity Wyner-Ziv video coding mode based on a spatial reduction framework
abstract
A spatial-resolution reduction based framework for incorporation of a Wyner-Ziv frame coding mode in existing video codecs is presented, to enable a mode of operation with low encoding complexity. The core Wyner-Ziv frame coder works on the Laplacian residual of a lower-resolution frame encoded by a regular codec at reduced resolution. The quantized transform coefficients of the residual frame are mapped to cosets to reduce the bit-rate. A detailed rate-distortion analysis and procedure for obtaining the optimal parameters based on a realistic statistical model for the transform coefficients and the side information is also presented. The decoder iteratively conducts motion-based side-information generation and coset decoding, to gradually refine the estimate of the frame. Preliminary results are presented for application to the H.263+ video codec.
Debargha Mukherjee, Bruno Macchiavello, Ricardo L. de Queiroz
VCIP3
2007 Segmentation-Driven Compound Document Coding Based on H.264/AVC-INTRA
abstract
In this paper, we explore H.264/AVC operating in intraframe mode to compress a mixed image, i.e., composed of text, graphics, and pictures. Even though mixed contents (compound) documents usually require the use of multiple compressors, we apply a single compressor for both text and pictures. For that, distortion is taken into account differently between text and picture regions. Our approach is to use a segmentation-driven adaptation strategy to change the H.264/AVC quantization parameter on a macroblock by macroblock basis, i.e., we deviate bits from pictorial regions to text in order to keep text edges sharp. We show results of a segmentation driven quantizer adaptation method applied to compress documents. Our reconstructed images have better text sharpness compared to straight unadapted coding, at negligible visual losses on pictorial regions. Our results also highlight the fact that H.264/AVC-INTRA outperforms coders such as JPEG-2000 as a single coder for compound images.
Alexandre Zaghetto, Ricardo L. de Queiroz
IEEE Trans. Image Process.2
2006 A Model for the Electronic Representation of Bank Checks
abstract
The substitution of physical bank check exchange by electronic check image transfer brings agility, security and cost reduction to the clearing system. In this paper, we propose a model for the electronic representation of bank checks based on the mixed raster content (MRC) model for compression of color and gray-scale images. The binary image is sent first and only if necessary the other MRC image planes are sent to reconstruct the color image of the check. Careful subjective evaluation helped us indicate the best binarization technique and studies led us to use JPEG 2000 to compress the MRC image planes. Furthermore, we propose a new watermarking technique to embed a digital signature into the check image for protection.
Danilo Dias, Ricardo L. de Queiroz
ICIP2
2006 On Macroblock Partition for Motion Compensation
abstract
In the H.264/AVC video coding standard, motion compensation can be performed by partitioning macroblocks into square or rectangular sub-macroblocks in a quadtree decomposition. This paper studies a motion compensation method using wedges, i.e. partitioning macroblocks or sub-macroblocks into two regions by an arbitrary line segment. This technique allows the shapes of the divided regions to better match the boundaries between moving objects. However, there are a large number of ways to slice a block and searching exhaustively over all of them would be an extremely computer-intensive task. Thus, we propose a fast algorithm which detects the predominant edge orientations within a block in order to pre-select candidate wedge lines. Finally a comparison among macroblock partition methods is performed, which points to the higher performance of the wedge partition method.
Edson M. Hung, Ricardo L. de Queiroz, Debargha Mukherjee
ICIP2
2006 Pre-Processing for MRC Layers of Scanned Images
abstract
The mixed raster content (MRC) imaging model represents compound images as a superposition of layers. The model is very efficient for representing sharp text and graphics onto a background, but, since the mask layer is binary, it is difficult to deal with scanned data and soft edges. The edge transitions don't fully belong to the foreground neither to the background, and cause some "halo" to the object edges using the MRC model. We present a method to detect segmented soft edges and a method to correct and sharpen the image within the MRC model. An improved data filling algorithm for the redundant regions is also presented. The original "softness" and relative position of the edges are estimated and we present a method to attempt to recreate the edge softness effect at the decoder.
Ricardo L. de Queiroz
ICIP1
2006 JPEG compression history estimation for color images
abstract
We routinely encounter digital color images that were previously compressed using the Joint Photographic Experts Group (JPEG) standard. En route to the image's current representation, the previous JPEG compression's various settings-termed its JPEG compression history (CH)-are often discarded after the JPEG decompression step. Given a JPEG-decompressed color image, this paper aims to estimate its lost JPEG CH. We observe that the previous JPEG compression's quantization step introduces a lattice structure in the discrete cosine transform (DCT) domain. This paper proposes two approaches that exploit this structure to solve the JPEG Compression History Estimation (CHEst) problem. First, we design a statistical dictionary-based CHEst algorithm that tests the various CHs in a dictionary and selects the maximum a posteriori estimate. Second, for cases where the DCT coefficients closely conform to a 3-D parallelepiped lattice, we design a blind lattice-based CHEst algorithm. The blind algorithm exploits the fact that the JPEG CH is encoded in the nearly orthogonal bases for the 3-D lattice and employs novel lattice algorithms and recent results on nearly orthogonal lattice bases to estimate the CH. Both algorithms provide robust JPEG CHEst performance in practice. Simulations demonstrate that JPEG CHEst can be useful in JPEG recompression; the estimated CH allows us to recompress a JPEG-decompressed image with minimal distortion (large signal-to-noise-ratio) and simultaneously achieve a small file-size.
Ramesh Neelamani, Ricardo L. de Queiroz, Zhigang Fan 0001, Sanjeeb Dash, Richard G. Baraniuk
IEEE Trans. Image Process.2
2006 Color to gray and back: color embedding into textured gray images
abstract
We have developed a reversible method to convert color graphics and pictures to gray images. The method is based on mapping colors to low-visibility high-frequency textures that are applied onto the gray image. After receiving a monochrome textured image, the decoder can identify the textures and recover the color information. More specifically, the image is textured by carrying a subband (wavelet) transform and replacing bandpass subbands by the chrominance signals. The low-pass subband is the same as that of the luminance signal. The decoder performs a wavelet transform on the received gray image and recovers the chrominance channels. The intent is to print color images with black and white printers and to be able to recover the color information afterwards. Registration problems are discussed and examples are presented.
Ricardo L. de Queiroz, Karen M. Braun
IEEE Trans. Image Process.1
2005 Watermarking JBIG2 text region for image authentication
abstract
An authentication watermarking technique (AWT) inserts hidden data into an image in order to detect any accidental or malicious alteration to the image. AWT usually computes, using a secret-key, the message authentication code (MAC) of the image, and somehow inserts MAC into the image itself. The authentication verification is performed using the same secret-key used in the insertion. JBIG2 is an international standard for lossy and lossless binary image compression. It decomposes the image into several regions (text, halftone, line-art) and encodes each one using the most appropriate method. This paper proposes a novel data-hiding technique to embed the data into the text region of a JBIG2 file. We, then, use it to create an AWT for a JBIG2-encoded image. The embedded data can be extracted from either the JBIG2 file itself or the binary image obtained by decoding the JBIG2 file.
Sergio Vicente Denser Pamboukian, Hae Yong Kim, Ricardo L. de Queiroz
ICIP (2)3
2005 Color embedding into gray images
abstract
We present a reversible method to convert color graphics and pictures to gray images based on mapping colors to low-visibility high-frequency textures. From a monochrome textured image, the decoder can identify the textures and recover the color information. An image is textured by carrying a wavelet transform and replacing band-pass sub-bands by the chrominance images. The low-pass sub-band is the same as that of the luminance signal. The decoder performs a wavelet transform on the received gray image and recovers the chrominance channels. Registration problems are discussed and examples are presented.
Ricardo L. de Queiroz, Karen M. Braun
ICIP (3)1
2005 Spatially varying gray component replacement for image watermarking
abstract
We present a method to include watermarks in printed images. We spatially vary the GCR along the image in a manner that is not perceptible, and we employ an estimation method to detect such changes. The choice of GCR for a given pixel (or region) comprises an additional information channel that embeds a watermark or hidden information. For that, we estimate the RGB values of each pixel and the CMYK values on the paper by scanning the printed page. With that information, we can estimate which GCR strategy was used in a given region and retrieve the watermark message. We are only concerned with robustly estimating the GCR strategy. Promising results illustrate the method's potential to watermark printed images.
Ricardo L. de Queiroz, Karen M. Braun, Robert P. Loce
ICIP (1)1
2004 A public-key authentication watermarking for binary images
abstract
Authentication watermarking is the process that inserts hidden data into an object (image) in order to detect any fraudulent alteration perpetrated by a malicious hacker. In the literature, quite a small number of secure authentication methods are available for binary images. This paper proposes a new secure authentication watermarking method for binary images. It can detect any visually significant alteration while maintaining good visual quality. As usual, the security of the algorithm lies on the secrecy of a private-key. Only its owner can insert the correct watermark while anyone may verify the authenticity through the corresponding public-key. A possible application of the proposed technique is in Internet fax transmission, i.e. for legal authentication of documents routed outside the phone network.
Hae Yong Kim, Ricardo L. de Queiroz
ICIP2
2004 Alteration-Locating Authentication Watermarking for Binary Images
Hae Yong Kim, Ricardo L. de Queiroz
IWDW2
2003 Inverse halftoning by decision tree learning
abstract
Inverse halftoning is the process to retrieve a (gray) continuous-tone image from a halftone. Recently, machine-learning-based inverse halftoning techniques have been proposed. Decision-tree learning has been applied with success to various machine-learning applications for quite some time. In this paper, we propose to use decision-tree learning to solve the inverse halftoning problem. This allows us to reuse a number of algorithms already developed. Especially, the maximization of entropy gain is a powerful idea that makes the learning algorithm to automatically select the ideal window as the decision-tree is constructed. The new technique has generated gray images with PSNR numbers, which are several dB above those previously reported in the literature. Moreover, it possesses very fast implementation, lending itself useful for real time applications.
Hae Yong Kim, Ricardo L. de Queiroz
ICIP (2)2
2003 JPEG compression history estimation for color images
abstract
We routinely encounter digital color images that were previously JPEG-compressed. We aim to retrieve the various settings - termed JPEG compression history (CH) - employed during previous JPEG operations. This information is often discarded en-route to the image's current representation. The discrete cosine transform coefficient histograms of previously JPEG-compressed images exhibit near-periodic behavior due to quantization. We propose a statistical approach to exploit this structure and thereby estimate the image's CH. Using simulations, we first demonstrate the accuracy of our estimation. Further, we show that JPEG recompression performed by exploiting the estimated CH strikes an excellent file-size versus distortion tradeoff.
Ramesh Neelamani, Ricardo L. de Queiroz, Zhigang Fan 0001, Richard G. Baraniuk
ICIP (3)2
2003 Identification of bitmap compression history: JPEG detection and quantizer estimation
abstract
Sometimes image processing units inherit images in raster bitmap format only, so that processing is to be carried without knowledge of past operations that may compromise image quality (e.g., compression). To carry further processing, it is useful to not only know whether the image has been previously JPEG compressed, but to learn what quantization table was used. This is the case, for example, if one wants to remove JPEG artifacts or for JPEG re-compression. In this paper, a fast and efficient method is provided to determine whether an image has been previously JPEG compressed. After detecting a compression signature, we estimate compression parameters. Specifically, we developed a method for the maximum likelihood estimation of JPEG quantization steps. The quantizer estimation method is very robust so that only sporadically an estimated quantizer step size is off, and when so, it is by one value.
Zhigang Fan 0001, Ricardo L. de Queiroz
IEEE Trans. Image Process.2
2002 A new 3-D subband video coding technique
abstract
We present in this paper a new 3-D subband coding method with low complexity, high performance, and features that facilitate channel error resilience. Four variations of the coder are discussed: one employing an octave band decomposition: one using the discrete cosine transforms (DCT) and two using lapped transforms. The base coder uses generalized adaptive quantization to code the subbands. Good performance at high compression ratios is obtained without traditional motion compensated prediction. Furthermore, these performance improvements can be realized without the use of entropy coding, which can make the system attractive for operation in error prone environments. Experimental results have shown that the subband video coder and its variations are able to achieve significant performance improvements over MPEG-2.
Hong Man, Ricardo L. de Queiroz, Mark J. T. Smith
ICASSP2
2002 Reduced DCT approximations for low bit rate coding
abstract
Transform approximations are explored for speeding up the software compression of images and video. Faster approximations are used to replace the regular DCT whenever only a few DCT coefficients are actually encoded. In one instance, we employ a data analyzer to drive a bank of coders by switching to the fastest coder depending on the image contents or on the output demand. The compression speed is adapted to the image contents in order to homogenize the data production rate, i.e. smoother areas produce fewer bits than detailed ones, but faster. The approximations are applicable to high compression (or constant output rate) environments and where lower complexity is required.
Ricardo L. de Queiroz
ICIP (1)1
2002 DCT approximation for low bit rate coding using a conditional transform
abstract
A transform approximation is explored for speeding up the software compression of images and video. It is used to replace the regular DCT whenever only a few DCT coefficients are actually encoded. A conditional transform is proposed to transform only the data that it determines to be relevant. The approximation is applicable to environments combining the requirements of high compression and low complexity.
Ricardo L. de Queiroz
ICIP (1)1
2002 Improved transforms for the compression of color and multispectral images
abstract
In the compression of color or multispectral imagery, intra pixel color transforms are usually employed to decorrelate planes. It is usually thought that plane decorrelation (such as provided by the Karhunen-Love transform (KLT)) may lead to higher compression. Spatial correlation, however, is usually not considered. We developed a method to devise a pixel-wise color transform that takes into account spatial correlation and outperforms the KLT. This is done by considering space-color correlation. We aim at decorrelating the data across color planes, but at correlating the data spatially, so that spatial transforms can more easily decorrelate each of the color planes. Experiment results are shown to demonstrate the gains of the propose transform.
Ricardo L. de Queiroz
ICIP (2)1
2002 On the completeness of the lattice factorization for linear-phase perfect reconstruction filter banks
abstract
In this letter, we re-examine the completeness of the lattice factorization for M-channel linear-phase perfect reconstruction filter bank (LPPRFB) with filters of the same length L=KM as discussed by Tran et al. (see IEEE Trans. Signal Processing, vol.48, p.133-47, Jan. 2000). We point out that the assertion of completeness is incorrect. Examples are presented to show that the proposed lattice structure of Tran et al. is not complete when K>2. In addition, we verify that the lattice structure is complete only when K/spl les/2.
Lu Gan 0002, Kai-Kuang Ma, Truong Q. Nguyen, Trac D. Tran, Ricardo L. de Queiroz
IEEE Signal Process. Lett.5
2002 Three-dimensional subband coding techniques for wireless video communications
abstract
We present a new 3D subband coding framework that achieves a good balance between high compression performance and channel error resilience. Various data transform methods for video decorrelation were examined and compared. The coding stage of the algorithm is based on a generalized adaptive quantization framework which is applied to the 3D transformed coefficients. It features a simple coding structure based on quadtree coding and lattice vector quantization. In typical applications, good performance at high compression ratios is obtained, often without entropy coding. Because temporal decorrelation is absorbed by the transform, traditional motion compensated prediction becomes unnecessary, resulting in a significant computational advantage over standard video coders. Error-resilience is achieved through classifying the compressed data streams into separated sub-streams with different error sensitivity levels. This enables a good adaptation to different channel models according to their noise statistics and error-protection protocols. Experimental results have shown that the subband video coder is able to achieve highly competitive performance relative to MPEG-2 in both noiseless and noisy environments. Lapped transforms are shown experimentally to outperform other transforms in the 3D subband environment. The subband coding framework provides a practical solution for video communications over wireless channels, where efficiency, error resilience and computational simplicity are vital in providing superior quality of service.
Hong Man, Ricardo L. de Queiroz, Mark J. T. Smith
IEEE Trans. Circuits Syst. Video Technol.2
2002 Variable complexity DCT approximations driven by an HVQ-based analyzer
abstract
Transform approximations are explored for speeding up the software compression of images and video. Approximations are used to replace the regular discrete cosine transform (DCT) whenever only a few DCT coefficients are actually encoded. We employ a novel hierarchical vector quantization (HVQ)-based image analyzer to drive a bank of coders, thus switching to the fastest coder depending on the image contents: smoother areas would produce less bits than detailed ones but faster. The HVQ analyzer is fast and simple enough to not offset the complexity gains brought by the DCT approximations. The approximations in this paper are applicable to high compression environments, where lower complexity is also required.
Ricardo L. de Queiroz
IEEE Trans. Circuits Syst. Video Technol.1
2001 On filter banks with rational oversampling
abstract
Despite the great popularity of critically-decimated filter banks, oversampled filter banks are useful in applications where data expansion is not a problem. We studied oversampled filter banks and showed that for some popular classes of filter banks it is not possible to obtain perfect reconstruction with rational (non-integer) oversampling ratios. Nevertheless, it is always possible to oversample the analysis filter bank by an integer factor, ie, there will be a similar synthesis bank which would provide perfect reconstruction. The analysis is carried within a time-aliasing framework development to analyze non-critically decimated filter banks.
Ricardo von Borries, Ricardo L. de Queiroz, C. Sidney Burrus
ICASSP2
2001 Signal processing using LUT filters based on hierarchical VQ
abstract
Vector quantization (VQ) is a powerful tool in signal processing. Hierarchical VQ (HVQ) is a method to implement VQ completely based on look-up tables (LUT). In HVQ, both encoders and decoders are inherently simple and fast, since there are no searches over codebooks. We introduce an overlapped HVQ (OHVQ) method, in which the number of samples is preserved after each HVQ stage. After the last stage, each OHVQ code in a particular location in the signal maps to a block (vector) which approximates that neighbourhood in the original sequence. For this reason, OBVQ is used as a basis to create a LUT-based filter, ie, a spatial signal processor with very fast implementation. Preliminary analysis and image processing examples are shown demonstrating the efficiency of the proposed method.
Ricardo L. de Queiroz, Patrick Fleckenstein
ICASSP1
2001 Compression color space estimation of JPEG images using lattice basis reduction
abstract
Given a color image that was quantized in some hidden color space (termed compression color space) during previous JPEG compression, we aim to estimate this unknown compression color space from the image. This knowledge is potentially useful for color image enhancement and JPEG re-compression. JPEG quantizes the discrete cosine transform (DCT) coefficients of each color plane independently during compression. Consequently, the DCT coefficients of such a color image conform to a lattice. We exploit this special geometry using the lattice reduction algorithm used in cryptography to estimate the compression color space. Simulations verify that the proposed algorithm yields accurate compression space estimates.
Ramesh Neelamani, Ricardo L. de Queiroz, Richard G. Baraniuk
ICIP (1)2
2001 Approximating lapped transforms through unitary postprocessing
abstract
We deal with the approximation of a nonsquare lapped transform matrix by another matrix. This second matrix is constrained to be the product of two stages: one lapped transform and a unitary postprocessing step, in such a way that approximation is accomplished by modifying the postprocessing stage alone. Viewing both matrices as transforms, the goal is to minimize the variance of the output error (discrepancy) for given input signal statistics. An extension of the result to frequency responses of filter bank transfer matrices is also given along with an example to demonstrate the feasibility of the method. The example demonstrates the potential use of the method in modifying transforms which can be applied in audio compression.
Ricardo L. de Queiroz
IEEE Signal Process. Lett.1
2000 Maximum Likelihood Estimation of JPEG Quantization Table in the Identification of Bitmap Compression History
abstract
To process previously JPEG coded images the knowledge of the quantization table used in compression is sometimes required. This happens for example in JPEG artifact removal and in JPEG re-compression. However, the quantization table might not be known due to various reasons. A method is presented for the maximum likelihood estimation (MLE) of the JPEG quantization tables. An efficient method is also provided to identify if an image has been previously JPEG compressed.
Zhigang Fan 0001, Ricardo L. de Queiroz
ICIP2
2000 On Data-Filling Algorithms for MRC Layers
abstract
A previously adopted standard, the mixed raster content (MRC) imaging model, represents compound images as a superposition of layers. Since layers are superimposed, large regions may not be imaged onto the final raster. Thus, those regions are redundant. We focus on techniques to replace the redundant data, i.e. on data filling redundant regions based on non-redundant ones. We present techniques to minimize the rate and distortion achieved by MRC compression through data filling. We start with a general method, then narrowing the presentation to DCT based compression of planes. Iterative block filling algorithms are presented in both spatial and DCT domain.
Ricardo L. de Queiroz
ICIP1
2000 Optimizing Block-Threshold Segmentation for MRC Compression
abstract
Compound document images contain graphic or textual content along with pictures. They are a very common form of documents, found in magazines, brochures, Web-sites, etc. We focus our attention on the mixed raster content (MRC) multi-layer approach for compound image compression. We study block thresholding as a means to segment an image for MRC. An attempt is made to optimize the block-threshold in a rate-distortion sense. Rate-distortion curves are presented to demonstrate the performance of the proposed algorithm.
Ricardo L. de Queiroz, Zhigang Fan 0001, Trac D. Tran
ICIP1
2000 Very fast JPEG compression using hierarchical vector quantization
abstract
We derive a JPEG compliant image compressor, which is based on hierarchical vector quantization (HVQ). The goal is to reduce complexity while increasing compression speed. For each block, the DCT DC coefficient is encoded in the regular way, while the residual is mapped through HVQ to a precomputed bit stream corresponding to the compressed DCT AC coefficients. Approximation quality is generally good for high compression ratios. Color Fax is one possible application target for the proposed system.
Ricardo L. de Queiroz, Patrick Fleckenstein
IEEE Signal Process. Lett.1
2000 Optimizing block-thresholding segmentation for multilayer compression of compound images
abstract
Compound document images contain graphic or textual content along with pictures. They are a very common form of documents, found in magazines, brochures, Web sites, etc. We focus our attention on the mixed raster content (MRC) multilayer approach for compound image compression. We study block thresholding as a means to segment an image for MRC. An attempt is made to optimize the block threshold in a rate-distortion sense. Also, a fast algorithm is presented to approximate the optimized method. Extensive results are presented including rate-distortion curves, segmentation masks and reconstructed images, showing the performance of the proposed algorithm.
Ricardo L. de Queiroz, Zhigang Fan 0001, Trac D. Tran
IEEE Trans. Image Process.1
2000 Classified JPEG coding of mixed document images for printing
abstract
This paper presents a modified JPEG coder that is applied to the compression of mixed documents (containing text, natural images, and graphics) for printing purposes. The modified JPEG coder proposed in this paper takes advantage of the distinct perceptually significant regions in these documents to achieve higher perceptual quality than the standard JPEG coder. The region-adaptivity is performed via classified thresholding being totally compliant with the baseline standard. A computationally efficient classification algorithm is presented, and the improved performance of the classified JPEG coder is verified.
Marcia G. Ramos, Ricardo L. de Queiroz
IEEE Trans. Image Process.2
1999 Adaptive rate-distortion-based thresholding: application in JPEG compression of mixed images for printing
abstract
We propose a new technique for transform coding based on rate-distortion (RD) optimized thresholding (i.e. discarding) of wasteful coefficients. The novelty in this proposed algorithm is that the distortion measure is made adaptive. We apply the method to the compression of mixed documents (containing text, natural images, and graphics) using JPEG for printing. Although the human visual system's response to compression artifacts varies depending on the region, JPEG applies the same coding algorithm throughout the mixed document. This paper takes advantage of perceptual classification to improve the performance of the standard JPEG implementation via adaptive thresholding, while being compatible with the baseline standard. A computationally efficient classification algorithm is presented, and the improved performance of the classified JPEG coder is verified. Tests demonstrate the method's efficiency compared to regular JPEG and to JPEG using non-adaptive thresholding. The non-stationary nature of distortion perception is true for most signal classes and the same concept can be used elsewhere.
Marcia G. Ramos, Ricardo L. de Queiroz
ICASSP2
1999 Compression of Compound Documents
abstract
Compound (or mixed) document images contain graphic or textual content along with pictures. They are a very common form of documents, found in magazines, brochures, web-sites etc. Because of the very distinct nature of those two image classes (text/graphics vs. pictures), their compression invariably involves multiple compression systems and a region segmentation (classification) method. We review state-of-the-art technologies on the subject while focusing our attention on the mixed raster content (MRC) multi-layer approach. We also present new results on segmentation for MRC based on optimized rate-distortion-based block thresholding.
Ricardo L. de Queiroz
ICIP (1)1
1999 On independent color space transformations for the compression of CMYK images
abstract
Device and image-independent color space transformations for the compression of CMYK images were studied. A new transformation (to a YYCC color space) was developed and compared to known ones. Several tests were conducted leading to interesting conclusions. Among them, color transformations are not always advantageous over independent compression of CMYK color planes. Another interesting conclusion is that chrominance subsampling is rarely advantageous in this context. Also, it is shown that transformation to YYCC consistently outperforms the transformation to YCbCrK, while being competitive with the image-dependent KLT-based approach.
Ricardo L. de Queiroz
IEEE Trans. Image Process.1
1998 The generalized lapped biorthogonal transform
abstract
A lattice structure based on the singular value decomposition (SVD) is introduced. The lattice can be proven to use a minimal number of delay elements and to completely span a large class of M-channel linear phase perfect reconstruction filter banks (LPPRFB): all analysis and synthesis filters have the same FIR length of L=KM, sharing the same center of symmetry. The lattice also structurally enforces both linear phase and perfect reconstruction properties, is capable of providing fast and efficient implementation, and avoids the costly matrix inversion problem in the optimization process. From a block transform perspective, the new lattice represents a family of generalized lapped biorthogonal transforms (GLBT) with arbitrary integer overlapping factor K. The relaxation of the orthogonal constraint allows the GLBT to have significantly different analysis and synthesis basis functions which can then be tailored appropriately to fit a particular application. Several design examples are presented along with a high-performance GLBT-based progressive image coder to demonstrate the superiority of the new lapped transforms.
Trac D. Tran, Ricardo L. de Queiroz, Truong Q. Nguyen
ICASSP2
1998 The Variable-Length Generalized Lapped Biorthogonal Transform
Trac D. Tran, Ricardo L. de Queiroz, Truong Q. Nguyen
ICIP (3)2
1998 On unitary transform approximations
abstract
In this letter, we present a method to find a unitary transform which is the closest to a given square transform matrix in the sense of minimizing the variance of the output error (discrepancy) for given signal statistics. An extension of the result to frequency responses of filterbank transfer matrices is also given along with an example to demonstrate the feasibility of the method.
Ricardo L. de Queiroz
IEEE Signal Process. Lett.1
1998 Processing JPEG-compressed images and documents
abstract
As the Joint Photographic Experts Group (JPEG) has become an international standard for image compression, we present techniques that allow the processing of an image in the "JPEG-compressed" domain. The goal is to reduce memory requirements while increasing speed by avoiding decompression and space domain operations. In each case, an effort is made to implement the minimum number of JPEG basic operations. Techniques are presented for scaling, previewing, rotating, mirroring, cropping, recompressing, and segmenting JPEG-compressed data. While most of the results apply to any image, we focus on scanned documents as our primary image source.
Ricardo L. de Queiroz
IEEE Trans. Image Process.1
1998 Nonexpansive pyramid for image coding using a nonlinear filterbank
abstract
A nonexpansive pyramidal decomposition is proposed for low-complexity image coding. The image is decomposed through a nonlinear filterbank into low- and highpass signals and the recursion of the filterbank over the lowpass signal generates a pyramid resembling that of the octave wavelet transform. The structure itself guarantees perfect reconstruction and we have chosen nonlinear filters for performance reasons. The transformed samples are grouped into square blocks and used to replace the discrete cosine transform (DCT) in the Joint Photographic Expert Group (JPEG) coder. The proposed coder has some advantages over the DCT-based JPEG: computation is greatly reduced, image edges are better encoded, blocking is eliminated, and it allows lossless coding.
Ricardo L. de Queiroz, Dinei A. F. Florêncio, Ronald W. Schafer
IEEE Trans. Image Process.1
1997 Downscaled inverses for M-channel lapped transforms
abstract
Compressed images may be decompressed for devices using different resolutions. Full decompression and rescaling in the space domain is a very expensive method. We studied downscaled inverses where the image is decompressed partially and a reduced inverse transform is used to recover the image. We studied the design of fast inverses, for a given forward transform. General solutions are presented for M-channel FIR filter banks of which block and lapped transforms are a subset.
Ricardo L. de Queiroz, Reiner Eschbach
ICASSP1
1997 Processing JPEG-Compressed Images
abstract
We present techniques that allow the processing of an image in the "JPEG-compressed" domain. The goal is to reduce memory requirements while increasing the speed by avoiding decompression and space domain operations. An effort is made to implement the minimum number of JPEG basic operations. Techniques are presented for scaling, previewing, rotating, mirroring, cropping, recompressing, and segmenting JPEG-compressed data.
Ricardo L. de Queiroz
ICIP (2)1
1997 Segmentation of Compressed Documents
abstract
We present a novel technique for segmentation of a JPEG-compressed document based on block activity. The activity is measured as the number of bits spent to encode each block. Each number is mapped to a pixel brightness value in an auxiliary image which is then used for segmentation. We introduce the use of such an image and show an example of a simple segmentation algorithm, which was successfully applied to test documents. The desired region can be identified and cropped (or replaced) from the compressed data without decompressing the image.
Ricardo L. de Queiroz, Reiner Eschbach
ICIP (3)1
1997 Wavelet transforms in a JPEG-like image coder
abstract
The discrete wavelet transform (DWT) is incorporated into the JPEG baseline system for image coding. The discrete cosine transform (DCT) is replaced by an association of two-channel filter banks connected hierarchically. The JPEG block-scanning and quantization schemes are adopted while we use JPEG's entropy coder. The changes in scanning can be incorporated into the transform block in such a way that the only part that needs to be changed in a JPEG framework is to replace the DCT by the DWT. Objective results and reconstructed images are presented demonstrating that the proposed coder outperforms JPEG and approaches the performance of more sophisticated and complex wavelet coders. However, it does not require full-image buffering nor imposes a large complexity increase.
Ricardo L. de Queiroz, C. K. Choi, Young Huh, Kamisetty Ramamohan Rao
IEEE Trans. Circuits Syst. Video Technol.1
1997 Fast downscaled inverses for images compressed with M-channel lapped transforms
abstract
Compressed images may be decompressed and displayed or printed using different devices at different resolutions. Full decompression and rescaling in space domain is a very expensive method. We studied downscaled inverses where the image is decompressed partially, and a reduced inverse transform is used to recover the image. In this fashion, fewer transform coefficients are used and the synthesis process is simplified. We studied the design of fast inverses, for a given forward transform. General solutions are presented for M-channel finite impulse response (FIR) filterbanks, of which block and lapped transforms are a subset. Designs of faster inverses are presented for popular block and lapped transforms.
Ricardo L. de Queiroz, Reiner Eschbach
IEEE Trans. Image Process.1
1996 A pyramidal coder using a nonlinear filter bank
abstract
We propose a novel pyramidal coder, which resembles the general structure of a JPEG coder, but which uses a nonlinear transform to replace the DCT. The nonlinear transform is obtained by the hierarchical application of a median filter predictor at subsampled versions of the original signal. The transformed samples are grouped into square blocks and used to replace the DCT in the JPEG baseline coder. The proposed coder shows several advantages: computation is greatly reduced compared to the DCT, image edges are better encoded, blocking is eliminated, and it allows lossless coding. Objective comparisons show the superiority of the proposed coder against both baseline and lossless JPEG.
Ricardo L. de Queiroz, Dinei A. F. Florêncio
ICASSP1
1995 Optimal orthogonal boundary filter banks
abstract
The use of a paraunitary filter bank for image processing requires a special treatment at image boundaries to ensure perfect reconstruction and orthogonality of these regions. Using time-varying boundary filter banks, we will discuss a procedure that explores all degrees of freedom of the border filters in a method essentially independent of signal extensions, allowing us to design optimal boundary filter banks, while maintaining fast implementation algorithms.
Ricardo L. de Queiroz, Kamisetty Ramamohan Rao
ICASSP1
1995 Variable block size lapped transforms
abstract
A structure for implementing lapped transforms with time-varying block sizes is discussed. It allows full orthogonality in the transitions based on a factorization of the transfer matrix into orthogonal factors. Such an approach can be viewed as a sequence of stages with variable-block-size transforms intermediated by sample-shuffling (delay) stages. Details for a first order system are given and a design example is presented.
Ricardo L. de Queiroz, Kamisetty Ramamohan Rao
ICIP1
1995 On Orthogonal Transforms of Images Using Paraunitary Filter Banks
Ricardo L. de Queiroz, Kamisetty Ramamohan Rao
J. Vis. Commun. Image Represent.1
1995 Extended lapped transform in image coding
abstract
A modulated lapped transform with extended overlap (ELT) is investigated in image coding with the objective of verifying its potential to replace the discrete cosine transform (DCT) in specific applications. Some of the criteria utilized for the performance comparison are reconstructed image quality (both objective and subjective), reduction of blocking artifacts, robustness against transmission errors, and filtering (for scalability). Also, a fast implementation algorithm for finite-length-signals using symmetric extensions is developed specially for the ELT with overlap factor 2 (ELT-2). This comparison shows that ELT-2 is superior to both DCT and the lapped orthogonal transform (LOT).
Ricardo L. de Queiroz, Kamisetty Ramamohan Rao
IEEE Trans. Image Process.1
1994 On Symmetric Extensions, Orthogonal Transforms of Images, and Paraunitary Filter Banks
abstract
Periodic or symmetric extensions are commonly used for processing images and other finite-length signals with a paraunitary filter bank (PUFB). Unlike infinite-length signals, PUFBs applied to finite-length signals will not necessarily lead to an orthogonal system. The authors show that for symmetric extensions, orthogonality is only possible for special PUFBs based on linear-phase filters. They also discuss implementation issues.>
Ricardo L. de Queiroz, Kamisetty Ramamohan Rao
ICIP (1)1
1994 Generalized Linear-Phase Lapped Orthogonal Transforms
abstract
The general factorization of a linear-phase paraunitary filter bank (LPPUFB) is revisited and we introduce a class of lapped orthogonal transforms with extended overlap (GenLOT). In this formulation, the discrete cosine transform (DCT) is the order-1 GenLOT, the lapped orthogonal transform is the order-2 GenLOT, and so on, for any filter length which is an integer multiple of the block size. All GenLOTs are based on the DCT and have fast implementation algorithms. The degrees of freedom in the design of GenLOTs are described and design examples are presented along with some practical applications.>
Ricardo L. de Queiroz, Truong Q. Nguyen, Kamisetty Ramamohan Rao
ISCAS1
1993 Adaptive extended lapped transforms
Ricardo L. de Queiroz, Kamisetty Ramamohan Rao
ICASSP (3)1
1993 On adaptive wavelet packets
Ricardo L. de Queiroz, Kamisetty Ramamohan Rao
ISCAS1
1993 Improved Chen-Smith image coder
Eduardo M. Rubino, Henrique S. Malvar, Ricardo L. de Queiroz
ISCAS3
1992 Subband processing of finite length signals without border distortions
abstract
New equations are derived for the perfect reconstruction of the boundary regions of a subband processed image. These equations are valid for filter banks with an arbitrary number of filters, having arbitrary nonlinear phase, and allowing any border extension method in the analysis section. The filter bank is maximally decimated in a wide sense, i.e., the sum of samples among the subbands is equal to the sum of the image samples, without having to store extended subband signals.>
Ricardo L. de Queiroz
ICASSP1