VLDB 2026 Research / reviewers in the wild / expert
Eduardo Pavez
dblp:72/11328
· DBLP profile ↗
42ranked-venue papers
12as first author
35since 2021 · last 2026
0000-0001-8985-2872ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 10 first-author · 33 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Region-Adaptive Learned Hierarchical Encoding for 3D Gaussian Splatting DataabstractWe introduce Region-Adaptive Learned Hierarchical Encoding (RALHE) for 3D Gaussian Splatting (3DGS) data. While 3DGS has recently become popular for novel view synthesis, the size of trained models limits its deployment in bandwidth-constrained applications such as volumetric media streaming. To address this, we propose a learned hierarchical latent representation that builds upon the principles of “overfitted” learned image compression (e.g., Cool-Chic and C3) to efficiently encode 3DGS attributes. Unlike images, 3DGS data have irregular spatial distributions of Gaussians (geometry) and consist of multiple attributes (signals) defined on the irregular geometry. Our codec is designed to account for these differences between images and 3DGS. Specifically, we leverage the octree structure of the voxelized 3DGS geometry to obtain a hierarchical multi-resolution representation. Our approach overfits latents to each Gaussian attribute under a global rate constraint. These latents are decoded independently through a lightweight decoder network. To estimate the bitrate during training, we employ an autoregressive probability model that leverages octree-derived contexts from the 3D point structure. The multi-resolution latents, decoder, and autoregressive entropy coding networks are jointly optimized for each Gaussian attribute. Experiments on 3DGS models from the Synthetic-NeRF dataset demonstrate that the proposed RALHE compression framework achieves a rendering PSNR gain of up to 2 dB at low bitrates ($\leq 1 \text{MB}$) compared to the baseline 3DGS compression methods. Shashank N. Sridhara, Birendra Kathariya, Fangjun Pu, Peng Yin 0002, Eduardo Pavez, Antonio Ortega |
DCC | 5 |
| 2026 | Image Coding for Machines via Feature-Preserving Rate-Distortion OptimizationabstractMany images and videos are primarily processed by computer vision algorithms, involving only occasional human inspection. When this content requires compression before processing, e.g., in distributed applications, coding methods must optimize for both visual quality and downstream task performance. We first show that, given the features obtained from the original and the decoded images, an approach to reduce the effect of compression on a task loss is to perform rate-distortion optimization (RDO) using the distance between features as a distortion metric. However, optimizing directly such a rate-distortion trade-off requires an iterative workflow of encoding, decoding, and feature evaluation for each coding parameter, which is computationally impractical. We address this problem by simplifying the RDO formulation to make the distortion term computable using block-based encoders. We first apply Taylor's expansion to the feature extractor, recasting the feature distance as a quadratic metric with the Jacobian matrix of the neural network. Then, we replace the linearized metric with a block-wise approximation, which we call input-dependent squared error (IDSE). To reduce computational complexity, we approximate IDSE using Jacobian sketches. The resulting loss can be evaluated block-wise in the transform domain and combined with the sum of squared errors (SSE) to address both visual quality and computer vision performance. Simulations with AVC across multiple feature extractors and downstream neural networks show up to 10% bit-rate savings for the same computer vision accuracy compared to RDO based on SSE, with no decoder complexity overhead and just a 7% encoder complexity increase. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega |
IEEE Trans. Multim. | 2 |
| 2025 | Fast DCT+: A Family of Fast Transforms Based on Rank-One Updates of the Path GraphabstractThis paper develops fast graph Fourier transform (GFT) algorithms with O(nlogn) runtime complexity for rank-one updates of the path graph. We first show that several commonly-used audio and video coding transforms belong to this class of GFTs, which we denote by DCT+. Next, starting from an arbitrary generalized graph Laplacian and using rank-one perturbation theory, we provide a factorization for the GFT after perturbation. This factorization is our central result and reveals a progressive structure: we first apply the unperturbed Laplacian’s GFT and then multiply the result by a Cauchy matrix. By specializing this decomposition to path graphs and exploiting the properties of Cauchy matrices, we show that Fast DCT+ algorithms exist. We also demonstrate that progressivity can speed up computations in applications involving multiple transforms related by rank-one perturbations (e.g., video coding) when combined with pruning strategies. Our results can be extended to other graphs and rank-k perturbations. Runtime analyses show that Fast DCT+ provides computational gains over the naive method for graph sizes larger than 64, with runtime approximately equal to that of 8 DCTs. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega |
ICASSP | 2 |
| 2025 | Graph-based Signal Sampling with Adaptive Subspace Reconstruction for Spatially-irregular Sensor DataabstractChoosing an appropriate frequency definition and norm is critical in graph signal sampling and reconstruction. Most previous works define frequencies based on the spectral properties of the graph and use the same frequency definition and ℓ2-norm for optimization for all sampling sets. Our previous work demonstrated that using a sampling-set-dependent norm (and corresponding frequency definition) can address challenges in conventional bandlimited approximations for graph signals, particularly with model mismatches and irregularly distributed data. This work proposes a method for selecting sampling sets tailored to the sampling-set-adaptive GFT-based interpolation. When the graph models the inverse covariance of the data, we show that this adaptive GFT enables tracking bandlimited model mismatch error and its effect in reconstruction, leveraging the spectral folding property, analogous to aliasing error in classical DSP. We propose a sampling set selection algorithm to minimize the worst-case bandlimited model mismatch error. We consider partitioning a set of sensors sampling a continuous spatial process as an application. Our experiments show that sampling and reconstruction using sampling-set-adaptive GFT significantly outperform methods that used fixed GFTs and bandwidth-based criterion. Darukeesan Pakiyarajah, Eduardo Pavez, Antonio Ortega |
ICASSP | 2 |
| 2025 | No-Reference Point Cloud Quality Assessment Based on Graph Signal VariationabstractIn real-time applications utilizing point clouds, no-reference point cloud quality assessment (NR-PCQA) methods are essential to improve the accuracy of downstream tasks. For example, in point cloud denoising, NR-PCQA results can be benchmarks for determining the optimal parameters when reference data are unavailable. This paper presents an accurate and fast NR-PCQA method based on graph signal processing. First, we propose new features derived from graph signal variation (GSV) to train a support vector regression model. These features improve the correlation with subjective scores and the robustness against inaccurate graph construction. Second, we present a point selection technique based on graph edge weights that allows us to exclude less relevant points, which results in a precise PCQA. Third, we propose a diagonal scan-line graph (DSLG) construction with a superior tradeoff between accurate and fast graph construction. Our experiments demonstrate improved accuracy and computation time compared with conventional methods with three types of open datasets. Ryosuke Watanabe, Keisuke Nonaka, Eduardo Pavez, Tatsuya Kobayashi, Antonio Ortega |
ICASSP | 3 |
| 2025 | Rate-Distortion Optimization with Non-Reference Metrics for UGC CompressionabstractService providers must encode a large volume of noisy videos to meet the demand for user-generated content (UGC) in online video-sharing platforms. However, low-quality UGC challenges conventional codecs based on rate-distortion optimization (RDO) with full-reference metrics (FRMs). While effective for pristine videos, FRMs drive codecs to preserve artifacts when the input is degraded, resulting in suboptimal compression. A more suitable approach used to assess UGC quality is based on non-reference metrics (NRMs). However, RDO with NRMs as a measure of distortion requires an iterative workflow of encoding, decoding, and metric evaluation, which is computationally impractical. This paper overcomes this limitation by linearizing the NRM around the uncompressed video. The resulting cost function enables block-wise bit allocation in the transform domain by estimating the alignment of the quantization error with the gradient of the NRM. To avoid large deviations from the input, we add sum of squared errors (SSE) regularization. We derive expressions for both the SSE regularization parameter and the Lagrangian, akin to the relationship used for SSE-RDO. Experiments with images and videos show bitrate savings of more than 30% over SSE-RDO using the target NRM, with no decoder complexity overhead and minimal encoder complexity increase. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega, Neil Birkbeck, Balu Adsumilli |
ICIP | 3 |
| 2025 | Joint Optimization of Primary and Secondary Transforms Using Rate-Distortion Optimized Transform DesignabstractData-dependent transforms are increasingly being incorporated into next-generation video coding systems such as AVM, a codec under development by the Alliance for Open Media (AOM), and VVC. To circumvent the computational complexities associated with implementing non-separable data-dependent transforms, combinations of separable primary transforms and non-separable secondary transforms have been studied and integrated into video coding standards. These codecs often utilize rate-distortion optimized transforms (RDOT) to ensure that the new transforms complement existing transforms like the DCT and the ADST. In this work, we propose an optimization framework for jointly designing primary and secondary transforms from data through a rate-distortion optimized clustering. Primary transforms are assumed to follow a path-graph model, while secondary transforms are non-separable. We empirically evaluate our proposed approach using AVM residual data and demonstrate that 1) the joint clustering method achieves lower total RD cost in the RDOT design framework, and 2) jointly optimized separable path-graph transforms (SPGT) provide better coding efficiency compared to separable KLTs obtained from the same data. Darukeesan Pakiyarajah, Eduardo Pavez, Antonio Ortega, Debargha Mukherjee, Onur G. Guleryuz, Keng-Shih Lu, Kruthika Koratti Sivakumar |
ICIP | 2 |
| 2025 | Adaptive Voxelization for Transform Coding of 3D Gaussian Splatting DataabstractWe present a novel compression framework for 3D Gaussian splatting (3DGS) data that leverages transform coding tools originally developed for point clouds. Contrary to existing 3DGS compression methods, our approach can produce compressed 3DGS models at multiple bitrates in a computationally efficient way. Point cloud voxelization is a discretization technique that point cloud codecs use to improve coding efficiency while enabling the use of fast transform coding algorithms. We propose an adaptive voxelization algorithm tailored to 3DGS data, to avoid the inefficiencies introduced by uniform voxelization used in point cloud codecs. We ensure the positions of larger volume Gaussians are represented at high resolution, as these significantly impact rendering quality. Meanwhile, a low-resolution representation is used for dense regions with smaller Gaussians, which have a relatively lower impact on rendering quality. This adaptive voxelization approach significantly reduces the number of Gaussians and the bitrate required to encode the 3DGS data. After voxelization, many Gaussians are moved or eliminated. Thus, we propose to fine-tune/recolor the remaining 3DGS attributes with an initialization that can reduce the amount of retraining required. Experimental results on pre-trained datasets show that our proposed compression framework outperforms existing methods. Chenjunjie Wang, Shashank N. Sridhara, Eduardo Pavez, Antonio Ortega |
ICIP | 3 |
| 2025 | INT-DTT+: Low-Complexity Data-Dependent Transforms for Video Coding
Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega, Tsung-Wei Huang, Thuong Nguyen Canh, Guan-Ming Su, Peng Yin 0002 |
PCS | 2 |
| 2025 | Understanding encoder-decoder structures in machine learning using information measuresabstractWe present a theory of representation learning to model and understand the role of encoder–decoder design in machine learning (ML) from an information-theoretic angle. We use two main information concepts, information sufficiency (IS) and mutual information loss to represent predictive structures in machine learning. Our first main result provides a functional expression that characterizes the class of probabilistic models consistent with an IS encoder–decoder latent predictive structure. This result formally justifies the encoder–decoder forward stages many modern ML architectures adopt to learn latent (compressed) representations for classification. To illustrate IS as a realistic and relevant model assumption, we revisit some known ML concepts and present some interesting new examples: invariant, robust, sparse, and digital models. Furthermore, our IS characterization allows us to tackle the fundamental question of how much performance could be lost, using the cross entropy risk, when a given encoder–decoder architecture is adopted in a learning setting. Here, our second main result shows that a mutual information loss quantifies the lack of expressiveness attributed to the choice of a (biased) encoder–decoder ML design. Finally, we address the problem of universal cross-entropy learning with an encoder–decoder design where necessary and sufficiency conditions are established to meet this requirement. In all these results, Shannon’s information measures offer new interpretations and explanations for representation learning. Jorge F. Silva, Victor Faraggi, Camilo Ramírez, Alvaro Egaña, Eduardo Pavez |
Signal Process. | 5 |
| 2025 | Full reference point cloud quality assessment using support vector regression
Ryosuke Watanabe, Shashank N. Sridhara, Haoran Hong, Eduardo Pavez, Keisuke Nonaka, Tatsuya Kobayashi, Antonio Ortega |
Signal Process. Image Commun. | 4 |
| 2024 | Irregularity-Aware Bandlimited Approximation for Graph Signal InterpolationabstractIn most work to date, graph signal sampling and reconstruction algorithms are intrinsically tied to graph properties, assuming bandlimitedness and optimal sampling set choices. However, practical scenarios often defy these assumptions, leading to suboptimal performance. In the context of sampling and reconstruction, graph irregularities lead to varying contributions from sampled nodes for interpolation and differing levels of reliability for interpolated nodes. The existing graph Fourier transform (GFT)-based methods in the literature make bandlimited signal approximations without considering graph irregularities and the relative significance of nodes, resulting in suboptimal reconstruction performance under various mismatch conditions. In this paper, we leverage the GFT equipped with a specific inner product to address graph irregularities and account for the relative importance of nodes during the bandlimited signal approximation and interpolation process. Empirical evidence demonstrates that the proposed method outperforms other GFT-based approaches for bandlimited signal interpolation in challenging scenarios, such as sampling sets selected independently of the underlying graph, low sampling rates, and high noise levels. Darukeesan Pakiyarajah, Eduardo Pavez, Antonio Ortega |
ICASSP | 2 |
| 2024 | Fast Graph-Based Denoising For Point Cloud Color InformationabstractPoint clouds are utilized in various 3D applications such as cross-reality (XR) and realistic 3D displays. In some applications, e.g., for live streaming using a 3D point cloud, real-time point cloud denoising methods are required to enhance the visual quality. However, conventional high-precision denoising methods cannot be executed in real time for large-scale point clouds owing to the complexity of graph constructions with K nearest neighbors and noise level estimation. This paper proposes a fast graph-based denoising (FGBD) for a large-scale point cloud. First, high-speed graph construction is achieved by scanning a point cloud in various directions and searching adjacent neighborhoods on the scanning lines. Second, we propose a fast noise level estimation method using eigenvalues of the covariance matrix on a graph. Finally, we also propose a new low-cost filter selection method to enhance denoising accuracy to compensate for the degradation caused by the acceleration algorithms. In our experiments, we succeeded in reducing the processing time dramatically while maintaining accuracy relative to conventional denoising methods. Denoising was performed at 30fps, with frames containing approximately 1 million points. Ryosuke Watanabe, Keisuke Nonaka, Eduardo Pavez, Tatsuya Kobayashi, Antonio Ortega |
ICASSP | 3 |
| 2024 | Full-Reference Point Cloud Quality Assessment Using Spectral Graph WaveletsabstractPoint clouds in 3D applications frequently experience quality degradation during processing, e.g., scanning and compression. Reliable point cloud quality assessment (PCQA) is important for developing compression algorithms with good bitrate-quality trade-offs and techniques for quality improvement (e.g., denoising). This paper introduces a full-reference (FR) PCQA method utilizing spectral graph wavelets (SGWs). First, we propose novel SGW-based PCQA metrics that compare SGW coefficients of coordinate and color signals between reference and distorted point clouds. Second, we achieve accurate PCQA by integrating several conventional FR metrics and our SGW-based metrics using support vector regression. To our knowledge, this is the first study to introduce SGWs for PCQA. Experimental results demonstrate the proposed PCQA metric is more accurately correlated with subjective quality scores compared to conventional PCQA metrics. Ryosuke Watanabe, Keisuke Nonaka, Eduardo Pavez, Tatsuya Kobayashi, Antonio Ortega |
ICIP | 3 |
| 2024 | Feature-Preserving Rate-Distortion Optimization in Image Coding for MachinesabstractWith the increasing number of images and videos consumed by computer vision algorithms, compression methods are evolving to consider both perceptual quality and performance in downstream tasks. Traditional codecs can tackle this problem by performing rate-distortion optimization (RDO) to minimize the distance at the output of a feature extractor. However, neural network non-linearities can make the rate-distortion landscape irregular, leading to reconstructions with poor visual quality even for high bit rates. Moreover, RDO decisions are made block-wise, while the feature extractor requires the whole image to exploit global information. In this paper, we address these limitations in three steps. First, we apply Taylor's expansion to the feature extractor, recasting the metric as an input-dependent squared error involving the Jacobian matrix of the neural network. Second, we make a localization assumption to compute the metric block-wise. Finally, we use randomized dimensionality reduction techniques to approximate the Jacobian. The resulting expression is monotonic with the rate and can be evaluated in the transform domain. Simulations with AVC show that our approach provides bit-rate savings while preserving accuracy in downstream tasks with less complexity than using the feature distance directly. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega |
MMSP | 2 |
| 2024 | Color-Guided Flying Pixel Correction in Depth ImagesabstractWe present a novel method to correct flying pixels within data captured by Time-of-flight (ToF) sensors. Flying pixel (FP) artifacts occur when signals from foreground and background objects reach the same sensor pixel, leading to a confident yet incorrect depth estimation in space-floating between two objects. Commercial RGB-D cameras have a complementary setup consisting of ToF sensors to capture depth in addition to RGB cameras. We propose a novel method to correct FPs by leveraging the aligned RGB and depth image in such RGB-D cameras to estimate the true depth values of FPs. Our method defines a 3D neighborhood around each point, representing a “field of view” that mirrors the acquisition process of ToF cameras. We propose a two-step iterative correction algorithm in which the FPs are first identified. Then, we estimate the true depth value of FPs by solving a least-squares optimization problem. Experimental results show that our proposed algorithm estimates the depth value of FPs as accurately as other algorithms in the literature. Ekamresh Vasudevan, Shashank N. Sridhara, Eduardo Pavez, Antonio Ortega, Raghavendra Singh, Srinath Kalluri |
MMSP | 3 |
| 2024 | Adaptive Online Learning of Separable Path Graph Transforms for Intra-PredictionabstractCurrent video coding standards, including H.264/AVC, HEVC, and VVC, employ discrete cosine transform (DCT), discrete sine transform (DST), and secondary Karhunen-Loéve transforms (KLTs) to decorrelate the intra-prediction residuals. However, the efficiency of these transforms in decorrelation can be limited when the signal has a non-smooth and non-periodic structure, such as those occurring in textures with intricate patterns. This paper introduces a novel adaptive separable path graph-based transform (GBT) that can provide better decorrelation than the DCT for intra-predicted texture data. The proposed GBT is learned in an online scenario with sequential$K$-means clustering, which groups similar blocks during encoding and decoding to adaptively learn the GBT for the current block from previously reconstructed areas with similar characteristics. A signaling overhead is added to the bitstream of each coding block to indicate the usage of the proposed graph-based transform. We assess the performance of this method combined with H.264/AVC intra-coding tools and demonstrate that it can significantly outperform H.264/AVC DCT for intra-predicted texture data. Wen-Yang Lu, Eduardo Pavez, Antonio Ortega, Xin Zhao 0003, Shan Liu 0001 |
PCS | 2 |
| 2023 | Graph-Based Point Cloud Color Denoising with 3-Dimensional Patch-Based SimilarityabstractPoint clouds are utilized in many 3-D applications such as cross-reality (XR) and realistic 3-D display. They consist of a set of points with 3-D coordinates and associated color signals. These color signals are often perturbed by noise induced by the measurement errors of scanning devices. In this paper, we propose a point cloud denoising method for color signals. Since many conventional methods for point cloud color denoising are based on a low-pass filter in the graph spectral domain, denoising accuracy is affected by the choice of graph. We propose a graph construction method using 3-D patch-based similarity, in which the similarity is calculated with small 3-D patches around the connected points. This is in contrast with conventional graph construction methods for denoising, which are based on point properties such as pairwise point distances and differences in color. Second, we propose a low-pass filtering method where the frequency response is chosen automatically depending on the estimated noise level. Our experimental results show that our proposed method, 3-D patch-based similarity (3DPBS), achieves the best denoising accuracy compared with graph-based state-of-the-art methods. Ryosuke Watanabe, Keisuke Nonaka, Eduardo Pavez, Tatsuya Kobayashi, Antonio Ortega |
ICASSP | 3 |
| 2023 | Graph Wavelet-Based Point Cloud Geometric Denoising with Surface-Consistent Non-Negative Kernel RegressionabstractPoint cloud applications suffer from geometric noise caused by measurement errors induced by the point cloud acquisition system. We propose a novel graph construction method, surface-consistent non-negative kernel regression (SC-NNK), that can achieve more accurate denoising of geometry information in combination with spectral graph wavelet transforms (SGWTs). Unlike conventional graph construction methods such as the K-nearest neighbor (KNN), which have been adopted in previous SGWT-based geometry denoising methods, SC-NNK graphs consider geometrical and frequency characteristics to remove redundant edge connections from a KNN graph. In addition, we propose a novel noise level estimation method that achieves improved accuracy by detecting flat surfaces in point clouds, resulting in better wavelet shrinkage thresholds for denoising. Our experimental results show that the proposed method outperforms recent deep-learning-based and graph-based state-of-the-art denoising methods. Ryosuke Watanabe, Keisuke Nonaka, Eduardo Pavez, Tatsuya Kobayashi, Antonio Ortega |
ICASSP | 3 |
| 2023 | Rate-Distortion Optimization with Alternative References for UGC Video CompressionabstractUser generated content (UGC) refers to videos that are uploaded by users and shared over the Internet. UGC may have low quality due to noise and previous compression. When re-encoding UGC for streaming or downloading, a traditional video coding pipeline will perform rate-distortion (RD) optimization to choose coding parameters. However, in the UGC video coding case, since the input is not pristine, quality “saturation” (or even degradation) can be observed, i.e., increased bitrate only leads to improved representation of coding artifacts and noise present in the UGC input. In this paper, we study the saturation problem in UGC compression, where the goal is to identify and avoid during encoding, the coding parameters and rates that lead to quality saturation. We proposed a geometric criterion for saturation detection that works with rate-distortion optimization, and only requires a few frames from the UGC video. In addition, we show how to combine the proposed saturation detection method with existing video coding systems that implement rate-distortion optimization for efficient compression of UGC videos. Eduardo Pavez, Antonio Ortega, Balu Adsumilli |
ICASSP | 2 |
| 2023 | Image Coding Via Perceptually Inspired Graph LearningabstractMost codec designs rely on the mean squared error (MSE) as a fidelity metric in rate-distortion optimization, which allows to choose the optimal parameters in the transform domain but may fail to reflect perceptual quality. Alternative distortion metrics, such as the structural similarity index (SSIM), can be computed only pixel-wise, so they cannot be used directly for transform-domain bit allocation. Recently, the irregularity-aware graph Fourier transform (IAGFT) emerged as a means to include pixel-wise perceptual information in the transform design. This paper extends this idea by also learning a graph (and corresponding transform) for sets of blocks that share similar perceptual characteristics and are observed to differ statistically, leading to different learned graphs. We demonstrate the effectiveness of our method with both SSIM- and saliency-based criteria. We also propose a framework to derive separable transforms, including separable IAGFTs. An empirical evaluation based on the 5th CLIC dataset shows that our approach achieves improvements in terms of MS-SSIM with respect to existing methods. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega |
ICIP | 2 |
| 2022 | Laplacian Constrained Precision Matrix Estimation: Existence and High Dimensional ConsistencyabstractThis paper considers the problem of estimating high dimensional Laplacian constrained precision matrices by minimizing Stein’s loss. We obtain a necessary and sufficient condition for existence of this estimator, that consists on checking whether a certain data dependent graph is connected. We also prove consistency in the high dimensional setting under the symmetrized Stein loss. We show that the error rate does not depend on the graph sparsity, or other type of structure, and that Laplacian constraints are sufficient for high dimensional consistency. Our proofs exploit properties of graph Laplacians, the matrix tree theorem, and a characterization of the proposed estimator based on effective graph resistances. We validate our theoretical claims with numerical experiments. Eduardo Pavez |
AISTATS | 1 |
| 2022 | Fractional Motion Estimation for Point Cloud CompressionabstractMotivated by the success of fractional pixel motion in video coding, we explore the design of motion estimation with fractional-voxel resolution for compression of color attributes of dynamic 3D point clouds. Our proposed block-based fractional-voxel motion estimation scheme takes into account the fundamental differences between point clouds and videos, i.e., the irregularity of the distribution of voxels within a frame and across frames. We show that motion compensation can benefit from the higher resolution reference and more accurate displacements provided by fractional precision. Our proposed scheme significantly outperforms comparable methods that only use integer motion. The proposed scheme can be combined with and add sizeable gains to state-of-the-art systems that use transforms such as Region Adaptive Graph Fourier Transform and Region Adaptive Haar Transform. Haoran Hong, Eduardo Pavez, Antonio Ortega, Ryosuke Watanabe, Keisuke Nonaka |
DCC | 2 |
| 2022 | Graph-Based Point Cloud Denoising Using Shape-Aware Consistency For Free-Viewpoint VideoabstractWe propose a novel graph-based denoising method to correct the quantization error (step noise) arising in the process of generating the visual hull, a commonly used technique to synthesize free-viewpoint video. To reduce this step noise effectively, we propose two new notions of consistency, pixel value consistency and normal vector consistency. The resulting denoising method involves a first step of graph construction using the proposed consistency metrics, followed by graph filtering of the 3D point cloud coordinates. Our experiments show that our approach provides visually and quantitatively better performance than state-of-the-art methods. Keisuke Nonaka, Ryosuke Watanabe, Haruhisa Kato, Tatsuya Kobayashi, Eduardo Pavez, Antonio Ortega |
ICASSP | 5 |
| 2022 | Point Cloud Attribute Compression Via Chroma SubsamplingabstractWe introduce chroma subsampling for 3D point cloud attribute compression by proposing a novel technique to sample points irregularly placed in 3D space. While most current video compression standards use chroma subsampling, these chroma subsampling methods cannot be directly applied to 3D point clouds, given their irregularity and sparsity. In this work, we develop a framework to incorporate chroma subsampling into geometry-based point cloud encoders, such as region adaptive hierarchical transform (RAHT) and region adaptive graph Fourier transform (RAGFT). We propose different sampling patterns on a regular 3D grid to sample the points at different rates. We use a simple graph-based nearest neighbor interpolation technique to reconstruct the full resolution point cloud at the decoder end. Experimental results demonstrate that our proposed method provides significant coding gains with negligible impact on the reconstruction quality. For some sequences, we observe a bitrate reduction of 10-15% under the Bjontegaard metric. More generally, perceptual masking makes it possible to achieve larger bitrate reductions without visible changes in quality. Shashank N. Sridhara, Eduardo Pavez, Antonio Ortega, Ryosuke Watanabe, Keisuke Nonaka |
ICASSP | 2 |
| 2022 | Point Cloud Denoising Using Normal Vector-Based Graph Wavelet ShrinkageabstractMany applications that use point clouds, such as 3D immersive telepresence, suffer from geometric quality degradation. This noise may be caused by measurement errors of the capturing device or by the point cloud estimation method. In this paper, we propose a novel graph-based point cloud denoising approach using the spectral graph wavelet transform (SGWT) and graph wavelet shrinkage. Unlike conventional SGWT-based denoising methods, the proposed wavelet shrinkage thresholds are determined based on the normal vector at each point and are thus based on the local geometric structure of the point cloud. This approach avoids excessive wavelet shrinkage, which can lead to the loss of complex geometric structure. Experimental results show that the proposed method achieves the best accuracy as compared with recent deep-learning-based and graph-based state-of-the-art denoising methods. Ryosuke Watanabe, Keisuke Nonaka, Haruhisa Kato, Eduardo Pavez, Tatsuya Kobayashi, Antonio Ortega |
ICASSP | 4 |
| 2022 | Intra Prediction of Regular and Near-Regular Textures Via Graph-Based InpaintingabstractIntra prediction is an important technique to improve coding efficiency by exploiting the spatial redundancy present in typical video sequences. In video coding standards such as H.264/AVC, HEVC and VVC, directional predictors are utilized to generate prediction along a single direction within a block to be coded. However, these predictors fail to generate an accurate prediction when the block contains complex patterns such as periodic textures. In this paper, we propose a graph-based inpainting method that can handle both regular and near-regular textures. The proposed inpainting method utilizes a total variation model associated with the Laplacian matrix of a graph, whose edge weights are a function of pixel patch distance. We evaluate the performance of our proposed method as an additional prediction mode combined with the H.264/AVC coding standard. Experimental results show that the proposed method can significantly outperform H.264/AVC predictors in areas with high frequency periodic patterns. Wen-Yang Lu, Eduardo Pavez, Antonio Ortega, Debargha Mukherjee, Onur G. Guleryuz, Keng-Shih Lu |
ICIP | 2 |
| 2022 | Compression of User Generated Content Using Denoised ReferencesabstractVideo shared over the internet is commonly referred to as user generated content (UGC). UGC video may have low quality due to various factors including previous compression. UGC video is uploaded by users, and then it is re-encoded to be made available at various levels of quality. In a traditional video coding pipeline the encoder parameters are optimized to minimize a rate-distortion criterion, but when the input signal has low quality, this results in sub-optimal coding parameters optimized to preserve undesirable artifacts. In this paper we formulate the UGC compression problem as that of compression of a noisy/corrupted source. The noisy source coding theorem reveals that an optimal UGC compression system is comprised of optimal denoising of the UGC signal, followed by compression of the denoised signal. Since optimal denoising is unattainable and users may be against modification of their content, we propose encoding the UGC signal, and using denoised references only to compute distortion, so the encoding process can be guided towards perceptually better solutions. We demonstrate the effectiveness of the proposed strategy for JPEG compression of UGC images and videos. Eduardo Pavez, Enrique Perez, Antonio Ortega, Balu Adsumilli |
ICIP | 1 |
| 2022 | Motion Estimation And Filtered Prediction For Dynamic Point Cloud Attribute CompressionabstractIn point cloud compression, exploiting temporal redundancy for inter predictive coding is challenging because of the irregular geometry. This paper proposes an efficient block-based inter-coding scheme for color attribute compression. The scheme includes integer-precision motion estimation and an adaptive graph based in-loop filtering scheme for improved attribute prediction. The proposed block-based motion estimation scheme consists of an initial motion search that exploits geometric and color attributes, followed by a motion refinement that only minimizes color prediction error. To further improve color prediction, we propose a vertex-domain low-pass graph filtering scheme that can adaptively remove noise from predictors computed from motion estimation with different accuracy. Our experiments demonstrate significant coding gain over state-of-the-art coding methods. Haoran Hong, Eduardo Pavez, Antonio Ortega, Ryosuke Watanabe, Keisuke Nonaka |
PCS | 2 |
| 2021 | A Graph Learning Algorithm Based On Gaussian Markov Random Fields And Minimax Concave PenaltyabstractThis paper presents a graph learning framework to produce sparse and accurate graphs from network data. While our formulation is inspired by the graphical lasso, a key difference is the use of a nonconvex alternative of the ℓ1norm as well as a quadratic term to ensure overall convexity. Specifically, the weakly-convex minimax concave penalty (MCP) is used, which is given by subtracting the Huber function from the ℓ1norm, inducing a less-biased sparse solution than ℓ1. In our framework, the graph Laplacian is represented by a linear transform of the vector corresponding to its upper triangular part. Via a reformulation relying on the Moreau decomposition, the problem can be solved by the primal-dual splitting method. An admissible choice of parameters for provable convergence is presented. Numerical examples show that the proposed method significantly outperforms its ℓ1-based counterpart for sparse grid graphs. Tatsuya Koyakumaru, Masahiro Yukawa, Eduardo Pavez, Antonio Ortega |
ICASSP | 3 |
| 2021 | Spectral Folding And Two-Channel Filter-Banks On Arbitrary GraphsabstractIn the past decade, several multi-resolution representation theories for graph signals have been proposed. Bipartite filter-banks stand out as the most natural extension of time domain filter-banks, in part because perfect reconstruction, orthogonality and bi-orthogonality conditions in the graph spectral domain resemble those for traditional filter-banks. Therefore, many of the well known orthogonal and bi-orthogonal designs can be easily adapted for graph signals. A major limitation is that this framework can only be applied to the normalized Laplacian of bipartite graphs. In this paper we extend this theory to arbitrary graphs and positive semi-definite variation operators. Our approach is based on a different definition of the graph Fourier transform (GFT), where orthogonality is defined with respect to the Q inner product. We construct GFTs satisfying a spectral folding property, which allows us to easily construct orthogonal and bi-orthogonal perfect reconstruction filter-banks. We illustrate signal representation and computational efficiency of our filter-banks on 3D point clouds with hundreds of thousands of points. Eduardo Pavez, Benjamin Girault, Antonio Ortega, Philip A. Chou |
ICASSP | 1 |
| 2021 | Orthogonality and Zero DC Tradeoffs in Biorthogonal Graph FilterbanksabstractBiorthogonal graph wavelet filterbanks, also known as GraphBior, are one of the most popular graph transforms used in image compression, but up to now, they could be designed based on two known admissible fundamental matrices: i) the random walk Laplacian, which heavily penalizes low degree pixels, and ii) the normalized Laplacian, which lacks a zero-DC response. By exploiting a new extension of the admissibility condition in GraphBior we propose a new fundamental matrix with the goal of distributing the errors of GraphBior more uniformly across pixels with different node degrees. Furthermore the proposed matrix preserves high energy compaction linked to the zero-DC GraphBior variation. Dion Eustathios Olivier Tzamarias, Eduardo Pavez, Benjamin Girault, Antonio Ortega, Ian Blanes, Joan Serra-Sagristà |
ICASSP | 2 |
| 2021 | Multi-Resolution Intra-Predictive Coding Of 3d Point Cloud AttributesabstractWe propose an intra frame predictive strategy for compression of 3D point cloud attributes. Our approach is integrated with the region adaptive graph Fourier transform (RAGFT), a multi-resolution transform formed by a composition of localized block transforms, which produces a set of low pass (approximation) and high pass (detail) coefficients at multiple resolutions. Since the transform operations are spatially localized, RAGFT coefficients at a given resolution may still be correlated. To exploit this phenomenon, we propose an intraprediction strategy, in which decoded approximation coefficients are used to predict uncoded detail coefficients. The prediction residuals are then quantized and entropy coded. For the 8i dataset, we obtain gains up to 0. 5db as compared to intra predicted point cloud compresion based on the region adaptive Haar transform (RAHT). Eduardo Pavez, André L. Souto, Ricardo L. de Queiroz, Antonio Ortega |
ICIP | 1 |
| 2021 | Cylindrical Coordinates for Lidar Point Cloud CompressionabstractWe present an efficient voxelization method to encode the geometry and attributes of 3D point clouds obtained from autonomous vehicles. Due to the circular scanning trajectory of sensors, the geometry of LiDAR point clouds is inherently different from that of point clouds captured from RGBD cameras. Our method exploits these specific properties to representing points in cylindrical coordinates instead of conventional Cartesian coordinates. We demonstrate that Region Adaptive Hierarchical Transform (RAHT) can be extended to this setting, leading to attribute encoding based on a volumetric partition in cylindrical coordinates. Experimental results show that our proposed voxelization outperforms conventional approaches based on Cartesian coordinates for this type of data. We observe a significant improvement in attribute coding performance with 5-10% reduction in bitrate and octree representation with 35-45% reduction in bits. Shashank N. Sridhara, Eduardo Pavez, Antonio Ortega |
ICIP | 2 |
| 2021 | Covariance Matrix Estimation With Non Uniform and Data Dependent Missing ObservationsabstractIn this paper we study covariance estimation with missing data. We consider missing data mechanisms that can be independent of the data, or have a time varying dependency. Additionally, observed variables may have arbitrary (non uniform) and dependent observation probabilities. For each mechanism, we construct an unbiased estimator and obtain bounds for the expected value of their estimation error in operator norm. Our bounds are equivalent, up to constant and logarithmic factors, to state of the art bounds for complete and uniform missing observations. Furthermore, for the more general non uniform and dependent cases, the proposed bounds are new or improve upon previous results. Our error estimates depend on quantities we callscaled effective rank, which generalize the effective rank to account for missing observations. All the estimators studied in this work have the same asymptotic convergence rate (up to logarithmic factors). Eduardo Pavez, Antonio Ortega |
IEEE Trans. Inf. Theory | 1 |
| 2020 | Region Adaptive Graph Fourier Transform for 3D Point CloudsabstractWe introduce the Region Adaptive Graph Fourier Transform (RAGFT) for compression of 3D point cloud attributes. The RA-GFT is a multiresolution transform, formed by combining spatially localized block transforms. We assume the points are organized by a family of nested partitions represented by a rooted tree. At each resolution level, attributes are processed in clusters using block transforms. Each block transform produces a single approximation (DC) coefficient, and various detail (AC) coefficients. The DC coefficients are promoted up the tree to the next (lower resolution) level, where the process can be repeated until reaching the root. Since clusters may have a different numbers of points, each block transform must incorporate the relative importance of each coefficient. For this, we introduce the Q-normalized graph Laplacian, and propose using its eigenvectors as the block transform. The RA-GFT achieves better complexity-performance trade-offs than previous approaches. In particular, it outperforms the Region Adaptive Haar Transform (RAHT) by up to 2.5 dB, with a small complexity overhead. Eduardo Pavez, Benjamin Girault, Antonio Ortega, Philip A. Chou |
ICIP | 1 |
| 2018 | Active Covariance Estimation by Random Sub-Sampling of VariablesabstractWe study covariance matrix estimation for the case of partially observed random vectors, where different samples contain different subsets of vector coordinates. Each observation is the product of the variable of interest with a 0 - 1 Bernoulli random variable. We analyze an unbiased covariance estimator under this model, and derive an error bound that reveals relations between the sub-sampling probabilities and the entries of the covariance matrix. We apply our analysis in an active learning framework, where the expected number of observed variables is small compared to the dimension of the vector of interest, and propose a design of optimal sub-sampling probabilities and an active covariance matrix estimation algorithm. Eduardo Pavez, Antonio Ortega |
ICASSP | 1 |
| 2017 | Dynamic polygon cloud compressionabstractWe introduce a compressible representation of 3D geometry (including its attributes, such as color texture) intermediate between polygonal meshes and point clouds called a polygon cloud. Polygon clouds, compared to polygonal meshes, are more robust to live capture noise and artifacts. Furthermore, dynamic polygon clouds, compared to dynamic point clouds, are easier to compress, if certain challenges are addressed. In this paper, we propose methods for compressing dynamic polygon clouds using transform coding of color and motion residuals. We find that, compared to static polygon clouds and a fortiori static point clouds, dynamic polygon clouds can improve color compression by up to 2-3 dB in fidelity, and can improve geometry compression up to a factor of 2-5 in bit rate. Eduardo Pavez, Philip A. Chou |
ICASSP | 1 |
| 2017 | Learning separable transforms by inverse covariance estimationabstractOrthogonal transforms are one of the most important components of a video encoder system. They are applied to residual block images obtained as the difference between a target and its prediction. In this paper we propose a framework to design separable transforms from prediction residual statistics. We model the data as a 2D Gaussian Markov random field and approximate its inverse covariance by a matrix with a separable structure, thus explicitly constructing a separable orthonormal matrix that approximates the KLT. Our designed transforms can adapt to prediction residual statistics, have low complexity (compared to non separable transforms), require selecting few parameters and outperform hybrid DCT/ADST separable transform for intra coding of AV1 residuals. Eduardo Pavez, Antonio Ortega, Debargha Mukherjee |
ICIP | 1 |
| 2016 | Generalized Laplacian precision matrix estimation for graph signal processingabstractGraph signal processing models high dimensional data as functions on the vertices of a graph. This theory is constructed upon the interpretation of the eigenvectors of the Laplacian matrix as the Fourier transform for graph signals. We formulate the graph learning problem as a precision matrix estimation with generalized Laplacian constraints, and we propose a new optimization algorithm. Our formulation takes a covariance matrix as input and at each iteration updates one row/column of the precision matrix by solving a non-negative quadratic program. Experiments using synthetic data with generalized Laplacian precision matrix show that our method detects the nonzero entries and it estimates its values more precisely than the graphical Lasso. For texture images we obtain graphs whose edges follow the orientation. We show our graphs are more sparse than the ones obtained using other graph learning methods. Eduardo Pavez, Antonio Ortega |
ICASSP | 1 |
| 2015 | GTT: Graph template transforms with applications to image codingabstractThe Karhunen-Loeve transform (KLT) is known to be optimal for decorrelating stationary Gaussian processes, and it provides effective transform coding of images. Although the KLT allows efficient representations for such signals, the transform itself is completely data-driven and computationally complex. This paper proposes a new class of transforms called graph template transforms (GTTs) that approximate the KLT by exploiting a priori information known about signals represented by a graph-template. In order to construct a GTT (i) a design matrix leading to a class of transforms is defined, then (ii) a constrained optimization framework is employed to learn graphs based on given graph templates structuring a priori known information. Our experimental results show that some instances of the proposed GTTs can closely achieve the rate-distortion performance of KLT with significantly less complexity. Eduardo Pavez, Hilmi E. Egilmez, Yongzhe Wang, Antonio Ortega |
PCS | 1 |
| 2012 | Analysis and design of Wavelet-Packet Cepstral coefficients for automatic speech recognition
Eduardo Pavez, Jorge F. Silva |
Speech Commun. | 1 |