EDBT 2026 Demo / reviewers in the wild / expert
Samuel Fernández-Menduiña
dblp:268/1276
· DBLP profile ↗
10ranked-venue papers
10as first author
9since 2021 · last 2026
0000-0002-0123-607XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 first-author · 8 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Image Coding for Machines via Feature-Preserving Rate-Distortion OptimizationabstractMany images and videos are primarily processed by computer vision algorithms, involving only occasional human inspection. When this content requires compression before processing, e.g., in distributed applications, coding methods must optimize for both visual quality and downstream task performance. We first show that, given the features obtained from the original and the decoded images, an approach to reduce the effect of compression on a task loss is to perform rate-distortion optimization (RDO) using the distance between features as a distortion metric. However, optimizing directly such a rate-distortion trade-off requires an iterative workflow of encoding, decoding, and feature evaluation for each coding parameter, which is computationally impractical. We address this problem by simplifying the RDO formulation to make the distortion term computable using block-based encoders. We first apply Taylor's expansion to the feature extractor, recasting the feature distance as a quadratic metric with the Jacobian matrix of the neural network. Then, we replace the linearized metric with a block-wise approximation, which we call input-dependent squared error (IDSE). To reduce computational complexity, we approximate IDSE using Jacobian sketches. The resulting loss can be evaluated block-wise in the transform domain and combined with the sum of squared errors (SSE) to address both visual quality and computer vision performance. Simulations with AVC across multiple feature extractors and downstream neural networks show up to 10% bit-rate savings for the same computer vision accuracy compared to RDO based on SSE, with no decoder complexity overhead and just a 7% encoder complexity increase. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega |
IEEE Trans. Multim. | 1 |
| 2025 | Fast DCT+: A Family of Fast Transforms Based on Rank-One Updates of the Path GraphabstractThis paper develops fast graph Fourier transform (GFT) algorithms with O(nlogn) runtime complexity for rank-one updates of the path graph. We first show that several commonly-used audio and video coding transforms belong to this class of GFTs, which we denote by DCT+. Next, starting from an arbitrary generalized graph Laplacian and using rank-one perturbation theory, we provide a factorization for the GFT after perturbation. This factorization is our central result and reveals a progressive structure: we first apply the unperturbed Laplacian’s GFT and then multiply the result by a Cauchy matrix. By specializing this decomposition to path graphs and exploiting the properties of Cauchy matrices, we show that Fast DCT+ algorithms exist. We also demonstrate that progressivity can speed up computations in applications involving multiple transforms related by rank-one perturbations (e.g., video coding) when combined with pruning strategies. Our results can be extended to other graphs and rank-k perturbations. Runtime analyses show that Fast DCT+ provides computational gains over the naive method for graph sizes larger than 64, with runtime approximately equal to that of 8 DCTs. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega |
ICASSP | 1 |
| 2025 | Rate-Distortion Optimization with Non-Reference Metrics for UGC CompressionabstractService providers must encode a large volume of noisy videos to meet the demand for user-generated content (UGC) in online video-sharing platforms. However, low-quality UGC challenges conventional codecs based on rate-distortion optimization (RDO) with full-reference metrics (FRMs). While effective for pristine videos, FRMs drive codecs to preserve artifacts when the input is degraded, resulting in suboptimal compression. A more suitable approach used to assess UGC quality is based on non-reference metrics (NRMs). However, RDO with NRMs as a measure of distortion requires an iterative workflow of encoding, decoding, and metric evaluation, which is computationally impractical. This paper overcomes this limitation by linearizing the NRM around the uncompressed video. The resulting cost function enables block-wise bit allocation in the transform domain by estimating the alignment of the quantization error with the gradient of the NRM. To avoid large deviations from the input, we add sum of squared errors (SSE) regularization. We derive expressions for both the SSE regularization parameter and the Lagrangian, akin to the relationship used for SSE-RDO. Experiments with images and videos show bitrate savings of more than 30% over SSE-RDO using the target NRM, with no decoder complexity overhead and minimal encoder complexity increase. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega, Neil Birkbeck, Balu Adsumilli |
ICIP | 1 |
| 2025 | INT-DTT+: Low-Complexity Data-Dependent Transforms for Video Coding
Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega, Tsung-Wei Huang, Thuong Nguyen Canh, Guan-Ming Su, Peng Yin 0002 |
PCS | 1 |
| 2024 | Tracking Beyond the Unambiguous Range with Modulo Single-Photon LidarabstractIn single photon lidar (SPL), the laser repetition rate sets the maximum distance that can be recovered unambiguously. Conventional SPL extends this maximum recordable depth by reducing the repetition rate; however, the slower acquisition speed limits the number of received photons, which may be insufficient to track fast-moving objects. Inspired by recent successes in modulo sensing, we leverage the smoothness of typical trajectories to achieve long-range tracking beyond the unambiguous range. Although SPL naturally acquires modulo time-of-flight measurements, it introduces several challenges—including random sampling times, multiple noise sources, and absolute distance uncertainty—that are not addressed by the current modulo sensing literature. Hence, we propose an interpolation and denoising method that operates directly over the modulo samples. We further disambiguate the absolute distance based on the changing reflectivity fall-off. Monte Carlo simulations considering realistic trajectories under practical conditions show that, when properly unwrapped, the normalized mean squared error of our depth estimate decreases by over 20 dB with respect to a lidar setup whose repetition period leads to no ambiguity. Samuel Fernández-Menduiña, Joshua Rapp, Hassan Mansour, M. Greiff, Kieran Parsons |
ICASSP | 1 |
| 2024 | Feature-Preserving Rate-Distortion Optimization in Image Coding for MachinesabstractWith the increasing number of images and videos consumed by computer vision algorithms, compression methods are evolving to consider both perceptual quality and performance in downstream tasks. Traditional codecs can tackle this problem by performing rate-distortion optimization (RDO) to minimize the distance at the output of a feature extractor. However, neural network non-linearities can make the rate-distortion landscape irregular, leading to reconstructions with poor visual quality even for high bit rates. Moreover, RDO decisions are made block-wise, while the feature extractor requires the whole image to exploit global information. In this paper, we address these limitations in three steps. First, we apply Taylor's expansion to the feature extractor, recasting the metric as an input-dependent squared error involving the Jacobian matrix of the neural network. Second, we make a localization assumption to compute the metric block-wise. Finally, we use randomized dimensionality reduction techniques to approximate the Jacobian. The resulting expression is monotonic with the rate and can be evaluated in the transform domain. Simulations with AVC show that our approach provides bit-rate savings while preserving accuracy in downstream tasks with less complexity than using the feature distance directly. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega |
MMSP | 1 |
| 2023 | Image Coding Via Perceptually Inspired Graph LearningabstractMost codec designs rely on the mean squared error (MSE) as a fidelity metric in rate-distortion optimization, which allows to choose the optimal parameters in the transform domain but may fail to reflect perceptual quality. Alternative distortion metrics, such as the structural similarity index (SSIM), can be computed only pixel-wise, so they cannot be used directly for transform-domain bit allocation. Recently, the irregularity-aware graph Fourier transform (IAGFT) emerged as a means to include pixel-wise perceptual information in the transform design. This paper extends this idea by also learning a graph (and corresponding transform) for sets of blocks that share similar perceptual characteristics and are observed to differ statistically, leading to different learned graphs. We demonstrate the effectiveness of our method with both SSIM- and saliency-based criteria. We also propose a framework to derive separable transforms, including separable IAGFTs. An empirical evaluation based on the 5th CLIC dataset shows that our approach achieves improvements in terms of MS-SSIM with respect to existing methods. Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega |
ICIP | 1 |
| 2023 | Source camera attribution via PRNU emphasis: Towards a generalized multiplicative modelabstractThe photoresponse non-uniformity (PRNU) is a camera-specific pattern, which acts as unique fingerprint of any imaging sensor and thus is widely adopted to solve multimedia forensics problems such as device identification or forgery detection. Customarily, the theoretical analysis of this fingerprint relies on a multiplicative model for the denoising residuals. This setup assumes that the nonlinear mapping from scene irradiance to preprocessed luminance, that is, the composition of the Camera Response Function (CRF) with the digital preprocessing pipeline, is a gamma correction. However, this assumption seldom holds in practice. In this paper, we improve the multiplicative model by including the influence of this nonlinear mapping, termed PRNU emphasis, on the denoising residuals. On the theoretical side, we conduct first an exploratory analysis to show that the response of typical cameras deviates from a gamma correction. We also propose a regularized least squares estimator to measure this effect. On the practical side, we argue that the PRNU emphasis is especially beneficial for a source camera attribution problem with cropped images. We back our argument with an extensive empirical evaluation using different denoisers and both compressed and uncompressed images. This new model will pave the way to future PRNU estimators and detectors. Samuel Fernández-Menduiña, Fernando Pérez-González, Miguel Masciopinto |
Signal Process. Image Commun. | 1 |
| 2021 | On the information leakage quantification of camera fingerprint estimatesabstractAbstract Camera fingerprints based on sensor PhotoResponse Non-Uniformity (PRNU) have gained broad popularity in forensic applications due to their ability to univocally identify the camera that captured a certain image. The fingerprint of a given sensor is extracted through some estimation method that requires a few images known to be taken with such sensor. In this paper, we show that the fingerprints extracted in this way leak a considerable amount of information from those images used in the estimation, thus constituting a potential threat to privacy. We propose to quantify the leakage via two measures: one based on the Mutual Information, and another based on the output of a membership inference test. Experiments with practical fingerprint estimators on a real-world image dataset confirm the validity of our measures and highlight the seriousness of the leakage and the importance of implementing techniques to mitigate it. Some of these techniques are presented and briefly discussed. Samuel Fernández-Menduiña, Fernando Pérez-González |
EURASIP J. Inf. Secur. | 1 |
| 2020 | Temporal Localization of Non-Static Digital Videos Using the Electrical Network FrequencyabstractNon-static scenes represent one of the main barriers for retrieving the electrical network frequency (ENF) from digital videos recorded in realistic scenarios, since movement influences the luminance of pixels, hindering the recovery of the desired information. Aiming to mitigate the effects of changes in the scene, in this letter a video processing stage, that detects and combines those pixels that are not affected by movement, is proposed. The sequence obtained thereby is delivered to a PLL-based FM demodulator, that estimates the ENF taking into account its autoregressive nature. Additionally, a frame rate estimation algorithm is employed to avoid time-consuming trial and error loops in the extraction process. The performance of the system is tested by computing the probability of estimating correctly the time of recording of a given video, using ground-truth sequences obtained directly from the mains or a reliable database. Samuel Fernández-Menduiña, Fernando Pérez-González |
IEEE Signal Process. Lett. | 1 |