VLDB 2026 Research / reviewers in the wild / expert
Fernando Díaz-de-María
dblp:69/235
· DBLP profile ↗
74ranked-venue papers
2as first author
3since 2021 · last 2022
0000-0002-6437-914XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 2 first-authorArtificial intelligence and machine learning · 23 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
5 papers |
Image and video coding · 64% Image and video processing · 17% Audio and music processing · 15% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% | |
| Artificial intelligence
5 papers |
Speech recognition and synthesis · 52% Generative modeling · 21% Learning paradigms · 14% |
Topics — the 24 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video coding
video compression |
0.6 | 3 | 2016 | Complexity Control Based on a Fast Coding Unit Decision Method in the HEVC Video Coding Standard · IEEE Trans. Multim. 2016 Standard-Compliant Low-Pass Temporal Filter to Reduce the Perceived Flicker Artifact · IEEE Trans. Multim. 2014 Mode Decision-Based Algorithm for Complexity Control in H.264/AVC · IEEE Trans. Multim. 2013 |
Information retrieval
image retrieval |
0.5 | 2 | 2017 | Neighborhood Matching for Image Retrieval · IEEE Trans. Multim. 2017 A Generative Model for Concurrent Image Retrieval and ROI Segmentation · IEEE Trans. Multim. 2014 |
Information retrieval › image retrieval
large-scale image retrieval |
0.3 | 1 | 2017 | Neighborhood Matching for Image Retrieval · IEEE Trans. Multim. 2017 |
Information retrieval › image retrieval
spatial verification |
0.3 | 1 | 2017 | Neighborhood Matching for Image Retrieval · IEEE Trans. Multim. 2017 |
Image and video coding › video compression
coding unit partitioning |
0.2 | 1 | 2016 | Complexity Control Based on a Fast Coding Unit Decision Method in the HEVC Video Coding Standard · IEEE Trans. Multim. 2016 |
Image and video coding › video compression › video codec
HEVC |
0.2 | 1 | 2016 | Complexity Control Based on a Fast Coding Unit Decision Method in the HEVC Video Coding Standard · IEEE Trans. Multim. 2016 |
Machine learning › Generative modeling › generative model
probabilistic generative model |
0.2 | 1 | 2014 | A Generative Model for Concurrent Image Retrieval and ROI Segmentation · IEEE Trans. Multim. 2014 |
Image and video processing › video restoration
flicker removal |
0.2 | 1 | 2014 | Standard-Compliant Low-Pass Temporal Filter to Reduce the Perceived Flicker Artifact · IEEE Trans. Multim. 2014 |
Image and video processing › image sequence processing
temporal filtering |
0.2 | 1 | 2014 | Standard-Compliant Low-Pass Temporal Filter to Reduce the Perceived Flicker Artifact · IEEE Trans. Multim. 2014 |
Image and video coding › video coding standards
H.264/AVC |
0.2 | 1 | 2013 | Mode Decision-Based Algorithm for Complexity Control in H.264/AVC · IEEE Trans. Multim. 2013 |
Image and video coding › rate-distortion optimization
mode decision |
0.2 | 1 | 2013 | Mode Decision-Based Algorithm for Complexity Control in H.264/AVC · IEEE Trans. Multim. 2013 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.2 | 2 | 2011 | Data Balancing for Efficient Training of Hybrid ANN/HMM Automatic Speech Recognition Systems · IEEE Trans. Speech Audio Process. 2011 Recognizing voice over IP: a robust front-end for speech recognition on the world wide web · IEEE Trans. Multim. 2001 |
Audio and music processing › speech recognition
acoustic modeling |
0.1 | 1 | 2012 | Real-Time Robust Automatic Speech Recognition Using Compact Support Vector Machines · IEEE Trans. Speech Audio Process. 2012 |
Audio and music processing
speech recognition |
0.1 | 1 | 2012 | Real-Time Robust Automatic Speech Recognition Using Compact Support Vector Machines · IEEE Trans. Speech Audio Process. 2012 |
Natural language and speech › Speech recognition and synthesis
acoustic modeling |
0.1 | 1 | 2011 | Data Balancing for Efficient Training of Hybrid ANN/HMM Automatic Speech Recognition Systems · IEEE Trans. Speech Audio Process. 2011 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
hybrid HMM/ANN |
0.1 | 1 | 2011 | Data Balancing for Efficient Training of Hybrid ANN/HMM Automatic Speech Recognition Systems · IEEE Trans. Speech Audio Process. 2011 |
Machine learning › Learning paradigms › data balancing
training data balancing |
0.1 | 1 | 2011 | Data Balancing for Efficient Training of Hybrid ANN/HMM Automatic Speech Recognition Systems · IEEE Trans. Speech Audio Process. 2011 |
Machine learning › Probabilistic and Bayesian machine learning › clustering
sequence clustering |
0.1 | 1 | 2009 | A New Distance Measure for Model-Based Sequence Clustering · IEEE Trans. Pattern Anal. Mach. Intell. 2009 |
Multimedia systems and quality of experience › multimedia quality assessment
voice quality assessment |
0.1 | 1 | 2009 | Characterization of Healthy and Pathological Voice Through Measures Based on Nonlinear Dynamics · IEEE Trans. Speech Audio Process. 2009 |
Information theory › information measures › divergence measures
kullback-leibler divergence |
0.1 | 1 | 2009 | A New Distance Measure for Model-Based Sequence Clustering · IEEE Trans. Pattern Anal. Mach. Intell. 2009 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
robust speech recognition |
0.1 | 1 | 2005 | Recognizing GSM digital speech · IEEE Trans. Speech Audio Process. 2005 |
Audio and music processing › speech recognition
hidden markov model |
0.0 | 1 | 2012 | Real-Time Robust Automatic Speech Recognition Using Compact Support Vector Machines · IEEE Trans. Speech Audio Process. 2012 |
Machine learning › Representation and self-supervised learning › representation learning
feature extraction |
0.0 | 1 | 2005 | Recognizing GSM digital speech · IEEE Trans. Speech Audio Process. 2005 |
Internet architecture and protocols
voice over IP |
0.0 | 1 | 2001 | Recognizing voice over IP: a robust front-end for speech recognition on the world wide web · IEEE Trans. Multim. 2001 |
Methods — techniques the papers use, named apart from their topics
visual similarity matching · 0.4probabilistic generative model · 0.4geometric transformation modeling · 0.4spatial verification · 0.3keypoint matching · 0.3on-the-fly parameter estimation · 0.2hierarchical approach · 0.2spectral clustering · 0.2model selection · 0.2temporal low-pass filtering · 0.2on-the-fly filter strength estimation · 0.2gaussian statistics · 0.2binary hypothesis testing · 0.2adaptive thresholding · 0.2weighted least squares · 0.1support vector machine · 0.1compact semiparametric model · 0.1posterior probability scaling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | An interpretable CNN-based CAD system for skin lesion diagnosis
Javier López-Labraca, Iván González-Díaz 0001, Fernando Díaz-de-María, Alejandro Fueyo-Casado |
Artif. Intell. Medicine | 3 |
| 2022 | ACME: Automatic feature extraction for cell migration examination through intravital microscopy imagingabstractCell detection and tracking applied to in vivo fluorescence microscopy has become an essential tool in biomedicine to characterize 4D (3D space plus time) biological processes at the cellular level. Traditional approaches to cell motion analysis by microscopy imaging, although based on automatic frameworks, still require manual supervision at some points of the system. Hence, when dealing with a large amount of data, the analysis becomes incredibly time-consuming and typically yields poor biological information. In this paper, we propose a fully-automated system for segmentation, tracking and feature extraction of migrating cells within blood vessels in 4D microscopy imaging. Our system consists of a robust 3D convolutional neural network (CNN) for joint blood vessel and cell segmentation, a 3D tracking module with collision handling, and a novel method for feature extraction, which takes into account the particular geometry in the cell-vessel arrangement. Experiments on a large 4D intravital microscopy dataset show that the proposed system achieves a significantly better performance than the state-of-the-art tools for cell segmentation and tracking. Furthermore, we have designed an analytical method of cell behaviors based on the automatically extracted features, which supports the hypotheses related to leukocyte migration posed by expert biologists. This is the first time that such a comprehensive automatic analysis of immune cell migration has been performed, where the total population under study reaches hundreds of neutrophils and thousands of time instances. Miguel Molina-Moreno, Iván González-Díaz 0001, Jon Sicilia, Georgiana Crainiciuc, Miguel Palomino-Segura, Andrés Hidalgo, Fernando Díaz-de-María |
Medical Image Anal. | 7 |
| 2021 | Training deep retrieval models with noisy datasets: Bag exponential loss
Tomás Martínez-Cortés, Iván González-Díaz 0001, Fernando Díaz-de-María |
Pattern Recognit. | 3 |
| 2020 | Probabilistic Topic Model for Context-Driven Visual Attention UnderstandingabstractModern computer vision techniques have to deal with vast amounts of visual data, which implies a computational effort that has often to be accomplished in broad and challenging scenarios. The interest in efficiently solving these image and video applications has led researchers to develop methods to expertly drive the corresponding processing to conspicuous regions that either depend on the context or are based on specific requirements. In this paper, we propose a general hierarchical probabilistic framework, independent of the application scenario, and relied on the most outstanding psychological studies about attention and eye movements which support that guidance is not based directly on the information provided by early visual processes but on a contextual representation that arose from them. The approach defines the task of context-driven visual attention as a mixture of latent sub-tasks, which are, in turn, modeled as a combination of specific distributions associated to low-, mid-, and high-level spatio-temporal features. Learning from fixations gathered from human observers, we incorporate an intermediate level between feature extraction and visual attention estimation that enables to obtain comprehensively guiding representations. The experiments show how our proposal successfully learns particularly adapted hierarchical explanations of visual attention in diverse video genres, outperforming several leading models in the literature. Miguel Angel Fernandez-Torres, Iván González-Díaz 0001, Fernando Díaz-de-María |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | A multi-threshold approach and a realistic error measure for vanishing point detection in natural landscapes
Álvaro García-Faura, Fernando Fernández Martínez, Ricardo Kleinlein, Rubén San-Segundo-Hernández, Fernando Díaz-de-María |
Eng. Appl. Artif. Intell. | 5 |
| 2019 | Finding landmarks within settled areas using hierarchical density-based clustering and meta-data from publicly available images
Eduardo Pla-Sacristán, Iván González-Díaz 0001, Tomás Martínez-Cortés, Fernando Díaz-de-María |
Expert Syst. Appl. | 4 |
| 2019 | Efficient Scale-Adaptive License Plate Detection SystemabstractLicense plate detection is a common problem in traffic surveillance applications. Although some solutions have been proposed in the literature, their success is usually restricted to very specific scenarios, with their performance dropping in more demanding conditions. One of the main challenges to be addressed for this kind of systems is the varying scale of the license plates, which depends on the distance between the vehicles and the camera. Traditionally, systems have handled this issue by sequentially running single-scale detectors over a pyramid of images. This approach, although simplifies the training process, requires as many evaluations as considered scales, which leads to running times that grow linearly with the number of scales considered. In this paper, we propose a scale-adaptive deformable part-based model which, based on a well-known boosting algorithm, automatically models scale during the training phase by selecting the most prominent features at each scale and notably reduces the test detection time by avoiding the evaluation at different scales. In addition, our method incorporates an empirically constrained-deformation model that adapts to different levels of deformation shown by distinct local features within license plates. As shown in the experimental section, the proposed detector is robust and scale and perspective independent and can work in quite diverse scenarios. Experiments on two datasets show that the proposed method achieves a significantly better performance in comparison with other methods of the state of the art. Miguel Molina-Moreno, Iván González-Díaz 0001, Fernando Díaz-de-María |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Emotion and attention: Audiovisual models for group-level skin response recognition in short moviesabstractThe electrodermal activity (EDA) is a psychophysiological indicator which can be considered a somatic marker of the emotional and attentional reaction of subjects towards stimuli. EDA measurements are not biased by the cognitive process of giving an opinion or a score to characterize the subjective perception, and group-level EDA recordings integrate the reaction of the whole audience, thus reducing the signal noise. This paper contributes to the field of affective video content analysis, extending previous novel work on the use of EDA as ground truth for prediction algorithms. Here, we label short video clips according to the audience’s emotion (high vs. low) and attention (increasing vs. decreasing), derived from EDA records. Then, we propose a set of low-level audiovisual descriptors and train binary classifiers that predict the emotion and attention with 75% and 80% accuracy, respectively. These results, along with those of previous works, reinforce the usefulness of such low-level audiovisual descriptors to model video in terms of the induced affective response. Álvaro García-Faura, Alejandro Hernández-García, Fernando Fernández Martínez, Fernando Díaz-de-María, Rubén San-Segundo-Hernández |
Web Intell. | 4 |
| 2018 | Automatic Learning of Image Representations Combining Content and MetadataabstractContent-based image representation is a very challenging task if we restrict to their visual content. However, associated metadata (such as tags or geolocation) become a valuable source of complementary information that may help to enhance the current system performance. In this paper, we propose an automatic training framework that uses both image visual contents and metadata to fine tune deep Convolutional Neural Networks (CNNs), providing better image descriptors adapted to certain locations, such as cities or regions. Specifically, we propose to estimate some weak labels by combining visual- and location-related information and incorporate them to a novel loss-function over pairs of images. Our experiments on a landmark discovery task show that this novel training procedure enhances the performance up to a 55% over well-established CNN-based models and is free from overfitting. Tomás Martínez-Cortés, Iván González-Díaz 0001, Fernando Díaz-de-María |
ICIP | 3 |
| 2018 | Exploiting visual saliency for assessing the impact of car commercials upon viewers
Fernando Fernández Martínez, René Arnulfo García-Hernández, Miguel Angel Fernandez-Torres, Iván González-Díaz 0001, Álvaro García-Faura, Fernando Díaz-de-María |
Multim. Tools Appl. | 6 |
| 2018 | Enriched dermoscopic-structure-based cad system for melanoma diagnosis
Javier López-Labraca, Miguel Angel Fernandez-Torres, Iván González-Díaz 0001, Fernando Díaz-de-María, Ángel Pizarro |
Multim. Tools Appl. | 4 |
| 2018 | Directional Transforms for Video Coding Based on Lifting on GraphsabstractIn this paper, we describe and optimize a general scheme based on lifting transforms on graphs for video coding. A graph is constructed to represent the video signal. Each pixel becomes a node in the graph and links between nodes represent similarity between them. Therefore, spatial neighbors and temporal motion-related pixels can be linked, while nonsimilar pixels (e.g., pixels across an edge) may not be. Then, a lifting-based transform, in which filtering operations are performed using linked nodes, is applied to this graph, leading to a 3D (spatio-temporal) directional transform, which can be viewed as an extension of wavelet transforms for video. The design of the proposed scheme requires four main steps: 1) graph construction; 2) graph splitting; 3) filter design; and 4) extension of the transform to different levels of decomposition. We focus on the optimization of these steps in order to obtain an effective transform for video coding. Furthermore, based on this scheme, we propose a coefficient reordering method and an entropy coder leading to a complete video encoder that achieves better coding performance than a motion-compensated temporal filtering wavelet-based encoder, and a simple encoder derived from H.264/AVC that makes use of similar tools as our proposed encoder (reference software JM15.1 configured to use one reference frame, no subpixel motion estimation, and 16 × 16 inter and 4 × 4 intra modes). Eduardo Martínez-Enríquez, Jesús Cid-Sueiro, Fernando Díaz-de-María, Antonio Ortega |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Emotion and attention: predicting electrodermal activity through video visual descriptorsabstractThis paper contributes to the field of affective video content analysis through the novel employment of electrodermal activity (EDA) measurements as ground truth for machine learning algorithms. The variation of the electrical properties of the skin, known as EDA, is a psychophysiological indicator widely used in medicine, psychology and neuroscience which can be considered a somatic marker of the emotional and attentional reaction of subjects towards stimuli. One of its main advantages is that the recorded information is not biased by the cognitive process of giving an opinion or a score to characterize the subjective perception. In this work, we predict the levels of emotion and attention, derived from EDA records, by means of a small set of low-level visual descriptors computed from the video stimuli. Linear regression experiments show that our descriptors predict significantly well the sum of emotion and attention levels, reaching a coefficient of determination R2 = 0.25. This result sets a promising path for further research on the prediction of emotion and attention from videos using EDA. Alejandro Hernández-García, Fernando Fernández Martínez, Fernando Díaz-de-María |
WI | 3 |
| 2017 | Adaptive Lagrange multiplier estimation algorithm in HEVC
José Luís González-de-Suso, Eduardo Martínez-Enríquez, Fernando Díaz-de-María |
Signal Process. Image Commun. | 3 |
| 2017 | Bayesian adaptive algorithm for fast coding unit decision in the High Efficiency Video Coding (HEVC) standardabstractThe latest High Efficiency Video Coding standard (HEVC) provides a set of new coding tools to achieve a significantly higher coding efficiency than previous standards. In this standard, the pixels are first grouped into Coding Units (CU), then Prediction Units (PU), and finally Transform Units (TU). All these coding levels are organized into a quadtree-shaped arrangement that allows highly flexible data representation; however, they involve a very high computational complexity. In this paper, we propose an effective early CU depth decision algorithm to reduce the encoder complexity. Our proposal is based on a hierarchical approach, in which a hypothesis test is designed to make a decision at every CU depth, where the algorithm either produces an early termination or decides to evaluate the subsequent depth level. Moreover, the proposed method is able to adaptively estimate the parameters that define each hypothesis test, so that it adapts its behavior to the variable contents of the video sequences. The proposed method has been extensively tested, and the experimental results show that our proposal outperforms several state-of-the-art methods, achieving a significant reduction of the computational complexity (36.5% and 38.2% average reductions in coding time for two different encoder configurations) in exchange for very slight losses in coding performance (1.7% and 0.8% average bit rate increments). Amaya Jimenez-Moreno, Eduardo Martínez-Enríquez, Fernando Díaz-de-María |
Signal Process. Image Commun. | 3 |
| 2017 | Neighborhood Matching for Image RetrievalabstractIn the last few years, large-scale image retrieval has attracted a lot of attention from the multimedia community. Usual approaches addressing this task first generate an initial ranking of the reference images using fast approximations that do not take into consideration the spatial arrangement of local features in the image (e.g., the bag-of-words paradigm). The top positions of the rankings are then re-estimated with verification methods that deal with more complex information, such as the geometric layout of the image. This verification step allows pruning of many false positives at the expense of an increase in the computational complexity, which may prevent its application to large-scale retrieval problems. This paper describes a geometric method known as neighborhood matching (NM), which revisits the keypoint matching process by considering a neighborhood around each keypoint and improves the efficiency of a geometric verification step in the image search system. Multiple strategies are proposed and compared to incorporate NM into a large-scale image retrieval framework. A detailed analysis and comparison of these strategies and baseline methods have been investigated. The experiments show that the proposed method not only improves the computational efficiency, but also increases the retrieval performance and outperforms state-of-the-art methods in standard datasets, such as the Oxford 5 k and 105 k datasets, for which the spatial verification step has a significant impact on the system performance. Iván González-Díaz 0001, Murat Birinci, Fernando Díaz-de-María, Edward J. Delp |
IEEE Trans. Multim. | 3 |
| 2016 | Comparing visual descriptors and automatic rating strategies for video aesthetics prediction
Alejandro Hernández-García, Fernando Fernández Martínez, Fernando Díaz-de-María |
Signal Process. Image Commun. | 3 |
| 2016 | Complexity Control Based on a Fast Coding Unit Decision Method in the HEVC Video Coding StandardabstractThe emerging high-efficiency video coding standard achieves higher coding efficiency than previous standards by virtue of a set of new coding tools such as the quadtree coding structure. In this novel structure, the pixels are organized into coding units (CU), prediction units, and transform units, the sizes of which can be optimized at every level following a tree configuration. These tools allow highly flexible data representation; however, they incur a very high computational complexity. In this paper, we propose an effective complexity control (CC) algorithm based on a hierarchical approach. An early termination condition is defined at every CU size to determine whether subsequent CU sizes should be explored. The actual encoding times are also considered to satisfy the target complexity in real time. Moreover, all parameters of the algorithm are estimated on the fly to adapt its behavior to the video content, the encoding configuration, and the target complexity over time. The experimental results prove that our proposal is able to achieve a target complexity reduction of up to 60% with respect to full exploration, with notable accuracy and limited losses in coding performance. It was compared with a state-of-the-art CC method and shown to achieve a significantly better trade-off between coding complexity and efficiency as well as higher accuracy in reaching the target complexity. Furthermore, a comparison with a state-of-the-art complexity reduction method highlights the advantages of our CC framework. Finally, we show that the proposed method performs well when the target complexity varies over time. Amaya Jimenez-Moreno, Eduardo Martínez-Enríquez, Fernando Díaz-de-María |
IEEE Trans. Multim. | 3 |
| 2015 | Succeeding metadata based annotation scheme and visual tips for the automatic assessment of video aesthetic quality in car commercials
Fernando Fernández Martínez, Alejandro Hernández-García, Fernando Díaz-de-María |
Expert Syst. Appl. | 3 |
| 2015 | Temporal segmentation and keyframe selection methods for user-generated video search-based annotation
Iván González-Díaz 0001, Tomás Martínez-Cortés, Ascensión Gallardo-Antolín, Fernando Díaz-de-María |
Expert Syst. Appl. | 4 |
| 2015 | Two-level sliding-window VBR control algorithm for video on demand streamingabstractA two-level variable bit rate (VBR) control algorithm for hierarchical video coding, specifically tailored for the new High Efficiency Video Coding (HEVC) standard, is presented here. A long-term level monitors the current bit count along a sliding window of a few seconds, comprising several intra-periods (IPs) and shifted on an IP basis. This long-term view allows the accommodation of the naturally occurring rate variations at a slow pace, avoiding the annoying sharp quality changes commonly appearing when non-sliding window approaches are used. The bit excesses or defects observed at this level are evenly delivered to a short-term level mechanism that establishes target bit budgets for a narrower sliding window covering a single IP and shifting on a frame basis. At this level, an adequate quantization parameter is estimated to comply with the designated target bit rate. Recommended test conditions as well as two few minutes long video sequences with scene cuts have been used for the assessment of the proposed VBR controller. Comparisons with a state-of-the-art rate control algorithm have produced good results in terms of quality consistency, in exchange for moderate rate-distortion performance losses. Manuel de-Frutos-López, José Luís González-de-Suso, Sergio Sanz Rodríguez, Carmen Peláez-Moreno, Fernando Díaz-de-María |
Signal Process. Image Commun. | 5 |
| 2014 | A Bayesian model for brain tumor classification using clinical-based featuresabstractThis paper tackles the problem of automatic brain tumor classification from Magnetic Resonance Imaging (MRI) where, traditionally, general-purpose texture and shape features extracted from the Region of Interest (tumor) have become the usual parameterization of the problem. Two main contributions are made in this context. First, a novel set of clinical-based features that intend to model intuitions and expert knowledge of physicians is suggested. Second, a system is proposed that is able to fuse multiple individual scores (based on a particular MRI sequence and a pathological indicator present in that sequence) by using a Bayesian model that produces a global system decision. This approximation provides a quite flexible solution able to handle missing data, which becomes a very likely case in a realistic scenario where the number clinical tests varies from one patient to another. Furthermore, the Bayesian model provides extra information concerning the uncertainty of the final decision. Our experimental results prove that the use of clinical-based feature leads to a significant increment of performance in terms of Area Under the Curve (AUC) when compared to a state-of-the art reference. Furthermore, the proposed Bayesian fusion model clearly outperforms other fusion schemes, especially when few diagnostic tests are available. Tomás Martínez-Cortés, Miguel Angel Fernandez-Torres, Amaya Jimenez-Moreno, Iván González-Díaz 0001, Fernando Díaz-de-María, Juan Adan Guzmán-De-Villoria, Pilar Fernández |
ICIP | 5 |
| 2014 | Improved Method to Select the Lagrange Multiplier for Rate-Distortion Based Motion Estimation in Video CodingabstractThe motion estimation (ME) process used in the H.264/AVC reference software is based on minimizing a cost function that involves two terms (distortion and rate) that are properly balanced through a Lagrangian parameter, usually denoted as λmotion. In this paper we propose an algorithm to improve the conventional way of estimating λmotion and, consequently, the ME process. First, we show that the conventional estimation of λmotion turns out to be significantly less accurate when ME-compromising events, which make the ME process to perform poorly, happen. Second, with the aim of improving the coding efficiency in these cases, an efficient algorithm is proposed that allows the encoder to choose between three different values of λmotion for the Inter 16x16 partition size. To be more precise, for this partition size, the proposed algorithm allows the encoder to additionally test λmotion=0 and λmotion arbitrarily large, which corresponds to minimum distortion and minimum rate solutions, respectively. By testing these two extreme values, the algorithm avoids making large ME errors. The experimental results on video segments exhibiting this type of ME-compromising events reveal an average rate reduction of 2.20% for the same coding quality with respect to the JM15.1 reference software of H.264/AVC. The algorithm has been also tested in comparison with a state-of-the-art algorithm called context adaptive Lagrange multiplier. Additionally, two illustrative examples of the subjective performance improvement are provided. José Luís González-de-Suso, Amaya Jimenez-Moreno, Eduardo Martínez-Enríquez, Fernando Díaz-de-María |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | A Generative Model for Concurrent Image Retrieval and ROI SegmentationabstractThis paper proposes a probabilistic generative model that concurrently tackles the problems of image retrieval and region-of-interest (ROI) segmentation. Specifically, the proposed model takes into account several properties of the matching process between two objects in different images, namely: objects undergoing a geometric transformation, typical spatial location of the region of interest, and visual similarity. In this manner, our approach improves the reliability of detected true matches between any pair of images. Furthermore, by taking advantage of the links to the ROI provided by the true matches, the proposed method is able to perform a suitable ROI segmentation. Finally, the proposed method is able to work when there is more than one ROI in the query image. Our experiments on two challenging image retrieval datasets proved that our approach clearly outperforms the most prevalent approach for geometrically constrained matching and compares favorably to most of the state-of-the-art methods. Furthermore, the proposed technique concurrently provided very good segmentations of the ROI. Furthermore, the capability of the proposed method to take into account several objects-of-interest was also tested on three experiments: two of them concerning image segmentation and object detection in multi-object image retrieval tasks, and another concerning multiview image retrieval. These experiments proved the ability of our approach to handle scenarios in which more than one object of interest is present in the query. Iván González-Díaz 0001, Carlos E. Baz-Hormigos, Fernando Díaz-de-María |
IEEE Trans. Multim. | 3 |
| 2014 | Standard-Compliant Low-Pass Temporal Filter to Reduce the Perceived Flicker ArtifactabstractFlicker is a common video-compression-related temporal artifact. It occurs when co-located regions of consecutive frames are not encoded in a consistent manner, especially when Intra frames are periodically inserted at low and medium bit rates. In this paper we propose a flicker reduction method which aims to make the luminance changes between pixels in the same area of consecutive frames less noticeable. To this end, a temporal low-pass filtering is proposed that smooths these luminance changes on a block-by-block basis. The proposed method has some advantages compared to another state-of-the-art methods. It has been designed to be compliant with conventional video coding standards, i.e., to generate a bitstream that is decodable by any standard decoder implementation. The filter strength is estimated on-the-fly to limit the PSNR loss and thus the appearance of a noticeable blurring effect. The proposed method has been implemented on the H.264/AVC reference software and thoroughly assessed in comparison to a couple of state-of-the-art methods. The flicker reduction achieved by the proposed method (calculated using an objective measurement) is notably higher than that of compared methods: 18.78% versus 5.32% and 31.96% versus 8.34%, in exchange of some slight losses in terms of coding efficiency. In terms of subjective quality, the proposed method is perceived more than two times better than the compared methods. Amaya Jimenez-Moreno, Eduardo Martínez-Enríquez, Fernando Díaz-de-María |
IEEE Trans. Multim. | 4 |
| 2013 | Standard compliant flicker reduction method with PSNR loss controlabstractFlicker is a common video coding artifact that occurs especially at low and medium bit rates. In this paper we propose a temporal filter-based method to reduce flicker. The proposed method has been designed to be compliant with conventional video coding standards, i.e., to generate a bitstream that is decodable by any standard decoder implementation. The aim of the proposed method is to make the luminance changes between consecutive frames smoother on a block-by-block basis. To this end, a selective temporal low-pass filtering is proposed that smooths these luminance changes on flicker-prone blocks. Furthermore, since the low-pass filtering can incur in a noticeable blurring effect, an adaptive algorithm that allows for limiting the PSNR loss -and thus the blur- has also been designed. The proposed method has been extensively assessed on the reference software of the H.264/AVC video coding standard and compared to a state-of-the-art method. The experimental results show the effectiveness of the proposed method and prove that its performance is superior to that of the state-of-the-art method. Amaya Jimenez-Moreno, Eduardo Martínez-Enríquez, Fernando Díaz-de-María |
ICASSP | 3 |
| 2013 | Filter optimization and complexity reduction for video coding using graph-based transformsabstractThe basis functions of lifting transform on graphs are completely determined by finding a bipartition of the graph and defining the prediction and update filters to be used. In this work we consider the design of prediction filters that minimize the quadratic prediction error and therefore the energy of the detail coefficients, which will give rise to higher energy compaction. Then, to determine the graph bipartition, we propose a distributed maximum-cut algorithm that significantly reduces the computational cost with respect to the centralized version used in our previous work. The proposed techniques show improvements in coding performance and computational cost as compared to our previous work. Eduardo Martínez-Enríquez, Fernando Díaz-de-María, Jesús Cid-Sueiro, Antonio Ortega |
ICIP | 2 |
| 2013 | Mid-level feature set for specific event and anomaly detection in crowded scenesabstractIn this paper we propose a system for automatic detection of specific events and abnormal behaviors in crowded scenes. In particular, we focus on the parametrization by proposing a set of mid-level spatio-temporal features that successfully model the characteristic motion of typical events in crowd behaviors. Furthermore, due to the fact that some features are more suitable than others to model specific events of interest, we also present an automatic process for feature selection. Our experiments prove that the suggested feature set works successfully for both explicit event detection and distance-based anomaly detection tasks. The results on PETS for explicit event detection are generally better than those previously reported. Regarding anomaly detection, the proposed method performance is comparable to those of state-of-the-art method for PETS and substantially better than that reported for Web dataset. Fernando de-la-Calle-Silos, Iván González-Díaz 0001, Fernando Díaz-de-María |
ICIP | 3 |
| 2013 | A region-centered topic model for object discovery and category-based image segmentation
Iván González-Díaz 0001, Fernando Díaz-de-María |
Pattern Recognit. | 2 |
| 2013 | Mode Decision-Based Algorithm for Complexity Control in H.264/AVCabstractThe latest H.264/AVC video coding standard achieves high compression rates in exchange for high computational complexity. Nowadays, however, many application scenarios require the encoder to meet some complexity constraints. This paper proposes a novel complexity control method that relies on a hypothesis testing that can handle time-variant content and target complexities. Specifically, it is based on a binary hypothesis testing that decides, on a macroblock basis, whether to use a low- or a high-complexity coding model. Gaussian statistics are assumed so that the probability density functions involved in the hypothesis testing can be easily adapted. The decision threshold is also adapted according to the deviation between the actual and the target complexities. The proposed method is implemented on the H.264/AVC reference software JM10.2 and compared with a state-of-the-art method. Our experimental results prove that the proposed method achieves a better trade-off between complexity control and coding efficiency. Furthermore, it leads to a lower deviation from the target complexity. Amaya Jimenez-Moreno, Eduardo Martínez-Enríquez, Fernando Díaz-de-María |
IEEE Trans. Multim. | 3 |
| 2012 | Perceptually-aware bilateral filtering for quality improvement in low bit rate video codingabstractPerceptual coding has become of great interest in modern video coding due to the need for higher compression rates. Many previous works have been carried out to incorporate perceptual information to hybrid video encoders, either modifying the quantization parameter according to a certain perceptual resource allocation map or preprocessing video sequences for removing information that is not perceptually relevant. The first strategy is limited by the presence of blocking artifacts and the second one lacks of adaptation to video content. In this paper, a novel and simple approach is proposed, which performs a smart filtering prior to the encoding process preserving both the structural and motion information. The experiments prove that the use of proposed method implemented on an H.264 encoder significantly improves its perceptual quality for low bit rates. Manuel de-Frutos-López, Helen Medina-Chanca, Sergio Sanz Rodríguez, Carmen Peláez-Moreno, Fernando Díaz-de-María |
PCS | 5 |
| 2012 | A simplified subjective video quality assessment method based on signal detection theoryabstractA simplified protocol and associated metrics based on Signal Detection Theory (SDT) for subjective Video Quality Assessment (VQA) is proposed with the aim of filling the gap existing between the lack of discrimination abilities of objective Quality Estimates (specially when perceptually motivated processing methods are involved) and the costly normative subjective quality tests. The proposed protocol employs a reduced number of assessors and provides a quality ranking of the methods being evaluated. It is intended for providing the rapid experimental turn around necessary for developing algorithms. We have validated our proposal by corroborating with our test a well-known result for the video coding community: the quality benefits of including an in-loop deblocking filter. A software interface to design and administrate the test is also made publicly available. Manuel de-Frutos-López, Ana Belén Mejía-Ocaña, Sergio Sanz Rodríguez, Carmen Peláez-Moreno, Fernando Díaz-de-María, Zygmunt Pizlo |
PCS | 5 |
| 2012 | Real-Time Robust Automatic Speech Recognition Using Compact Support Vector MachinesabstractIn the last years, support vector machines (SVMs) have shown excellent performance in many applications, especially in the presence of noise. In particular, SVMs offer several advantages over artificial neural networks (ANNs) that have attracted the attention of the speech processing community. Nevertheless, their high computational requirements prevent them from being used in practice in automatic speech recognition (ASR), where ANNs have proven to be successful. The high complexity of SVMs in this context arises from the use of huge speech training databases with millions of samples and highly overlapped classes. This paper suggests the use of a weighted least squares (WLS) training procedure that facilitates the possibility of imposing a compact semiparametric model on the SVM, which results in a dramatic complexity reduction. Such a complexity reduction with respect to conventional SVMs, which is between two and three orders of magnitude, allows the proposed hybrid WLS-SVC/HMM system to perform real-time speech decoding on a connected-digit recognition task (SpeechDat Spanish database). The experimental evaluation of the proposed system shows encouraging performance levels in clean and noisy conditions, although further improvements are required to reach the maturity level of current context-dependent HMM-based recognizers. Rubén Solera-Ureña, Ana I. García-Moral, Carmen Peláez-Moreno, Manel Martínez-Ramón, Fernando Díaz-de-María |
IEEE Trans. Speech Audio Process. | 5 |
| 2012 | In-Layer Multibuffer Framework for Rate-Controlled Scalable Video CodingabstractTemporal scalability is supported in scalable video coding (SVC) by means of hierarchical prediction structures, where the higher layers can be ignored for frame rate reduction. Nevertheless, this kind of scalability is not totally exploited by the rate control (RC) algorithms since the hypothetical reference decoder (HRD) requirement is only satisfied for the highest frame rate substream of every dependence (spatial or coarse grain scalability) layer. In this paper, we propose a novel RC approach that aims to deliver several HRD-compliant temporal resolutions within a particular dependence layer. Instead of using the common SVC encoder configuration consisting of a dependence layer per each temporal resolution, a compact configuration that does not require additional dependence layers for providing different HRD-compliant temporal resolutions is proposed. Specifically, the proposed framework for rate-controlled SVC uses a set of virtual buffers within a dependence layer so that their levels can be simultaneously controlled for overflow and underflow prevention while minimizing the reconstructed video distortion of the corresponding substreams. This in-layer multibuffer approach has been built on the top of a baseline H.264/SVC RC algorithm for variable bit rate applications. The experimental results show that our proposal achieves a good performance in terms of mean quality, quality consistency, and buffer control using a reduced number of layers. Sergio Sanz Rodríguez, Fernando Díaz-de-María |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Video encoder based on lifting transforms on graphsabstractWe propose a complete video encoder based on directional "non- separable" transforms that allow spatial and temporal correlation to be jointly exploited. These lifting-based wavelet transforms are applied on graphs that link pixels in a video sequence based on motion information. In this paper, we first consider a low complexity version of this transform, which can operate on subgraphs without significant loss in performance. We then study coefficient reordering techniques that lead to a more realistic and efficient encoder that the one we presented in our earlier work. Our proposed technique shows encouraging results as compared to a comparable scheme based on the DCT transform. Eduardo Martínez-Enríquez, Fernando Díaz-de-María, Antonio Ortega |
ICIP | 2 |
| 2011 | Rate control initialization algorithm for scalable video codingabstractIn this paper we propose a novel rate control initialization algorithm for real-time H.264/scalable video coding. In particular, a two-step approach is proposed. First, the initial quantization parameter (QP) for each layer is determined by means of a parametric rate-quantization (R-Q) modeling that depends on the layer identifier (base or enhancement) and on the type of scalability (spatial or quality). Second, an intra-frame QP refinement method that allows for adapting the initial QP value when needed is carried out over the three first coded frames in order to take into consideration both the buffer control and the spatio-temporal complexity of the scene. The experimental results show that the proposed R-Q modeling for initial QP estimation, in combination with the intra-frame QP refinement method, provide a good performance in terms of visual quality and buffer control, achieving remarkably similar results to those achieved by using ideal initial QP values. Sergio Sanz Rodríguez, Fernando Díaz-de-María |
ICIP | 2 |
| 2011 | State-space dynamics distance for clustering sequential data
Dario García-García, Emilio Parrado-Hernández, Fernando Díaz-de-María |
Pattern Recognit. | 3 |
| 2011 | Data Balancing for Efficient Training of Hybrid ANN/HMM Automatic Speech Recognition SystemsabstractHybrid speech recognizers, where the estimation of the emission pdf of the states of hidden Markov models (HMMs), usually carried out using Gaussian mixture models (GMMs), is substituted by artificial neural networks (ANNs) have several advantages over the classical systems. However, to obtain performance improvements, the computational requirements are heavily increased because of the need to train the ANN. Departing from the observation of the remarkable skewness of speech data, this paper proposes sifting out the training set and balancing the amount of samples per class. With this method, the training time has been reduced 18 times while obtaining performances similar to or even better than those with the whole database, especially in noisy environments. However, the application of these reduced sets is not straightforward. To avoid the mismatch between training and testing conditions created by the modification of the distribution of the training data, a proper scaling of the a posteriori probabilities obtained and a resizing of the context window need to be performed as demonstrated in this paper. Ana I. García-Moral, Rubén Solera-Ureña, Carmen Peláez-Moreno, Fernando Díaz-de-María |
IEEE Trans. Speech Audio Process. | 4 |
| 2011 | A Two-Level Classification-Based Approach to Inter Mode Decision in H.264/AVCabstractThe H.264/AVC standard achieves a high coding efficiency compared to previous standards. However, this gain is accomplished at great computational cost, with mode decision being one of the most demanding subsystems. In this paper, a two-level classification-based approach to the inter mode decision problem is proposed. A first classifier detects SKIP/Direct modes, while a second one is able to decide whether to use a large (16 × 16, 16 × 8, and 8 × 16) or a small mode (8 × 8, 8 × 4, 4 × 8, and 4 × 4). The suggested classifiers are binary and linear, and the input features in the classifiers have been carefully selected. A novel cost function that pays more attention to the most critical samples during the classifier training process has been designed. The experimental results show an average computational savings of 60% of the total encoding time with respect to JM10.2 over a comprehensive variety of sequences and formats. This is achieved with negligible degradation in rate-distortion performance and compares favorably with state-of-the-art fast mode decision methods. Furthermore, the proposed method has been successfully assessed at different levels of complexity reduction. Eduardo Martínez-Enríquez, Amaya Jimenez-Moreno, Miguel Angel-Pellon, Fernando Díaz-de-María |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2011 | RBF-Based QP Estimation Model for VBR Control in H.264/SVCabstractIn this paper, we propose a novel variable bit rate (VBR) controller for real-time H.264/scalable video coding (SVC) applications. The proposed VBR controller relies on the fact that consecutive pictures within the same scene often exhibit similar degrees of complexity, and consequently should be encoded using similar quantization parameter (QP) values for the sake of quality consistency. In order to prevent unnecessary QP fluctuations, the proposed VBR controller allows for just an incremental variation of QP with respect to that of the previous picture, focusing on the design of an effective method for estimating this QP variation. The implementation in H.264/SVC requires to locate a rate controller at each dependency layer (spatial or coarse grain scalability). In particular, the QP increment estimation at each layer is computed by means of a radial basis function (RBF) network that is specially designed for this purpose. Furthermore, the RBF network design process was conceived to provide an effective solution for a wide range of practical real-time VBR applications for scalable video content delivery. In order to assess the proposed VBR controller, two real-time application scenarios were simulated: mobile live streaming and Internet protocol television broadcast. It was compared to constant QP encoding and a recently proposed constant bit rate (CBR) controller for H.264/SVC. The experimental results show that the proposed method achieves remarkably consistent quality, outperforming the reference CBR controller in the two scenarios for all the spatio-temporal resolutions considered. Sergio Sanz Rodríguez, Fernando Díaz-de-María |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | A two-level sliding-window VBR controller for real-time hierarchical video codingabstractIn this paper, a novel rate control algorithm for real-time VBR hierarchical video coding is proposed. The algorithm works at two levels that are called long- and short-term levels. The long-term level aims at ensuring that the bit count does not exceed the maximum allowed amount for a few-second long window. To this end, it considers a sliding window spanning several GOPs, which is shifted on a GOP basis. In doing so, it avoids the potentially sharp adjustments at the end of the GOP that usually happen in non-sliding approaches. The short-term level aims to provide a proper QP adaptation to fit the target bit budget, which is dictated by the long-term level. It also uses a sliding window, which in this case extends over one GOP. The proposed algorithm has been assessed in realistic conditions for a variety of video sequences. It has been compared to both a constant quality and CBR hierarchical approaches, showing an excellent performance in terms of both rate-distortion and PSNR variation. Manuel de-Frutos-López, Oscar del-Ama-Esteban, Sergio Sanz Rodríguez, Fernando Díaz-de-María |
ICIP | 4 |
| 2010 | RBF-based VBR controller for real-time H.264/SVC video codingabstractIn this paper we propose a novel VBR controller for real-time H.264/SVC video coding. Since consecutive pictures within the same scene often exhibit similar degrees of complexity, the proposed VBR controller allows for just an incremental variation of QP with respect to that of the previous picture, so preventing unnecessary QP fluctuations. For this purpose, an RBF network has been carefully designed to estimate the QP increment at each dependency (spatial or CGS) layer. A mobile live streaming application scenario was simulated to assess the performance of the proposed VBR controller, which was compared to a recently proposed CBR controller for H.264/SVC. The experimental results show a remarkably consistent quality, notably outperforming the reference CBR controller. Sergio Sanz Rodríguez, Fernando Díaz-de-María |
PCS | 2 |
| 2010 | Uncertainty decoding on Frequency Filtered parameters for robust ASR
Jesús Vicente-Peña, Fernando Díaz-de-María |
Speech Commun. | 2 |
| 2010 | The synergy between bounded-distance HMM and spectral subtraction for robust speech recognition
Jesús Vicente-Peña, Fernando Díaz-de-María, W. Bastiaan Kleijn |
Speech Commun. | 2 |
| 2010 | An improved fast mode decision algorithm for intraprediction in H.264/AVC video coding
Manuel de-Frutos-López, Daniel Orellana-Quirós, Jose Carlos Pujol-Alcolado, Fernando Díaz-de-María |
Signal Process. Image Commun. | 4 |
| 2010 | Cauchy-Density-Based Basic Unit Layer Rate Controller for H.264/AVCabstractThe rate control problem has been extensively studied in parallel to the development of the different video coding standards. The bit allocation via Cauchy-density-based rate-distortion modeling of the discrete cosine transform coefficients has proved to be one of the most accurate solutions at picture level. Nevertheless, in some specific applications operating in real-time low-delay environments, a basic unit (BU) layer is recommended in order to provide a good tradeoff between picture quality and delay control. In this letter, a novel BU bit allocation for H.264/advanced video coding is proposed based on a simplified Cauchy probability density function source modeling. The experimental results are twofold: 1) the proposed rate control algorithm (RCA) achieves an average peak signal-to-noise ratio improvement of 0.28 dB respect to a well-known BU layer RCA, while maintaining a similar buffer occupancy evolution, and 2) it achieves to notably reduce the buffer occupancy fluctuations respect to a well-known picture layer RCA, while maintaining similar quality levels. Sergio Sanz Rodríguez, Oscar del-Ama-Esteban, Manuel de-Frutos-López, Fernando Díaz-de-María |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | A spatially aware generative model for image classification, topic discovery and segmentationabstractFor the last few years bag-of-words models have been succesfully applied to the information retrieval field. However their application to visual content suffers from an important shortcoming: they model images as sets of unordered visual words rather than consider their spatial and geometric layout. Visual information is highly organized along the dimensions of an image and algorithms should make use of this to enhance the performance of several visual processing tasks. In this paper, a generative model is proposed that fuses both the local information obtained from visual words and the global geometric layout given by a previous segmentation of the image. Furthermore, the model considers inter-region influences so topics can spread along the image and, thus, generate final segmentations in which regions represent semantic concepts. The proposed model is succesfully tested on three different tasks. Iván González-Díaz 0001, Dario García-García, Fernando Díaz-de-María |
ICIP | 3 |
| 2009 | A hierarchical classification-based approach to Inter Mode Decision in H.264/AVCabstractThe H.264/AVC standard achieves a high coding efficiency compared with previous standards. However, it does so at a very high computational cost, with motion estimation being one of the most demanding subsystems. In this paper a hierarchical classificationbased approach to the Inter Mode Decision (MD) problem is proposed. A first classifier detects SKIP/Direct modes while a second one is able to decide whether using a large (16×16, 16×8 and 8×16) or a small mode (8×8, 8×4, 4×8 and 4×4). The same procedure is applied at SubMacroblock level. The suggested classifiers are binary and linear. The input features that feed both classifiers have been carefully selected. A novel cost function that pays more attention to the most critical samples during the classifier training process has been designed. The results are very promising: a 64 % computational saving of the total encoding time with respect to JM10.2 is achieved with negligible degradation in the rate-distortion performance. Eduardo Martínez-Enríquez, Fernando Díaz-de-María |
ICME | 2 |
| 2009 | Sequence Segmentation via Clustering of SubsequencesabstractWe propose a new algorithm for sequence segmentation based on recent advances in semi-parametric sequence clustering. This approach implies the use of model-based distance measures between sequences, as well as a variant of spectral clustering specially tailored for segmentation. The method is highly flexible since it allows for the use of any probabilistic generative model for the individual segments. The performance of the proposed algorithm is demonstrated using both a synthetic dataset and a speaker segmentation task. Dario García-García, Emilio Parrado-Hernández, Fernando Díaz-de-María |
ICMLA | 3 |
| 2009 | A novel fast inter mode decision in H.264/AVC based on a regionalized hypothesis testingabstractThe H.264/AVC standard achieves a high coding efficiency compared to previous standards, but in exchange for a very high computational cost. This paper focuses on the mode decision subsystem, which is a critical one from the computational cost point of view. In particular, a hierarchical early-termination-based approach to the inter mode decision problem is proposed, in such a way that a cascade of early stops are considered, going from the most likely modes to the less likely ones. The main novelty of the proposed algorithm relies on the use of a regionalized hypothesis testing methodology for making the decisions on every potential early stop. The proposed method has been assessed on a large set of QCIF video sequences for several qualities, achieving a quite notable improvement in terms of total encoding time reduction in relation to a well-established mode decision method. Specifically, the time reduction is 43.43 % and 21.33 % for IPPP and IBBP patterns, respectively, without any significant quality loss. Eduardo Martínez-Enríquez, Amaya Jimenez-Moreno, Fernando Díaz-de-María |
PCS | 3 |
| 2009 | Low-complexity VBR controller for spatial-CGS and temporal scalable video codingabstractThis paper presents a rate control (RC) algorithm for the scalable extension of the H.264/AVC video coding standard. The proposed rate controller is designed for real-time video streaming with buffer constraint. Since a large buffer delay and bit rate variation are allowed in this kind of applications, our proposal reduces the quantization parameter (QP) fluctuation to provide consistent visual quality bit streams to receivers with a variety of spatio-temporal resolutions and processing capabilities. The low computational cost is another characteristic of the described method, since a simple lookup table is used to regulate the QP variation on a frame basis. Sergio Sanz Rodríguez, Fernando Díaz-de-María, Mehdi Rezaei |
PCS | 2 |
| 2009 | A Cauchy-density-based rate controller for H.264/AVC in low-delay environmentsabstractThe accuracy of the Cauchy probability density function for modeling of the discrete cosine transform coefficient distribution has already been proved for the frame layer of the rate control subsystem of a hybrid video coder. Nevertheless, in some specific applications operating in real-time low-delay environments, a basic unit layer is recommended in order to provide a good trade-off between quality and delay control. In this paper, a novel basic unit bit allocation for H.264/AVC is proposed based on a simplified Cauchy probability density function source modeling. The experimental results show that the proposed algorithm improves the average peak signal-to-noise ratio in 0.28 and 0.35 dB with respect to two well-known rate control schemes, while maintaining similar peak signal-to-noise ratio standard deviation and buffer occupancy evolution. Oscar del-Ama-Esteban, Sergio Sanz Rodríguez, Manuel de-Frutos-López, Fernando Díaz-de-María |
PCS | 4 |
| 2009 | A New Distance Measure for Model-Based Sequence ClusteringabstractWe review the existing alternatives for defining model-based distances for clustering sequences and propose a new one based on the Kullback-Leibler divergence. This distance is shown to be especially useful in combination with spectral clustering. For improved performance in real-world scenarios, a model selection scheme is also proposed. Dario García-García, Emilio Parrado-Hernández, Fernando Díaz-de-María |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Characterization of Healthy and Pathological Voice Through Measures Based on Nonlinear DynamicsabstractIn this paper, we propose to quantify the quality of the recorded voice through objective nonlinear measures. Quantification of speech signal quality has been traditionally carried out with linear techniques since the classical model of voice production is a linear approximation. Nevertheless, nonlinear behaviors in the voice production process have been shown. This paper studies the usefulness of six nonlinear chaotic measures based on nonlinear dynamics theory in the discrimination between two levels of voice quality: healthy and pathological. The studied measures are first- and second-order Renyi entropies, the correlation entropy and the correlation dimension. These measures were obtained from the speech signal in the phase-space domain. The values of the first minimum of mutual information function and Shannon entropy were also studied. Two databases were used to assess the usefulness of the measures: a multiquality database composed of four levels of voice quality (healthy voice and three levels of pathological voice); and a commercial database (MEEI Voice Disorders) composed of two levels of voice quality (healthy and pathological voices). A classifier based on standard neural networks was implemented in order to evaluate the measures proposed. Global success rates of 82.47% (multiquality database) and 99.69% (commercial database) were obtained. Patricia Henríquez Rodríguez, Jesús B. Alonso, Miguel A. Ferrer, Carlos Manuel Travieso-González, Juan Ignacio Godino-Llorente, Fernando Díaz-de-María |
IEEE Trans. Speech Audio Process. | 6 |
| 2008 | Incorporating spatio-temporal mid-level features in a region segmentation algorithm for video sequencesabstractSegmentation algorithms traditionally employ low-level features to divide images into different regions that show a certain degree of homogeneity. However, low-level features, spatial or temporal, are not always reliable when processing real-world video sequences, because of issues like illuminations or complex backgrounds. Furthermore, real world objects can be composed of different regions with heterogeneous features. Although the inclusion of motion can mitigate some of these effects, many problems are still present. This paper proposes the utilization of some spatio-temporal mid-level features that are related, on the one hand, to geometric properties of real objects and, on the other, to well-known motion patterns. Specifically, the proposed algorithm uses a mid-level module that controls the subsequent segmentation using these kinds of features. Some experiments and evaluations show that the inclusion of mid-level features can help to obtain perceptually more meaningful segmentations, thus resulting in regions that are closer to semantic concepts. Iván González-Díaz 0001, Kevin McGuinness, Tomasz Adamek, Noel E. O'Connor, Fernando Díaz-de-María |
ICIP | 5 |
| 2008 | Adaptive Multipattern Fast Block-Matching Algorithm Based on Motion Classification TechniquesabstractIn most video coding standards, motion estimation becomes the most time-consuming subsystem. Consequently, in the last few years, a great deal of effort has been devoted to the research of novel algorithms capable of saving computations with minimal effects on the coding quality. Adaptive algorithms and particularly multipattern solutions, have evolved as the most robust general-purpose solutions owing to two main reasons: 1) real video sequences usually exhibit a wide-range of motion content, from uniform to random, and 2) a vast amount of coding applications have appeared demanding different degrees of coding quality. In this study, we propose an adaptive algorithm, called motion classification-based search (MCS), which makes use of an especially tailored classifier that detects some motion cues and chooses the search pattern that best fits them. The MCS has been experimentally assessed for a comprehensive set of selected video sequences and qualities. Our experimental results show that MCS notably reduces the computational cost up to 55% and 84% in search points, with respect to two state-of-the-art methods, while maintaining the quality. Iván González-Díaz 0001, Fernando Díaz-de-María |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Improved Motion Classification Techniques for Adaptive Multi-Pattern Fast Block-Matching AlgorithmabstractIn several video coding standards, such as H.264, motion estimation becomes the most-time consuming subsystem. Therefore, recently research on video coding focuses on the development of novel algorithms able to save computations with minimal effects over the coding distortion. Due to the fact that real video sequences usually exhibit a wide-range of motion content, from uniform to random, and to the vast amount of coding applications demanding different degrees of coding quality, adaptive algorithms have revealed as the most robust general purpose solutions. In particular, multi-pattern algorithms can adapt to video contents as well as to required coding quality by means of the use of a set of heterogeneus search patterns, each one adapting better to particular motion and quality requirements. This paper applies some improvements to the Motion Classification based Search, an adaptive multi-pattern algorithm based on motion classification techniques. Our experimental results show that MCS notably reduces the computational cost with respect to some well-known algorithms while maintaining the quality. Iván González-Díaz 0001, Fernando Díaz-de-María |
ICIP (2) | 2 |
| 2007 | A Fast Motion-Cost Based Algorithm for H.264/AVC Inter Mode DecisionabstractThe H.264/AVC standard achieves a high coding efficiency compared to previous standards. However, the encoder complexity results in very high computational cost due to motion estimation and macroblock mode decisions. In this paper we propose a fast mode decision for low computational complexity applications for which the rate distortion optimization mode decision becomes unacceptable. The proposed pruned mode decision method consists in a motion-cost based early termination algorithm and saves about 50% encoding time with negligible quality loss. Eduardo Martínez-Enríquez, Manuel de-Frutos-López, Jose Carlos Pujol-Alcolado, Fernando Díaz-de-María |
ICIP (5) | 4 |
| 2007 | Robust ASR using Support Vector Machines
Rubén Solera-Ureña, Darío Martín-Iglesias, Ascensión Gallardo-Antolín, Carmen Peláez-Moreno, Fernando Díaz-de-María |
Speech Commun. | 5 |
| 2006 | Individual on-line variance adaptation of frequency filtered parameters for robust ASRabstractIn this paper we address the problem of robust speech recognition. We propose a new method based on the individual variance adaptation of frequency filtered parameters to reduce the deleterious effects of additive narrow-band noise. The method can be interpreted as a spectral weighting that assigns increased importance to the most reliable spectral components, typically the spectral peaks. The experiments confirm that the suggested method results in significantly improved recognition rates for additive narrow-band noise. Jesús Vicente-Peña, Fernando Díaz-de-María, W. Bastiaan Kleijn |
INTERSPEECH | 2 |
| 2006 | A comparison of front-ends for bitstream-based ASR over IP
Carmen Peláez-Moreno, Ascensión Gallardo-Antolín, Diego F. Gómez-Cajas, Fernando Díaz-de-María |
Signal Process. | 4 |
| 2006 | Band-pass filtering of the time sequences of spectral parameters for robust wireless speech recognition
Jesús Vicente-Peña, Ascensión Gallardo-Antolín, Carmen Peláez-Moreno, Fernando Díaz-de-María |
Speech Commun. | 4 |
| 2005 | Design of a voice-enabled interface for real-time access to stock exchange from a PDA through GPRS
Darío Martín-Iglesias, Yago Pereiro-Estevan, Ana I. García-Moral, Ascensión Gallardo-Antolín, Fernando Díaz-de-María |
INTERSPEECH | 5 |
| 2005 | Recognizing GSM digital speechabstractThe Global System for Mobile (GSM) environment encompasses three main problems for automatic speech recognition (ASR) systems: noisy scenarios, source coding distortion, and transmission errors. The first one has already received much attention; however, source coding distortion and transmission errors must be explicitly addressed. In this paper, we propose an alternative front-end for speech recognition over GSM networks. This front-end is specially conceived to be effective against source coding distortion and transmission errors. Specifically, we suggest extracting the recognition feature vectors directly from the encoded speech (i.e., the bitstream) instead of decoding it and subsequently extracting the feature vectors. This approach offers two significant advantages. First, the recognition system is only affected by the quantization distortion of the spectral envelope. Thus, we are avoiding the influence of other sources of distortion as a result of the encoding-decoding process. Second, when transmission errors occur, our front-end becomes more effective since it is not affected by errors in bits allocated to the excitation signal. We have considered the half and the full-rate standard codecs and compared the proposed front-end with the conventional approach in two ASR tasks, namely, speaker-independent isolated digit recognition and speaker-independent continuous speech recognition. In general, our approach outperforms the conventional procedure, for a variety of simulated channel conditions. Furthermore, the disparity increases as the network conditions worsen. Ascensión Gallardo-Antolín, Carmen Peláez-Moreno, Fernando Díaz-de-María |
IEEE Trans. Speech Audio Process. | 3 |
| 2003 | Linear equalization of the modulation spectra: a novel approach for noisy speech recognitionabstractThe paper tackles the problem of noisy speech recognition. In particular, we present a novel approach to the design of filters for processing the modulation spectrum, that we call linear equalization. We postulate that, as long as the distortion of the spectral parameters due to noise can be modeled as linear, an advantageous solution consists of estimating this linear perturbation system and designing its inverse system (the equalizer). Our experimental results show that the proposed method is very effective for three of the five types of noise considered. Fernando Díaz-de-María, Jesús Vicente-Peña, Ascensión Gallardo-Antolín, Carmen Peláez-Moreno |
ICASSP (2) | 1 |
| 2002 | An Application of SVM to Lost Packets Reconstruction in Voice-Enabled Services
Carmen Peláez-Moreno, Emilio Parrado-Hernández, Ascensión Gallardo-Antolín, Adrián Zambrano-Miranda, Fernando Díaz-de-María |
ICANN | 5 |
| 2002 | Influence of transmission errors on ASR systems
Carmen Peláez-Moreno, Ascensión Gallardo-Antolín, Jesús Vicente-Peña, Fernando Díaz-de-María |
INTERSPEECH | 4 |
| 2001 | A robust front-end for ASR over IP snd GSM networks: an integrated scenario
Ascensión Gallardo-Antolín, Carmen Peláez-Moreno, Fernando Díaz-de-María |
INTERSPEECH | 3 |
| 2001 | Recognizing voice over IP: a robust front-end for speech recognition on the world wide webabstractThe Internet Protocol (IP) environment poses two relevant sources of distortion to the speech recognition problem: lossy speech coding and packet loss. In this paper, we propose a new front-end for speech recognition over IP networks. Specifically, we suggest extracting the recognition feature vectors directly from the encoded speech (i.e., the bit stream) instead of decoding it and subsequently extracting the feature vectors. This approach offers two significant benefits. First, the recognition system is only affected by the quantization distortion of the spectral envelope. Thus, we are avoiding the influence of other sources of distortion due to the encoding-decoding process. Second, when packet loss occurs, our front-end becomes more effective since it is not constrained to the error handling mechanism of the codec. We have considered the ITU G.723.1 standard codec, which is one of the most preponderant coding algorithms in voice over IP (VoIP) and compared the proposed front-end with the conventional approach in two automatic speech recognition (ASR) tasks, namely, speaker-independent isolated digit recognition and speaker-independent continuous speech recognition. In general, our approach outperforms the conventional procedure, for a variety of simulated packet loss rates. Furthermore, the improvement is higher as network conditions worsen. Carmen Peláez-Moreno, Ascensión Gallardo-Antolín, Fernando Díaz-de-María |
IEEE Trans. Multim. | 3 |
| 1999 | Avoiding distortions due to speech coding and transmission errors in GSM ASR tasksabstractWe have extended our previous research on a new approach to automatic speech recognition (ASR) in the GSM environment. Instead of recognizing from the decoded speech signal, our system works from the digital speech representation used by the GSM encoder. We have compared the performance of a conventional system and the one we propose on a speaker independent, isolated-digit ASR task. For the half and full-rate GSM codecs, from our results, we conclude that the proposed approach is much more effective in coping with the coding distortion and transmission errors. Furthermore, in clean speech conditions, our approach does not impoverish the recognition performance, even recognizing from GSM digital speech, in comparison with a conventional system working on unencoded speech. Ascensión Gallardo-Antolín, Fernando Díaz-de-María, Francisco J. Valverde-Albacete |
ICASSP | 2 |
| 1999 | Backward adaptive RBF-based hybrid predictors for CELP-type coders at medium bit-rates
Carmen Peláez-Moreno, Fernando Díaz-de-María |
EUROSPEECH | 2 |
| 1998 | Recognition from GSM digital speech
Ascensión Gallardo-Antolín, Fernando Díaz-de-María, Francisco J. Valverde-Albacete |
ICSLP | 2 |
| 1996 | A new inverse filter criterion for blind deconvolution of spiky signals using Gaussian mixturesabstractThis paper presents a new Bussgang-type technique for blind deconvolution of spiky signals. Based on a Gaussian mixture model for the spiky signal, the method obtains a deconvolution filter and a zero-memory nonlinearity to estimate the signal. A new updating procedure for the mixture parameters (and, therefore, for the nonlinear estimator) is included in the algorithm: it allows to apply the algorithm without any prior knowledge about the signal and noise. A simulation example illustrates the performance of the proposed method. Ignacio Santamaría, Carlos Pantaleón, Fernando Díaz-de-María, Antonio Artés-Rodríguez |
ICASSP | 3 |
| 1995 | Nonlinear prediction for speech coding using radial basis functionsabstractRadial basis functions (RBF) networks constitute an interesting option for dealing with nonlinear prediction of speech because they provide a regularized solution. They can guarantee the stability of the corresponding synthesis scheme; consequently, they are used in code excited nonlinear prediction (CENP) coders. This approach is presented, and some simulations examples show its advantage in the prediction performance. The practical implementations of CENP coders are also addressed. Fernando Díaz-de-María, Fernando R. Figueiras-Vidal |
ICASSP | 1 |