VLDB 2026 Research / reviewers in the wild / expert
Gloria Haro
dblp:h/GloriaHaro · also Gloria Haro Ortega
· DBLP profile ↗
29ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0002-8194-8092ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 16 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Critical Assessment of Visual Sound Source Localization Models Including Negative AudioabstractThe task of Visual Sound Source Localization (VSSL) involves identifying the location of sound sources in visual scenes, integrating audio-visual data for enhanced scene understanding. Despite advancements in state-of-the-art (SOTA) models, we observe three critical flaws: i) The evaluation of the models is mainly focused in sounds produced by objects that are visible in the image, ii) The evaluation often assumes a prior knowledge of the size of the sounding object, and iii) No universal threshold for localization in real-world scenarios is established, as previous approaches only consider positive examples without accounting for both positive and negative cases. In this paper, we introduce extended test sets and new metrics designed to complete the current standard evaluation of VSSL models by testing them in scenarios where none of the objects in the image corresponds to the audio input, i.e. a negative audio. We consider three types of negative audio: silence, noise and offscreen. Our analysis reveals that numerous SOTA models fail to appropriately adjust their predictions based on audio input, suggesting that these models may not be leveraging audio information as intended. Additionally, we provide a comprehensive analysis of the range of maximum values in the estimated audio-visual similarity maps, in both positive and negative audio cases, and show that most of the models are not discriminative enough, making them unfit to choose a universal threshold appropriate to perform sound localization without any a priori information of the sounding object, that is, object size and visibility. Xavier Juanola, Gloria Haro, Magdalena Fuentes |
ICASSP | 2 |
| 2023 | Speech inpainting: Context-based speech synthesis guided by videoabstractAudio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorporate the visual modality to enhance an acoustic speech signal or even restore missing audio information. Specifically, this paper focuses on the problem of audio-visual speech inpainting, which is the task of synthesizing the speech in a corrupted audio segment in a way that it is consistent with the corresponding visual content and the uncorrupted audio context. We present an audio-visual transformer-based deep learning model that leverages visual cues that provide information about the content of the corrupted audio. It outperforms the previous state-of-the-art audio-visual model and audio-only baselines. We also show how visual features extracted with AV-HuBERT, a large audiovisual transformer for speech recognition, are suitable for synthesizing speech. Juan F. Montesinos, Daniel Michelsanti, Gloria Haro, Zheng-Hua Tan, Jesper Jensen 0001 |
INTERSPEECH | 3 |
| 2022 | VoViT: Low Latency Graph-Based Audio-Visual Voice Separation Transformer
Juan F. Montesinos, Venkatesh S. Kadandale, Gloria Haro |
ECCV (37) | 3 |
| 2022 | VocaLiST: An Audio-Visual Synchronisation Model for Lips and VoicesabstractComunicació presentada a Interspeech 2022, celebrat del 18 al 22 de setembre de 2022 a Inchon, Corea del Sud. Venkatesh S. Kadandale, Juan F. Montesinos, Gloria Haro |
INTERSPEECH | 3 |
| 2021 | A cappella: Audio-visual Singing Voice Separation
Juan F. Montesinos, Venkatesh S. Kadandale, Gloria Haro |
BMVC | 3 |
| 2021 | DIP-VBTV: A Color Image Restoration Model Combining a Deep Image Prior and a Vector Bundle Total VariationabstractIn this paper, we introduce a new variational model for color image restoration, called DIP-VBTV, which combines two priors: a deep image prior (DIP), which assumes that the restored image can be generated through a neural network, and a vector bundle total variation (VBTV), which generalizes the vectorial total variation (VTV) on vector bundles. VBTV is determined by a geometric triplet: a Riemannian metric on the base manifold, a covariant derivative, and a metric on the vector bundle. Whereas the VTV prior encourages the restored images to be piecewise constant, the VBTV prior encourages them to be piecewise parallel with respect to a covariant derivative. For well-chosen geometric triplets, we show that the minimization of VBTV encourages the solutions of the restoration model to share some visual content with the clean image. Then, we show in experiments that DIP-VBTV benefits from this property by outperforming DIP-VTV and state-of-the-art unsupervised methods. It demonstrates the relevance of combining DIP and VBTV priors. Thomas Batard, Gloria Haro, Coloma Ballester |
SIAM J. Imaging Sci. | 2 |
| 2021 | Conditioned Source Separation for Musical Instrument PerformancesabstractIn music source separation, the number of sources may vary for each piece and some of the sources may belong to the same family of instruments, thus sharing timbral characteristics and making the sources more correlated. This leads to additional challenges in the source separation problem. This paper proposes a source separation method for multiple musical instruments sounding simultaneously and explores how much additional information apart from the audio stream can lift the quality of source separation. We explore conditioning techniques at different levels of a primary source separation network and utilize two extra modalities of data, namely presence or absence of instruments in the mixture, and the corresponding video stream data. Olga Slizovskaia, Gloria Haro, Emilia Gómez |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Always Look On The Bright Side Of The Field: Merging Pose And Contextual Data To Estimate Orientation Of Soccer PlayersabstractAlthough orientation has proven to be a key skill of soccer players in order to succeed in a broad spectrum of plays, body orientation is a yet-little-explored area in sports analytics' research. Despite being an inherently ambiguous concept, player orientation can be defined as the projection (2D) of the normal vector placed in the center of the upper-torso of players (3D). This research presents a novel technique to obtain player orientation from monocular video recordings by mapping pose parts (shoulders and hips) in a 2D field by combining OpenPose with a super-resolution network, and merging the obtained estimation with contextual information (ball position). Results have been validated with players-held EPTS devices, obtaining a median error of 27 degrees/player. Moreover, three novel types of orientation maps are proposed in order to make raw orientation data easy to visualize and understand, thus allowing further analysis at team- or player-level. Adrià Arbués Sangüesa, Adrián Martín, Javier Fernández 0006, Gloria Haro, Coloma Ballester |
ICIP | 5 |
| 2020 | Vocoder-Based Speech Synthesis from Silent VideosabstractBoth acoustic and visual information influence human perception of speech. For this reason, the lack of audio in a video sequence determines an extremely low speech intelligibility for untrained lip readers. In this paper, we present a way to synthesise speech from the silent video of a talker using deep learning. The system learns a mapping function from raw video frames to acoustic features and reconstructs the speech with a vocoder synthesis algorithm. To improve speech reconstruction performance, our model is also trained to predict text information in a multi-task learning fashion and it is able to simultaneously reconstruct and recognise speech in real time. The results in terms of estimated speech quality and intelligibility show the effectiveness of our method, which exhibits an improvement over existing video-to-speech approaches. Daniel Michelsanti, Olga Slizovskaia, Gloria Haro, Emilia Gómez, Zheng-Hua Tan, Jesper Jensen 0001 |
INTERSPEECH | 3 |
| 2020 | Multi-channel U-Net for Music Source SeparationabstractA fairly straightforward approach for music source separation is to train independent models, wherein each model is dedicated for estimating only a specific source. Training a single model to estimate multiple sources generally does not perform as well as the independent dedicated models. However, Conditioned U-Net (C-U-Net) uses a control mechanism to train a single model for multi-source separation and attempts to achieve a performance comparable to that of the dedicated models. We propose a multi-channel U-Net (M-U-Net) trained using a weighted multi-task loss as an alternative to the C-U-Net. We investigate two weighting strategies for our multi-task loss: 1) Dynamic Weighted Average (DWA), and 2) Energy Based Weighting (EBW). DWA determines the weights by tracking the rate of change of loss of each task during training. EBW aims to neutralize the effect of the training bias arising from the difference in energy levels of each of the sources in a mixture. Our methods provide three-fold advantages compared to C-U-Net: 1) Fewer effective training iterations per epoch, 2) Fewer trainable network parameters (no control parameters), and 3) Faster processing at inference. Our methods achieve performance comparable to that of C-U-Net and the dedicated U-Nets at a much lower training cost. Venkatesh S. Kadandale, Juan F. Montesinos, Gloria Haro, Emilia Gómez |
MMSP | 3 |
| 2020 | Solos: A Dataset for Audio-Visual Music AnalysisabstractIn this paper, we present a new dataset of music performance videos which can be used for training machine learning methods for multiple tasks such as audio-visual blind source separation and localization, cross-modal correspondences, cross-modal generation and, in general, any audio-visual self-supervised task. These videos, gathered from YouTube, consist of solo musical performances of 13 different instruments. Compared to previously proposed audio-visual datasets, Solos is cleaner since a big amount of its recordings are auditions and manually checked recordings, ensuring there is no background noise nor effects added in the video post-processing. Besides, it is, up to the best of our knowledge, the only dataset that contains the whole set of instruments present in the URMP dataset, a high-quality dataset of 44 audio-visual recordings of multi-instrument classical music pieces with individual audio tracks. URMP was intented to be used for source separation, thus, we evaluate the performance on the URMP dataset of two different source-separation models trained on Solos. The dataset is publicly available at https://juanfmontesinos.github.io/Solos/. Juan F. Montesinos, Olga Slizovskaia, Gloria Haro |
MMSP | 3 |
| 2019 | Deep Single Image Camera Calibration With Radial DistortionabstractSingle image calibration is the problem of predicting the camera parameters from one image. This problem is of importance when dealing with images collected in uncontrolled conditions by non-calibrated cameras, such as crowd-sourced applications. In this work we propose a method to predict extrinsic (tilt and roll) and intrinsic (focal length and radial distortion) parameters from a single image. We propose a parameterization for radial distortion that is better suited for learning than directly predicting the distortion parameters. Moreover, predicting additional heterogeneous variables exacerbates the problem of loss balancing. We propose a new loss function based on point projections to avoid having to balance heterogeneous loss terms. Our method is, to our knowledge, the first to jointly estimate the tilt, roll, focal length, and radial distortion parameters from a single image. We thoroughly analyze the performance of the proposed method and the impact of the improvements and compare with previous approaches for single image radial distortion correction. Manuel Lopez, Roger Marí, Pau Gargallo, Yubin Kuang, Javier González 0001, Gloria Haro |
CVPR | 6 |
| 2019 | End-to-end Sound Source Separation Conditioned on Instrument LabelsabstractCan we perform an end-to-end music source separation with a variable number of sources using a deep learning model? This paper presents an extension of the Wave-U-Net [1] model which allows end-to-end monaural source separation with a non-fixed number of sources. Furthermore, we propose multiplicative conditioning with instrument labels at the bottleneck of the Wave-U-Net and show its effect on the separation results. This approach can be further extended to other types of conditioning such as audio-visual source separation and score-informed source separation. Olga Slizovskaia, Leo Kim, Gloria Haro, Emilia Gómez |
ICASSP | 3 |
| 2018 | L1 Patch-Based Image Partitioning into Homogeneous Textured RegionsabstractThis paper proposes a novel patch-based variational segmentation method that considers adaptive patches to characterize, in an affine invariant way, the local structure of each homogeneous texture region of the image and thus being capable of grouping the same kind of texture regardless of differences in the point of view or suffered perspective distortion. The patches are computed using an affine covariant structure tensor defined at every pixel of the image domain, so that they can automatically adapt its shape and size. They are used in a segmentation model that uses an L1-norm fidelity term and fuzzy membership functions, which is solved by an alternating scheme. The output of the method is a partition of the image in regions with homogeneous texture together with a patch representative of the texture of each region. Maria Oliver, Gloria Haro, Vadim Fedorov, Coloma Ballester |
ICASSP | 2 |
| 2018 | Motion Inpainting by an Image-Based Geodesic AMLE MethodabstractThis work presents an automatic method for optical flow inpainting. Given a video, each frame domain is endowed with a Riemannian metric based on the video pixel values. The missing optical flow is recovered by solving the Absolutely Minimizing Lipschitz Extension (AMLE) partial differential equation on the Riemannian manifold. An efficient numerical algorithm is proposed using eikonal operators for nonlinear elliptic partial differential equations on a finite graph. The choice of the metric is discussed and the method is applied to optical flow inpainting and sparse-to-dense optical flow estimation, achieving top-tier performance in terms of End-Point-Error (EPE). Maria Oliver, Lara Raad, Coloma Ballester, Gloria Haro |
ICIP | 4 |
| 2017 | Spatio-temporal binary video inpainting via threshold dynamicsabstractWe propose a new variational method for the completion of moving shapes through binary video inpainting that works by smoothly recovering the objects into an inpainting hole. We solve it by a simple dynamic shape analysis algorithm based on threshold dynamics. The model takes into account the optical flow and motion occlusions. The resulting inpainting algorithm diffuses the available information along the space and the visible trajectories of the pixels in time. We show its performance with examples from the Sintel dataset, which contains complex object motion and occlusions. Maria Oliver, Roberto P. Palomares, Coloma Ballester, Gloria Haro |
ICASSP | 4 |
| 2017 | Musical Instrument Recognition in User-generated Videos using a Multimodal Convolutional Neural Network ArchitectureabstractThis paper presents a method for recognizing musical instruments in user-generated videos. Musical instrument recognition from music signals is a well-known task in the music information retrieval (MIR) field, where current approaches rely on the analysis of the good-quality audio material. This work addresses a real-world scenario with several research challenges, i.e. the analysis of user-generated videos that are varied in terms of recording conditions and quality and may contain multiple instruments sounding simultaneously and background noise. Our approach does not only focus on the analysis of audio information, but we exploit the multimodal information embedded in the audio and visual domains. In order to do so, we develop a Convolutional Neural Network (CNN) architecture which combines learned representations from both modalities at a late fusion stage. Our approach is trained and evaluated on two large-scale video datasets: YouTube-8M and FCVID. The proposed architectures demonstrate state-of-the-art results in audio and video object recognition, provide additional robustness to missing modalities, and remains computationally cheap to train. Olga Slizovskaia, Emilia Gómez, Gloria Haro |
ICMR | 3 |
| 2014 | A Rotation-Invariant Regularization Term for Optical Flow Related Problems
Roberto P. Palomares, Gloria Haro, Coloma Ballester |
ACCV (5) | 2 |
| 2014 | Shape from silhouette consensus and photo-consistencyabstractWe propose a 3D reconstruction algorithm based on silhouettes and color images. It is robust to inconsistent silhouettes, often common in real applications due to occlusions, errors in the background subtraction, noise or even calibration errors. The recovery of the shape that best fits the available data is formulated as a continuous energy minimization problem. The energy is based on the error between the silhouettes and the shape plus a regularization term based on a photo-consistency measure that places the surface at photo-consistent locations. The visibility is modeled as a function of the shape. The proposed photo-consistency measure takes visibility into account, although the presented variational framework can use different photo-consistency computations. Gloria Haro |
ICIP | 1 |
| 2012 | Shape from Silhouette Consensus
Gloria Haro |
Pattern Recognit. | 1 |
| 2012 | Enhanced foreground segmentation and tracking combining Bayesian background, shadow and foreground modeling
Jaime Gallego, Montse Pardàs, Gloria Haro |
Pattern Recognit. Lett. | 3 |
| 2012 | Photographing Paintings by Image FusionabstractThis paper addresses the problem of obtaining a quality photograph of a painting by multi-image fusion methods. The problem is particularly challenging because of the uncontrolled illumination conditions and of the destructive reflection speckle present in most photographs of paintings. A fully automatic image processing chain is described that, starting from several bursts of a painting taken under different angles, permits one to obtain the best possible result by eliminating highlights and motion blur by robust statistics, reducing noise by fusion, and compensating optical distortion in the registration process. This image fusion method is applicable to photographs of a painting taken with a hand-held camera without any particular setup. It works under bad lighting conditions and eliminates motion blur, even when the painting is protected by a glass screen creating structured reflections of the room. The careful discussion of each step of the processing chain also permits one to review and discuss the efficiency of the image fusion tools recently proposed in the literature and insert several new ones in the chain. Gloria Haro, Antoni Buades, Jean-Michel Morel |
SIAM J. Imaging Sci. | 1 |
| 2010 | Shape from incomplete silhouettes based on the reprojection error
Gloria Haro, Montse Pardàs |
Image Vis. Comput. | 1 |
| 2009 | Bayesian foreground segmentation and tracking using pixel-wise background model and region based foreground modelabstractIn this paper we present a segmentation system for monocular video sequences with static camera that aims at foreground/background separation and tracking. We propose to combine a simple pixel-wise model for the background with a general purpose region based model for the foreground. The background is modeled using one Gaussian per pixel, thus achieving a precise and easy to update model. The foreground is modeled using a Gaussian mixture model with feature vectors consisting of the spatial (x, y) and colour (r, g, b) components. The spatial components of this model are updated using the expectation maximization algorithm after the classification of each frame. The background model is formulated in the 5 dimensional feature space in order to be able to apply a maximum a posteriori framework for the classification. The classification is done using a graph cut algorithm that allows taking into account neighborhood information. The results presented in the paper show the improvement of the system in situations where the foreground objects have similar colors to those of the background. Jaime Gallego, Montse Pardàs, Gloria Haro |
ICIP | 3 |
| 2008 | On geometric variational models for inpainting surface holes
Vicent Caselles, Gloria Haro, Guillermo Sapiro, Joan Verdera |
Comput. Vis. Image Underst. | 2 |
| 2008 | Translated Poisson Mixture Model for Stratification Learning
Gloria Haro, Gregory Randall, Guillermo Sapiro |
Int. J. Comput. Vis. | 1 |
| 2007 | Regularized Mixed Dimensionality and Density Learning in Computer VisionabstractA framework for the regularized estimation of nonuniform dimensionality and density in high dimensional data is introduced in this work. This leads to learning stratifications, that is, mixture of manifolds representing different characteristics and complexities in the data set. The basic idea relies on modeling the high dimensional sample points as a process of Poisson mixtures, with regularizing restrictions and spatial continuity constraints. Theoretical asymptotic results for the model are presented as well. The presentation of the framework is complemented with artificial and real examples showing the importance of regularized stratification learning in computer vision applications. Gloria Haro, Gregory Randall, Guillermo Sapiro |
CVPR | 1 |
| 2006 | Stratification Learning: Detecting Mixed Density and Dimensionality in High Dimensional Point CloudsabstractThe study of point cloud data sampled from a stratification, a collection of manifolds with possible different dimensions, is pursued in this paper. We present a technique for simultaneously soft clustering and estimating the mixed dimensionality and density of such structures. The framework is based on a maximum likelihood estimation of a Poisson mixture model. The presentation of the approach is completed with artificial and real examples demonstrating the importance of extending manifold learning to stratification learning. Gloria Haro, Gregory Randall, Guillermo Sapiro |
NIPS | 1 |
| 2006 | Visual Acuity in Day for Night
Gloria Haro, Marcelo Bertalmío, Vicent Caselles |
Int. J. Comput. Vis. | 1 |