EDBT 2026 Demo / reviewers in the wild / expert
Valentina Sanguineti
dblp:239/5615
· DBLP profile ↗
6ranked-venue papers
4as first author
3since 2021 · last 2023
0000-0001-7995-6205ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Audio and music processing · 64% Multimedia analysis and retrieval · 28% Computational photography and imaging · 8% | |
| Artificial intelligence
2 papers |
3D vision · 54% Representation and self-supervised learning · 46% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Multimedia analysis and retrieval › audio-visual learning
audio-visual scene understanding |
0.6 | 1 | 2022 | Unsupervised Synthetic Acoustic Image Generation for Audio-Visual Scene Understanding · IEEE Trans. Image Process. 2022 |
Computer vision › 3D vision
multimodal scene understanding |
0.5 | 1 | 2021 | Audio-Visual Localization by Synthetic Acoustic Image Generation · AAAI 2021 |
Audio and music processing › sound source localization
audio-visual sound source localization |
0.5 | 1 | 2021 | Audio-Visual Localization by Synthetic Acoustic Image Generation · AAAI 2021 |
Audio and music processing
sound source localization |
0.5 | 1 | 2021 | Audio-Visual Localization by Synthetic Acoustic Image Generation · AAAI 2021 |
Machine learning › Representation and self-supervised learning › representation learning › sequence representation learning
audio representation learning |
0.4 | 1 | 2020 | Leveraging Acoustic Images for Effective Self-supervised Audio Representation Learning · ECCV (22) 2020 |
Computational photography and imaging
acoustic imaging |
0.2 | 1 | 2022 | Unsupervised Synthetic Acoustic Image Generation for Audio-Visual Scene Understanding · IEEE Trans. Image Process. 2022 |
Audio and music processing
spatial audio |
0.2 | 1 | 2022 | Unsupervised Synthetic Acoustic Image Generation for Audio-Visual Scene Understanding · IEEE Trans. Image Process. 2022 |
Methods — techniques the papers use, named apart from their topics
variational autoencoder · 1.6u-net · 1.6deep architecture · 1.0self-supervised learning · 0.9cross-modal conditioning · 0.6adversarial model · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Audio-Visual Inpainting: Reconstructing Missing Visual Information with SoundabstractWe tackle audio-visual inpainting, the problem of completing an image in such a way to be consistent with the sound associated to the scene. To this end, we propose a multimodal, audio-visual inpainting method (AVIN), and show how to leverage sound to reconstruct semantically consistent images. AVIN is a 2-stage algorithm, which first learns the scene semantics and reconstructs low resolution images based on a conditional probability distribution of pixels in the space conditioned to audio, and then refines such result with a GAN-based network to increase the resolution of the reconstructed image. We show that AVIN is able to recover the original content, especially in the hard cases where the missing area heavily degrades the scene semantics: it can perform cross-modal generation whenever no visual context is observed at all, reconstructing visual data from sound only. Code will be made available upon acceptance. Valentina Sanguineti, Sanket Kumar Thakur, Pietro Morerio, Alessio Del Bue, Vittorio Murino |
ICASSP | 1 |
| 2022 | Unsupervised Synthetic Acoustic Image Generation for Audio-Visual Scene UnderstandingabstractAcoustic images are an emergent data modality for multimodal scene understanding. Such images have the peculiarity of distinguishing the spectral signature of the sound coming from different directions in space, thus providing a richer information as compared to that derived from single or binaural microphones. However, acoustic images are typically generated by cumbersome and costly microphone arrays which are not as widespread as ordinary microphones. This paper shows that it is still possible to generate acoustic images from off-the-shelf cameras equipped with only a single microphone and how they can be exploited for audio-visual scene understanding. We propose three architectures inspired by Variational Autoencoder, U-Net and adversarial models, and we assess their advantages and drawbacks. Such models are trained to generate spatialized audio by conditioning them to the associated video sequence and its corresponding monaural audio track. Our models are trained using the data collected by a microphone array as ground truth. Thus they learn to mimic the output of an array of microphones in the very same conditions. We assess the quality of the generated acoustic images considering standard generation metrics and different downstream tasks (classification, cross-modal retrieval and sound localization). We also evaluate our proposed models by considering multimodal datasets containing acoustic images, as well as datasets containing just monaural audio signals and RGB video frames. In all of the addressed downstream tasks we obtain notable performances using the generated acoustic data, when compared to the state of the art and to the results obtained using real acoustic images as input. Valentina Sanguineti, Pietro Morerio, Alessio Del Bue, Vittorio Murino |
IEEE Trans. Image Process. | 1 |
| 2021 | Audio-Visual Localization by Synthetic Acoustic Image GenerationabstractAcoustic images constitute an emergent data modality for multimodal scene understanding. Such images have the peculiarity to distinguish the spectral signature of sounds coming from different directions in space, thus providing richer information than the one derived from mono and binaural microphones. However, acoustic images are typically generated by cumbersome microphone arrays, which are not as widespread as ordinary microphones mounted on optical cameras. To exploit this empowered modality while using standard microphones and cameras we propose to leverage the generation of synthetic acoustic images from common audio-video data for the task of audio-visual localization. The generation of synthetic acoustic images is obtained by a novel deep architecture, based on Variational Autoencoder and U-Net models, which is trained to reconstruct the ground truth spatialized audio data collected by a microphone array, from the associated video and its corresponding monaural audio signal. Namely, the model learns how to mimic what an array of microphones can produce in the same conditions. We assess the quality of the generated synthetic acoustic images on the task of unsupervised sound source localization in a qualitative and quantitative manner, while also considering standard generation metrics. Our model is evaluated by considering both multimodal datasets containing acoustic images, used for the training, and unseen datasets containing just monaural audio signals and RGB frames, showing to reach more accurate localization results as compared to the state of the art. Valentina Sanguineti, Pietro Morerio, Alessio Del Bue, Vittorio Murino |
AAAI | 1 |
| 2020 | Leveraging Acoustic Images for Effective Self-supervised Audio Representation Learning
Valentina Sanguineti, Pietro Morerio, Niccolò Pozzetti, Danilo Greco, Marco Cristani, Vittorio Murino |
ECCV (22) | 1 |
| 2020 | Audio-Visual Model Distillation Using Acoustic ImagesabstractIn this paper, we investigate how to learn rich and robust feature representations for audio classification from visual data and acoustic images, a novel audio data modality. Former models learn audio representations from raw signals or spectral data acquired by a single microphone, with remarkable results in classification and retrieval. However, such representations are not so robust towards variable environmental sound conditions. We tackle this drawback by exploiting a new multimodal labeled action recognition dataset acquired by a hybrid audio-visual sensor that provides RGB video, raw audio signals, and spatialized acoustic data, also known as acoustic images, where the visual and acoustic images are aligned in space and synchronized in time. Using this richer information, we train audio deep learning models in a teacher-student fashion. In particular, we distill knowledge into audio networks from both visual and acoustic image teachers. Our experiments suggest that the learned representations are more powerful and have better generalization capabilities than the features learned from models trained using just single-microphone audio data. Andrés F. Pérez, Valentina Sanguineti, Pietro Morerio, Vittorio Murino |
WACV | 2 |
| 2018 | Learning Switching Models for Abnormality Detection for Autonomous DrivingabstractWe present an approach to learn a model to estimate the dynamical states at continuous and discrete inference levels when trajectory information is available. We learn from sparse data a probabilistic switching model that generates trajectories associated with a stationary plan of an agent. The learned generative model is used within a Markov Jump Linear System (MJLSs) to switch among set of space dependent linear filters that analyze new trajectories and detect deviations from the learned model based on internal innovation measurements. We show examples of application of the proposed approach to learn filters for evaluating deviations from a reference human driving task execution that includes static and dynamic obstacle avoidance. Mohamad Baydoun, Damian Campo, Valentina Sanguineti, Lucio Marcenaro, Andrea Cavallaro, Carlo S. Regazzoni |
FUSION | 3 |