VLDB 2026 Research / reviewers in the wild / expert
Sebastià Vicenc Amengual Garí
dblp:270/4740
· DBLP profile ↗
13ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-9801-1108ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Audiovisual Realism in MR: Investigating the Effects of Room Acoustics on Co-Presence with Photorealistic AvatarsabstractWith advances in spatial computing and growth in commercially available head-mounted displays, mixed reality (MR) is an emerging context for social experiences and collaborative interaction. While the role of visual representations in enhancing user experiences has been extensively studied, the contribution of audio has comparatively received little attention. In this exploratory, within-subjects study, we investigate the effects of audio quality on co-presence and associated experiential outcomes during avatar-mediated conversations in MR. In dyads, participants engaged in semi-structured conversation under anechoic and reverberant audio conditions while embodying photorealistic avatars. Results shed a nuanced light on the potential of audio in facilitating co-presence. We conclude by discussing implications for design and future research on audio within audiovisual MR environments. Cyan DeVeaux, Elizabeth H. Hall, Andy J. Shaw, Andrew Frederick Francl, Paulus van Horne, Frank M. Nieuwenhuizen, Sebastià Vicenc Amengual Garí, Madeline Huberth |
VR | 7 |
| 2025 | Hearing Anywhere in Any EnvironmentabstractIn mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR) estimation, most existing methods are limited to the single environment on which they are trained, lacking the ability to generalize to new rooms with different geometries and surface materials. We aim to develop a unified model capable of reconstructing the spatial acoustic experience of any environment with minimum additional measurements. To this end, we present xRIR, a framework for cross-room RIR prediction. The core of our generalizable approach lies in combining a geometric feature extractor, which captures spatial context from panorama depth images, with a RIR encoder that extracts detailed acoustic features from only a few reference RIR samples. To evaluate our method, we introduce AcousticRooms, a new dataset featuring high-fidelity simulation of over 300,000 RIRs from 260 rooms. Experiments show that our method strongly outperforms a series of baselines. Furthermore, we successfully perform sim-to-real transfer by evaluating our model on four real-world environments, demonstrating the generalizability of our approach and the realism of our dataset. Xiulong Liu 0002, Anurag Kumar 0003, Paul Calamia, Sebastià Vicenc Amengual Garí, Calvin Murdock, Ishwarya Ananthabhotla, Philip W. Robinson, Eli Shlizerman, Vamsi K. Ithapu, Ruohan Gao |
CVPR | 4 |
| 2024 | Binaural Room Transfer Function Interpolation Via System InversionabstractThis paper is concerned with the spatial interpolation of Binaural Room Transfer Functions (BRTFs). The proposed method is a binaural extension of Room Transfer Function (RTF) interpolation methods framed as inverse problems, and is based on a parametric representation of the sound field using either Plane Waves (PWs) or Equivalent Sources (ESs). Once the parameters are obtained via system inversion, the BRTFs can be synthesised at any other position. Four combinations of acoustic models (PW or ES) and regularisation functions (Tikhonov and l1-norm) are tested. The proposed method is shown to have a good performance below 1 kHz, comparable to standard pressure-based RTF interpolation. Using PWs with l1-norm regularisation produces optimal solutions, resulting in a Normalised Mean Squared Error (NMSE) of −15 dB at 700 Hz using 25 BRTF measurements. The important case where the listener’s Head-Related Transfer Function (HRTF) is unknown is also tested, revealing that using a different HRTF for inversion and synthesis did not yield a significant drop in performance. It is hypothesised that this is due to the small variations between individual HRTF magnitudes within the operating frequency range of this class of models. Amal Emthyas, Sebastià Vicenc Amengual Garí, Enzo De Sena |
ICASSP | 2 |
| 2024 | Blind Identification of Binaural Room Impulse Responses From Smart GlassesabstractSmart glasses are increasingly recognized as a key medium for augmented reality, offering a hands-free platform with integrated microphones and non-ear-occluding loudspeakers to seamlessly mix virtual sound sources into the real-world acoustic scene. To convincingly integrate virtual sound sources, the room acoustic rendering of the virtual sources must match the real-world acoustics. Information about a user's acoustic environment however is typically not available. This work uses a microphone array in a pair of smart glasses to blindly identify binaural room impulse responses (BRIRs) from a few seconds of speech in the real-world environment. The proposed method uses dereverberation and beamforming to generate a pseudo reference signal that is used by a multichannel Wiener filter to estimate room impulse responses which are then converted to BRIRs. The multichannel room impulse responses can be used to estimate room acoustic parameters which is shown to outperform baseline algorithms in the estimation of reverberation time and direct-to-reverberant energy ratio. Results from a listening experiment further indicate that the estimated BRIRs often reproduce the real-world room acoustics perceptually more convincingly than measured BRIRs from other rooms of similar size. Thomas Deppisch, Nils Meyer-Kahlen, Sebastià Vicenc Amengual Garí |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | The R3VIVAL Dataset: Repository of Room Responses and 360 Videos of a Variable Acoustics LababstractThis paper presents a dataset of spatial room impulse responses (SRIRs) and 360° stereoscopic video captures of a variable acoustics laboratory. A total of 34 source positions are measured with 8 different acoustic panel configurations, resulting in a total of 272 SRIRs. The source positions are arranged in 30° increments at concentric circles of radius 1.5, 2, and 3 m measured with a directional studio monitor, as well as 4 extra positions at the room corners measured with an omnidirectional source. The receiver is a 7 channel open microphone array optimized for its use with the Spatial Decomposition Method (SDM). The 8 acoustic configurations are achieved by setting a subset of the panels to their absorptive configuration in 5 steps (0%, 25%, 50%, 75%, 100% of the panels), as well as 3 configurations in which entire walls are set to their absorptive configuration (right, right/back, right/back/left). Video captures of the laboratory and a second room are obtained using a 360° stereoscopic camera with a resolution of 4096 × 2160 pixels, covering the same source/receiver combinations. Furthermore, we present an acoustic analysis of both time-energy and spatio-temporal parameters showcasing the differences in the measured configurations. The dataset, together with spatial analysis and rendering scripts, is publicly released in a GitHub repository1. Florian Klein 0003, Sebastià Vicenc Amengual Garí |
ICASSP | 2 |
| 2023 | Direct and Residual Subspace Decomposition of Spatial Room Impulse ResponsesabstractPsychoacoustic experiments have shown that directional properties of the direct sound, salient reflections, and the late reverberation of an acoustic room response can have a distinct influence on the auditory perception of a given room. Spatial room impulse responses (SRIRs) capture those properties and thus are used for direction-dependent room acoustic analysis and virtual acoustic rendering. This work proposes a subspace method that decomposes SRIRs into a direct part, which comprises the direct sound and the salient reflections, and a residual, to facilitate enhanced analysis and rendering methods by providing individual access to these components. The proposed method is based on the generalized singular value decomposition and interprets the residual as noise that is to be separated from the other components of the reverberation. Large generalized singular values are attributed to the direct part, which is then obtained as a low-rank approximation of the SRIR. By advancing from the end of the SRIR toward the beginning while iteratively updating the residual estimate, the method adapts to spatio-temporal variations of the residual. The method is evaluated using a spatio-spectral error measure and simulated SRIRs of different rooms, microphone arrays, and ratios of direct sound to residual energy. The proposed method creates lower errors than existing approaches in all tested scenarios, including a scenario with two simultaneous reflections. A case study with measured SRIRs shows the applicability of the method under real-world acoustic conditions. A reference implementation is provided. Thomas Deppisch, Sebastià Vicenc Amengual Garí, Paul Calamia, Jens Ahrens |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Spherical Harmonic Decomposition of a Sound Field Using Microphones on a Circumferential Contour Around a Non-Spherical BaffleabstractSpherical harmonic (SH) representations of sound fields are usually obtained from microphone arrays with rigid spherical baffles whereby the microphones are distributed over the entire surface of the baffle. We present a method that overcomes the requirement for the baffle to be spherical. Furthermore, the microphones can be placed along a circumferential contour around the baffle. This greatly reduces the required number of microphones for a given spatial resolution compared to conventional spherical arrays. Our method is based on the analytical solution for SH decomposition based on observations along the equator of a rigid sphere that we presented recently. It incorporates a calibration stage in which the microphone signals due to a suitable set of calibration sound fields are projected onto the SH decomposition of those same sound fields on the surface of a notional rigid sphere by means of a linear filtering operation. The filter coefficients are computed from the calibration data via a least/squares fit. We present an evaluation of the method based on the application of binaural rendering of the SH decomposition of the signals from an 18/element array that uses a human head as the baffle and that provides 8th ambisonic order. We analyse the accuracy and robustness of our method based on simulated data as well as based on measured data from a prototype. Jens Ahrens, Hannes Helmholz, David L. Alon, Sebastià Vicenc Amengual Garí |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | The Far-Field Equatorial Array for Binaural RenderingabstractWe present a method for obtaining a spherical harmonic representation of a sound field based on a microphone array along the equator of a rigid spherical scatterer. The two-dimensional plane wave de-composition of the incoming sound field is computed from the microphone signals. The influence of the scatterer is removed under the assumption of distant sound sources, and the result is converted to a spherical harmonic (SH) representation, which in turn can be rendered binaurally. The approach requires an order of magnitude fewer microphones compared to conventional spherical arrays that operate at the same SH order at the expense of not being able to accurately represent non-horizontally-propagating sound fields. Although the scattering removal is not perfect at high frequencies at low harmonic orders, numerical evaluation demonstrates the effectiveness of the approach. Jens Ahrens, Hannes Helmholz, David L. Alon, Sebastià Vicenc Amengual Garí |
ICASSP | 4 |
| 2021 | Room Impulse Response Interpolation from a Sparse Set of Measurements Using a Modal ArchitectureabstractIn augmented reality applications, where room geometries and material properties are not readily available, it is desirable to get a representation of the sound field in a room from a limited set of available room impulse response measurements. In this paper, we propose a novel method for 2D interpolation of room modes from a sparse set of RIR measurements that are non-uniformly sampled within a space. We first obtain the mode parameters of a measured room. Using the common-acoustical pole theory, the mode frequencies and decay rates are kept constant over space, and a unique set of mode amplitudes is obtained for each measurement location. Based on the general solution to the Helmholtz equation, these mode amplitudes are modeled as periodic functions of 2D spatial location. For low frequency room modes, the model parameters are found with sequential non-linear least squares. Results show accurate spatial interpolation of perceptually relevant low frequency modes in rooms with simple geometries having non-rigid walls. Orchisama Das, Paul Calamia, Sebastià Vicenc Amengual Garí |
ICASSP | 3 |
| 2021 | Audio-Visual Floorplan ReconstructionabstractGiven only a few glimpses of an environment, how much can we infer about its entire floorplan? Existing methods can map only what is visible or immediately apparent from context, and thus require substantial movements through a space to fully map it. We explore how both audio and visual sensing together can provide rapid floorplan reconstruction from limited viewpoints. Audio not only helps sense geometry outside the camera’s field of view, but it also reveals the existence of distant freespace (e.g., a dog barking in another room) and suggests the presence of rooms not visible to the camera (e.g., a dishwasher humming in what must be the kitchen to the left). We introduce AV-Map, a novel multi-modal encoder-decoder framework that reasons jointly about audio and vision to reconstruct a floorplan from a short input video sequence. We train our model to predict both the interior structure of the environment and the associated rooms’ semantic labels. Our results on 85 large real-world environments show the impact: with just a few glimpses spanning 26% of an area, we can estimate the whole area with 66% accuracy—substantially better than the state of the art approach for extrapolating visual maps. Senthil Purushwalkam, Sebastià Vicenc Amengual Garí, Vamsi K. Ithapu, Carl Schissler, Philip W. Robinson, Abhinav Gupta 0001, Kristen Grauman |
ICCV | 2 |
| 2021 | Effects of Additive Noise in Binaural Rendering of Spherical Microphone Array SignalsabstractAdditive noise produced by the recording hardware will contribute to streamed signals from spherical microphone arrays under practical conditions. For the application of binaural reproduction and under the assumption that the noise is uncorrelated between the array channels, the spectral properties and the overall level of the rendered noise in the ear signals have been shown to be strongly influenced by the configuration of the array as well as of the processing pipeline. In a previous investigation, we determined the audibility thresholds for changes in the rendered noise due to listener head rotations as a function of the differences in noise level of individual array channels. In this article, we calibrate the instrumental metric of Composite Loudness Level to the perceptual data and predict audibility of changes in the additive noise due to head rotations for a broad set of array configurations and distributions of the noise levels across array channels. We demonstrate that some types of microphone layouts can produce audible variations even if the noise level is equal in all channels. This is particularly the case for sampling grids that exhibit negative quadrature weights such as the Lebedev and Fliege-Maier grids for some spherical harmonic orders. The analysis of configurations with unevenly distributed noise contributions show that the influence of the noise from individual array channels is determined by the proximity of their virtual location to the relative trajectory of the ears. Hannes Helmholz, David L. Alon, Sebastià Vicenc Amengual Garí, Jens Ahrens |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | SoundSpaces: Audio-Visual Navigation in 3D Environments
Changan Chen, Unnat Jain, Carl Schissler, Sebastià Vicenc Amengual Garí, Ziad Al-Halah, Vamsi K. Ithapu, Philip W. Robinson, Kristen Grauman |
ECCV (6) | 4 |
| 2020 | Evaluation of Sensor Self-Noise In Binaural Rendering of Spherical Microphone Array SignalsabstractSpherical microphone arrays are used to capture spatial sound fields, which can then be rendered via headphones. We use the Real-Time Spherical Array Renderer (ReTiSAR) to analyze and auralize the propagation of sensor self-noise through the processing pipeline. An instrumental evaluation confirms a strong global influence of different array and rendering parameters on the spectral balance and the overall level of the rendered noise. The character of the noise is direction independent in the case of spatially uniformly distributed noise. However, timbre of the rendered self-noise changes with head orientation in the case of spatially non-uniform noise. We determine audibility thresholds of the coloration artifact during head rotations for different array configurations in a perceptual user study. Hannes Helmholz, Jens Ahrens, David L. Alon, Sebastià Vicenc Amengual Garí, Ravish Mehra |
ICASSP | 4 |