Paul Calamia

dblp:210/6221 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
14since 2021 · last 2025
0000-0002-0401-6996ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 11 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021
YearPublicationVenuePosition
2025 Hearing Anywhere in Any Environment
abstract
In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR) estimation, most existing methods are limited to the single environment on which they are trained, lacking the ability to generalize to new rooms with different geometries and surface materials. We aim to develop a unified model capable of reconstructing the spatial acoustic experience of any environment with minimum additional measurements. To this end, we present xRIR, a framework for cross-room RIR prediction. The core of our generalizable approach lies in combining a geometric feature extractor, which captures spatial context from panorama depth images, with a RIR encoder that extracts detailed acoustic features from only a few reference RIR samples. To evaluate our method, we introduce AcousticRooms, a new dataset featuring high-fidelity simulation of over 300,000 RIRs from 260 rooms. Experiments show that our method strongly outperforms a series of baselines. Furthermore, we successfully perform sim-to-real transfer by evaluating our model on four real-world environments, demonstrating the generalizability of our approach and the realism of our dataset.
Xiulong Liu 0002, Anurag Kumar 0003, Paul Calamia, Sebastià Vicenc Amengual Garí, Calvin Murdock, Ishwarya Ananthabhotla, Philip W. Robinson, Eli Shlizerman, Vamsi K. Ithapu, Ruohan Gao
CVPR3
2023 Chat2Map: Efficient Scene Mapping from Multi-Ego Conversations
abstract
Can conversational videos captured from multiple egocentric viewpoints reveal the map of a scene in a cost-efficient way? We seek to answer this question by proposing a new problem: efficiently building the map of a previously unseen 3D environment by exploiting shared information in the egocentric audio-visual observations of participants in a natural conversation. Our hypothesis is that as multiple people (“egos”) move in a scene and talk among themselves, they receive rich audio-visual cues that can help uncover the unseen areas of the scene. Given the high cost of continuously processing egocentric visual streams, we further explore how to actively coordinate the sampling of visual information, so as to minimize redundancy and reduce power use. To that end, we present an audio-visual deep reinforcement learning approach that works with our shared scene mapper to selectively turn on the camera to efficiently chart out the space. We evaluate the approach using a state-of-the-art audio-visual simulator for 3D scenes as well as real-world video. Our model outperforms previous state-of-the-art mapping methods, and achieves an excellent cost-accuracy tradeoff. Project: http://vision.cs.utexas.edu/projects/chat2map.
Sagnik Majumder, Hao Jiang 0007, Pierre Moulon, Ethan Henderson, Paul Calamia, Kristen Grauman, Vamsi K. Ithapu
CVPR5
2023 Towards Improved Room Impulse Response Estimation for Speech Recognition
abstract
We propose a novel approach for blind room impulse response (RIR) estimation systems in the context of a downstream application scenario, far-field automatic speech recognition (ASR). We first draw the connection between improved RIR estimation and improved ASR performance, as a means of evaluating neural RIR estimators. We then propose a generative adversarial network (GAN) based architecture that encodes RIR features from reverberant speech and constructs an RIR from the encoded features, and uses a novel energy decay relief loss to optimize for capturing energy-based properties of the input reverberant speech. We show that our model outperforms the state-of-the-art baselines on acoustic benchmarks (by 17% on the energy decay relief and 22% on an early-reflection energy metric), as well as in an ASR evaluation task (by 6.9% in word error rate).
Anton Ratnarajah, Ishwarya Ananthabhotla, Vamsi K. Ithapu, Pablo Hoffmann, Dinesh Manocha, Paul Calamia
ICASSP6
2023 Computational modeling of auditory brainstem responses derived from modified speech
Tzu-Han Zoe Cheng, Paul Calamia
INTERSPEECH2
2023 Direct and Residual Subspace Decomposition of Spatial Room Impulse Responses
abstract
Psychoacoustic experiments have shown that directional properties of the direct sound, salient reflections, and the late reverberation of an acoustic room response can have a distinct influence on the auditory perception of a given room. Spatial room impulse responses (SRIRs) capture those properties and thus are used for direction-dependent room acoustic analysis and virtual acoustic rendering. This work proposes a subspace method that decomposes SRIRs into a direct part, which comprises the direct sound and the salient reflections, and a residual, to facilitate enhanced analysis and rendering methods by providing individual access to these components. The proposed method is based on the generalized singular value decomposition and interprets the residual as noise that is to be separated from the other components of the reverberation. Large generalized singular values are attributed to the direct part, which is then obtained as a low-rank approximation of the SRIR. By advancing from the end of the SRIR toward the beginning while iteratively updating the residual estimate, the method adapts to spatio-temporal variations of the residual. The method is evaluated using a spatio-spectral error measure and simulated SRIRs of different rooms, microphone arrays, and ratios of direct sound to residual energy. The proposed method creates lower errors than existing approaches in all tested scenarios, including a scenario with two simultaneous reflections. A case study with measured SRIRs shows the applicability of the method under real-world acoustic conditions. A reference implementation is provided.
Thomas Deppisch, Sebastià Vicenc Amengual Garí, Paul Calamia, Jens Ahrens
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Visual Acoustic Matching
abstract
We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, the goal is to re-synthesize the audio to match the target room acoustics as suggested by its visible geometry and materials. To address this novel task, we propose a cross-modal transformer model that uses audio-visual attention to inject visual properties into the audio and generate realistic audio output. In addition, we devise a self-supervised training objective that can learn acoustic matching from in-the-wild Web videos, despite their lack of acoustically mismatched audio. We demonstrate that our approach successfully translates human speech to a variety of real-world environments depicted in images, outperforming both traditional acoustic matching and more heavily supervised baselines.
Changan Chen, Ruohan Gao, Paul Calamia, Kristen Grauman
CVPR3
2022 TPARN: Triple-Path Attentive Recurrent Network for Time-Domain Multichannel Speech Enhancement
abstract
In this work, we propose a new model called triple-path attentive recurrent network (TPARN) for multichannel speech enhancement in the time domain. TPARN extends a single-channel dual-path network to a multichannel network by adding a third path along the spatial dimension. First, TPARN processes speech signals from all channels independently using a dual-path attentive recurrent network (ARN), which is a recurrent neural network (RNN) augmented with self-attention. Next, an ARN is introduced along the spatial dimension for spatial context aggregation. TPARN is designed as a multiple-input and multiple-output architecture to enhance all input channels simultaneously. Experimental results demonstrate the superiority of TPARN over existing state-of-the-art approaches.
Ashutosh Pandey 0004, Buye Xu, Anurag Kumar 0003, Jacob Donley, Paul Calamia, DeLiang Wang
ICASSP5
2022 Multichannel Speech Enhancement Without Beamforming
abstract
Deep neural networks are often coupled with traditional spatial filters, such as MVDR beamformers for effectively exploiting spatial information. Even though single-stage end-to-end supervised models can obtain impressive enhancement, combining them with a traditional beamformer and a DNN-based post-filter in a multistage processing provides additional improvements. In this work, we propose a two-stage strategy for multi-channel speech enhancement that does not require a traditional beamformer for additional performance. First, we propose a novel attentive dense convolutional network (ADCN) for estimating real and imaginary parts of complex spectrogram. ADCN obtains state-of-the-art results among single-stage models. Next, we use ADCN with a recently proposed triple-path attentive recurrent network (TPARN) for estimating waveform samples. The proposed strategy uses two insights; first, using different approaches in two stages; and second, using a stronger model in the first stage. We illustrate the efficacy of our strategy by evaluating multiple models in a two-stage approach with and without a traditional beamformer.
Ashutosh Pandey 0004, Buye Xu, Anurag Kumar 0003, Jacob Donley, Paul Calamia, DeLiang Wang
ICASSP5
2022 Time-domain Ad-hoc Array Speech Enhancement Using a Triple-path Network
abstract
Deep neural networks (DNNs) are very effective for multichannel speech enhancement with fixed array geometries.However, it is not trivial to use DNNs for ad-hoc arrays with unknown order and placement of microphones.We propose a novel triplepath network for ad-hoc array processing in the time domain.The key idea in the network design is to divide the overall processing into spatial processing and temporal processing and use self-attention for spatial processing.Using self-attention for spatial processing makes the network invariant to the order and the number of microphones.The temporal processing is done independently for all channels using a recently proposed dual-path attentive recurrent network.The proposed network is a multiple-input multiple-output architecture that can simultaneously enhance signals at all microphones.Experimental results demonstrate the excellent performance of the proposed approach.Further, we present analysis to demonstrate the effectiveness of the proposed network in utilizing multichannel information even from microphones at far locations.
Ashutosh Pandey 0004, Buye Xu, Anurag Kumar 0003, Jacob Donley, Paul Calamia, DeLiang Wang
INTERSPEECH5
2022 SAQAM: Spatial Audio Quality Assessment Metric
abstract
Audio quality assessment is critical for assessing the perceptual realism of sounds.However, the time and expense of obtaining "gold standard" human judgments limit the availability of such data.For AR&VR, good perceived sound quality and localizability of sources are among the key elements to ensure complete immersion of the user.Our work introduces SAQAM which uses a multi-task learning framework to assess listening quality (LQ) and spatialization quality (SQ) between any given pair of binaural signals without using any subjective data.We model LQ by training on a simulated dataset of triplet human judgments, and SQ by utilizing activation-level distances from networks trained for direction of arrival (DOA) estimation.We show that SAQAM correlates well with human responses across four diverse datasets.Since it is a deep network, the metric is differentiable, making it suitable as a loss function for other tasks.For example, simply replacing an existing loss with our metric yields improvement in a speech-enhancement network.
Pranay Manocha, Anurag Kumar 0003, Buye Xu, Anjali Menon, Israel D. Gebru, Vamsi K. Ithapu, Paul Calamia
INTERSPEECH7
2022 SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning
abstract
We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from arbitrary microphone locations. Together with existing 3D visual assets, it supports an array of audio-visual research tasks, such as audio-visual navigation, mapping, source localization and separation, and acoustic matching. Compared to existing resources, SoundSpaces 2.0 has the advantages of allowing continuous spatial sampling, generalization to novel environments, and configurable microphone and material properties. To our knowledge, this is the first geometry-based acoustic simulation that offers high fidelity and realism while also being fast enough to use for embodied learning. We showcase the simulator's properties and benchmark its performance against real-world audio measurements. In addition, we demonstrate two downstream tasks---embodied navigation and far-field automatic speech recognition---and highlight sim2real performance for the latter. SoundSpaces 2.0 is publicly available to facilitate wider research for perceptual systems that can both see and hear.
Changan Chen, Carl Schissler, Sanchit Garg, Philip Kobernik, Alexander Clegg, Paul Calamia, Dhruv Batra, Philip W. Robinson, Kristen Grauman
NeurIPS6
2021 Room Impulse Response Interpolation from a Sparse Set of Measurements Using a Modal Architecture
abstract
In augmented reality applications, where room geometries and material properties are not readily available, it is desirable to get a representation of the sound field in a room from a limited set of available room impulse response measurements. In this paper, we propose a novel method for 2D interpolation of room modes from a sparse set of RIR measurements that are non-uniformly sampled within a space. We first obtain the mode parameters of a measured room. Using the common-acoustical pole theory, the mode frequencies and decay rates are kept constant over space, and a unique set of mode amplitudes is obtained for each measurement location. Based on the general solution to the Helmholtz equation, these mode amplitudes are modeled as periodic functions of 2D spatial location. For low frequency room modes, the model parameters are found with sequential non-linear least squares. Results show accurate spatial interpolation of perceptually relevant low frequency modes in rooms with simple geometries having non-rigid walls.
Orchisama Das, Paul Calamia, Sebastià Vicenc Amengual Garí
ICASSP2
2021 Robustness of Acoustic Rake Filters in Minimum Variance Beamforming
abstract
Acoustic rake filters perform coherent summation of the early room reflections using beamforming, with the aim of improving beamforming performance. This concept has been investigated for speech enhancement applications, improving noise reduction and late reverberation attenuation. Current studies typically assume that the parameters of the early reflections, such as the direction-of-arrival, delay and amplitude, are known in advance in the rake filter design. This work presents a novel investigation of the acoustic rake filter in a more practical context, focusing on the minimum variance distortionless response (MVDR) formulation. First, the sensitivity of the filter performance to perturbations in the reflection parameters is derived analytically, and investigated numerically using Monte Carlo simulations with a spherical microphone array. Then, an end-to-end example of rake filtering in a blind scenario is presented, where the reflection parameters are estimated from speech signals without any prior information. This example demonstrates for the first time the use of rake filtering in a realistic scenario.
Ran Weisman, Tom Shlomo, Vladimir Tourbabin, Paul Calamia, Boaz Rafaely
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Fast diffraction pathfinding for dynamic sound propagation
abstract
In the context of geometric acoustic simulation, one of the more perceptually important yet difficult to simulate acoustic effects is diffraction, a phenomenon that allows sound to propagate around obstructions and corners. A significant bottleneck in real-time simulation of diffraction is the enumeration of high-order diffraction propagation paths in scenes with complex geometry (e.g. highly tessellated surfaces). To this end, we present a dynamic geometric diffraction approach that consists of an extensive mesh preprocessing pipeline and complementary runtime algorithm. The preprocessing module identifies a small subset of edges that are important for diffraction using a novel silhouette edge detection heuristic. It also extends these edges with planar diffraction geometry and precomputes a graph data structure encoding the visibility between the edges. The runtime module uses bidirectional path tracing against the diffraction geometry to probabilistically explore potential paths between sources and listeners, then evaluates the intensities for these paths using the Uniform Theory of Diffraction. It uses the edge visibility graph and the A* pathfinding algorithm to robustly and efficiently find additional high-order diffraction paths. We demonstrate how this technique can simulate 10th-order diffraction up to 568 times faster than the previous state of the art, and can efficiently handle large scenes with both high geometric complexity and high numbers of sources.
Carl Schissler, Gregor Mückl, Paul Calamia
ACM Trans. Graph.3
2020 Spatial Covariance Matrix Estimation for Reverberant Speech with Application to Speech Enhancement
Ran Weisman, Vladimir Tourbabin, Paul Calamia, Boaz Rafaely
INTERSPEECH3