EDBT 2026 Demo / reviewers in the wild / expert
Hannes Gamper
dblp:42/9856
· DBLP profile ↗
38ranked-venue papers
8as first author
21since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 5 first-author · 18 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Audio Entailment: Assessing Deductive Reasoning for Audio UnderstandingabstractRecent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable performance in tasks including Text-to-Audio Retrieval, Captioning, and Question Answering. However, their ability to engage in more complex open-ended tasks, like Interactive Question-Answering, requires proficiency in logical reasoning- a skill not yet benchmarked. We introduce the novel task of Audio Entailment to evaluate an ALM's deductive reasoning ability. This task assesses whether a text description (hypothesis) of audio content can be deduced from an audio recording (premise), with potential conclusions being entailment, neutral, or contradiction, depending on the sufficiency of the evidence. We create two datasets for this task with audio recordings sourced from two audio captioning datasets-AudioCaps and Clotho-and hypotheses generated using Large Language Models (LLMs). We benchmark state-of-the-art ALMs and find deficiencies in logical reasoning with both zero-shot and linear probe evaluations. Finally, we propose "caption-before-reason", an intermediate step of captioning that improves the Zero-Shot and linear-probe performance of ALMs by an absolute 6% and 3%, respectively. Soham Deshmukh, Hazim T. Bukhari, Benjamin Elizalde, Hannes Gamper, Rita Singh, Bhiksha Raj |
AAAI | 5 |
| 2025 | Make Some Noise: Towards LLM audio reasoning and generation using sound tokensabstractIntegrating audio comprehension and generation into large language models (LLMs) remains challenging due to the continuous nature of audio and the resulting high sampling rates. Here, we introduce a novel approach that combines Variational Quantization with Conditional Flow Matching to convert audio into ultra-low bitrate discrete tokens of 0.23kpbs, allowing for seamless integration with text tokens in LLMs. We fine-tuned a pretrained text-based LLM using Low-Rank Adaptation (LoRA) to assess its effectiveness in achieving true multimodal capabilities, i.e., audio comprehension and generation. Our tokenizer outperforms a traditional VQ-VAE across various datasets with diverse acoustic events. Despite the substantial loss of fine-grained details through audio tokenization, our multimodal LLM trained with discrete tokens achieves competitive results in audio comprehension with state-of-the-art methods, though audio generation is poor. Our results highlight the need for larger, more diverse datasets and improved evaluation metrics to advance multimodal LLM performance. Shivam Mehta, Nebojsa Jojic, Hannes Gamper |
ICASSP | 3 |
| 2025 | Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality AssessmentabstractIn this paper, we investigate distillation and pruning methods to reduce model size for non-intrusive speech quality assessment based on self-supervised representations. Our experiments build on XLS-R-SQA, a speech quality assessment model using wav2vec 2.0 XLS-R embeddings. We retrain this model on a large compilation of mean opinion score datasets, encompassing over 100,000 labeled clips. For distillation, using this model as a teacher, we generate pseudo-labels on unlabeled degraded speech signals and train student models of varying sizes. For pruning, we use a data-driven strategy. While data-driven pruning performs better at larger model sizes, distillation on unlabeled data is more effective for smaller model sizes. Distillation can halve the gap between the baseline’s correlation with ground-truth MOS labels and that of the XLS-R-based teacher model, while reducing model size by two orders of magnitude compared to the teacher model. Benjamin Stahl, Hannes Gamper |
ICASSP | 2 |
| 2025 | Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio DistanceabstractThe complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a study on Music Emotion Recognition (MER) and Emotional Music Generation (EMG), employing diverse audio encoders alongside Frechet Audio Distance (FAD), a reference-free evaluation metric. Our study begins with a benchmark evaluation of MER, highlighting the limitations of using a single audio encoder and the disparities observed across different measurements. We then propose assessing MER performance using FAD derived from multiple encoders to provide a more objective measure of musical emotion. Furthermore, we introduce an enhanced EMG approach designed to improve both the variability and prominence of generated musical emotion, thereby enhancing its realism. Additionally, we investigate the differences in realism between the emotions conveyed in real and synthetic music, comparing our EMG model against two baseline models. Experimental results underscore the issue of emotion bias in both MER and EMG and demonstrate the potential of using FAD and diverse audio encoders to evaluate musical emotion more objectively and effectively. Yuanchao Li, Azalea Gui, Dimitra Emmanouilidou, Hannes Gamper |
ICME | 4 |
| 2024 | Adapting Frechet Audio Distance for Generative Music EvaluationabstractThe growing popularity of generative music models underlines the need for perceptually relevant, objective music quality metrics. The Frechet Audio Distance (FAD) is commonly used for this purpose even though its correlation with perceptual quality is understudied. We show that FAD performance may be hampered by sample size bias, poor choice of audio embeddings, or the use of biased or low-quality reference sets. We propose reducing sample size bias by extrapolating scores towards an infinite sample size. Through comparisons with MusicCaps labels and a listening test we identify audio embeddings and music reference sets that yield FAD scores well-correlated with acoustic and musical quality. Our results suggest that per-song FAD can be useful to identify outlier samples and predict perceptual quality for a range of music sets and generative models. Finally, we release a toolkit that allows adapting FAD for generative music evaluation. Azalea Gui, Hannes Gamper, Sebastian Braun, Dimitra Emmanouilidou |
ICASSP | 2 |
| 2024 | PAM: Prompting Audio-Language Models for Audio Quality Assessment
Soham Deshmukh, Dareen Alharthi, Benjamin Elizalde, Hannes Gamper, Mahmoud Al Ismail, Rita Singh, Bhiksha Raj, Huaming Wang |
INTERSPEECH | 4 |
| 2023 | Spatialized Audio and Hybrid Video Conferencing: Where Should Voices be Positioned for People in the Room and Remote Headset Users?abstractHybrid video calls include attendees in a conference room with loudspeakers and remote attendees using headsets, each with different options for rendering sound spatially. Two studies explored the listener experience with spatial audio in video calls. One study examined the in-room experience using loudspeakers, comparing among spatialization algorithms spreading voices out horizontally. A second study compared varying degrees of horizontal separation of binaurally rendered voices for a remote participant using a headset. In-room participants preferred the widest spatialization over monophonic, stereo, and stereo-binary audio in metrics related to intelligibility and helpfulness. Remote participants preferred different widths of the audio stage depending on the number of voices. In both studies, rendering sound spatially increased performance in speech stream identification. Results indicate spatial audio benefits for in-room and remote attendees in video calls, although the in-room attendees accepted a wider audio stage than remote users. Jeremy Hyrkas, Andrew D. Wilson, John C. Tang, Hannes Gamper, Hong Sodoma, Lev Tankelevitch, Kori Inkpen, Shreya Chappidi, Brennan Jones |
CHI | 4 |
| 2023 | Speech MOS Multi-Task Learning and Rater Bias CorrectionabstractPerceptual speech quality is an important performance metric for teleconferencing applications. The mean opinion score (MOS) is standardized for the perceptual evaluation of speech quality and is obtained by asking listeners to rate the quality of a speech sample. Recently, there has been increasing research interest in developing models for estimating MOS blindly. Here we propose a multitask framework to include additional labels and data in training to improve the performance of a blind MOS estimation model. Experimental results indicate that the proposed model can be trained to jointly estimate MOS, reverberation time (T60), and clarity (C50) by combining two disjoint data sets in training, one containing only MOS labels and the other containing only T60 and C50 labels. Furthermore, we use a semi-supervised framework to combine two MOS data sets in training, one containing only MOS labels (per ITU-T Recommendation P.808), and the other containing separate scores for speech signal, background noise, and overall quality (per ITU-T Recommendation P.835). Finally, we present preliminary results for addressing individual rater bias in the MOS labels. Haleh Akrami, Hannes Gamper |
ICASSP | 2 |
| 2022 | Effect of Noise Suppression Losses on Speech Distortion and ASR PerformanceabstractDeep learning based speech enhancement has made rapid development towards improving quality, while models are becoming more compact and usable for real-time on-the-edge inference. However, the speech quality scales directly with the model size, and small models are often still unable to achieve sufficient quality. Furthermore, the introduced speech distortion and artifacts greatly harm speech quality and intelligibility, and often significantly degrade automatic speech recognition (ASR) rates. In this work, we shed light on the success of the spectral complex compressed mean squared error (MSE) loss, and how its magnitude and phase-aware terms are related to the speech distortion vs. noise reduction trade off. We further investigate integrating pre-trained reference-less predictors for mean opinion score (MOS) and word error rate (WER), and pre-trained embeddings on ASR and sound event detection. Our analyses reveal that none of the pre-trained networks added significant performance over the strong spectral loss. Sebastian Braun, Hannes Gamper |
ICASSP | 2 |
| 2022 | ICASSP 2022 Acoustic Echo Cancellation ChallengeabstractThe ICASSP 2022 Acoustic Echo Cancellation Challenge is intended to stimulate research in acoustic echo cancellation (AEC), which is an important area of speech enhancement and still a top issue in audio communication. This is the third AEC challenge and it is enhanced by including mobile scenarios, adding speech recognition word accuracy rate as a metric, and making the audio 48 kHz. We open source two large datasets to train AEC models under both single talk and double talk scenarios. These datasets consist of recordings from more than 10,000 real audio devices and human speakers in real environments, as well as a synthetic dataset. We open source an online subjective test framework and provide an online objective metric service for researchers to quickly test their results. The winners of this challenge were selected based on the average Mean Opinion Score (MOS) achieved across all scenarios and the word accuracy rate. Ross Cutler, Ando Saabas, Tanel Pärnamaa, Marju Purin, Hannes Gamper, Sebastian Braun, Karsten Sørensen, Robert Aichner |
ICASSP | 5 |
| 2022 | Icassp 2022 Deep Noise Suppression ChallengeabstractThe Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. This is the 4th DNS challenge, with the previous editions held at INTERSPEECH 2020 [1], ICASSP 2021 [2], and INTERSPEECH 2021 [3]. We open-source datasets and test sets for researchers to train their deep noise suppression models, as well as a subjective evaluation framework based on ITU-T P.835 to rate and rank-order the challenge entries. We provide access to DNS-MOS P.835 and word accuracy (WAcc) APIs to challenge participants to help with iterative model improvements. In this challenge, we introduced the following changes: (i) Included mobile device scenarios in the blind test set; (ii) Included a personalized noise suppression track with baseline; (iii) Added WAcc as an objective metric; (iv) Included DNSMOS P.835; (v) Made the training datasets and test sets fullband (48 kHz). We use an average of WAcc and subjective scores P.835 SIG, BAK, and OVRL to get the final score for ranking the DNS models. We believe that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-world scenarios. Harishchandra Dubey, Vishak Gopal, Ross Cutler, Ashkan Aazami, Sergiy Matusevych, Sebastian Braun, Sefik Emre Eskimez, Manthan Thakker, Takuya Yoshioka, Hannes Gamper, Robert Aichner |
ICASSP | 10 |
| 2022 | Predicting label distribution improves non-intrusive speech quality estimation
Abu Zaher Md Faridee, Hannes Gamper |
INTERSPEECH | 2 |
| 2022 | Adaptive Radiation Survey Using an Autonomous Robot Executing LiDAR Scans in the Large Hadron Collider
Hannes Gamper, David Forkel, Alejandro Díaz Rosales, Jorge Playán Garai, Carlos Veiga Almagro, Luca Rosario Buonocore, Eloise Matheson, Mario Di Castro |
ISRR | 1 |
| 2021 | Decoding Music Attention from "EEG Headphones": A User-Friendly Auditory Brain-Computer InterfaceabstractPeople enjoy listening to music as part of their life. This makes music an excellent choice for designing a user-friendly brain-computer interface (BCI) for long-term use. We propose a novel BCI system using music stimuli that relies on brain signals collected via Smartfones, an EEG recording device integrated into a pair of headphones. In a user study of the proposed system, participants were asked to pay attention to one of three musical instruments playing simultaneously from separate spatial directions. We used a stimulus reconstruction method to decode attention from EEG signals. Results show that the proposed system can achieve good decoding accuracy (>70%) while providing superior user-friendliness compared to a traditional EEG setup. Winko W. An, Barbara G. Shinn-Cunningham, Hannes Gamper, Dimitra Emmanouilidou, David Johnston, Mihai Jalobeanu, Edward Cutrell, Andrew D. Wilson, Kuan-Jung Chiang, Ivan Tashev |
ICASSP | 3 |
| 2021 | Towards Efficient Models for Real-Time Deep Noise SuppressionabstractWith recent research advancements, deep learning models are be-coming attractive and powerful choices for speech enhancement in real-time applications. While state-of-the-art models can achieve outstanding results in terms of speech quality and background noise reduction, the main challenge is to obtain compact enough models, which are resource efficient during inference time. An important but often neglected aspect for data-driven methods is that results can be only convincing when tested on real-world data and evaluated with useful metrics. In this work, we investigate reasonably small recurrent and convolutional-recurrent network architectures for speech enhancement, trained on a large dataset considering also reverberation. We show interesting tradeoffs between computational complexity and the achievable speech quality, measured on real recordings using a highly accurate MOS estimator. It is shown that the achievable speech quality is a function of network complexity, and show which models have better tradeoffs. Sebastian Braun, Hannes Gamper, Chandan K. A. Reddy, Ivan Tashev |
ICASSP | 2 |
| 2021 | ICASSP 2021 Deep Noise Suppression ChallengeabstractThe Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH 2020 where we open-sourced training and test datasets for researchers to train their noise suppression models. We also open-sourced a subjective evaluation framework and used the tool to evaluate and select the final winners. Many researchers from academia and industry made significant contributions to push the field forward. We also learned that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-time conditions. In this challenge, we expanded both our training and test datasets. Clean speech in the training set has increased by 200% with the addition of singing voice, emotion data, and non-English languages. The test set has increased by 100% with the addition of singing, emotional, non-English (tonal and non-tonal) languages, and, personalized DNS test clips. There are two tracks with focus on (i) real-time denoising, and (ii) real-time personalized DNS. We present the challenge results at the end. Chandan K. A. Reddy, Harishchandra Dubey, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan 0003 |
ICASSP | 6 |
| 2021 | ICASSP 2021 Acoustic Echo Cancellation Challenge: Datasets, Testing Framework, and ResultsabstractThe ICASSP 2021 Acoustic Echo Cancellation Challenge is intended to stimulate research in the area of acoustic echo cancellation (AEC), which is an important part of speech enhancement and still a top issue in audio communication and conferencing systems. Many recent AEC studies report good performance on synthetic datasets where the train and test samples come from the same underlying distribution. However, the AEC performance often degrades significantly on real recordings. Also, most of the conventional objective metrics such as echo return loss enhancement (ERLE) and perceptual evaluation of speech quality (PESQ) do not correlate well with subjective speech quality tests in the presence of background noise and reverberation found in realistic environments. In this challenge, we open source two large datasets to train AEC models under both single talk and double talk scenarios. These datasets consist of recordings from more than 2,500 real audio devices and human speakers in real environments, as well as a synthetic dataset. We open source two large test sets, and we open source an online subjective test framework for researchers to quickly test their results. The winners of this challenge will be selected based on the average Mean Opinion Score (MOS) achieved across all different single talk and double talk scenarios. Kusha Sridhar, Ross Cutler, Ando Saabas, Tanel Pärnamaa, Markus Loide, Hannes Gamper, Sebastian Braun, Robert Aichner, Sriram Srinivasan 0003 |
ICASSP | 6 |
| 2021 | Design Optimization of a Manipulator for CERN's Future Circular Collider (FCC)
Hannes Gamper, Hubert Gattringer, Andreas Müller 0002, Mario Di Castro |
ICINCO | 1 |
| 2021 | INTERSPEECH 2021 Acoustic Echo Cancellation Challenge
Ross Cutler, Ando Saabas, Tanel Pärnamaa, Markus Loide, Sten Sootla, Marju Purin, Hannes Gamper, Sebastian Braun, Karsten Sørensen, Robert Aichner, Sriram Srinivasan 0003 |
Interspeech | 7 |
| 2021 | INTERSPEECH 2021 Deep Noise Suppression ChallengeabstractThe Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH and ICASSP 2020. We open-sourced training and test datasets for the wideband scenario. We also open-sourced a subjective evaluation framework based on ITU-T standard P.808, which was also used to evaluate participants of the challenge. Many researchers from academia and industry made significant contributions to push the field forward, yet even the best noise suppressor was far from achieving superior speech quality in challenging scenarios. In this version of the challenge organized at INTERSPEECH 2021, we are expanding both our training and test datasets to accommodate full band scenarios. The two tracks in this challenge will focus on real-time denoising for (i) wide band, and(ii) full band scenarios. We are also making available a reliable non-intrusive objective speech quality metric called DNSMOS for the participants to use during their development phase. Chandan K. A. Reddy, Harishchandra Dubey, Kazuhito Koishida, Arun Asokan Nair, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan 0003 |
Interspeech | 8 |
| 2021 | A Fast Forest Reverberator Using Single Scattering CylindersabstractSimulating forest acoustics has important applications for rendering forest sound scenes in mixed and virtual reality, developing wildlife monitoring systems that use microphone arrays distributed in a forest, or as an artistic effect. Previously proposed methods for forest impulse response (IR) synthesis are limited to small or sparse forests because of their cubic asymptotic complexity with respect to the number of trees. Here we propose a simple and efficient parametric forest IR generation algorithm that relies on a multitude of single scattering cylinders to approximate scattering caused by tree trunks. The proposed method was compared to measured forest IRs in terms of the IR echo density, energy decay, reverberation time (T60), and clarity (C50). Experimental results indicate that the proposed algorithm generates forest reverb with acoustic characteristics similar to real forest IRs at a low computational cost. Shoken Kaneko, Hannes Gamper |
MMSP | 2 |
| 2020 | Fast Acoustic Scattering Using Convolutional Neural NetworksabstractDiffracted scattering and occlusion are important acoustic effects in interactive auralization and noise control applications, typically requiring expensive numerical simulation. We propose training a convolutional neural network to map from a convex scatterer's cross-section to a 2D slice of the resulting spatial loudness distribution. We show that employing a full-resolution residual network for the resulting image-to-image regression problem yields spatially detailed loudness fields with a root-mean-squared error of less than 1 dB, at over 100x speedup compared to full wave simulation. Ziqi Fan, Vibhav Vineet, Hannes Gamper, Nikunj Raghuvanshi |
ICASSP | 3 |
| 2020 | Predicting Word Error Rate for Reverberant SpeechabstractReverberation negatively impacts the performance of automatic speech recognition (ASR). Prior work on quantifying the effect of reverberation has shown that clarity (C50), a parameter that can be estimated from the acoustic impulse response, is correlated with ASR performance. In this paper we propose predicting ASR performance in terms of the word error rate (WER) directly from acoustic parameters via a polynomial, sigmoidal, or neural network fit, as well as blindly from reverberant speech samples using a convolutional neural network (CNN). We carry out experiments on two state-of-the-art ASR models and a large set of acoustic impulse responses (AIRs). The results confirm C50 and C80 to be highly correlated with WER, allowing WER to be predicted with the proposed fitting approaches. The proposed non-intrusive CNN model outperforms C50-based WER prediction, indicating that WER can be estimated blindly, i.e., directly from the reverberant speech samples without knowledge of the acoustic parameters. Hannes Gamper, Dimitra Emmanouilidou, Sebastian Braun, Ivan Tashev |
ICASSP | 1 |
| 2020 | Blind C50 estimation from single-channel speech using a convolutional neural networkabstractThe early-to-late reverberation energy ratio is an important parameter describing the acoustic properties of an environment. C50, i.e., the ratio between the first 50 ms and the remaining late energy, affects the perceived clarity and intelligibility of speech, and can be used as a design parameter in mixed reality applications or to predict the performance of speech recognition systems. While established methods exist to derive C50 from impulse response measurements, such measurements are rarely available in practice. Recently, methods have been proposed to estimate C50 blindly from reverberant speech signals. Here, a convolutional neural network (CNN) architecture with a long short-term memory (LSTM) layer is proposed to estimate C50 blindly. The CNN-LSTM operates directly on the spectrogram of variable-length, noisy, reverberant utterances. A feature comparison indicates that log Mel spectrogram features with a frame size of 128 samples achieve the best performance with an average root-mean-square error of about 2.7 dB, outperforming previously proposed blind C50 estimators. Hannes Gamper |
MMSP | 1 |
| 2019 | Non-intrusive Speech Quality Assessment Using Neural NetworksabstractEstimating the perceived quality of an audio signal is critical for many multimedia and audio processing systems. Providers strive to offer optimal and reliable services in order to increase the user quality of experience (QoE). In this work, we present an investigation of the applicability of neural networks for non-intrusive audio quality assessment. We propose three neural network-based approaches for mean opinion score (MOS) estimation. We compare our results to three instrumental measures: the perceptual evaluation of speech quality (PESQ), the ITU-T Recommendation P.563, and the speech-to-reverberation energy ratio. Our evaluation uses a speech dataset contaminated with convolutive and additive noise, labeled using a crowd-based QoE evaluation, evaluated with Pearson correlation with MOS labels, and mean-squared-error of the estimated MOS. Our proposed approaches outperform the aforementioned instrumental measures, with a fully connected deep neural network using Mel-frequency features providing the best correlation (0.87) and the lowest mean squared error (0.15). Anderson R. Avila, Hannes Gamper, Chandan K. A. Reddy, Ross Cutler, Ivan Tashev, Johannes Gehrke |
ICASSP | 2 |
| 2019 | Blind Room Volume Estimation from Single-channel Noisy SpeechabstractRecent work on acoustic parameter estimation indicates that geometric room volume can be useful for modeling the character of an acoustic environment. However, estimating volume from audio signals remains a challenging problem. Here we propose using a convolutional neural network model to estimate the room volume blindly from reverberant single-channel speech signals in the presence of noise. The model is shown to produce estimates within approximately a factor of two to the true value, for rooms ranging in size from small offices to large concert halls. Andrea F. Genovese, Hannes Gamper, Ville Pulkki, Nikunj Raghuvanshi, Ivan Tashev |
ICASSP | 2 |
| 2019 | Improving Binaural Ambisonics Decoding by Spherical Harmonics Domain Tapering and Coloration CompensationabstractA powerful and flexible approach to record or encode a spatial sound scene is through spherical harmonics (SHs), or Ambisonics. An SH-encoded scene can be rendered binaurally by applying SH-encoded head-related transfer functions (HRTFs). Limitations of the recording equipment or computational constraints dictate the spatial reproduction accuracy, thus rendering might suffer from spatial degradation as well as coloration. This paper studies the effect of tapering the SH representation of a binaurally rendered sound field in conjunction with its spectral equalization. The proposed approach is shown to reduce coloration and thus improves perceived audio quality. Christoph Hold, Hannes Gamper, Ville Pulkki, Nikunj Raghuvanshi, Ivan Tashev |
ICASSP | 2 |
| 2019 | A Sparsity Measure for Echo Density Growth in General EnvironmentsabstractWe study the detailed temporal evolution of echo density in impulse responses for applications in acoustic analysis and rendering on general environments. For this purpose, we propose a smooth sorted density measure that yields an intuitive trend of echo density growth with time. This is fitted with a general power-law model motivated from theoretical considerations. We validate the framework against theory on simple room geometries and present experiments on measured and numerically simulated impulse responses in complex scenes. Our results show that the growth power of echo density is a promising statistical parameter that shows noticeable, consistent differences between indoor and outdoor responses, meriting further study. Helena Peic Tukuljac, Ville Pulkki, Hannes Gamper, Keith W. Godin, Ivan Tashev, Nikunj Raghuvanshi |
ICASSP | 3 |
| 2018 | Spatial Audio Feature Discovery with Convolutional Neural NetworksabstractThe advent of mixed reality consumer products brings about a pressing need to develop and improve spatial sound rendering techniques for a broad user base. Despite a large body of prior work, the precise nature and importance of various sound localization cues and how they should be personalized for an individual user to improve localization performance is still an open research problem. Here we propose training a convolutional neural network (CNN) to classify the elevation angle of spatially rendered sounds and employing Layer-wise Relevance Propagation (LRP) on the trained CNN model. LRP provides saliency maps that can be used to identify spectral features used by the network for classification. These maps, in addition to the convolution filters learned by the CNN, are discussed in the context of listening tests reported in the literature. The proposed approach could potentially provide an avenue for future studies on modeling and personalization of head-related transfer functions (HRTFs). Etienne Thuillier, Hannes Gamper, Ivan Tashev |
ICASSP | 2 |
| 2017 | Interaural time delay personalisation using incomplete head scansabstractWhen using a set of generic head-related transfer functions (HRTFs) for spatial sound rendering, personalisation can be considered to minimise localisation errors. This typically involves tuning the characteristics of the HRTFs or a parametric model according to the listener's anthropometry. However, measuring anthropometric features directly remains a challenge in practical applications, and the mapping between anthropometric and acoustic features is an open research problem. Here we propose matching a face template to a listener's head scan or depth image to extract anthropometric information. The deformation of the template is used to personalise the interaural time differences (ITDs) of a generic HRTF set. The proposed method is shown to outperform reference methods when used with high-resolution 3-D scans. Experiments with single-frame depth images indicate that the method is applicable to lower resolution or partial scans which are quicker and easier to obtain than full 3-D scans. These results suggest that the proposed method may be a viable option for ITD personalisation in practical applications. Hannes Gamper, David Johnston, Ivan Tashev |
ICASSP | 1 |
| 2016 | Applications of 3D spherical transforms to personalization of head-related transfer functionsabstractHead-related transfer functions (HRTFs) depend on the shape of the human head and ears, motivating HRTF personalization methods that detect and exploit morphological similarities between subjects in an HRTF database and a new user. Prior work determined similarity from sets of morphological parameters. Here we propose a non-parametric morphological similarity based on a harmonic expansion of head scans. Two 3D spherical transforms are explored for this task, and an appropriate shape similarity metric is defined. A case study focusing on personalisation of interaural time differences (ITDs) is conducted by applying this similarity metric on a database of 3D head scans. Archontis Politis, Mark R. P. Thomas, Hannes Gamper, Ivan Tashev |
ICASSP | 3 |
| 2016 | BFGUI: An interactive tool for the synthesis and analysis of microphone array beamformersabstractMicrophone arrays are beneficial for distant speech capture because the signals they capture can be exploited with beamforming to suppress noise and reverberation. The theory for the design and analysis of microphone arrays is well established, however the performance of a microphone array beamformer is often subject to conflicting criteria that need to be assessed manually. This paper describes BFGUI, a interactive graphical tool for MATLAB, for simulating microphone arrays and synthesizing beamformers, and whose parameters can be modified and performance metrics monitored in real-time. Primarily aimed at teaching and research, this tool provides the user with an intuitive insight into the effects of microphone types, number and geometry, and the influence of design constraints such as regularization and white noise gain on derived metrics. The resulting directivity pattern, directivity index and front-back ratio are examples of such metrics. Multiple analytic microphone models are supported and external measured microphone directivity patterns can also be loaded. The designs can be then exported in a variety of formats for processing of real-world data. Mark R. P. Thomas, Hannes Gamper, Ivan Tashev |
ICASSP | 2 |
| 2016 | Synthesis of Device-Independent Noise Corpora for Realistic ASR EvaluationabstractIn order to effectively evaluate the accuracy of automatic speech recognition (ASR) with a novel capture device, it is important to create a realistic test data corpus that is representative of real-world noise conditions. Typically, this involves either recording the output of a device under test (DUT) in a noisy environment, or synthesizing an environment over loudspeakers in a way that simulates realistic signal-to-noise ratios (SNRs), reverberation times, and spatial noise distributions. Here we propose a method that aims at combining the realism of in-situ recordings with the convenience and repeatability of synthetic corpora. A device-independent spatial recording containing noise and speech is combined with the measured directivity pattern of a DUT to generate a synthetic test corpus for evaluating the performance of an ASR system. This is achieved by a spherical harmonic decomposition of both the sound field and the DUT’s directivity patterns. Experimental results suggest that the proposed method can be a viable alternative to costly and cumbersome device-dependent measurements. The proposed simulation method predicted the SNR of the DUT response to within about 3 dB and the word error rate (WER) to within about 20%, across a range of test SNRs, target source directions, and noise types. Hannes Gamper, Mark R. P. Thomas, Lyle Corbin, Ivan Tashev |
INTERSPEECH | 1 |
| 2016 | Discussion
Dayana Ribas González, Emmanuel Vincent 0001, John H. L. Hansen, Emma Jokinen, Mirco Ravanelli, Hannes Gamper, Fred Richardson |
INTERSPEECH | 6 |
| 2015 | Estimation of multipath propagation delays and interaural time differences from 3-D head scansabstractThe estimation of acoustic propagation delays from a sound source to a listener's ear entrances is useful for understanding and visualising the wave propagation along the surface of the head, and necessary for individualised spatial sound rendering. The interaural time difference (ITD) is of particular research interest, as it constitutes one of the main localisation cues exploited by the human auditory system. Here, an approach is proposed that employs ray tracing on a 3-D head scan to estimate and visualise the propagation delays and ITDs from a sound source to a subject's ear entrances. Experimental results indicate that the proposed approach is computationally efficient, and performs equally well or better than optimally tuned parametric ITD models, with a mean absolute ITD estimation error of about 14μs. Hannes Gamper, Mark R. P. Thomas, Ivan Tashev |
ICASSP | 1 |
| 2015 | Dereverberation sweet spot dilation with combined channel equalization and beamformingabstractBeamforming and channel equalizers can be formulated as optimal multichannel filter-and-sum operations with different objective criteria. It has been shown in previous studies that the combination of both concepts under a common framework can yield results that combine both the spatial robustness of beamforming and the dereverberation performance of channel equalization. This paper introduces an additional method for leveraging both approaches that exploits channel estimates in a wanted spatial location and derives robustness from knowledge of the array geometry alone. Experiments with an objective assessment of speech quality as a function of source perturbation reveal that the proposed technique can be viewed as a sweet spot dilator when compared with the MINT channel equalizer. Mark R. P. Thomas, Hannes Gamper, Ivan Tashev |
ICASSP | 2 |
| 2013 | Sound sample detection and numerosity estimation using auditory displayabstractThis article investigates the effect of various design parameters of auditory information display on user performance in two basic information retrieval tasks. We conducted a user test with 22 participants in which sets of sound samples were presented. In the first task, the test participants were asked to detect a given sample among a set of samples. In the second task, the test participants were asked to estimate the relative number of instances of a given sample in two sets of samples. We found that the stimulus onset asynchrony (SOA) of the sound samples had a significant effect on user performance in both tasks. For the sample detection task, the average error rate was about 10% with an SOA of 100 ms. For the numerosity estimation task, an SOA of at least 200 ms was necessary to yield average error rates lower than 30%. Other parameters, including the samples' sound type (synthesized speech or earcons) and spatial quality (multichannel loudspeaker or diotic headphone playback), had no substantial effect on user performance. These results suggest that diotic, or indeed monophonic, playback with appropriately chosen SOA may be sufficient in practical applications for users to perform the given information retrieval tasks, if information about the sample location is not relevant. If location information was provided through spatial playback of the samples, test subjects were able to simultaneously detect and localize a sample with reasonable accuracy. Hannes Gamper, Christina Dicke, Mark Billinghurst, Kai Puolamäki |
ACM Trans. Appl. Percept. | 1 |
| 2011 | Analyzing Emotional Semantics of Abstract Art Using Low-Level Image Features
He Zhang 0009, Eimontas Augilius, Timo Honkela, Jorma Laaksonen, Hannes Gamper, Henok Alene |
IDA | 5 |