Shabnam Ghaffarzadegan

dblp:117/8486 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0001-5784-8908ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Towards Few-Shot Training-Free Anomaly Sound Detection
Ho-Hsiang Wu, Abinaya Kumar, Luca Bondi, Shabnam Ghaffarzadegan, Juan Pablo Bello
INTERSPEECH5
2024 Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting Networks?
abstract
In the design of traffic monitoring solutions for optimizing the urban mobility infrastructure, acoustic vehicle counting models have received attention due to their cost effectiveness and energy efficiency. Although deep learning has proven effective for visual traffic monitoring, its use has not been thoroughly investigated in the audio domain, likely due to real-world data scarcity. In this work, we propose a novel approach to acoustic vehicle counting by developing: i) a traffic noise simulation framework to synthesize realistic vehicle pass-by events; ii) a strategy to mix synthetic and real data to train a deep-learning model for traffic counting. The proposed system is capable of simultaneously counting cars and commercial vehicles driving on a two-lane road, and identifying their direction of travel under moderate traffic density conditions. With only 24 hours of labeled real-world traffic noise, we are able to improve counting accuracy on real-world data from 63% to 88% for cars and from 86% to 94% for commercial vehicles.
Stefano Damiano, Luca Bondi, Shabnam Ghaffarzadegan, Andre Guntoro, Toon van Waterschoot
ICASSP3
2024 CLAP4Emo: ChatGPT-Assisted Speech Emotion Retrieval with Natural Language Supervision
abstract
Speech emotion retrieval is an important technique for large-scale and high-quality data collection. Conventional approach using ensemble of classification models might limit the retrieved emotion diversity and/or underperform in out-of-domain acoustic conditions. Natural language is diverse and agnostic to specific acoustic concepts, embedding a huge potential for developing language-based speech emotion retrieval system. In this paper we introduce CLAP4Emo, a novel framework to retrieve emotional speech via natural language prompts based on contrastive language-audio pretraining. To compensate for the absence of training captions in existing public datasets, we propose a systematic framework that applies ChatGPT to generate emotion captions. The experimental results demonstrate that our method can effectively improve the retrieved sample diversity while maintaining high precision across five benchmark datasets. By leveraging large language models, we establish a connection between audio and language for emotion description, culminating in an intuitive and interactive retrieval system. We release the generated emotion captions at: https://github.com/boschresearch/soundsee-emo-caps
Shabnam Ghaffarzadegan, Luca Bondi, Abinaya Kumar, Samarjit Das, Ho-Hsiang Wu
ICASSP2
2024 Sound of Traffic: A Dataset for Acoustic Traffic Identification and Counting
Shabnam Ghaffarzadegan, Luca Bondi, Abinaya Kumar, Ho-Hsiang Wu, Hans-Georg Horst, Samarjit Das
INTERSPEECH1
2023 Active Learning for Abnormal Lung Sound Data Curation and Detection in Asthma
Shabnam Ghaffarzadegan, Luca Bondi, Ho-Hsiang Wu, Sirajum Munir, Kelly J. Shields, Samarjit Das, Joseph Aracri
INTERSPEECH1
2023 Background Domain Switch: A Novel Data Augmentation Technique for Robust Sound Event Detection
Luca Bondi, Shabnam Ghaffarzadegan
INTERSPEECH3
2021 Unsupervised Discriminative Learning of Sounds for Audio Event Classification
abstract
Recent progress in network-based audio event classification has shown the benefit of pre-training models on visual data such as ImageNet. While this process allows knowledge transfer across different domains, training a model on large-scale visual datasets is time consuming. On several audio event classification benchmarks, we show a fast and effective alternative that pre-trains the model unsupervised, only on audio data and yet delivers on-par performance with ImageNet pre-training. Furthermore, we show that our discriminative audio learning can be used to transfer knowledge across audio datasets and optionally include ImageNet pre-training.
Sascha Hornauer, Kenneth Li 0002, Stella X. Yu, Shabnam Ghaffarzadegan, Liu Ren 0001
ICASSP4
2020 An Ontology-Aware Framework for Audio Event Classification
abstract
Recent advancements in audio event classification often ignore the structure and relation between the label classes available as prior information. This structure can be defined by ontology and augmented in the classifier as a form of domain knowledge. To capture such dependencies between the labels, we propose an ontology-aware neural network containing two components: feed-forward ontology layers and graph convolutional networks (GCN). The feed-forward ontology layers capture the intra-dependencies of labels between different levels of ontology. On the other hand, GCN mainly models interdependency structure of labels within an ontology level. The framework is evaluated on two benchmark datasets for single-label and multi-label audio event classification tasks. The results demonstrate the proposed solutions efficacy to capture and explore the ontology relations and improve the classification performance.
Shabnam Ghaffarzadegan
ICASSP2
2020 Towards Domain Invariant Heart Sound Abnormality Detection Using Learnable Filterbanks
abstract
OBJECTIVE: Cardiac auscultation is the most practiced non-invasive and cost-effective procedure for the early diagnosis of heart diseases. While machine learning based systems can aid in automatically screening patients, the robustness of these systems is affected by numerous factors including the stethoscope/sensor, environment, and data collection protocol. This article studies the adverse effect of domain variability on heart sound abnormality detection and develops strategies to address this problem. METHODS: We propose a novel Convolutional Neural Network (CNN) layer, consisting of time-convolutional (tConv) units, that emulate Finite Impulse Response (FIR) filters. The filter coefficients can be updated via backpropagation and be stacked in the front-end of the network as a learnable filterbank. RESULTS: On publicly available multi-domain datasets, the proposed method surpasses the top-scoring systems found in the literature for heart sound abnormality detection (a binary classification task). We utilized sensitivity, specificity, F-1 score and Macc (average of sensitivity and specificity) as performance metrics. Our systems achieved relative improvements of up to 11.84% in terms of MAcc, compared to state-of-the-art methods. CONCLUSION: The results demonstrate the effectiveness of the proposed learnable filterbank CNN architecture in achieving robustness towards sensor/domain variability in PCG signals. SIGNIFICANCE: The proposed methods pave the way for deploying automated cardiac screening systems in diversified and underserved communities.
Ahmed Imtiaz Humayun, Shabnam Ghaffarzadegan, Md. Istiaq Ansari, Zhe Feng 0003, Taufiq Hasan
IEEE J. Biomed. Health Informatics2
2019 A Real-Time Audio Monitoring Framework with Limited Data for Constrained Devices
abstract
An effective and non-invasive audio monitoring system needs to be capable of simultaneous real-time detection of multiple audio events in many different environments, and locally executable on resource constrained devices, such as, smart microphones. A major challenge in this research domain is having limited available annotated data. This paper presents a novel framework to generate robust detection models of environmental and human audio events with limited available data. The framework presents the generation of a large synthetic dataset using limited data for any audio event, a novel computationally efficient feature modeling technique, named Audio2Vec, that is robust against environmental variations, and identifies and exploits the syntactic relation between audio states represented by the features and the targeted audio events. The presented framework achieves 10.3% higher F-1 scores compared to the best baseline approaches. To demonstrate the effectiveness of the framework we implemented a real-time audio monitoring system that simultaneously detects 10 audio events on a Raspberry Pi 3B and evaluate it in real home and in-car settings, that achieve F-1 scores of 0.96 and 0.956, respectively.
Asif Salekin, Shabnam Ghaffarzadegan, Zhe Feng 0003, John A. Stankovic
DCOSS2
2018 End-to-End Neural Network Based Automated Speech Scoring
abstract
In recent years, machine learning models for automated speech scoring systems were mainly built using data-driven approaches with handcrafted features as one of the main components. However, the remarkable successes of deep learning (DL) technology in a variety of machine learning tasks has demonstrated its effectiveness in extracting features. Although there have been some efforts in utilizing DL technology for the automated speech scoring task, a thorough investigation of learning useful features is still missing. In this paper, we propose an end-to-end solution that consists of using deep neural network models to encode both lexical and acoustical cues to learn predictive features automatically. Experiments also confirm the effectiveness of our proposed solution compared to conventional methods based on handcrafted features.
Lei Chen 0004, Jidong Tao, Shabnam Ghaffarzadegan, Yao Qian
ICASSP3
2018 An Ensemble of Transfer, Semi-supervised and Supervised Learning Methods for Pathological Heart Sound Classification
abstract
In this work, we propose an ensemble of classifiers to distinguish between various degrees of abnormalities of the heart using Phonocardiogram (PCG) signals acquired using digital stethoscopes in a clinical setting, for the INTERSPEECH 2018 Computational Paralinguistics (ComParE) Heart Beats SubChallenge. Our primary classification framework constitutes a convolutional neural network with 1D-CNN time-convolution (tConv) layers, which uses features transferred from a model trained on the 2016 Physionet Heart Sound Database. We also employ a Representation Learning (RL) approach to generate features in an unsupervised manner using Deep Recurrent Autoencoders and use Support Vector Machine (SVM) and Linear Discriminant Analysis (LDA) classifiers. Finally, we utilize an SVM classifier on a high-dimensional segment-level feature extracted using various functionals on short-term acoustic features, i.e., Low-Level Descriptors (LLD). An ensemble of the three different approaches provides a relative improvement of 11.13% compared to our best single sub-system in terms of the Unweighted Average Recall (UAR) performance metric on the evaluation dataset.
Ahmed Imtiaz Humayun, Md. Tauhiduzzaman Khan, Shabnam Ghaffarzadegan, Zhe Feng 0003, Taufiq Hasan
INTERSPEECH3
2017 Occupancy Detection in Commercial and Residential Environments Using Audio Signal
Shabnam Ghaffarzadegan, Attila Reiss, Mirko Ruhs, Robert Dürichen, Zhe Feng 0003
INTERSPEECH1
2017 Acoustic Scene Classification Using a CNN-SuperVector System Trained with Auditory and Spectrogram Image Features
Rakib Hyder, Shabnam Ghaffarzadegan, Zhe Feng 0003, John H. L. Hansen, Taufiq Hasan
INTERSPEECH2
2016 Exploring deep learning architectures for automatically grading non-native spontaneous speech
abstract
We investigate two deep learning architectures reported to have superior performance in ASR over the conventional GMM system, with respect to automatic speech scoring. We use an approximately 800-hour large-vocabulary non-native spontaneous English corpus to build three ASR systems. One system is in GMM, and two are in deep learning architectures - namely, DNN and Tandem with bottleneck features. The evaluation results show that the both deep learning systems significantly outperform the GMM ASR. These ASR systems are used as the front-end in building an automated speech scoring system. To examine the effectiveness of the deep learning ASR systems for automated scoring, another non-native spontaneous speech corpus is used to train and evaluate the scoring models. Using deep learning architectures, ASR accuracies drop significantly on the scoring corpus, whereas the performance of the scoring systems get closer to human raters, and consistently better than the GMM one. Compared to the DNN ASR, the Tandem performs slightly better on the scoring speech while it is a little less accurate on the ASR evaluation dataset. Furthermore, given the results of the improved scoring performance while using fewer scoring features, the Tandem system shows more robustness for scoring task than the DNN one.
Jidong Tao, Shabnam Ghaffarzadegan, Lei Chen 0004, Klaus Zechner
ICASSP2
2016 Generative Modeling of Pseudo-Whisper for Robust Whispered Speech Recognition
abstract
Whisper is a common means of communication used to avoid disturbing individuals or to exchange private information. As a vocal style, whisper would be an ideal candidate for human-handheld/computer interactions in open-office or public area scenarios. Unfortunately, current speech technology is predominantly focused on modal (neutral) speech and completely breaks down when exposed to whisper. One of the major barriers for successful whisper recognition engines is the lack of available large transcribed whispered speech corpora. This study introduces two strategies that require only a small amount of untranscribed whisper samples to produce excessive amounts of whisper-like (pseudo-whisper) utterances from easily accessible modal speech recordings. Once generated, the pseudo-whisper samples are used to adapt modal acoustic models of a speech recognizer toward whisper. The first strategy is based on Vector Taylor Series (VTS) where a whisper “background” model is first trained to capture a rough estimate of global whisper characteristics from a small amount of actual whisper data. Next, that background model is utilized in the VTS to establish specific broad phone classes' (unvoiced/voiced phones) transformations from each input modal utterance to its pseudo-whispered version. The second strategy generates pseudo-whisper samples by means of denoising autoencoders (DAE). Two generative models are investigated-one produces pseudo-whisper cepstral features on a frame-by-frame basis, while the second generates pseudo-whisper statistics for whole phone segments. It is shown that word error rates of a TIMIT-trained speech recognizer are considerably reduced for a whisper recognition task with a constrained lexicon after adapting the acoustic model toward the VTS or DAE pseudo-whisper samples, compared to model adaptation on an available small whisper set.
Shabnam Ghaffarzadegan, Hynek Boril, John H. L. Hansen
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Generative modeling of pseudo-target domain adaptation samples for whispered speech recognition
abstract
The lack of available large corpora of transcribed whispered speech is one of the major roadblocks for development of successful whisper recognition engines. Our recent study has introduced a Vector Taylor Series (VTS) approach to pseudo-whisper sample generation which requires availability of only a small number of real whispered utterances to produce large amounts of whisper-like samples from easily accessible transcribed neutral recordings. The pseudo-whisper samples were found particularly effective in adapting a neutral-trained recognizer to whisper. Our current study explores the use of denoising autoencoders (DAE) for pseudo-whisper sample generation. Two types of generative models are investigated - one which produces pseudo-whispered cepstral vectors on a frame basis and another which generates pseudo-whisper statistics of whole phone segments. It is shown that the DAE approach considerably reduces word error rates of the baseline system as well as the system adapted on real whisper samples. The DAE approach provides competitive results to the VTS-based method while cutting its computational overhead nearly in half.
Shabnam Ghaffarzadegan, Hynek Boril, John H. L. Hansen
ICASSP1
2015 Leveraging automatic speech recognition in cochlear implants for improved speech intelligibility under reverberation
abstract
Despite recent advancements in digital signal processing technology for cochlear implant (CI) devices, there still remains a significant gap between speech identification performance of CI users in reverberation compared to that in anechoic quiet conditions. Alternatively, automatic speech recognition (ASR) systems have seen significant improvements in recent years resulting in robust speech recognition in a variety of adverse environments, including reverberation. In this study, we exploit advancements seen in ASR technology for alternative formulated solutions to benefit CI users. Specifically, an ASR system is developed using multicondition training on speech data with different reverberation characteristics (e.g., T60values), resulting in low word error rates (WER) in reverberant conditions. A speech synthesizer is then utilized to generate speech waveforms from the output of the ASR system, from which the synthesized speech is presented to CI listeners. The effectiveness of this hybrid recognition-synthesis CI strategy is evaluated under moderate to highly reverberant conditions (i.e., T60= 0.3, 0.6, 0.8, and 1.0s) using speech material extracted from the TIMIT corpus. Experimental results confirm the effectiveness of multi-condition training on performance of the ASR system in reverberation, which consequently results in substantial speech intelligibility gains for CI users in reverberant environments.
Oldooz Hazrati Yadkoori, Shabnam Ghaffarzadegan, John H. L. Hansen
ICASSP2
2014 UT-Vocal Effort II: Analysis and constrained-lexicon recognition of whispered speech
abstract
This study focuses on acoustic variations in speech introduced by whispering, and proposes several strategies to improve robustness of automatic speech recognition of whispered speech with neutral-trained acoustic models. In the analysis part, differences in neutral and whispered speech captured in the UT-Vocal Effort II corpus are studied in terms of energy, spectral slope, and formant center frequency and bandwidth distributions in silence, voiced, and unvoiced speech signal segments. In the part dedicated to speech recognition, several strategies involving front-end filter bank redistribution, cepstral dimensionality reduction, and lexicon expansion for alternative pronunciations are proposed. The proposed neutral-trained system employing redistributed filter bank and reduced features provides a 7.7 % absolute WER reduction over the baseline system trained on neutral speech, and a 1.3 % reduction over a baseline system with whisper-adapted acoustic models.
Shabnam Ghaffarzadegan, Hynek Boril, John H. L. Hansen
ICASSP1
2014 Model and feature based compensation for whispered speech recognition
abstract
This study proposes model and feature based strategies for au-tomatic whispered speech recognition. Our goal is to compensate for the mismatch between neutral-trained recognizer models and parameters of whispered speech. We propose a pseudo-whisper generation from neutral speech samples for efficient acoustic model adaptation. The scheme is based on the popular Vector Tay-lor Series (VTS) algorithm. In the first step, a ‘background ’ model capturing a rough estimate of the target whispered speech charac-teristics from a small amount of whispered data is trained. Second, the target background model is utilized in the VTS strategy to es-tablish broad phone classes (consonants and vowels) transforma-tions for individual neutral utterances and transform them towards whisper. Finally, these pseudo-whisper samples are used to adapt neutral recognizer models towards whisper. This approach is eval-uated together with Vocal Tract Length Normalization (VTLN) and Shift frequency transforms and show to greatly benefit recog-nition performance compared to a traditional whisper-adaptation approach. The absolute WER on the closed speakers whisper sce-nario has been reduced from 17.3 % to 8.4 % and the open speakers scenario from 27.7 % to 17.5 %. Index Terms: whispered speech recognition, Vector Taylor Series, vocal length normalization
Shabnam Ghaffarzadegan, Hynek Boril, John H. L. Hansen
INTERSPEECH1
2012 Efficient Frequency Domain Implementation of Noncausal Multichannel Blind Deconvolution for Convolutive Mixtures of Speech
abstract
Multichannel blind deconvolution (MCBD) algorithms are known to suffer from an extensive computational complexity problem, which makes them impractical for blind source separation (BSS) of speech and audio signals. This problem is even more serious with noncausal MCBD algorithms that must be used in many frequently occurring BSS setups. In this paper, we propose a novel frequency domain algorithm for the efficient implementation of noncausal multichannel blind deconvolution. A block-wise formulation is first developed for filtering and adaptation of filter coefficients. Based on this formulation, we present a modified overlap-save procedure for noncausal filtering in the frequency domain. We also derive update equations for training both causal and anti-causal filters in the frequency domain. Our evaluations indicate that the proposed frequency domain implementation reduces the computational requirements of the algorithm by a factor of more than 100 for typical filter lengths used in blind speech separation. The algorithm is employed successfully for the separation of speech mixtures in a reverberant room. Simulation results demonstrate the superior performance of the proposed algorithm over causal MCBD algorithms in many potential source and microphone positions. It is shown that in BSS problems, causal MCBD algorithms with center-spike initialization do not always converge to a delayed form of the desired noncausal solution, further revealing the need for an efficient noncausal MCBD algorithm.
Seyedmahdad Mirsamadi, Shabnam Ghaffarzadegan, Hamid Sheikhzadeh, Seyed Mohammad Ahadi, Amir Hossein Rezaie
IEEE Trans. Speech Audio Process.2