Giampiero Salvi

dblp:39/4545 · DBLP profile ↗
← Back
48ranked-venue papers
8as first author
16since 2021 · last 2025
0000-0002-3323-5311ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 6 first-author · 11 since 2021Systems, architecture and hardware · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 Effects of Prosodic Information on Dialect Classification Using Whisper Features
Phoebe Parsons, Heming Strømholt Bremnes, Knut Kvale, Torbjørn Svendsen, Giampiero Salvi
INTERSPEECH5
2025 S-HR-VQVAE: Sequential Hierarchical Residual Learning Vector Quantized Variational Autoencoder for Video Prediction
abstract
We address the video prediction task by putting forth a novel model that combines (i) a novel hierarchical residual learning vector quantized variational autoencoder (HR-VQVAE), and (ii) a novel autoregressive spatiotemporal predictive model (AST-PM). We refer to this approach as a sequential hierarchical residual learning vector quantized variational autoencoder (S-HR-VQVAE). By leveraging the intrinsic capabilities of HR-VQVAE at modeling still images with a parsimonious representation, combined with the AST-PM's ability to handle spatiotemporal information, S-HR-VQVAE can better deal with major challenges in video prediction. These include learning spatiotemporal information, handling high dimensional data, combating blurry prediction, and implicit modeling of physical characteristics. Extensive experimental results on four challenging tasks, namely KTH Human Action, TrafficBJ, Human3.6 M, and Kitti, demonstrate that our model compares favorably against state-of-the-art video prediction techniques both in quantitative and qualitative evaluations despite a much smaller model size. Finally, we boost S-HR-VQVAE by proposing a novel training method to jointly estimate the HR-VQVAE and AST-PM parameters.
Mohammad Adiban, Kalin Stefanov, Sabato Marco Siniscalchi, Giampiero Salvi
IEEE Trans. Multim.4
2024 Collecting Linguistic Resources for Assessing Children's Pronunciation of Nordic Languages
abstract
This paper reports on the experience collecting a number of corpora of Nordic languages spoken by children. The aim of the data collection is providing annotated data to develop and evaluate computer assisted pronunciation assessment systems both for non-native children learning a Nordic language (L2) and for L1 children with speech sound disorder (SSD). The paper presents the challenges encountered recording and annotating data for Finnish, Swedish and Norwegian, as well as the ethical considerations related with making this data publicly available. We hope that sharing this experience will encourage others to collect similar data for other languages. Of the different data collections, we were able to make the Norwegian corpus publicly available in the hope that it will serve as a reference in pronunciation assessment research.
Anne Marte Haug Olstad, Anna-Riikka Smolander, Sofia Strömbergsson, Sari Ylinen, Minna Lehtonen, Mikko Kurimo, Yaroslav Getman, Tamás Grósz, Xinwei Cao, Torbjørn Svendsen, Giampiero Salvi
LREC/COLING11
2024 A Framework for Phoneme-Level Pronunciation Assessment Using CTC
Xinwei Cao, Zijian Fan, Torbjørn Svendsen, Giampiero Salvi
INTERSPEECH4
2024 Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions
abstract
This work is concerned with devising a robust Parkinson’s (PD) disease detector from speech in real-world operating conditions using (i) foundational models, and (ii) speech enhancement (SE) methods. To this end, we first fine-tune several foundational-based models on the standard PC-GITA (s-PCGITA) clean data. Our results demonstrate superior performance to previously proposed models. Second, we assess the generalization capability of the PD models on the extended PCGITA (e-PC-GITA) recordings, collected in real-world operative conditions, and observe a severe drop in performance moving from ideal to real-world conditions. Third, we align training and testing conditions applaying off-the-shelf SE techniques on e-PC-GITA, and a significant boost in performance is observed only for the foundational-based models. Finally, combining the two best foundational-based models trained on s-PCGITA, namely WavLM Base and Hubert Base, yielded top performance on the enhanced e-PC-GITA
Moreno La Quatra, Maria Francesca Turco, Torbjørn Svendsen, Giampiero Salvi, Juan Rafael Orozco-Arroyave, Sabato Marco Siniscalchi
INTERSPEECH4
2023 Using Modified Adult Speech as Data Augmentation for Child Speech Recognition
abstract
Data augmentation is a technique which enhances the size and quality of training data such that deep learning or machine learning models can achieve better performance. This paper proposes a novel way of applying data augmentation for child speech recognition in the low data resource scenario. Data augmentation is achieved by modifying existing adult speech signals. The procedure consists of two main parts, resampling, and time scaling. The experiment involves both speech from children aged from kindergarten to grade 10, and adults’ speech. We test the proposed method using both a TDNN-HMM and a GMM-HMM acoustic model. The results show that the proposed data augmentation scheme achieves a relative 7.95% reduction of WERs compared with 4.56% relative reduction when using a traditional bilinear frequency warping approach.
Zijian Fan, Xinwei Cao, Giampiero Salvi, Torbjørn Svendsen
ICASSP3
2023 An Analysis of Goodness of Pronunciation for Child Speech
Xinwei Cao, Zijian Fan, Torbjørn Svendsen, Giampiero Salvi
INTERSPEECH4
2023 Perceptual and Task-Oriented Assessment of a Semantic Metric for ASR Evaluation
Janine Rugayan, Giampiero Salvi, Torbjørn Svendsen
INTERSPEECH2
2023 A step-by-step training method for multi generator GANs with application to anomaly detection and cybersecurity
abstract
Cyber attacks and anomaly detection are problems where the data is often highly unbalanced towards normal observations. Furthermore, the anomalies observed in real applications may be significantly different from the ones contained in the training data. It is, therefore, desirable to study methods that are able to detect anomalies only based on the distribution of the normal data. To address this problem, we propose a novel objective function for generative adversarial networks (GANs), referred to as STEP-GAN. STEP-GAN simulates the distribution of possible anomalies by learning a modified version of the distribution of the task-specific normal data. It leverages multiple generators in a step-by-step interaction with a discriminator in order to capture different modes in the data distribution. The discriminator is optimized to distinguish not only between normal data and anomalies but also between the different generators, thus encouraging each generator to model a different mode in the distribution. This reduces the well-known mode collapse problem in GAN models considerably. We tested our method in the areas of power systems and network traffic control systems (NTCSs) using two publicly available highly imbalanced datasets, ICS (Industrial Control System) security dataset and UNSW-NB15, respectively. In both application domains, STEP-GAN outperforms the state-of-the-art systems as well as the two baseline systems we implemented as a comparison. In order to assess the generality of our model, additional experiments were carried out on seven real-world numerical datasets for anomaly detection in a variety of domains. In all datasets, the number of normal samples is significantly more than that of abnormal samples. Experimental results show that STEP-GAN outperforms several semi-supervised methods while being competitive with supervised methods.
Mohammad Adiban, Sabato Marco Siniscalchi, Giampiero Salvi
Neurocomputing3
2023 NAAQA: A Neural Architecture for Acoustic Question Answering
abstract
The goal of the Acoustic Question Answering (AQA) task is to answer a free-form text question about the content of an acoustic scene. It was inspired by the Visual Question Answering (VQA) task. In this paper, based on the previously introduced CLEAR dataset, we propose a new benchmark for AQA, namely CLEAR2, that emphasizes the specific challenges of acoustic inputs. These include handling of variable duration scenes, and scenes built with elementary sounds that differ between training and test set. We also introduce NAAQA, a neural architecture that leverages specific properties of acoustic inputs. The use of 1D convolutions in time and frequency to process 2D spectro-temporal representations of acoustic content shows promising results and enables reductions in model complexity. We show that time coordinate maps augment temporal localization capabilities which enhance performance of the network by ∼ 17 percentage points. On the other hand, frequency coordinate maps have little influence on this task. NAAQA achieves 79.5% of accuracy on the AQA task with ∼ four times fewer parameters than the previously explored VQA model. We evaluate the performance of NAAQA on an independent data set reconstructed from DAQA. We also test the addition of a MALiMo module in our model on both CLEAR2 and DAQA. We provide a detailed analysis of the results for the different question types. We release the code to produce CLEAR2 as well as NAAQA to foster research in this newly emerging machine learning task.
Jérôme Abdelnour, Jean Rouat, Giampiero Salvi
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Hierarchical Residual Learning Based Vector Quantized Variational Autoencoder for Image Reconstruction and Generation
Mohammad Adiban, Kalin Stefanov, Sabato Marco Siniscalchi, Giampiero Salvi
BMVC4
2022 wav2vec2-based Speech Rating System for Children with Speech Sound Disorder
abstract
The computational resources were provided by Aalto ScienceIT. This work was supported by NordForsk through the funding to Technology-enhanced foreign and second-language learning of Nordic languages, project number 103893.
Yaroslav Getman, Ragheb Al-Ghezi, Katja Voskoboinik, Tamás Grósz, Mikko Kurimo, Giampiero Salvi, Torbjørn Svendsen, Sofia Strömbergsson
INTERSPEECH6
2022 Semantically Meaningful Metrics for Norwegian ASR Systems
abstract
Evaluation metrics are important for quanitfying the performance of Automatic Speech Recognition (ASR) systems. However, the widely used word error rate (WER) captures errors at the word-level only and weighs each error equally, which makes it insufficient to discern ASR system performance for downstream tasks such as Natural Language Understanding (NLU) or information retrieval. We explore in this paper a more robust and discriminative evaluation metric for Norwegian ASR systems through the use of semantic information modeled by a transformer-based language model. We propose Aligned Semantic Distance (ASD) which employs dynamic programming to quantify the similarity between the reference and hypothesis text. First, embedding vectors are generated using the NorBERT model. Afterwards, the minimum global distance of the optimal alignment between these vectors is obtained and normalized by the sequence length of the reference embedding vector. In addition, we present results using Semantic Distance (SemDist), and compare them with ASD. Results show that for the same WER, ASD and SemDist values can vary significantly, thus, exemplifying that not all recognition errors can be considered equally important. We investigate the resulting data, and present examples which demonstrate the nuances of both metrics in evaluating various transcription errors.
Janine Rugayan, Torbjørn Svendsen, Giampiero Salvi
INTERSPEECH3
2022 Acoustic-to-Articulatory Mapping With Joint Optimization of Deep Speech Enhancement and Articulatory Inversion Models
abstract
We investigate the problem of speaker independent acoustic-to-articulatory inversion (AAI) in noisy conditions within the deep neural network (DNN) framework. In contrast with recent results in the literature, we argue that a DNN vector-to-vector regression front-end for speech enhancement (DNN-SE) can play a key role in AAI when used to enhance spectral features prior to AAI back-end processing. We experimented with single- and multi-task training strategies for the DNN-SE block finding the latter to be beneficial to AAI. Furthermore, we show that coupling DNN-SE producing enhanced speech features with an AAI trained on clean speech outperforms a multi-condition AAI (AAI-MC) when tested on noisy speech. We observe a 15% relative improvement in the Pearson's correlation coefficient (PCC) between our system and AAI-MC at 0dB signal-to-noise ratio on the Haskins corpus. Our approach also compares favourably against using a conventional DSP approach to speech enhancement (MMSE with IMCRA) in the front-end. Finally, we demonstrate the utility of articulatory inversion in a downstream speech application. We report significant WER improvements on an automatic speech recognition task in mismatched conditions based on the Wall Street Journal corpus (WSJ) when leveraging articulatory information estimated by AAI-MC system over spectral-alone speech features.
Abdolreza Sabzi Shahrebabaki, Giampiero Salvi, Torbjørn Svendsen, Sabato Marco Siniscalchi
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 STEP-GAN: A One-Class Anomaly Detection Model with Applications to Power System Security
abstract
Smart grid systems (SGSs), and in particular power systems, play a vital role in today’s urban life. The security of these grids is now threatened by adversaries that use false data injection (FDI) to produce a breach of availability, integrity, or confidential principles of the system. We propose a novel structure for the multi-generator generative adversarial network (GAN) to address the challenges of detecting adversarial attacks. We modify the GAN objective function and the training procedure for the malicious anomaly detection task. The model only requires normal operation data to be trained, making it cheaper to deploy and robust against unseen attacks. Moreover, the model operates on the raw input data, eliminating the need for feature extraction. We show that the model reduces the well-known mode collapse problem of GAN-based systems, it has low computational complexity and considerably outperforms the baseline system (OCAN) with about 55% in terms of accuracy on a freely available cyber attack dataset.
Mohammad Adiban, Arash Safari, Giampiero Salvi
ICASSP3
2021 A DNN Based Speech Enhancement Approach to Noise Robust Acoustic-to-Articulatory Inversion
abstract
In this work, we investigate the problem of speaker independent acoustic-to-articulatory inversion (AAI) in noisy condition within the deep neural network (DNN) framework. We claim that DNN vector-to-vector regression for speech enhancement (DNN-SE) can play a key role in AAI when used in a front-end stage to enhance speech features before AAI backend processing. Our claim contrasts recent literature reporting a drop in AAI accuracy on MMSE enhanced data and thereby sheds some light on the opportunities offered by DNN-SE in robust speech applications. We have also tested single- and multitask training strategies of the DNN-SE block and experimentally found the latter to be beneficial to AAI. Moreover, DNN-SE coupled with an AAI deep system tested on enhanced speech can outperform a multi-condition AAI deep system tested on noisy speech. We assess our approach on the Haskins corpus using the Pearson's correlation coefficient (PCC). A 15% relative PCC improvement is observed over a multi-condition AAI system at 0dB signal-to-noise ratio (SNR). Our approach also compares favorably against using a conventional DSP approach, namely MMSE with IMCRA, in the front-end stage.
Abdolreza Sabzi Shahrebabaki, Sabato Marco Siniscalchi, Giampiero Salvi, Torbjørn Svendsen
ISCAS3
2020 Spatial Bias in Vision-Based Voice Activity Detection
abstract
We develop and evaluate models for automatic vision-based voice activity detection (VAD) in multiparty human-human interactions that are aimed at complementing acoustic VAD methods. We provide evidence that this type of vision-based VAD models are susceptible to spatial bias in the dataset used for their development; the physical settings of the interaction, usually constant throughout data acquisition, determines the distribution of head poses of the participants. Our results show that when the head pose distributions are significantly different in the train and test sets, the performance of the vision-based VAD models drops significantly. This suggests that previously reported results on datasets with a fixed physical configuration may overestimate the generalization capabilities of this type of models. We also propose a number of possible remedies to the spatial bias, including data augmentation, input masking and dynamic features, and provide an in-depth analysis of the visual cues used by the developed vision-based VAD models.
Kalin Stefanov, Mohammad Adiban, Giampiero Salvi
ICPR3
2020 Transfer Learning of Articulatory Information Through Phone Information
abstract
Articulatory information has been argued to be useful for several speech tasks.However, in most practical scenarios this information is not readily available.We propose a novel transfer learning framework to obtain reliable articulatory information in such cases.We demonstrate its reliability both in terms of estimating parameters of speech production and its ability to enhance the accuracy of an end-to-end phone recognizer.Articulatory information is estimated from speaker independent phonemic features, using a small speech corpus, with electromagnetic articulography (EMA) measurements.Next, we employ a teacher-student model to learn estimation of articulatory features from acoustic features for the targeted phone recognition task.Phone recognition experiments, demonstrate that the proposed transfer learning approach outperforms the baseline transfer learning system acquired directly from an acousticto-articulatory (AAI) model.The articulatory features estimated by the proposed method, in conjunction with acoustic features, improved the phone error rate (PER) by 6.7% and 6% on the TIMIT core test and development sets, respectively, compared to standalone static acoustic features.Interestingly, this improvement is slightly higher than what is obtained by static+dynamic acoustic features, but with a significantly less.Adding articulatory features on top of static+dynamic acoustic features yields a small but positive PER improvement.
Abdolreza Sabzi Shahrebabaki, Negar Olfati, Sabato Marco Siniscalchi, Giampiero Salvi, Torbjørn Svendsen
INTERSPEECH4
2020 Sequence-to-Sequence Articulatory Inversion Through Time Convolution of Sub-Band Frequency Signals
abstract
We propose a new acoustic-to-articulatory inversion (AAI) sequence-to-sequence neural architecture, where spectral sub-bands are independently processed in time by 1-dimensional (1-D) convolutional filters of different sizes. The learned features maps are then combined and processed by a recurrent block with bi-directional long short-term memory (BLSTM) gates for preserving the smoothly varying nature of the articulatory trajectories. Our experimental evidence shows that, on a speaker dependent AAI task, in spite of the reduced number of parameters, our model demonstrates better root mean squared error (RMSE) and Pearson's correlation coefficient (PCC) than a both a BLSTM model and an FC-BLSTM model where the first stages are fully connected layers. In particular, the average RMSE goes from 1.401 when feeding the filterbank features directly into the BLSTM, to 1.328 with the FC-BLSTM model, and to 1.216 with the proposed method. Similarly, the average PCC increases from 0.859 to 0.877, and 0.895, respectively. On a speaker independent AAI task, we show that our convolutional features outperform the original filterbank features, and can be combined with phonetic features bringing independent information to the solution of the problem. To the best of the authors' knowledge, we report the best results on the given task and data.
Abdolreza Sabzi Shahrebabaki, Sabato Marco Siniscalchi, Giampiero Salvi, Torbjørn Svendsen
INTERSPEECH3
2019 Active Mini-Batch Sampling Using Repulsive Point Processes
abstract
The convergence speed of stochastic gradient descent (SGD) can be improved by actively selecting mini-batches. We explore sampling schemes where similar data points are less likely to be selected in the same mini-batch. In particular, we prove that such repulsive sampling schemes lower the variance of the gradient estimator. This generalizes recent work on using Determinantal Point Processes (DPPs) for mini-batch diversification (Zhang et al., 2017) to the broader class of repulsive point processes. We first show that the phenomenon of variance reduction by diversified sampling generalizes in particular to non-stationary point processes. We then show that other point processes may be computationally much more efficient than DPPs. In particular, we propose and investigate Poisson Disk sampling—frequently encountered in the computer graphics community—for this task. We show empirically that our approach improves over standard SGD both in terms of convergence speed as well as final model performance.
Cheng Zhang 0005, A. Cengiz Öztireli, Stephan Mandt, Giampiero Salvi
AAAI4
2019 Modeling of Human Visual Attention in Multiparty Open-World Dialogues
abstract
This study proposes, develops, and evaluates methods for modeling the eye-gaze direction and head orientation of a person in multiparty open-world dialogues, as a function of low-level communicative signals generated by his/hers interlocutors. These signals include speech activity, eye-gaze direction, and head orientation, all of which can be estimated in real time during the interaction. By utilizing these signals and novel data representations suitable for the task and context, the developed methods can generate plausible candidate gaze targets in real time. The methods are based on Feedforward Neural Networks and Long Short-Term Memory Networks. The proposed methods are developed using several hours of unrestricted interaction data and their performance is compared with a heuristic baseline method. The study offers an extensive evaluation of the proposed methods that investigates the contribution of different predictors to the accurate generation of candidate gaze targets. The results show that the methods can accurately generate candidate gaze targets when the person being modeled is in a listening state. However, when the person being modeled is in a speaking state, the proposed methods yield significantly lower performance.
Kalin Stefanov, Giampiero Salvi, Dimosthenis Kontogiorgos, Hedvig Kjellström, Jonas Beskow
ACM Trans. Hum. Robot Interact.2
2017 Cepstral and Entropy Analyses in Vowels Excerpted from Continuous Speech of Dysphonic and Control Speakers
abstract
There is a growing interest in Cepstral and Entropy analyses of voice samples for defining a vocal health indicator, due to their reliability in investigating both regular and irregular voice signals. The purpose of this study is to determine whether the Cepstral Peak Prominence Smoothed (CPPS) and Sample Entropy (SampEn) could differentiate dysphonic speakers from normal speakers in vowels excerpted from readings and to compare their discrimination power. Results are reported for 33 patients and 31 controls, who read a standardized phonetically balanced passage while wearing a head mounted microphone. Vowels were excerpted from recordings using Automatic Speech Recognition and, after obtaining a measure for each vowel, individual distributions and their descriptive statistics were considered for CPPS and SampEn. The Receiver Operating Curve analysis revealed that the mean of the distributions was the parameter with the highest discrimination power for both CPPS and SampEn. CPPS showed a higher diagnostic precision than SampEn, exhibiting an Area Under Curve (AUC) of 0.85 compared to 0.72. A negative correlation between the parameters was found (Spearman; = −0.61), with higher SampEn corresponding to lower CPPS. The automatic method used in this study could provide support to voice monitorings in clinic and during individual's daily activities.
Antonella Castellana, Andreas Selamtzis, Giampiero Salvi, Alessio Carullo, Arianna Astolfi
INTERSPEECH3
2015 Detecting repetitions in spoken dialogue systems using phonetic distances
abstract
This paper addresses the problem of automatic detection of re-peated turns in Spoken Dialogue Systems. Repetitions can be a symptom of problematic communication between users and systems. Such repetitions are often due to speech recognition errors, which in turn makes it hard to use speech recognition to detect repetitions. We present an approach to detect rep-etition using the phonetic distance to find the best alignment between turns in the same dialogue. The alignment score ob-tained is combined with different features to improve repeti-tion detection. To evaluate the method proposed we compare several alignment techniques from edit distance to DTW-based distance, previously used in Spoken-Term detection tasks. We also compare two different methods to compute the phonetic distance: the first one using the phoneme sequence, and the second one using the distance between the phone posterior vec-tors. Two different datasets were used in this evaluation: a bus-schedule information system (in English) and a call routing system (in Swedish). The results show that approaches using phoneme distances over-perform approaches using Levenshtein distances between ASR outputs for repetition detection. Index Terms: spoken dialogue systems, repetition detection, phonetic distance
José Lopes 0001, Giampiero Salvi, Gabriel Skantze, Alberto Abad, Joakim Gustafson, Fernando Batista, Raveesh Meena, Isabel Trancoso
INTERSPEECH2
2014 Pattern discovery in continuous speech using Block Diagonal Infinite HMM
abstract
We propose the application of a recently introduced inference method, the Block Diagonal Infinite Hidden Markov Model (BDiHMM), to the problem of learning the topology of a Hidden Markov Model (HMM) from continuous speech in an unsupervised way. We test the method on the TiDigits continuous digit database and analyse the emerging patterns corresponding to the blocks of states inferred by the model. We show how the complexity of these patterns increases with the amount of observations and number of speakers. We also show that the patterns correspond to sub-word units that constitute stable and discriminative representations of the words contained in the speech material.
Niklas Vanhainen, Giampiero Salvi
ICASSP2
2014 Audio-visual classification and detection of human manipulation actions
abstract
Humans are able to merge information from multiple perceptional modalities and formulate a coherent representation of the world. Our thesis is that robots need to do the same in order to operate robustly and autonomously in an unstructured environment. It has also been shown in several fields that multiple sources of information can complement each other, overcoming the limitations of a single perceptual modality. Hence, in this paper we introduce a data set of actions that includes both visual data (RGB-D video and 6DOF object pose estimation) and acoustic data. We also propose a method for recognizing and segmenting actions from continuous audio-visual data. The proposed method is employed for extensive evaluation of the descriptive power of the two modalities, and we discuss how they can be used jointly to infer a coherent interpretation of the recorded action.
Alessandro Pieropan, Giampiero Salvi, Karl Pauwels, Hedvig Kjellström
IROS2
2014 The WaveSurfer Automatic Speech Recognition Plugin
Giampiero Salvi, Niklas Vanhainen
LREC1
2014 Free Acoustic and Language Models for Large Vocabulary Continuous Speech Recognition in Swedish
Niklas Vanhainen, Giampiero Salvi
LREC2
2013 A gaze-based method for relating group involvement to individual engagement in multimodal multiparty dialogue
abstract
This paper is concerned with modelling individual engagement and group involvement as well as their relationship in an eight-party, mutimodal corpus. We propose a number of features (presence, entropy, symmetry and maxgaze) that summarise different aspects of eye-gaze patterns and allow us to describe individual as well as group behaviour in time. We use these features to define similarities between the subjects and we compare this information with the engagement rankings the subjects expressed at the end of each interactions about themselves and the other participants. We analyse how these features relate to four classes of group involvement and we build a classifier that is able to distinguish between those classes with 71\% of accuracy.
Catharine Oertel, Giampiero Salvi
ICMI2
2013 On mispronunciation analysis of individual foreign speakers using auditory periphery models
Christos Koniaris, Giampiero Salvi, Olov Engwall
Speech Commun.2
2013 Semi-supervised methods for exploring the acoustics of simple productive feedback
Daniel Neiberg, Giampiero Salvi, Joakim Gustafson
Speech Commun.2
2012 Auditory and Dynamic Modeling Paradigms to Detect L2 Mispronunciations
abstract
This paper expands our previous work on automatic pronunciation error detection that exploits knowledge from psychoacoustic auditory models. The new system has two additional important features, i.e., auditory and acoustic processing of the temporal cues of the speech signal, and classification feedback from a trained linear dynamic model. We also perform a pronunciation analysis by considering the task as a classification problem. Finally, we evaluate the proposed methods conducting a listening test on the same speech material and compare the judgment of the listeners and the methods. The automatic analysis based on spectro-temporal cues is shown to have the best agreement with the human evaluation, particularly with that of language teachers, and with previous plenary linguistic studies. Index Terms: L2 pronunciation error, auditory model, linear dynamic model, distortion measure, phoneme.
Christos Koniaris, Olov Engwall, Giampiero Salvi
INTERSPEECH3
2012 Word Discovery with Beta Process Factor Analysis
abstract
We propose the application of a recently developed nonparametric Bayesian method for factor analysis to the problem of word discovery from continuous speech. The method, based on Beta Process priors, has a number of advantages compared to previously proposed methods, such as Non-negative Matrix Factorisation (NMF). Beta Process Factor Analysis (BPFA) is able to estimate the size of the basis, and therefore the number of recurring patterns, or word candidates, found in the data. We compare the results obtained with BPFA and NMF on the TIDigits database, showing that our method is capable of not only finding the correct words, but also the correct number of words. We also show that the method can infer the approximate number of words for different vocabulary sizes by testing on randomly generated sequences of words. Index Terms: word discovery, beta process factor analysis, Bayesian nonparametric method, non-negative matrix factorisation
Niklas Vanhainen, Giampiero Salvi
INTERSPEECH2
2012 Language Bootstrapping: Learning Word Meanings From Perception-Action Association
abstract
We address the problem of bootstrapping language acquisition for an artificial system similarly to what is observed in experiments with human infants. Our method works by associating meanings to words in manipulation tasks, as a robot interacts with objects and listens to verbal descriptions of the interactions. The model is based on an affordance network, i.e., a mapping between robot actions, robot perceptions, and the perceived effects of these actions upon objects. We extend the affordance model to incorporate spoken words, which allows us to ground the verbal symbols to the execution of actions and the perception of the environment. The model takes verbal descriptions of a task as the input and uses temporal co-occurrence to create links between speech utterances and the involved objects, actions, and effects. We show that the robot is able form useful word-to-meaning associations, even without considering grammatical structure in the learning process and in the presence of recognition errors. These word-to-meaning associations are embedded in the robot's own understanding of its actions. Thus, they can be directly used to instruct the robot to perform tasks and also allow to incorporate context in the speech recognition task. We believe that the encouraging results with our approach may afford robots with a capacity to acquire language descriptors in their operation's environment as well as to shed some light as to how this challenging process develops with human infants.
Giampiero Salvi, Luis Montesano, Alexandre Bernardino, José Santos-Victor
IEEE Trans. Syst. Man Cybern. Part B1
2011 Using Imitation to Learn Infant-Adult Acoustic Mappings
abstract
Abstract This paper discusses a model which conceptually demonstrateshow infants could learn the normalization between infant-adultacoustics. Themodelproposesthatthemappingcanbeinferredfrom the topological correspondences between the adult and in-fant acoustic spaces, that are clustered separately in an unsu-pervised manner. The model requires feedback from the adultin order to select the right topology for clustering, which is acrucial aspect of the model. The feedback is in terms of anoverall rating of the imitation effort by the infant, rather thana frame-by-frame correspondence. Using synthetic, but contin-uous speech data, we demonstrate that clusters, which have agood topological correspondence, are perceived to be similarby a phonetically trained listener.IndexTerms: infantspeechacquisition, unsupervisedlearning,self organizing maps 1. Introduction An infant learning to communicate with speech has been asource of intrigue and interest, both from the point of view ofpsychology and medicine, and from the point of view of speechresearch. Understanding this phenomenon is especially diffi-cult, given that infants do not remember the process once theygrowupandtheremaybelimitedmeansofcommunicatingwithinfants while the process of learning takes place.There are several challenges an infant faces when tryingto acquire the ability to speak. The first and the most impor-tant challenge is learning the sensorimotor mapping betweenthe acoustics and the articulatory configurations [1, 2, 3]. Mostofthetheoriesintheabovestudiesmakeuseofthephenomenonofbabblingtoexplainthisprocess. Secondly, aninfantlearningto speak needs to learn how to categorize the different soundsthat he/she produces or hears from the adults [4, 5].An infant also needs to understand how to correlate the dif-ferentsoundsproducedduringbabblingtothesoundspresentinadult speech which is the main focus of our study. There havebeen several proposals about which acoustic correlates are in-variant between infant and adult speech (e.g., [6]). These mea-sureswereevaluatedby[7]and[8]andwereshowntobeusefulinacquiringtheabilitytoimitatethesoundsproducedbyadults.However, are these invariant acoustic features that normalizeadult speech, intrinsically and instinctively known to children,or do they learnt it? While most studies have assumed this to beinnate to children, Plummer
Gopal Ananthakrishnan, Giampiero Salvi
INTERSPEECH2
2010 Cluster analysis of differential spectral envelopes on emotional speech
abstract
This paper reports on the analysis of the spectral variation of emotional speech. Spectral envelopes of time aligned speech frames are compared between emotionally neutral and active utterances. Statistics are computed over the resulting differential spectral envelopes for each phoneme. Finally, these statistics are classified using agglomerative hierarchical clustering and a measure of dissimilarity between statistical distributions and the resulting clusters are analysed. The results show that there are systematic changes in spectral envelopes when going from neutral to sad or happy speech, and those changes depend on the valence of the emotional content (negative, positive) as well as on the phonetic properties of the sounds such as voicing and place of articulation.
Giampiero Salvi, Fabio Tesser, Enrico Zovato, Piero Cosi
INTERSPEECH1
2009 Affordance based word-to-meaning association
abstract
This paper presents a method to associate meanings to words in manipulation tasks. We base our model on an affordance network, i.e., a mapping between robot actions, robot perceptions and the perceived effects of these actions upon objects. We extend the affordance model to incorporate words. Using verbal descriptions of a task, the model uses temporal co-occurrence to create links between speech utterances and the involved objects, actions and effects. We show that the robot is able form useful word-to-meaning associations, even without considering grammatical structure in the learning process and in the presence of recognition errors. These word-to-meaning associations are embedded in the robot's own understanding of its actions. Thus they can be directly used to instruct the robot to perform tasks and also allow to incorporate context in the speech recognition task.
Verica Krunic, Giampiero Salvi, Alexandre Bernardino, Luis Montesano, José Santos-Victor
ICRA2
2009 Virtual speech reading support for hard of hearing in a domestic multi-media setting
abstract
In this paper we present recent results on the development of the SynFace lip synchronized talking head towards multilinguality, varying signal conditions and noise robustness in the Hearing at Home project. We then describe the large scale hearing impaired user studies carried out for three languages. The user tests focus on measuring the gain in Speech Reception Threshold in Noise when using SynFace, and on measuring the effort scaling when using SynFace by hearing impaired people. Preliminary analysis of the results does not show significant gain in SRT or in effort scaling. But looking at inter-subject variability, it is clear that many subjects benefit from SynFace especially with speech with stereo babble noise.
Samer Al Moubayed, Jonas Beskow, Anne-Marie Öster, Giampiero Salvi, Björn Granström, Nic van Son, Ellen Ormel
INTERSPEECH4
2008 Hearing at home - communication support in home environments for hearing impaired persons
abstract
The Hearing at Home (HaH) project focuses on the needs of hearing-impaired people in home environments. The project is researching and developing an innovative media-center solution for hearing support, with several integrated features that support perception of speech and audio, such as individual loudness amplification, noise reduction, audio classification and event detection, and the possibility to display an animated talking head providing real-time speechreading support. In this paper we provide a brief project overview and then describe some recent results related to the audio classifier and the talking head. As the talking head expects clean speech input, an audio classifier has been developed for the task of classifying audio signals as clean speech, speech in noise or other. The mean accuracy of the classifier was 82%. The talking head (based on technology from the SynFace project) has been adapted for German, and a small speech-in-noise intelligibility experiment was conducted where sentence recognition rates increased from 3 % to 17 % when the talking head was present. Index Terms: Sound classification, speech processing, hearing impairment, communication support, talking heads, speechreading 1.
Jonas Beskow, Björn Granström, Peter Nordqvist, Samer Al Moubayed, Giampiero Salvi, Tobias Herzke, Arne Schulz
INTERSPEECH5
2006 User Evaluation of the SYNFACE Talking Head Telephone
Eva Agelfors, Jonas Beskow, Inger Karlsson, Jo Kewley, Giampiero Salvi
ICCHP5
2006 Dynamic behaviour of connectionist speech recognition with strong latency constraints
Giampiero Salvi
Speech Commun.1
2006 Segment boundary detection via class entropy measurements in connectionist phoneme recognition
Giampiero Salvi
Speech Commun.1
2005 Ecological language acquisition via incremental model-based clustering
abstract
We analyse the behaviour of Incremental Model-Based Clustering on child-directed speech data, and suggest a possible use of this method to describe the acquisition of phonetic classes by an infant. The effects of two factors are analysed, namely the number of coefficients describing the speech signal, and the frame length of the incremental clustering procedure. The results show that, although the number of predicted clusters vary in different conditions, the classifications obtained are essentially consistent. Different classifications were compared using the variation of information measure. 1.
Giampiero Salvi
INTERSPEECH1
2005 Advances in regional accent clustering in Swedish
abstract
The regional pronunciation variation in Swedish is analysed on a large database. Statistics over each phoneme and for each region of Sweden are computed using the EM algorithm in a hidden Markov model framework to overcome the difficulties of transcribing the whole set of data at the phonetic level. The model representations obtained this way are compared using a distance measure in the space spanned by the model parameters, and hierarchical clustering. The regional variants of each phoneme may group with those of any other phoneme, on the basis of their acoustic properties. The log likelihood of the data given the model is shown to display interesting properties regarding the choice of number of clusters, given a particular level of details. Discriminative analysis is used to find the parameters that most contribute to the separation between groups, adding an interpretative value to the discussion. Finally a number of examples are given on some of the phenomena that are revealed by examining the clustering tree. 1.
Giampiero Salvi
INTERSPEECH1
2004 SYNFACE - A Talking Head Telephone for the Hearing-Impaired
Jonas Beskow, Inger Karlsson, Jo Kewley, Giampiero Salvi
ICCHP4
2003 SYNFACE - a talking face telephone
abstract
The SYNFACE project has as its primary goal to facilitate for \nhearing-impaired people to use an ordinary telephone. This \nwill be achieved by using a talking face connected to the \ntelephone. The incoming speech signal will govern the speech \nmovements of the talking face, hence the talking face will \nprovide lip-reading support for the user. \nThe project will define the visual speech information that \nsupports lip-reading, and develop techniques to derive this \ninformation from the acoustic speech signal in near real time \nfor three different languages: Dutch, English and Swedish. \nThis requires the development of automatic speech \nrecognition methods that detect information in the acoustic \nsignal that correlates with the speech movements. This \ninformation will govern the speech movements in a synthetic \nface and synchronise them with the acoustic speech signal. \nA prototype system is being constructed. The prototype \ncontains results achieved so far in SYNFACE. This system \nwill be tested and evaluated for the three languages by \nhearing-impaired users. \nSYNFACE is an IST project (IST-2001-33327) with \npartners from the Netherlands, UK and Sweden. SYNFACE \nbuilds on experiences gained in the Swedish Teleface project.
Inger Karlsson, Andrew Faulkner, Giampiero Salvi
INTERSPEECH3
2003 Using accent information in ASR models for Swedish
abstract
In this study accent information is used in an attempt to improve acoustic models for automatic speech recognition (ASR). First, accent dependent Gaussian models were trained independently. The Bhattacharyya distance was then used in conjunction with agglomerative hierarchical clustering to define optimal strategies for merging those models. The resulting allophonic classes were analyzed and compared with the phonetic literature. Finally, accent "aware" models were built, in which the parametric complexity for each phoneme corresponds to the degree of variability across accent areas and to the amount of training data available for it. The models were compared to models with the same, but evenly spread, overall complexity showing in some cases a slight improvement in recognition accuracy.
Giampiero Salvi
INTERSPEECH1
2000 A Noise Robust Multilingual Reference Recogniser Based on Speechdat(II)
abstract
An important aspect of noise robustness of automatic speech recognisers (ASR) is the proper handling of non-speech acoustic events. The present paper describes further improvements of an already existing reference recogniser towards achieving such kind of robustness. The reference recogniser applied is the COST 249 SpeechDat reference recogniser, which is a fully automatic, language-independent training procedure for building a phonetic recogniser (http://www.telenor.no/fou/prosjekter/taletek/refrec). The reference recogniser relies on the HTK toolkit and a SpeechDat (II) compatible database, and is designed to serve as a reference system in multilingual speech recognition research. The paper describes version 0.96 of the reference recogniser which take into account labelled non-speech acoustic events during training and provides robustness against these during testing. Results are presented on small and medium vocabulary recognition for six languages.
Børge Lindberg, Finn Tore Johansen, Narada D. Warakagoda, Gunnar Lehtinen, Zdravko Kacic, Andrej Zgank, Kjell Elenius, Giampiero Salvi
INTERSPEECH8
2000 The COST 249 SpeechDat Multilingual Reference Recogniser
Finn Tore Johansen, Narada D. Warakagoda, Børge Lindberg, Gunnar Lehtinen, Zdravko Kacic, Andrej Zgank, Kjell Elenius, Giampiero Salvi
LREC8