VLDB 2026 Research / reviewers in the wild / expert
Maurizio Omologo
dblp:06/3413
· DBLP profile ↗
87ranked-venue papers
7as first author
7since 2021 · last 2022
0000-0003-0879-0548ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 77 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 48 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Caching Networks: Capitalizing on Common Speech for ASRabstractWe introduce Caching Networks (CachingNets), a speech recognition network architecture capable of delivering faster, more accurate decoding by leveraging common speech patterns. By explicitly incorporating select sentences unique to each user into the network’s design, we show how to train the model as an extension of the popular sequence transducer architecture through a multitask learning procedure. We further propose and experiment with different phrase caching policies, which are effective for virtual voice-assistant (VA) applications, to complement the architecture. Our results demonstrate that by pivoting between different inference strategies on the fly, CachingNets can deliver significant performance improvements. Specifically, on an industrial-scale, VA ASR task, we observe up to 7.4% relative word error rate (WER) and 11% sentence error rate (SER) improvements with accompanied latency gains. Anastasios Alexandridis, Grant P. Strimel, Ariya Rastrow, Pavel Kveton, Maurizio Omologo, Siegfried Kunzmann, Athanasios Mouchtaris |
ICASSP | 6 |
| 2022 | A Neural Prosody Encoder for End-to-End Dialogue Act ClassificationabstractDialogue act classification (DAC) is a critical task for spoken language understanding in dialogue systems. Prosodic features such as energy and pitch have been shown to be useful for DAC. Despite their importance, little research has explored neural approaches to integrate prosodic features into end-to-end (E2E) DAC models which infer dialogue acts directly from audio signals. In this work, we propose an E2E neural architecture that takes into account the need for characterizing prosodic phenomena co-occurring at different levels inside an utterance. A novel part of this architecture is a learnable gating mechanism that assesses the importance of prosodic features and selectively retains core information necessary for E2E DAC. Our proposed model improves DAC accuracy by 1.07% absolute across three publicly available benchmark datasets. Dillon Knox, Martin Radfar, Grant P. Strimel, Nathan Susanj, Athanasios Mouchtaris, Maurizio Omologo |
ICASSP | 9 |
| 2022 | Overlapped Speech Detection and speaker counting using distant microphone arrays
Samuele Cornell, Maurizio Omologo, Stefano Squartini, Emmanuel Vincent 0001 |
Comput. Speech Lang. | 2 |
| 2022 | Audio-Visual Tracking of Concurrent SpeakersabstractAudio-visual tracking of an unknown number of concurrent speakers in 3D is a challenging task, especially when sound and video are collected with a compact sensing platform. In this paper, we propose a tracker that builds on generative and discriminative audio-visual likelihood models formulated in a particle filtering framework. We localize multiple concurrent speakers with a de-emphasized acoustic map assisted by the image detection-derived 3D video observations. The 3D multi-modal observations are either assigned to existing tracks for discriminative likelihood computation or used to initialize new tracks. The generative likelihoods rely on color distribution of the target and the de-emphasized acoustic map value. Experiments on AV16.3 and CAV3D datasets show that the proposed tracker outperforms the uni-modal trackers and the state-of-the-art approaches both in 3D and on the image plane. Xinyuan Qian 0001, Alessio Brutti, Oswald Lanz, Maurizio Omologo, Andrea Cavallaro |
IEEE Trans. Multim. | 4 |
| 2021 | Context-Aware Transformer Transducer for Speech RecognitionabstractEnd-to-end (E2E) automatic speech recognition (ASR) systems often have difficulty recognizing uncommon words, that appear infrequently in the training data. One promising method, to improve the recognition accuracy on such rare words, is to latch onto personalized/contextual information at inference. In this work, we present a novel context-aware transformer transducer (CATT) network that improves the state-of-the-art transformer-based ASR system by taking advantage of such contextual signals. Specifically, we propose a multi-head attention-based context-biasing network, which is jointly trained with the rest of the ASR sub-networks. We explore different techniques to encode contextual data and to create the final attention context vectors. We also leverage both BLSTM and pretrained BERT based models to encode contextual data and guide the network training. Using an in-house far-field dataset, we show that CATT, using a BERT based context encoder, improves the word error rate of the baseline transformer transducer and outperforms an existing deep contextual model by 24.2% and 19.4% respectively. Feng-Ju Chang, Martin Radfar, Athanasios Mouchtaris, Maurizio Omologo, Ariya Rastrow, Siegfried Kunzmann |
ASRU | 5 |
| 2021 | Multi-Channel Transformer Transducer for Speech RecognitionabstractMulti-channel inputs offer several advantages over singlechannel, to improve the robustness of on-device speech recognition systems.Recent work on multi-channel transformer, has proposed a way to incorporate such inputs into end-to-end ASR for improved accuracy.However, this approach is characterized by a high computational complexity, which prevents it from being deployed in on-device systems.In this paper, we present a novel speech recognition model, Multi-Channel Transformer Transducer (MCTT), which features end-to-end multi-channel training, low computation cost, and low latency so that it is suitable for streaming decoding in on-device speech recognition.In a far-field in-house dataset, our MCTT outperforms stagewise multi-channel models with transformer-transducer up to 6.01% relative WER improvement (WERR).In addition, MCTT outperforms the multi-channel transformer up to 11.62% WERR, and is 15.8 times faster in terms of inference speed.We further show that we can improve the computational cost of MCTT by constraining the future and previous context in attention computations. Feng-Ju Chang, Martin Radfar, Athanasios Mouchtaris, Maurizio Omologo |
Interspeech | 4 |
| 2021 | Phonetically Induced Subwords for End-to-End Speech Recognition
Vasileios Papadourakis, Athanasios Mouchtaris, Maurizio Omologo |
Interspeech | 5 |
| 2020 | Detecting and Counting Overlapping Speakers in Distant Speech ScenariosabstractInternational audience Samuele Cornell, Maurizio Omologo, Stefano Squartini, Emmanuel Vincent 0001 |
INTERSPEECH | 2 |
| 2020 | DiPCo - Dinner Party CorpusabstractWe present a speech data corpus that simulates a "dinner party" scenario taking place in an everyday home environment. The corpus was created by recording multiple groups of four Amazon employee volunteers having a natural conversation in English around a dining table. The participants were recorded by a single-channel close-talk microphone and by five far-field 7-microphone array devices positioned at different locations in the recording room. The dataset contains the audio recordings and human labeled transcripts of a total of 10 sessions with a duration between 15 and 45 minutes. The corpus was created to advance in the field of noise robust and distant speech processing and is intended to serve as a public research and benchmarking data set. Maarten Van Segbroeck, Ahmed Zaid, Ksenia Kutsenko, Cirenia Huerta, Tinh Nguyen, Xuewen Luo, Björn Hoffmeister, Jan Trmal, Maurizio Omologo, Roland Maas |
INTERSPEECH | 9 |
| 2019 | Accurate Target Annotation in 3D from Multimodal StreamsabstractAccurate annotation is fundamental to quantify the performance of multi-sensor and multi-modal object detectors and trackers. However, invasive or expensive instrumentation is needed to automatically generate these annotations. To mitigate this problem, we present a multi-modal approach that leverages annotations from reference streams (e.g. individual camera views) and measurements from unannotated additional streams (e.g. audio) to infer 3D trajectories through an optimization. The core of our approach is a multi-modal extension of Bundle Adjustment with a cross-modal correspondence detection that selectively uses measurements in the optimization. We apply the proposed approach to fully annotate a new multi-modal and multi-view dataset for multi-speaker 3D tracking. Oswald Lanz, Alessio Brutti, Alessio Xompero, Xinyuan Qian 0001, Maurizio Omologo, Andrea Cavallaro |
ICASSP | 5 |
| 2019 | Multi-Speaker Tracking From an Audio-Visual Sensing DeviceabstractCompact multi-sensor platforms are portable and thus desirable for robotics and personal-assistance tasks. However, compared to physically distributed sensors, the size of these platforms makes person tracking more difficult. To address this challenge, we propose a novel 3-D audio-visual people tracker that exploits visual observations (object detections) to guide the acoustic processing by constraining the acoustic likelihood on the horizontal plane defined by the predicted height of a speaker. This solution allows the tracker to estimate, with a small microphone array, the distance of a sound. Moreover, we apply a color-based visual likelihood on the image plane to compensate for misdetections. Finally, we use a 3-D particle filter and greedy data association to combine visual observations, color-based, and acoustic likelihoods to track the position of multiple simultaneous speakers. We compare the proposed multimodal 3-D tracker against two state-of-the-art methods on the AV16.3 dataset and on a newly collected dataset with co-located sensors, which we make available to the research community. Experimental results show that our multimodal approach outperforms the other methods both in 3-D and on the image plane. Xinyuan Qian 0001, Alessio Brutti, Oswald Lanz, Maurizio Omologo, Andrea Cavallaro |
IEEE Trans. Multim. | 4 |
| 2018 | 3D Mouth Tracking from a Compact Microphone Array Co-Located with a cameraabstractWe address the 3D audio-visual mouth tracking problem when using a compact platform with co-located audio-visual sensors, without a depth camera. In particular, we propose a multi-modal particle filter that combines a face detector and 3D hypothesis mapping to the image plane. The audio likelihood computation is assisted by video, which relies on a GCC-PHAT based acoustic map. By combining audio and video inputs, the proposed approach can cope with a reverberant and noisy environment, and can deal with situations when the person is occluded, outside the Field of View (FoV), or not facing the sensors. Experimental results show that the proposed tracker is accurate both in 3D and on the image plane. Xinyuan Qian 0001, Alessio Xompero, Andrea Cavallaro, Alessio Brutti, Oswald Lanz, Maurizio Omologo |
ICASSP | 6 |
| 2018 | Cepstral distance based channel selection for distant speech recognition
Cristina Guerrero, Georgina Tryfou, Maurizio Omologo |
Comput. Speech Lang. | 3 |
| 2018 | Automatic context window composition for distant speech recognition
Mirco Ravanelli, Maurizio Omologo |
Speech Commun. | 2 |
| 2017 | 3D audio-visual speaker tracking with an adaptive particle filterabstractWe propose an audio-visual fusion algorithm for 3D speaker tracking from a localised multi-modal sensor platform composed of a camera and a small microphone array. After extracting audio-visual cues from individual modalities we fuse them adaptively using their reliability in a particle filter framework. The reliability of the audio signal is measured based on the maximum Global Coherence Field (GCF) peak value at each frame. The visual reliability is based on colour-histogram matching with detection results compared with a reference image in the RGB space. Experiments on the AV16.3 dataset show that the proposed adaptive audio-visual tracker outperforms both the individual modalities and a classical approach with fixed parameters in terms of tracking accuracy. Xinyuan Qian 0001, Alessio Brutti, Maurizio Omologo, Andrea Cavallaro |
ICASSP | 3 |
| 2017 | A network of deep neural networks for Distant Speech RecognitionabstractDespite the remarkable progress recently made in distant speech recognition, state-of-the-art technology still suffers from a lack of robustness, especially when adverse acoustic conditions characterized by non-stationary noises and reverberation are met. A prominent limitation of current systems lies in the lack of matching and communication between the various technologies involved in the distant speech recognition process. The speech enhancement and speech recognition modules are, for instance, often trained independently. Moreover, the speech enhancement normally helps the speech recognizer, but the output of the latter is not commonly used, in turn, to improve the speech enhancement. To address both concerns, we propose a novel architecture based on a network of deep neural networks, where all the components are jointly trained and better cooperate with each other thanks to a full communication scheme between them. Experiments, conducted using different datasets, tasks and acoustic conditions, revealed that the proposed framework can overtake other competitive solutions, including recent joint training approaches. Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, Yoshua Bengio |
ICASSP | 3 |
| 2017 | A reassigned based singing voice pitch contour extraction methodabstractAlthough there are many systems concerned with melody extraction from polyphonic music, there are certain limitations stemming from the spectral processing that are yet to be overpassed. In this paper, we propose a novel method to create sets of melodic pitch contours which are shown to contain harmonic information critical for a melody extraction system. The proposed approach exploits interesting characteristics of the reassigned spectrogram and computes a new representation which comprises a set of points in the time-frequency domain, weighted according to their dominance, in terms of harmonic content. The experimental results show that the proposed method is a valid approach to the detection of time-frequency points that are related to the melodic content of music signals. Moreover, the quality of the acquired melodic pitch contours is proved through a comparison with those extracted by a state-of-the-art melody extraction system. Georgina Tryfou, Maurizio Omologo |
ICASSP | 2 |
| 2017 | Improving Speech Recognition by Revising Gated Recurrent UnitsabstractSpeech recognition is largely taking advantage of deep learning, showing that substantial benefits can be obtained by modern Recurrent Neural Networks (RNNs). The most popular RNNs are Long Short-Term Memory (LSTMs), which typically reach state-of-the-art performance in many tasks thanks to their ability to learn long-term dependencies and robustness to vanishing gradients. Nevertheless, LSTMs have a rather complex design with three multiplicative gates, that might impair their efficient implementation. An attempt to simplify LSTMs has recently led to Gated Recurrent Units (GRUs), which are based on just two multiplicative gates. This paper builds on these efforts by further revising GRUs and proposing a simplified architecture potentially more suitable for speech recognition. The contribution of this work is two-fold. First, we suggest to remove the reset gate in the GRU design, resulting in a more efficient single-gate architecture. Second, we propose to replace tanh with ReLU activations in the state update equations. Results show that, in our implementation, the revised architecture reduces the per-epoch training time with more than 30% and consistently improves recognition performance across different tasks, input features, and noisy conditions when compared to a standard GRU. Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, Yoshua Bengio |
INTERSPEECH | 3 |
| 2017 | Audio Source Separation in Reverberant Environments Using β-Divergence-Based Nonnegative FactorizationabstractIn Gaussian model-based multichannel audio source separation, the likelihood of observed mixtures of source signals is parametrized by source spectral variances and by associated spatial covariance matrices. These parameters are estimated by maximizing the likelihood through an expectation-maximization algorithm and used to separate the signals by means of multichannel Wiener filtering. We propose to estimate these parameters by applying nonnegative factorization based on prior information on source variances. In the nonnegative factorization, spectral basis matrices can be defined as the prior information. The matrices can be either extracted or indirectly made available through a redundant library that is trained in advance. In a separate step, applying nonnegative tensor factorization, two algorithms are proposed in order to either extract or detect the basis matrices that best represent the power spectra of the source signals in the observed mixtures. The factorization is achieved by minimizing the β-divergence through multiplicative update rules. The sparsity of factorization can be controlled by tuning the value of β. Experiments show that sparsity, rather than the value assigned to β in the training, is crucial in order to increase the separation performance. The proposed method was evaluated in several mixing conditions. It provides better separation quality with respect to other comparable algorithms. Mahmoud Fakhry, Piergiorgio Svaizer, Maurizio Omologo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Channel Selection for Distant Speech Recognition Exploiting Cepstral Distance
Cristina Guerrero, Georgina Tryfou, Maurizio Omologo |
INTERSPEECH | 3 |
| 2016 | Realistic Multi-Microphone Data Simulation for Distant Speech RecognitionabstractThe availability of realistic simulated corpora is of key importance for the future progress of distant speech recognition technology. The reliability, flexibility and low computational cost of a data simulation process may ultimately allow researchers to train, tune and test different techniques in a variety of acoustic scenarios, avoiding the laborious effort of directly recording real data from the targeted environment. In the last decade, several simulated corpora have been released to the research community, including the data-sets distributed in the context of projects and international challenges, such as CHiME and REVERB. These efforts were extremely useful to derive baselines and common evaluation frameworks for comparison purposes. At the same time, in many cases they highlighted the need of a better coherence between real and simulated conditions. In this paper, we examine this issue and we describe our approach to the generation of realistic corpora in a domestic context. Experimental validation, conducted in a multi-microphone scenario, shows that a comparable performance trend can be observed with both real and simulated data across different recognition frameworks, acoustic models, as well as multi-microphone processing techniques. Mirco Ravanelli, Piergiorgio Svaizer, Maurizio Omologo |
INTERSPEECH | 3 |
| 2016 | Batch-normalized joint training for DNN-based distant speech recognitionabstractImproving distant speech recognition is a crucial step towards flexible human-machine interfaces. Current technology, however, still exhibits a lack of robustness, especially when adverse acoustic conditions are met. Despite the significant progress made in the last years on both speech enhancement and speech recognition, one potential limitation of state-of-the-art technology lies in composing modules that are not well matched because they are not trained jointly. Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, Yoshua Bengio |
SLT | 3 |
| 2015 | Boosted acoustic model learning and hypotheses rescoring on the CHiME-3 taskabstractSpeech recognition in a realistic noisy environment using multiple microphones is the focal point of the third CHiME challenge. Over the baseline ASR system provided for this challenge, we apply state of the art algorithms for boosting acoustic model learning and hypothesis rescoring to improve the final output. To this aim, we first use the automatic transcription of each channel to re-train the acoustic model for that channel and then we apply linear language model rescoring to find a better solution in the n-best list. LM rescoring is performed using an efficient set of N-gram and Recurrent Neural Network LM (RNNLM) trained on a wisely-selected text set. In the experiments, we show that the proposed approach improves not only the individual channel transcription, but also the enhanced channels produced by MVDR and delay-and-sum beamforming. Shahab Jalalvand, Daniele Falavigna, Marco Matassoni, Piergiorgio Svaizer, Maurizio Omologo |
ASRU | 5 |
| 2015 | The DIRHA-ENGLISH corpus and related tasks for distant-speech recognition in domestic environmentsabstractThis paper introduces the contents and the possible usage of the DIRHA-ENGLISH multi-microphone corpus, recently realized under the EC DIRHA project. The reference scenario is a domestic environment equipped with a large number of microphones and microphone arrays distributed in space. The corpus is composed of both real and simulated material, and it includes 12 US and 12 UK English native speakers. Each speaker uttered different sets of phonetically-rich sentences, newspaper articles, conversational speech, keywords, and commands. From this material, a large set of 1-minute sequences was generated, which also includes typical domestic background noise as well as inter/intra-room reverberation effects. Dev and test sets were derived, which represent a very precious material for different studies on multi-microphone speech processing and distant-speech recognition. Various tasks and corresponding Kaldi recipes have already been developed. The paper reports a first set of baseline results obtained using different techniques, including Deep Neural Networks (DNN), aligned with the state-of-the-art at international level. Mirco Ravanelli, Luca Cristoforetti, Roberto Gretter, Marco Pellin, Alessandro Sosi, Maurizio Omologo |
ASRU | 6 |
| 2015 | Audio source separation using a redundant library of source spectral bases for non-negative tensor factorizationabstractThis work proposes a solution to the problem of under-determined audio source separation using pre-trained redundant source-based prior information. In local Gaussian modeling of a mixing process, an observed mixture is modeled by a Gaussian distribution parameterized by source variances and spatial covariance matrices. The separation is performed by estimating the parameters, and applying Wiener filtering on the observed mixture. We propose, in a training phase, to build a redundant library of spectral basis matrices of all probable source power spectra, applying non-negative tensor factorization (NTF). In the testing phase, the matrices that match the observed mixture are detected using NTF. With the help of the detected matrices, a maximum likelihood algorithm is proposed in order to iteratively estimate the parameters of the model, exploiting the spatial redundancy of the observed mixture and using NTF. The proposed algorithm proves more flexibility and efficiency with respect to a baseline algorithm used as a reference. Mahmoud Fakhry, Piergiorgio Svaizer, Maurizio Omologo |
ICASSP | 3 |
| 2015 | A multi-channel corpus for distant-speech interaction in presence of known interferencesabstractThis paper describes a new corpus of multi-channel audio data designed to study and develop distant-speech recognition systems able to cope with known interfering sounds propagating in an environment. The corpus consists of both real and simulated signals and of a corresponding detailed annotation. An extensive set of speech recognition experiments was conducted using three different Acoustic Echo Cancellation (AEC) techniques to establish baseline results for future reference. The AEC techniques were applied both to single distant microphone input signals and beamformed signals generated using two state-of-the-art beamforming techniques. We show that the speech recognition performance using the different techniques is comparable for both the simulated and real data, demonstrating the usefulness of this corpus for speech research. We also show that a significant improvement in speech recognition performance can be obtained by combining state-of-the-art AEC and beamforming techniques, compared to using a single distant microphone input. Erich Zwyssig, Mirco Ravanelli, Piergiorgio Svaizer, Maurizio Omologo |
ICASSP | 4 |
| 2015 | Contaminated speech training methods for robust DNN-HMM distant speech recognitionabstractDespite the significant progress made in the last years, state-of-the-art speech recognition technologies provide a satisfactory performance only in the close-talking condition. Robustness of distant speech recognition in adverse acoustic conditions, on the other hand, remains a crucial open issue for future applications of human-machine interaction. To this end, several advances in speech enhancement, acoustic scene analysis as well as acoustic modeling, have recently contributed to improve the state-of-the-art in the field. One of the most effective approaches to derive a robust acoustic modeling is based on using contaminated speech, which proved helpful in reducing the acoustic mismatch between training and testing conditions. In this paper, we revise this classical approach in the context of modern DNN-HMM systems, and propose the adoption of three methods, namely, asymmetric context windowing, close-talk based supervision, and close-talk based pre-training. The experimental results, obtained using both real and simulated data, show a significant advantage in using these three methods, overall providing a 15% error rate reduction compared to the baseline systems. The same trend in performance is confirmed either using a high-quality training set of small size, and a large one. Mirco Ravanelli, Maurizio Omologo |
INTERSPEECH | 2 |
| 2014 | On the selection of the impulse responses for distant-speech recognition based on contaminated speech training
Mirco Ravanelli, Maurizio Omologo |
INTERSPEECH | 2 |
| 2014 | The DIRHA simulated corpus
Luca Cristoforetti, Mirco Ravanelli, Maurizio Omologo, Alessandro Sosi, Alberto Abad, Martin Hagmüller, Petros Maragos |
LREC | 3 |
| 2013 | Geometric contamination for GMM/UBM speaker verification in reverberant environments
Alessio Brutti, Maurizio Omologo |
INTERSPEECH | 2 |
| 2013 | Embedding speech recognition to control lights
Alessandro Sosi, Fabio Brugnara, Luca Cristoforetti, Marco Matassoni, Mirco Ravanelli, Maurizio Omologo |
INTERSPEECH | 6 |
| 2013 | An environment aware ML estimation of acoustic radiation pattern with distributed microphone pairs
Alessio Brutti, Maurizio Omologo, Piergiorgio Svaizer |
Signal Process. | 2 |
| 2012 | A probabilistic approach to simultaneous extraction of beats and downbeatsabstractThis paper focuses on the automatic extraction of beat structure from a musical piece. A novel statistical approach to modeling beat sequences based on the application of Hidden Markov Models (HMM) is introduced. The resulting beat labels are obtained by running the Viterbi decoder and subsequent lattice rescoring. For the observation vectors we propose a new feature set that is based on the impulsive and harmonic components of the reassigned spectrogram. Different components of observation vectors have been investigated for their efficiency. The main advantage of the proposed approach is the absence of imposed deterministic rules. All the parameters are learned from the training data, and the experimental results show the efficiency of the proposed schema. Maksim Khadkevich, Thomas Fillon, Gaël Richard, Maurizio Omologo |
ICASSP | 4 |
| 2012 | Enhanced multidimensional spatial functions for unambiguous localization of multiple sparse acoustic sourcesabstractThe Steered Response Power with PHAT transform (SRP-PHAT) or Global Coherence Field (GCF), has become a standard method for acoustic source localization, thanks to their simplicity, computational inexpensiveness and robustness against mid-high reverberation. However, originally formulated for the single source localization case, it does not apply satisfactorily to the multiple source case. In this paper, we analyze the structure of the spatial function and reshape it according to a generic multidimensional metric. We show that traditional functions are based on the L1 norm which is prone to generate ambiguous locations with high likelihood (i.e. ghosts). A more generic multidimensional kernel based on higher norms and on a partitioned representation of the cross-power spectrum is introduced, which better exploits the source sparseness in the discrete time-frequency domain. Evaluation results over simulated data show that the new spatial functions considerably improve the detection of multiple competing sources in both spatial and multidimensional TDOA domains. Francesco Nesta, Maurizio Omologo |
ICASSP | 2 |
| 2012 | Semi-Blind Model Adaptation using Piece-wise Energy Decay Curve for Large Reverberant EnvironmentsabstractThis work presents semi-blind acoustic model adaptation based on a piece-wise energy decay curve. The dual slope representation of the piece-wise curve accurately captures the early and late reflection decay that helps in precisely modeling the smearing effect caused due to reverberation. The slopes are estimated in a semi-blind fashion, late reflection slope is estimated blindly by finding the highest likelihood obtained after matching the test features with Gaussian mixture models trained on reverberant data, while the early reflection slope is empirically computed. Adaptation using piece-wise decay curve leads to robust acoustic models consequently improving the recognition performance. The approach is tested on connected digits recognition task in a lecture room with various large reverberation times. The performance is compared with the exponential decay approach and incremental MLLR, where the proposed technique is found to be robust and consistent across all the cases. Abdul Waheed Mohammed, Marco Matassoni, Hari Krishna Maganti, Maurizio Omologo |
INTERSPEECH | 4 |
| 2012 | Generalized State Coherence Transform for Multidimensional TDOA Estimation of Multiple SourcesabstractAccording to the physical meaning of the frequency-domain blind source separation (FD-BSS), each mixing matrix estimated by independent component analysis (ICA) contains information on the physical acoustic propagation related to each source and then can be used for localization purposes. In this paper, we analyze the Generalized State Coherence Transform (GSCT) which is a non-linear transform of the space represented by the whole demixing matrices. The transform enables an accurate estimation of the propagation time-delay of multiple sources in multiple dimensions. Furthermore, it is shown that with appropriate nonlinearities and a statistical model for the reverberation, GSCT can be considered an approximated kernel density estimator of the acoustic propagation time-delay. Experimental results confirm the good properties of the transform and its effectiveness in addressing multiple source TDOA detection (e.g., 2-D TDOA estimation of several sources with only three microphones). Francesco Nesta, Maurizio Omologo |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Time-frequency reassigned features for automatic chord recognitionabstractThis paper addresses feature extraction for automatic chord recognition systems. Most chord recognition systems use chroma features as a front-end and some kind of classifier (HMM, SVM or template matching). The vast majority of feature extraction approaches are based on mapping frequency bins from spectrum orconstant-Q spectrum to chroma bins. In this work a set of new chroma features that are based on the time-frequency reassignment (TFR) technique is investigated. The proposed feature set was evaluated on the commonly used Beatles dataset and proved to be efficient for the chord recognition task, outperforming standard chroma. Maksim Khadkevich, Maurizio Omologo |
ICASSP | 2 |
| 2011 | Approximated kernel density estimation for multiple TDOA detectionabstractThe Generalized State Coherence Transform (GSCT) has been recently proposed as an efficient tool for the estimation of multidimensional TDOA of multiple sources. The transform defines a multivariate likelihood of the TDOA through a non-linear integration of complex-valued states, representing the acoustic propagation of multiple sources. In the previous works the non-linearity was heuristically motivated leading to a difficult interpretation of the resulting likelihoods and of a correct choice of the parameters. Modeling the time-delays of the acoustic propagation of multiple sources with a multivariate multimodal distribution, a non-parametric kernel density estimator may be derived, which intrinsically accounts for spatial aliasing. From the theoretical analysis it follows that with an appropriate frequency-dependent non-linearity the GSCT likelihood approximates the true kernel density. Theoretical discussion is confirmed by experimental results which show that the proposed nonlinearity dramatically improves both resolution and smoothness of unidimensional and bidimensional likelihoods. Francesco Nesta, Maurizio Omologo |
ICASSP | 2 |
| 2011 | Convolutive BSS of Short Mixtures by ICA Recursively Regularized Across FrequenciesabstractThis paper proposes a new method of frequency-domain blind source separation (FD-BSS), able to separate acoustic sources in challenging conditions. In frequency-domain BSS, the time-domain signals are transformed into time-frequency series and the separation is generally performed by applying independent component analysis (ICA) at each frequency envelope. When short signals are observed and long demixing filters are required, the number of time observations for each frequency is limited and the variance of the ICA estimator increases due to the intrinsic statistical bias. Furthermore, common methods used to solve the permutation problem fail, especially with sources recorded under highly reverberant conditions. We propose a recursively regularized implementation of the ICA (RR-ICA) that overcomes the mentioned problem by exploiting two types of deterministic knowledge: 1) continuity of the demixing matrix across frequencies; 2) continuity of the time-activity of the sources. The recursive regularization propagates the statistics of the sources across frequencies reducing the effect of statistical bias and the occurrence of permutations. Experimental results on real-data show that the algorithm can successfully perform a fast separation of short signals (e.g., 0.5-1s), by estimating long demixing filters to deal with highly reverberant environments (e.g., ms). Francesco Nesta, Piergiorgio Svaizer, Maurizio Omologo |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Cooperative Wiener-ICA for source localization and Separation by distributed microphone arraysabstractDuring the last decade, distributed microphone arrays have been proposed in order to increase accuracy and spatial coverage of speaker localization systems operating in large and reverberant rooms. In principle, the framework provided by a distributed microphone network can also be applied effectively when using Blind Source Separation (BSS). Separation is commonly performed by processing the signals sampled at closely spaced microphones in a single adaptation step, for example by means of Independent Component Analysis (ICA). When the microphone spacing or the distance between source and microphones increase, the separation performance reduces due to spatial aliasing effects and to a reduced spatial coherence at microphones. In this paper we propose a new method, here referred to as Cooperative Wiener ICA (CW-ICA), which is able to apply BSS to signals acquired by a network of distributed microphone arrays. Different ICA adaptations are applied to the signals recorded by each array and are interconnected in order to constrain each adaptation to converge to a solution related to the same physical interpretation. A preliminary analysis on a network of two arrays shows that the proposed method can be applied successfully to source separation and localization tasks. Francesco Nesta, Maurizio Omologo |
ICASSP | 2 |
| 2010 | Experiments on distant-talking speaker verification in TV scenarioabstractIn this work text-independent speaker verification (SV) in a distant-talking noisy scenario is addressed: users can interact with a TV-system able to understand vocal commands and verify simultaneously the identity of the speaker. The main issues with SV under this scenario are related to reverberation, interfering sound sources (TV output) and usually very short utterances; as a consequence, an increasing confusability among speakers models can be observed. To partially cope with this, we propose a system that exploits the processing of signals acquired by a microphone array and a phonetic class segmentation in unsupervised modality. Comparing the proposed system with a GMM-UBM based system we demonstrate the effectiveness of the approach on data acquired with a real prototype. Christian Zieger, Marco Matassoni, Maurizio Omologo |
ICASSP | 3 |
| 2009 | Robust two-channel TDOA estimation for multiple speaker localization by using recursive ICA and a state coherence transformabstractA novel method is presented for a robust two channel multiple time difference of arrival (TDOA) estimation for multispeaker localization which can provide satisfactory performance even in highly reverberant environment. The method is based on a recursive frequency-domain independent component analysis (ICA) and on a novel state coherence transform (SCT). Exploiting the phase coherence of the demixing matrices obtained in the ICA stage the SCT is able to generate envelopes with clear peaks in the corresponding maximum-likelihood TDOAs. The SCT envelopes are computed independently in each time-block and accurate multiple TDOAs are estimated by means of a time-frequency sparse representation of the sources. The method has been applied to real data obtained by recording many sources in a room with a reverberation time of 700 ms. Experimental results show that an accurate localization of 7 closely-spaced sources is possible given only few seconds of data even in the case of low SNR. Experiments also show the advantage of using the proposed solution rather than the well-known GCC-PHAT. Francesco Nesta, Piergiorgio Svaizer, Maurizio Omologo |
ICASSP | 3 |
| 2008 | Localization of multiple speakers based on a two step acoustic map analysisabstractAn interface for distant-talking control of home devices requires the possibility of identifying the positions of multiple users. Acoustic maps, based either on global coherence field (GCF) or oriented global coherence field (OGCF), have already been exploited successfully to determine position and head orientation of a single speaker. This paper proposes a new method using acoustic maps to deal with the case of two simultaneous speakers. The method is based on a two step analysis of a coherence map: first the dominant speaker is localized; then the map is modified by compensating for the effects due to the first speaker and the position of the second speaker is detected. Simulations were carried out to show how an appropriate analysis of OGCF and GCF maps allows one to localize both speakers. Experiments proved the effectiveness of the proposed solution in a linear microphone array set up. Alessio Brutti, Maurizio Omologo, Piergiorgio Svaizer |
ICASSP | 2 |
| 2008 | Acoustic event classification using a distributed microphone network with a GMM/SVM combined algorithmabstractThis work proposes a system for acoustic event classification using signals acquired by a Distributed Microphone Network (DMN). The system is based on the combination of Gaussian Mixture Models (GMM) and Support Vector Machines (SVM). The acoustic event list includes both speech and non-speech events typical of seminars and meetings. The robustness of the system was investigated by considering two scenarios characterized by different types of trained models and testing conditions. Experimental results were obtained by using real-world data collected at two sites. The results in terms of classification error rate show that in each scenario the proposed system outperforms any single classifier based system. Index Terms: audio classification, distributed microphone network, GMM, SVM. Christian Zieger, Maurizio Omologo |
INTERSPEECH | 2 |
| 2008 | Combination of clean and contaminated GMM/SVM for far-field text-independent speaker verificationabstractThis paper addresses the problem of speaker verification under reverberant conditions, using only the signal acquired by a single distant microphone. The proposed system combines four different subsystems. Two of them are Gaussian Mixture Model (GMM) based and the other two are Support Vector Machine (SVM) based. The subsystems that use the same type of classifier differ in terms of models: one is trained with clean speech and the other is trained with noisy and reverberant speech obtained through the contamination of the clean data with the measured impulse responses of the room. The results show that the proposed system outperforms each single subsystem under matched or mismatched conditions. Index Terms: speaker verification, reverberation, GMM, SVM. 1. Christian Zieger, Maurizio Omologo |
INTERSPEECH | 2 |
| 2008 | WOZ Acoustic Data Collection for Interactive TV
Alessio Brutti, Luca Cristoforetti, Walter Kellermann, Lutz Marquardt, Maurizio Omologo |
LREC | 5 |
| 2007 | Classification of Acoustic Maps to Determine Speaker Position and Orientation from a Distributed Microphone NetworkabstractAcoustic maps created on the basis of the signals acquired by distributed networks of microphones allow to identify position and orientation of an active talker in an enclosure. In adverse situations of high background noise, high reverberation or unavailability of direct paths to the microphones, localization may fail. This paper proposes a novel approach to talker localization and estimation of head orientation based on the classification of global coherence field (GCF) or oriented GCF maps. Preliminary experiments with data obtained by simulated propagation as well as with data acquired in a real room show that the match with precalculated map models provides a robust behavior in adverse conditions. Alessio Brutti, Maurizio Omologo, Piergiorgio Svaizer, Christian Zieger |
ICASSP (4) | 2 |
| 2007 | Adaptive weighting of microphone arrays for distant-talking F0 and voiced/unvoiced estimationabstractThis paper introduces a new technique of multi-microphone processing which aims to provide features for the extraction of fundamental frequency and for the classification of voiced/unvoiced segments in distant-talking speech. A multichannel periodicity function (MPF) is derived from an adaptive weighting of normalized and compressed magnitude spectra. This function highlights periodic clues of the given speech signals, even under noisy and reverberant conditions. The resulting MPF features are then exploited for voiced/unvoiced classification based on Hidden Markov Models. Experiments, conducted both on simulated data and on real seminar recordings based on a network of reversed T-shaped arrays, showed the robustness of the proposed technique. Federico Flego, Christian Zieger, Maurizio Omologo |
INTERSPEECH | 3 |
| 2006 | Speaker localization based on oriented global coherence fieldabstractAbstract This paper proposes a new speaker localization method that isbased on a preliminary estimation of the head orientation. The ba-sic information on which the estimation is accomplished is calledOriented Global Coherence Field (OGCF).The new algorithm is shown to be significantly more robustthan the traditional ones so far explored. Its robustness is also dueto an effective speech activity detection, implicitly performed bya thresholding technique applied to OGCF information. To showthe performance of the proposed system, experiments were con-ducted on the NIST RT-05 Spring Evaluation source localizationtask, which is based on real recordings of lectures in noisy andreverberant environments. Index Terms : speaker localization, head orientation, microphonearrays, global coherence field. 1. Introduction Since 1990, several Speaker LOCalization (SLOC) techniqueshave been proposed as reported in [1, 2]. Most of the traditionalSLOC techniques are based on the estimation of time differencesof wavefront arrival at each sensor and on a consequent applica-tion of geometrical information to infer the acoustic source posi-tions. One of the most common techniques for Time Delay Es-timation (TDE) is based on Generalized Cross-Correlation PhaseTransform (GCC-PHAT) [3, 4]. Other effective SLOC techniquesare based on a preliminary computation of an acoustic map, asfor instance the Global Coherence Field (GCF) [5] representation,fromwhichthemostlikelysourcepositionisderivedthroughmax-imization in space.This paper aims at describing a new SLOC method that wasconceived starting from the effectiveness of the Oriented GlobalCoherence Field(OGCF), introduced in [6], which allows to char-acterize the orientation of an active speaker’s head with a satis-factory accuracy (in terms of angle error) even under reverberantconditions. By exploiting OGCF information, one can also derivemore robust speaker position estimates, since they are mostly re-lated to the propagation of a direct wavefront from a given point.On the other hand, previous SLOC techniques did not deal withthe way the sound is being radiated from a hypothesized positionin space.Theproposed method requires touseadistributed microphonenetwork similar to those available in the laboratories involved inthe EC CHIL Alessio Brutti, Maurizio Omologo, Piergiorgio Svaizer |
INTERSPEECH | 2 |
| 2006 | Multi-microphone periodicity function for robust F0 estimation in real noisy and reverberant environmentsabstractAbstract This paper outlines a new method to extract F0 from distant-talking speech signals acquired by a microphone network, whichexploits the redundancy across the signals proceeding from eachmicrophone, by jointly processing the different contributes. Tothis purpose, a multi-microphone periodicity function is derivedfrom the magnitude spectrum computed on each microphone sig-nal. This function allows to estimate F0 reliably, even under re-verberant conditions, without the need of any post-processing orsmoothing technique. Experiments, conducted on real lectures,showed that the proposed frequency-domain algorithm is moresuitable than other time-domain based ones. IndexTerms : speechanalysis,fundamentalfrequencyestimation,multi-microphone processing, distant-talking interaction. 1. Introduction In the CHIL project, various signal processing techniques are be-ing investigated that aim to address challenging problems amongwhich acoustic event classification, speakerlocalization and track-ing, distant-talking speech recognition, speech activity detection,speaker identification and verification [1].One way to pursue all these objectives is that of deriving amodel of the source (e.g. the speaker) from the given multi-microphone data. To this purpose, a Distributed Microphone Net-work (DMN) is used, which consists in a generic set of micro-phones localized in space without any specific geometry.In this work we address the problem of deriving a robust es-timation of the fundamental frequency F0 from the variety of sig-nals recorded through the microphone network. Speech signalsrecorded by microphones placed far from a talker are severely de-graded by both background noise and reverberation, which de-pends on spatial relationships among the microphones and thetalker, as well as on the scenario acoustic characteristics.Estimating F0 independently for each microphone signal andapplying then majority vote or other fusion based methods mayrepresent a possible approach. Another way to perform F0 es-timation is to extend to the multi-microphone case a paradigmthat works for a single microphone close-talking case. A time-domain F0 extraction algorithm based on Weighted Autocorre-lation (WAUTOC) [2] was experimented in the past [3], whichshowed good performance on a real multi-microphone databaseof distant-talking speech sequences reproduced in an office envi-ronment. In particular, the resulting multi-microphone WAUTOCtechnique offers the advantage of obtaining better performance Federico Flego, Maurizio Omologo |
INTERSPEECH | 2 |
| 2006 | Speech Recognition in Reverberant Environments Using Remote MicrophonesabstractThis paper addresses distant-talking speech recognition by means of remote sensors in a reverberant room. Recognition performances is investigated for different ways of initializing, steering, and optimizing the related beamformer. Results show how much critical that front-end processing may be in such a challenging setup, according to the different positions and orientations of the speaker Luca Giulio Brayda, Christian Wellekens, Marco Matassoni, Maurizio Omologo |
ISM | 4 |
| 2005 | Automatic Speech Activity Detection, Source Localization, and Speech Recognition on the Chil Seminar CorpusabstractTo realize the long-term goal of ubiquitous computing, technological advances in multi-channel acoustic analysis are needed in order to solve several basic problems, including speaker localization and tracking, speech activity detection (SAD) and distant-talking automatic speech recognition (ASR). The European Commission integrated project CHIL, “ Computers in the Human Interaction Loop”, aims to make significant advances in these three technologies. In this work, we report the results of our initial automatic source localization, speech activity detection, and speech recognition experiments on the CHIL seminar corpus, which is comprised of spontaneous speech collected by both near- and far-field microphones. In addition to the audio sensors, the seminars were also recorded by calibrated video cameras. This simultaneous audio-visual data capture enables the realistic evaluation of component technologies as was never possible with earlier data bases. Dusan Macho, Jaume Padrell, Alberto Abad, Climent Nadeu, Javier Hernando, John W. McDonough, Matthias Wölfel, Ulrich Klee, Maurizio Omologo, Alessio Brutti, Piergiorgio Svaizer, Gerasimos Potamianos, Stephen M. Chu |
ICME | 9 |
| 2005 | Oriented global coherence field for the estimation of the head orientation in smart rooms equipped with distributed microphone arrays
Alessio Brutti, Maurizio Omologo, Piergiorgio Svaizer |
INTERSPEECH | 2 |
| 2004 | Weighted autocorrelation-based F0 estimation for distant-talking interaction with a distributed microphone networkabstractA distant-talking scenario is addressed, where a distributed microphone network provides multi-channel input sequences to process for speaker modeling purposes. Possible related applications are speaker tracking and distant-talking speech recognition, given a noisy and reverberant environment with one or more speakers. The paper investigates on the use of a multi-channel version of a weighted autocorrelation (WAUTOC) based F0 estimation method, with the purpose of deriving a common excitation model. Experiments conducted on a real database show the advantages and the robustness of the proposed method in extracting the fundamental frequency with no regard about the microphone and talker position as well as head orientation. Luca Armani, Maurizio Omologo |
ICASSP (1) | 2 |
| 2004 | On the use of a weighted autocorrelation based fundamental frequency estimation for a multidimensional speech inputabstractThe problem of computing the fundamental frequency F0 in an accurate way is a known and still partially unsolved problem, especially given a noisy speech input. In this work, a distanttalking scenario is addressed, where a distributed microphone network provides multi-channel input sequences to process for speaker modeling purposes. Given this context, one may process in an independent way each channel and then apply a majority vote or other fusion methods. Otherwise, the redundancy across the channels can be exploited jointly by processing the different signals to obtain a more reliable and robust F0 estimation. The paper investigates the use of a multi-channel version of a Weighted Autocorrelation(WAUTOC)-based F0 estimation technique. A postprocessing corrective step is introduced to improve the resulting F0 accuracy. Experiments conducted on a real database show the advantages and the robustness of the proposed method in extracting the fundamental frequency with no regard about the microphone and talker position as well as the head orientation. Federico Flego, Luca Armani, Maurizio Omologo |
INTERSPEECH | 3 |
| 2003 | Use of parallel recognizers for robust in-car speech interactionabstractThis paper refers to an activity under way at the speech recognition technology level for the development of a hands-free dialogue interaction system in the car environment. The use of a set of HMM recognizers, running in parallel, is being investigated in order to ensure low complexity, modularity, fast response, and to allow a real-time reconfiguration of the language models and grammars according to the policy indicated by natural language understanding and dialogue manager modules. A corpus of spontaneous speech interactions was collected using the Wizard-of-Oz method in a real driving situation with a microphone placed far from the driver. The use of parallel recognition units, each specialized on a given geographical domain, was explored using the resulting real corpus. Experiments show the advantage of selecting the recognized sentence according to the maximum likelihood among the active units when compared to the use of a single language model based on a very large vocabulary. Luca Cristoforetti, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer |
ICASSP (1) | 3 |
| 2003 | Use of a CSP-based voice activity detector for distant-talking ASR
Luca Armani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer |
INTERSPEECH | 3 |
| 2002 | On the joint use of noise reduction and MLLR adaptation for in-car hands-free speech recognitionabstractThis paper refers to an activity under way at the speech recognition technology level for the development of a hands-free dialogue interaction system in the car environment. The work here presented concerns the use of two noise reduction techniques, as well as of MLLR adaptation, for recognition error reduction in low and medium complexity tasks, namely connected digits and spelling with or without bigram/trigram statistical constraints. Experiments are based on the use of SpeechDat Car database, a corpus collected under real noisy conditions. Results show the additive improvements in performance, obtained by adopting noise reduction techniques and MLLR adaptation. Marco Matassoni, Maurizio Omologo, Alfiero Santarelli, Piergiorgio Svaizer |
ICASSP | 2 |
| 2002 | Hidden Markov model training with contaminated speech material for distant-talking speech recognition
Marco Matassoni, Maurizio Omologo, Diego Giuliani, Piergiorgio Svaizer |
Comput. Speech Lang. | 2 |
| 2001 | Use of real and contaminated speech for training of a hands-free in-car speech recognizerabstractA database of in-car speech for the Italian language was collected under the European projects SpeechDatCar and VODIS II. It consists of 600 sessions recorded under various noise and driving conditions and includes close-talk signals and far microphone signals for hands-free interaction. This paper describes some recognition experiments on two tasks conceived on a portion of this database: connected digit sequences and isolated command words. Recognition rate achieved by means of HMMs trained on real in-car speech is compared with that accomplished by a speech contamination approach, which aims at simulating in-car data starting from a clean speech corpus. Recognition performance is also analyzed as a function of the different noise conditions and of the consequent SNR at the far microphones. Finally, the effect of HMM adaptation is investigated in order to tune the recognizer on the conditions of the various sessions. Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer |
INTERSPEECH | 2 |
| 2000 | Hands-free speech recognition using a filtered clean corpus and incremental HMM adaptationabstractA challenging scenario is addressed in which a hands-free speech recognizer operates in a noisy office environment with incremental model adaptation functionalities. The use of a single far microphone as well as that of a microphone array input are investigated. In a previous work it was shown that the acoustic mismatch, remaining after the application of microphone array processing, can be further reduced by conditioning hidden Markov models to operating acoustic conditions. Conditioned HMMs are models trained using the filtered version of a clean corpus, which is speech material better representing noisy real environments. Afterwards, conditioned models are used as initial models for unsupervised incremental adaptation. Experimental results of connected digit recognition show that models trained with filtered clean speech allows to obtain better recognition performance than models trained with clean speech. Furthermore, results show a significant performance increase when incremental adaptation is applied, even after recognition of few utterances. Marco Matassoni, Maurizio Omologo, Diego Giuliani |
ICASSP | 2 |
| 2000 | Annotation of a Multichannel Noisy Speech Corpus
Luca Cristoforetti, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer, Enrico Zovato |
LREC | 3 |
| 1999 | Training of HMM with filtered speech material for hands-free recognitionabstractThis paper addresses the problem of hands-free speech recognition in a noisy office environment. An array of six omnidirectional microphones and a corresponding time delay compensation module are used to provide a beamformed signal as input to a HMM-based recognizer. Training of HMMs is performed either using a clean speech database or using a filtered version of the same database. Filtering consists in a convolution with the acoustic impulse response between the speaker and microphone, to reproduce the reverberation effect. Background noise is summed to provide the desired SNR. The paper shows that the new models trained on these data perform better than the baseline ones. Furthermore, the paper investigates on maximum likelihood linear regression (MLLR) adaptation of the new models. It is shown that a further performance improvement is obtained, allowing to reach a 98.7% WRR in a connected digit recognition task, when the talker is at 1.5 m distance from the array. Diego Giuliani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer |
ICASSP | 3 |
| 1998 | Experiments of HMM adaptation for hands-free connected digit recognitionabstractA scenario concerning hands-free connected digit recognition in a noisy office environment is investigated. An array of six omnidirectional microphones and a corresponding time delay compensation module are used to provide a beamformed signal as input to a hidden Markov model (HMM) based recognizer. Two different techniques of phone HMM adaptation have been considered, to reduce the mismatch between training and test conditions. Adaptation material and test material were collected in two different sessions. Results show that a digit accuracy close to 98% can be achieved when the talker is at 1.5 m distance from the array. This result has to be compared with 99.5% accuracy obtained by using a close-talk microphone. Diego Giuliani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer |
ICASSP | 3 |
| 1998 | Environmental conditions and acoustic transduction in hands-free speech recognition
Maurizio Omologo, Piergiorgio Svaizer, Marco Matassoni |
Speech Commun. | 1 |
| 1997 | Microphone array based speech recognition with different talker-array positionsabstractThe use of a microphone array for hands-free continuous speech recognition in noisy and reverberant environment is investigated. An array of eight omnidirectional microphones was placed at different angles and distances from the talker. A time delay compensation module was used to provide a beamformed signal as input to a hidden Markov model (HMM) based recognizer. A phone HMM adaptation, based on a small amount of phonetically rich sentences, further improved the recognition rate obtained by applying only beamforming. These results were confirmed both by experiments conducted in a noisy and reverberant environment and by simulations. In the latter case, different conditions were recreated by using the image method to reproduce synthetic versions of the array microphone signals. Maurizio Omologo, Marco Matassoni, Piergiorgio Svaizer, Diego Giuliani |
ICASSP | 1 |
| 1997 | Acoustic source location in a three-dimensional space using crosspower spectrum phaseabstractA microphone array can be used to locate a dominant acoustic source in a given environment. This capability is successfully employed to locate an active talker in teleconferencing or other multi-speaker applications. In this work the source location is obtained in two steps: (1) a time difference of arrival (TDOA) computation between the signals of the array; (2) an "optimal" source location based on the interchannel delay estimates and on a geometrical description of the sensor arrangement. The crosspower spectrum phase technique was used for TDOA estimation, while a maximum likelihood approach was followed to derive the source coordinates. Source location experiments in a three-dimensional space were performed by means of an array of 8 microphones. For this purpose both a loudspeaker and a real talker were used to collect data in a large noisy and reverberant room. Piergiorgio Svaizer, Marco Matassoni, Maurizio Omologo |
ICASSP | 3 |
| 1997 | Automatic diphone extraction for an Italian text-to-speech synthesis system
Bianca Angelini, Claudia Barolo, Daniele Falavigna, Maurizio Omologo, Stefano Sandri |
EUROSPEECH | 4 |
| 1997 | Use of different microphone array configurations for hands-free speech recognition in noisy and reverberant environmentabstractIn this work hands-free continuous speech recognition based on microphone arrays is investigated. A set of experiments was carried out using arrays having different numbers of omnidirectional microphones as well as different configurations. Both real and simulated array signals, generated by means of the image method, were used. An enhanced input to a recognizer based on Hidden Markov Models was obtained by a time delay compensation module providing a beamformed signal. HMM adaptation was used to realign the recognizer acoustic modeling to the given acoustic condition. Diego Giuliani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer |
EUROSPEECH | 3 |
| 1997 | Use of the crosspower-spectrum phase in acoustic event locationabstractThe article reports on the use of crosspower-spectrum phase (CSP) analysis as an accurate time delay estimation (TDE) technique. It is used in a microphone array system for the location of acoustic events in noisy and reverberant environments. A corresponding coherence measure (CM) and its graphical representation are introduced to show the TDE accuracy. Using a two-microphone pair array, real experiments show less than a 10 cm average location error in a 6 m/spl times/6 m area. Maurizio Omologo, Piergiorgio Svaizer |
IEEE Trans. Speech Audio Process. | 1 |
| 1996 | Acoustic source location in noisy and reverberant environment using CSP analysisabstractA linear four microphone array can be employed for acoustic event location in a real environment using an accurate time delay estimation. This paper refers to the use of a specific technique, based on crosspower spectrum phase (CSP) analysis, that yielded accurate location performance. The behavior of this technique is investigated under different noise and reverberation conditions. Real experiments as well as simulations were conducted to analyze a wide variety of situations. Results show system robustness at quite critical environmental conditions. Maurizio Omologo, Piergiorgio Svaizer |
ICASSP | 1 |
| 1996 | Experiments of speech recognition in a noisy and reverberant environment using a microphone array and HMM
Diego Giuliani, Maurizio Omologo, Piergiorgio Svaizer |
ICSLP | 2 |
| 1995 | Hands free continuous speech recognition in noisy environment using a four microphone arrayabstractThis paper describes advances in the use of HMM based technology for speaker independent continuous speech recognition, in noisy environment, under hands free interaction mode. For this purpose an array of four omnidirectional microphones is employed as the acquisition system. The processing of phase information in the cross-power spectrum provides the capability both of locating the talker position and of reconstructing an enhanced speech spectrum. Two enhancement techniques are described, that provide recognition improvement in the case of clean input speech as well as under different adverse conditions. The results refer to the use of a new multichannel corpus, collected in a real environment by a microphone array as well as a close-talk microphone. Diego Giuliani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer |
ICASSP | 3 |
| 1995 | Robust continuous speech recognition using a microphone array
Diego Giuliani, Marco Matassoni, Maurizio Omologo, Piergiorgio Svaizer |
EUROSPEECH | 3 |
| 1994 | Acoustic event localization using a crosspower-spectrum phase based techniqueabstractLinear microphone arrays can be employed for acoustic event localization in a noisy environment using time delay estimation. Three techniques are investigated that allow delay estimation, namely normalized cross correlation, LMS adaptive filters, crosspower-spectrum phase: they are combined with a bidimensional representation, the coherence measure, in order to emphasize information that can be exploited for estimating position of both non-moving and moving acoustic sources. To compare the given techniques, different acoustic sources were considered, that generated events in different positions in space. Expressing performance in terms of accuracy of the wavefront direction angle, experiments showed that the crosspower-spectrum phase based technique outperforms the other two. This technique provided very promising preliminary results also in terms of source position estimation.> Maurizio Omologo, Piergiorgio Svaizer |
ICASSP (2) | 1 |
| 1994 | Speaker independent continuous speech recognition using an acoustic-phonetic Italian corpusabstractThe objective of this paper is to describe the activity that is being carried out at IRST laboratories for the development of an HMM-based speaker independent continuous speech recognition system for the Italian language. The recognition system is trained and tested using the acoustic-phonetic continuous speech portion of the APASCI corpus. Acoustic modeling is based on the use of Continuous Density HMMs with gaussian mixture observation densities. As a baseline, a set of 38 Context Independent Units was evaluated using different numbers of mixture components. Then, two other classes of Context Dependent Unit sets were considered, that provide different performance and system complexity. Performance, expressed in terms of Phone loop recognition accuracy and Word loop recognition accuracy, shows an improvement using both of these classes of unit sets, with respect to the baseline. I. INTRODUCTION A baseline of a speaker independent continuous speech recognition system for the Italian ... Bianca Angelini, Fabio Brugnara, Daniele Falavigna, Diego Giuliani, Roberto Gretter, Maurizio Omologo |
ICSLP | 6 |
| 1994 | Talker localization and speech recognition using a microphone array and a cross-powerspectrum phase analysisabstractMismatch in training and testing conditions reduces considerably the performance of a speaker-independent HMM-based continuous speech recognizer. Compensation of this mismatch can avoid the complex and time-consuming retraining of the recognizer. This paper describes an acquisition system based on a four omnidirectional microphone array that was employed to reproduce a "beamformed" version of the original acoustic messages acquired in a noisy and reverberant environment, with a talker-microphone distance of one meter. In this preliminary activity, some simple noise compensation techniques (i.e. a Mean Spectrum based Enhancement and a Cepstrum Mean Subtraction) were incorporated in this preprocessing stage to obtain an enhanced version of the given utterance. Feeding a clean-condition trained continuous speech recognizer with enhanced signals led to a significant improvement of performance, if compared to the use of unprocessed single-microphone signals as input. I. INTRODUCTION Perfo... Diego Giuliani, Maurizio Omologo, Piergiorgio Svaizer |
ICSLP | 2 |
| 1993 | Automatic segmentation and labeling of English and Italian speech databases
Bianca Angelini, Fabio Brugnara, Daniele Falavigna, Diego Giuliani, Roberto Gretter, Maurizio Omologo |
EUROSPEECH | 6 |
| 1993 | A baseline of a speaker independent continuous speech recognizer of Italian
Bianca Angelini, Fabio Brugnara, Daniele Falavigna, Diego Giuliani, Roberto Gretter, Maurizio Omologo |
EUROSPEECH | 6 |
| 1993 | Talker localization and speech enhancement in a noisy environment using a microphone array based acquisition systemabstractThis paper deals with the use of linear microphone arrays for detection, localization and enhancement of a generic acoustic message produced in a noisy environment. A CrosspowerSpectrum Phase based analysis and a Coherence Measure representation are presented, that allow an accurate time delay estimation employed for the acoustic source position hypothesis. Preliminary results in terms of source localization accuracy are given. Once source position is estimated, an enhanced version of the original acoustic message is derived, that can represent the input for a speech recognition system. Keywords: Microphone Arrays, Talker Localization, Speech Enhancement. 1. INTRODUCTION Automatic speech recognizer performance often degrades drastically when employed in conditions that are different with respect to those for which they were designed. One reason for this is environmental noise. Another reason can be inconsistent use of the acoustic transducers during message acquisition. For example,... Maurizio Omologo, Piergiorgio Svaizer |
EUROSPEECH | 1 |
| 1993 | Automatic segmentation and labeling of speech based on Hidden Markov Models
Fabio Brugnara, Daniele Falavigna, Maurizio Omologo |
Speech Commun. | 3 |
| 1992 | A family of parallel hidden Markov modelsabstractStochastic signal models represent a powerful tool for automatic speech recognition. A particular type of stochastic modeling based on first-order hidden Markov models (HMMs), has been increasingly popular, because it has a solid theoretical basis and offers practical advantages. The authors extend the standard HMM theory to parallel hidden Markov models (PHMMs). The parallel model consists of two statistically related HMMs. This configuration has mixture densities of HMM observations whose weights can be made variable depending on the probability of other HMMs being in certain states. This allows one to dynamically adapt observation statistics to acoustic contexts. Some preliminary experiments have been carried out in order to compare the PHMMs with standard HMMs and the results are presented.> Fabio Brugnara, Renato De Mori, Diego Giuliani, Maurizio Omologo |
ICASSP | 4 |
| 1992 | A HMM-based system for automatic segmentation and labeling of speech
Fabio Brugnara, Daniele Falavigna, Maurizio Omologo |
ICSLP | 3 |
| 1992 | Improved connected digit recognition using spectral variation functions
Fabio Brugnara, Renato De Mori, Diego Giuliani, Maurizio Omologo |
ICSLP | 4 |
| 1991 | A parallel HMM approach to speech recognition
Fabio Brugnara, Renato De Mori, Diego Giuliani, Maurizio Omologo |
EUROSPEECH | 4 |
| 1991 | A preliminary statistical evaluation of manual and automatic segmentation discrepancies
Piero Cosi, Daniele Falavigna, Maurizio Omologo |
EUROSPEECH | 3 |
| 1989 | The computation and some spectral considerations on line spectrum pairs (LSP)
Maurizio Omologo |
EUROSPEECH | 1 |