EDBT 2026 Demo / reviewers in the wild / expert
Alessio Brutti
dblp:77/6379
· DBLP profile ↗
46ranked-venue papers
12as first author
23since 2021 · last 2026
0000-0003-4146-3071ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 9 first-author · 19 since 2021Artificial intelligence and machine learning · 23 · 6 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Phonetic-based Ranking for Improved Pseudo-Labeling in Low-Resource ASR
Marco Matassoni, Roberto Gretter, Falavigna Daniele, Mohamed Nabih Ali, Alessio Brutti, Matteo Negri, Mauro Cettolo, Marco Gaido, Sara Papi, Luisa Bentivogli |
LREC | 5 |
| 2025 | EFL-PEFT: A communication Efficient Federated Learning framework using PEFT sparsification for ASRabstractFederated Learning (FL) has garnered substantial interest in training different speech-based tasks (e.g. automatic speech recognition (ASR), and other speech classification tasks): recently, fine-tuning pre-trained self-supervised models for different speech-based tasks has shown promising performance and been successfully applied in FL settings. Nevertheless, fine-tuning these architectures is computationally burdensome and not affordable in several real-time settings. Moreover, the communication costs of transferring all the model parameters for the aggregation stage is critically high. As an alternative approach, parameter-efficient fine-tuning (PEFT) approaches provide promising performance without changing the backbone of the pre-trained model. PEFT has been fruitfully applied, in a variety of flavours, for ASR in central training configurations while only few works investigate its use in FL settings. In this paper, we consolidate the use of PEFT for ASR with pre-trained models, demonstrating that it enables efficient FL reducing the amount of parameters to share with respect to full fine-tuning. We also explore combining PEFT with sparsification methods to further reduce communication cost by transmitting only a fraction of the adapter parameters. Additionally, we show that agglomerating adapters using "FedAvg" is compatible with differential privacy, aligning with trends observed in other domains. Our proposed approach is supported by experimental analysis on ASR using two public datasets, as well as on intent classification tasks. Mohamed Nabih Ali, Daniele Falavigna, Alessio Brutti |
ICASSP | 3 |
| 2025 | Large Language Models are Strong Audio-Visual Speech Recognition LearnersabstractMultimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped with (automatic) speech recognition (ASR) abilities by just concatenating the audio tokens, computed with an audio encoder, and the text tokens to achieve state-of-the-art results. On the contrary, tasks like visual and audio-visual speech recognition (VSR/AVSR), which also exploit noise-invariant lip movement information, have received little or no attention. To bridge this gap, we propose Llama-AVSR, a new MLLM with strong audio-visual speech recognition capabilities. It leverages pre-trained audio and video encoders to produce modality-specific tokens which, together with the text tokens, are processed by a pre-trained LLM (e.g., Llama3.1-8B) to yield the resulting response in an auto-regressive fashion. Llama-AVSR requires a small number of trainable parameters as only modality-specific projectors and LoRA modules are trained whereas the multi-modal encoders and LLM are kept frozen. We evaluate our proposed approach on LRS3, the largest public AVSR benchmark, and we achieve new state-of-the-art results for the tasks of ASR and AVSR with a WER of 0.79% and 0.77%, respectively. To bolster our results, we investigate the key factors that underpin the effectiveness of Llama-AVSR: the choice of the pre-trained encoders and LLM, the efficient integration of LoRA modules, and the optimal performance-efficiency trade-off obtained via modality-aware compression rates. Umberto Cappellazzo, Honglie Chen, Pingchuan Ma 0001, Stavros Petridis, Daniele Falavigna, Alessio Brutti, Maja Pantic |
ICASSP | 7 |
| 2025 | Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
Umberto Cappellazzo, Stavros Petridis, Daniele Falavigna, Alessio Brutti |
INTERSPEECH | 5 |
| 2025 | Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource LanguagesabstractLarge language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks.However, their applicability is still less explored in low-resource settings.This work investigates the use of Speech LLMs for lowresource Automatic Speech Recognition using the SLAM-ASR framework, where a trainable lightweight projector connects a speech encoder and a LLM.Firstly, we assess training data volume requirements to match Whisper-only performance, reemphasizing the challenges of limited data.Secondly, we show that leveraging mono-or multilingual projectors pretrained on high-resource languages reduces the impact of data scarcity, especially with small training sets.Using multilingual LLMs (EuroLLM, Salamandra) with whisper-large-v3-turbo, we evaluate performance on several public benchmarks, providing insights for future research on optimizing Speech LLMs for lowresource languages and multilinguality. Seraphina Fong, Marco Matassoni, Alessio Brutti |
INTERSPEECH | 3 |
| 2025 | An Effective Training Framework for Light-Weight Automatic Speech Recognition Models
Abdul Hannan, Alessio Brutti, Shah Nawaz, Mubashir Noman |
INTERSPEECH | 2 |
| 2025 | Granary: Speech Recognition and Translation Dataset in 25 European Languages
Nithin Rao Koluguri, Monica Sekoyan, George Zelenfroynd, Sasha Meister, Shuoyang Ding, Sofia Kostandian, He Huang 0012, Nikolay Karpov, Jagadeesh Balam, Vitaly Lavrukhin, Yifan Peng 0003, Sara Papi, Marco Gaido, Alessio Brutti, Boris Ginsburg |
INTERSPEECH | 14 |
| 2025 | Automatic detection of speech sound disorders in German-speaking children: augmenting the data with typically developed speech
Darline Monika Marx, Marco Matassoni, Alessio Brutti |
INTERSPEECH | 3 |
| 2024 | MOSEL: 950, 000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU LanguagesabstractMarco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih, Matteo Negri. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti, Mauro Cettolo, Roberto Gretter, Marco Matassoni, Mohamed Nabih Ali, Matteo Negri |
EMNLP | 4 |
| 2024 | Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters
Umberto Cappellazzo, Daniele Falavigna, Alessio Brutti |
INTERSPEECH | 3 |
| 2024 | End-to-end integration of speech separation and voice activity detection for low-latency diarization of telephone conversations
Giovanni Morrone, Samuele Cornell, Luca Serafini, Enrico Zovato, Alessio Brutti, Stefano Squartini |
Speech Commun. | 5 |
| 2023 | An Investigation of the Combination of Rehearsal and Knowledge Distillation in Continual Learning for Spoken Language Understanding
Umberto Cappellazzo, Daniele Falavigna, Alessio Brutti |
INTERSPEECH | 3 |
| 2023 | Sequence-Level Knowledge Distillation for Class-Incremental End-to-End Spoken Language Understanding
Umberto Cappellazzo, Muqiao Yang, Daniele Falavigna, Alessio Brutti |
INTERSPEECH | 4 |
| 2023 | Direct enhancement of pre-trained speech embeddings for speech processing in noisy conditions
Mohamed Nabih Ali, Alessio Brutti, Daniele Falavigna |
Comput. Speech Lang. | 2 |
| 2023 | An experimental review of speaker diarization methods with application to two-speaker conversational telephone speech recordingsabstractWe performed an experimental review of current diarization systems for the conversational telephone speech (CTS) domain. In detail, we considered a total of eight different algorithms belonging to clustering-based, end-to-end neural diarization (EEND), and speech separation guided diarization (SSGD) paradigms. We studied the inference-time computational requirements and diarization accuracy on four CTS datasets with different characteristics and languages. We found that, among all methods considered, EEND-vector clustering (EEND-VC) offers the best trade-off in terms of computing requirements and performance. More in general, EEND models have been found to be lighter and faster in inference compared to clustering-based methods. However, they also require a large amount of diarization-oriented annotated data. In particular EEND-VC performance in our experiments degraded when the dataset size was reduced, whereas self-attentive EEND (SA-EEND) was less affected. We also found that SA-EEND gives less consistent results among all the datasets compared to EEND-VC, with its performance degrading on long conversations with high speech sparsity. Clustering-based diarization systems, and in particular VBx, instead have more consistent performance compared to SA-EEND but are outperformed by EEND-VC. The gap with respect to this latter is reduced when overlap-aware clustering methods are considered. SSGD is the most computationally demanding method, but it could be convenient if speech recognition has to be performed. Its performance is close to SA-EEND but degrades significantly when the training and inference data characteristics are less matched. Luca Serafini, Samuele Cornell, Giovanni Morrone, Enrico Zovato, Alessio Brutti, Stefano Squartini |
Comput. Speech Lang. | 5 |
| 2022 | End-to-End Low Resource Keyword Spotting Through Character Recognition and Beam-Search Re-ScoringabstractThis paper describes an end-to-end approach to perform keyword spotting with a pre-trained acoustic model that uses recurrent neural networks and connectionist temporal classification loss. Our approach is specifically designed for low-resource keyword spotting tasks where extremely small amounts of in-domain data are available to train the system. The pre-trained model, largely used in ASR tasks, is fine-tuned on in-domain audio recordings. In inference the model output is matched against the set of predefined keywords using a beam-search re-scoring based on the edit distance.We demonstrate that this approach significantly outperforms the best state-of-the art systems on a well known keyword spotting benchmark, namely "google speech commands". Moreover, com-pared against state-of-the-art methods, our proposed approach is extremely robust in case of limited in domain training material. We show that a very small performance reduction is observed when fine tuning with a very small fraction (around 5%) of the training set.We report an extensive set of experiments on two keyword spotting tasks, varying training sizes and correlating keyword classification accuracy with character error rates provided by the system. We also report an ablation study to assess on the contribution of the out-of-domain pre-training and of the beam-search re-scoring. Ephrem Tibebe Mekonnen, Alessio Brutti, Daniele Falavigna |
ICASSP | 2 |
| 2022 | Scalable Neural Architectures for End-to-End Environmental Sound ClassificationabstractSound Event Detection (SED) is a complex task simulating human ability to recognize what is happening in the surrounding from auditory signals only. This technology is a crucial asset in many applications such as smart cities. Here, urban sounds can be detected and processed by embedded devices in an Internet of Things (IoT) to identify meaningful events for municipalities or law enforcement. However, while current deep learning techniques for SED are effective, they are also resource- and power-hungry, thus not appropriate for pervasive battery-powered devices. In this paper, we propose novel neural architectures based on PhiNets for real-time acoustic event detection on microcontroller units. The proposed models are easily scalable to fit the hardware requirements and can operate both on spectrograms and waveforms. In particular, our architectures achieve state-of-the-art performance on UrbanSound8K in spectrogram classification (around 77%) with extreme compression factors (99.8%) with respect to current state-of-the-art architectures. Francesco Paissan, Alberto Ancilotto, Alessio Brutti, Elisabetta Farella |
ICASSP | 3 |
| 2022 | Is Cross-Attention Preferable to Self-Attention for Multi-Modal Emotion Recognition?abstractHumans express their emotions via facial expressions, voice intonation and word choices. To infer the nature of the underlying emotion, recognition models may use a single modality, such as vision, audio, and text, or a combination of modalities. Generally, models that fuse complementary information from multiple modalities outperform their uni-modal counterparts. However, a successful model that fuses modalities requires components that can effectively aggregate task-relevant information from each modality. As cross-modal attention is seen as an effective mechanism for multi-modal fusion, in this paper we quantify the gain that such a mechanism brings compared to the corresponding self-attention mechanism. To this end, we implement and compare a cross-attention and a self-attention model. In addition to attention, each model uses convolutional layers for local feature extraction and recurrent layers for global sequential modelling. We compare the models using different modality combinations for a 7-class emotion classification task using the IEMOCAP dataset. Experimental results indicate that albeit both models improve upon the state-of-the-art in terms of weighted and unweighted accuracy for tri- and bi-modal configurations, their performance is generally statistically comparable. The code to replicate the experiments is available at https://github.com/smartcameras/SelfCrossAttn Vandana Rajan, Alessio Brutti, Andrea Cavallaro |
ICASSP | 2 |
| 2022 | Enhancing Embeddings for Speech Classification in Noisy Conditions
Mohamed Nabih Ali, Alessio Brutti, Daniele Falavigna |
INTERSPEECH | 2 |
| 2022 | Low-Latency Speech Separation Guided Diarization for Telephone ConversationsabstractIn this paper, we carry out an analysis on the use of speech separation guided diarization (SSGD) in telephone conversations. SSGD performs diarization by separating the speakers signals and then applying voice activity detection on each estimated speaker signal. In particular, we compare two low-latency speech separation models. Moreover, we show a post-processing algorithm that significantly reduces the false alarm errors of a SSGD pipeline. We perform our experiments on two datasets: Fisher Corpus Part 1 and CALLHOME, evaluating both separation and diarization metrics. Notably, our SSGD DPRNN-based online model achieves 11.1% DER on CALL-HOME, comparable with most state-of-the-art end-to-end neural diarization models despite being trained on an order of magnitude less data and having considerably lower latency, i.e., 0.1 vs. 10 seconds. We also show that the separated signals can be readily fed to a speech recognition back-end with performance close to the oracle source signals. Giovanni Morrone, Samuele Cornell, Desh Raj, Luca Serafini, Enrico Zovato, Alessio Brutti, Stefano Squartini |
SLT | 6 |
| 2022 | Audio-Visual Tracking of Concurrent SpeakersabstractAudio-visual tracking of an unknown number of concurrent speakers in 3D is a challenging task, especially when sound and video are collected with a compact sensing platform. In this paper, we propose a tracker that builds on generative and discriminative audio-visual likelihood models formulated in a particle filtering framework. We localize multiple concurrent speakers with a de-emphasized acoustic map assisted by the image detection-derived 3D video observations. The 3D multi-modal observations are either assigned to existing tracks for discriminative likelihood computation or used to initialize new tracks. The generative likelihoods rely on color distribution of the target and the de-emphasized acoustic map value. Experiments on AV16.3 and CAV3D datasets show that the proposed tracker outperforms the uni-modal trackers and the state-of-the-art approaches both in 3D and on the image plane. Xinyuan Qian 0001, Alessio Brutti, Oswald Lanz, Maurizio Omologo, Andrea Cavallaro |
IEEE Trans. Multim. | 2 |
| 2021 | Robust Latent Representations Via Cross-Modal Translation and AlignmentabstractMulti-modal learning relates information across observation modalities of the same physical phenomenon to leverage complementary information. Most multi-modal machine learning methods require that all the modalities used for training are also available for testing. This is a limitation when signals from some modalities are unavailable or severely degraded. To address this limitation, we aim to improve the testing performance of uni-modal systems using multiple modalities during training only. The proposed multi-modal training framework uses cross-modal translation and correlation-based latent space alignment to improve the representations of a worse performing (or weaker) modality. The translation from the weaker to the better performing (or stronger) modality generates a multi-modal intermediate encoding that is representative of both modalities. This encoding is then correlated with the stronger modality representation in a shared latent space. We validate the proposed framework on the AVEC 2016 dataset (RECOLA) for continuous emotion recognition and show the effectiveness of the framework that achieves state-of- the-art (uni-modal) performance for weaker modalities. Vandana Rajan, Alessio Brutti, Andrea Cavallaro |
ICASSP | 2 |
| 2021 | Learning to Rank Microphones for Distant Speech RecognitionabstractFully exploiting ad-hoc microphone networks for distant speech recognition is still an open issue. Empirical evidence shows that being able to select the best microphone leads to significant improvements in recognition without any additional effort on front-end processing. Current channel selection techniques either rely on signal, decoder or posterior-based features. Signal-based features are inexpensive to compute but do not always correlate with recognition performance. Instead decoder and posterior-based features exhibit better correlation but require substantial computational resources. In this work, we tackle the channel selection problem by proposing MicRank, a learning to rank framework where a neural network is trained to rank the available channels using directly the recognition performance on the training set. The proposed approach is agnostic with respect to the array geometry and type of recognition back-end. We investigate different learning to rank strategies using a synthetic dataset developed on purpose and the CHiME-6 data. Results show that the proposed approach is able to considerably improve over previous selection techniques, reaching comparable and in some instances better performance than oracle signal-based measures. Samuele Cornell, Alessio Brutti, Marco Matassoni, Stefano Squartini |
Interspeech | 2 |
| 2020 | Supervised Online Diarization with Sample Mean Loss for Multi-Domain DataabstractRecently, a fully supervised speaker diarization approach was proposed (UIS-RNN) which models speakers using multiple instances of a parameter-sharing recurrent neural network. In this paper we propose qualitative modifications to the model that significantly improve the learning efficiency and the overall diarization performance. In particular, we introduce a novel loss function, we called Sample Mean Loss and we present a better modelling of the speaker turn behaviour, by devising an analytical expression to compute the probability of a new speaker joining the conversation. In addition, we demonstrate that our model can be trained on fixed-length speech segments, removing the need for speaker change information in inference. Using x-vectors as input features, we evaluate our proposed approach on the multi-domain dataset employed in the DIHARD-II challenge: our online method improves with respect to the original UIS-RNN and achieves similar performance to an offline agglomerative clustering baseline using PLDA scoring. Enrico Fini, Alessio Brutti |
ICASSP | 2 |
| 2019 | Accurate Target Annotation in 3D from Multimodal StreamsabstractAccurate annotation is fundamental to quantify the performance of multi-sensor and multi-modal object detectors and trackers. However, invasive or expensive instrumentation is needed to automatically generate these annotations. To mitigate this problem, we present a multi-modal approach that leverages annotations from reference streams (e.g. individual camera views) and measurements from unannotated additional streams (e.g. audio) to infer 3D trajectories through an optimization. The core of our approach is a multi-modal extension of Bundle Adjustment with a cross-modal correspondence detection that selectively uses measurements in the optimization. We apply the proposed approach to fully annotate a new multi-modal and multi-view dataset for multi-speaker 3D tracking. Oswald Lanz, Alessio Brutti, Alessio Xompero, Xinyuan Qian 0001, Maurizio Omologo, Andrea Cavallaro |
ICASSP | 2 |
| 2019 | Neural Network Distillation on IoT Platforms for Sound Event DetectionabstractIn most classification tasks, wide and deep neural networks perform and generalize better than their smaller counterparts, in particular when they are exposed to large and heterogeneous training sets. However, in the emerging field of Internet of Things memory footprint and energy budget pose severe limits on the size and complexity of the neural models that can be implemented on embedded devices. The Student-Teacher approach is an attractive strategy to distill knowledge from a large network into smaller ones, that can fit on low-energy low-complexity embedded IoT platforms. In this paper, we consider the outdoor sound event detection task as a use case. Building upon the VGGish network, we investigate different distillation strategies to substantially reduce the classifier's size and computational cost with minimal performance losses. Experiments on the UrbanSound8K dataset show that extreme compression factors (up to 4.2 · 10−4 for parameters and 1.2 · 10−3 for operations with respect to VGGish) can be achieved, limiting the accuracy degradation from 75% to 70%. Finally, we compare different embedded platforms to analyze the trade-off between available resources and achievable accuracy. Gianmarco Cerutti, Rahul Prasad, Alessio Brutti, Elisabetta Farella |
INTERSPEECH | 3 |
| 2019 | ConflictNET: End-to-End Learning for Speech-Based Conflict Intensity EstimationabstractComputational paralinguistics aims to infer human emotions, personality traits and behavioural patterns from speech signals. In particular, verbal conflict is an important example of human-interaction behaviour, whose detection would enable monitoring and feedback in a variety of applications. The majority of methods for detection and intensity estimation of verbal conflict apply off-the-shelf classifiers/regressors to generic hand-crafted acoustic features. Generating conflict-specific features requires refinement steps and the availability of metadata, such as the number of speakers and their speech overlap duration. Moreover, most techniques treat feature extraction and regression as independent modules, which require separate training and parameter tuning. To address these limitations, we propose the first end-to-end convolutional-recurrent neural network architecture that learns conflict-specific features directly from raw speech waveforms, without using explicit domain knowledge or metadata. Additionally, to selectively focus the model on portions of speech containing verbal conflict instances, we include a global attention interface that learns the alignment between layers of the recurrent network. Experimental results on the SSPNet Conflict Corpus show that our end-to-end architecture achieves state-of-the-art performance in terms of Pearson Correlation Coefficient. Vandana Rajan, Alessio Brutti, Andrea Cavallaro |
IEEE Signal Process. Lett. | 2 |
| 2019 | Multi-Speaker Tracking From an Audio-Visual Sensing DeviceabstractCompact multi-sensor platforms are portable and thus desirable for robotics and personal-assistance tasks. However, compared to physically distributed sensors, the size of these platforms makes person tracking more difficult. To address this challenge, we propose a novel 3-D audio-visual people tracker that exploits visual observations (object detections) to guide the acoustic processing by constraining the acoustic likelihood on the horizontal plane defined by the predicted height of a speaker. This solution allows the tracker to estimate, with a small microphone array, the distance of a sound. Moreover, we apply a color-based visual likelihood on the image plane to compensate for misdetections. Finally, we use a 3-D particle filter and greedy data association to combine visual observations, color-based, and acoustic likelihoods to track the position of multiple simultaneous speakers. We compare the proposed multimodal 3-D tracker against two state-of-the-art methods on the AV16.3 dataset and on a newly collected dataset with co-located sensors, which we make available to the research community. Experimental results show that our multimodal approach outperforms the other methods both in 3-D and on the image plane. Xinyuan Qian 0001, Alessio Brutti, Oswald Lanz, Maurizio Omologo, Andrea Cavallaro |
IEEE Trans. Multim. | 2 |
| 2018 | 3D Mouth Tracking from a Compact Microphone Array Co-Located with a cameraabstractWe address the 3D audio-visual mouth tracking problem when using a compact platform with co-located audio-visual sensors, without a depth camera. In particular, we propose a multi-modal particle filter that combines a face detector and 3D hypothesis mapping to the image plane. The audio likelihood computation is assisted by video, which relies on a GCC-PHAT based acoustic map. By combining audio and video inputs, the proposed approach can cope with a reverberant and noisy environment, and can deal with situations when the person is occluded, outside the Field of View (FoV), or not facing the sensors. Experimental results show that the proposed tracker is accurate both in 3D and on the image plane. Xinyuan Qian 0001, Alessio Xompero, Andrea Cavallaro, Alessio Brutti, Oswald Lanz, Maurizio Omologo |
ICASSP | 4 |
| 2017 | 3D audio-visual speaker tracking with an adaptive particle filterabstractWe propose an audio-visual fusion algorithm for 3D speaker tracking from a localised multi-modal sensor platform composed of a camera and a small microphone array. After extracting audio-visual cues from individual modalities we fuse them adaptively using their reliability in a particle filter framework. The reliability of the audio signal is measured based on the maximum Global Coherence Field (GCF) peak value at each frame. The visual reliability is based on colour-histogram matching with detection results compared with a reference image in the RGB space. Experiments on the AV16.3 dataset show that the proposed adaptive audio-visual tracker outperforms both the individual modalities and a classical approach with fixed parameters in terms of tracking accuracy. Xinyuan Qian 0001, Alessio Brutti, Maurizio Omologo, Andrea Cavallaro |
ICASSP | 2 |
| 2017 | Optimizing DNN Adaptation for Recognition of Enhanced Speech
Marco Matassoni, Alessio Brutti, Daniele Falavigna |
INTERSPEECH | 2 |
| 2017 | Online Cross-Modal Adaptation for Audio-Visual Person Identification With Wearable CamerasabstractWe propose an audio-visual target identification approach for egocentric data with cross-modal model adaptation. The proposed approach blindly and iteratively adapts the time-dependent models of each modality to varying target appearance and environmental conditions using the posterior of the other modality. The adaptation is unsupervised and performed online; thus, models can be improved as new unlabeled data become available. In particular, accurate models do not deteriorate when a modality is underperforming thanks to an appropriate selection of the parameters in the adaptation. Importantly, unlike traditional audio-visual integration methods, the proposed approach is also useful for temporal intervals during which only one modality is available or when different modalities are used for different tasks. We evaluate the proposed method in an end-to-end multimodal person identification application with two challenging real-world datasets and show that the proposed approach successfully adapts models in presence of mild mismatch. We also show that the proposed approach is beneficial to other multimodal score fusion algorithms. Alessio Brutti, Andrea Cavallaro |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2016 | A Phase-Based Time-Frequency Masking for Multi-Channel Speech Enhancement in Domestic Environments
Alessio Brutti, Antigoni Tsiami, Athanasios Katsamanis, Petros Maragos |
INTERSPEECH | 1 |
| 2016 | On the relationship between Early-to-Late Ratio of Room Impulse Responses and ASR performance in reverberant environments
Alessio Brutti, Marco Matassoni |
Speech Commun. | 1 |
| 2015 | Multi-channel speaker verification based on total variability modelling
Maria Joana Correia, Alessio Brutti, Alberto Abad |
INTERSPEECH | 2 |
| 2014 | On the use of Early-To-Late Reverberation ratio for ASR in reverberant environmentsabstractThis work presents an analysis of distant-talking speech recognition in a variety of reverberant conditions, correlating ASR performance to the acoustic characteristics of a given propagation channel. In particular we show how, for a digit recognition task, the ASR accuracy is directly related to the Early-to-Late Reverberation ratio of the room impulse response, capturing in a single parameter the reverberation properties of a given channel independently of the setup. Consequently, this measure can be successfully considered for acoustic model training either selecting the most suitable model for a given spatial configuration, or defining the subset of RIRs to be used for the creation of multi-condition models. Experimental results on simulated data as well as on data generated with real impulse responses support our claims. Alessio Brutti, Marco Matassoni |
ICASSP | 1 |
| 2013 | Geometric contamination for GMM/UBM speaker verification in reverberant environments
Alessio Brutti, Maurizio Omologo |
INTERSPEECH | 1 |
| 2013 | Tracking of multidimensional TDOA for multiple sources with distributed microphone pairs
Alessio Brutti, Francesco Nesta |
Comput. Speech Lang. | 1 |
| 2013 | An environment aware ML estimation of acoustic radiation pattern with distributed microphone pairs
Alessio Brutti, Maurizio Omologo, Piergiorgio Svaizer |
Signal Process. | 1 |
| 2009 | Acoustic Based Surveillance System for Intrusion DetectionabstractThis paper describes a surveillance system for intrusion detection which is based only on information derived from the processing of audio signals acquired by a distributed microphone network (DMN). In particular the system exploits different acoustic features and estimates of acoustic event positions in order to detect intrusion and reject possible false alarms that may be generated by sound sources inside and outside the monitored room. An evaluation has been conducted in order to measure the performance in terms of false alarms and missed alarms in presence of acoustic events produced inside and outside a test room.The obtained results are very promising and encouraging for future works aimed at improving the actual system accuracy. Christian Zieger, Alessio Brutti, Piergiorgio Svaizer |
AVSS | 2 |
| 2008 | Localization of multiple speakers based on a two step acoustic map analysisabstractAn interface for distant-talking control of home devices requires the possibility of identifying the positions of multiple users. Acoustic maps, based either on global coherence field (GCF) or oriented global coherence field (OGCF), have already been exploited successfully to determine position and head orientation of a single speaker. This paper proposes a new method using acoustic maps to deal with the case of two simultaneous speakers. The method is based on a two step analysis of a coherence map: first the dominant speaker is localized; then the map is modified by compensating for the effects due to the first speaker and the position of the second speaker is detected. Simulations were carried out to show how an appropriate analysis of OGCF and GCF maps allows one to localize both speakers. Experiments proved the effectiveness of the proposed solution in a linear microphone array set up. Alessio Brutti, Maurizio Omologo, Piergiorgio Svaizer |
ICASSP | 1 |
| 2008 | WOZ Acoustic Data Collection for Interactive TV
Alessio Brutti, Luca Cristoforetti, Walter Kellermann, Lutz Marquardt, Maurizio Omologo |
LREC | 1 |
| 2007 | Classification of Acoustic Maps to Determine Speaker Position and Orientation from a Distributed Microphone NetworkabstractAcoustic maps created on the basis of the signals acquired by distributed networks of microphones allow to identify position and orientation of an active talker in an enclosure. In adverse situations of high background noise, high reverberation or unavailability of direct paths to the microphones, localization may fail. This paper proposes a novel approach to talker localization and estimation of head orientation based on the classification of global coherence field (GCF) or oriented GCF maps. Preliminary experiments with data obtained by simulated propagation as well as with data acquired in a real room show that the match with precalculated map models provides a robust behavior in adverse conditions. Alessio Brutti, Maurizio Omologo, Piergiorgio Svaizer, Christian Zieger |
ICASSP (4) | 1 |
| 2006 | Speaker localization based on oriented global coherence fieldabstractAbstract This paper proposes a new speaker localization method that isbased on a preliminary estimation of the head orientation. The ba-sic information on which the estimation is accomplished is calledOriented Global Coherence Field (OGCF).The new algorithm is shown to be significantly more robustthan the traditional ones so far explored. Its robustness is also dueto an effective speech activity detection, implicitly performed bya thresholding technique applied to OGCF information. To showthe performance of the proposed system, experiments were con-ducted on the NIST RT-05 Spring Evaluation source localizationtask, which is based on real recordings of lectures in noisy andreverberant environments. Index Terms : speaker localization, head orientation, microphonearrays, global coherence field. 1. Introduction Since 1990, several Speaker LOCalization (SLOC) techniqueshave been proposed as reported in [1, 2]. Most of the traditionalSLOC techniques are based on the estimation of time differencesof wavefront arrival at each sensor and on a consequent applica-tion of geometrical information to infer the acoustic source posi-tions. One of the most common techniques for Time Delay Es-timation (TDE) is based on Generalized Cross-Correlation PhaseTransform (GCC-PHAT) [3, 4]. Other effective SLOC techniquesare based on a preliminary computation of an acoustic map, asfor instance the Global Coherence Field (GCF) [5] representation,fromwhichthemostlikelysourcepositionisderivedthroughmax-imization in space.This paper aims at describing a new SLOC method that wasconceived starting from the effectiveness of the Oriented GlobalCoherence Field(OGCF), introduced in [6], which allows to char-acterize the orientation of an active speaker’s head with a satis-factory accuracy (in terms of angle error) even under reverberantconditions. By exploiting OGCF information, one can also derivemore robust speaker position estimates, since they are mostly re-lated to the propagation of a direct wavefront from a given point.On the other hand, previous SLOC techniques did not deal withthe way the sound is being radiated from a hypothesized positionin space.Theproposed method requires touseadistributed microphonenetwork similar to those available in the laboratories involved inthe EC CHIL Alessio Brutti, Maurizio Omologo, Piergiorgio Svaizer |
INTERSPEECH | 1 |
| 2005 | Automatic Speech Activity Detection, Source Localization, and Speech Recognition on the Chil Seminar CorpusabstractTo realize the long-term goal of ubiquitous computing, technological advances in multi-channel acoustic analysis are needed in order to solve several basic problems, including speaker localization and tracking, speech activity detection (SAD) and distant-talking automatic speech recognition (ASR). The European Commission integrated project CHIL, “ Computers in the Human Interaction Loop”, aims to make significant advances in these three technologies. In this work, we report the results of our initial automatic source localization, speech activity detection, and speech recognition experiments on the CHIL seminar corpus, which is comprised of spontaneous speech collected by both near- and far-field microphones. In addition to the audio sensors, the seminars were also recorded by calibrated video cameras. This simultaneous audio-visual data capture enables the realistic evaluation of component technologies as was never possible with earlier data bases. Dusan Macho, Jaume Padrell, Alberto Abad, Climent Nadeu, Javier Hernando, John W. McDonough, Matthias Wölfel, Ulrich Klee, Maurizio Omologo, Alessio Brutti, Piergiorgio Svaizer, Gerasimos Potamianos, Stephen M. Chu |
ICME | 10 |
| 2005 | Oriented global coherence field for the estimation of the head orientation in smart rooms equipped with distributed microphone arrays
Alessio Brutti, Maurizio Omologo, Piergiorgio Svaizer |
INTERSPEECH | 1 |