EDBT 2026 Demo / reviewers in the wild / expert
Samuele Cornell
dblp:251/3685
· DBLP profile ↗
38ranked-venue papers
9as first author
33since 2021 · last 2026
0000-0002-5358-1844ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 6 first-author · 25 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Recent trends in distant conversational speech recognition: A review of CHiME-7 and 8 DASR challenges
Samuele Cornell, Christoph Böddeker, Taejin Park, He Huang 0012, Desh Raj, Matthew Wiesner, Yoshiki Masuyama, Xuankai Chang, Zhongqiu Wang 0001, Stefano Squartini, L. Paola García-Perera, Shinji Watanabe 0001 |
Comput. Speech Lang. | 1 |
| 2026 | An end-to-end integration of speech separation and recognition with self-supervised learning representationabstractMulti-speaker automatic speech recognition (ASR) has gained growing attention in a wide range of applications, including conversation analysis and human–computer interaction. Speech separation and enhancement (SSE) and single-speaker ASR have witnessed remarkable performance improvements with the rapid advances in deep learning. Complex spectral mapping predicts the short-time Fourier transform (STFT) coefficients of each speaker and has achieved promising results in several SSE benchmarks. Meanwhile, self-supervised learning representation (SSLR) has demonstrated its significant advantage in single-speaker ASR. In this work, we push forward the performance of multi-speaker ASR under noisy reverberant conditions by integrating powerful SSE, SSL, and ASR models in an end-to-end manner. We systematically investigate both monaural and multi-channel SSE methods and various feature representations. Our experiments demonstrate the advantages of recently proposed complex spectral mapping and SSLRs in multi-speaker ASR. The experimental results also confirm that end-to-end fine-tuning with an ASR criterion is important to achieve state-of-the-art word error rates (WERs) even with powerful pre-trained models. Moreover, we show the performance trade-off between SSE and ASE and mitigate it with a multi-task learning framework with both SSE and ASR criteria. Yoshiki Masuyama, Xuankai Chang, Wangyou Zhang, Samuele Cornell, Nobutaka Ono, Yanmin Qian, Shinji Watanabe 0001 |
Comput. Speech Lang. | 4 |
| 2025 | Less is More: Data Curation Matters in Scaling Speech EnhancementabstractThe vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling speech enhancement data. We focus on a critical factor: prevalent quality issues in “clean” training labels within large-scale datasets. This work re-examines this phenomenon and demonstrates that, within large-scale training sets, prioritizing high-quality training data is more important than merely expanding the data volume. Experimental findings suggest that models trained on a carefully curated subset of 700 hours can outperform models trained on the 2,500 -hour full dataset. This outcome highlights the crucial role of data curation in scaling speech enhancement systems effectively. Chenda Li, Wangyou Zhang, Wei Wang 0010, Robin Scheibler, Kohei Saijo, Samuele Cornell, Yihui Fu, Marvin Sach, Zhaoheng Ni, Anurag Kumar 0003, Tim Fingscheidt, Shinji Watanabe 0001, Yanmin Qian |
ASRU | 6 |
| 2025 | URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement CompetitionabstractThe Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS to avoid this issue, they often suffer from insufficient training data. Recognizing that the comparison of speech enhancement (SE) systems prioritizes a reliable system comparison over absolute scores, we propose URGENT-PK, a novel ranking approach leveraging pairwise comparisons. URGENT-PK takes homologous enhanced speech pairs as input to predict relative quality rankings. This pairwise paradigm efficiently utilizes limited training data, as all pairwise permutations of multiple systems constitute a training instance. Experiments across multiple open test sets demonstrate URGENT-PK’s superior system-level ranking performance over state-of-the-art baselines, despite its simple network architecture and limited training data. Chenda Li, Wei Wang 0010, Wangyou Zhang, Samuele Cornell, Marvin Sach, Robin Scheibler, Kohei Saijo, Yihui Fu, Zhaoheng Ni, Anurag Kumar 0003, Tim Fingscheidt, Shinji Watanabe 0001, Yanmin Qian |
ASRU | 5 |
| 2025 | Investigation of Spatial Self-Supervised Learning and Its Application to Target Speaker Speech RecognitionabstractIn this paper, we investigate spatial self-supervised learning for target speaker speech recognition. Neural separation models can be trained in a self-supervised manner by using only multichannel mixture signals. Such a framework is typically based on a physics-informed generative model, widely studied in blind source separation (BSS). Our study focuses on the application of spatial self-supervised learning for distant speech recognition (DSR). Multi-talker DSR systems often rely on a BSS method called guided source separation (GSS), which separates target speech utterances from multichannel mixtures and their speaker activities. Its performance, however, would be limited by its assumption that each time-frequency (TF) bin contains only one sound source or noise. To address this limitation, factor models have been studied by assuming each TF bin as a sum of all the sources. We investigate both the classic BSS and self-supervised methods based on factor models by evaluating them extensively on multiple different DSR scenarios in the CHiME-8 DASR challenge. We show that the classic BSS methods based on factor models suffer from initialization sensitivity, while self-supervised learning mitigated this problem and slightly outperformed GSS. The best-performing method and audio samples are available online at https://ybando.jp/projects/icassp2025/. Yoshiaki Bando, Samuele Cornell, Satoru Fukayama, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2025 | Interspeech 2025 URGENT Speech Enhancement Challenge
Kohei Saijo, Wangyou Zhang, Samuele Cornell, Robin Scheibler, Chenda Li, Zhaoheng Ni, Anurag Kumar 0003, Marvin Sach, Yihui Fu, Wei Wang 0010, Tim Fingscheidt, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2025 | Scalable Spontaneous Speech Dataset (SSSD): Crowdsourcing Data Collection to Promote Dialogue Research
Zaid Sheikh, Shuichiro Shimizu, Siddhant Arora, Jiatong Shi, Samuele Cornell, Shinji Watanabe 0001 |
INTERSPEECH | 5 |
| 2025 | Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
Wangyou Zhang, Kohei Saijo, Samuele Cornell, Robin Scheibler, Chenda Li, Zhaoheng Ni, Anurag Kumar 0003, Marvin Sach, Wei Wang 0010, Yihui Fu, Shinji Watanabe 0001, Tim Fingscheidt, Yanmin Qian |
INTERSPEECH | 3 |
| 2025 | ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric EstimationabstractSpeech signal analysis poses significant challenges, particularly in tasks such as speech quality evaluation and profiling, where the goal is to predict multiple perceptual and objective metrics. For instance, metrics like PESQ (Perceptual Evaluation of Speech Quality), STOI (Short-Time Objective Intelligibility), and MOS (Mean Opinion Score) each capture different aspects of speech quality. However, these metrics often have different scales, assumptions, and dependencies, making joint estimation non-trivial. To address these issues, we introduce ARECHO (Autoregressive Evaluation via Chain-based Hypothesis Optimization), a chain-based, versatile evaluation system for speech assessment grounded in autoregressive dependency modeling. ARECHO is distinguished by three key innovations: (1) a comprehensive speech information tokenization pipeline; (2) a dynamic classifier chain that explicitly captures inter-metric dependencies; and (3) a two-step confidence-oriented decoding algorithm that enhances inference reliability. Experiments demonstrate that ARECHO significantly outperforms the baseline framework across diverse evaluation scenarios, including enhanced speech analysis, speech generation evaluation, and noisy speech evaluation. Furthermore, its dynamic dependency modeling improves interpretability by capturing inter-metric relationships. Across tasks, ARECHO offers reference-free evaluation using its dynamic classifier chain to support subset queries (single or multiple metrics) and reduces error propagation via confidence-oriented decoding. Jiatong Shi, Bo-Hao Su, Hye-Jin Shim, Jinchuan Tian, Samuele Cornell, Siddhant Arora, Shinji Watanabe 0001 |
NeurIPS | 6 |
| 2024 | One Model to Rule Them All ? Towards End-to-End Joint Speaker Diarization and Speech RecognitionabstractThis paper presents a novel framework for joint speaker diarization (SD) and automatic speech recognition (ASR), named SLIDAR (sliding-window diarization-augmented recognition). SLIDAR can process arbitrary length inputs and can handle any number of speakers, effectively solving "who spoke what, when" concurrently. SLIDAR leverages a sliding window approach and consists of an end-to-end diarization-augmented speech transcription (E2E DAST) model which provides, locally, for each window: transcripts, diarization and speaker embeddings. The E2E DAST model is based on an encoder-decoder architecture and leverages recent techniques such as serialized output training and "Whisper-style" prompting. The local outputs are then combined to get the final SD+ASR result by clustering the speaker embeddings to get global speaker identities. Experiments performed on monaural recordings from the AMI corpus confirm the effectiveness of the method in both close-talk and far-field speech scenarios. Samuele Cornell, Jee-Weon Jung, Shinji Watanabe 0001, Stefano Squartini |
ICASSP | 1 |
| 2024 | Resource-Efficient Separation TransformerabstractTransformers have recently achieved state-of-the-art performance in speech separation. These models, however, are computationally demanding and require a lot of learnable parameters. This paper explores Transformer-based speech separation with a reduced computational cost. Our main contribution is the development of the Resource-Efficient Separation Transformer (RE-SepFormer), a self-attention-based architecture that reduces the computational burden in two ways. First, it uses non-overlapping blocks in the latent space. Second, it operates on compact latent summaries calculated from each chunk. The RE-SepFormer reaches a competitive performance on the popular WSJ0-2Mix and WHAM! datasets in both causal and non-causal settings. Remarkably, it scales significantly better than the previous Transformer-based architectures in terms of memory and inference time, making it more suitable for processing long mixtures. Luca Della Libera, Cem Subakan, Mirco Ravanelli, Samuele Cornell, Frédéric Lepoutre, François Grondin |
ICASSP | 4 |
| 2024 | URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
Wangyou Zhang, Robin Scheibler, Kohei Saijo, Samuele Cornell, Chenda Li, Zhaoheng Ni, Jan Pirklbauer, Marvin Sach, Shinji Watanabe 0001, Tim Fingscheidt, Yanmin Qian |
INTERSPEECH | 4 |
| 2024 | Diffusion-Based Generative Modeling With Discriminative Guidance for Streamable Speech EnhancementabstractDiffusion-based generative models (DGMs) have recently attracted attention in speech enhancement (SE) research as previous works showed a remarkable generalization capability. However, DGMs are also computationally intensive, since they usually require many iterations in the reverse diffusion process (RDP), making them impractical for streaming SE systems. In this paper, we propose to use scores estimated from discriminative models in the first steps of the RDP. These discriminative-based scores require only one forward pass with the discriminative model for multiple RDP steps, thus greatly reducing computations. This approach also allows for performance improvements. We show that choosing an appropriate number of discriminative guidance steps can result in an overall model with better performance than generative and discriminative models. Furthermore, we propose a novel streamable time-domain generative model with an algorithmic latency of 50 ms, which has no significant performance degradation compared to offline models. Chenda Li, Samuele Cornell, Shinji Watanabe 0001, Yanmin Qian |
SLT | 2 |
| 2024 | ESPnet-EZ: Python-Only ESPnet For Easy Fine-Tuning And IntegrationabstractWe introduce ESPnet-EZ, an extension of the open-source speech processing toolkit ESPnet, aimed at quick and easy development of speech models. ESPnet-EZ focuses on two major aspects: (i) easy fine-tuning and inference of existing ESPnet models on various tasks and (ii) easy integration with popular deep neural network frameworks such as PyTorch-Lightning, Hugging Face transformers and datasets, and Lhotse. By replacing ESPnet design choices inherited from Kaldi with a Python-only, Bash-free interface, we dramatically reduce the effort required to build, debug, and use a new model. For example, to fine-tune a speech foundation model, ESPnet-EZ, compared to ESPnet, reduces the number of newly written code by 2.7x and the amount of dependent code by 6.7x while dramatically reducing the Bash script dependencies. The codebase of ESPnet-EZ is publicly available.1 Masao Someki, Kwanghee Choi, Siddhant Arora, Samuele Cornell, Jionghao Han, Yifan Peng 0003, Jiatong Shi, Vaibhav Srivastav, Shinji Watanabe 0001 |
SLT | 5 |
| 2024 | A Graph-Based Neural Approach to Linear Sum Assignment ProblemsabstractLinear assignment problems are well-known combinatorial optimization problems involving domains such as logistics, robotics and telecommunications. In general, obtaining an optimal solution to such problems is computationally infeasible even in small settings, so heuristic algorithms are often used to find near-optimal solutions. In order to attain the right assignment permutation, this study investigates a general-purpose learning strategy that uses a bipartite graph to describe the problem structure and a message passing Graph Neural Network (GNN) model to learn the correct mapping. Comparing the proposed structure with two existing DNN solutions, simulation results show that the proposed approach significantly improves classification accuracy, proving to be very efficient in terms of processing time and memory requirements, due to its inherent parameter sharing capability. Among the many practical uses that require solving allocation problems in everyday scenarios, we decided to apply the proposed approach to address the scheduling of electric smart meters access within an electricity distribution smart grid infrastructure, since near-real-time energy monitoring is a key element of the green transition that has become increasingly important in recent times. The results obtained show that the proposed graph-based solver, although sub-optimal, exhibits the highest scalability, compared with other state-of-the-art heuristic approaches. To foster the reproducibility of the results, we made the code available at https://github.com/aircarlo/GNN_LSAP. Carlo Aironi, Samuele Cornell, Stefano Squartini |
Int. J. Neural Syst. | 2 |
| 2024 | End-to-end integration of speech separation and voice activity detection for low-latency diarization of telephone conversations
Giovanni Morrone, Samuele Cornell, Luca Serafini, Enrico Zovato, Alessio Brutti, Stefano Squartini |
Speech Commun. | 2 |
| 2023 | Multi-Channel Speaker Extraction with Adversarial Training: The Wavlab Submission to The Clarity ICASSP 2023 Grand ChallengeabstractIn this work we detail our submission to the Clarity ICASSP 2023 grand challenge, in which participants have to develop a strong target speech enhancement system for hearing-aid (HA) devices in noisy-reverberant environments. Our system builds on our previous submission at the Second Clarity Enhancement Challenge (CEC2): iNeuBe-X, which consists in an iterative neural/conventional beamforming enhancement pipeline, guided by an enrollment utterance from the target speaker. This model, which won by a large margin the CEC2, is an extension of the state-of-the-art TF-GridNet model for multi-channel, streamable target-speaker speech enhancement. Here, this approach is extended and further improved by leveraging generative adversarial training, which we show proves especially useful when the training data is limited. Using only the official 6k training scenes data, our best model achieves 0.80 hearing-aid speech perception index (HASPI) and 0.41 hearing-aid speech quality index (HASQI) scores on the synthetic evaluation set. However, our model generalized poorly on the semi-real evaluation set. This highlights the fact that our community should focus more on real-world evaluation and less on fully synthetic datasets. Samuele Cornell, Zhongqiu Wang 0001, Yoshiki Masuyama, Shinji Watanabe 0001, Manuel Pariente, Nobutaka Ono, Stefano Squartini |
ICASSP | 1 |
| 2023 | Performance Above All? Energy Consumption vs. Performance, a Study on Sound Event Detection with Heterogeneous DataabstractInternational audience Romain Serizel, Samuele Cornell, Nicolas Turpault |
ICASSP | 2 |
| 2023 | TF-GRIDNET: Making Time-Frequency Domain Models Great Again for Monaural Speaker SeparationabstractWe propose TF-GridNet, a novel multi-path deep neural network (DNN) operating in the time-frequency (T-F) domain, for monaural talker-independent speaker separation in anechoic conditions. The model stacks several multi-path blocks, each consisting of an intra-frame spectral module, a sub-band temporal module, and a full-band self-attention module, to leverage local and global spectro-temporal information for separation. The model is trained to perform complex spectral mapping, where the real and imaginary (RI) components of the input mixture are stacked as input features to predict target RI components. Besides using the scale-invariant signal-to-distortion ratio (SI-SDR) loss for model training, we include a novel loss term to encourage separated sources to add up to the input mixture. Without using dynamic mixing, we obtain 23.4 dB SI-SDR improvement (SI-SDRi) on the WSJ0-2mix dataset, outperforming the previous best by a large margin. Zhongqiu Wang 0001, Samuele Cornell, Shukjae Choi, Younglo Lee, Byeong-Yeol Kim, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2023 | FNeural Speech Enhancement with Very Low Algorithmic Latency and Complexity via Integrated full- and sub-band ModelingabstractWe propose FSB-LSTM, a novel long short-term memory (LSTM) based architecture that integrates full- and sub-band (FSB) modeling, for single- and multi-channel speech enhancement in the short-time Fourier transform (STFT) domain. The model maintains an information highway to flow an over-complete input representation through multiple FSB-LSTM modules. Each FSB-LSTM module consists of a full-band block to model spectro-temporal patterns at all frequencies and a sub-band block to model patterns within each sub-band, where each of the two blocks takes a down-sampled representation as input and returns an up-sampled discriminative representation to be added to the block input via a residual connection. The model is designed to have a low algorithmic complexity, a small run-time buffer and a very low algorithmic latency, at the same time producing a strong enhancement performance on a noisy-reverberant speech enhancement task even if the hop size is as low as 2 ms. Zhongqiu Wang 0001, Samuele Cornell, Shukjae Choi, Younglo Lee, Byeong-Yeol Kim, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2023 | An experimental review of speaker diarization methods with application to two-speaker conversational telephone speech recordingsabstractWe performed an experimental review of current diarization systems for the conversational telephone speech (CTS) domain. In detail, we considered a total of eight different algorithms belonging to clustering-based, end-to-end neural diarization (EEND), and speech separation guided diarization (SSGD) paradigms. We studied the inference-time computational requirements and diarization accuracy on four CTS datasets with different characteristics and languages. We found that, among all methods considered, EEND-vector clustering (EEND-VC) offers the best trade-off in terms of computing requirements and performance. More in general, EEND models have been found to be lighter and faster in inference compared to clustering-based methods. However, they also require a large amount of diarization-oriented annotated data. In particular EEND-VC performance in our experiments degraded when the dataset size was reduced, whereas self-attentive EEND (SA-EEND) was less affected. We also found that SA-EEND gives less consistent results among all the datasets compared to EEND-VC, with its performance degrading on long conversations with high speech sparsity. Clustering-based diarization systems, and in particular VBx, instead have more consistent performance compared to SA-EEND but are outperformed by EEND-VC. The gap with respect to this latter is reduced when overlap-aware clustering methods are considered. SSGD is the most computationally demanding method, but it could be convenient if speech recognition has to be performed. Its performance is close to SA-EEND but degrades significantly when the training and inference data characteristics are less matched. Luca Serafini, Samuele Cornell, Giovanni Morrone, Enrico Zovato, Alessio Brutti, Stefano Squartini |
Comput. Speech Lang. | 2 |
| 2023 | Exploring Self-Attention Mechanisms for Speech SeparationabstractTransformers have enabled impressive improvements in deep learning. They often outperform recurrent and convolutional models in many tasks while taking advantage of parallel processing. Recently, we proposed the SepFormer, which obtains state-of-the-art performance in speech separation with the WSJ0-2/3 Mix datasets. This paper studies in-depth Transformers for speech separation. In particular, we extend our previous findings on the SepFormer by providing results on more challenging noisy and noisy-reverberant datasets, such as LibriMix, WHAM!, and WHAMR!. Moreover, we extend our model to perform speech enhancement and provide experimental evidence on denoising and dereverberation tasks. Finally, we investigate, for the first time in speech separation, the use of efficient self-attention mechanisms such as Linformers, Lonformers, and ReFormers. We found that they reduce memory requirements significantly. For example, we show that the Reformer-based attention outperforms the popular Conv-TasNet model on the WSJ0-2Mix dataset while being faster at inference and comparable in terms of memory consumption. Cem Subakan, Mirco Ravanelli, Samuele Cornell, François Grondin, Mirko Bronzi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | TF-GridNet: Integrating Full- and Sub-Band Modeling for Speech SeparationabstractWe propose TF-GridNet for speech separation. The model is a novel deep neural network (DNN) integrating full- and sub-band modeling in the time-frequency (T-F) domain. It stacks several blocks, each consisting of an intra-frame full-band module, a sub-band temporal module, and a cross-frame self-attention module. It is trained to perform complex spectral mapping, where the real and imaginary (RI) components of input signals are stacked as features to predict target RI components. We first evaluate it on monaural anechoic speaker separation. Without using data augmentation and dynamic mixing, it obtains a state-of-the-art 23.5 dB improvement in scale-invariant signal-to-distortion ratio (SI-SDR) on WSJ0-2mix, a standard dataset for two-speaker separation. To show its robustness to noise and reverberation, we evaluate it on monaural reverberant speaker separation using the SMS-WSJ dataset and on noisy-reverberant speaker separation using WHAMR!, and obtain state-of-the-art performance on both datasets. We then extend TF-GridNet to multi-microphone conditions through multi-microphone complex spectral mapping, and integrate it into a two-DNN system with a beamformer in between (named as MISO-BF-MISO in earlier studies), where the beamformer proposed in this paper is a novel multi-frame Wiener filter computed based on the outputs of the first DNN. State-of-the-art performance is obtained on the multi-channel tasks of SMS-WSJ and WHAMR!. Besides speaker separation, we apply the proposed algorithms to speech dereverberation and noisy-reverberant speech enhancement. State-of-the-art performance is obtained on a dereverberation dataset and on the dataset of the recent L3DAS22 multi-channel speech enhancement challenge. Zhongqiu Wang 0001, Samuele Cornell, Shukjae Choi, Younglo Lee, Byeong-Yeol Kim, Shinji Watanabe 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Learning Filterbanks for End-to-End Acoustic BeamformingabstractRecent work on monaural source separation has shown that performance can be increased by using fully learned filterbanks with short windows. On the other hand it is widely known that, for conventional beamforming techniques, performance increases with long analysis windows. This applies also to most hybrid neural beamforming methods which rely on a deep neural network (DNN) to estimate the spatial covariance matrices. In this work we try to bridge the gap between these two worlds and explore fully end-to-end hybrid neural beamforming in which, instead of using the Short-Time-Fourier Transform, also the analysis and synthesis filterbanks are learnt jointly with the DNN. In detail, we explore two different types of learned filterbanks: fully learned and analytic. We perform a detailed analysis using the recent Clarity Challenge data and show that by using learnt filterbanks it is possible to surpass oracle-mask based beamforming for short windows. Samuele Cornell, Manuel Pariente, François Grondin, Stefano Squartini |
ICASSP | 1 |
| 2022 | Towards Low-Distortion Multi-Channel Speech Enhancement: The ESPNET-Se Submission to the L3DAS22 ChallengeabstractThis paper describes our submission to the L3DAS22 Challenge Task 1, which consists of speech enhancement with 3D Ambisonic microphones. The core of our approach combines Deep Neural Network (DNN) driven complex spectral mapping with linear beamformers such as the multi-frame multi-channel Wiener filter. Our proposed system has two DNNs and a linear beamformer in between. Both DNNs are trained to perform complex spectral mapping, using a combination of waveform and magnitude spectrum losses. The estimated signal from the first DNN is used to drive a linear beamformer, and the beamforming result, together with this enhanced signal, are used as extra inputs for the second DNN which refines the estimation. Then, from this new estimated signal, the linear beamformer and second DNN are run iteratively. The proposed method was ranked first in the challenge, achieving, on the evaluation set, a ranking metric of 0.984, versus 0.833 of the challenge baseline. Yen-Ju Lu, Samuele Cornell, Xuankai Chang, Wangyou Zhang, Chenda Li, Zhaoheng Ni, Zhongqiu Wang 0001, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2022 | Real-M: Towards Speech Separation on Real MixturesabstractIn recent years, deep learning based source separation has achieved impressive results. Most studies, however, still evaluate separation models on synthetic datasets, while the performance of state-of-the-art techniques on in-the-wild speech data remains an open question. This paper contributes to fill this gap in two ways. First, we release the REAL-M dataset, a crowd-sourced corpus of real-life mixtures. Secondly, we address the problem of performance evaluation of real-life mixtures, where the ground truth is not available. We bypass this issue by carefully designing a blind Scale-Invariant Signal-to-Noise Ratio (SI-SNR) neural estimator. Through a user study, we show that our estimator reliably evaluates the separation performance on real mixtures, i.e. we observe that the performance predictions of the SI-SNR estimator correlate well with human opinions. Moreover, when evaluating popular speech separation models, we observe that the performance trends predicted by our estimator on the REAL-M dataset closely follow the performance trends achieved on synthetic benchmarks. Cem Subakan, Mirco Ravanelli, Samuele Cornell, François Grondin |
ICASSP | 3 |
| 2022 | ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and UnderstandingabstractThis paper presents recent progress on integrating speech separation and enhancement (SSE) into the ESPnet toolkit.Compared with the previous ESPnet-SE work, numerous features have been added, including recent state-of-the-art speech enhancement models with their respective training and evaluation recipes.Importantly, a new interface has been designed to flexibly combine speech enhancement front-ends with other tasks, including automatic speech recognition (ASR), speech translation (ST), and spoken language understanding (SLU).To showcase such integration, we performed experiments on carefully designed synthetic datasets for noisy-reverberant multichannel ST and SLU tasks, which can be used as benchmark corpora for future research.In addition to these new tasks, we also use CHiME-4 and WSJ0-2Mix to benchmark multiand single-channel SE approaches.Results show that the integration of SE front-ends with back-end tasks is a promising research direction even for tasks besides ASR, especially in the multi-channel scenario.The code is available online at https://github.com/ESPnet/ESPnet.The multichannel ST and SLU datasets, which are another contribution of this work, are released on HuggingFace. Yen-Ju Lu, Xuankai Chang, Chenda Li, Wangyou Zhang, Samuele Cornell, Zhaoheng Ni, Yoshiki Masuyama, Brian Yan, Robin Scheibler, Zhongqiu Wang 0001, Yu Tsao 0001, Yanmin Qian, Shinji Watanabe 0001 |
INTERSPEECH | 5 |
| 2022 | Implicit Acoustic Echo Cancellation for Keyword Spotting and Device-Directed Speech DetectionabstractIn many speech-enabled human-machine interactions, user speech can overlap with the device playback audio. In these instances, the performance of tasks such as keyword-spotting (KWS) and device-directed speech detection (DDD) can degrade significantly. To address this problem, we propose an implicit acoustic echo cancellation (iAEC) framework where a neural network is trained to exploit the additional information from a reference microphone channel to learn to ignore the interfering signal and improve detection performance. We study this framework for the tasks of KWS and DDD on, respectively, an augmented version of Google Speech Commands v2 and a real-world Alexa device dataset. Notably, we show a 56% reduction in false-reject rate for the DDD task during device playback conditions. We also show comparable or superior performance over a strong end-to-end neural echo cancellation baseline for the KWS task with two order of magnitude less computational requirements. Samuele Cornell, Thomas Balestri, Thibaud Senechal |
SLT | 1 |
| 2022 | End-to-End Integration of Speech Recognition, Dereverberation, Beamforming, and Self-Supervised Learning RepresentationabstractSelf-supervised learning representation (SSLR) has demonstrated its significant effectiveness in automatic speech recognition (ASR), mainly with clean speech. Recent work pointed out the strength of integrating SSLR with single-channel speech enhancement for ASR in noisy environments. This paper further advances this integration by dealing with multi-channel input. We propose a novel end-to-end architecture by integrating dereverberation, beamforming, SSLR, and ASR within a single neural network. Our system achieves the best performance reported in the literature on the CHiME-4 6-channel track with a word error rate (WER) of 1.77%. While the WavLM-based strong SSLR demonstrates promising results by itself, the end-to-end integration with the weighted power minimization distortionless response beamformer, which simultaneously performs dereverberation and denoising, improves WER significantly. Its effectiveness is also validated on the REVERB dataset. Yoshiki Masuyama, Xuankai Chang, Samuele Cornell, Shinji Watanabe 0001, Nobutaka Ono |
SLT | 3 |
| 2022 | Low-Latency Speech Separation Guided Diarization for Telephone ConversationsabstractIn this paper, we carry out an analysis on the use of speech separation guided diarization (SSGD) in telephone conversations. SSGD performs diarization by separating the speakers signals and then applying voice activity detection on each estimated speaker signal. In particular, we compare two low-latency speech separation models. Moreover, we show a post-processing algorithm that significantly reduces the false alarm errors of a SSGD pipeline. We perform our experiments on two datasets: Fisher Corpus Part 1 and CALLHOME, evaluating both separation and diarization metrics. Notably, our SSGD DPRNN-based online model achieves 11.1% DER on CALL-HOME, comparable with most state-of-the-art end-to-end neural diarization models despite being trained on an order of magnitude less data and having considerably lower latency, i.e., 0.1 vs. 10 seconds. We also show that the separated signals can be readily fed to a speech recognition back-end with performance close to the oracle source signals. Giovanni Morrone, Samuele Cornell, Desh Raj, Luca Serafini, Enrico Zovato, Alessio Brutti, Stefano Squartini |
SLT | 2 |
| 2022 | Overlapped Speech Detection and speaker counting using distant microphone arrays
Samuele Cornell, Maurizio Omologo, Stefano Squartini, Emmanuel Vincent 0001 |
Comput. Speech Lang. | 1 |
| 2021 | Attention Is All You Need In Speech SeparationabstractRecurrent Neural Networks (RNNs) have long been the dominant architecture in sequence-to-sequence learning. RNNs, however, are inherently sequential models that do not allow parallelization of their computations. Transformers are emerging as a natural alternative to standard RNNs, replacing recurrent computations with a multi-head attention mechanism.In this paper, we propose the SepFormer, a novel RNN-free Transformer-based neural network for speech separation. The Sep-Former learns short and long-term dependencies with a multi-scale approach that employs transformers. The proposed model achieves state-of-the-art (SOTA) performance on the standard WSJ0-2/3mix datasets. It reaches an SI-SNRi of 22.3 dB on WSJ0-2mix and an SI-SNRi of 19.5 dB on WSJ0-3mix. The SepFormer inherits the parallelization advantages of Transformers and achieves a competitive performance even when downsampling the encoded representation by a factor of 8. It is thus significantly faster and it is less memory-demanding than the latest speech separation systems with comparable performance. Cem Subakan, Mirco Ravanelli, Samuele Cornell, Mirko Bronzi, Jianyuan Zhong |
ICASSP | 3 |
| 2021 | Learning to Rank Microphones for Distant Speech RecognitionabstractFully exploiting ad-hoc microphone networks for distant speech recognition is still an open issue. Empirical evidence shows that being able to select the best microphone leads to significant improvements in recognition without any additional effort on front-end processing. Current channel selection techniques either rely on signal, decoder or posterior-based features. Signal-based features are inexpensive to compute but do not always correlate with recognition performance. Instead decoder and posterior-based features exhibit better correlation but require substantial computational resources. In this work, we tackle the channel selection problem by proposing MicRank, a learning to rank framework where a neural network is trained to rank the available channels using directly the recognition performance on the training set. The proposed approach is agnostic with respect to the array geometry and type of recognition back-end. We investigate different learning to rank strategies using a synthetic dataset developed on purpose and the CHiME-6 data. Results show that the proposed approach is able to considerably improve over previous selection techniques, reaching comparable and in some instances better performance than oracle signal-based measures. Samuele Cornell, Alessio Brutti, Marco Matassoni, Stefano Squartini |
Interspeech | 1 |
| 2020 | Filterbank Design for End-to-end Speech SeparationabstractSingle-channel speech separation has recently made great progress thanks to learned filterbanks as used in ConvTasNet. In parallel, parameterized filterbanks have been proposed for speaker recognition where only center frequencies and bandwidths are learned. In this work, we extend real-valued learned and parameterized filterbanks into complex-valued analytic filterbanks and define a set of corresponding representations and masking strategies. We evaluate these filterbanks on a newly released noisy speech separation dataset (WHAM). The results show that the proposed analytic learned filterbank consistently outperforms the real-valued filterbank of ConvTasNet. Also, we validate the use of parameterized filterbanks and show that complex-valued representations and masks are beneficial in all conditions. Finally, we show that the STFT achieves its best performance for 2 ms windows. Manuel Pariente, Samuele Cornell, Antoine Deleforge, Emmanuel Vincent 0001 |
ICASSP | 2 |
| 2020 | A Novel Adversarial Training Scheme for Deep Neural Network based Speech EnhancementabstractIn this work, we propose a novel representation-learning technique for Deep Learning-based Speech Enhancement algorithms inspired by Domain-Adversarial training. A gradient reversal layer and an additional network are employed, only at training time, to explicitly enforce a representation that is orthogonal to the additive noise in the input signal. We show that such learning scheme, which can be applied easily to most mask-based Deep Neural Network Speech Enhancement approaches, is able to improve the denoising performance when used in conjunction with scale-invariant signal-to-distortion ratio loss and allows to reach state-of-the-art performance with no computational overhead at run-time. In particular, on the commonly used VoiceBank-DEMAND benchmarking dataset, we improve signal-to-distortion ratio and signal-to-noise ratio over the nonadversarial model and CSIG, COVL and CBAK over other, state-of-the art, adversarial training techniques. Samuele Cornell, Emanuele Principi, Stefano Squartini |
IJCNN | 1 |
| 2020 | Who Cried When: Infant Cry Diarization with Dilated Fully-Convolutional Neural NetworksabstractIn this paper, we address the problem of the concurrent detection of multiple infant cries by using microphones located in the cribs of a Neonatal Intensive Care Unit (NICU). We term this task as infant cry diarization in resemblance with the "speaker diarization" task related to the speech signal: instead of determining "who spoke when", here the problem is determining "who cried when". The proposed algorithm consists of a fully-convolutional neural network (Conv-DetNet) that processes simultaneously all the audio signals acquired from the microphone in each crib and detects if the infants cried or not. The neural network takes as input Log-Mel coefficients and it is composed of stacked dilated convolutional blocks with increasing dilation factors. Each block is composed of pointwise and depthwise convolutional layers that replace standard convolutions with a mathematically equivalent but more efficient operation. The architecture has been compared to its single-channel equivalent and to single and multi-channel architectures presented in a previous work, composed of standard convolutional layers and fully-connected layers. The experiments have been conducted on a synthetic dataset that simulates the acoustic environment of the Salesi Hospital NICU located in Ancona (Italy). The results have been evaluated in terms of Area Under Precision-Recall Curve (PRC-AUC) and they showed that the proposed multi-channel Conv-DetNet achieves the highest performance with a PRC-AUC equal to 87.58%, outperforming all the comparative methods. Marco Severini, Emanuele Principi, Samuele Cornell, Leonardo Gabrielli, Stefano Squartini |
IJCNN | 3 |
| 2020 | Detecting and Counting Overlapping Speakers in Distant Speech ScenariosabstractInternational audience Samuele Cornell, Maurizio Omologo, Stefano Squartini, Emmanuel Vincent 0001 |
INTERSPEECH | 1 |
| 2020 | Asteroid: The PyTorch-Based Audio Source Separation Toolkit for ResearchersabstractThis paper describes Asteroid, the PyTorch-based audio source separation toolkit for researchers. Inspired by the most successful neural source separation systems, it provides all neural building blocks required to build such a system. To improve reproducibility, Kaldi-style recipes on common audio source separation datasets are also provided. This paper describes the software architecture of Asteroid and its most important features. By showing experimental results obtained with Asteroid's recipes, we show that our implementations are at least on par with most results reported in reference papers. The toolkit is publicly available at https://github.com/mpariente/asteroid . Manuel Pariente, Samuele Cornell, Joris Cosentino, Sunit Sivasankaran, Efthymios Tzinis, Jens Heitkaemper, Michel Olvera, Fabian-Robert Stöter, Mathieu Hu, Juan M. Martín-Doñas, David Ditter, Ariel Frank, Antoine Deleforge, Emmanuel Vincent 0001 |
INTERSPEECH | 2 |