VLDB 2026 Research / reviewers in the wild / expert
Catalin Zorila
dblp:243/6620
· DBLP profile ↗
18ranked-venue papers
5as first author
15since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Geodesic Interpolation of Frame-Wise Speaker Embeddings for the Diarization of Meeting ScenariosabstractWe propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partially overlapping speech. To this end, a geodesic distance loss is used that enforces the embeddings computed from regions with two active speakers to lie on the shortest path on a sphere between the points given by the d-vectors of each of the active speakers. Using those frame-wise speaker embeddings in clustering-based diarization outperforms segment-level clustering-based diarization systems such as VBx and Spectral Clustering. By extending our approach to a mixture-model-based diarization, the performance can be further improved, approaching the diarization error rates of diarization systems that use a dedicated overlap detection, and outperforming these systems when also employing an additional overlap detection. Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2023 | Towards a Unified End-to-End Language Understanding System for Speech and Text InputsabstractEnd-to-end (E2E) spoken language understanding (SLU) systems facilitate mapping speech inputs directly to semantic outputs, eliminating the need for modular processing of speech-to-text and text-to-semantics sub-tasks using separate models. However, they are now limited to processing speech inputs only, and are not flexible to deal with plain texts. In this paper, we propose an E2E spoken and natural language understanding (SNLU) system that can handle both speech and text within a unified architecture. The system follows the Mask-CTC non-autoregressive approach, and the input flexibility is acquired by partially sharing the decoder between SLU and NLU tasks. Experiments on the SLURP dataset show that the proposed architecture achieves similar performance to using separate E2E SLU and NLU modules, but with relatively 43.7 % less model parameters. We also explore the use of pre-trained speech and language models into the SNLU system, and show that they further improve the performance. Mohan Li, Catalin Zorila, Cong-Thanh Do, Rama Sanand Doddipatla |
ASRU | 2 |
| 2023 | Frame-Wise and Overlap-Robust Speaker Embeddings for Meeting DiarizationabstractUsing a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even for segments with speech overlap, the frame-wise embeddings serve as an appropriate representation of the input speech signal for an end-to-end neural meeting diarization (EEND) system. We show in experiments that this representation helps mitigate a well-known problem of EEND systems: when increasing the number of speakers the diarization performance drop is significantly reduced. We also introduce block-wise processing to be able to diarize arbitrarily long meetings. Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2023 | On the Effectiveness of Monoaural Target Source Extraction for Distant end-to-end Automatic Speech RecognitionabstractRecent work on enhancement has shown that frequency domain methods may outperform the time domain approaches, while most of the prior art is focused on reporting objective enhancement metrics on simulated noisy data or use less modern hybrid acoustic models for evaluation. In this paper we investigate the effectiveness of target source extraction for improving the robustness of end-to-end automatic speech recognition in noisy and reverberant conditions. A frequency domain source extraction approach is introduced and compared against a state-of-the-art time domain method using several publicly available simulated and real noisy speech test sets. The results show that the frequency domain method outperforms the time domain one only for simulated conditions, and that it is more stable to window size variations. The experiments also indicate that remixing the unprocessed signal with the enhanced speech (referred to as speaker/source reinforcement) yields similar or better results than by using a matched acoustic model retrained using distortions introduced by enhancement. Catalin Zorila, Rama Sanand Doddipatla |
ICASSP | 1 |
| 2023 | A Teacher-Student Approach for Extracting Informative Speaker Embeddings From Speech MixturesabstractWe introduce a monaural neural speaker embeddings extractor that computes an embedding for each speaker present in a speech mixture.To allow for supervised training, a teacherstudent approach is employed: the teacher computes the target embeddings from each speaker's utterance before the utterances are added to form the mixture, and the student embedding extractor is then tasked to reproduce those embeddings from the speech mixture at its input.The system much more reliably verifies the presence or absence of a given speaker in a mixture than a conventional speaker embedding extractor, and even exhibits comparable performance to a multi-channel approach that exploits spatial information for embedding extraction.Further, it is shown that a speaker embedding computed from a mixture can be used to check for the presence of that speaker in another mixture. Tobias Cord-Landwehr, Christoph Böddeker, Catalin Zorila, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
INTERSPEECH | 3 |
| 2022 | Transformer-Based Streaming ASR with Cumulative AttentionabstractIn this paper, we propose an online attention mechanism, known as cumulative attention (CA), for streaming Transformer-based automatic speech recognition (ASR). Inspired by monotonic chunk-wise attention (MoChA) and head-synchronous decoder-end adaptive computation steps (HS-DACS) algorithms, CA triggers the ASR outputs based on the acoustic information accumulated at each encoding timestep, where the decisions are made using a trainable device, referred to as halting selector. In CA, all the attention heads of the same decoder layer are synchronised to have a unified halting position. This feature effectively alleviates the problem caused by the distinct behaviour of individual heads, which may otherwise give rise to severe latency issues as encountered by MoChA. The ASR experiments conducted on AIShell-1 and Librispeech datasets demonstrate that the proposed CA-based Transformer system can achieve on par or better performance with significant reduction in latency during inference, when compared to other streaming Transformer systems in literature. Mohan Li, Shucong Zhang, Catalin Zorila, Rama Sanand Doddipatla |
ICASSP | 3 |
| 2022 | Speaker Reinforcement Using Target Source Extraction for Robust Automatic Speech RecognitionabstractImproving the accuracy of single-channel automatic speech recognition (ASR) in noisy conditions is challenging. Strong speech enhancement front-ends are available, however, they typically require that the ASR model is retrained to cope with the processing artifacts. In this paper we explore a speaker reinforcement strategy for improving recognition performance without retraining the acoustic model (AM). This is achieved by remixing the enhanced signal with the unprocessed input to alleviate the processing artifacts. We evaluate the proposed approach using a DNN speaker extraction based speech denoiser trained with a perceptually motivated loss function. Results show that (without AM retraining) our method yields about 23% and 25% relative accuracy gains compared with the unprocessed for the monoaural simulated and real CHiME-4 evaluation sets, respectively, and outperforms a state-of-the-art reference method. Catalin Zorila, Rama Sanand Doddipatla |
ICASSP | 1 |
| 2022 | Self-regularised Minimum Latency Training for Streaming Transformer-based Speech Recognition
Mohan Li, Rama Sanand Doddipatla, Catalin Zorila |
INTERSPEECH | 3 |
| 2022 | On monoaural speech enhancement for automatic recognition of real noisy speech using mixture invariant trainingabstractIn this paper, we explore an improved framework to train a monoaural neural enhancement model for robust speech recognition. The designed training framework extends the existing mixture invariant training criterion to exploit both unpaired clean speech and real noisy data. It is found that the unpaired clean speech is crucial to improve quality of separated speech from real noisy speech. The proposed method also performs remixing of processed and unprocessed signals to alleviate the processing artifacts. Experiments on the single-channel CHiME-3 real test sets show that the proposed method improves significantly in terms of speech recognition performance over the enhancement system trained either on the mismatched simulated data in a supervised fashion or on the matched real data in an unsupervised fashion. Between 16% and 39% relative WER reduction has been achieved by the proposed system compared to the unprocessed signal using end-to-end and hybrid acoustic models without retraining on distorted data. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
INTERSPEECH | 2 |
| 2022 | End-to-End Neural Based Modification of Noisy Speech for Speech-in-Noise Intelligibility ImprovementabstractIntelligibility of speech can be significantly reduced when it is presented in adverse near-end listening conditions, like background noise. Multiple approaches have been suggested to improve the perception of speech in such conditions. However, most of these approaches were designed to work with clean input speech. Therefore, they have serious limitations when deployed in real world applications like telephony and hearing aids, where noisy input speech is quite common. In this paper we present an end-to-end neural network approach for the above problem, which effectively reduces the input noise and improves the intelligibility for listeners in adverse conditions. To that end, a convolutional neural network topology with variable dilation factors is proposed and evaluated both in a causal and a non-causal configuration using raw speech as input. A Teacher-Student training strategy is employed, where the Teacher is a well-established speech-in-noise intelligibility enhancer based on spectral shaping followed by dynamic range compression (SSDRC). The evaluation is performed both objectively using the speech intelligibility in bits metric (SIIB), and subjectively on the Greek Harvard corpus. A noise robust multi-band version of SSDRC was used as a baseline. Compared with the baseline, at 0 dB input SNR, the suggested neural network system achieved about 380% and 230% relative SIIB improvements in fluctuating and stationary backgrounds, respectively. Subjectively, the suggested model increased listeners’ keyword correct rate in stationary noise from 25% to 60% at 0 dB input SNR, and from about 52% to 75% at 5 dB input SNR, compared with the baseline. P. V. Muhammed Shifas, Catalin Zorila, Yannis Stylianou |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Head-Synchronous Decoding for Transformer-Based Streaming ASRabstractOnline Transformer-based automatic speech recognition (ASR) systems have been extensively studied due to the increasing demand for streaming applications. Recently proposed Decoder-end Adaptive Computation Steps (DACS) algorithm for online Transformer ASR was shown to achieve state-of-the-art performance and outperform other existing methods. However, like any other online approach, the DACS-based attention heads in each of the Transformer decoder layers operate independently (or asynchronously) and lead to diverged attending positions. Since DACS employs a truncation threshold to determine the halting position, some of the attention weights are cut off untimely and might impact the stability and precision of decoding. To overcome these issues, here we propose a head-synchronous (HS) version of the DACS algorithm, where the boundary of attention is jointly detected by all the DACS heads in each decoder layer. ASR experiments on Wall Street Journal (WSJ), AIShell-1 and Lib- rispeech show that the proposed method consistently outperforms vanilla DACS and achieves state-of-the-art performance. We will also demonstrate that HS-DACS has reduced decoding cost when compared to vanilla DACS. Mohan Li, Catalin Zorila, Rama Sanand Doddipatla |
ICASSP | 2 |
| 2021 | Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning MechanismabstractIn this paper, we present a novel multi-channel speech extraction system to simultaneously extract multiple clean individual sources from a mixture in noisy and reverberant environments. The proposed method is built on an improved multi-channel time-domain speech separation network which employs speaker embeddings to identify and extract multiple targets without label permutation ambiguity. To efficiently inform the speaker information to the extraction model, we propose a new speaker conditioning mechanism by designing an additional speaker branch for receiving external speaker embeddings. Experiments on 2-channel WHAMR! data show that the proposed system improves by 9% relative the source separation performance over a strong multi-channel baseline, and it increases the speech recognition accuracy by more than 16% relative over the same baseline. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
ICASSP | 2 |
| 2021 | Teacher-Student MixIT for Unsupervised and Semi-Supervised Speech SeparationabstractIn this paper, we introduce a novel semi-supervised learning framework for end-to-end speech separation. The proposed method first uses mixtures of unseparated sources and the mixture invariant training (MixIT) criterion to train a teacher model. The teacher model then estimates separated sources that are used to train a student model with standard permutation invariant training (PIT). The student model can be fine-tuned with supervised data, i.e., paired artificial mixtures and clean speech sources, and further improved via model distillation. Experiments with single and multi channel mixtures show that the teacher-student training resolves the over-separation problem observed in the original MixIT method. Further, the semisupervised performance is comparable to a fully-supervised separation system trained using ten times the amount of supervised data. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
Interspeech | 2 |
| 2021 | Transformer-Based Online Speech Recognition with Decoder-end Adaptive Computation StepsabstractTransformer-based end-to-end (E2E) automatic speech recognition (ASR) systems have recently gained wide popularity, and are shown to outperform E2E models based on recurrent structures on a number of ASR tasks. However, like other E2E models, Transformer ASR also requires the full input sequence for calculating the attentions on both encoder and decoder, leading to increased latency and posing a challenge for online ASR. The paper proposes Decoder-end Adaptive Computation Steps (DACS) algorithm to address the issue of latency and facilitate online ASR. The proposed algorithm streams the decoding of Transformer ASR by triggering an output after the confidence acquired from the encoder states reaches a certain threshold. Unlike other monotonic attention mechanisms that risk visiting the entire encoder states for each output step, the paper introduces a maximum look-ahead step into the DACS algorithm to prevent from reaching the end of speech too fast. A Chunkwise en-coder is adopted in our system to handle real-time speech inputs. The proposed online Transformer ASR system has been evaluated on Wall Street Journal (WSJ) and AIShell-1 datasets, yielding 5.5% word error rate (WER) and 7.1% character error rate (CER) respectively, with only a minor decay in performance when compared to the offline systems. Mohan Li, Catalin Zorila, Rama Sanand Doddipatla |
SLT | 2 |
| 2021 | An Investigation into the Multi-channel Time Domain Speaker Extraction NetworkabstractThis paper presents an investigation into the effectiveness of spatial features for improving time-domain speaker extraction systems. A two-dimensional Convolutional Neural Network (CNN) based encoder is proposed to capture the spatial information within the multichannel input, which are then combined with the spectral features of a single channel extraction network. Two variants of target speaker extraction methods were tested, one which employs a pre-trained i-vector system to compute a speaker embedding (System A), and one which employs a jointly trained neural network to extract the embeddings directly from time domain enrolment signals (System B). The evaluation was performed on the spatialized WSJ0-2mix dataset using the Signal-to-Distortion Ratio (SDR) metric, and ASR accuracy. In the anechoic condition, more than 10 dB and 7 dB absolute SDR gains were achieved when the 2-D CNN spatial encoder features were included with Systems A and B, respectively. The performance gains in reverberation were lower, however, we have demonstrated that retraining the systems by applying dereverberation preprocessing can significantly boost both the target speaker extraction and ASR performances. Catalin Zorila, Mohan Li, Rama Sanand Doddipatla |
SLT | 1 |
| 2020 | On End-to-end Multi-channel Time Domain Speech Separation in Reverberant EnvironmentsabstractThis paper introduces a new method for multi-channel time domain speech separation in reverberant environments. A fully-convolutional neural network structure has been used to directly separate speech from multiple microphone recordings, with no need of conventional spatial feature extraction. To reduce the influence of reverberation on spatial feature extraction, a dereverberation pre-processing method has been applied to further improve the separation performance. A spatialized version of wsj0-2mix dataset has been simulated to evaluate the proposed system. Both source separation and speech recognition performance of the separated signals have been evaluated objectively. Experiments show that the proposed fully-convolutional network improves the source separation metric and the word error rate (WER) by more than 13% and 50% relative, respectively, over a reference system with conventional features. Applying dereverberation as pre-processing to the proposed system can further reduce the WER by 29% relative using an acoustic model trained on clean and reverberated data. Jisi Zhang, Catalin Zorila, Rama Sanand Doddipatla, Jon Barker |
ICASSP | 2 |
| 2019 | An Investigation into the Effectiveness of Enhancement in ASR Training and Test for Chime-5 Dinner Party TranscriptionabstractDespite the strong modeling power of neural network acoustic models, speech enhancement has been shown to deliver additional word error rate improvements if multi-channel data is available. However, there has been a longstanding debate whether enhancement should also be carried out on the ASR training data. In an extensive experimental evaluation on the acoustically very challenging CHiME-5 dinner party data we show that: (i) cleaning up the training data can lead to substantial error rate reductions, and (ii) enhancement in training is advisable as long as enhancement in test is at least as strong as in training. This approach stands in contrast and delivers larger gains than the common strategy reported in the literature to augment the training database with additional artificially degraded speech. Together with an acoustic model topology consisting of initial CNN layers followed by factorized TDNN layers we achieve with 41.6 % and 43.2 % WER on the DEV and EVAL test sets, respectively, a new single-system state-of-the-art result on the CHiME-5 data. This is a 8 % relative improvement compared to the best word error rate published so far for a speech recognizer without system combination. Catalin Zorila, Christoph Böddeker, Rama Sanand Doddipatla, Reinhold Häb-Umbach |
ASRU | 1 |
| 2019 | On Reducing the Effect of Speaker Overlap for Chime-5abstractThe CHiME-5 speech separation and recognition challenge was recently shown to pose a difficult task for the current automatic speech recognition systems. Speaker overlap was one of the main difficulties of the challenge. The presence of noise, reverberation and the moving speakers have made the traditional source separation methods ineffective in improving the recognition accuracy. In this paper we have explored several enhancement strategies aimed to reduce the effect of speaker overlap for CHiME-5 without performing source separation. One is based on discarding the overlap segments using the speaker diarisation information from the challenge, another one is a neural network driven automatic gain control enhancement aimed to improve the previous speaker diarisation information, and the last one is based on optimal multi-array data selection. State-of-the-art acoustic models were used to perform the ASR experiments. Results have shown that proposed automatic gain control method yields word error rate (WER) reductions between 2% and 3% absolute on the development set of CHiME-5. Catalin Zorila, Rama Sanand Doddipatla |
ICASSP | 1 |