Shlomo E. Chazan

dblp:170/0036 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-6614-6112ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Measuring the Effect of Transcription Noise on Downstream Language Understanding Tasks
abstract
With the increasing prevalence of recorded human speech, spoken language understanding (SLU) is essential for its efficient processing.In order to process the speech, it is commonly transcribed using automatic speech recognition technology.This speech-to-text transition introduces errors into the transcripts, which subsequently propagate to downstream NLP tasks, such as dialogue summarization.While it is known that transcript noise affects downstream tasks, a general-purpose and systematic approach to analyzing its effects across different noise severities and types has not been addressed.We propose a configurable framework for assessing task models in diverse noisy settings, and for examining the impact of transcript-cleaning techniques.The framework facilitates the investigation of task model behavior, which can in turn support the development of effective SLU solutions.We exemplify the utility of our framework on three SLU tasks and four task models, offering insights regarding the effect of transcript noise on tasks in general and models in particular.For instance, we find that task models can tolerate a certain level of noise, and are affected differently by the types of errors in the transcript. 1
Ori Shapira, Shlomo E. Chazan, Amir David Nissan Cohen
ACL (1)2
2025 Automatic Detection of Domain Shifts in Speech Enhancement Systems Using Confidence-Based Metrics
abstract
Introducing a domain shift, such as a change in language or environment, to a well-trained speech enhancement system can cause severe performance degradation. Most current research assumes that a domain shift has already been detected and focuses on either supervised or unsupervised domain adaptation techniques. Here, we address the problem of automatically detecting when a domain shift has occurred. We present a domain shift detection method based on monitoring the confidence of a network that predicts the quality of enhanced speech. The experimental results show that our method can effectively detect a domain mismatch between the training and test sets.
Lior Frankel, Shlomo E. Chazan, Jacob Goldberger
ICASSP2
2025 Pull It Together: Reducing the Modality Gap in Contrastive Learning
Amit Sofer, Yoav Goldman, Shlomo E. Chazan
INTERSPEECH3
2024 C-CLAPA: Improving Text-Audio Cross Domain Retrieval with Captioning and Augmentations
abstract
In this paper, we introduce Captioning decoder Contrastive Language-Audio Pretraining with data Augmantation (C-CLAPA), a new Audio-Text model for the Cross Domain Retrieval (CDR) task. The model’s backbone is comprised of two encoders, one for the text and the other for the audio. The embedding vectors from the different modalities are commonly trained with a contrastive-loss. In our approach, a captioning decoder is also used to generate a text-description from the embedding vector of the audio sample. This decoder is used to ensure that the audio embedding encapsulates text information, and is used only on training stage. Data preparations including filtering, augmentations and text generation utilizing Large Language Models (LLMs), are used to extend the current training dataset. The proposed model is finally trained using a curriculum training procedure. In this approach, we train the model on datasets with increasing quality. In our empirical investigation, we provide compelling evidence that our model significantly surpasses the current State Of The Art (SOTA) models on the available benchmarks. Ablation analysis provides empirical evidence showcasing the advantages in the proposed architectural design as well as the efficacy of the employed data processing methodology.
Amit Sofer, Shlomo E. Chazan
ICASSP2
2024 Domain Adaptation Using Suitable Pseudo Labels for Speech Enhancement and Dereverberation
abstract
Speech enhancement and dereverberation approaches based on neural networks are designed to learn a transformation from noisy to clean speech using supervised learning. However, networks trained in this way may fail to effectively handle languages, types of noise, or acoustic environments that were not included in the training data. To tackle this issue, the present study centers around unsupervised domain adaptation, specifically addressing scenarios characterized by substantial domain gaps. In this scenario, we have noisy speech data from the new domain, but the corresponding clean speech data is unavailable. We propose an adaptation method based on domainadversarial training followed by iterative self-training, where the estimated speech is used as pseudo labels, and the target samples are gradually introduced to the network based on their similarity to the source domain. The self-training also utilizes labeled samples from the source domain which are similar to the target domain. The experimental results show that our method effectively mitigates the domain mismatch between the training and test sets, thus outperforming the current baselines.
Lior Frenkel, Shlomo E. Chazan, Jacob Goldberger
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Don't Be So Sure! Boosting ASR Decoding via Confidence Relaxation
abstract
Automatic Speech Recognition (ASR) systems frequently use a search-based decoding strategy aiming to find the best attainable transcript by considering multiple candidates. One prominent speech recognition decoding heuristic is beam search, which seeks the transcript with the greatest likelihood computed using the predicted distribution. While showing substantial performance gains in various tasks, beam search loses some of its effectiveness when the predicted probabilities are highly confident, i.e., the predicted distribution is massed for a single or very few classes. We show that recently proposed Self-Supervised Learning (SSL)-based ASR models tend to yield exceptionally confident predictions that may hamper beam search from truly considering a diverse set of candidates. We perform a layer analysis to reveal and visualize how predictions evolve, and propose a decoding procedure that improves the performance of fine-tuned ASR models. Our proposed approach does not require further training beyond the original fine-tuning, nor additional model parameters. In fact, we find that our proposed method requires significantly less inference computation than current approaches. We propose aggregating the top M layers, potentially leveraging useful information encoded in intermediate layers, and relaxing model confidence. We demonstrate the effectiveness of our approach by conducting an empirical study on varying amounts of labeled resources and different model sizes, showing consistent improvements in particular when applied to low-resource scenarios.
Tomer Wullach, Shlomo E. Chazan
AAAI2
2023 Optimized Tokenization for Transcribed Error Correction
abstract
The challenges facing speech recognition systems, such as variations in pronunciations, adverse audio conditions, and the scarcity of labeled data, emphasize the necessity for a postprocessing step that corrects recurring errors.Previous research has shown the advantages of employing dedicated error correction models, yet training such models requires large amounts of labeled data which is not easily obtained.To overcome this limitation, synthetic transcribedlike data is often utilized, however, bridging the distribution gap between transcribed errors and synthetic noise is not trivial.In this paper, we demonstrate that the performance of correction models can be significantly increased by training solely using synthetic data.Specifically, we empirically show that: (1) synthetic data generated using the error distribution derived from a set of transcribed data outperforms the common approach of applying random perturbations; (2) applying language-specific adjustments to the vocabulary of a BPE tokenizer strike a balance between adapting to unseen distributions and retaining knowledge of transcribed errors.We showcase the benefits of these key observations, and evaluate our approach using multiple languages, speech recognition systems and prominent speech recognition datasets.
Tomer Wullach, Shlomo E. Chazan
EMNLP2
2023 Domain Adaptation for Speech Enhancement in a Large Domain Gap
Lior Frenkel, Jacob Goldberger, Shlomo E. Chazan
INTERSPEECH3
2021 Speech Enhancement with Mixture of Deep Experts with Clean Clustering Pre-Training
abstract
In this study we present a mixture of deep experts (MoDE) neural-network architecture for single microphone speech enhancement. Our architecture comprises a set of deep neural networks (DNNs), each of which is an ‘expert’ in a different speech spectral pattern such as phoneme. A gating DNN is responsible for the latent variables which are the weights assigned to each expert’s output given a speech segment. The experts estimate a mask from the noisy input and the final mask is then obtained as a weighted average of the experts’ estimates, with the weights determined by the gating DNN. A soft spectral attenuation, based on the estimated mask, is then applied to enhance the noisy speech signal. As a byproduct, we gain reduction at the complexity in test time. We show that the experts specialization allows better robustness to unfamiliar noise types.1
Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot
ICASSP1
2021 Single Channel Voice Separation for Unknown Number of Speakers Under Reverberant and Noisy Settings
abstract
We present a unified network for voice separation of an unknown number of speakers. The proposed approach is composed of several separation heads optimized together with a speaker classification branch. The separation is carried out in the time domain, together with parameter sharing between all separation heads. The classification branch estimates the number of speakers while each head is specialized in separating a different number of speakers. We evaluate the proposed model under both clean and noisy reverberant settings. Results suggest that the proposed approach is superior to the baseline model by a significant margin. Additionally, we present a new noisy and reverberant dataset of up to five different speakers speaking simultaneously.
Shlomo E. Chazan, Lior Wolf, Eliya Nachmani, Yossi Adi
ICASSP1
2020 K-Autoencoders Deep Clustering
abstract
In this study we propose a deep clustering algorithm that extends the k-means algorithm. Each cluster is represented by an autoencoder instead of a single centroid vector. Each data point is associated with the autoencoder which yields the minimal reconstruction error. The optimal clustering is found by learning a set of autoencoders that minimize the global reconstruction mean-square error loss. The network architecture is a simplified version of a previous method that is based on mixture-of-experts. The proposed method is evaluated on standard image corpora and performs on par with state-of-the-art methods which are based on much more complicated network architectures.
Yaniv Opochinsky, Shlomo E. Chazan, Sharon Gannot, Jacob Goldberger
ICASSP2
2020 A Composite DNN Architecture for Speech Enhancement
abstract
In speech enhancement, the use of supervised algorithms in the form of deep neural networks (DNNs) has become tremendously popular in recent years. The target function of the DNN (and the associated estimators) is often either a masking function applied to the noisy spectrum, or the clean log-spectrum. In this work, we show that both separate cost functions are unsuitable for dealing with narrowband noise, and propose a new composite estimator in the log-spectrum domain. The new technique relies on a single DNN that outputs both a masking function and an estimated log-spectrum. Both outputs are used for the composite enhancement. The proposed estimator demonstrates superior performance for speech utterances contaminated by additive narrowband noise, while maintaining the enhancement quality of the baseline algorithms for wideband noise.
Yochai Yemini, Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot
ICASSP2
2018 DNN-Based Concurrent Speakers Detector and its Application to Speaker Extraction with LCMV Beamforming
abstract
In this paper, we present a new control mechanism for LCMV beamforming. Application of the LCMV beamformer to speaker separation tasks requires accurate estimates of its building blocks, e.g. the noise spatial cross-power spectral density (cPSD) matrix and the relative transfer function (RTF) of all sources of interest. An accurate classification of the input frames to various speaker activity patterns can facilitate such an estimation procedure. We propose a DNN-based concurrent speakers detector (CSD) to classify the noisy frames. The CSD, trained in a supervised manner using a DNN, classifies noisy frames into three classes: 1) all speakers are inactive - used for estimating the noise spatial cPSD matrix; 2) a single speaker is active - used for estimating the RTF of the active speaker; and 3) more than one speaker is active - discarded for estimation purposes. Finally, using the estimated blocks, the LCMV beamformer is constructed and applied for extracting the desired speaker from a noisy mixture of speakers.
Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot
ICASSP1
2016 A Hybrid Approach for Speech Enhancement Using MoG Model and Neural Network Phoneme Classifier
abstract
In this paper, we present a single-microphone speech enhancement algorithm. A hybrid approach is proposed merging the generative mixture of Gaussians (MoG) model and the discriminative deep neural network (DNN). The proposed algorithm is executed in two phases, the training phase, which does not recur, and the test phase. First, the noise-free speech log-power spectral density is modeled as an MoG, representing the phoneme-based diversity in the speech signal. A DNN is then trained with phoneme labeled database of clean speech signals for phoneme classification with mel-frequency cepstral coefficients as the input features. In the test phase, a noisy utterance of an untrained speech is processed. Given the phoneme classification results of the noisy speech utterance, a speech presence probability (SPP) is obtained using both the generative and discriminative models. SPP-controlled attenuation is then applied to the noisy speech while simultaneously, the noise estimate is updated. The discriminative DNN maintains the continuity of the speech and the generative phoneme-based MoG preserves the speech spectral structure. Extensive experimental study using real speech and noise signals is provided. We also compare the proposed algorithm with alternative speech enhancement algorithms. We show that we obtain a significant improvement over previous methods in terms of speech quality measures. Finally, we analyze the contribution of all components of the proposed algorithm indicating their combined importance.
Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot
IEEE ACM Trans. Audio Speech Lang. Process.1