EDBT 2026 Demo / reviewers in the wild / expert
Yael Segal-Feldman
dblp:239/6443 · also Yael Segal
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-2513-0752ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Transcription: Mechanistic Interpretability in ASRabstractInterpretability methods have recently gained significant attention, particularly in the context of large language models, enabling insights into linguistic representations, error detection, and model behaviors such as hallucinations and repetitions. However, these techniques remain underexplored in automatic speech recognition (ASR), despite their potential to advance both the performance and interpretability of ASR systems. In this work, we adapt and systematically apply established interpretability methods such as logit lens, linear probing, and activation patching, to examine how acoustic and semantic information evolves across layers in ASR systems. Our experiments reveal previously unknown internal dynamics, including specific encoder-decoder interactions responsible for repetition hallucinations and semantic biases encoded deep within acoustic representations. These insights demonstrate the benefits of extending and applying interpretability techniques to speech recognition, opening promising directions for future research on improving model transparency and robustness. Neta Glazer, Yael Segal-Feldman, Hilit Segev, Aviv Shamsian, Asaf Buchnick, Gill Hetz, Ethan Fetaya, Joseph Keshet, Aviv Navon |
AAAI | 2 |
| 2026 | Open-vocabulary keyword spotting with hyper-matched filters for small footprint devicesabstractOpen-vocabulary keyword spotting (KWS) refers to the task of detecting words or terms within speech recordings, regardless of whether they were included in the training data. This paper introduces an open-vocabulary keyword spotting model with state-of-the-art detection accuracy for small-footprint devices. The model is composed of a speech encoder, a target keyword encoder, and a detection network. The speech encoder is either a tiny Whisper or a tiny Conformer. The target keyword encoder is implemented as a hyper-network that takes the desired keyword as a character string and generates a unique set of weights for a convolutional layer, which can be considered as a keyword-specific matched filter. The detection network uses the matched-filter weights to perform a keyword-specific convolution, which guides the cross-attention mechanism of a Perceiver module in determining whether the target term appears in the recording. The results indicate that our system achieves state-of-the-art detection performance and generalizes effectively to out-of-domain conditions, including second-language (L2) speech. Notably, our smallest model, with just 4.2 million parameters, matches or outperforms models that are several times larger, demonstrating both efficiency and robustness. Yael Segal-Feldman, Ann R. Bradlow, Matthew Goldrick 0001, Joseph Keshet |
Comput. Speech Lang. | 1 |
| 2025 | Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASRabstractLarge transformer-based models have significant potential for speech transcription and translation. Their self-attention mechanisms and parallel processing enable them to capture complex patterns and dependencies in audio sequences. However, this potential comes with challenges, as these large and computationally intensive models lead to slow inference speeds. Various optimization strategies have been proposed to improve performance, including efficient hardware utilization and algorithmic enhancements. In this paper, we introduce Whisper-Medusa, a novel approach designed to enhance processing speed with minimal impact on Word Error Rate (WER). The proposed model extends the OpenAI’s Whisper architecture by predicting multiple tokens per iteration, resulting in a 50% reduction in latency. We showcase the effectiveness of Whisper-Medusa across different learning setups and datasets. Yael Segal-Feldman, Aviv Shamsian, Aviv Navon, Gill Hetz, Joseph Keshet |
ICASSP | 1 |
| 2025 | FlowTSE: Target Speaker Extraction with Flow Matching
Aviv Navon, Aviv Shamsian, Yael Segal-Feldman, Neta Glazer, Gil Hetz, Joseph Keshet |
INTERSPEECH | 3 |
| 2025 | Enhancing analysis of diadochokinetic speech using deep neural networks
Yael Segal-Feldman, Kasia Hitczenko, Matthew Goldrick 0001, Adam Buchwald, Angela Roberts 0001, Joseph Keshet |
Comput. Speech Lang. | 1 |
| 2024 | HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing
Arnon Turetzky, Or Tal, Yael Segal-Feldman, Yehoshua Dissen, Ella Zeldes, Amit Roth, Eyal Cohen, Yosi Shrem, Bronya Roni Chernyak, Olga Seleznova, Joseph Keshet, Yossi Adi |
INTERSPEECH | 3 |
| 2022 | DeepFry: Identifying Vocal Fry Using Deep Neural NetworksabstractVocal fry or creaky voice refers to a voice quality characterized by irregular glottal opening and low pitch. It occurs in diverse languages and is prevalent in American English, where it is used not only to mark phrase finality, but also sociolinguistic factors and affect. Due to its irregular periodicity, creaky voice challenges automatic speech processing and recognition systems, particularly for languages where creak is frequently used. This paper proposes a deep learning model to detect creaky voice in fluent speech. The model is composed of an encoder and a classifier trained together. The encoder takes the raw waveform and learns a representation using a convolutional neural network. The classifier is implemented as a multi-headed fully-connected network trained to detect creaky voice, voicing, and pitch, where the last two are used to refine creak prediction. The model is trained and tested on speech of American English speakers, annotated for creak by trained phoneticians. We evaluated the performance of our system using two encoders: one is tailored for the task, and the other is based on a state-of-the-art unsupervised representation. Results suggest our best-performing system has improved recall and F1 scores compared to previous methods on unseen data. Bronya Roni Chernyak, Talia Ben Simon, Yael Segal-Feldman, Jeremy Steffman, Eleanor Chodroff, Jennifer Cole 0001, Joseph Keshet |
INTERSPEECH | 3 |
| 2022 | DDKtor: Automatic Diadochokinetic Speech AnalysisabstractDiadochokinetic speech tasks (DDK), in which participants repeatedly produce syllables, are commonly used as part of the assessment of speech motor impairments.These studies rely on manual analyses that are time-intensive, subjective, and provide only a coarse-grained picture of speech.This paper presents two deep neural network models that automatically segment consonants and vowels from unannotated, untranscribed speech.Both models work on the raw waveform and use convolutional layers for feature extraction.The first model is based on an LSTM classifier followed by fully connected layers, while the second model adds more convolutional layers followed by fully connected layers.These segmentations predicted by the models are used to obtain measures of speech rate and sound duration.Results on a young healthy individuals dataset show that our LSTM model outperforms the current state-of-the-art systems and performs comparably to trained human annotators.Moreover, the LSTM model also presents comparable results to trained human annotators when evaluated on unseen older individuals with Parkinson's Disease dataset. Yael Segal-Feldman, Kasia Hitczenko, Matthew Goldrick 0001, Adam Buchwald, Angela Roberts 0001, Joseph Keshet |
INTERSPEECH | 1 |
| 2021 | CNN-Based Spoken Term Detection and Localization without Dynamic ProgrammingabstractIn this paper, we propose a spoken term detection algorithm for simultaneous prediction and localization of in-vocabulary and out-of-vocabulary terms within an audio segment. The proposed algorithm infers whether a term was uttered within a given speech signal or not by predicting the word embeddings of various parts of the speech signal and comparing them to the word embedding of the desired term. The algorithm utilizes an existing embedding space for this task and does not need to train a task-specific embedding space. At inference the algorithm simultaneously predicts all possible locations of the target term and does not need dynamic programming for optimal search. We evaluate our system on several spoken term detection tasks on read speech corpora. Tzeviya Fuchs, Yael Segal-Feldman, Joseph Keshet |
ICASSP | 2 |
| 2021 | Pitch Estimation by Multiple Octave DecodersabstractPitch estimation is an essential task in audio processing due to its key role in many speech and music applications. Still, accurately predicting a continuous value from a high range of pitch frequencies is a challenging task. Inspired by the success of signal processing filterbank methods, we propose a novel deep architecture for accurate pitch estimation. The proposed method is composed of an encoder and multiple decoders. The encoder is implemented by a convolutional neural network that provides a good representation of the raw audio signal, and its output is fed into a set of decoders. Each decoder predicts the pitch value within a specific frequency band and is implemented by a fully-connected neural network. Such a construction allows each decoder to specialize in a particular frequency regime, which turns into a more accurate estimation of pitch values for music and speech signals. Yael Segal-Feldman, May Arama-Chayoth, Joseph Keshet |
IEEE Signal Process. Lett. | 1 |
| 2019 | SpeechYOLO: Detection and Localization of Speech ObjectsabstractIn this paper, we propose to apply object detection methods from the vision domain on the speech recognition domain, by treating audio fragments as objects. More specifically, we present SpeechYOLO, which is inspired by the YOLO algorithm for object detection in images. The goal of SpeechYOLO is to localize boundaries of utterances within the input signal, and to correctly classify them. Our system is composed of a convolutional neural network, with a simple least-mean-squares loss function. We evaluated the system on several keyword spotting tasks, that include corpora of read speech and spontaneous speech. Our system compares favorably with other algorithms trained for both localization and classification. Yael Segal-Feldman, Tzeviya Fuchs, Joseph Keshet |
INTERSPEECH | 1 |