EDBT 2026 Demo / reviewers in the wild / expert
Joseph Keshet
dblp:45/4451 · also Yossi Keshet
· DBLP profile ↗
66ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0003-2332-5783ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 51 · 5 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 6 first-author · 19 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Transcription: Mechanistic Interpretability in ASRabstractInterpretability methods have recently gained significant attention, particularly in the context of large language models, enabling insights into linguistic representations, error detection, and model behaviors such as hallucinations and repetitions. However, these techniques remain underexplored in automatic speech recognition (ASR), despite their potential to advance both the performance and interpretability of ASR systems. In this work, we adapt and systematically apply established interpretability methods such as logit lens, linear probing, and activation patching, to examine how acoustic and semantic information evolves across layers in ASR systems. Our experiments reveal previously unknown internal dynamics, including specific encoder-decoder interactions responsible for repetition hallucinations and semantic biases encoded deep within acoustic representations. These insights demonstrate the benefits of extending and applying interpretability techniques to speech recognition, opening promising directions for future research on improving model transparency and robustness. Neta Glazer, Yael Segal-Feldman, Hilit Segev, Aviv Shamsian, Asaf Buchnick, Gill Hetz, Ethan Fetaya, Joseph Keshet, Aviv Navon |
AAAI | 8 |
| 2026 | Open-vocabulary keyword spotting with hyper-matched filters for small footprint devicesabstractOpen-vocabulary keyword spotting (KWS) refers to the task of detecting words or terms within speech recordings, regardless of whether they were included in the training data. This paper introduces an open-vocabulary keyword spotting model with state-of-the-art detection accuracy for small-footprint devices. The model is composed of a speech encoder, a target keyword encoder, and a detection network. The speech encoder is either a tiny Whisper or a tiny Conformer. The target keyword encoder is implemented as a hyper-network that takes the desired keyword as a character string and generates a unique set of weights for a convolutional layer, which can be considered as a keyword-specific matched filter. The detection network uses the matched-filter weights to perform a keyword-specific convolution, which guides the cross-attention mechanism of a Perceiver module in determining whether the target term appears in the recording. The results indicate that our system achieves state-of-the-art detection performance and generalizes effectively to out-of-domain conditions, including second-language (L2) speech. Notably, our smallest model, with just 4.2 million parameters, matches or outperforms models that are several times larger, demonstrating both efficiency and robustness. Yael Segal-Feldman, Ann R. Bradlow, Matthew Goldrick 0001, Joseph Keshet |
Comput. Speech Lang. | 4 |
| 2025 | WhisperNER: Unified Open Named Entity and Speech RecognitionabstractIntegrating named entity recognition (NER) with automatic speech recognition (ASR) can significantly enhance transcription accuracy and enrich its content. We introduce WhisperNER, a novel model that facilitates joint speech transcription and entity recognition. WhisperNER supports opentype NER, enabling recognition of various entities during inference. Building on recent advancements in open NER research, we augment a large synthetic dataset with synthetic speech samples. This approach enables us to train WhisperNER on numerous examples with various NER tags. During training, the model is prompted with NER labels and optimized to produce the transcribed utterance alongside the corresponding tagged entities. For evaluation, we generate synthetic speech for commonly used NER benchmarks and annotate existing ASR datasets with open NER tags. Our experiments show that WhisperNER outperforms natural baselines in both out-of-domain open-type NER and supervised fine-tuning. Gil Ayache, Menachem Pirchi, Aviv Navon, Aviv Shamsian, Gill Hetz, Joseph Keshet |
ASRU | 6 |
| 2025 | Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASRabstractLarge transformer-based models have significant potential for speech transcription and translation. Their self-attention mechanisms and parallel processing enable them to capture complex patterns and dependencies in audio sequences. However, this potential comes with challenges, as these large and computationally intensive models lead to slow inference speeds. Various optimization strategies have been proposed to improve performance, including efficient hardware utilization and algorithmic enhancements. In this paper, we introduce Whisper-Medusa, a novel approach designed to enhance processing speed with minimal impact on Word Error Rate (WER). The proposed model extends the OpenAI’s Whisper architecture by predicting multiple tokens per iteration, resulting in a 50% reduction in latency. We showcase the effectiveness of Whisper-Medusa across different learning setups and datasets. Yael Segal-Feldman, Aviv Shamsian, Aviv Navon, Gill Hetz, Joseph Keshet |
ICASSP | 5 |
| 2025 | FlowTSE: Target Speaker Extraction with Flow Matching
Aviv Navon, Aviv Shamsian, Yael Segal-Feldman, Neta Glazer, Gil Hetz, Joseph Keshet |
INTERSPEECH | 6 |
| 2025 | Spectral Analysis of Diffusion Models with Application to Schedule DesignabstractDiffusion models (DMs) have emerged as powerful tools for modeling complex data distributions and generating realistic new samples. Over the years, advanced architectures and sampling methods have been developed to make these models practically usable. However, certain synthesis process decisions still rely on heuristics without a solid theoretical foundation.
In our work, we offer a novel analysis of the DM's inference process, introducing a comprehensive frequency response perspective. Specifically, by relying on Gaussianity assumption, we present the inference process as a closed-form spectral transfer function, capturing how the generated signal evolves in response to the initial noise. We demonstrate how the proposed analysis can be leveraged to design a noise schedule that aligns effectively with the characteristics of the data. The spectral perspective also provides insights into the underlying dynamics and sheds light on the relationship between spectral properties and noise schedule structure. Our results lead to scheduling curves that are dependent on the spectral content of the data, offering a theoretical justification for some of the heuristics taken by practitioners. Roi Benita, Miki Elad, Joseph Keshet |
NeurIPS | 3 |
| 2025 | Enhancing analysis of diadochokinetic speech using deep neural networks
Yael Segal-Feldman, Kasia Hitczenko, Matthew Goldrick 0001, Adam Buchwald, Angela Roberts 0001, Joseph Keshet |
Comput. Speech Lang. | 6 |
| 2024 | Open-Vocabulary Keyword-Spotting with Adaptive Instance NormalizationabstractOpen vocabulary keyword spotting is a crucial and challenging task in automatic speech recognition (ASR) that focuses on detecting user-defined keywords within a spoken utterance. Keyword spotting methods commonly map the audio utterance and keyword into a joint embedding space to obtain some affinity score. In this work, we propose AdaKWS, a novel method for keyword spotting in which a text encoder is trained to output keyword-conditioned normalization parameters. These parameters are used to process the auditory input. We provide an extensive evaluation using challenging and diverse multi-lingual benchmarks and show significant improvements over recent keyword spotting and ASR baselines. Furthermore, we study the effectiveness of our approach on low-resource languages that were unseen during the training. The results demonstrate a substantial performance improvement compared to baseline methods. Aviv Navon, Aviv Shamsian, Neta Glazer, Gill Hetz, Joseph Keshet |
ICASSP | 5 |
| 2024 | DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform GenerationabstractDiffusion models have recently been shown to be relevant for high-quality speech generation. Most work has been focused on generating spectrograms, and as such, they further require a subsequent model to convert the spectrogram to a waveform (i.e., a vocoder). This work proposes a diffusion probabilistic end-to-end model for generating a raw speech waveform. The proposed model is autoregressive, generating overlapping frames sequentially, where each frame is conditioned on a portion of the previously generated one. Hence, our model can effectively synthesize an unlimited speech duration while preserving high-fidelity synthesis and temporal coherence. We implemented the proposed model for unconditional and conditional speech generation, where the latter can be driven by an input sequence of phonemes, amplitudes, and pitch values. Working on the waveform directly has some empirical advantages. Specifically, it allows the creation of local acoustic behaviors, like vocal fry, which makes the overall waveform sounds more natural. Furthermore, the proposed diffusion model is stochastic and not deterministic; therefore, each inference generates a slightly different waveform variation, enabling abundance of valid realizations. Experiments show that the proposed model generates speech with superior quality compared with other state-of-the-art neural speech generation systems. Roi Benita, Michael Elad, Joseph Keshet |
ICLR | 3 |
| 2024 | Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
Yehoshua Dissen, Shiry Yonash, Israel Cohen, Joseph Keshet |
INTERSPEECH | 4 |
| 2024 | Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment
Rotem Rousso, Eyal Cohen, Joseph Keshet, Eleanor Chodroff |
INTERSPEECH | 3 |
| 2024 | Keyword-Guided Adaptation of Automatic Speech Recognition
Aviv Shamsian, Aviv Navon, Neta Glazer, Gill Hetz, Joseph Keshet |
INTERSPEECH | 5 |
| 2024 | HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing
Arnon Turetzky, Or Tal, Yael Segal-Feldman, Yehoshua Dissen, Ella Zeldes, Amit Roth, Eyal Cohen, Yosi Shrem, Bronya Roni Chernyak, Olga Seleznova, Joseph Keshet, Yossi Adi |
INTERSPEECH | 11 |
| 2022 | DeepFry: Identifying Vocal Fry Using Deep Neural NetworksabstractVocal fry or creaky voice refers to a voice quality characterized by irregular glottal opening and low pitch. It occurs in diverse languages and is prevalent in American English, where it is used not only to mark phrase finality, but also sociolinguistic factors and affect. Due to its irregular periodicity, creaky voice challenges automatic speech processing and recognition systems, particularly for languages where creak is frequently used. This paper proposes a deep learning model to detect creaky voice in fluent speech. The model is composed of an encoder and a classifier trained together. The encoder takes the raw waveform and learns a representation using a convolutional neural network. The classifier is implemented as a multi-headed fully-connected network trained to detect creaky voice, voicing, and pitch, where the last two are used to refine creak prediction. The model is trained and tested on speech of American English speakers, annotated for creak by trained phoneticians. We evaluated the performance of our system using two encoders: one is tailored for the task, and the other is based on a state-of-the-art unsupervised representation. Results suggest our best-performing system has improved recall and F1 scores compared to previous methods on unseen data. Bronya Roni Chernyak, Talia Ben Simon, Yael Segal-Feldman, Jeremy Steffman, Eleanor Chodroff, Jennifer Cole 0001, Joseph Keshet |
INTERSPEECH | 7 |
| 2022 | Self-supervised Speaker DiarizationabstractOver the last few years, deep learning has grown in popularity for speaker verification, identification, and diarization.Inarguably, a significant part of this success is due to the demonstrated effectiveness of their speaker representations.These, however, are heavily dependent on large amounts of annotated data and can be sensitive to new domains.This study proposes an entirely unsupervised deep-learning model for speaker diarization.Specifically, the study focuses on generating highquality neural speaker representations without any annotated data, as well as on estimating secondary hyperparameters of the model without annotations.The speaker embeddings are represented by an encoder trained in a self-supervised fashion using pairs of adjacent segments assumed to be of the same speaker.The trained encoder model is then used to self-generate pseudo-labels to subsequently train a similarity score between different segments of the same call using probabilistic linear discriminant analysis (PLDA) and further to learn a clustering stopping threshold.We compared our model to state-of-the-art unsupervised as well as supervised baselines on the CallHome benchmarks.According to empirical results, our approach outperforms unsupervised methods when only two speakers are present in the call, and is only slightly worse than recent supervised models. Yehoshua Dissen, Felix Kreuk, Joseph Keshet |
INTERSPEECH | 3 |
| 2022 | Unsupervised Word Segmentation using K Nearest Neighbors
Tzeviya Fuchs, Yedid Hoshen, Joseph Keshet |
INTERSPEECH | 3 |
| 2022 | DDKtor: Automatic Diadochokinetic Speech AnalysisabstractDiadochokinetic speech tasks (DDK), in which participants repeatedly produce syllables, are commonly used as part of the assessment of speech motor impairments.These studies rely on manual analyses that are time-intensive, subjective, and provide only a coarse-grained picture of speech.This paper presents two deep neural network models that automatically segment consonants and vowels from unannotated, untranscribed speech.Both models work on the raw waveform and use convolutional layers for feature extraction.The first model is based on an LSTM classifier followed by fully connected layers, while the second model adds more convolutional layers followed by fully connected layers.These segmentations predicted by the models are used to obtain measures of speech rate and sound duration.Results on a young healthy individuals dataset show that our LSTM model outperforms the current state-of-the-art systems and performs comparably to trained human annotators.Moreover, the LSTM model also presents comparable results to trained human annotators when evaluated on unseen older individuals with Parkinson's Disease dataset. Yael Segal-Feldman, Kasia Hitczenko, Matthew Goldrick 0001, Adam Buchwald, Angela Roberts 0001, Joseph Keshet |
INTERSPEECH | 6 |
| 2022 | Formant Estimation and Tracking using Probabilistic Heat-MapsabstractFormants are the spectral maxima that result from acoustic resonances of the human vocal tract, and their accurate estimation is among the most fundamental speech processing problems.Recent work has been shown that those frequencies can accurately be estimated using deep learning techniques.However, when presented with a speech from a different domain than that in which they have been trained on, these methods exhibit a decline in performance, limiting their usage as generic tools.The contribution of this paper is to propose a new network architecture that performs well on a variety of different speaker and speech domains.Our proposed model is composed of a shared encoder that gets as input a spectrogram and outputs a domain-invariant representation.Then, multiple decoders further process this representation, each responsible for predicting a different formant while considering the lower formant predictions.An advantage of our model is that it is based on heatmaps that generate a probability distribution over formant predictions.Results suggest that our proposed model better represents the signal over various domains and leads to better formant frequency tracking and estimation. Yosi Shrem, Felix Kreuk, Joseph Keshet |
INTERSPEECH | 3 |
| 2022 | Correcting Mispronunciations in Speech using Spectrogram InpaintingabstractLearning a new language involves constantly comparing speech productions with reference productions from the environment.Early in speech acquisition, children make articulatory adjustments to match their caregivers' speech.Grownup learners of a language tweak their speech to match the tutor reference.This paper proposes a method to synthetically generate correct pronunciation feedback given incorrect production.Furthermore, our aim is to generate the corrected production while maintaining the speaker's original voice.The system prompts the user to pronounce a phrase.The speech is recorded, and the samples associated with the inaccurate phoneme are masked with zeros.This waveform serves as an input to a speech generator, implemented as a deep learning inpainting system with a U-net architecture, and trained to output a reconstructed speech.The training set is composed of unimpaired proper speech examples, and the generator is trained to reconstruct the original proper speech.We evaluated the performance of our system on phoneme replacement of minimal pair words of English as well as on children with pronunciation disorders.Results suggest that human listeners slightly prefer our generated speech over a smoothed replacement of the inaccurate phoneme with a production of a different speaker. Talia Ben Simon, Felix Kreuk, Faten Awwad, Jacob T. Cohen, Joseph Keshet |
INTERSPEECH | 5 |
| 2022 | A Baseline for Detecting Out-of-Distribution Examples in Image CaptioningabstractImage captioning research achieved breakthroughs in recent years by developing neural models that can generate diverse and high-quality descriptions for images drawn from the same distribution as training images. However, when facing out-of-distribution (OOD) images, such as corrupted images, or images containing unknown objects, the models fail in generating relevant captions. Gal-Lev Shalev, Gabi Shalev, Joseph Keshet |
ACM Multimedia | 3 |
| 2022 | Speech Time-Scale Modification With GANsabstractWhile listening to spoken content, it is often desired to vary the speech rate while preserving the speaker’s timbre and pitch. To date, advanced signal processing techniques are used to address this task, but it still remains a challenge to maintain a high speech quality at all time-scales. Inspired by the success of speech generation using Generative Adversarial Networks (GANs), we propose a novel unsupervised learning algorithm for time-scale modification (TSM) of speech, called ScalerGAN. The model is trained using a set of speech utterances, where no time-scales are provided. The ScalerGAN algorithm is composed of a generator that gets as input speech with the desired rate and outputs a time-adjusted speech; a discriminator that works on various spectrum scales; and a decoder that converts the time-adjusted signal back to the original rate to maintain consistency. Using an A/B test and conditional A/B test, human listeners were asked to compare ScalerGAN with other state-of-the-art TSM methods. The results showed that the speech quality of ScalerGAN outperforms all other methods. Eyal Cohen, Felix Kreuk, Joseph Keshet |
IEEE Signal Process. Lett. | 3 |
| 2021 | Fairness in the Eyes of the Data: Certifying Machine-Learning ModelsabstractWe present a framework that allows to certify the fairness degree of a model based on an interactive and privacy-preserving test. The framework verifies any trained model, regardless of its training process and architecture. Thus, it allows us to evaluate any deep learning model on multiple fairness definitions empirically. We tackle two scenarios, where either the test data is privately available only to the tester or is publicly known in advance, even to the model creator. We investigate the soundness of the proposed approach using theoretical analysis and present statistical guarantees for the interactive test. Finally, we provide a cryptographic technique to automate fairness testing and certified inference with only black-box access to the model at hand while hiding the participants' sensitive data. Shahar Segal, Yossi Adi, Benny Pinkas, Carsten Baum, Chaya Ganesh, Joseph Keshet |
AIES | 6 |
| 2021 | CNN-Based Spoken Term Detection and Localization without Dynamic ProgrammingabstractIn this paper, we propose a spoken term detection algorithm for simultaneous prediction and localization of in-vocabulary and out-of-vocabulary terms within an audio segment. The proposed algorithm infers whether a term was uttered within a given speech signal or not by predicting the word embeddings of various parts of the speech signal and comparing them to the word embedding of the desired term. The algorithm utilizes an existing embedding space for this task and does not need to train a task-specific embedding space. At inference the algorithm simultaneously predicts all possible locations of the target term and does not need dynamic programming for optimal search. We evaluate our system on several spoken term detection tasks on read speech corpora. Tzeviya Fuchs, Yael Segal-Feldman, Joseph Keshet |
ICASSP | 3 |
| 2021 | Pitch Estimation by Multiple Octave DecodersabstractPitch estimation is an essential task in audio processing due to its key role in many speech and music applications. Still, accurately predicting a continuous value from a high range of pitch frequencies is a challenging task. Inspired by the success of signal processing filterbank methods, we propose a novel deep architecture for accurate pitch estimation. The proposed method is composed of an encoder and multiple decoders. The encoder is implemented by a convolutional neural network that provides a good representation of the raw audio signal, and its output is fed into a set of decoders. Each decoder predicts the pitch value within a specific frequency band and is implemented by a fully-connected neural network. Such a construction allows each decoder to specialize in a particular frequency regime, which turns into a more accurate estimation of pitch values for music and speech signals. Yael Segal-Feldman, May Arama-Chayoth, Joseph Keshet |
IEEE Signal Process. Lett. | 3 |
| 2020 | Phoneme Boundary Detection Using Learnable Segmental FeaturesabstractPhoneme boundary detection plays an essential first step for a variety of speech processing applications such as speaker diarization, speech science, keyword spotting, etc. In this work, we propose a neural architecture coupled with a parameterized structured loss function to learn segmental representations for the task of phoneme boundary detection. First, we evaluated our model when the spoken phonemes were not given as input. Results on the TIMIT and Buckeye corpora suggest that the proposed model is superior to the baseline models and reaches state-of-the-art performance in terms of F1 and R-value. We further explore the use of phonetic transcription as additional supervision and show this yields minor improvements in performance but substantially better convergence rates. We additionally evaluate the model on a He-brew corpus and demonstrate such phonetic supervision can be beneficial in a multi-lingual setting. Felix Kreuk, Yaniv Sheena, Joseph Keshet, Yossi Adi |
ICASSP | 3 |
| 2020 | Hide and Speak: Towards Deep Neural Networks for Speech SteganographyabstractSteganography is the science of hiding a secret message within an ordinary public message, which is referred to as Carrier. Traditionally, digital signal processing techniques, such as least significant bit encoding, were used for hiding messages. In this paper, we explore the use of deep neural networks as steganographic functions for speech data. We showed that steganography models proposed for vision are less suitable for speech, and propose a new model that includes the short-time Fourier transform and inverse-short-time Fourier transform as differentiable layers within the network, thus imposing a vital constraint on the network outputs. We empirically demonstrated the effectiveness of the proposed method comparing to deep learning based on several speech datasets and analyzed the results quantitatively and qualitatively. Moreover, we showed that the proposed approach could be applied to conceal multiple messages in a single carrier using multiple decoders or a single conditional decoder. Lastly, we evaluated our model under different channel distortions. Qualitative experiments suggest that modifications to the carrier are unnoticeable by human listeners and that the decoded messages are highly intelligible. Felix Kreuk, Yossi Adi, Bhiksha Raj, Rita Singh, Joseph Keshet |
INTERSPEECH | 5 |
| 2020 | Self-Supervised Contrastive Learning for Unsupervised Phoneme SegmentationabstractWe propose a self-supervised representation learning model for the task of unsupervised phoneme boundary detection. The model is a convolutional neural network that operates directly on the raw waveform. It is optimized to identify spectral changes in the signal using the Noise-Contrastive Estimation principle. At test time, a peak detection algorithm is applied over the model outputs to produce the final boundaries. As such, the proposed model is trained in a fully unsupervised manner with no manual annotations in the form of target boundaries nor phonetic transcriptions. We compare the proposed approach to several unsupervised baselines using both TIMIT and Buckeye corpora. Results suggest that our approach surpasses the baseline models and reaches state-of-the-art performance on both data sets. Furthermore, we experimented with expanding the training set with additional examples from the Librispeech corpus. We evaluated the resulting model on distributions and languages that were not seen during the training phase (English, Hebrew and German) and showed that utilizing additional untranscribed data is beneficial for model performance. Felix Kreuk, Joseph Keshet, Yossi Adi |
INTERSPEECH | 2 |
| 2020 | Minimal Modifications of Deep Neural Networks using VerificationabstractDeep neural networks (DNNs) are revolutionizing the way complex systems are de- signed, developed and maintained. As part of the life cycle of DNN-based systems, there is often a need to modify a DNN in subtle ways that affect certain aspects of its behav- ior, while leaving other aspects of its behavior unchanged (e.g., if a bug is discovered and needs to be fixed, without altering other functionality). Unfortunately, retraining a DNN is often difficult and expensive, and may produce a new DNN that is quite different from the original. We leverage recent advances in DNN verification and propose a technique for modifying a DNN according to certain requirements, in a way that is provably minimal, does not require any retraining, and is thus less likely to affect other aspects of the DNN’s behavior. Using a proof-of-concept implementation, we demonstrate the usefulness and potential of our approach in addressing two real-world needs: (i) measuring the resilience of DNN watermarking schemes; and (ii) bug repair in already-trained DNNs. Ben Goldberger, Guy Katz, Yossi Adi, Joseph Keshet |
LPAR | 4 |
| 2020 | Online prediction of time series with assumed behavior
Ariel Rosenfeld, Moshe Cohen, Sarit Kraus, Joseph Keshet |
Eng. Appl. Artif. Intell. | 4 |
| 2019 | SpeechYOLO: Detection and Localization of Speech ObjectsabstractIn this paper, we propose to apply object detection methods from the vision domain on the speech recognition domain, by treating audio fragments as objects. More specifically, we present SpeechYOLO, which is inspired by the YOLO algorithm for object detection in images. The goal of SpeechYOLO is to localize boundaries of utterances within the input signal, and to correctly classify them. Our system is composed of a convolutional neural network, with a simple least-mean-squares loss function. We evaluated the system on several keyword spotting tasks, that include corpora of read speech and spontaneous speech. Our system compares favorably with other algorithms trained for both localization and classification. Yael Segal-Feldman, Tzeviya Fuchs, Joseph Keshet |
INTERSPEECH | 3 |
| 2019 | Dr.VOT: Measuring Positive and Negative Voice Onset Time in the WildabstractVoice Onset Time (VOT), a key measurement of speech for basic research and applied medical studies, is the time between the onset of a stop burst and the onset of voicing. When the voicing onset precedes burst onset the VOT is negative; if voicing onset follows the burst, it is positive. In this work, we present a deep-learning model for accurate and reliable measurement of VOT in naturalistic speech. The proposed system addresses two critical issues: it can measure positive and negative VOT equally well, and it is trained to be robust to variation across annotations. Our approach is based on the structured prediction framework, where the feature functions are defined to be RNNs. These learn to capture segmental variation in the signal. Results suggest that our method substantially improves over the current state-of-the-art. In contrast to previous work, our Deep and Robust VOT annotator, Dr.VOT, can successfully estimate negative VOTs while maintaining state-of-the-art performance on positive VOTs. This high level of performance generalizes to new corpora without further retraining. Index Terms: structured prediction, multi-task learning, adversarial training, recurrent neural networks, sequence segmentation. Yosi Shrem, Matthew Goldrick 0001, Joseph Keshet |
INTERSPEECH | 3 |
| 2018 | Fooling End-To-End Speaker Verification With Adversarial ExamplesabstractAutomatic speaker verification systems are increasingly used as the primary means to authenticate costumers. Recently, it has been proposed to train speaker verification systems using end-to-end deep neural models. In this paper, we show that such systems are vulnerable to adversarial example attacks. Adversarial examples are generated by adding a peculiar noise to original speaker examples, in such a way that they are almost indistinguishable, by a human listener. Yet, the generated waveforms, which sound as speaker A can be used to fool such a system by claiming as if the waveforms were uttered by speaker B. We present white-box attacks on a deep end-to-end network that was either trained on YOHO or NTIMIT. We also present two black-box attacks. In the first one, we generate adversarial examples with a system trained on NTIMIT and perform the attack on a system that trained on YOHO. In the second one, we generate the adversarial examples with a system trained using Mel-spectrum features and perform the attack on a system trained using MFCCs. Our results show that one can significantly decrease the accuracy of a target system even when the adversarial examples are generated with different system potentially using different features. Felix Kreuk, Yossi Adi, Moustapha Cissé, Joseph Keshet |
ICASSP | 4 |
| 2018 | Out-of-Distribution Detection using Multiple Semantic Label RepresentationsabstractDeep Neural Networks are powerful models that attained remarkable results on a variety of tasks. These models are shown to be extremely efficient when training and test data are drawn from the same distribution. However, it is not clear how a network will act when it is fed with an out-of-distribution example. In this work, we consider the problem of out-of-distribution detection in neural networks. We propose to use multiple semantic dense representations instead of sparse representation as the target label. Specifically, we propose to use several word representations obtained from different corpora or architectures as target labels. We evaluated the proposed model on computer vision, and speech commands detection tasks and compared it to previous methods. Results suggest that our method compares favorably with previous work. Besides, we present the efficiency of our approach for detecting wrongly classified and adversarial examples. Gabi Shalev, Yossi Adi, Joseph Keshet |
NeurIPS | 3 |
| 2018 | Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by Backdooring
Yossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas, Joseph Keshet |
USENIX Security Symposium | 5 |
| 2017 | Sequence segmentation using joint RNN and structured prediction modelsabstractWe describe and analyze a simple and effective algorithm for sequence segmentation applied to speech processing tasks. We propose a neural architecture that is composed of two modules trained jointly: a recurrent neural network (RNN) module and a structured prediction model. The RNN outputs are considered as feature functions to the structured model. The overall model is trained with a structured loss function which can be designed to the given segmentation task. We demonstrate the effectiveness of our method by applying it to two simple tasks commonly used in phonetic studies: word segmentation and voice onset time segmentation. Results suggest the proposed model is superior to previous methods, obtaining state-of-the-art results on the tested datasets. Yossi Adi, Joseph Keshet, Emily Cibelli, Matthew Goldrick 0001 |
ICASSP | 2 |
| 2017 | Learning Similarity Functions for Pronunciation VariationsabstractA significant source of errors in Automatic Speech Recognition (ASR) systems is due to pronunciation variations which occur in spontaneous and conversational speech. Usually ASR systems use a finite lexicon that provides one or more pronunciations for each word. In this paper, we focus on learning a similarity function between two pronunciations. The pronunciations can be the canonical and the surface pronunciations of the same word or they can be two surface pronunciations of different words. This task generalizes problems such as lexical access (the problem of learning the mapping between words and their possible pronunciations), and defining word neighborhoods. It can also be used to dynamically increase the size of the pronunciation lexicon, or in predicting ASR errors. We propose two methods, which are based on recurrent neural networks, to learn the similarity function. The first is based on binary classification, and the second is based on learning the ranking of the pronunciations. We demonstrate the efficiency of our approach on the task of lexical access using a subset of the Switchboard conversational speech corpus. Results suggest that on this task our methods are superior to previous methods which are based on graphical Bayesian methods. Einat Naaman, Yossi Adi, Joseph Keshet |
INTERSPEECH | 3 |
| 2017 | Automatic Measurement of Pre-AspirationabstractPre-aspiration is defined as the period of glottal friction occurring in sequences of vocalic/consonantal sonorants and phonetically voiceless obstruents. We propose two machine learning methods for automatic measurement of pre-aspiration duration: a feedforward neural network, which works at the frame level; and a structured prediction model, which relies on manually designed feature functions, and works at the segment level. The input for both algorithms is a speech signal of an arbitrary length containing a single obstruent, and the output is a pair of times which constitutes the pre-aspiration boundaries. We train both models on a set of manually annotated examples. Results suggest that the structured model is superior to the frame-based model as it yields higher accuracy in predicting the boundaries and generalizes to new speakers and new languages. Finally, we demonstrate the applicability of our structured prediction algorithm by replicating linguistic analysis of pre-aspiration in Aberystwyth English with high correlation. Yaniv Sheena, Mísa Hejná, Yossi Adi, Joseph Keshet |
INTERSPEECH | 4 |
| 2017 | Houdini: Fooling Deep Structured Visual and Speech Recognition Models with Adversarial ExamplesabstractGenerating adversarial examples is a critical step for evaluating and improving the robustness of learning machines. So far, most existing methods only work for classification and are not designed to alter the true performance measure of the problem at hand. We introduce a novel flexible approach named Houdini for generating adversarial examples specifically tailored for the final performance measure of the task considered, be it combinatorial and non-decomposable. We successfully apply Houdini to a range of applications such as speech recognition, pose estimation and semantic segmentation. In all cases, the attacks based on Houdini achieve higher success rate than those based on the traditional surrogates used to train the models while using a less perceptible adversarial perturbation. Moustapha Cissé, Yossi Adi, Natalia Neverova, Joseph Keshet |
NIPS | 4 |
| 2016 | Online Prediction of Exponential Decay Time Series with Human-Agent ApplicationabstractExponential decay time series are prominent in many fields. In some applications, the time series behavior can change over time due to a change in the user's preferences or a change of environment. In this paper we present an innovative online learning algorithm, which we name Exponentron, for the prediction of exponential decay time series. We state a regret bound for our setting, which theoretically compares the performance of our online algorithm relative to the performance of the best batch prediction mechanism, which can be chosen in hindsight from a class of hypotheses after observing the entire time series. In experiments with synthetic and real-world data sets, we found that the proposed algorithm compares favorably with the classic time series prediction methods by providing up to 41% improvement in prediction accuracy. Furthermore, we used the proposed algorithm for the design of a novel automated agent for the improvement of the communication process between a driver and its automotive climate control system. Throughout extensive human study with 24 drivers we show that our agent improves the communication process and increases drivers' satisfaction, exemplifying the Exponentron's applicative benefit. Ariel Rosenfeld, Joseph Keshet, Claudia V. Goldman, Sarit Kraus |
ECAI | 2 |
| 2016 | The relationship of voice onset time and Voice Offset Time to physical ageabstractIn a speech signal, Voice Onset Time (VOT) is the period between the release of a plosive and the onset of vocal cord vibrations in the production of the following sound. Voice Offset Time (VOFT), on the other hand, is the period between the end of a voiced sound and the release of the following plosive. Traditionally, VOT has been studied across multiple disciplines and has been related to many factors that influence human speech production, including physical, physiological and psychological characteristics of the speaker. The mechanism of extraction of VOT has however been largely manual, and studies have been carried out over small ensembles of individuals under very controlled conditions, usually in clinical settings. Studies of VOFT follow similar trends, but are more limited in scope due to the inherent difficulty in the extraction of VOFT from speech signals. In this paper we use a structured-prediction based mechanism for the automatic computation of VOT and VOFT. We show that for specific combinations of plosives and vowels, these are re-latable to the physical age of the speaker. The paper also highlights the ambiguities in the prediction of age from VOT and VOFT, and consequently in the use of these measures in forensic analysis of voice. Rita Singh, Joseph Keshet, Deniz Gençaga, Bhiksha Raj |
ICASSP | 2 |
| 2016 | Automatic Measurement of Voice Onset Time and Prevoicing Using Recurrent Neural NetworksabstractVoice onset time (VOT) is defined as the time difference between the onset of the burst and the onset of voicing. When voicing begins preceding the burst, the stop is called prevoiced, and the VOT is negative. When voicing begins following the burst the VOT is positive. While most of the work on automatic measurement of VOT has focused on positive VOT mostly evident in American English, in many languages the VOT can be negative. We propose an algorithm that estimates if the stop is prevoiced, and measures either positive or negative VOT, respectively. More specifically, the input to the algorithm is a speech segment of an arbitrary length containing a single stop consonant, and the output is the time of the burst onset, the duration of the burst, and the time of the prevoicing onset with a confidence. Manually labeled data is used to train a recurrent neural network that can model the dynamic temporal behavior of the input signal, and outputs the events' onset and duration. Results suggest that the proposed algorithm is superior to the current state-of-the-art both in terms of the VOT measurement and in terms of prevoicing detection. Yossi Adi, Joseph Keshet, Olga Dmitrieva, Matthew Goldrick 0001 |
INTERSPEECH | 2 |
| 2016 | Formant Estimation and Tracking Using Deep Learning
Yehoshua Dissen, Joseph Keshet |
INTERSPEECH | 2 |
| 2016 | StructED: Risk Minimization in Structured PredictionabstractStructured tasks are distinctive: each task has its own measure of performance, such as the word error rate in speech recognition, the BLEU score in machine translation, the NDCG score in information retrieval, or the intersection-over-union score in visual object segmentation. This paper presents StructED, a software package for learning structured prediction models with training methods that aimed at optimizing the task measure of performance. The package was written in Java and released under the MIT license. It can be downloaded from adiyoss.github.io/StructED. Yossi Adi, Joseph Keshet |
J. Mach. Learn. Res. | 2 |
| 2013 | Discriminative articulatory models for spoken term detection in low-resource conversational settingsabstractWe study spoken term detection (STD) - the task of determining whether and where a given word or phrase appears in a given segment of speech - using articulatory feature-based pronunciation models. The models are motivated by the requirements of STD in low-resource settings, in which it may not be feasible to train a large-vocabulary continuous speech recognition system, as well as by the need to address pronunciation variation in conversational speech. Our STD system is trained to maximize the expected area under the receiver operating characteristic curve, often used to evaluate STD performance. In experimental evaluations on the Switchboard corpus, we find that our approach outperforms a baseline HMM-based system across a number of training set sizes, as well as a discriminative phone-based model in some settings. Rohit Prabhavalkar, Karen Livescu, Eric Fosler-Lussier, Joseph Keshet |
ICASSP | 4 |
| 2013 | Predicting Human Strategic Decisions Using Facial Expressions
Noam Peled, Moshe Bitan, Joseph Keshet, Sarit Kraus |
IJCAI | 3 |
| 2013 | Learning Efficient Random Maximum A-Posteriori Predictors with Non-Decomposable Loss FunctionsabstractIn this work we develop efficient methods for learning random MAP predictors for structured label problems. In particular, we construct posterior distributions over perturbations that can be adjusted via stochastic gradient methods. We show that every smooth posterior distribution would suffice to define a smooth PAC-Bayesian risk bound suitable for gradient methods. In addition, we relate the posterior distributions to computational properties of the MAP predictors. We suggest multiplicative posteriors to learn super-modular potential functions that accompany specialized MAP predictors such as graph-cuts. We also describe label-augmented posterior models that can use efficient MAP approximations, such as those arising from linear program relaxations. Tamir Hazan, Subhransu Maji, Joseph Keshet, Tommi S. Jaakkola |
NIPS | 3 |
| 2012 | Discriminative Pronunciation Modeling: A Large-Margin, Feature-Rich Approach
Hao Tang 0002, Joseph Keshet, Karen Livescu |
ACL (1) | 2 |
| 2012 | Automatic Measurement of Positive and Negative Voice Onset TimeabstractPrevious work on automatic VOT measurement has focused on positive-valued VOT. However, in many languages VOT can be either positive or negative (“prevoiced”). We present a discriminative algorithm that simultaneously decides whether a stop is prevoiced and measures its VOT. The algorithm operates on feature functions designed to locate the burst and voicing onsets in the positive and negative VOT cases. Tested on a database of positive- and negative-VOT voiced stops, the algorithm predicts prevoicing with>90 % accuracy, and gives good agreement between automatic and manual measurements. Index Terms: voice onset time, automatic phonetic measurement, discriminative methods, structured prediction Katharine Henry, Morgan Sonderegger, Joseph Keshet |
INTERSPEECH | 3 |
| 2011 | PAC-Bayesian approach for minimization of phoneme error rateabstractWe describe a new approach for phoneme recognition which aims at minimizing the phoneme error rate. Building on structured prediction techniques, we formulate the phoneme recognizer as a linear combination of feature functions. We state a PAC-Bayesian generalization bound, which gives an upper-bound on the expected phoneme error rate in terms of the empirical phoneme error rate. Our algorithm is derived by finding the gradient of the PAC-Bayesian bound and minimizing it by stochastic gradient descent. The resulting algorithm is iterative and easy to implement. Experiments on the TIMIT corpus show that our method achieves the lowest phoneme error rate compared to other discriminative and generative models with the same expressive power. Joseph Keshet, David A. McAllester, Tamir Hazan |
ICASSP | 1 |
| 2011 | Direct Error Rate Minimization of Hidden Markov ModelsabstractWe explore discriminative training of HMM parameters that directly minimizes the expected error rate. In discriminative training one is interested in training a system to minimize a desired error function, like word error rate, phone error rate, or frame error rate. We review a recent method (McAllester, Hazan and Keshet, 2010), which introduces an analytic expression for the gradient of the expected error-rate. The analytic expression leads to a perceptron-like update rule, which is adapted here for training of HMMs in an online fashion. While the proposed method can work with any type of the error function used in speech recognition, we evaluated it on phoneme recognition of TIMIT, when the desired error function used for training was frame error rate. Except for the case of GMM with a single mixture per state, the proposed update rule provides lower error rates, both in terms of frame error rate and phone error rate, than other approaches, including MCE and large margin. Index Terms: hidden Markov models, online learning, direct error minimization, discriminative training, automatic speech recognition, minimum phone error, minimum frame error 1. Joseph Keshet, Chih-Chieh Cheng, Mark Stoehr, David A. McAllester |
INTERSPEECH | 1 |
| 2011 | A GPU-tailored approach for training kernelized SVMsabstractWe present a method for efficiently training binary and multiclass kernelized SVMs on a Graphics Processing Unit (GPU). Our methods apply to a broad range of kernels, including the popular Gaus- sian kernel, on datasets as large as the amount of available memory on the graphics card. Our approach is distinguished from earlier work in that it cleanly and efficiently handles sparse datasets through the use of a novel clustering technique. Our optimization algorithm is also specifically designed to take advantage of the graphics hardware. This leads to different algorithmic choices then those preferred in serial implementations. Our easy-to-use library is orders of magnitude faster then existing CPU libraries, and several times faster than prior GPU approaches. Andrew Cotter, Nathan Srebro, Joseph Keshet |
KDD | 3 |
| 2011 | Generalization Bounds and Consistency for Latent Structural Probit and Ramp LossabstractWe consider latent structural versions of probit loss and ramp loss. We show that these surrogate loss functions are consistent in the strong sense that for any feature map (finite or infinite dimensional) they yield predictors approaching the infimum task loss achievable by any linear predictor over the given features. We also give finite sample generalization bounds (convergence rates) for these loss functions. These bounds suggest that probit loss converges more rapidly. However, ramp loss is more easily optimized and may ultimately be more practical. David A. McAllester, Joseph Keshet |
NIPS | 2 |
| 2010 | Automatic discriminative measurement of voice onset timeabstractCrammer, K., Dekel, O., Keshet, J., Shalev-Shwartz, S., and Singer, Y. (2006). Online passive-aggressivealgorithms. The Journal of Machine Learning Research, 7:551–585.Das, S. and Hansen, J. (2004). Detection of Voice Onset Time (VOT) for unvoiced stops (/p/,/t/,/k/) using the TeagerEnergy Operator (TEO) for automatic detection of accented English. In Proc. 6th NORSIG, pp. 344–347.Fischer, E. and Goberman, A. (2010). Voice onset time in Parkinson disease. J. Comm. Disorders, 43:21–34.Kazemzadeh, A., Tepperman, J., Silva, J., You, H., Lee, S., Alwan, A., and Narayanan, S. (2006). Automaticdetection of voice onset time contrasts for use in pronunciation assessment. In Proc. INTERSPEECH,pp. 721–724.Keshet, J., Shalev-Shwartz, S., Singer, Y., and Chazan, D. (2005). Phoneme alignment based on discriminativelearning. In Proc. INTERSPEECH, pp. 2961–2964.Keshet, J., Shalev-Shwartz, S., Singer, Y., and Chazan, D. (2007). A large margin algorithm for speech-to-phonemeand music-to-score alignment. IEEE Trans. Audio, Speech, Language Process., 15(8):2373–2382.Kuhl, P. and Miller, J. (1978). Speech perception by the chinchilla: Identification functions for synthetic VOT stimuli.J. Acoust. Soc. America, 63(3):905–917.Morris, R., McCrea, C., and Herring, K. (2008b). Voice onset time differences between adult males and females:Isolated syllables. J. Phonetics, 36(2):308–317.Stouten, V. and van Hamme, H. (2009). Automatic voice onset time estimation from reassignment spectra. SpeechCommunication, 51(12):1194–1205.Taskar, B., Guestrin, C., and Koller, D. (2004). Max-margin Markov networks. Advances in Neural InformationProcessing Systems, 16.Theodore, R., Miller, J., and DeSteno, D. (2009b). Individual talker differences in voice-onset-time: Contextualinfluences. J. Acoust. Soc. America, 125:3974–3982.Tsochantaridis, I., Hofmann, T., Joachims, T., and Altun, Y. (2004). Support vector machine learning forinterdependent and structured output spaces. In Proc. 21st ICML.Yao, Y. (2007). Closure duration and VOT of word-initial voiceless plosives in English in spontaneous connectedspeech. UC Berkeley Phonology Lab Annual Report, pp. 183–225. Morgan Sonderegger, Joseph Keshet |
INTERSPEECH | 2 |
| 2010 | Direct Loss Minimization for Structured PredictionabstractIn discriminative machine learning one is interested in training a system to optimize a certain desired measure of performance, or loss. In binary classification one typically tries to minimizes the error rate. But in structured prediction each task often has its own measure of performance such as the BLEU score in machine translation or the intersection-over-union score in PASCAL segmentation. The most common approaches to structured prediction, structural SVMs and CRFs, do not minimize the task loss: the former minimizes a surrogate loss with no guarantees for task loss and the latter minimizes log loss independent of task loss. The main contribution of this paper is a theorem stating that a certain perceptron-like learning rule, involving features vectors derived from loss-adjusted inference, directly corresponds to the gradient of task loss. We give empirical results on phonetic alignment of a standard test set from the TIMIT corpus, which surpasses all previously reported results on this problem. David A. McAllester, Tamir Hazan, Joseph Keshet |
NIPS | 3 |
| 2009 | Robust discriminative keyword spotting for emotionally colored spontaneous speech using bidirectional LSTM networksabstractIn this paper we propose a new technique for robust keyword spotting that uses bidirectional long short-term memory (BLSTM) recurrent neural nets to incorporate contextual information in speech decoding. Our approach overcomes the drawbacks of generative HMM modeling by applying a discriminative learning procedure that non-linearly maps speech features into an abstract vector space. By incorporating the outputs of a BLSTM network into the speech features, it is able to make use of past and future context for phoneme predictions. The robustness of the approach is evaluated on a keyword spotting task using the HUMAINE sensitive artificial listener (SAL) database, which contains accented, spontaneous, and emotionally colored speech. The test is particularly stringent because the system is not trained on the SAL database, but only on the TIMIT corpus of read speech. We show that our method prevails over a discriminative keyword spotter without BLSTM-enhanced feature functions, which in turn has been proven to outperform HMM-based techniques. Martin Wöllmer, Florian Eyben, Joseph Keshet, Alex Graves, Björn W. Schuller, Gerhard Rigoll |
ICASSP | 3 |
| 2009 | Bounded Kernel-Based Online Learning
Francesco Orabona, Joseph Keshet, Barbara Caputo |
J. Mach. Learn. Res. | 2 |
| 2009 | Discriminative keyword spotting
Joseph Keshet, David Grangier, Samy Bengio |
Speech Commun. | 1 |
| 2008 | The projectron: a bounded kernel-based PerceptronabstractWe present a discriminative online algorithm with a bounded memory growth, which is based on the kernel-based Perceptron. Generally, the required memory of the kernel-based Perceptron for storing the online hypothesis is not bounded. Previous work has been focused on discarding part of the instances in order to keep the memory bounded. In the proposed algorithm the instances are not discarded, but projected onto the space spanned by the previous online hypothesis. We derive a relative mistake bound and compare our algorithm both analytically and empirically to the state-of-the-art Forgetron algorithm (Dekel et al, 2007). The first variant of our algorithm, called Projectron, outperforms the Forgetron. The second variant, called Projectron++, outperforms even the Perceptron. Francesco Orabona, Joseph Keshet, Barbara Caputo |
ICML | 2 |
| 2008 | Support Vector Machines with a Reject OptionabstractWe consider the problem of binary classification where the classifier may abstain instead of classifying each observation. The Bayes decision rule for this setup, known as Chow's rule, is defined by two thresholds on posterior probabilities. From simple desiderata, namely the consistency and the sparsity of the classifier, we derive the double hinge loss function that focuses on estimating conditional probabilities only in the vicinity of the threshold points of the optimal decision rule. We show that, for suitable kernel machines, our approach is universally consistent. We cast the problem of minimizing the double hinge loss as a quadratic program akin to the standard SVM optimization problem and propose an active set method to solve it efficiently. We finally provide preliminary experimental results illustrating the interest of our constructive approach to devising loss functions. Yves Grandvalet, Alain Rakotomamonjy, Joseph Keshet, Stéphane Canu |
NIPS | 3 |
| 2007 | A Large Margin Algorithm for Speech-to-Phoneme and Music-to-Score AlignmentabstractWe describe and analyze a discriminative algorithm for learning to align an audio signal with a given sequence of events that tag the signal. We demonstrate the applicability of our method for the tasks of speech-to-phoneme alignment (ldquoforced alignmentrdquo) and music-to-score alignment. In the first alignment task, the events that tag the speech signal are phonemes while in the music alignment task, the events are musical notes. Our goal is to learn an alignment function whose input is an audio signal along with its accompanying event sequence and its output is a timing sequence representing the actual start time of each event in the audio signal. Generalizing the notion of separation with a margin used in support vector machines for binary classification, we cast the learning task as the problem of finding a vector in an abstract inner-product space. To do so, we devise a mapping of the input signal and the event sequence along with any possible timing sequence into an abstract vector space. Each possible timing sequence therefore corresponds to an instance vector and the predicted timing sequence is the one whose projection onto the learned prediction vector is maximal. We set the prediction vector to be the solution of a minimization problem with a large set of constraints. Each constraint enforces a gap between the projection of the correct target timing sequence and the projection of an alternative, incorrect, timing sequence onto the vector. Though the number of constraints is very large, we describe a simple iterative algorithm for efficiently learning the vector and analyze the formal properties of the resulting learning algorithm. We report experimental results comparing the proposed algorithm to previous studies on speech-to-phoneme and music-to-score alignment, which use hidden Markov models. The results obtained in our experiments using the discriminative alignment algorithm are comparable to results of state-of-the-art systems. Joseph Keshet, Shai Shalev-Shwartz, Yoram Singer, Dan Chazan |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Discriminative kernel-based phoneme sequence recognitionabstractAbstract. We describe a new method for phoneme sequence recognition given a speech utterance. In contrast to HMM-based approaches, our method uses a kernel-based discriminative training procedure in which the learning process is tailored to the goal of minimizing the Levenshtein distance between the predicted phoneme sequence and the correct sequence. The phoneme sequence predictor is devised by mapping the speech utterance along with a proposed phoneme sequence to a vector-space endowed with an inner-product that is realized by a Mercer kernel. Building on large margin techniques for predicting whole sequences, we are able to devise a learning algorithm which distills to separating the correct phoneme sequence from all other sequences. We describe an iterative algorithm for learning the phoneme sequence recognizer and further describe an efficient implementation of it. We present initial encouraging experimental results with the TIMIT and compare the proposed method to an HMM-based approach. 2 IDIAP–RR 06-14 1 Joseph Keshet, Shai Shalev-Shwartz, Samy Bengio, Yoram Singer, Dan Chazan |
INTERSPEECH | 1 |
| 2006 | Online Passive-Aggressive AlgorithmsabstractWe present a family of margin based online learning algorithms for various prediction tasks. In particular we derive and analyze algorithms for binary and multiclass categorization, regression, uniclass prediction and sequence prediction. The update steps of our different algorithms are all based on analytical solutions to simple constrained optimization problems. This unified view allows us to prove worst-case loss bounds for the different algorithms and for the various decision problems based on a single lemma. Our bounds on the cumulative loss of the algorithms are relative to the smallest loss that can be attained by any fixed hypothesis, and as such are applicable to both realizable and unrealizable settings. We demonstrate some of the merits of the proposed algorithms in a series of experiments with synthetic and real data sets. Koby Crammer, Ofer Dekel, Joseph Keshet, Shai Shalev-Shwartz, Yoram Singer |
J. Mach. Learn. Res. | 3 |
| 2005 | Phoneme alignment based on discriminative learningabstractWe propose a new paradigm for aligning a phoneme sequence of a speech utterance with its acoustical signal counterpart. In contrast to common HMM-based approaches, our method employs a discriminative learning procedure in which the learning phase is tightly coupled with the alignment task at hand. The alignment function we devise is based on mapping the input acousticsymbolic representations of the speech utterance along with the target alignment into an abstract vector space. We suggest a specific mapping into the abstract vector-space which utilizes standard speech features (e.g. spectral distances) as well as confidence outputs of a framewise phoneme classifier. Building on techniques used for large margin methods for predicting whole sequences, our alignment function distills to a classifier in the abstract vector-space which separates correct alignments from incorrect ones. We describe a simple iterative algorithm for learning the alignment function and discuss its formal properties. Experiments with the TIMIT corpus show that our method outperforms the current state-of-the-art approaches. Joseph Keshet, Shai Shalev-Shwartz, Yoram Singer, Dan Chazan |
INTERSPEECH | 1 |
| 2004 | Large margin hierarchical classificationabstractWe present an algorithmic framework for supervised classification learning where the set of labels is organized in a predefined hierarchical structure. This structure is encoded by a rooted tree which induces a metric over the label set. Our approach combines ideas from large margin kernel methods and Bayesian analysis. Following the large margin principle, we associate a prototype with each label in the tree and formulate the learning task as an optimization problem with varying margin constraints. In the spirit of Bayesian methods, we impose similarity requirements between the prototypes corresponding to adjacent labels in the hierarchy. We describe new online and batch algorithms for solving the constrained optimization problem. We derive a worst case loss-bound for the online algorithm and provide generalization analysis for its batch counterpart. We demonstrate the merits of our approach with a series of experiments on synthetic, text and speech data. Ofer Dekel, Joseph Keshet, Yoram Singer |
ICML | 2 |
| 2002 | Kernel Design Using BoostingabstractThe focus of the paper is the problem of learning kernel operators from empirical data. We cast the kernel design problem as the construction of an accurate kernel from simple (and less accurate) base kernels. We use the boosting paradigm to perform the kernel construction process. To do so, we modify the booster so as to accommodate kernel operators. We also devise an efficient weak-learner for simple kernels that is based on generalized eigen vector decomposition. We demonstrate the effective- ness of our approach on synthetic data and on the USPS dataset. On the USPS dataset, the performance of the Perceptron algorithm with learned kernels is systematically better than a fixed RBF kernel. 1 Introduction and problem Setting The last decade brought voluminous amount of work on the design, analysis and experi- mentation of kernel machines. Algorithm based on kernels can be used for various ma- chine learning tasks such as classification, regression, ranking, and principle component analysis. The most prominent learning algorithm that employs kernels is the Support Vec- tor Machines (SVM) [1, 2] designed for classification and regression. A key component in a kernel machine is a kernel operator which computes for any pair of instances their inner-product in some abstract vector space. Intuitively and informally, a kernel operator is a means for measuring similarity between instances. Almost all of the work that em- ployed kernel operators concentrated on various machine learning problems that involved a predefined kernel. A typical approach when using kernels is to choose a kernel before learning starts. Examples to popular predefined kernels are the Radial Basis Functions and the polynomial kernels (see for instance [1]). Despite the simplicity required in modifying a learning algorithm to a “kernelized” version, the success of such algorithms is not well understood yet. More recently, special efforts have been devoted to crafting kernels for specific tasks such as text categorization [3] and protein classification problems [4]. Our work attempts to give a computational alternative to predefined kernels by learning kernel operators from data. We start with a few definitions. Let X be an instance space. . An explicit way to describe K A kernel is an inner-product operator K : X (cid:2) X ! is via a mapping (cid:30) : X ! H from X to an inner-products space H such that K(x; x0) = (cid:30)(x)(cid:1)(cid:30)(x0). Given a kernel operator and a finite set of instances S = fxi; yigm i=1, the kernel matrix (a.k.a the Gram matrix) is the matrix of all possible inner-products of pairs from S, Ki;j = K(xi; xj). We therefore refer to the general form of K as the kernel operator and to the application of the kernel operator to a set of pairs of instances as the kernel matrix. The specific setting of kernel design we consider assumes that we have access to a base kernel learner and we are given a target kernel K ? manifested as a kernel ma- trix on a set of examples. Upon calling the base kernel learner it returns a kernel op- erator denote Kj. The goal thereafter is to find a weighted combination of kernels ^K(x; x0) = Pj (cid:11)jKj(x; x0) that is similar, in a sense that will be defined shortly, to the target kernel, ^K (cid:24) K ?. Cristianini et al. [5] in their pioneering work on kernel target alignment employed as the notion of similarity the inner-product between the kernel ma- trices < K; K 0 >F =Pm i;j=1 K(xi; xj)K 0(xi; xj). Given this definition, they defined the kernel-similarity, or alignment, to be the above inner-product normalized by the norm of each kernel, ^A(S; ^K; K ?) = (cid:16)< ^K; K ? >F(cid:17) =q< ^K; ^K >F < K ?; K ? >F ; where S is, as above, a finite sample of m instances. Put another way, the kernel alignment Cris- tianini et al. employed is the cosine of the angle between the kernel matrices where each matrix is “flattened” into a vector of dimension m2. Therefore, this definition implies that the alignment is bounded above by 1 and can attain this value iff the two kernel matrices are identical. Given a (column) vector of m labels y where yi 2 f(cid:0)1; +1g is the label of the instance xi, Cristianini et al. used the outer-product of y as the the target kernel, K ? = yyT . Therefore, an optimal alignment is achieved if ^K(xi; xj) = yiyj. Clearly, if such a kernel is used for classifying instances from X , then the kernel itself suffices to construct an excellent classifier f : X ! f(cid:0)1; +1g by setting, f (x) = sign(yiK(xi; x)) where (xi; yi) is any instance-label pair. Cristianini et al. then devised a procedure that works with both labelled and unlabelled examples to find a Gram matrix which attains a good alignment with K ? on the labelled part of the matrix. While this approach can clearly construct powerful kernels, a few problems arise from the notion of kernel alignment they employed. For instance, a kernel operator such that the sign(K(xi; xj)) is equal to yiyj but its magnitude, jK(xi; xj)j, is not necessarily 1, might achieve a poor alignment score while it can constitute a classifier whose empirical loss is zero. Furthermore, the task of finding a good kernel when it is not always possible to find a kernel whose sign on each pair of instances is equal to the products of the labels (termed the soft-margin case in [5, 6]) becomes rather tricky. We thus propose a different approach which attempts to overcome some of the difficulties above. Like Cristianini et al. we assume that we are given a set of labelled instances S = f(xi; yi) j xi 2 X ; yi 2 f(cid:0)1; +1g; i = 1; : : : ; mg : We are also given a set of unlabelled examples ~S = f~xig ~m i=1. If such a set is not provided we can simply use the labelled in- stances (without the labels themselves) as the set ~S. The set ~S is used for constructing the primitive kernels that are combined to constitute the learned kernel ^K. The labelled set is used to form the target kernel matrix and its instances are used for evaluating the learned kernel ^K. This approach, known as transductive learning, was suggested in [5, 6] for kernel alignment tasks when the distribution of the instances in the test data is different from that of the training data. This setting becomes in particular handy in datasets where the test data was collected in a different scheme than the training data. We next discuss the notion of kernel goodness employed in this paper. This notion builds on the objective function that several variants of boosting algorithms maintain [7, 8]. We therefore first discuss in brief the form of boosting algorithms for kernels. 2 Using Boosting to Combine Kernels Numerous interpretations of AdaBoost and its variants cast the boosting process as a pro- cedure that attempts to minimize, or make small, a continuous bound on the classification error (see for instance [9, 7] and the references therein). A recent work by Collins et al. [8] unifies the boosting process for two popular loss functions, the exponential-loss (denoted henceforth as ExpLoss) and logarithmic-loss (denoted as LogLoss) that bound the empir- Input: Labelled and unlabelled sets of examples: S = f(xi; yi)gm Initialize: K 0 (all zeros matrix) For t = 1; 2; : : : ; T : i=1 Koby Crammer, Joseph Keshet, Yoram Singer |
NIPS | 2 |
| 2001 | Plosive spotting with margin classifiersabstractThis paper presents a novel algorithm for precise spotting of plosives. The algorithm is based on a pattern matching technique implemented with margin classifiers, such as support vector machines (SVM). A special hierarchical treatment to overcome the problem of fricative and false silence detection is presented. It uses the loss-based multi-class decisions. Furthermore, a method for smoothing the overall decisions by sequential linear programming is described. The proposed algorithm was tested on the TIMIT corpus, which produced a very high spotting accuracy. The algorithm presented here is applied to plosives detection, but can easily be adapted to any class of phonemes. 1. Joseph Keshet, Dan Chazan, Ben-Zion Bobrovsky |
INTERSPEECH | 1 |