EDBT 2026 Demo / reviewers in the wild / expert
Hagai Aronowitz
dblp:81/6320
· DBLP profile ↗
53ranked-venue papers
28as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 27 first-author · 10 since 2021Artificial intelligence and machine learning · 34 · 18 first-author · 5 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Granite-speech: open-source speech-aware LLMs with strong English ASR capabilitiesabstractGranite-speech LLMs are compact and efficient speech language models specifically designed for English ASR1and automatic speech translation (AST). The models were trained by modality aligning granite-3.3-instruct to speech on publicly available open-source corpora. Comprehensive benchmarking on English ASR shows that they outperform several competitors’ models that were trained on orders of magnitude more proprietary data, and they keep pace on English-to-X AST for major European languages, Japanese, and Mandarin. The speech-specific components are: a conformer acoustic encoder using block attention and self-conditioning trained with connectionist temporal classification, a windowed query-transformer speech modality adapter used to do temporal downsampling of the acoustic embeddings and map them to the LLM text embedding space, and LoRA adapters to further fine-tune the text LLM. The models are freely available on HuggingFace2under a permissive Apache 2.0 license.1The latest models (revision 3.3.2) support multilingual ASR in English, French, German, Spanish and Portuguese and bidirectional speech translation to and from English. This paper covers the initial English-only release.2https://huggingface.co/ibm-granite/granite-speech-3.3-2b (and…-8b). George Saon, Avihu Dekel, Alexi Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Jeff Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas 0001, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Zvi Kons |
ASRU | 12 |
| 2025 | Speech Synthesis From Continuous Features Using Per-Token Latent DiffusionabstractWe present SALAD, a zero-shot text-to-speech (TTS) autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuous representations for the next time step. We compare our approach against a discrete variant of SALAD as well as publicly available zero-shot TTS systems, and conduct a comprehensive analysis of discrete versus continuous modeling techniques. Our results show that SALAD achieves superior intelligibility while matching the speech quality and speaker similarity of ground-truth audio. Arnon Turetzky, Avihu Dekel, Nimrod Shabtay, Slava Shechtman, David Haws, Hagai Aronowitz, Ron Hoory, Yossi Adi |
ASRU | 6 |
| 2025 | A Non-autoregressive Model for Joint STT and TTSabstractIn this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimodal framework capable of handling the speech and text modalities as input either individually or together. The proposed model can also be trained with unpaired speech or text data owing to its multimodal nature. We further propose an iterative refinement strategy to improve the STT and TTS performance of our model such that the partial hypothesis at the output can be fed back to the input of our model, thus iteratively improving both STT and TTS predictions. We show that our joint model can effectively perform both STT and TTS tasks, outperforming the STT-specific baseline in all tasks and performing competitively with the TTS-specific baseline across a wide range of evaluation metrics. Vishal Sunder, Brian Kingsbury, George Saon, Samuel Thomas 0001, Slava Shechtman, Hagai Aronowitz, Eric Fosler-Lussier, Luis A. Lastras |
ICASSP | 6 |
| 2025 | Exploring the Limits of Conformer CTC-Encoder for Speech Emotion Recognition using Large Language Models
Edmilson da Silva Morais, Hagai Aronowitz, Aharon Satt, Ron Hoory, Avihu Dekel, Brian Kingsbury, George Saon |
INTERSPEECH | 2 |
| 2025 | Spoken Question Answering for Visual Queries
Nimrod Shabtay, Zvi Kons, Avihu Dekel, Hagai Aronowitz, Ron Hoory, Assaf Arbelle |
INTERSPEECH | 4 |
| 2023 | Modeling Turn-Taking in Human-To-Human Spoken Dialogue Datasets Using Self-Supervised FeaturesabstractSelf-supervised pre-trained models have consistently delivered state-of-art results in the fields of natural language and speech processing. However, we argue that their merits for modeling Turn-Taking for spoken dialogue systems still need further investigation. Due to that, in this paper we intro-duce a modular End-to-End system based on an Upstream + Downstream architecture paradigm, which allows easy use/integration of a large variety of self-supervised features to model the specific Turn-Taking task of End-of-Turn Detection (EOTD). Several architectures to model the EOTD task using audio-only, text-only and audio+text modalities are presented, and their performance and robustness are carefully evaluated for three different human-to-human spoken dialogue datasets. The proposed model not only achieves SOTA results for EOTD, but also brings light to the possibility of powerful and well fine-tuned self-supervised models to be successfully used for a wide variety Turn-Taking tasks. Edmilson da Silva Morais, Matheus Damasceno, Hagai Aronowitz, Aharon Satt, Ron Hoory |
ICASSP | 3 |
| 2022 | Towards A Common Speech Analysis EngineabstractRecent innovations in self-supervised representation learning have led to remarkable advances in natural language processing. That said, in the speech processing domain, self-supervised representation learning-based systems are not yet considered state-of-the-art.We propose leveraging recent advances in self-supervised-based speech processing to create a common speech analysis engine. Such an engine should be able to handle multiple speech processing tasks, using a single architecture, to obtain state-of-the-art accuracy. The engine must also enable support for new tasks with small training datasets. Beyond that, a common engine should be capable of supporting distributed training with client in-house private data.We present the architecture for a common speech analysis engine based on the HuBERT self-supervised speech representation. Based on experiments, we report our results for language identification and emotion recognition on the standard evaluations NIST-LRE 07 and IEMOCAP. Our results surpass the state-of-the-art performance reported so far on these tasks.We also analyzed our engine on the emotion recognition task using reduced amounts of training data and show how to achieve improved results. Hagai Aronowitz, Itai Gat, Edmilson da Silva Morais, Weizhong Zhu, Ron Hoory |
ICASSP | 1 |
| 2022 | Speaker Normalization for Self-Supervised Speech Emotion RecognitionabstractLarge speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristics. These shortcuts usually harm a model’s ability to generalize. To address this challenge, we propose a gradient-based adversary learning framework that learns a speech emotion recognition task while normalizing speaker characteristics from the feature representation. We demonstrate the efficacy of our method on both speaker-independent and speaker-dependent settings and obtain new state-of-the-art results on the challenging IEMOCAP dataset. Itai Gat, Hagai Aronowitz, Weizhong Zhu, Edmilson da Silva Morais, Ron Hoory |
ICASSP | 2 |
| 2022 | Speech Emotion Recognition Using Self-Supervised FeaturesabstractSelf-supervised pre-trained features have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in the field of speech emotion recognition (SER) still need further investigation. In this paper we introduce a modular End-to-End (E2E) SER system based on an Upstream + Downstream architecture paradigm, which allows easy use/integration of a large variety of self-supervised features. Several SER experiments for predicting categorical emotion classes from the IEMOCAP dataset are performed. These experiments investigate interactions among fine-tuning of self-supervised feature models, aggregation of frame-level features into utterance-level features and back-end classification networks. The proposed monomodal speech-only based system not only achieves SOTA results, but also brings light to the possibility of powerful and well fine-tuned self-supervised acoustic features that reach results similar to the results achieved by SOTA multimodal systems using both Speech and Text modalities. Edmilson da Silva Morais, Ron Hoory, Weizhong Zhu, Itai Gat, Matheus Damasceno, Hagai Aronowitz |
ICASSP | 6 |
| 2022 | Extending RNN-T-based speech recognition systems with emotion and language classificationabstractSpeech transcription, emotion recognition, and language identification are usually considered to be three different tasks.Each one requires a different model with a different architecture and training process.We propose using a recurrent neural network transducer (RNN-T)-based speech-to-text (STT) system as a common component that can be used for emotion recognition and language identification as well as for speech recognition.Our work extends the STT system for emotion classification through minimal changes, and shows successful results on the IEMOCAP and MELD datasets.In addition, we demonstrate that by adding a lightweight component to the RNN-T module, it can also be used for language identification.In our evaluations, this new classifier demonstrates state-of-the-art accuracy for the NIST-LRE-07 dataset. Zvi Kons, Hagai Aronowitz, Edmilson da Silva Morais, Matheus Damasceno, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, George Saon |
INTERSPEECH | 2 |
| 2020 | Context and Uncertainty Modeling for Online Speaker Change DetectionabstractSpeaker change detection is often addressed as a key component in speaker diarization systems. In this work we focus on online speaker change detection as a standalone task which is required for online closed captioning of broadcast television. Contrary to related works, we do not operate on frame-level features such as MFCC. Instead, we leverage state-of-the-art speaker recognition-based technology by modeling sequences of pretrained speaker embeddings (x-vectors) using a deep neural network. We explicitly address two types of uncertainties. The first one is uncertainty in embedding point estimate which is due to short and varying segment duration. The second type is uncertainty in which context segments are relevant to representing the speaker talking right before the hypothesized speaker change. We also show the robustness of affinity matrix-representation for speaker change detection. Our methods provide very significant accuracy improvements compared to several baselines including a recently published end-to-end system. Hagai Aronowitz, Weizhong Zhu |
ICASSP | 1 |
| 2020 | New Advances in Speaker Diarization
Hagai Aronowitz, Weizhong Zhu, Masayuki Suzuki, Gakuto Kurata, Ron Hoory |
INTERSPEECH | 1 |
| 2020 | Siamese X-Vector Reconstruction for Domain Adapted Speaker RecognitionabstractWith the rise of voice-activated applications, the need for speaker recognition is rapidly increasing. The x-vector, an embedding approach based on a deep neural network (DNN), is considered the state-of-the-art when proper end-to-end training is not feasible. However, the accuracy significantly decreases when recording conditions (noise, sample rate, etc.) are mismatched, either between the x-vector training data and the target data or between enrollment and test data. We introduce the Siamese x-vector Reconstruction (SVR) for domain adaptation. We reconstruct the embedding of a higher quality signal from a lower quality counterpart using a lean auxiliary Siamese DNN. We evaluate our method on several mismatch scenarios and demonstrate significant improvement over the baseline. Shai Rozenberg, Hagai Aronowitz, Ron Hoory |
INTERSPEECH | 2 |
| 2018 | Robust Audiovisual Liveness Detection for Biometric Authentication Using Deep Joint Embedding and Dynamic Time WarpingabstractWe address the problem of liveness detection in audiovisual recordings for preventing spoofing attacks in biometric authentication systems. We assume that liveness is detected from a recording of a speaker saying a predefined phrase and that another recording of the same phrase is a priori available, a setting, which is common in text-dependent authentication systems. We propose to measure liveness by comparing between alignments of audio and video to the a priori recorded sequence using dynamic time warping. The alignments are computed in a joint feature space to which audio and video are embedded using deep convolutional neural networks. We investigate the robustness of the proposed algorithm across datasets by training and testing it on different datasets. Experimental results demonstrate that the proposed algorithm generalizes well across datasets providing improved performance compared to competing methods. Amit Aides, David Dov, Hagai Aronowitz |
ICASSP | 3 |
| 2017 | Inter dataset variability modeling for speaker recognitionabstractWe introduce a novel approach of addressing inter-dataset variability in the context of speaker recognition in a mismatched condition under the JHU-2013 domain adaptation challenge (DAC) framework. Previously, we took a subspace removal approach for inter-dataset variability compensation (IDVC) of within speaker variability. In this work we substitute subspace removal with incorporation of the variability into the Probabilistic Linear Discriminant Analysis (PLDA) model. We do that by introducing a novel optimality criterion which is minimizing the expected square error in estimation of the log-likelihood ratio of target trials when dataset-dependent PLDA models are replaced by a dataset independent PLDA model. The result we obtain is a correction term for the commonly estimated within speaker variability matrix. The correction term represents the normalized inter-dataset variability of the within speaker variability matrices. The proposed method outperforms the extended IDVC method on the DAC. Hagai Aronowitz |
ICASSP | 1 |
| 2017 | Speaker recognition using common passphrases in RedDotsabstractIn this paper we report our work on the recently collected text dependent speaker recognition dataset named RedDots, with a focus on the common passphrase condition. We first investigate an out-of-the-box approach. We then report several strategies to train on RedDots itself using up to 40 speakers for training. The GMM-NAP framework is used as a baseline. We report the following novelties: First, we demonstrate the use of bagging for improved accuracy. Second, we estimate the EER of a passphrase using metadata only. Third, the estimated EERs are used for improved score normalization. Finally we report an analysis of system sensitivity to the duration between enrollment and testing (template aging). Hagai Aronowitz |
ICASSP | 1 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 17 |
| 2016 | Speaker recognition using matched filtersabstractNowadays state-of-the-art speaker recognition systems obtain quite accurate results for both text-independent and text-dependent tasks as long as they are trained on a fair amount of development data from the target domain, and as long as the target data is clean. In this work we investigate the use of matched filters for speaker recognition in the framework of a small in-domain development data. We show how a matched filter can be optimized to maximize SNR (signal to noise ratio) when the noise component includes both intra-speaker variability and center/mean hyper-parameter variability. The proposed method generalizes our previous method named score stabilization and obtains significant speaker recognition error reductions. Hagai Aronowitz |
ICASSP | 1 |
| 2016 | Audio enhancing with DNN autoencoder for speaker recognitionabstractIn this paper we present a design of a DNN-based autoencoder for speech enhancement and its use for speaker recognition systems for distant microphones and noisy data. We started with augmenting the Fisher database with artificially noised and reverberated data and trained the autoencoder to map noisy and reverberated speech to its clean version. We use the autoencoder as a preprocessing step in the later stage of modelling in state-of-the-art text-dependent and text-independent speaker recognition systems. We report relative improvements up to 50% for the text-dependent system and up to 48% for the text-independent one. With text-independent system, we present a more detailed analysis on various conditions of NIST SRE 2010 and PRISM suggesting that the proposed preprocessig is a promising and efficient way to build a robust speaker recognition system for distant microphone and noisy data. Oldrich Plchot, Lukás Burget, Hagai Aronowitz, Pavel Matejka |
ICASSP | 3 |
| 2016 | Text-Dependent Audiovisual Synchrony Detection for Spoofing Detection in Mobile Person Recognition
Amit Aides, Hagai Aronowitz |
INTERSPEECH | 2 |
| 2015 | Exploiting supervector structure for speaker recognition trained on a small development set
Hagai Aronowitz |
INTERSPEECH | 1 |
| 2015 | Score stabilization for speaker recognition trained on a small development set
Hagai Aronowitz |
INTERSPEECH | 1 |
| 2015 | The reddots data collection for speaker recognitionabstractde niveau recherche, publiés ou non, émanant des établissements d'enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Kong-Aik Lee, Anthony Larcher, Guangsen Wang, Patrick Kenny, Niko Brümmer, David A. van Leeuwen, Hagai Aronowitz, Marcel Kockmann, Carlos Vaquero, Bin Ma 0001, Haizhou Li 0001, Themos Stafylakis, Jahangir Alam 0001, Albert Swart, Javier Perez |
INTERSPEECH | 7 |
| 2014 | Inter dataset variability compensation for speaker recognitionabstractRecently satisfactory results have been obtained in NIST speaker recognition evaluations. These results are mainly due to accurate modeling of a very large development dataset provided by LDC. However, for many realistic scenarios the use of this development dataset is limited due to a dataset mismatch. In such cases, collection of a large enough dataset is infeasible. In this work we analyze the sources of degradation for a particular setup in the context of an i-vector PLDA system and conclude that the main source for degradation is an i-vector dataset shift. As a remedy, we introduce inter dataset variability compensation (IDVC) to explicitly compensate for dataset shift in the i-vector space. This is done using the nuisance attribute projection (NAP) method. Using IDVC we managed to reduce error dramatically by more than 50% for the domain mismatch setup. Hagai Aronowitz |
ICASSP | 1 |
| 2014 | Multi-modal biometrics for mobile authenticationabstractUser authentication in the context of a secure transaction needs to be continuously evaluated for the risks associated with the transaction authorization. The situation becomes even more critical when there are regulatory compliance requirements. Need for such systems have grown dramatically with the introduction of smart mobile devices which make it far easier for the user to complete such transaction quickly but with a huge exposure to risk. Biometrics can play a very significant role in addressing such problems as a key indicator of the user identity and thus reducing the risk of fraud. While unimodal biometrics authentication systems are being increasingly experimented by mainstream mobile system manufacturers (e.g., fingerprint in iOS), we explore various opportunities of reducing risk in a multimodal biometrics system. The multimodal system is based on fusion of several biometrics combined with a policy manager. A new biometric modality: chirography which is based on user writing on multi-touch screens using their finger is introduced. Coupling with chirography, we also use two other biometrics: face and voice. Our fusion strategy is based on inter-modality score level fusion that takes into account a voice quality measure. The proposed system has been evaluated on an in-house database that reflects the latest smart mobile devices. On this database, we demonstrate a very high accuracy multi-modal authentication system reaching an EER of 0.1% in an office environment and an EER of 0.5% in challenging noisy environments. Hagai Aronowitz, Orith Toledo-Ronen, Sivan Harary, Amir B. Geva, Shay Ben-David, Asaf Rendel, Ron Hoory, Nalini K. Ratha, Sharath Pankanti, David Nahamoo |
IJCB | 1 |
| 2014 | Domain adaptation for text dependent speaker verificationabstractRecently we have investigated the use of state-of-the-art textdependent speaker verification algorithms for user authentication and obtained satisfactory results mainly by using a fair amount of text-dependent development data from the target domain. In this work we investigate the ability to build high accuracy text-dependent systems using no data at all from the target domain. Instead of using target domain data, we use resources such as TIMIT, Switchboard, and NIST data. We introduce several techniques addressing both lexical mismatch and channel mismatch. These techniques include synthesizing a universal background model according to lexical content, automatic filtering of irrelevant phonetic content, exploiting information in residual supervectors (usually discarded in the i-vector framework), and inter dataset variability modeling. These techniques reduce verification error significantly, and also improve accuracy when target domain data is available. Hagai Aronowitz, Asaf Rendel |
INTERSPEECH | 1 |
| 2013 | Diffusion maps for PLDA-based speaker verificationabstractDuring the last few years, i-vectors have become an important component in most state-of-the-art speaker recognition systems. I-vector extraction is based on an assumption that GMM supervectors reside on a low dimensional space, which is modeled using Factor Analysis. In this paper we replace the above assumption with an assumption that the GMM supervectors reside on a low dimensional manifold and propose to use Diffusion Maps to learn that manifold. The learnt manifold implies a mapping of spoken sessions into a modified i-vector space which we call d-vector space. D-vectors can further be processed using standard techniques such as LDA, WCCN, cosine distance scoring or Probabilistic Linear Discriminant Analysis (PLDA). We demonstrate the usefulness of our approach on the telephone core conditions of NIST 2010, and obtain significant error reduction. Oren Barkan, Hagai Aronowitz |
ICASSP | 2 |
| 2013 | Fast High Dimensional Vector Multiplication Face RecognitionabstractThis paper advances descriptor-based face recognition by suggesting a novel usage of descriptors to form an over-complete representation, and by proposing a new metric learning pipeline within the same/not-same framework. First, the Over-Complete Local Binary Patterns (OCLBP) face representation scheme is introduced as a multi-scale modified version of the Local Binary Patterns (LBP) scheme. Second, we propose an efficient matrix-vector multiplication-based recognition system. The system is based on Linear Discriminant Analysis (LDA) coupled with Within Class Covariance Normalization (WCCN). This is further extended to the unsupervised case by proposing an unsupervised variant of WCCN. Lastly, we introduce Diffusion Maps (DM) for non-linear dimensionality reduction as an alternative to the Whitened Principal Component Analysis (WPCA) method which is often used in face recognition. We evaluate the proposed framework on the LFW face recognition dataset under the restricted, unrestricted and unsupervised protocols. In all three cases we achieve very competitive results. Oren Barkan, Jonathan Weill, Lior Wolf, Hagai Aronowitz |
ICCV | 4 |
| 2013 | On leveraging conversational data for building a text dependent speaker verification system
Hagai Aronowitz, Oren Barkan |
INTERSPEECH | 1 |
| 2013 | Voice transformation-based spoofing of text-dependent speaker verification systems
Zvi Kons, Hagai Aronowitz |
INTERSPEECH | 2 |
| 2012 | Efficient approximated i-vector extractionabstractI-vectors are currently widely used by state-of-the-art speech processing systems for tasks such as speaker verification and language identification. A shortcoming of i-vector-based systems is that the i-vector extraction process is computationally expensive. In this paper we propose an efficient method to extract i-vectors approximately. The method normalizes the GMM counts to be similar across sessions. We validate our method empirically for the speaker verification task on five different datasets, both text independent and text dependent. A significant speedup was obtained with a very small degradation in accuracy compared to the standard exact method. Hagai Aronowitz, Oren Barkan |
ICASSP | 1 |
| 2012 | Confidence for Speaker Diarization using PCA Spectral Ratio
Orith Toledo-Ronen, Hagai Aronowitz |
INTERSPEECH | 2 |
| 2011 | Speech processing and retrieval in a personal memory aid system for the elderlyabstractThe paper presents a new application of automatic speech processing in the Ambient Assisted Living area, developed in the course of a three year research project. Recording and automatic processing of spoken conversations plays a major role in this solution enabling effective search in a personal audio archive and fast browsing of conversations. Processing of elderly conversational speech recorded by a distant PDA microphone poses a great challenge. The speech processing flow includes transcription, speaker tracking and combined indexing and search of spoken terms and participating speakers identity extracted from the audio. We present the entire application and individual speech processing components as well as evaluation results of the individual components and of the end-to-end spoken information retrieval solution. Alexander Sorin, Hagai Aronowitz, Jonathan Mamou, Orith Toledo-Ronen, Ron Hoory, Michael Kuritzky, Yael Erez, Bhuvana Ramabhadran, Abhinav Sethy |
ICASSP | 2 |
| 2011 | Speaker Diarization Using a priori Acoustic Information
Hagai Aronowitz |
INTERSPEECH | 1 |
| 2011 | New Developments in Joint Factor Analysis for Speaker Verification
Hagai Aronowitz, Oren Barkan |
INTERSPEECH | 1 |
| 2011 | New Developments in Voice Biometrics for User Authentication
Hagai Aronowitz, Ron Hoory, Jason W. Pelecanos, David Nahamoo |
INTERSPEECH | 1 |
| 2011 | Implicit Segmentation in Two-Wire Speaker Recognition
Yosef A. Solewicz, Hagai Aronowitz |
INTERSPEECH | 2 |
| 2011 | Towards Goat Detection in Text-Dependent Speaker Verification
Orith Toledo-Ronen, Hagai Aronowitz, Ron Hoory, Jason W. Pelecanos, David Nahamoo |
INTERSPEECH | 2 |
| 2010 | Efficient score normalization for speaker recognitionabstractScore normalization is an important component in most speech classification tasks including speaker recognition. State-of-the-art scoring approaches use both T-norm and Z-norm. This paper addresses the following goals: better understanding of existing score normalization methods, reducing the need for explicit score normalization, and improving the computational efficiency of score normalization. In addition, the importance of score normalization for speaker identification is demonstrated, and accuracy is improved considerably using various normalization techniques. Hagai Aronowitz, Vanessia Aronowitz |
ICASSP | 1 |
| 2009 | Two-wire nuisance attribute projection
Yosef A. Solewicz, Hagai Aronowitz |
INTERSPEECH | 2 |
| 2008 | Recent advances in the IBM GALE Mandarin transcription systemabstractThis paper describes the system and algorithmic developments in the automatic transcription of Mandarin broadcast speech made at IBM in the second year of the DARPA GALE program. Technical advances over our previous system include improved acoustic models using embedded tone modeling, and a new topic-adaptive language model (LM) rescoring technique based on dynamically generated LMs. We present results on three community-defined test sets designed to cover both the broadcast news and the broadcast conversation domain. It is shown that our new baseline system attains a 15.4% relative reduction in character error rate compared with our previous GALE evaluation system. And a further 13.6% improvement over the baseline is achieved with the two described techniques. Selina M. Chu, Hong-Kwang Jeff Kuo, Lidia Mangu, Yi Y. Liu 0002, Yong Qin 0001, Qin Shi 0001, Shilei Zhang, Hagai Aronowitz |
ICASSP | 8 |
| 2008 | Online vocabulary adaptation using contextual information and information retrievalabstractThis paper presents an algorithm for automatic online vocabulary adaptation based on contextual information and information retrieval. Experiments are presented on a transcription task of spoken annotations of business cards recorded by a hand-held device. Contextual information is used to trigger web search which is used to adapt the vocabulary for a given business card. Finally, the language model for the adapted vocabulary is modified by taking into account the relative value of each context information source. On the business card task, the proposed algorithm reduces 75% of the out-of-vocabulary rate and 16% of the word error rate. Hagai Aronowitz |
INTERSPEECH | 1 |
| 2008 | Speaker recognition in two-wire test sessionsabstractThis paper deals with the task of speaker recognition in fourwire training and two-wire testing conditions. Instead of performing blind speaker diarization before the recognition stage, we directly perform the recognition on the nonsegmented (or imperfectly diarized) speech. We present an analysis of the problem with respect to three different speaker recognition systems and propose improved recognition techniques both in the frame domain and in the model domain. The proposed techniques reduce error rate significantly. Furthermore, the developed techniques may be also beneficial in conjunction with an imperfect blind diarization stage. Index Terms: speaker recognition, two-wire, summed channel Hagai Aronowitz, Yosef A. Solewicz |
INTERSPEECH | 1 |
| 2007 | Segmental Modeling for Audio SegmentationabstractTrainable speech/non-speech segmentation and music detection algorithms usually consist of a frame based scoring phase combined with a smoothing phase. This paper suggests a framework in which both phases are explicitly unified in a segment based classifier. We suggest a novel segment based generative model in which audio segments are modeled as supervectors and each class (speech, silence, music) is modeled by a distribution over the supervector space. Segmental speech classes can then be modeled by generative models such as GMMs or can be classified by SVMs. Our suggested framework leads to a significant reduction in error rate. Hagai Aronowitz |
ICASSP (4) | 1 |
| 2007 | Speaker recognition using kernel-PCA and intersession variability modeling
Hagai Aronowitz |
INTERSPEECH | 1 |
| 2007 | Trainable speaker diarization
Hagai Aronowitz |
INTERSPEECH | 1 |
| 2007 | Efficient Speaker Recognition Using Approximated Cross Entropy (ACE)abstractTechniques for efficient speaker recognition are presented. These techniques are based on approximating Gaussian mixture modeling (GMM) likelihood scoring using approximated cross entropy (ACE). Gaussian mixture modeling is used for representing both training and test sessions and is shown to perform speaker recognition and retrieval extremely efficiently without any notable degradation in accuracy compared to classic GMM-based recognition. In addition, a GMM compression algorithm is presented. This algorithm decreases considerably the storage needed for speaker retrieval. Hagai Aronowitz, David Burshtein |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | A Session-GMM Generative Model Using Test Utterance Gaussian Mixture Modeling for Speaker VerificationabstractTest utterance parameterization (TUP) using Gaussian mixture models (GMMs) has recently been shown to be beneficial for speaker indexing due to its computational efficiency and identical accuracy compared to classic GMM-based recognizers. We show that TUP can also lead to more accurate speaker recognition. On the NIST-2004 evaluation corpus, recognition error rate was reduced by 8% compared to the classic GMM-based algorithm. Furthermore, we introduce a novel generative statistical model for generation of test utterances by speakers. This model is incorporated naturally into the TUP framework and improves speaker recognition accuracy. On the NIST-2004 evaluation corpus, recognition error rate was reduced by 15% compared to the classic GMM-based algorithm. Hagai Aronowitz, David Burshtein, Amihood Amir |
ICASSP (1) | 1 |
| 2005 | Efficient speaker identification and retrievalabstractIn this paper we present techniques for efficient speaker recognition of a large population of speakers and for efficient speaker retrieval in large audio archives. We deal with aspects of both time and storage. We use Gaussian mixture modeling (GMM) for representing both train and test sessions and show how to perform speaker recognition and retrieval efficiently with only a small degradation in accuracy compared to classic GMM based recognition. We present techniques for achieving a dramatic acceleration of both tasks. Finally, we present a GMM compression algorithm that decreases considerably the storage needed for speaker retrieval. 1. Hagai Aronowitz, David Burshtein |
INTERSPEECH | 1 |
| 2005 | Modeling intra-speaker variability for speaker recognitionabstractIn this paper we present a speaker recognition algorithm that models explicitly intra-speaker inter-session variability. Such variability may be caused by changing speaker characteristics (mood, fatigue, etc.), channel variability or noise variability. We define a session-space in which each session (either train or test session) is a vector. We then calculate a rotation of the session-space for which the estimated intra-speaker subspace is isolated and can be modeled explicitly. We evaluated our technique on the NIST-2004 speaker recognition evaluation corpus, and compared it to a GMM baseline system. Results indicate significant reduction in error rate. 1. Hagai Aronowitz, Dror Irony, David Burshtein |
INTERSPEECH | 1 |
| 2005 | A distance measure between GMMs based on the unscented transform and its application to speaker recognitionabstractThis paper proposes a dissimilarity measure between two Gaussian mixture models (GMM). Computing a distance measure between two GMMs that were learned from speech segments is a key element in speaker verification, speaker segmentation and many other related applications. A natural measure between two distributions is the Kullback-Leibler divergence. However, it cannot be analytically computed in the case of GMM. We propose an accurate and efficiently computed approximation of the KL-divergence. The method is based on the unscented transform which is usually used to obtain a better alternative to the extended Kalman filter. The suggested distance is evaluated in an experimental setup of speakers data-set. The experimental results indicate that our proposed approximations outperform previously suggested methods. 1. Jacob Goldberger, Hagai Aronowitz |
INTERSPEECH | 2 |
| 2004 | Speaker indexing in audio archives using test utterance Gaussian mixture modelingabstractSpeaker Indexing has recently emerged as an important task due to the rapidly growing volume of audio archives. Current filtration techniques still suffer from problems both in accuracy and efficiency. The major reason for the drawbacks of existing solutions is the use of inaccurate anchor models. The contribution of this paper is two-fold. On the theoretical side, a new method is developed for simulating GMM scoring. This enables to fit a GMM not only to every target speaker but also to every test utterance, and then compute the likelihood of the test call using these GMMs instead of using the original data. The second contribution of this paper is in harnessing this GMM simulation to achieve very efficient speaker indexing in terms of both search time and index size. Results on the SPIDRE corpus show that our approach maintains the accuracy of the conventional GMM algorithm. 1. Hagai Aronowitz, David Burshtein, Amihood Amir |
INTERSPEECH | 1 |
| 2004 | Text independent speaker recognition using speaker dependent word spottingabstractThis paper is motivated by the fact that text dependent speaker recognition is inherently more accurate than text independent speaker recognition. In this work we assign models to frequent words spoken by a speaker and spot them in a test call. In this way, text-dependent speaker recognition technology can be used for text independent tasks. The approach we take is to use DTW (Dynamic Time Warp) word spotting to find words in the test that resemble words in the train set. Results on the SPIDRE corpus show that using a combined DTW spotter based system and a GMM system improves performance significantly. For very low false acceptance rate (0.1%) misdetection was reduced from 32.2% to 23.3 % (28 % reduction). For low false acceptance rate (1%) misdetection was reduced from 28.9 % to 21.1 % (27% reduction). 1. Hagai Aronowitz, David Burshtein, Amihood Amir |
INTERSPEECH | 1 |