EDBT 2026 Demo / reviewers in the wild / expert
Ron Hoory
dblp:50/2590
· DBLP profile ↗
47ranked-venue papers
1as first author
12since 2021 · last 2025
0009-0006-1327-5160ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 27 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Granite-speech: open-source speech-aware LLMs with strong English ASR capabilitiesabstractGranite-speech LLMs are compact and efficient speech language models specifically designed for English ASR1and automatic speech translation (AST). The models were trained by modality aligning granite-3.3-instruct to speech on publicly available open-source corpora. Comprehensive benchmarking on English ASR shows that they outperform several competitors’ models that were trained on orders of magnitude more proprietary data, and they keep pace on English-to-X AST for major European languages, Japanese, and Mandarin. The speech-specific components are: a conformer acoustic encoder using block attention and self-conditioning trained with connectionist temporal classification, a windowed query-transformer speech modality adapter used to do temporal downsampling of the acoustic embeddings and map them to the LLM text embedding space, and LoRA adapters to further fine-tune the text LLM. The models are freely available on HuggingFace2under a permissive Apache 2.0 license.1The latest models (revision 3.3.2) support multilingual ASR in English, French, German, Spanish and Portuguese and bidirectional speech translation to and from English. This paper covers the initial English-only release.2https://huggingface.co/ibm-granite/granite-speech-3.3-2b (and…-8b). George Saon, Avihu Dekel, Alexi Brooks, Tohru Nagano, Abraham Daniels, Aharon Satt, Ashish R. Mittal, Brian Kingsbury, David Haws, Edmilson da Silva Morais, Gakuto Kurata, Hagai Aronowitz, Ibrahim Ibrahim, Hong-Kwang Jeff Kuo, Kate Soule, Luis A. Lastras, Masayuki Suzuki, Ron Hoory, Samuel Thomas 0001, Sashi Novitasari, Takashi Fukuda, Vishal Sunder, Zvi Kons |
ASRU | 18 |
| 2025 | Speech Synthesis From Continuous Features Using Per-Token Latent DiffusionabstractWe present SALAD, a zero-shot text-to-speech (TTS) autoregressive model operating over continuous speech representations. SALAD utilizes a per-token diffusion process to refine and predict continuous representations for the next time step. We compare our approach against a discrete variant of SALAD as well as publicly available zero-shot TTS systems, and conduct a comprehensive analysis of discrete versus continuous modeling techniques. Our results show that SALAD achieves superior intelligibility while matching the speech quality and speaker similarity of ground-truth audio. Arnon Turetzky, Avihu Dekel, Nimrod Shabtay, Slava Shechtman, David Haws, Hagai Aronowitz, Ron Hoory, Yossi Adi |
ASRU | 7 |
| 2025 | Exploring the Limits of Conformer CTC-Encoder for Speech Emotion Recognition using Large Language Models
Edmilson da Silva Morais, Hagai Aronowitz, Aharon Satt, Ron Hoory, Avihu Dekel, Brian Kingsbury, George Saon |
INTERSPEECH | 4 |
| 2025 | Spoken Question Answering for Visual Queries
Nimrod Shabtay, Zvi Kons, Avihu Dekel, Hagai Aronowitz, Ron Hoory, Assaf Arbelle |
INTERSPEECH | 5 |
| 2024 | Speak While You Think: Streaming Speech Synthesis During Text GenerationabstractLarge Language Models (LLMs) demonstrate impressive capabilities, yet interaction with these models is mostly facilitated through text. Using Text-To-Speech to synthesize LLM outputs typically results in notable latency, which is impractical for fluent voice conversations. We propose LLM2Speech, an architecture to synthesize speech while text is being generated by an LLM which yields significant latency reduction. LLM2Speech mimics the predictions of a non-streaming teacher model while limiting the exposure to future context in order to enable streaming. It exploits the hidden embeddings of the LLM, a by-product of the text generation that contains informative semantic context. Experimental results show that LLM2Speech maintains the teacher’s quality while reducing the latency to enable natural conversations. Avihu Dekel, Slava Shechtman, Raul Fernandez, David Haws, Zvi Kons, Ron Hoory |
ICASSP | 6 |
| 2024 | Creating an African American-Sounding TTS: Guidelines, Technical Challenges, and Surprising EvaluationsabstractRepresentations of AI agents in user interfaces and robotics are predominantly White, not only in terms of facial and skin features, but also in the synthetic voices they use. In this paper we explore some unexpected challenges in the representation of race we found in the process of developing an U.S. English Text-to-Speech (TTS) system aimed to sound like an educated, professional, regional accent-free African American woman. The paper starts by presenting the results of focus groups with African American IT professionals where guidelines and challenges for the creation of a representative and appropriate TTS system were discussed and gathered, followed by a discussion about some of the technical difficulties faced by the TTS system developers. We then describe two studies with U.S. English speakers where the participants were not able to attribute the correct race to the African American TTS voice while overwhelmingly correctly recognizing the race of a White TTS system of similar quality. A focus group with African American IT workers not only confirmed the representativeness of the African American voice we built, but also suggested that the surprising recognition results may have been caused by the inability or the latent prejudice from non-African Americans to associate educated, non-vernacular, professionally-sounding voices to African American people. Claudio S. Pinhanez, Raul Fernandez, Marcelo Grave, Julio Nogima, Ron Hoory |
IUI | 5 |
| 2023 | Modeling Turn-Taking in Human-To-Human Spoken Dialogue Datasets Using Self-Supervised FeaturesabstractSelf-supervised pre-trained models have consistently delivered state-of-art results in the fields of natural language and speech processing. However, we argue that their merits for modeling Turn-Taking for spoken dialogue systems still need further investigation. Due to that, in this paper we intro-duce a modular End-to-End system based on an Upstream + Downstream architecture paradigm, which allows easy use/integration of a large variety of self-supervised features to model the specific Turn-Taking task of End-of-Turn Detection (EOTD). Several architectures to model the EOTD task using audio-only, text-only and audio+text modalities are presented, and their performance and robustness are carefully evaluated for three different human-to-human spoken dialogue datasets. The proposed model not only achieves SOTA results for EOTD, but also brings light to the possibility of powerful and well fine-tuned self-supervised models to be successfully used for a wide variety Turn-Taking tasks. Edmilson da Silva Morais, Matheus Damasceno, Hagai Aronowitz, Aharon Satt, Ron Hoory |
ICASSP | 5 |
| 2022 | Towards A Common Speech Analysis EngineabstractRecent innovations in self-supervised representation learning have led to remarkable advances in natural language processing. That said, in the speech processing domain, self-supervised representation learning-based systems are not yet considered state-of-the-art.We propose leveraging recent advances in self-supervised-based speech processing to create a common speech analysis engine. Such an engine should be able to handle multiple speech processing tasks, using a single architecture, to obtain state-of-the-art accuracy. The engine must also enable support for new tasks with small training datasets. Beyond that, a common engine should be capable of supporting distributed training with client in-house private data.We present the architecture for a common speech analysis engine based on the HuBERT self-supervised speech representation. Based on experiments, we report our results for language identification and emotion recognition on the standard evaluations NIST-LRE 07 and IEMOCAP. Our results surpass the state-of-the-art performance reported so far on these tasks.We also analyzed our engine on the emotion recognition task using reduced amounts of training data and show how to achieve improved results. Hagai Aronowitz, Itai Gat, Edmilson da Silva Morais, Weizhong Zhu, Ron Hoory |
ICASSP | 5 |
| 2022 | Speaker Normalization for Self-Supervised Speech Emotion RecognitionabstractLarge speech emotion recognition datasets are hard to obtain, and small datasets may contain biases. Deep-net-based classifiers, in turn, are prone to exploit those biases and find shortcuts such as speaker characteristics. These shortcuts usually harm a model’s ability to generalize. To address this challenge, we propose a gradient-based adversary learning framework that learns a speech emotion recognition task while normalizing speaker characteristics from the feature representation. We demonstrate the efficacy of our method on both speaker-independent and speaker-dependent settings and obtain new state-of-the-art results on the challenging IEMOCAP dataset. Itai Gat, Hagai Aronowitz, Weizhong Zhu, Edmilson da Silva Morais, Ron Hoory |
ICASSP | 5 |
| 2022 | A New Data Augmentation Method for Intent Classification Enhancement and its Application on Spoken Conversation Datasets
Zvi Kons, Aharon Satt, Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Boaz Carmeli, Ron Hoory, Brian Kingsbury |
ICASSP | 6 |
| 2022 | Speech Emotion Recognition Using Self-Supervised FeaturesabstractSelf-supervised pre-trained features have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in the field of speech emotion recognition (SER) still need further investigation. In this paper we introduce a modular End-to-End (E2E) SER system based on an Upstream + Downstream architecture paradigm, which allows easy use/integration of a large variety of self-supervised features. Several SER experiments for predicting categorical emotion classes from the IEMOCAP dataset are performed. These experiments investigate interactions among fine-tuning of self-supervised feature models, aggregation of frame-level features into utterance-level features and back-end classification networks. The proposed monomodal speech-only based system not only achieves SOTA results, but also brings light to the possibility of powerful and well fine-tuned self-supervised acoustic features that reach results similar to the results achieved by SOTA multimodal systems using both Speech and Text modalities. Edmilson da Silva Morais, Ron Hoory, Weizhong Zhu, Itai Gat, Matheus Damasceno, Hagai Aronowitz |
ICASSP | 2 |
| 2021 | RNN Transducer Models for Spoken Language UnderstandingabstractWe present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available annotations are SLU labels and their values, and a more restrictive case where transcripts are available but not corresponding audio. We show how RNN-T SLU models can be developed starting from pre-trained automatic speech recognition (ASR) systems, followed by an SLU adaptation step. In settings where real audio data is not available, artificially synthesized speech is used to successfully adapt various SLU models. When evaluated on two SLU data sets, the ATIS corpus and a customer call center data set, the proposed models closely track the performance of other E2E models and achieve state-of-the-art results. Samuel Thomas 0001, Hong-Kwang Jeff Kuo, George Saon, Zoltán Tüske, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory |
ICASSP | 8 |
| 2020 | Leveraging Unpaired Text Data for Training End-To-End Speech-to-Intent SystemsabstractTraining an end-to-end (E2E) neural network speech-to-intent (S2I) system that directly extracts intents from speech requires large amounts of intent-labeled speech data, which is time consuming and expensive to collect. Initializing the S2I model with an ASR model trained on copious speech data can alleviate data sparsity. In this paper, we attempt to leverage NLU text resources. We implemented a CTC-based S2I system that matches the performance of a state-of-the-art, traditional cascaded SLU system. We performed controlled experiments with varying amounts of speech and text training data. When only a tenth of the original data is available, intent classification accuracy degrades by 7.6% absolute. Assuming we have additional text-to-intent data (without speech) available, we investigated two techniques to improve the S2I system: (1) transfer learning, in which acoustic embeddings for intent classification are tied to fine-tuned BERT text embeddings; and (2) data augmentation, in which the text-to-intent data is converted into speech-to-intent data using a multi-speaker text-to-speech system. The proposed approaches recover 80% of performance lost due to using limited intent-labeled speech. Hong-Kwang Jeff Kuo, Samuel Thomas 0001, Zvi Kons, Kartik Audhkhasi, Brian Kingsbury, Ron Hoory, Michael Picheny |
ICASSP | 7 |
| 2020 | New Advances in Speaker Diarization
Hagai Aronowitz, Weizhong Zhu, Masayuki Suzuki, Gakuto Kurata, Ron Hoory |
INTERSPEECH | 5 |
| 2020 | End-to-End Spoken Language Understanding Without Full TranscriptsabstractAn essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develop end-to-end (E2E) spoken language understanding systems that directly convert speech input to semantic entities and investigate if these E2E SLU models can be trained solely on semantic entity annotations without word-for-word transcripts. Training such models is very useful as they can drastically reduce the cost of data collection. We created two types of such speech-to-entities models, a CTC model and an attention-based encoder-decoder model, by adapting models trained originally for speech recognition. Given that our experiments involve speech input, these systems need to recognize both the entity label and words representing the entity value correctly. For our speech-to-entities experiments on the ATIS corpus, both the CTC and attention models showed impressive ability to skip non-entity words: there was little degradation when trained on just entities versus full transcripts. We also explored the scenario where the entities are in an order not necessarily related to spoken order in the utterance. With its ability to do re-ordering, the attention model did remarkably well, achieving only about 2% degradation in speech-to-bag-of-entities F1 score. Hong-Kwang Jeff Kuo, Zoltán Tüske, Samuel Thomas 0001, Kartik Audhkhasi, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory, Luis A. Lastras |
INTERSPEECH | 9 |
| 2020 | Siamese X-Vector Reconstruction for Domain Adapted Speaker RecognitionabstractWith the rise of voice-activated applications, the need for speaker recognition is rapidly increasing. The x-vector, an embedding approach based on a deep neural network (DNN), is considered the state-of-the-art when proper end-to-end training is not feasible. However, the accuracy significantly decreases when recording conditions (noise, sample rate, etc.) are mismatched, either between the x-vector training data and the target data or between enrollment and test data. We introduce the Siamese x-vector Reconstruction (SVR) for domain adaptation. We reconstruct the embedding of a higher quality signal from a lower quality counterpart using a lean auxiliary Siamese DNN. We evaluate our method on several mismatch scenarios and demonstrate significant improvement over the baseline. Shai Rozenberg, Hagai Aronowitz, Ron Hoory |
INTERSPEECH | 3 |
| 2020 | Principal Style Components: Expressive Style Control and Cross-Speaker Transfer in Neural TTS
Alexander Sorin, Slava Shechtman, Ron Hoory |
INTERSPEECH | 3 |
| 2019 | High Quality, Lightweight and Adaptable TTS Using LPCNetabstractWe present a lightweight adaptable neural TTS system with high quality output. The system is composed of three separate neural network blocks: prosody prediction, acoustic feature prediction and Linear Prediction Coding Net as a neural vocoder. This system can synthesize speech with close to natural quality while running 3 times faster than real-time on a standard CPU. The modular setup of the system allows for simple adaptation to new voices with a small amount of data. We first demonstrate the ability of the system to produce high quality speech when trained on large, high quality datasets. Following that, we demonstrate its adaptability by mimicking unseen voices using 5 to 20 minutes long datasets with lower recording quality. Large scale Mean Opinion Score quality and similarity tests are presented, showing that the system can adapt to unseen voices with quality gap of 0.12 and similarity gap of 3% compared to natural speech for male voices and quality gap of 0.35 and similarity of gap of 9 % for female voices. Zvi Kons, Slava Shechtman, Alexander Sorin, Carmel Rabinovitz, Ron Hoory |
INTERSPEECH | 5 |
| 2018 | Word Emphasis Prediction for Expressive Text to Speech
Yosi Mass, Slava Shechtman, Moran Mordechay, Ron Hoory, Oren Sar Shalom, Guy Lev, David Konopnicki |
INTERSPEECH | 4 |
| 2018 | The IBM Virtual Voice Creator
Alexander Sorin, Slava Shechtman, Zvi Kons, Ron Hoory, Shay Ben-David, Joe Pavitt, Shai Rozenberg, Carmel Rabinovitz, Tal Drory |
INTERSPEECH | 4 |
| 2018 | Neural TTS Voice ConversionabstractRecently, speaker adaptation of neural TTS models received significant interest, and several studies focusing on this topic have been published. All of them explore an adaptation of an initial multi-speaker model trained on a corpus containing from tens to hundreds of individual speaker voices.In this work we focus on a challenging task of TTS voice conversion where an initial system is trained on a single-speaker data and then need to be adapted to a variety of external speaker voices. The TTS voice conversion setup represents a very important use case. Transcribed multi-speaker datasets might be unavailable for many languages while any TTS technology provider is expected to have at least one suitable single-speaker dataset per supported language.We present a neural TTS system comprising separate prosody generator and synthesizer DNN models. The system is trained on a high quality proprietary male speaker dataset. We show that the system models can be converted to a variety of external male and female ordinary voices and an extremely expressive artist's voice and present crowd-base subjective evaluation results. Zvi Kons, Slava Shechtman, Alexander Sorin, Ron Hoory, Carmel Rabinovitz, Edmilson da Silva Morais |
SLT | 4 |
| 2017 | Voice-transformation-based data augmentation for prosodic classificationabstractIn this work we explore data-augmentation techniques for the task of improving the performance of a supervised recurrent-neural-network classifier tasked with predicting prosodic-boundary and pitch-accent labels. The technique is based on applying voice transformations to the training data that modify the pitch baseline and range, as well as the vocal-tract and vocal-source characteristics of the speakers to generate further training examples. We demonstrate the validity of the approach by improving performance when the amount of base labeled examples is small (showing reductions in the range of 7%–12% for reduced-data conditions) as well as in terms of its generalization to speakers unseen in the training set (showing a relative reduction in the error rate of 8.74% and 4.75%, on the average, for boundaries and accent tasks respectively, in leave-one-speaker-out validation). Raul Fernandez, Andrew Rosenberg, Alexander Sorin, Bhuvana Ramabhadran, Ron Hoory |
ICASSP | 5 |
| 2017 | Weakly-Supervised Phrase Assignment from Text in a Speech-Synthesis System Using Noisy Labels
Asaf Rendel, Raul Fernandez, Zvi Kons, Andrew Rosenberg, Ron Hoory, Bhuvana Ramabhadran |
INTERSPEECH | 5 |
| 2017 | Efficient Emotion Recognition from Speech Using Deep Learning on Spectrograms
Aharon Satt, Shai Rozenberg, Ron Hoory |
INTERSPEECH | 3 |
| 2016 | Using continuous lexical embeddings to improve symbolic-prosody prediction in a text-to-speech front-endabstractThe prediction of symbolic prosodic categories from text is an important, but challenging, natural-language processing task given the various ways in which an input can be realized, and the fact that knowledge about what features determine this realization is incomplete or inaccessible to the model. In this work, we look at augmenting baseline features with lexical representations that are derived from text, providing continuous embeddings of the lexicon in a lower-dimensional space. Although learned in an unsupervised fashion, such features capture semantic and syntactic properties that make them amenable for prosody prediction. We deploy various embedding models on prominence- and phrase-break prediction tasks, showing substantial gains, particularly for prominence prediction. Asaf Rendel, Raul Fernandez, Ron Hoory, Bhuvana Ramabhadran |
ICASSP | 3 |
| 2015 | Using deep bidirectional recurrent neural networks for prosodic-target prediction in a unit-selection text-to-speech system
Raul Fernandez, Asaf Rendel, Bhuvana Ramabhadran, Ron Hoory |
INTERSPEECH | 4 |
| 2014 | Multi-modal biometrics for mobile authenticationabstractUser authentication in the context of a secure transaction needs to be continuously evaluated for the risks associated with the transaction authorization. The situation becomes even more critical when there are regulatory compliance requirements. Need for such systems have grown dramatically with the introduction of smart mobile devices which make it far easier for the user to complete such transaction quickly but with a huge exposure to risk. Biometrics can play a very significant role in addressing such problems as a key indicator of the user identity and thus reducing the risk of fraud. While unimodal biometrics authentication systems are being increasingly experimented by mainstream mobile system manufacturers (e.g., fingerprint in iOS), we explore various opportunities of reducing risk in a multimodal biometrics system. The multimodal system is based on fusion of several biometrics combined with a policy manager. A new biometric modality: chirography which is based on user writing on multi-touch screens using their finger is introduced. Coupling with chirography, we also use two other biometrics: face and voice. Our fusion strategy is based on inter-modality score level fusion that takes into account a voice quality measure. The proposed system has been evaluated on an in-house database that reflects the latest smart mobile devices. On this database, we demonstrate a very high accuracy multi-modal authentication system reaching an EER of 0.1% in an office environment and an EER of 0.5% in challenging noisy environments. Hagai Aronowitz, Orith Toledo-Ronen, Sivan Harary, Amir B. Geva, Shay Ben-David, Asaf Rendel, Ron Hoory, Nalini K. Ratha, Sharath Pankanti, David Nahamoo |
IJCB | 8 |
| 2014 | Exploring modulation spectrum features for speech-based depression level classification
Elif Bozkurt, Orith Toledo-Ronen, Alexander Sorin, Ron Hoory |
INTERSPEECH | 4 |
| 2014 | Prosody contour prediction with long short-term memory, bi-directional, deep recurrent neural networksabstractDeep Neural Networks (DNNs) have been shown to provide state-of-the-art performance over other baseline models in the task of predicting prosodic targets from text in a speechsynthesis system. However, prosody prediction can be affected by an interaction of short- and long-term contextual factors that a static model that depends on a fixed-size context window can fail to properly capture. In this work, we look at a recurrent formulation of neural networks (RNNs) that are deep in time and can store state information from an arbitrarily large input history when making a prediction. We show that RNNs provide improved performance over DNNs of comparable size in terms of various objective metrics for a variety of prosodic streams (notably, a relative reduction of about 6% in F0 mean-square error accompanied by a relative increase of about 14% in F0 variance), as well as in terms of perceptual quality assessed through mean-opinion-score listening tests. Index Terms: speech synthesis,text-to-speech, prosody prediction, recurrent neural networks, deep learning Raul Fernandez, Asaf Rendel, Bhuvana Ramabhadran, Ron Hoory |
INTERSPEECH | 4 |
| 2014 | Speech-based automatic and robust detection of very early dementiaabstractWe provide evidence to the potential use of simple spoken tasks for automatic assessment of very early dementia. Timely detection of dementia is required for effective psychological treatment and to enable patients to participate in new drug therapy research. The technology enables automatic, cheap, remote and wide-scale screening of dementia, typically a costly and complex procedure. It can aid clinicians in the diagnosis of very early dementia, as well as assessing the disease progression. We describe the spoken tasks, and their respective languageindependent vocal feature extraction, followed by classification accuracy evaluation. We use recordings from over 60 persons, diagnosed as healthy-control (CTRL) / mildcognitive-impairment (MCI) / early-stage-Alzheimer-disease and early-mixed-dementia (AD). We present a new data regularization technique to overcome data sparseness due to the limited data set size. Next, we present a comprehensive statistical analysis, showing that the suggested classifier generalizes, and revealing the role and the statistical importance of the different spoken tasks and their respective vocal features. We demonstrate classification accuracy of about 80% for CTRL vs. MCI and MCI vs. AD, and 87% for CTRL vs. AD, all shown to generalize. This provides an evidence for potential use for automatic detection of very early dementia. Aharon Satt, Ron Hoory, Alexandra König, Pauline Aalten, Philippe H. Robert |
INTERSPEECH | 2 |
| 2013 | F0 contour prediction with a deep belief network-Gaussian process hybrid modelabstractIn this work we look at using non-parametric, exemplar-based regression for the prediction of prosodic contour targets from textual features in a speech synthesis system. We investigate the performance of Gaussian Process regression on this task when the covariance kernel operates on a variety of input feature spaces. In particular, we consider non-linear features extracted via Deep Belief Networks. We motivate the use of this hybrid model by considering the initial deep-layer model as a feature extractor that can summarize high-level structure from the raw inputs to improve the regression of an exemplar-based model in the second part of the approach. By looking at both objective metrics and perceptual listening tests, we evaluate these proposals against each other, and against the standard clustering-tree techniques implemented in parametric synthesis for the prediction of prosodic targets. Raul Fernandez, Asaf Rendel, Bhuvana Ramabhadran, Ron Hoory |
ICASSP | 4 |
| 2012 | Towards automatic phonetic segmentation for TTSabstractPhonetic segmentation is an important step in the development of a concatenative TTS voice. This paper introduces a segmentation process consisting of two phases. First, forced alignment is performed using an HMM-GMM model. The resulting segmentation is then locally refined using an SVM based boundary model. Both the models are derived from multi-speaker data using a speaker adaptive training procedure. Evaluation results are obtained on the TIMIT corpus and on a proprietary single-speaker TTS corpus. Asaf Rendel, Alexander Sorin, Ron Hoory, Andrew P. Breen |
ICASSP | 3 |
| 2011 | Speech processing and retrieval in a personal memory aid system for the elderlyabstractThe paper presents a new application of automatic speech processing in the Ambient Assisted Living area, developed in the course of a three year research project. Recording and automatic processing of spoken conversations plays a major role in this solution enabling effective search in a personal audio archive and fast browsing of conversations. Processing of elderly conversational speech recorded by a distant PDA microphone poses a great challenge. The speech processing flow includes transcription, speaker tracking and combined indexing and search of spoken terms and participating speakers identity extracted from the audio. We present the entire application and individual speech processing components as well as evaluation results of the individual components and of the end-to-end spoken information retrieval solution. Alexander Sorin, Hagai Aronowitz, Jonathan Mamou, Orith Toledo-Ronen, Ron Hoory, Michael Kuritzky, Yael Erez, Bhuvana Ramabhadran, Abhinav Sethy |
ICASSP | 5 |
| 2011 | New Developments in Voice Biometrics for User Authentication
Hagai Aronowitz, Ron Hoory, Jason W. Pelecanos, David Nahamoo |
INTERSPEECH | 2 |
| 2011 | Improved Spoken Query Transcription Using Co-Occurrence InformationabstractSpoken queries are a natural medium for searching the Mobile Web. Language modeling for voice search recognition offers different challenges compared to more conventional speech applications. The challenges arise from the fact that spoken queries are usually a set of keywords and do not have a syntactic and grammatical structure. This paper describes a cooccurrence based approach to improve the accuracy of voice queries automatic transcription. With the right choice of scoring function and co-occurrence level, we show that co-occurrence information gives a 2% relative accuracy improvement over a state of the art system. Jonathan Mamou, Abhinav Sethy, Bhuvana Ramabhadran, Ron Hoory, Paul Vozila |
INTERSPEECH | 4 |
| 2011 | Towards Goat Detection in Text-Dependent Speaker Verification
Orith Toledo-Ronen, Hagai Aronowitz, Ron Hoory, Jason W. Pelecanos, David Nahamoo |
INTERSPEECH | 3 |
| 2006 | High Quality Sinusoidal Modeling of Wideband Speech for the Purposes of Speech Synthesis and ModificationabstractThis paper describes an efficient sinusoidal modeling framework for high quality wide band (WB) speech synthesis and modification. This technique may serve as a basis for speech compression in the context of small footprint concatenative Text to Speech systems. In addition, it is a useful representation for voice transformation and morphing purposes, e.g., simultaneous pitch modification and spectral envelope warping. The conventional sinusoidal modeling is enhanced with an adaptive frequency dithering mechanism, based on a degree of voicing analysis. Considerable reduction of the amount of model parameters is achieved by high band phase extension. The proposed model is evaluated and compared to the alternative STRAIGHT framework [1]. Being simpler and considerably more efficient than STRAIGHT, it outperforms it in speech quality for both speech reconstruction and transformation. Dan Chazan, Ron Hoory, Ariel Sagi, Slava Shechtman, Alexander Sorin, Zhiwei Shuang, Raimo Bakis |
ICASSP (1) | 2 |
| 2006 | Spoken document retrieval from call-center conversationsabstractWe are interested in retrieving information from conversational speech corpora, such as call-center data. This data comprises spontaneous speech conversations with low recording quality, which makes automatic speech recognition (ASR) a highly difficult task. For typical call-center data, even state-of-the-art large vocabulary continuous speech recognition systems produce a transcript with word error rate of 30% or higher. In addition to the output transcript, advanced systems provide word confusion networks (WCNs), a compact representation of word lattices associating each word hypothesis with its posterior probability. Our work exploits the information provided by WCNs in order to improve retrieval performance. In this paper, we show that the mean average precision (MAP) is improved using WCNs compared to the raw word transcripts. Finally, we analyze the effect of increasing ASR word error rate on search effectiveness. We show that MAP is still reasonable even under extremely high error rate. Jonathan Mamou, David Carmel, Ron Hoory |
SIGIR | 3 |
| 2005 | Automatic analysis of call-center conversationsabstractWe describe a system for automating call-center analysis and monitoring. Our system integrates transcription of incoming calls with analysis of their content; for the analysis, we introduce a novel method of estimating the domain-specific importance of conversation fragments, based on divergence of corpus statistics. Combining this method with Information Retrieval approaches, we provide knowledge-mining tools both for the call-center agents and for administrators of the center. Gilad Mishne, David Carmel, Ron Hoory, Alexey Roytman, Aya Soffer |
CIKM | 3 |
| 2005 | Small footprint concatenative text-to-speech synthesis system using complex spectral envelope modelingabstractIn this paper we present a method for speech modeling and its utilization in IBM’s small footprint concatenative text-to-speech system. The method is based on frequency-domain, complex spectral envelope modeling, where the phase component plays a crucial role in attaining high quality speech synthesis. The modeling scheme presented enables low bit rate compression of the amplitude and phase information and low-complexity reconstruction of high quality speech with wide range pitch modification. Listening tests conducted for the overall text-to-speech system show a major improvement in MOS, compared to a previous, MFCC-based, system. 1. Dan Chazan, Ron Hoory, Zvi Kons, Ariel Sagi, Slava Shechtman, Alexander Sorin |
INTERSPEECH | 2 |
| 2004 | The ETSI extended distributed speech recognition (DSR) standards: server-side speech reconstructionabstractIn this paper we present work that has been carried out in developing the ETSI Extended DSR standards ES 202 211 and ES 202 212. These standards extend the previous ETSI DSR standards: basic front-end ES 201 108 and advanced (noise robust) front-end ES 202 050 respectively. The extensions enable enhanced tonal language recognition as well as server-side speech reconstruction capability. This paper discusses the server-side speech reconstruction whereas a companion paper discusses the front-end extension and tonal language recognition. Experimental results show that the reconstructed speech produced by the standards is highly intelligible under clean and noisy background conditions with the DRT (diagnostic rhyme test) and TT (transcription test) scores meeting or exceeding the objective values corresponding to the USA DoD (Department of Defence) federal standard MELP (mixed-excitation linear predictive) coder operating at 2400 bit/s. Tenkasi Ramabadran, Alexander Sorin, Michael J. McLaughlin, Dan Chazan, David Pearce 0002, Ron Hoory |
ICASSP (1) | 6 |
| 2004 | The ETSI extended distributed speech recognition (DSR) standards: client side processing and tonal language recognition evaluationabstractWe present work that has been carried out in developing the ETSI extended DSR standards ES 202 211 and ES 202 212 (2003). These standards extend the previous ETSI DSR standards: basic front-end ES 201 108 and advanced (noise robust) front-end ES 202 050 respectively. The extensions enable enhanced tonal language recognition as well as server-side speech reconstruction capability. The paper discusses the client-side estimation of pitch and voicing class parameters whereas a companion paper discusses the server-side speech reconstruction. Experimental results show enhancement of tonal language recognition rates of proprietary recognition engines, when the standard extensions are used. Alexander Sorin, Tenkasi Ramabadran, Dan Chazan, Ron Hoory, Michael J. McLaughlin, David Pearce 0002 |
ICASSP (1) | 4 |
| 2002 | Reducing the footprint of the IBM trainable speech synthesis systemabstractThis paper presents a novel approach for concatenative speech synthesis. This approach enables reduction of the dataset size of a concatenative text-to-speech system, namely the IBM trainable speech synthesis system, by more than an order of magnitude. A spectral acoustic feature based speech representation is used for computing a cost function during segment selection as well as for speech generation. Initial results indicate that even with a dataset size of a few megabytes it is possible to achieve quality which is significantly higher than existing small footprint formant based synthesizers. Dan Chazan, Ron Hoory, Zvi Kons, Dorel Silberstein, Alexander Sorin |
INTERSPEECH | 2 |
| 2001 | Efficient periodicity extraction based on sine-wave representation and its application to pitch determination of speech signalsabstractThis paper presents a novel low-complexity method for extracting periodicity of signals based on their sine-wave representation. In this representation, the signal is modeled as a finite sum of sine-waves, with time-varying amplitudes, phases and frequencies. We describe how one can modify the familiar spectral-comb analysis method to obtain a guaranteed and effective procedure to find the fundamental-frequency which gives the best harmonic approximation of the signal spectrum. The search is efficiently carried out in the frequency domain. The procedure obtains a successive refinement of possible pitch values which are consistent with an increasing number of sine wave components. Other pitch intervals are pruned at an early stage of the search. The advantage of this algorithm is its high accuracy achieved at a relatively low complexity. We also briefly describe one possible application in the area of pitch determination of speech signals. Dan Chazan, Meir Tzur, Ron Hoory, Gilad Cohen |
INTERSPEECH | 3 |
| 2000 | Speech reconstruction from mel frequency cepstral coefficients and pitch frequencyabstractThis paper presents a novel low complexity, frequency domain algorithm for reconstruction of speech from the mel-frequency cepstral coefficients (MFCC), commonly used by speech recognition systems, and the pitch frequency values. The reconstruction technique is based on the sinusoidal speech representation. A set of sine-wave frequencies is derived using the pitch frequency and voicing decisions, and synthetic phases are then assigned to each respective sine wave. The sine-wave amplitudes are generated by sampling a linear combination of frequency domain basis functions. The basis function gains are determined such that the mel-frequency binned spectrum of the reconstructed speech is similar to the mel-frequency binned spectrum, obtained from the original MFCC vector by IDCT and antilog operations. Natural sounding, good quality intelligible speech is obtained by this procedure. Dan Chazan, Ron Hoory, Gilad Cohen, Meir Tzur |
ICASSP | 2 |
| 2000 | Conversational networking: conversational protocols for transport, coding, and control
Stéphane H. Maes, Dan Chazan, Gilad Cohen, Ron Hoory |
INTERSPEECH | 4 |
| 1994 | Speech synthesis for a specific speaker based on a labeled speech databaseabstractThis paper proposes a new text-to-speech synthesis technique, for producing continuous, natural sounding speech of a specific speaker. The synthesis technique is based on selecting short speech frames from a phoneme-labeled speech database. The selection procedure involves minimization of a distortion criterion, by a dynamic programming algorithm. The proposed scheme is more flexible than many existing schemes using fixed speech segments, such as diphones. It results in a more natural synthesized speech. An efficient speech representation is used to express simply and accurately the spectral continuity of speech. A further improvement in the database search mechanism and in database size was obtained by sectioning the speech phonemes into "steady-states" and "transitions". The resulting synthesized speech quality, is satisfactory and preserves the natural voice of the speaker. Ron Hoory, Dan Chazan |
ICPR (3) | 1 |