Evelina Bakhturina

dblp:211/6749 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
11since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021
YearPublicationVenuePosition
2025 HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset
Ryan Langman, Xuesong Yang, Paarth Neekhara, Shehzeen Hussain, Edresson Casanova, Evelina Bakhturina, Jason Li 0007
INTERSPEECH6
2024 A Chat about Boring Problems: Studying GPT-Based Text Normalization
abstract
Text normalization - the conversion of text from written to spoken form - is traditionally assumed to be an ill-formed task for language modeling. In this work, we argue otherwise. We empirically show the capacity of Large-Language Models (LLM) for text normalization in few-shot scenarios. Combining self-consistency reasoning with linguistic-informed prompt engineering, we find LLM-based text normalization to achieve error rates approximately 40% lower than production-level normalization systems. Further, upon error analysis, we note key limitations in the conventional design of text normalization tasks. We create a new taxonomy of text normalization errors and apply it to results from GPT-3.5-Turbo and GPT-4.0. Through this new framework, we identify strengths and weaknesses of LLM-based TN, opening opportunities for future work.
Yang Zhang 0089, Travis M. Bartley, Mariana Graterol-Fuenmayor, Vitaly Lavrukhin, Evelina Bakhturina, Boris Ginsburg
ICASSP5
2024 Retrieval meets Long Context Large Language Models
abstract
Extending the context window of large language models (LLMs) is getting popular recently, while the solution of augmenting LLMs with retrieval has existed for years. The natural questions are: i) Retrieval-augmentation versus long context window, which one is better for downstream tasks? ii) Can both methods be combined to get the best of both worlds? In this work, we answer these questions by studying both solutions using two state-of-the-art pretrained LLMs, i.e., a proprietary 43B GPT and Llama2-70B. Perhaps surprisingly, we find that LLM with 4K context window using simple retrieval-augmentation at generation can achieve comparable performance to finetuned LLM with 16K context window via positional interpolation on long context tasks, while taking much less computation. More importantly, we demonstrate that retrieval can significantly improve the performance of LLMs regardless of their extended context window sizes. Our best model, retrieval-augmented Llama2-70B with 32K context window, outperforms GPT-3.5-turbo-16k and Davinci003 in terms of average score on nine long context tasks including question answering, query-based summarization, and in-context few-shot learning tasks. It also outperforms its non-retrieval Llama2-70B-32k baseline by a margin, while being much faster at generation. Our study provides general insights on the choice of retrieval-augmentation versus long context extension of LLM for practitioners.
Peng Xu 0008, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu 0001, Zihan Liu 0001, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, Bryan Catanzaro
ICLR8
2023 LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of End-to-End ASR Models
abstract
Traditional automatic speech recognition (ASR) models output lower-cased words without punctuation marks, which reduces readability and necessitates a subsequent text processing model to convert ASR transcripts into a proper format. Simultaneously, the development of end-to-end ASR models capable of predicting punctuation and capitalization presents several challenges, primarily due to limited data availability and shortcomings in the existing evaluation methods, such as inadequate assessment of punctuation prediction. In this paper, we introduce a LibriSpeech-PC benchmark designed to assess the punctuation and capitalization prediction capabilities of end-to-end ASR models. The benchmark includes a LibriSpeech-PC dataset with restored punctuation and capitalization, a novel evaluation metric called Punctuation Error Rate (PER) that focuses on punctuation marks, and initial baseline models. All code, data, and models are publicly available.
Aleksandr Meister, Matvei Novikov, Nikolay Karpov, Evelina Bakhturina, Vitaly Lavrukhin, Boris Ginsburg
ASRU4
2023 SpellMapper: A non-autoregressive neural spellchecker for ASR customization with candidate retrieval based on n-gram mappings
Alexandra Antonova, Evelina Bakhturina, Boris Ginsburg
INTERSPEECH2
2023 P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech Prompting
abstract
While recent large-scale neural codec language models have shown significant improvement in zero-shot TTS by training on thousands of hours of data, they suffer from drawbacks such as a lack of robustness, slow sampling speed similar to previous autoregressive TTS methods, and reliance on pre-trained neural codec representations. Our work proposes P-Flow, a fast and data-efficient zero-shot TTS model that uses speech prompts for speaker adaptation. P-Flow comprises a speech-prompted text encoder for speaker adaptation and a flow matching generative decoder for high-quality and fast speech synthesis. Our speech-prompted text encoder uses speech prompts and text input to generate speaker-conditional text representation. The flow matching generative decoder uses the speaker-conditional output to synthesize high-quality personalized speech significantly faster than in real-time. Unlike the neural codec language models, we specifically train P-Flow on LibriTTS dataset using a continuous mel-representation. Through our training method using continuous speech prompts, P-Flow matches the speaker similarity performance of the large-scale zero-shot TTS models with two orders of magnitude less training data and has more than 20$\times$ faster sampling speed. Our results show that P-Flow has better pronunciation and is preferred in human likeness and speaker similarity to its recent state-of-the-art counterparts, thus defining P-Flow as an attractive and desirable alternative. We provide audio samples on our demo page: [https://research.nvidia.com/labs/adlr/projects/pflow](https://research.nvidia.com/labs/adlr/projects/pflow)
Sungwon Kim 0001, Kevin J. Shih, Rohan Badlani, João Felipe Santos, Evelina Bakhturina, Mikyas T. Desta, Rafael Valle, Sungroh Yoon, Bryan Catanzaro
NeurIPS5
2022 Thutmose Tagger: Single-pass neural model for Inverse Text Normalization
abstract
Inverse text normalization (ITN) is an essential post-processing step in automatic speech recognition (ASR).It converts numbers, dates, abbreviations, and other semiotic classes from the spoken form generated by ASR to their written forms.One can consider ITN as a Machine Translation task and use neural sequence-tosequence models to solve it.Unfortunately, such neural models are prone to hallucinations that could lead to unacceptable errors.To mitigate this issue, we propose a single-pass token classifier model that regards ITN as a tagging task.The model assigns a replacement fragment to every input token or marks it for deletion or copying without changes.We present a method of dataset preparation, based on granular alignment of ITN examples.The proposed model is less prone to hallucination errors.The model is trained on the Google Text Normalization dataset and achieves state-of-the-art sentence accuracy on both English and Russian test sets.One-to-one correspondence between tags and input words improves the interpretability of the model's predictions, simplifies debugging, and allows for post-processing corrections.The model is simpler than sequence-to-sequence models and easier to optimize in production settings.The model and the code to prepare the dataset is published as part of NeMo project 1 .
Alexandra Antonova, Evelina Bakhturina, Boris Ginsburg
INTERSPEECH2
2022 Shallow Fusion of Weighted Finite-State Transducer and Language Model for Text Normalization
abstract
Text normalization (TN) systems in production are largely rule-based using weighted finite-state transducers (WFST).However, WFST-based systems struggle with ambiguous input when the normalized form is context-dependent.On the other hand, neural text normalization systems can take context into account but they suffer from unrecoverable errors and require labeled normalization datasets, which are hard to collect.We propose a new hybrid approach that combines the benefits of rule-based and neural systems.First, a non-deterministic WFST outputs all normalization candidates, and then a neural language model picks the best one -similar to shallow fusion for automatic speech recognition.While the WFST prevents unrecoverable errors, the language model resolves contextual ambiguity.The approach is easy to extend and we show it is effective.It achieves comparable or better results than existing state-of-theart TN models.
Evelina Bakhturina, Yang Zhang 0089, Boris Ginsburg
INTERSPEECH1
2021 Hi-Fi Multi-Speaker English TTS Dataset
abstract
This paper introduces a new multi-speaker English dataset for training text-to-speech models.The dataset is based on Lib-riVox audiobooks and Project Gutenberg texts, both in the public domain.The new dataset contains about 292 hours of speech from 10 speakers with at least 17 hours per speaker sampled at 44.1 kHz.To select speech samples with high quality, we considered audio recordings with a signal bandwidth of at least 13 kHz and a signal-to-noise ratio (SNR) of at least 32 dB.The dataset is publicly released at "http://www.openslr.org/109/".
Evelina Bakhturina, Vitaly Lavrukhin, Boris Ginsburg, Yang Zhang 0089
Interspeech1
2021 NeMo (Inverse) Text Normalization: From Development to Production
Yang Zhang 0089, Evelina Bakhturina, Boris Ginsburg
Interspeech2
2021 NeMo Inverse Text Normalization: From Development to Production
abstract
Inverse text normalization (ITN) converts spoken-domain automatic speech recognition (ASR) output into written-domain text to improve the readability of the ASR output. Many state-of-the-art ITN systems use hand-written weighted finite-state transducer(WFST) grammars since this task has extremely low tolerance to unrecoverable errors. We introduce an open-source Python WFST-based library for ITN which enables a seamless path from development to production. We describe the specification of ITN grammar rules for English, but the library can be adapted for other languages. It can also be used for written-to-spoken text normalization. We evaluate the NeMo ITN library using a modified version of the Google Text normalization dataset.
Yang Zhang 0089, Evelina Bakhturina, Kyle Gorman, Boris Ginsburg
Interspeech2
2020 BioMegatron: Larger Biomedical Domain Language Model
abstract
Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, Raghav Mani. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Hoo-Chang Shin, Yang Zhang 0089, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, Raghav Mani
EMNLP (1)3