VLDB 2026 Research / reviewers in the wild / expert
Diamantino Caseiro
dblp:45/3872
· DBLP profile ↗
37ranked-venue papers
8as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 25 · 6 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Optimizing Large-Scale Context Retrieval for End-to-End ASR
Diamantino Caseiro, Kandarp Joshi, Christopher Li, Pat Rondon, Zelin Wu, Petr Zadrazil, Lillian Zhou |
INTERSPEECH | 2 |
| 2024 | Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm
Zelin Wu, Diamantino Caseiro, Tsendsuren Munkhdalai, Khe Chai Sim, Pat Rondon, Golan Pundak, Gan Song, Rohit Prabhavalkar, Zhong Meng, Ding Zhao, Tara Sainath, Yanzhang He, Pedro J. Moreno 0001 |
INTERSPEECH | 3 |
| 2023 | Contextual Spelling Correction with Large Language ModelsabstractContextual Spelling Correction (CSC) models are used to improve automatic speech recognition (ASR) quality given userspecific context. Typically, context is modeled as a large set of text spans to compare against a given ASR hypothesis using some distance measure (text, phonetic, or neural embedding). In this work we propose a CSC system based on a single Large Language Model (LLM) adapted with prompt tuning. Our approach is shown to be data efficient, and does not require dedicated serving. Our system exhibits advanced contextualization capabilities, such as support for phonetic spellings, cross-lingual scripts, and context specified as topics, with little to no data engineering. On voice assistant datasets, our system achieves $7.8 \%$ absolute word error rate reduction from a reference ASR system with relevant context and improving upon other contextualization solutions. Finally, we test our system in a prompt-injection attack scenario and report vulnerabilities and mitigations. Gan Song, Zelin Wu, Golan Pundak, Angad Chandorkar, Kandarp Joshi, Xavier Velez, Diamantino Caseiro, Ben Haynor, Nikhil Siddhartha, Pat Rondon, Khe Chai Sim |
ASRU | 7 |
| 2023 | Improving Contextual Biasing with Text InjectionabstractIn this work, we present a model-based approach to improving contextual biasing that improves quality without drastically increasing model computation during inference. Specifically, we look at injecting text data during training which is representative of contextually-relevant context that will be seen at inference, using a modality-matching text injection method known as JOIST. As JOIST injects text data directly into the E2E model, there is no additional model computation during inference, which is a big difference compared to most model-based biasing techniques. We find that our proposed approach, when combined with an FST-based context model, improves recognition of contacts between 5–15% relative. Tara N. Sainath, Rohit Prabhavalkar, Diamantino Caseiro, Pat Rondon, Cyril Allauzen |
ICASSP | 3 |
| 2021 | Improving Entity Recall in Automatic Speech Recognition with Neural EmbeddingsabstractAutomatic speech recognition (ASR) systems often have difficulty recognizing long-tail entities such as contact names and local restaurant names, which usually do not occur, or occur infrequently, in the system’s training data. In this work, we present a method which uses learned text embeddings and nearest neighbor retrieval within a large database of entity embeddings to correct misrecognitions. Our text embeddings are produced by a neural network trained so that the embeddings of acoustically confusable phrases have low cosine distances. Given the embedding of the text of a potential entity misrecognition and a precomputed database containing entities and their corresponding embeddings, we use fast, scalable nearest neighbor retrieval algorithms to find candidate corrections within the database. The inserted candidates are then scored using a function of the original text’s cost in the lattice and the distance between the embedding of the original text and the embedding of the candidate correction. Using this lattice augmentation techique, we demonstrate a 46% reduction in word error rate (WER) and 46% reduction in oracle word error rate (OWER) on an evaluation set with popular film queries. Christopher Li, Pat Rondon, Diamantino Caseiro, Leonid Velikovich, Xavier Velez, Petar S. Aleksic |
ICASSP | 3 |
| 2021 | An Efficient Streaming Non-Recurrent On-Device End-to-End Model with Improvements to Rare-Word Modeling
Tara N. Sainath, Yanzhang He, Arun Narayanan, Rami Botros, Ruoming Pang, David Rybach, Cyril Allauzen, Ehsan Variani, James Qin, Quoc-Nam Le-The, Shuo-Yiin Chang, Bo Li 0028, Anmol Gulati, Chung-Cheng Chiu, Diamantino Caseiro, Wei Li 0133, Qiao Liang 0001, Pat Rondon |
Interspeech | 16 |
| 2020 | Mixed Case Contextual ASR Using Capitalization Masks
Diamantino Caseiro, Pat Rondon, Quoc-Nam Le The, Petar S. Aleksic |
INTERSPEECH | 1 |
| 2018 | Entropy Based Pruning of Backoff Maxent Language Models with Contextual FeaturesabstractIn this paper, we present a pruning technique for maximum entropy (MaxEnt) language models. It is based on computing the exact entropy loss when removing each feature from the model, and it explicitly supports backoff features by replacing each removed feature with its backoff. The algorithm computes the loss on the training data, so it is not restricted to models with n-gram like features, allowing models with any feature, including long range skips, triggers, and contextual features such as device location. Results on the I-billion word corpus show large perplexity improvements relative for frequency pruned models of comparable size. Automatic speech recognition (ASR) experiments show word error rate improvements in a large-scale cloud based mobile ASR system for Italian. Tongzhou Chen, Diamantino Caseiro, Pat Rondon |
ICASSP | 2 |
| 2017 | Effectively Building Tera Scale MaxEnt Language Models Incorporating Non-Linguistic Signals
Fadi Biadsy, Mohammadreza Ghodsi, Diamantino Caseiro |
INTERSPEECH | 3 |
| 2017 | Sparse Non-Negative Matrix Language Modeling: Maximum Entropy Flexibility on the Cheap
Ciprian Chelba, Diamantino Caseiro, Fadi Biadsy |
INTERSPEECH | 2 |
| 2013 | Multiple parallel hidden layers and other improvements to recurrent neural network language modelingabstractRecurrent neural network language modeling (RNNLM) have been shown to outperform most other advanced language modeling techniques, however, it suffers from high computational complexity. In this paper, we present techniques for building faster and more accurate RNNLMs. In particular, we show that Brown clustering of the vocabulary is much more effective than other techniques. We also present an algorithm for converting an ensemble of RNNLMs into a single model that can be further tuned or adapted. The resulting models have significantly lower perplexity than single models with the same number of parameters. An error rate reduction of 5.9% was observed on a state of the art multi-pass voice-mail to text ASR system using RNNLMs trained with the proposed algorithm. Diamantino Caseiro, Andrej Ljolje |
ICASSP | 1 |
| 2012 | A general discriminative training algorithm for speech recognition using weighted finite-state transducersabstractIn this paper, we present a general algorithmic framework based on WFSTs for implementing a variety of discriminative training methods, such as MMI, MCE, and MPE/MWE. In contrast to the ordinary word lattices, the transducer-based lattices are more amenable to representing and manipulating the underlying hypothesis space and have a finer granularity at the HMM-state level. The transducers are processed into a two-layer hierarchy: at a high level, it is analogous to a word lattice, and each word transition embodies an HMM-state subgraph for that word at a lower level. This hierarchy combined with the appropriate customization of the transducers leads to a flexible implementation for all of the training criteria being discussed. The effectiveness of the framework is verified on two speech recognition tasks: Resource Management, and AT&T SCANMail, an internal voicemail-to-text task. Yong Zhao 0008, Andrej Ljolje, Diamantino Caseiro, Biing-Hwang Juang |
ICASSP | 3 |
| 2011 | Speech recognition modeling advances for mobile voice searchabstractThis paper reports on the development and advances in automatic speech recognition for the AT&T Speak4it®voice-search application. With Speak4it as real-life example, we show the effectiveness of acoustic model (AM) and language model (LM) estimation (adaptation and training) on relatively small amounts of application field-data. We then introduce algorithmic improvements concerning the use of sentence length in LM, of non-contextual features in AM decision-trees, and of the Teager energy in the acoustic front-end. The combination of these algorithms, integrated into the AT&T Watson recognizer, yields substantial accuracy improvements. LM and AM estimation on field-data samples increases the word accuracy from 66.4% to 77.1%, a relative word error reduction of 32%. The algorithmic improvements increase the accuracy to 79.7%, an additional 11.3% relative error reduction. Enrico Bocchieri, Diamantino Caseiro, Dimitrios Dimitriadis |
ICASSP | 2 |
| 2011 | An alternative front-end for the AT&T WATSON LV-CSR systemabstractIn previously published work, we have proposed a novel feature extraction algorithm, based on the Teager-Kaiser energy estimates, that approximates human auditory characteristics and that is more robust to sub-band noise than the mean-square estimates of standard MFCCs. We refer to the novel features as Teager energy cepstrum coefficients (TECC). Herein, we study the TECC performance under additive noise and suggest how to predict the noisy TECC deviations by estimating the subband SNR values. Then, we report on the effectiveness of the TECCs when they are used hi the acoustic front-end of the state-of-the-art AT&T WATSON large-vocabulary recognizer. The TECC front-end is tested in the real-life voice-search Speak4it application for mobile devices. It provides a 6% relative word error rate reduction w.r.t. the MFCC front-end, using the same high performance language model, lexicon and acoustic model training. Dimitrios Dimitriadis, Enrico Bocchieri, Diamantino Caseiro |
ICASSP | 3 |
| 2011 | Semantic data selection for vertical business voice searchabstractLocal business voice search is a popular application for mobile phones, where hands-free interaction and speed are critical to users. However, speech recognition accuracy is still not satisfactory when the number of businesses and locations is extended nationwide. For mobile users, searching a local business directory is often related to the fulfillment of specific tasks “on-the-move”, such as finding a restaurant, a movie theater, or a retailer chain. Restricting the local search to specific domains improves the quality of search results. In this paper, we present a new approach to data selection for bootstrapping and optimizing language models for vertical business sectors by exploiting semantic knowledge encoded in the business database and in the business category taxonomy. We demonstrate that, in the case of queries in the restaurant domain and without collecting new data, speech recognition word accuracy improves by 9.5% relative when compared with a generic local business language model. Giuseppe Di Fabbrizio, Diamantino Caseiro, Amanda Stent |
ICASSP | 2 |
| 2011 | SpeechForms: From Web to Speech and BackabstractThis paper describes SpeechForms, a system that uses novel techniques to automatically identify form element semantics and form element content, and to semi-automatically generate language models that allow users to fill out each web form element by voice. Preliminary experimental results show that simple per-element language models are faster and may be more accurate than statistical n-gram language models trained on large amounts of web text data. Index Terms: language modeling, form understanding, information retrieval Luciano Barbosa, Diamantino Caseiro, Giuseppe Di Fabbrizio |
INTERSPEECH | 2 |
| 2011 | Your Mobile Virtual Assistant Just Got Smarter!abstractA Mobile Virtual Assistant (MVA) is a communication agent that recognizes and understands free speech, and performs actions such as retrieving information and completing transactions. One essential characteristic of MVAs is their ability to learn and adapt without supervision. This paper describes our ongoing research in developing more intelligent MVAs that recognize and understand very large vocabulary speech input across a variety of tasks. In particular, we present our architecture for unsupervised acoustic and language model adaptation. Experimental results show that unsupervised acoustic model learning approaches the performance of supervised learning when adapting on 40-50 device-specific utterances. Unsupervised language model learning results in an 8% absolute drop in word error rate. Mazin Gilbert, Iker Arizmendi, Enrico Bocchieri, Diamantino Caseiro, Vincent Goffin, Andrej Ljolje, Mike Phillips, Chao Wang 0018, Jay G. Wilpon |
INTERSPEECH | 4 |
| 2011 | Visual Voice Mail to Text on the iPhone/iPad
Andrej Ljolje, Vincent Goffin, Diamantino Caseiro, Taniya Mishra, Mazin Gilbert |
INTERSPEECH | 3 |
| 2010 | Use of geographical meta-data in ASR language and acoustic modelsabstractThe query distribution, in the speech recognition applications of directory assistance (DA) and voice-search, depends on the customer's location. This motivates the research on query models conditioned on the user location, here denoted as local models. We describe and test our methods for the estimation of local models with various degrees of spacial “granularity”, for the recognition of city-state (sub-task of DA) and for the recognition of business listings, spoken over iPhones in a nation-wide business-listing voice-search service. Our local language models improve the accuracy of city-state by 2.4% absolute (32% relative error reduction), and of voice-search by 2.2% (7% relative). Enrico Bocchieri, Diamantino Caseiro |
ICASSP | 2 |
| 2010 | WFST compression for automatic speech recognition
Diamantino Caseiro |
INTERSPEECH | 1 |
| 2009 | Geo-Centric Language Models for Local Business Voice Search
Amanda Stent, Ilija Zeljkovic, Diamantino Caseiro, Jay G. Wilpon |
HLT-NAACL | 3 |
| 2008 | Broadcast news subtitling system in PortugueseabstractThe subtitling of broadcast news programs are starting to become a very interesting application due to the technological advances in automatic speech recognition and associated technologies. However, to build this kind of systems, several advances are necessary both in terms of the technological components and on main blocks integration. In this paper, we are presenting the overall architecture of a subtitling system running daily at RTP (the Portuguese public broadcast company). The goal is to integrate our components in a system for the subtitling of RTP programs. The global system includes the subtitling of recorded and direct programs. João Paulo da Silva Neto, Hugo Meinedo, Márcio Viveiros, Renato Cassaca, Ciro Martins, Diamantino Caseiro |
ICASSP | 6 |
| 2008 | Building a Golden Collection of Parallel Multi-Language Word Alignment
João Graça, Joana Paulo Pardal, Luísa Coheur, Diamantino Caseiro |
LREC | 4 |
| 2008 | Recovering capitalization and punctuation marks for automatic speech recognition: Case study for Portuguese broadcast news
Fernando Batista, Diamantino Caseiro, Nuno J. Mamede, Isabel Trancoso |
Speech Commun. | 2 |
| 2007 | The Titech large vocabulary WFST speech recognition systemabstractIn this paper we present evaluations on the large vocabulary speech decoder we are currently developing at Tokyo Institute of Technology. Our goal is to build a fast, scalable, flexible decoder to operate on weighted finite state transducer (WFST) search spaces. Even though the development of the decoder is still in its infancy we have already implemented a impressive feature set and are achieving good accuracy and speed on a large vocabulary spontaneous speech task. We have developed a technique to allow parts of the decoder to be run on the graphics processor, this can lead to a very significant speed up. Paul R. Dixon, Diamantino Caseiro, Tasuku Oonishi, Sadaoki Furui |
ASRU | 2 |
| 2007 | Recovering punctuation marks for automatic speech recognitionabstractThis paper shows results of recovering punctuation over speech transcriptions for a Portuguese broadcast news corpus. The approach is based on maximum entropy models and uses word, part-of-speech, time and speaker information. The contribution of each type of feature is analyzed individually. Separate results for each focus condition are given, making it possible to analyze the differences of performance between planned and spontaneous speech. Index Terms: rich transcription, punctuation recovery, sentence boundary detection, maximum entropy. Fernando Batista, Diamantino Caseiro, Nuno J. Mamede, Isabel Trancoso |
INTERSPEECH | 2 |
| 2006 | Spoken language technologies applied to digital talking booksabstractDigital Talking Books (DTBs) offer to visually impaired users an evolution of analogue talking books that mimics the interaction possibilities of print books. This paper describes a new DTB player which tries to improve the usability and accessibility of current players, through the combination of the possibilities offered by multimodal interaction and interface adaptability, and the integration of several language processing components. Besides the potential for a greater enjoyment of the reader in general, these modifications also pave the way to the use of DTBs in different domains, from e-inclusion to e-learning applications. Index Terms: digital talking books, Portuguese. 1. Isabel Trancoso, Carlos Duarte, António Joaquim Serralheiro, Diamantino Caseiro, Luís Carriço, Céu Viana |
INTERSPEECH | 4 |
| 2006 | Recognition of classroom lectures in european portugueseabstractClassroom lectures may be very challenging for automatic speech recognizers, because the vocabulary may be very specific and the speaking style very spontaneous. Our first experiments using a recognizer trained for Broadcast News resulted in word error rates near 60%, clearly confirming the need for adaptation to the specific topic of the lectures, on one hand, and for better strategies for handling spontaneous speech. This paper describes our efforts in these two directions: the different domain adaptation steps that lowered the error rate to 45%, with very little transcribed adaptation material, and the exploratory study of spontaneous speech phenomena in European Portuguese, namely concerning filled pauses. Index Terms: spontaneous speech recognition, Portuguese. 1. Isabel Trancoso, Ricardo Nunes, Luís Neves, Céu Viana, Helena Moniz, Diamantino Caseiro, Ana Isabel Mata |
INTERSPEECH | 6 |
| 2006 | A specialized on-the-fly algorithm for lexicon and language model compositionabstractThis paper presents an algorithm for the composition of weighted finite-state transducers which is specially tailored to speech recognition applications: it composes the lexicon with the language model while simultaneously optimizing the resulting transducer. Furthermore, it performs these computations "on-the-fly" to allow easier management of the tradeoff between offline and online computation and memory. The algorithm is exact for local knowledge integration and optimization operations such as composition and determinization. Minimization and pushing operations are approximated. Our results have confirmed the efficiency of these approximations Diamantino Caseiro, Isabel Trancoso |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Finite-state transducer inference for a speech-input Portuguese-to-English machine translation systemabstractStatistical techniques and grammatical inference have been used for dealing with automatic speech recognition with success, and can also be used for speech-to-speech machine translation. In this paper, new advances on a method for finite-state transducer inference are presented. This method has been tested experimentally in a speech-input translation task using a recognizer that allows a flexible use of models by means of efficient algorithms for on-thefly transducer composition. These are the first reported results of a speech-to-speech translation task involving European Portuguese input that we know of. David Picó, Francisco Casacuberta, Diamantino Caseiro, Isabel Trancoso |
INTERSPEECH | 4 |
| 2005 | Aligning and recognizing spoken books in different varieties of PortugueseabstractThis paper tries to present digital spoken books as a useful diagnostic tool for detecting alignment and recognition problems and for studying the porting of these technologies to different varieties of the same language- Portuguese, in our case. We summarize the main differences between European and Brazilian Portuguese (EP/BP) and describe how they affect the GtoP system. Despite the small size of our parallel spoken book corpus in the two varieties, our preliminary experiments confirmed our expectations in terms of the effectiveness of an EP-trained aligner used on BP spoken books. They also confirmed the inadequacy of an EP Broadcast News recognizer tested over literary contents, and the expected degradation in recognition scores caused by using that recognizer on a BP spoken book. Pronunciation adaptation was tested by adding variants derived by the BP GtoP system to our EP lexicon, resulting in a very small improvement in terms of recognition scores. Isabel Trancoso, António Joaquim Serralheiro, Céu Viana, Diamantino Caseiro |
INTERSPEECH | 4 |
| 2003 | A tail-sharing WFST composition algorithm for large vocabulary speech recognitionabstractThis paper presents an algorithm for approximating minimization in the context of the weighted finite-state transducers approach to large vocabulary speech recognition. The algorithm is designed for the integration of the lexicon with the language model and performs composition, determinization and pushing in one step. Furthermore, it uses tail-sharing in order to approximate minimization. Our results show that it is a good approximation to explicit minimization, with the added advantage that it can be used "on-the-fly" in a dynamic decoder. Diamantino Caseiro, Isabel Trancoso |
ICASSP (1) | 1 |
| 2003 | Towards a repository of digital talking booksabstractConsiderable effort has been devoted at# # # to increase and broaden our speech and text data resources. Digital Talking Books (DTB), comprising both speech and text data are, as such, an invaluable asset as multimedia resources. Furthermore, those DTB have been under a speech-to-text alignment procedure, either word or phone-based, to increase their potential in research activities. This paper thus describes the motivation and the method that we used to accomplish this goal for aligning DTBs. This alignment allows specific access interfaces for persons with special needs, and also tools for easily detecting and indexing units (words, sentences, topics) in the spoken books. The alignment tool was implemented in a Weighted Finite State Transducer framework, which provides an efficient way to combine different types of knowledge sources, such as alternative pronunciation rules. With this tool, a 2-hour long spoken book was aligned in a single step in much less than real time. Last but not least, new browsing interfaces, allowing improved access and data retrieval to and from the DTBs, are described in this paper. António Joaquim Serralheiro, Isabel Trancoso, Diamantino Caseiro, Teresa Chambel, Luís Carriço, Nuno Guimarães |
INTERSPEECH | 3 |
| 2002 | Using dynamic WFST composition for recognizing broadcast newsabstractOur first application of weighted finite state transducers to the recognition of broadcast news provided us with an interesting framework to study several problems related to the optimization of the search space. The paper starts by describing how the use of our lexicon and language model "on-the-fly" composition algorithm is crucial in extending the transducer approach to large systems. We present an efficient representation for WFSTs, that allowed us to reduce runtime memory requirements, and discuss several types of language model optimizations, including a context-sharing algorithm. Experimental results obtained with the broadcast news corpus collected for European Portuguese illustrate the impact of the various possible optimizations of the components on the performance of the system. Diamantino Caseiro, Isabel Trancoso |
INTERSPEECH | 1 |
| 2001 | On integrating the lexicon with the language modelabstractThe goal of this work was to develop an algorithm for the integration of the lexicon with the language model which would be computationally efficient in terms of memory requirements, even in the case of large trigram models. Two specialized versions of the algorithm for transducer composition were implemented. The first one is basically a composition algorithm that uses the precomputed set of the output labels that can be reached from a particular epsilon edge of the lexicon Diamantino Caseiro, Isabel Trancoso |
INTERSPEECH | 1 |
| 2000 | Phonetic vocoder assessmentabstractThe efficiency of phonetic vocoders stems from the fact that the only transmitted information is the index of the recognised units and the corresponding prosodic parameters. Hence, speaker recognisability is one of the main issues in this class of coders. Our approach to minimise this drawback was to include some speaker adaptation capability. The purpose of this paper is two-folded: on one hand, to describe the recognisability and intelligibility tests that were performed with our phonetic vocoder with and without speaker adaptation; on the other hand, to present our recent developments of this coder, using the SpeechDat corpus for Portuguese, that includes telephone calls from 5000 speakers. This allowed us to generate improved HMM models, codebooks, and quantization tables, and to investigate the performance of the coder in non-clean environments and with a much wider speaker population. Carlos M. Ribeiro, Isabel Trancoso, Diamantino Caseiro |
INTERSPEECH | 3 |
| 1998 | Spoken language identification using the speechdat corpusabstractCurrent language identification systems vary significantly in their complexity. The systems that use higher level linguistic information have the best performance. Nevertheless, that information is hard to collect for each new language. The system presented in this paper is easily extendable to new languages because it uses very little linguistic information. In fact, the presented system needs only one language specific phone recogniser (in our case the Portuguese one), and is trained with speech from each of the other languages. With the SpeechDat-M corpus, with 6 European languages (English, French, German, Italian, Portuguese and Spanish) our system achieved an identification rate of 83.4% on 5-second utterances, this result shows an improvement of 5% over our previous version, mainly through the use of a neural network classifier. Both the baseline and the full system were implemented in realtime. 1. INTRODUCTION When designing a language identification system, we face the prob... Diamantino Caseiro, Isabel Trancoso |
ICSLP | 1 |