EDBT 2026 Demo / reviewers in the wild / expert
Tanel Alumäe
dblp:78/2247
· DBLP profile ↗
35ranked-venue papers
19as first author
12since 2021 · last 2026
0000-0001-5083-1556ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 16 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 18 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Estonian Native Large Language Model BenchmarkabstractThe availability of LLM benchmarks for the Estonian language is limited, and a comprehensive evaluation comparing the performance of different LLMs on Estonian tasks has yet to be conducted. We introduce a new benchmark for evaluating LLMs in Estonian, based on seven diverse datasets. These datasets assess general and domain-specific knowledge, understanding of Estonian grammar and vocabulary, summarization abilities, contextual comprehension, and more. The datasets are all generated from native Estonian sources without using machine translation. We compare the performance of base models, instruction-tuned open-source models, and commercial models. Our evaluation includes 6 base models and 26 instruction-tuned models. To assess the results, we employ both human evaluation and LLM-as-a-judge methods. Human evaluation scores showed moderate to high correlation with benchmark evaluations, depending on the dataset. Claude 3.7 Sonnet, used as an LLM judge, demonstrated strong alignment with human ratings, indicating that top-performing LLMs can effectively support the evaluation of Estonian-language models. Helena Grete Lillepalu, Tanel Alumäe |
LREC | 2 |
| 2026 | Design choices for PixIT-based speaker-attributed ASR: Team ToTaTo at the NOTSOFAR-1 challenge
Joonas Kalda, Séverin Baroudi, Martin Lebourdais, Clément Pagés, Ricard Marxer, Tanel Alumäe, Hervé Bredin |
Comput. Speech Lang. | 6 |
| 2025 | TalTech Systems for the PROCESS Signal Processing Grand ChallengeabstractThe PROCESS Challenge aims to detect cognitive decline, including early stages like mild cognitive impairment, through spontaneous speech. This paper describes TalTech’s systems prepared for the challenge that applied machine learning models incorporating multimodal features to address both regression and classification tasks. For regression, the Lasso model achieved an RMSE of 2.54 on the test set, achieving 2nd place in the challenge. For classification, the XGBoost model achieved a macro F1 score of 0.61, placing 6th. These results demonstrate the potential of integrating diverse speech-based features and predictive modeling for scalable, early detection of cognitive decline. Erik Illaste, Tanel Alumäe |
ICASSP | 2 |
| 2025 | TalTech Systems for the Interspeech 2025 ML-SUPERB 2.0 Challenge
Tanel Alumäe, Artem Fedorchenko |
INTERSPEECH | 1 |
| 2025 | Diarization-Guided Multi-Speaker EmbeddingsabstractInternational audience Joonas Kalda, Clément Pagés, Tanel Alumäe, Hervé Bredin |
INTERSPEECH | 3 |
| 2025 | Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion
Ajinkya Kulkarni, Sandipana Dowerah, Tanel Alumäe, Mathew Magimai-Doss |
INTERSPEECH | 3 |
| 2024 | TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024abstractInternational audience Joonas Kalda, Tanel Alumäe, Martin Lebourdais, Hervé Bredin, Séverin Baroudi, Ricard Marxer |
INTERSPEECH | 2 |
| 2023 | Dialect Adaptation and Data Augmentation for Low-Resource ASR: Taltech Systems for the Madasr 2023 ChallengeabstractThis paper describes Tallinn University of Technology (TalTech) systems developed for the ASRU MADASR 2023 Challenge. The challenge focuses on automatic speech recognition of dialect-rich Indian languages with limited training audio and text data. TalTech participated in two tracks of the challenge: Track 1 that allowed using only the provided training data and Track 3 which allowed using additional audio data. In both tracks, we relied on wav2vec 2.0 models. Our methodology diverges from the traditional procedure of finetuning pretrained wav2vec 2.0 models in two key points: firstly, through the implementation of the aligned data augmentation technique to enhance the linguistic diversity of the training data, and secondly, via the application of deep prefix tuning for dialect adaptation of wav2vec 2.0 models. In both tracks, our approach yielded significant improvements over the provided baselines, achieving the lowest word error rates across all participating teams. Tanel Alumäe, Jiaming Kong, Daniil Robnikov |
ASRU | 1 |
| 2023 | Exploring the Impact of Pretrained Models and Web-Scraped Data for the 2022 NIST Language Recognition Evaluation
Tanel Alumäe, Kunnar Kukk, Viet Bac Le, Claude Barras, Abdelkhalek Messaoudi, Waad Ben Kheder |
INTERSPEECH | 1 |
| 2022 | Improving Language Identification of Accented SpeechabstractLanguage identification from speech is a common preprocessing step in many spoken language processing systems.In recent years, this field has seen fast progress, mostly due to the use of self-supervised models pretrained on multilingual data and the use of large training corpora.This paper shows that for speech with a non-native or regional accent, the accuracy of spoken language identification systems drops dramatically, and that the accuracy of identifying the language is inversely correlated with the strength of the accent.We also show that using the output of a lexicon-free speech recognition system of the particular language helps to improve language identification performance on accented speech by a large margin, without sacrificing accuracy on native speech.We obtain relative error rate reductions ranging from to 35 to 63% over the state-of-the-art model across several non-native speech datasets. Kunnar Kukk, Tanel Alumäe |
INTERSPEECH | 2 |
| 2021 | Combining Hybrid and End-to-End Approaches for the OpenASR20 Challenge
Tanel Alumäe, Jiaming Kong |
Interspeech | 1 |
| 2021 | VOXLINGUA107: A Dataset for Spoken Language RecognitionabstractThis paper investigates the use of automatically collected web audio data for the task of spoken language recognition. We generate semi-random search phrases from language-specific Wikipedia data that are then used to retrieve videos from YouTube for 107 languages. Speech activity detection and speaker diarization are used to extract segments from the videos that contain speech. Post-filtering is used to remove segments from the database that are likely not in the given language, increasing the proportion of correctly labeled segments to 98%, based on crowd-sourced verification. The size of the resulting training set (VoxLingua107) is 6628 hours (62 hours per language on the average) and it is accompanied by an evaluation set of 1609 verified utterances. We use the data to build language recognition models for several spoken language identification tasks. Experiments show that using the automatically retrieved training data gives competitive results to using hand-labeled proprietary datasets. The dataset is publicly available1. Jörgen Valk, Tanel Alumäe |
SLT | 2 |
| 2020 | Robust Training of Vector Quantized Bottleneck ModelsabstractIn this paper we demonstrate methods for reliable and efficient training of discrete representation using Vector-Quantized Variational Auto-Encoder models (VQ-VAEs). Discrete latent variable models have been shown to learn nontrivial representations of speech, applicable to unsupervised voice conversion and reaching state-of-the-art performance on unit discovery tasks. For unsupervised representation learning, they became viable alternatives to continuous latent variable models such as the Variational Auto-Encoder (VAE). However, training deep discrete variable models is challenging, due to the inherent non-differentiability of the discretization operation. In this paper we focus on VQ-VAE, a state-of-the-art discrete bottleneck model shown to perform on par with its continuous counterparts. It quantizes encoder outputs with on-line k-means clustering. We show that the codebook learning can suffer from poor initialization and non-stationarity of clustered encoder outputs. We demonstrate that these can be successfully overcome by increasing the learning rate for the codebook and periodic date-dependent codeword re-initialization. As a result, we achieve more robust training across different tasks, and significantly increase the usage of latent codewords even for large codebooks. This has practical benefit, for instance, in unsupervised representation learning, where large codebooks may lead to disentanglement of latent representations. Adrian Lancucki, Jan Chorowski, Guillaume Sanchez, Ricard Marxer, Nanxin Chen, Hans J. G. A. Dolfing, Sameer Khurana, Tanel Alumäe, Antoine Laurent |
IJCNN | 8 |
| 2020 | The TalTech Systems for the Short-Duration Speaker Verification Challenge 2020
Tanel Alumäe, Jörgen Valk |
INTERSPEECH | 1 |
| 2019 | Recognition of Creaky Voice from Emergency Calls
Lauri Tavi, Tanel Alumäe |
INTERSPEECH | 2 |
| 2018 | Training Speaker Recognition Models with Recording-Level LabelsabstractIn this paper, we investigate training speaker recognition models using coarse-grained speaker labels provided only at the recording level. The approach is based on the recently proposed weakly supervised training method that allows to train a speaker recognition deep neural network using a special cost function that doesn't need segment-level annotations. Experiments are conducted on the VoxCeleb corpus. We show that without using any reference segment-level labeling, the method can achieve 1% speaker recognition error rate on the official VoxCeleb closed set speaker recognition test set, as opposed to 5.4% that was previously reported. By training a x-vector based speaker verification system on the resegmented and relabeled VoxCeleb corpus, we can achieve 4.57% EER on the VoxCeleb speaker verification test set which is a 17% relative improvement over the best system that uses the official VoxCeleb speaker annotations. Tanel Alumäe |
SLT | 1 |
| 2017 | The 2016 BBN Georgian telephone speech keyword spotting systemabstractIn this paper we describe the 2016 BBN conversational telephone speech keyword spotting system; the culmination of four years of research and development under the IARPA Babel program. The system was constructed in response to the NIST Open Keyword Search (OpenKWS) evaluation of 2016. We present our technological breakthroughs in building top-performing keyword spotting processing systems for new languages, in the face of limited transcribed speech, noisy conditions, and limited system build time of one week. Tanel Alumäe, Damianos Karakos, William Hartmann, Roger Hsiao, Le Zhang 0002, Long Nguyen 0001, Stavros Tsakalidis, Richard M. Schwartz |
ICASSP | 1 |
| 2017 | Analysis of keyword spotting performance across IARPA babel languagesabstractWith the completion of the IARPA Babel program, it is possible to systematically analyze the performance of speech recognition systems across a wide variety of languages. We select 16 languages from the dataset and compare performance using a deep neural network-based acoustic model. The focus is on keyword spotting using the actual term-weighted value (ATWV) metric. We demonstrate that ATWV is keyword dependent, and that this must be accounted for in any cross-language analysis. Further, we show that while performance across languages does not track with any particular feature of the language, it is correlated with inter-annotator agreement. William Hartmann, Damianos Karakos, Roger Hsiao, Le Zhang 0002, Tanel Alumäe, Stavros Tsakalidis, Richard M. Schwartz |
ICASSP | 5 |
| 2017 | Implementation of a Radiology Speech Recognition System for Estonian Using Open Source Software
Tanel Alumäe, Andrus Paats, Ivo Fridolin, Einar Meister |
INTERSPEECH | 1 |
| 2016 | Improved Multilingual Training of Stacked Neural Network Acoustic Models for Low Resource Languages
Tanel Alumäe, Stavros Tsakalidis, Richard M. Schwartz |
INTERSPEECH | 1 |
| 2016 | Sage: The New BBN Speech Processing Platform
Roger Hsiao, Ralf Meermeier, Tim Ng, Zhongqiang Huang, Maxwell Jordan, Enoch Kan, Tanel Alumäe, Jan Silovský, William Hartmann, Francis Keith, Omer Lang, Man-Hung Siu, Owen Kimball |
INTERSPEECH | 7 |
| 2016 | Bidirectional Recurrent Neural Network with Attention Mechanism for Punctuation Restoration
Ottokar Tilk, Tanel Alumäe |
INTERSPEECH | 2 |
| 2015 | LSTM for punctuation restoration in speech transcripts
Ottokar Tilk, Tanel Alumäe |
INTERSPEECH | 2 |
| 2014 | Neural network phone duration model for speech recognition
Tanel Alumäe |
INTERSPEECH | 1 |
| 2013 | Multi-domain neural network language modelabstractThe paper describes a neural network language model that jointly models language in many related domains. In addition to the traditional layers of a neural network language model, the proposed model also trains a vector of factors for each domain in the training data that are used to modulate the connections from the projection layer to the hidden layer. The model is found to outperform simple neural network language models as well as domain-adapted maximum entropy language models in perplexity evaluation and speech recognition experiments. Index Terms: neural network language model, language model adaptation 1. Tanel Alumäe |
INTERSPEECH | 1 |
| 2013 | Phone duration modeling using clustering of rich contextsabstractThis paper describes a phone duration model applied to speech recognition. The model is based on a decision tree that finds clusters of phones in various contexts that tend to have similar durations. Wide contexts with rich linguistic and phonetic features are used. To better model varying and non-stationary speaking rates, the contextual features also include the observed duration values of previous phones. For each resulting phone cluster, a log-normal distribution of duration is estimated. The resulting decision tree and the log-normal distributions are used to calculate likelihoods of phone durations in N-best lists. Experiments on two Estonian recognition tasks show a small but significant improvement in speech recognition accuracy. Index Terms: duration modeling, speech recognition, decision trees Tanel Alumäe, Rena Nemoto |
INTERSPEECH | 1 |
| 2012 | Maximum Entropy Language Model Adaptation for Mobile Speech InputabstractThis paper describes unsupervised adaptation of language model for many related target domains. In mobile speech input, subject and vocabulary of the language depend highly on the usage context. We use automatically transcribed speech data to select a subset from the language model training data for building a maximum entropy model adapted to speech input. This model is further adapted for most popular mobile applications. When used in interpolation with the background N-gram model, the adapted models give over 10 % relative word error rate reduction in Estonian mobile speech input experiments. Index Terms: language model adaptation, maximum entropy, mobile speech input Tanel Alumäe, Kaarel Kaljurand |
INTERSPEECH | 1 |
| 2012 | A Hierarchical Dirichlet Process Model for Joint Part-of-Speech and Morphology Induction
Kairit Sirts, Tanel Alumäe |
HLT-NAACL | 2 |
| 2011 | TSAB - Web Interface for Transcribed Speech Collections
Tanel Alumäe, Ahti Kitsik |
INTERSPEECH | 1 |
| 2010 | Efficient estimation of maximum entropy language models with n-gram features: an SRILM extensionabstractWe present an extension to the SRILM toolkit for training maximum entropy language models with N -gram features. The extension uses a hierarchical parameter estimation procedure [1] for making the training time and memory consumption feasible for moderately large training data (hundreds of millions of words). Experiments on two speech recognition tasks indicate that the models trained with our implementation perform equally to or better than N -gram models built with interpolated Kneser-Ney discounting. Tanel Alumäe, Mikko Kurimo |
INTERSPEECH | 1 |
| 2008 | Comparison of Different Modeling Units for Language Model Adaptation for Inflected Languages
Tanel Alumäe |
CICLing | 1 |
| 2007 | LSA-based language model adaptation for highly inflected languagesabstractThis paper presents a language model topic adaptation framework for highly inflected languages. In such languages, subword units are used as basic units for language modeling. Since such units carry little semantic information, they are not very suitable for topic adaptation. We propose to lemmatize the corpus of training documents before constructing a latent topic model. To adapt language model, we use few lemmatized training sentences to find a set of documents that are semantically close to the current document. Fast marginal adaptation of subword trigram language model is used for adapting the background model. Experiments on a set of Estonian test texts show that the proposed approach gives a 19 % decrease in language model perplexity. A statistically significant decrease in perplexity is observed already when using just two sentences for adaptation. We also show that the model employing lemmatization gives consistently better results than the unlemmatized model. Index Terms: speech recognition, language model adaptation, LSA, inflected languages Tanel Alumäe, Toomas Kirt |
INTERSPEECH | 1 |
| 2006 | Sentence-Adapted Factored Language Model for Transcribing Estonian SpeechabstractThis work presents a 2-pass recognition method for highly inflected agglutinative languages based on an Estonian large vocabulary recognition task. Morphemes are used as basic recognition units in a standard trigram language model in the first pass. The recognized morphemes are reconstructed back to words using hidden event language model for compound word detection. In the second pass, the vocabulary from N-best sentence candidates from the first pass is used to create an adaptive sentence-specific word-based language model which is applied for rescoring the N-best hypotheses. The sentence specific language model is based on the factored language model paradigm and estimates word probabilities based on the preceding two words and part-of-speech tags. The method achieves a 7.3% relative word error rate improvement over the baseline system that is used in the first pass Tanel Alumäe |
ICASSP (1) | 1 |
| 2006 | Unlimited vocabulary speech recognition for agglutinative languages
Mikko Kurimo, Antti Puurula, Ebru Arisoy, Vesa Siivola, Teemu Hirsimäki, Janne Pylkkönen, Tanel Alumäe, Murat Saraclar |
HLT-NAACL | 7 |
| 2004 | Large vocabulary continuous speech recognition for estonian using morpheme classes
Tanel Alumäe |
INTERSPEECH | 1 |