EDBT 2026 Demo / reviewers in the wild / expert
Ville Hautamäki
dblp:75/5713
· DBLP profile ↗
71ranked-venue papers
9as first author
17since 2021 · last 2025
0000-0002-5885-0003ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 7 first-author · 13 since 2021Artificial intelligence and machine learning · 40 · 8 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Importance of Representation in Imitating Human-Like GameplayabstractBeing able to synthesize human-like game play is highly useful for automating play-testing, generating naturally behaving bots, and for assisting game design in general. However, for complex games that require long-term planning this remains a major challenge. Imitation learning (IL) algorithms offer one approach for learning to replicate human-like behavior by training models that imitate the behavior observed in human demonstrations, and in fact much of the research in IL is done in context of games due to fast feedback loops and relatively easy access to the required data. One recent approach employing IL is Video PreTraining (VPT) that leverages representation models trained on massive unlabeled video collection of humans playing a complex game, specifically Minecraft. Despite promising results on a complex task, in this work we empirically demonstrate a failure mode of VPT as a representation learner in Minecraft, showing how the pretrained model is insufficient for providing useful representations for new tasks. We then explain how the fundamental challenge remains even in a substantially simplified game environment. We argue that for many practitioners finetuning the representation models for the tasks of interest is unfeasible, making the overall approach of limited interest for use in game design applications. Ville Tanskanen, Arto Klami, Ville Hautamäki |
CoG | 3 |
| 2025 | Interpreting Deep Neural Network-Based Receiver Under Varying Signal-To-Noise RatiosabstractWe propose a novel method for interpreting neural networks, focusing on convolutional neural network-based receiver model. The method identifies which unit or units of the model contain most (or least) information about the channel parameter(s) of the interest, providing insights at both global and local levels—with global explanations aggregating local ones. Experiments on link-level simulations demonstrate the method’s effectiveness in identifying units that contribute most (and least) to signal-to-noise ratio processing. Although we focus on a radio receiver model, the method generalizes to other neural network architectures and applications, offering robust estimation even in high-dimensional settings. Marko Tuononen, Dani Korpi, Ville Hautamäki |
ICASSP | 3 |
| 2025 | Continuous Learning for Children's ASR: Overcoming Catastrophic Forgetting with Elastic Weight Consolidation and Synaptic Intelligence
Edem Ahadzi, Vishwanath Pratap Singh, Tomi Kinnunen, Ville Hautamäki |
INTERSPEECH | 4 |
| 2025 | Zero-shot World Models via Search in MemoryabstractWorld Models have vastly permeated the field of Reinforcement Learning. Their ability to model the transition dynamics of an environment have led to tremendous improvements in sample efficiency for online RL. Among them, the most notorious example is Dreamer, a model that learns to act in a diverse set of image-based environments. In this paper, we leverage similarity search and stochastic representations to approximate a world model without a training procedure. We establish a comparison with PlaNet, a well-established world model of the Dreamer family. We evaluate the models on the quality of latent reconstruction and on the perceived similarity of the reconstructed image, on both next-step and long horizon dynamics prediction. The results of our study demonstrate that a search-based world model is comparable to a training based one in both cases. Notably, our model shows stronger performance in long-horizon prediction with respect to the baseline on a range of visually different environments. Federico Malato, Ville Hautamäki |
NeurIPS | 2 |
| 2025 | Improving Numerical Stability of Normalized Mutual Information Estimator on High DimensionsabstractMutual information provides a powerful, general-purpose metric for quantifying the amount of shared information between variables. Estimating normalized mutual information using a k-Nearest Neighbor (k-NN) based approach involves the calculation of the scaling-invariant k-NN radius. Calculation of the radius suffers from numerical overflow when the joint dimensionality of the data becomes high, typically in the range of several hundred dimensions. To address this issue, we propose a logarithmic transformation technique that improves the numerical stability of the radius calculation in high-dimensional spaces. By applying the proposed transformation during the calculation of the radius, numerical overflow is avoided, and precision is maintained. Proposed transformation is validated through both theoretical analysis and empirical evaluation, demonstrating its ability to stabilize the calculation without compromising precision, increasing bias, or adding significant computational overhead, while also helping to maintain estimator variance. Marko Tuononen, Ville Hautamäki |
IEEE Signal Process. Lett. | 2 |
| 2024 | Online Adaptation for Enhancing Imitation Learning PoliciesabstractImitation learning enables autonomous agents to learn from human examples, without the need for a reward signal. Still, if the provided dataset does not encapsulate the task correctly, or when the task is too complex to be modeled, such agents fail to reproduce the expert policy. We propose to recover from these failures through online adaptation. Our approach combines the action proposal coming from a pre-trained policy with relevant experience recorded by an expert. The combination results in an adapted action that closely follows the expert. Our experiments show that an adapted agent performs better than its pure imitation learning counterpart. Notably, adapted agents can achieve reasonable performance even when the base, non-adapted policy catastrophically fails. Federico Malato, Ville Hautamäki |
CoG | 2 |
| 2024 | Gradient Weighting for Speaker Verification in Extremely Low Signal-to-Noise RatioabstractSpeaker verification is hampered by background noise, particularly at extremely low Signal-to-Noise Ratio (SNR) under 0 dB. It is difficult to suppress noise without introducing unwanted artifacts, which adversely affects speaker verification. We proposed the mechanism called Gradient Weighting (Grad-W), which dynamically identifies and reduces artifact noise during prediction. The mechanism is based on the property that the gradient indicates which parts of the input the model is paying attention to. Specifically, when the speaker network focuses on a region in the denoised utterance but not on the clean counterpart, we consider it artifact noise and assign higher weights for this region during optimization of enhancement. We validate it by training an enhancement model and testing the enhanced utterance on speaker verification. The experimental results show that our approach effectively reduces artifact noise, improving speaker verification across various SNR levels. Kong-Aik Lee, Ville Hautamäki, Meng Ge, Haizhou Li 0001 |
ICASSP | 3 |
| 2024 | Zero-Shot Imitation Policy Via Search In Demonstration DatasetabstractBehavioral cloning uses a dataset of demonstrations to learn a policy. To overcome computationally expensive training procedures and address the policy adaptation problem, we propose to use latent spaces of pre-trained foundation models to index a demonstration dataset, instantly access similar relevant experiences, and copy behavior from these situations. Actions from a selected similar situation can be performed by the agent until representations of the agent’s current situation and the selected experience diverge in the latent space. Thus, we formulate our control problem as a dynamic search problem over a dataset of experts’ demonstrations. We test our approach on BASALT MineRL-dataset in the latent representation of a Video Pre-Training model. We compare our model to state-of-the-art, Imitation Learning-based Minecraft agents. Our approach can effectively recover meaningful demonstrations and show human-like behavior of an agent in the Minecraft environment in a wide variety of scenarios. Experimental results reveal that performance of our search-based approach clearly wins in terms of accuracy and perceptual evaluation over learning-based models. Federico Malato, Florian Leopold, Andrew Melnik, Ville Hautamäki |
ICASSP | 4 |
| 2024 | ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2vec2.0 Based ASR
Vishwanath Pratap Singh, Federico Malato, Ville Hautamäki, Md. Sahidullah, Tomi Kinnunen |
INTERSPEECH | 3 |
| 2024 | Meta-Learning Approaches For Improving Detection of Unseen Speech DeepfakesabstractCurrent speech deepfake detection approaches perform satisfactorily against known adversaries; however, generalization to unseen attacks remains an open challenge. The proliferation of speech deepfakes on social media underscores the need for systems that can generalize to unseen attacks not observed during training. We address this problem from the perspective of meta-learning, aiming to learn attack-invariant features to adapt to unseen attacks with very few samples available. This approach is promising since generating of a high-scale training dataset is often expensive or infeasible. Our experiments demonstrated an improvement in the Equal Error Rate (EER) from 21.67% to 10.42% on the InTheWild dataset, using just 96 samples from the unseen dataset. Continuous few-shot adaptation ensures that the system remains up-to-date. Ivan Kukanov, Janne Laakkonen, Tomi Kinnunen, Ville Hautamäki |
SLT | 4 |
| 2023 | Self-Supervised Training of Speaker Encoder With Multi-Modal Diverse Positive PairsabstractWe study a novel neural speaker encoder and its training strategies for speaker recognition without using any identity labels. The speaker encoder is trained to extract a fixed dimensional speaker embedding from a spoken utterance of variable length. Contrastive learning is a typical self-supervised learning technique. However, the contrastive learning of the speaker encoder depends very much on the sampling strategy of positive and negative pairs. It is common that we sample a positive pair of segments from the same utterance. Unfortunately, such a strategy, denoted as poor-man's positive pairs (PPP), lacks the necessary diversity. In this work, we propose a multi-modal contrastive learning technique with novel sampling strategies. By cross-referencing between speech and face data, we find diverse positive pairs (DPP) for contrastive learning, thus improving the robustness of speaker encoder. We train the speaker encoder on the VoxCeleb2 dataset without any speaker labels, and achieve an equal error rate (EER) of 2.89%, 3.17% and 6.27% under the proposed progressive clustering strategy, and an EER of 1.44%, 1.77% and 3.27% under the two-stage learning strategy with pseudo labels, on the three test sets of VoxCeleb1. This novel solution outperforms the state-of-the-art self-supervised learning methods by a large margin, at the same time, achieves comparable results with the supervised learning counterpart. We also evaluate our self-supervised learning technique on the LRS2 and LRW datasets, where speaker information is unavailable. All experiments suggest that the proposed neural architecture and sampling strategies are robust across datasets. Ruijie Tao, Kong-Aik Lee, Rohan Kumar Das, Ville Hautamäki, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | GAN-Aimbots: Using Machine Learning for Cheating in First Person ShootersabstractPlaying games with cheaters is not fun, and in a multi-billion-dollar video game industry with hundreds of millions of players, game developers aim to improve the security and, consequently, the user experience of their games by preventing cheating. Both traditional software-based methods and statistical systems have been successful in protecting against cheating, but recent advances in the automatic generation of content, such as images or speech, threaten the video game industry; they could be used to generate artificial gameplay indistinguishable from that of legitimate human players. To better understand this threat, we begin by reviewing the current state of multiplayer video game cheating, and then proceed to build a proof-of-concept method, GAN-Aimbot. By gathering data from various players in a first-person shooter game we show that the method improves players performance while remaining hidden from automatic and manual protection mechanisms. By sharing this work we hope to raise awareness on this issue and encourage further research into protecting the gaming communities. Anssi Kanervisto, Tomi Kinnunen, Ville Hautamäki |
IEEE Trans. Games | 3 |
| 2022 | Self-Supervised Speaker Recognition with Loss-Gated LearningabstractIn self-supervised learning for speaker recognition, pseudo labels are useful as the supervision signals. It is a known fact that a speaker recognition model doesn’t always benefit from pseudo labels due to their unreliability. In this work, we observe that a speaker recognition network tends to model the data with reliable labels faster than those with unreliable labels. This motivates us to study a loss-gated learning (LGL) strategy, which extracts the reliable labels through the fitting ability of the neural network during training. With the proposed LGL, our speaker recognition model obtains a 46.3% performance gain over the system without it. Further, the proposed self-supervised speaker recognition with LGL trained on the VoxCeleb2 dataset without any labels achieves an equal error rate of 1.66% on the VoxCeleb1 original test set. Ruijie Tao, Kong-Aik Lee, Rohan Kumar Das, Ville Hautamäki, Haizhou Li 0001 |
ICASSP | 4 |
| 2022 | Optimizing Tandem Speaker Verification and Anti-Spoofing SystemsabstractAs automatic speaker verification (ASV) systems are vulnerable to spoofing attacks, they are typically used in conjunction with spoofing countermeasure (CM) systems to improve security. For example, the CM can first determine whether the input is human speech, then the ASV can determine whether this speech matches the speakers identity. The performance of such a tandem system can be measured with a tandem detection cost function (t-DCF). However, ASV and CM systems are usually trained separately, using different metrics and data, which does not optimize their combined performance. In this work, we propose to optimize the tandem system directly by creating a differentiable version of t-DCF and employing techniques from reinforcement learning. The results indicate that these approaches offer better outcomes than finetuning, with our method providing a 20\% relative improvement in the t-DCF in the ASVSpoof19 dataset in a constrained setting. Anssi Kanervisto, Ville Hautamäki, Tomi Kinnunen, Junichi Yamagishi |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Voxceleb Enrichment for Age and Gender RecognitionabstractVoxCeleb datasets are widely used in speaker recognition studies. Our work serves two purposes. First, we provide speaker age labels and (an alternative) annotation of speaker gender. Second, we demonstrate the use of this metadata by constructing age and gender recognition models with different features and classifiers. We query different celebrity databases and apply consensus rules to derive age and gender labels. We also compare the original VoxCeleb gender labels with our labels to identify records that might be mislabeled in the original VoxCeleb data. On modeling side, we design a comprehensive study of multiple features and models for recognizing gender and age. Our best system, using i-vector features, achieved an F1-score of 0.9829 for gender recognition task using logistic regression, and the lowest mean absolute error (MAE) in age regression, 9.443 years, is obtained with ridge regression. This indicates challenge in age estimation from in-the-wild style speech data. Khaled Hechmi, Trung Ngo Trong, Ville Hautamäki, Tomi Kinnunen |
ASRU | 3 |
| 2021 | PL-EESR: Perceptual Loss Based End-to-End Robust Speaker Representation ExtractionabstractSpeech enhancement aims to improve the perceptual quality of the speech signal by suppression of the background noise. However, excessive suppression may lead to speech distortion and speaker information loss, which degrades the performance of speaker embedding extraction. To alleviate this problem, we propose an end-to-end deep learning framework, dubbed PL-EESR, for robust speaker representation extraction. This framework is optimized based on the feedback of the speaker identification task and the high-level perceptual deviation between the raw speech signal and its noisy version. We conducted speaker verification tasks in both noisy and clean environment respectively to evaluate our system. Compared to the baseline, our method shows better performance in both clean and noisy environments, which means our method can not only enhance the speaker relative information but also avoid adding distortions. Kong-Aik Lee, Ville Hautamäki, Haizhou Li 0001 |
ASRU | 3 |
| 2021 | Distilling Reinforcement Learning Tricks for Video GamesabstractReinforcement learning (RL) research focuses on general solutions that can be applied across different domains. This results in methods that RL practitioners can use in almost any domain. However, recent studies often lack the engineering steps (“tricks”) which may be needed to effectively use RL, such as reward shaping, curriculum learning, and splitting a large task into smaller chunks. Such tricks are common, if not necessary, to achieve state-of-the-art results and win RL competitions. To ease the engineering efforts, we distill descriptions of tricks from state-of-the-art results and study how well these tricks can improve a standard deep Q-learning agent. The long-term goal of this work is to enable combining proven RL methods with domain-specific tricks by providing a unified software framework and accompanying insights in multiple domains. Anssi Kanervisto, Christian Scheller, Yanick Schraner, Ville Hautamäki |
CoG | 4 |
| 2020 | Benchmarking End-to-End Behavioural Cloning on Video GamesabstractBehavioural cloning, where a computer is taught to perform a task based on demonstrations, has been successfully applied to various video games and robotics tasks, with and without reinforcement learning. This also includes end-to-end approaches, where a computer plays a video game like humans do: by looking at the image displayed on the screen, and sending keystrokes to the game. As a general approach to playing video games, this has many inviting properties: no need for specialized modifications to the game, no lengthy training sessions and the ability to re-use the same tools across different games. However, related work includes game-specific engineering to achieve the results. We take a step towards a general approach and study the general applicability of behavioural cloning on twelve video games, including six modern video games (published after 2010), by using human demonstrations as training data. Our results show that these agents cannot match humans in raw performance but do learn basic dynamics and rules. We also demonstrate how the quality of the data matters, and how recording data from humans is subject to a state-action mismatch, due to human reflexes. Anssi Kanervisto, Joonas Pussinen, Ville Hautamäki |
CoG | 3 |
| 2020 | Action Space Shaping in Deep Reinforcement LearningabstractReinforcement learning (RL) has been successful in training agents in various learning environments, including video games. However, such work modifies and shrinks the action space from the game's original. This is to avoid trying "pointless" actions and to ease the implementation. Currently, this is mostly done based on intuition, with little systematic research supporting the design decisions. In this work, we aim to gain insight on these action space modifications by conducting extensive experiments in video game environments. Our results show how domain-specific removal of actions and discretization of continuous actions can be crucial for successful learning. With these insights, we hope to ease the use of RL in new environments, by clarifying what action-spaces are easy to learn. Anssi Kanervisto, Christian Scheller, Ville Hautamäki |
CoG | 3 |
| 2020 | From Video Game to Real Robot: The Transfer Between Action SpacesabstractDeep reinforcement learning has proven to be successful for learning tasks in simulated environments, but applying same techniques for robots in real-world domain is more challenging, as they require hours of training. To address this, transfer learning can be used to train the policy first in a simulated environment and then transfer it to physical agent. As the simulation never matches reality perfectly, the physics, visuals and action spaces by necessity differ between these environments to some degree. In this work, we study how general video games can be directly used instead of fine-tuned simulations for the sim-to-real transfer. Especially, we study how the agent can learn the new action space autonomously, when the game actions do not match the robot actions. Our results show that the different action space can be learned by re-training only part of neural network and we obtain above 90% mean success rate in simulation and robot experiments. Janne Karttunen, Anssi Kanervisto, Ville Kyrki, Ville Hautamäki |
ICASSP | 4 |
| 2020 | Maximal Figure-of-Merit Framework to Detect Multi-Label Phonetic Features for Spoken Language RecognitionabstractBottleneck features (BNFs) generated with a deep neural network (DNN) have proven to boost spoken language recognition accuracy over basic spectral features significantly. However, BNFs are commonly extracted using language-dependent tied-context phone states as learning targets. Moreover, BNFs are less phonetically expressive than the output layer in a DNN, which is usually not used as a speech feature because of its very high dimensionality hindering further post-processing. In this article, we put forth a novel deep learning framework to overcome all of the above issues and evaluate it on the 2017 NIST Language Recognition Evaluation (LRE) challenge. We use manner and place of articulation as speech attributes, which lead to low-dimensional “universal” phonetic features that can be defined across all spoken languages. To model the asynchronous nature of the speech attributes while capturing their intrinsic relationships in a given speech segment, we introduce a new training scheme for deep architectures based on a Maximal Figure of Merit (MFoM) objective. MFoM introduces non-differentiable metrics into the backpropagation-based approach, which is elegantly solved in the proposed framework. The experimental evidence collected on the recent NIST LRE 2017 challenge demonstrates the effectiveness of our solution. In fact, the performance of speech language recognition (SLR) systems based on spectral features is improved for more than 5% absolute Cavg. Finally, the F1 metric can be brought from 77.6% up to 78.1% by combining the conventional baseline phonetic BNFs with the proposed articulatory attribute features. Ivan Kukanov, Trung Ngo Trong, Ville Hautamäki, Sabato Marco Siniscalchi, Valerio Mario Salerno, Kong-Aik Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | ToriLLE: Learning Environment for Hand-to-Hand CombatabstractWe present Toribash Learning Environment (To-riLLE), a learning environment for machine learning agents based on the video game Toribash. Toribash is a MuJoCo-like environment of two humanoid characters fighting each other hand-to-hand, controlled by changing actuation modes of the joints. Competitive nature of Toribash as well its focused domain provide a platform for evaluating self-play methods, and evaluating machine learning agents against human players. In this paper we describe the environment with ToriLLE's capabilities and limitations, and experimentally show its applicability as a learning environment with baseline and human experiments. The source code of the environment and conducted experiments can be found at https://github.com/Miffyli/ToriLLE. Anssi Kanervisto, Ville Hautamäki |
CoG | 2 |
| 2019 | Who Do I Sound like? Showcasing Speaker Recognition Technology by Youtube Voice SearchabstractThe popularization of science can often be disregarded by scientists as it may be challenging to put highly sophisticated research into words that general public can understand. This work aims to help presenting speaker recognition research to public by proposing a publicly appealing concept for showcasing recognition systems. We leverage data from YouTube and use it in a large-scale voice search web application that finds the celebrity voices that best match to the user's voice. The concept was tested in a public event as well as "in the wild" and the received feedback was mostly positive. The i-vector based speaker identification back end was found to be fast (665 ms per request) and had a high identification accuracy (93%) for the YouTube target speakers. To help other researchers to develop the idea further, we share the source codes of the web platform used for the demo at https://github.com/bilalsoomro/speech-demo-platform. Ville Vestman, Bilal Soomro, Anssi Kanervisto, Ville Hautamäki, Tomi Kinnunen |
ICASSP | 4 |
| 2019 | I4U Submission to NIST SRE 2018: Leveraging from a Decade of Shared ExperiencesabstractThe I4U consortium was established to facilitate a joint entry to NIST speaker recognition evaluations (SRE). The latest edition of such joint submission was in SRE 2018, in which the I4U submission was among the best-performing systems. SRE'18 also marks the 10-year anniversary of I4U consortium into NIST SRE series of evaluation. The primary objective of the current paper is to summarize the results and lessons learned based on the twelve sub-systems and their fusion submitted to SRE'18. It is also our intention to present a shared view on the advancements, progresses, and major paradigm shifts that we have witnessed as an SRE participant in the past decade from SRE'08 to SRE'18. In this regard, we have seen, among others, a paradigm shift from supervector representation to deep speaker embedding, and a switch of research challenge from channel compensation to domain adaptation. Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Hitoshi Yamamoto, Koji Okabe, Ville Vestman, Jing Huang 0019, Guo-Hong Ding, Hanwu Sun, Anthony Larcher, Rohan Kumar Das, Haizhou Li 0001, Mickael Rouvier, Pierre-Michel Bousquet, Wei Rao 0002, Qing Wang 0039, Fahimeh Bahmaninezhad, Héctor Delgado, Massimiliano Todisco |
INTERSPEECH | 2 |
| 2019 | Towards Debugging Deep Neural Networks by Generating Speech UtterancesabstractDeep neural networks (DNN) are able to successfully process and classify speech utterances. However, understanding the reason behind a classification by DNN is difficult. One such debugging method used with image classification DNNs is activation maximization, which generates example-images that are classified as one of the classes. In this work, we evaluate applicability of this method to speech utterance classifiers as the means to understanding what DNN "listens to". We trained a classifier using the speech command corpus and then use activation maximization to pull samples from the trained model. Then we synthesize audio from features using WaveNet vocoder for subjective analysis. We measure the quality of generated samples by objective measurements and crowd-sourced human evaluations. Results show that when combined with the prior of natural speech, activation maximization can be used to generate examples of different classes. Based on these results, activation maximization can be used to start opening up the DNN black-box in speech tasks. Bilal Soomro, Anssi Kanervisto, Trung Ngo Trong, Ville Hautamäki |
INTERSPEECH | 4 |
| 2018 | Maximal Figure-of-Merit Embedding for Multi-Label Audio ClassificationabstractThis work tackles the problem of the domestic audio tagging or environmental sound classification, where one audio recording can contain one or more acoustic events and a recognizer should output all of those tags. A baseline model for this task is a convolutional recurrent neural network (CRNN) with sigmoid output nodes optimized using the binary cross-entropy objective. Traditional error metrics, such as classification error, are not suitable for this type of task. In this work, we show that the maximal figure-of-merit (MFoM) framework helps to separate the multi-label classes in terms of equal error rate (EER). We embed MFoM into the deep learning objective function and gain more than 9% relative improvement, compared to the baseline model with binary cross-entropy. Ivan Kukanov, Ville Hautamäki, Kong-Aik Lee |
ICASSP | 2 |
| 2017 | Effects of gender information in text-independent and text-dependent speaker verificationabstractIt is well-known that for speaker recognition task, gender-dependent acoustic modeling performs better than gender-independent modeling. The practice is to use the gender ground-truth and to train gender-dependent models. However, such information is not necessarily available, especially if speakers are remotely enrolled. A way to overcome this is to use a gender classification system, which introduces an additional layer of uncertainty. To date, such uncertainty has not been studied. We implement two gender classifier systems and test them with two different corpora and speaker verification systems. We find that estimated gender information can improve speaker verification accuracy over gender-independent methods. Our detailed analysis suggests that gender estimation should have a sufficiently high accuracy to yield improvements in speaker verification performance. Anssi Kanervisto, Ville Vestman, Md. Sahidullah, Ville Hautamäki, Tomi Kinnunen |
ICASSP | 4 |
| 2017 | RedDots replayed: A new replay spoofing attack corpus for text-dependent speaker verification researchabstractThis paper describes a new database for the assessment of automatic speaker verification (ASV) vulnerabilities to spoofing attacks. In contrast to other recent data collection efforts, the new database has been designed to support the development of replay spoofing countermeasures tailored towards the protection of text-dependent ASV systems from replay attacks in the face of variable recording and playback conditions. Derived from the re-recording of the original RedDots database, the effort is aligned with that in text-dependent ASV and thus well positioned for future assessments of replay spoofing countermeasures, not just in isolation, but in integration with ASV. The paper describes the database design and re-recording, a protocol and some early spoofing detection results. The new “RedDots Replayed” database is publicly available through a creative commons license. Tomi Kinnunen, Md. Sahidullah, Mauro Falcone, Luca Costantini, Rosa González Hautamäki, Dennis Alexander Lehmann Thomsen, Achintya Kumar Sarkar, Zheng-Hua Tan, Héctor Delgado, Massimiliano Todisco, Nicholas W. D. Evans, Ville Hautamäki, Kong-Aik Lee |
ICASSP | 12 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 2 |
| 2017 | Acoustical and perceptual study of voice disguise by age modification in speaker verification
Rosa González Hautamäki, Md. Sahidullah, Ville Hautamäki, Tomi Kinnunen |
Speech Commun. | 3 |
| 2016 | Variation in Spoken North Sami Language
Kristiina Jokinen, Trung Ngo Trong, Ville Hautamäki |
INTERSPEECH | 3 |
| 2016 | Utterance Verification for Text-Dependent Speaker Recognition: A Comparative Assessment Using the RedDots CorpusabstractText-dependent automatic speaker verification naturally calls for the simultaneous verification of speaker identity and spoken content. These two tasks can be achieved with automatic speaker verification (ASV) and utterance verification (UV) technologies. While both have been addressed previously in the literature, a treatment of simultaneous speaker and utterance verification with a modern, standard database is so far lacking. This is despite the burgeoning demand for voice biometrics in a plethora of practical security applications. With the goal of improving overall verification performance, this paper reports different strategies for simultaneous ASV and UV in the context of short-duration, text-dependent speaker verification. Experiments performed on the recently released RedDots corpus are reported for three different ASV systems and four different UV systems. Results show that the combination of utterance verification with automatic speaker verification is (almost) universally beneficial with significant performance improvements being observed. Tomi Kinnunen, Md. Sahidullah, Ivan Kukanov, Héctor Delgado, Massimiliano Todisco, Achintya Kumar Sarkar, Nicolai Bæk Thomsen, Ville Hautamäki, Nicholas W. D. Evans, Zheng-Hua Tan |
INTERSPEECH | 8 |
| 2016 | The 2015 NIST Language Recognition Evaluation: The Shared View of I2R, Fantastic4 and SingaMSabstractTechnical report for NIST LRE 2015 Workshop Kong-Aik Lee, Haizhou Li 0001, Li Deng 0001, Ville Hautamäki, Wei Rao 0002, Anthony Larcher, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Aleksandr Sizov, Jianshu Chen, Ivan Kukanov, Amir Hossein Poorjam, Trung Ngo Trong, Chenglin Xu, Haihua Xu 0001, Bin Ma 0001, Chng Eng Siong, Sylvain Meignier |
INTERSPEECH | 4 |
| 2016 | Robust Speaker Recognition with Combined Use of Acoustic and Throat Microphone SpeechabstractAccuracy of automatic speaker recognition (ASV) systems degrades severely in the presence of background noise. In this paper, we study the use of additional side information provided by a body-conducted sensor, throat microphone. Throat microphone signal is much less affected by background noise in comparison to acoustic microphone signal. This makes throat microphones potentially useful for feature extraction or speech activity detection. This paper, firstly, proposes a new prototype system for simultaneous data-acquisition of acoustic and throat microphone signals. Secondly, we study the use of this additional information for both speech activity detection, feature extraction and fusion of the acoustic and throat microphone signals. We collect a pilot database consisting of 38 subjects including both clean and noisy sessions. We carry out speaker verification experiments using Gaussian mixture model with universal background model (GMM-UBM) and i-vector based system. We have achieved considerable improvement in recognition accuracy even in highly degraded conditions. Md. Sahidullah, Rosa González Hautamäki, Dennis Alexander Lehmann Thomsen, Tomi Kinnunen, Zheng-Hua Tan, Ville Hautamäki, Robert Parts, Martti Pitkänen |
INTERSPEECH | 6 |
| 2016 | Deep learning with maximal figure-of-merit cost to advance multi-label speech attribute detectionabstractIn this work, we are interested in boosting speech attribute detection by formulating it as a multi-label classification task, and deep neural networks (DNNs) are used to design speech attribute detectors. A straightforward way to tackle the speech attribute detection task is to estimate DNN parameters using the mean squared error (MSE) loss function and employ a sigmoid function in the DNN output nodes. A more principled way is nonetheless to incorporate the micro-F1 measure, which is a widely used metric in the multi-label classification, into the DNN loss function to directly improve the metric of interest at training time. Micro-F1 is not differentiable, yet we overcome such a problem by casting our task under the maximal figure-of-merit (MFoM) learning framework. The results demonstrate that our MFoM approach consistently outperforms the baseline systems. Ivan Kukanov, Ville Hautamäki, Sabato Marco Siniscalchi, Kehuang Li |
SLT | 2 |
| 2016 | i-Vector Modeling of Speech Attributes for Automatic Foreign Accent RecognitionabstractWe propose a unified approach to automatic foreign accent recognition. It takes advantage of recent technology advances in both linguistics and acoustics based modeling techniques in automatic speech recognition (ASR) while overcoming the issue of a lack of a large set of transcribed data often required in designing state-of-the-art ASR systems. The key idea lies in defining a common set of fundamental units “universally” across all spoken accents such that any given spoken utterance can be transcribed with this set of “accent-universal” units. In this study, we adopt a set of units describing manner and place of articulation as speech attributes. These units exist in most spoken languages and they can be reliably modeled and extracted to represent foreign accent cues. We also propose an i-vector representation strategy to model the feature streams formed by concatenating these units. Testing on both the Finnish national foreign language certificate (FSD) corpus and the English NIST 2008 SRE corpus, the experimental results with the proposed approach demonstrate a significant system performance improvement with p-value over those with the conventional spectrum-based techniques. We observed up to a 15% relative error reduction over the already very strong i-vector accented recognition system when only manner information is used. Additional improvement is obtained by adding place of articulation clues along with context information. Furthermore, diagnostic information provided by the proposed approach can be useful to the designers to further enhance the system performance. Hamid Behravan, Ville Hautamäki, Sabato Marco Siniscalchi, Tomi Kinnunen, Chin-Hui Lee 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Boosting universal speech attributes classification with deep neural network for foreign accent characterizationabstractWe have recently proposed a universal acoustic characterisation to foreign accent recognition, in which any spoken foreign accent was described in terms of a common set of fundamental speech attributes. Although experimental evidence demonstrated the feasibility of our approach, we belive that speech attributes, namely manner and place of articulation, can be better modelled by a deep neural network. In this work, we propose the use of deep neural network trained on telephone bandwidth material from different languages to improve the proposed universal acoustic characterisation. We demonstrate that deeper neural architectures enhance the attribute classification accuracy. Furthermore, we show that improvements in attribute classification carry over to foreign accent recognition by producing a 21% relative improvement over previous baseline on spoken Finnish, and a 5.8% relative improvement on spoken English Ville Hautamäki, Sabato Marco Siniscalchi, Hamid Behravan, Valerio Mario Salerno, Ivan Kukanov |
INTERSPEECH | 1 |
| 2015 | Factors affecting i-vector based foreign accent recognition: A case study in spoken Finnish
Hamid Behravan, Ville Hautamäki, Tomi Kinnunen |
Speech Commun. | 2 |
| 2015 | Automatic versus human speaker verification: The case of voice mimicry
Rosa González Hautamäki, Tomi Kinnunen, Ville Hautamäki, Anne-Maria Laukkanen |
Speech Commun. | 3 |
| 2014 | Introducing attribute features to foreign accent recognitionabstractWe propose a hybrid approach to foreign accent recognition combining both phonotactic and spectral based systems by treating the problem as a spoken language recognition task. We extract speech attribute features that represent speech and acoustic cues reflecting foreign accents of a speaker to obtain feature streams that are modeled with the i-vector methodology. Testing on the Finnish Language Proficiency exam corpus, we find our proposed technique to achieve a significant performance improvement over the state-of-the-art systems using only spectral based features. Hamid Behravan, Ville Hautamäki, Sabato Marco Siniscalchi, Tomi Kinnunen, Chin-Hui Lee 0001 |
ICASSP | 2 |
| 2014 | An i-vector based descriptor for alphabetical gesture recognitionabstractAn i-vector approach to extracting features for video camera based gesture recognition is proposed. Conventional low-level raw features, such as position, speed, and acceleration, are low-dimensional feature representations which often suffer from measurement noise and thus are not highly discriminative. High-level features, such as Fourier descriptor, usually take a global transformation on the whole raw features of a gesture, but local statistical information is seldom considered. Moreover, compared with speech recordings, video cameras used to capture data are often at a low frame rate such that it is challenging for proper modeling and recognition. In this paper, we show that the proposed i-vector framework can handle both local statistical information and sparse trajectory representations more efficiently under the sparse data scenarios for an in-car hand-gesturing English letter recognition system. Experimental results confirm the effectiveness of the proposed i-vector features, which can reduce the letter error rate by as much as 36-44% relatively from the results obtained with the conventional location based raw features. You-Chi Cheng, Ville Hautamäki, Zhen Huang 0001, Kehuang Li, Chin-Hui Lee 0001 |
ICASSP | 2 |
| 2014 | Dialect levelling in Finnish: a universal speech attribute approachabstractWe adopt automatic language recognition methods to study dialect levelling - a phenomenon that leads to reduced structural differences among dialects in a given spoken language. In terms of dialect characterisation, levelling is a nuisance variable that adversely affects recognition accuracy: The more similar two dialects are, the harder it is to set them apart. We address levelling in Finnish regional dialects using a new SAPU (Satakunta in Speech) corpus containing material from Satakunta (South-Western Finland) between 2007 and 2013. To define a compact and universal set of sound units to characterize dialects, we adopt speech attributes features, namely manner and place of articulation. It will be shown that speech attribute distributions can indeed characterise differences among dialects. Experiments with an i-vector system suggest that (1) the attribute features achieve higher dialect recognition accuracy and (2) they are less sensitive against age-related levelling in comparison to traditional spectral approach. Hamid Behravan, Ville Hautamäki, Sabato Marco Siniscalchi, Elie Khoury 0001, Tommi Kurki, Tomi Kinnunen, Chin-Hui Lee 0001 |
INTERSPEECH | 2 |
| 2013 | Foreign accent detection from spoken Finnish using i-vectorsabstractI-vector based recognition is a well-established technique in state-of-the-art speaker and language recognition but its use in dialect and accent classification has received less attention. We represent an experimental study of i-vector based dialect classi-fication, with a special focus on foreign accent detection from spoken Finnish. Using the CallFriend corpus, we first study how recognition accuracy is affected by the choices of vari-ous i-vector system parameters, such as the number of Gaus-sians, i-vector dimensionality and reduction method. We then apply the same methods on the Finnish national foreign lan-guage certificate (FSD) corpus and compare the results to tra-ditional Gaussian mixture model- universal background model (GMM-UBM) recognizer. The results, in terms of equal error rate, indicate that i-vectors outperform GMM-UBM as one ex-pects. We also notice that in foreign accent detection, 7 out of 9 accents were more accurately detected by Gaussian scoring than by cosine scoring. Index Terms: Dialect recognition, foreign accent recognition, i-vector, GMM-UBM, Finnish language Hamid Behravan, Ville Hautamäki, Tomi Kinnunen |
INTERSPEECH | 2 |
| 2013 | Minimax i-vector extractor for short duration speaker verificationabstractTotal variability modeling, based on i-vector extraction of converting a variable-length sequence of feature vectors into a fixed-length i-vector, is currently an adopted parametrization technique for state of-the-art speaker verification systems. However, when the number of the feature vectors is low, uncertainty in the i-vector representation as a point estimate of the linear-Gaussian model is understandably problematic. It is known that the zeroth and first order sufficient statistics, given the hyperparameters, completely characterize the extracted i-vectors. In this study we propose to use a minimax strategy to estimate the sufficient statistics in order to increase the robustness of the extracted i-vectors. We show by experiments that the proposed minimax technique can improve over the baseline system from 9.89 % to 7.99 % on the NIST SRE 2010 8conv-10sec task. Ville Hautamäki, You-Chi Cheng, Padmanabhan Rajan, Chin-Hui Lee 0001 |
INTERSPEECH | 1 |
| 2013 | Merging human and automatic system decisions to improve speaker recognition performanceabstractHuman judgment is the final authority in forensic speaker recognition, but the use of modern speaker verification systems with accurate algorithms to perform the task under various circumstances has a huge potential to help the expert. The ultimate goal is to improve the accuracy of automatic systems when challenging data is provided and find a methodology for human-aided speaker recognition systems. This work presents an evaluation of speaker recognition carried out by human listeners and a gender dependent i-vector recognizer with a strategy for fusion of the decision process. Our experiments with HASR 2010 and HASR 2012 data indicate complementarity in Rosa González Hautamäki, Ville Hautamäki, Padmanabhan Rajan, Tomi Kinnunen |
INTERSPEECH | 2 |
| 2013 | I-vectors meet imitators: on vulnerability of speaker verification systems against voice mimicryabstractVoice imitation is mimicry of another speaker’s voice characteristics and speech behavior. Professional voice mimicry can create entertaining, yet realistic sounding target speaker renditions. As mimicry tends to exaggerate prosodic, idiosyncratic and lexical behavior, it is unclear how modern spectral-feature automatic speaker verification systems respond to mimicry “attacks”. We study the vulnerability of two well-known speaker recognition systems, traditional Gaussian mixture model – universal background model (GMM-UBM) and a state-of-the-art i-vector classifier with cosine scoring. The material consists of one professional Finnish imitator impersonating five wellknown Finnish public figures. In a carefully controlled setting, mimicry attack does slightly increase the false acceptance rate for the i-vector system, but generally this is not alarmingly large in comparison to voice conversion or playback attacks. Index Terms: Voice imitation, speaker recognition, mimicry attack Rosa González Hautamäki, Tomi Kinnunen, Ville Hautamäki, Timo Leino, Anne-Maria Laukkanen |
INTERSPEECH | 3 |
| 2013 | Automatic regularization of cross-entropy cost for speaker recognition fusionabstract\n Contains fulltext :\n 116325.pdf (author's version ) (Open Access)\n Ville Hautamäki, Kong-Aik Lee, David A. van Leeuwen, Rahim Saeidi, Anthony Larcher, Tomi Kinnunen, Taufiq Hasan, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, John H. L. Hansen, Benoit G. B. Fauve |
INTERSPEECH | 1 |
| 2013 | A blind segmentation approach to acoustic event detection based on i-vectorabstractWe propose a new blind segmentation approach to acous-tic event detection (AED) based on i-vectors. Conventional approaches to AED often required well-segmented data with non-overlapping boundaries for competing events. Inspired by block-based automatic image annotation in image retrieval tasks, we blindly segment audio streams into equal-length pieces, label the underlying observed acoustic events with mul-tiple categories and with no event boundary information, extract i-vector for them, and perform classification using support vec-tor machine and maximal figure-of-merit based classifiers. Ex-periments on various sets of audio data show promising results with an average of 8 % absolute gain in F1 over the conventional hidden Markov model based approach. An enhanced robustness at different noise levels is also observed. The key to the suc-cess lies in the enhanced discrimination power offered by the i-vector representation of the acoustic data. Index Terms: acoustic event detection, i-vector, blind segmen-tation, support vector machine, maximal figure-of-merit Zhen Huang 0001, You-Chi Cheng, Kehuang Li, Ville Hautamäki, Chin-Hui Lee 0001 |
INTERSPEECH | 4 |
| 2013 | Effect of multicondition training on i-vector PLDA configurations for speaker recognitionabstractThe i-vector representation and PLDA classifier have shown state-of-the-art performance for speaker recognition systems. The availability of more than one enrollment utterance for a speaker allows a variety of configurations which can be used to enhance robustness to noise. The well-known technique of multicondition training can be utilized at different stages of the system, including enrollment and classifier training. We also study the effect of mismatched training, averaging and length normalization. Our study indicates that multicondition training of the PLDA model, and if possible the enrollment i-vectors are the most important to achieve good performance in noisy evaluation data. Index Terms: Speaker verification, i-vector, PLDA, multicondition training Padmanabhan Rajan, Tomi Kinnunen, Ville Hautamäki |
INTERSPEECH | 3 |
| 2013 | I4u submission to NIST SRE 2012: a large-scale collaborative effort for noise-robust speaker verificationabstractI4U is a joint entry of nine research Institutes and Universities across 4 continents to NIST SRE 2012. It started with a brief discussion during the Odyssey 2012 workshop in Singapore. An online discussion group was soon set up, providing a discussion platform for different issues surrounding NIST SRE’12. Noisy test segments, uneven multi-session training, variable enrollment duration, and the issue of open-set identification were actively discussed leading to various solutions integrated to the I4U submission. The joint submission and several of its 17 sub-systems were among top-performing systems. We summarize the lessons learnt from this large-scale effort. Rahim Saeidi, Kong-Aik Lee, Tomi Kinnunen, Tawfik Hasan, Benoit G. B. Fauve, Pierre-Michel Bousquet, Elie Khoury 0001, Pablo Luis Sordo Martinez, Jia Min Karen Kua, Chang Huai You, Hanwu Sun, Anthony Larcher, Padmanabhan Rajan, Ville Hautamäki, Cemal Hanilçi, Billy Braithwaite, Rosa González Hautamäki, Seyed Omid Sadjadi, Gang Liu 0001, Hynek Boril, Navid Shokouhi, Driss Matrouf, Laurent El Shafey, Pejman Mowlaee, Julien Epps, Tharmarajah Thiruvaran, David A. van Leeuwen, Bin Ma 0001, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Sébastien Marcel, John S. D. Mason, Eliathamby Ambikairajah |
INTERSPEECH | 14 |
| 2013 | Sparse Classifier Fusion for Speaker VerificationabstractState-of-the-art speaker verification systems take advantage of a number of complementary base classifiers by fusing them to arrive at reliable verification decisions. In speaker verification, fusion is typically implemented as a weighted linear combination of the base classifier scores, where the combination weights are estimated using a logistic regression model. An alternative way for fusion is to use classifier ensemble selection, which can be seen as sparse regularization applied to logistic regression. Even though score fusion has been extensively studied in speaker verification, classifier ensemble selection is much less studied. In this study, we extensively study a sparse classifier fusion on a collection of twelve I4U spectral subsystems on the NIST 2008 and 2010 speaker recognition evaluation (SRE) corpora. Ville Hautamäki, Tomi Kinnunen, Filip Sedlak, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Random swap EM algorithm for Gaussian mixture models
Qinpei Zhao, Ville Hautamäki, Ismo Kärkkäinen, Pasi Fränti |
Pattern Recognit. Lett. | 2 |
| 2011 | Classifier subset selection and fusion for speaker verificationabstractState-of-the-art speaker verification systems consists of a number of complementary subsystems whose outputs are fused, to arrive at more accurate and reliable verification decision. In speaker verification, fusion is typically implemented as a linear combination of the subsystem scores. Parameters of the linear model are commonly estimated using the logistic regression method, as implemented in the popular FoCal toolkit. In this paper, we study simultaneous use of classifier selection and fusion. We study four alternative fusion strategies, three score warping techniques, and provide interesting experimental bounds on optimal classifier subset selection. Detailed experiments are carried out on the NIST 2008 and 2010 SRE corpora. Filip Sedlak, Tomi Kinnunen, Ville Hautamäki, Kong-Aik Lee, Haizhou Li 0001 |
ICASSP | 3 |
| 2011 | RSEM: An Accelerated Algorithm on Repeated EMabstractExpectation maximization (EM) algorithm, being a gradient ascent algorithm depends highly on the initialization. Repeating EM multiple times with different initial solutions and taking the best result is used to attack this problem. However, the solution space is searched inefficiently in Repeated EM, because after each restart it can take a long time to converge without any guarantee that it leads to an improved solution. A random swap EM algorithm utilizes random swap strategy to improve the problem in a more efficient way. In this paper, a theoretical and experimental comparison between RSEM and REM is conducted. Based on GMM estimation theory, it is proved that RSEM reaches the optimal result faster than REM with high probability. It is also shown experimentally that RSEM speeds up REM from 9% to 63%. A study in color-texture images demonstrates an application of EM algorithms in a segmentation task. Qinpei Zhao, Ville Hautamäki, Pasi Fränti |
ICIG | 2 |
| 2011 | Regularized Logistic Regression Fusion for Speaker VerificationabstractFusion of the base classifiers is seen as the way to achieve stateof-the art performance in the speaker verfication systems. Standard approach is to pose the fusion problem as the linear binary classification task. Most successful loss function in speaker verification fusion has been the weighted logistic regression popularized by the FoCal toolkit. However, it is known that optimizing logistic regression can overfit severely without appropriate regularization. In addition, subset classifier selection can be achieved by using an external 0/1 loss function on the best subset. In this work, we propose to use LASSO based regularization on the FoCal cost function to achive improved performance and classifier subset selection method integrated into one optimization task. Proposed method is able to achieve 51 % relative improvement in Actual DCF over the FoCal baseline. Index Terms: logistic regression, regularization, compressed sensing, linear fusion, speaker verification Ville Hautamäki, Kong-Aik Lee, Tomi Kinnunen, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2011 | Spoken Language Recognition in the Latent Topic SimplexabstractInternational audience Kong-Aik Lee, Chang Huai You, Ville Hautamäki, Anthony Larcher, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2010 | Location-based search engine for multimedia phonesabstractLocation-based search engine is an alternative approach for information retrieval to traditional location-based services based on fixed databases. This is a relatively new concept that aims at utilizing the location of user but without restricting to any fixed location-based service. In this paper, we outline a prototype solution for multimedia mobile phones based on web search, ad-hoc georeferencing, prefix tree structure and gazetteer. Experimental results show that the proposed solution finds search results that have higher or equal mean relevance than that of the GoogleMaps and YellowPages. Pasi Fränti, Andrei Tabarcea, Juha Kuittinen, Ville Hautamäki |
ICME | 4 |
| 2010 | Approaching human listener accuracy with modern speaker verificationabstractBeing able to recognize people from their voice is a natural ability that we take for granted. Recent advances have shown significant improvement in automatic speaker recognition performance. Besides being able to process large amount of data in a fraction of time required by human, automatic systems are now able to deal with diverse channel effects. The goal of this paper is to examine how state-of-the-art automatic system performs in comparison with human listeners, and to investigate the strategy for human-assisted form of automatic speaker recognition, which is useful in forensic investigation. We set up an experimental protocol using data from the NIST SRE 2008 core set. A total of 36 listeners have participated in the listening experiments from three sites, namely Australia, Finland and Singapore. State-of-the-art automatic system achieved 20 % error rate, whereas fusion of human listeners achieved 22%. 1. Ville Hautamäki, Tomi Kinnunen, Mohaddeseh Nosratighods, Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2010 | Towards long-range prosodic attribute modeling for language recognitionabstractAs a high-level feature, prosody may be an effective feature when it is modeled over longer ranges than the typical range of a syllable. This paper is about language recognition with the high-level prosodic attributes. It studies two important issues of long-range modeling, namely the data scarcity handling method, and the model which properly describes prosodic boundary events. Illustrated by NIST language recognition evaluation (LRE) 2009, long-range modeling is shown to bring a 7.2% relative improvement to a prosodic language detector. Score fusion between the long-range prosodic system and a phonotactic system gives an EER of 3.07%. Exploiting boundary N -grams is the main contributing factor to global EER reduction, while different long-range prosodic modeling factors benefit the detection of different languages. Analysis reveals the evidence of language-specific long-range prosodic attributes, which sheds light on robust long-range modeling methods for language recognition. Index Terms: language recognition, prosody, long-range modeling Raymond W. M. Ng, Cheung-Chi Leung, Ville Hautamäki, Tan Lee, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2010 | Ad-hoc Georeferencing of Web-pages using Street-name Prefix Trees
Andrei Tabarcea, Ville Hautamäki, Pasi Fränti |
WEBIST (1) | 2 |
| 2009 | Comparing maximum a posteriori vector quantization and Gaussian mixture models in speaker verificationabstractGaussian mixture model - universal background model (GMM-UBM) is a standard reference classifier in speaker verification. We have proposed a simplified model using vector quantization (VQ-UBM). In this study, we extensively compare these two classifiers on NIST 2005, 2006 and 2008 SRE corpora, while having a standard discriminative classifier (GLDS-SVM) as a reference point. We focus on parameter setting for N-top scoring, model order, and performance for different amounts of training data. The most interesting result, against a general belief, is that GMM-UBM yields better results for short segments whereas VQ-UBM is good for long utterances. The results also suggest that maximum likelihood training of the UBM is sub-optimal, and hence, alternative ways to train the UBM should be considered. Tomi Kinnunen, Juhani Saastamoinen, Ville Hautamäki, Mikko Vinni, Pasi Fränti |
ICASSP | 3 |
| 2009 | Random swap EM algorithm for finite mixture models in image segmentationabstractThe expectation-maximization (EM) algorithm is a popular tool in estimating model parameters, especially mixture models. As the EM algorithm is a hill-climbing approach, problems such as local maxima, plateau and ridges may appear. In the case of mixture models, these problems involve the initialization of the algorithm and the structure of the data set. We propose a random swap EM algorithm (RSEM) to overcome these problems in Gaussian mixture models. Random swaps are repeatedly performed in our method, which can break the configuration of the local maxima and other problems. Compared to the strategies in other methods, the proposed algorithm has relative improvements on log-likelihood value in most cases and less variance than other algorithms. We also apply RSEM to the image segmentation problem. Qinpei Zhao, Ville Hautamäki, Ismo Kärkkäinen, Pasi Fränti |
ICIP | 2 |
| 2009 | Comparative evaluation of maximum a Posteriori vector quantization and gaussian mixture models in speaker verification
Tomi Kinnunen, Juhani Saastamoinen, Ville Hautamäki, Mikko Vinni, Pasi Fränti |
Pattern Recognit. Lett. | 3 |
| 2008 | Knee Point Detection in BIC for Detecting the Number of Clusters
Qinpei Zhao, Ville Hautamäki, Pasi Fränti |
ACIVS | 2 |
| 2008 | Probabilistic clustering by random swap algorithmabstractWe formulate probabilistic clustering method based on a sequence of random swaps of cluster centroids. We show that the algorithm has linear dependency on the number of data vectors, quadratic on the number of clusters, and inverse dependency on the dimensionality. Each halving of the probability of failure (e.g. from 1% to 0.5%) is achieved at the cost of only linear increase in the processing time. Pasi Fränti, Olli Virmajoki, Ville Hautamäki |
ICPR | 3 |
| 2008 | Time-series clustering by approximate prototypesabstractClustering time-series data poses problems, which do not exist in traditional clustering in Euclidean space. Specifically, cluster prototype needs to be calculated, where common solution is to use cluster medoid. In this work, we define an optimal prototype as an optimization problem and propose a local search solution to it. We experimentally compare different time-series clustering methods and find out that the proposed prototype with agglomerative clustering followed by k-means algorithm provides best clustering accuracy. Ville Hautamäki, Pekka Nykänen, Pasi Fränti |
ICPR | 1 |
| 2008 | Text-independent speaker recognition using graph matching
Ville Hautamäki, Tomi Kinnunen, Pasi Fränti |
Pattern Recognit. Lett. | 1 |
| 2008 | Maximum a Posteriori Adaptation of the Centroid Model for Speaker VerificationabstractMaximum a posteriori adapted Gaussian mixture model (GMM-MAP) is widely used in speaker verification. GMMs have three sets of parameters to be adapted: means, covariances, and weights. However, practice has shown that it is sufficient to adapt the means only. Motivated by this, we formulate maximum a posteriori vector quantization (VQ-MAP) procedure which stores and adapts the mean vectors (centroids) only. Experiments on the NIST 2001 and NIST 2006 corpora indicate that VQ-MAP gives comparable accuracy with GMM-MAP with simpler implementation and faster adaptation. Ville Hautamäki, Tomi Kinnunen, Ismo Kärkkäinen, Juhani Saastamoinen, Marko Tuononen, Pasi Fränti |
IEEE Signal Process. Lett. | 1 |
| 2006 | Fast Agglomerative Clustering Using a k-Nearest Neighbor GraphabstractWe propose a fast agglomerative clustering method using an approximate nearest neighbor graph for reducing the number of distance calculations. The time complexity of the algorithm is improved from O(tauN2) to O(tauNlogN) at the cost of a slight increase in distortion; here, tau denotes the number of nearest neighbor updates required at each iteration. According to the experiments, a relatively small neighborhood size is sufficient to maintain the quality close to that of the full search. Pasi Fränti, Olli Virmajoki, Ville Hautamäki |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2003 | Fast PNN-based Clustering Using K-nearest Neighbor GraphabstractSearch for nearest neighbor is the main source of computation in most clustering algorithms. We propose the use of nearest neighbor graph for reducing the number of candidates. The number of distance calculations per search can be reduced from O(N) to O(k) or where N is the number of clusters, and k is the number of neighbors in the graph. We apply the proposed scheme within agglomerative clustering algorithm known as the PNN algorithm. Pasi Fränti, Olli Virmajoki, Ville Hautamäki |
ICDM | 3 |
| 2003 | On the fusion of dissimilarity-based classifiers for speaker identificationabstractIn this work, we describe a speaker identification system that uses multiple supplementary information sources for computing a combined match score for the unknown speaker. Each speaker profile in the database consists of multiple feature vector sets that can vary in their scale, dimensionality, and the number of vectors. The evidence from a given feature set is weighted by its reliability that is set in a priori fashion. The confidence of the identification result is also estimated. The system is evaluated with a corpus of 110 Finnish speakers. The evaluated feature sets include mel-cepstrum, LPC-cepstrum, dynamic cepstrum, long-term averaged spectrum of /A/ vowel, and F0. Tomi Kinnunen, Ville Hautamäki, Pasi Fränti |
INTERSPEECH | 2 |