Tasnima Sadekova

dblp:270/4697 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Improved Sampling Algorithms for Lévy-Itô Diffusion Models
abstract
Lévy-Itô denoising diffusion models relying on isotropic α-stable noise instead of Gaussian distribution have recently been shown to improve performance of conventional diffusion models in image generation on imbalanced datasets while performing comparably in the standard settings. However, the stochastic algorithm of sampling from such models consists in solving the stochastic differential equation describing only an approximate inverse of the process of adding α-stable noise to data which may lead to suboptimal performance. In this paper, we derive a parametric family of stochastic differential equations whose solutions have the same marginal densities as those of the forward diffusion and show that the appropriate choice of the parameter values can improve quality of the generated images when the number of reverse diffusion steps is small. Also, we demonstrate that Lévy-Itô diffusion models are applicable to diverse domains and show that a well-trained text-to-speech Lévy-Itô model may have advantages over standard diffusion models on highly imbalanced datasets.
Vadim Popov, Assel Yermekova, Tasnima Sadekova, Artem Khrapov, Mikhail Sergeevich Kudinov
ICLR3
2024 PitchFlow: adding pitch control to a Flow-matching based TTS model
Tasnima Sadekova, Mikhail A. Kudinov, Vadim Popov, Assel Yermekova, Artem Khrapov
INTERSPEECH1
2023 Optimal Transport in Diffusion Modeling for Conversion Tasks in Audio Domain
abstract
Diffusion models have recently become a popular generative modeling framework in various domains because of their high-quality sampling capabilities. Lately, it has been hypothesized that optimally trained diffusion models supplied with specific differential equation solvers provide a solution to the optimal transport problem between the data distribution and the prior distribution. In this paper, we empirically show that applying the optimal transport point of view on diffusion modeling allows making a good choice of a noise sample the reverse diffusion starts generating from. We consider two audio-related tasks: voice conversion and timbre transfer. In the former, we improve upon the recent state-of-the-art model and demonstrate that the optimal transport helps us to keep the prosody of the source utterances significantly better than the vanilla diffusion-based model does. As for timbre transfer, we propose the novel diffusion model capable of many-to-many timbre transfer performing on par with common algorithms in terms of the overall music quality.
Vadim Popov, Amantur Amatov, Mikhail A. Kudinov, Vladimir Gogoryan, Tasnima Sadekova, Ivan Vovk
ICASSP5
2023 Exploiting Emotion Information in Speaker Embeddings for Expressive Text-to-Speech
Zein Shaheen, Tasnima Sadekova, Yulia Matveeva, Alexandra Shirshova, Mikhail A. Kudinov
INTERSPEECH2
2022 Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, Mikhail Sergeevich Kudinov, Jiansheng Wei
ICLR4
2022 A Unified System for Voice Cloning and Voice Conversion through Diffusion Probabilistic Modeling
Tasnima Sadekova, Vladimir Gogoryan, Ivan Vovk, Vadim Popov, Mikhail A. Kudinov, Jiansheng Wei
INTERSPEECH1
2022 Fast Grad-TTS: Towards Efficient Diffusion-Based Speech Generation on CPU
Ivan Vovk, Tasnima Sadekova, Vladimir Gogoryan, Vadim Popov, Mikhail A. Kudinov, Jiansheng Wei
INTERSPEECH2
2021 Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
abstract
Recently, denoising diffusion probabilistic models and generative score matching have shown high potential in modelling complex data distributions while stochastic calculus has provided a unified point of view on these techniques allowing for flexible inference schemes. In this paper we introduce Grad-TTS, a novel text-to-speech model with score-based decoder producing mel-spectrograms by gradually transforming noise predicted by encoder and aligned with text input by means of Monotonic Alignment Search. The framework of stochastic differential equations helps us to generalize conventional diffusion probabilistic models to the case of reconstructing data from noise with different parameters and allows to make this reconstruction flexible by explicitly controlling trade-off between sound quality and inference speed. Subjective human evaluation shows that Grad-TTS is competitive with state-of-the-art text-to-speech approaches in terms of Mean Opinion Score.
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, Mikhail A. Kudinov
ICML4
2020 Gaussian Lpcnet for Multisample Speech Synthesis
abstract
LPCNet vocoder has recently been presented to TTS community and is now gaining increasing popularity due to its effectiveness and high quality of the speech synthesized with it. In this work, we present a modification of LPCNet that is 1.5x faster, has twice less non-zero parameters and synthesizes speech of the same quality. Such enhancement is possible mostly due to two features that we introduce into the original architecture: the proposed vocoder is designed to generate 16-bit signal instead of 8-bit μ-companded signal, and it predicts two consecutive excitation values at a time independently of each other. To show that these modifications do not lead to quality degradation we train models for five different languages and perform extensive human evaluation.
Vadim Popov, Mikhail A. Kudinov, Tasnima Sadekova
ICASSP3
2020 Fast and Lightweight On-Device TTS with Tacotron2 and LPCNet
Vadim Popov, Stanislav Kamenev, Mikhail A. Kudinov, Sergey Repyevsky, Tasnima Sadekova, Vitalii Bushaev, Vladimir Kryzhanovskiy, Denis Parkhomenko
INTERSPEECH5