EDBT 2026 Demo / reviewers in the wild / expert
Ante Jukic
dblp:132/1276
· DBLP profile ↗
16ranked-venue papers
6as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and InferenceabstractLarge language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and inference, especially for autoregressive models. To address this challenge, we present the Low Frame-rate Speech Codec (LFSC): a neural audio codec that leverages finite scalar quantization and adversarial training with large speech language models to achieve high-quality audio compression with a 1.89 kbps bitrate and 21.5 frames per second. We demonstrate that our novel codec can make the inference of LLM-based text-to-speech models around three times faster while improving intelligibility and producing quality comparable to previous models. Edresson Casanova, Ryan Langman, Paarth Neekhara, Shehzeen Hussain, Jason Li 0007, Subhankar Ghosh, Ante Jukic, Sang-gil Lee |
ICASSP | 7 |
| 2025 | Modifying Flow Matching for Generative Speech EnhancementabstractDiffusion-based generative models have been shown to be highly effective in various speech enhancement tasks. This work presents an analysis of a flow matching-based framework for generative speech enhancement as a simpler alternative to diffusion. Four different modifications to flow matching are proposed, employing an informed prior, a data prediction loss, deterministic inference, and early stopping. The proposed variants are evaluated on speech denoising, demonstrating performance comparable to a previous state-of-the art model using the same data setup. Through ablation studies, an efficient deterministic one-step inference configuration is proposed, which does not require any advanced training techniques such as pre-training or distillation. The proposed variants are also evaluated on speech dereverberation, demonstrating that stochastic inference without informed prior is preferable for this task. Roman Korostik, Rauf Nasretdinov, Ante Jukic |
ICASSP | 3 |
| 2025 | Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and RestorationabstractThis paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any vocoders for time-domain signal reconstruction. As a result, our model simplifies the synthesis process and removes the quality upper-bound introduced by any mel-spectrogram vocoder compared to prior work SpeechFlow. The proposed method is evaluated on multiple speech restoration tasks, including speech denoising, bandwidth extension, codec artifact removal, and target speaker extraction. In all scenarios, finetuning our pretrained model results in superior performance over strong baselines. Notably, in the target speaker extraction task, our model outperforms existing systems, including those leveraging SSL-pretrained encoders like WavLM. The code and the pretrained checkpoints are publicly available in the NVIDIA NeMo framework. Pin-Jui Ku, Alexander H. Liu, Roman Korostik, Sung-Feng Huang, Szu-Wei Fu, Ante Jukic |
ICASSP | 6 |
| 2025 | Robust Speech Recognition with Schrödinger Bridge-Based Speech EnhancementabstractIn this work, we investigate application of generative speech enhancement to improve the robustness of ASR models in noisy and reverberant conditions. We employ a recently-proposed speech enhancement model based on Schrödinger bridge, which has been shown to perform well compared to diffusion-based approaches. We analyze the impact of model scaling and different sampling methods on the ASR performance. Furthermore, we compare the considered model with predictive and diffusion-based baselines and analyze the speech recognition performance when using different pre-trained ASR models. The proposed approach significantly reduces the word error rate, reducing it by approximately 40% relative to the unprocessed speech signals and by approximately 8% relative to a similarly-sized predictive approach. Rauf Nasretdinov, Roman Korostik, Ante Jukic |
ICASSP | 3 |
| 2025 | NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
Edresson Casanova, Paarth Neekhara, Ryan Langman, Shehzeen Hussain, Subhankar Ghosh, Xuesong Yang, Ante Jukic, Jason Li 0007, Boris Ginsburg |
INTERSPEECH | 7 |
| 2025 | Universal Speech Enhancement with Regression and Generative Mamba
Rong Chao, Rauf Nasretdinov, Yu-Chiang Frank Wang, Ante Jukic, Szu-Wei Fu, Yu Tsao 0001 |
INTERSPEECH | 4 |
| 2025 | VoiceNoNG: Robust High-Quality Speech Editing Model without Hallucinations
Sung-Feng Huang, Heng-Cheng Kuo, Zhehuai Chen, Xuesong Yang, Pin-Jui Ku, Ante Jukic, Chao-Han Huck Yang, Yu Tsao 0001, Yu-Chiang Frank Wang, Hung-yi Lee, Szu-Wei Fu |
INTERSPEECH | 6 |
| 2025 | Unified Semi-Supervised Pipeline for Automatic Speech Recognition
Nune Tadevosyan, Nikolay Karpov, Andrei Andrusenko, Vitaly Lavrukhin, Ante Jukic |
INTERSPEECH | 5 |
| 2024 | Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
Kunal Dhawan, Nithin Rao Koluguri, Ante Jukic, Ryan Langman, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 3 |
| 2024 | Schrödinger Bridge for Generative Speech Enhancement
Ante Jukic, Roman Korostik, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 1 |
| 2017 | Adaptive Speech Dereverberation Using Constrained Sparse Multichannel Linear PredictionabstractIn this letter, we present an adaptive speech dereverberation method based on constrained sparse multichannel linear prediction (MCLP), minimizing the mixed ℓ2,pnorm of the desired component. In order to prevent overestimation of the undesired reverberant component, possibly leading to severe distortions of the output, we propose to use a statistical model for late reverberation to limit the power of the MCLP-based estimate. The resulting constrained optimization problem is solved by using the alternating direction method of multipliers, resulting in two variants of the dereverberation algorithm. Simulation results show that the proposed constraint increases the robustness with respect to parameter selection and improves the usability for dynamic scenarios in comparison to the unconstrained method. Ante Jukic, Toon van Waterschoot, Simon Doclo |
IEEE Signal Process. Lett. | 1 |
| 2016 | Robust sparsity-promoting acoustic multi-channel equalization for speech dereverberationabstractThis paper presents a novel signal-dependent method to increase the robustness of acoustic multi-channel equalization techniques against room impulse response (RIR) estimation errors. Aiming at obtaining an output signal which better resembles a clean speech signal, we propose to extend the acoustic multi-channel equalization cost function with a penalty function which promotes sparsity of the output signal in the short-time Fourier transform domain. Two conventionally used sparsity-promoting penalty functions are investigated, i.e., the l0-norm and the l1-norm, and the sparsity-promoting filters are iteratively computed using the alternating direction method of multipliers. Simulation results for several RIR estimation errors show that incorporating a sparsity-promoting penalty function significantly increases the robustness, with the l1-norm penalty function outperforming the l0-norm penalty function. Ina Kodrasi, Ante Jukic, Simon Doclo |
ICASSP | 2 |
| 2015 | Multi-channel linear prediction-based speech dereverberation with low-rank power spectrogram approximationabstractIn many acoustic conditions the recorded speech signals may be severely affected by reverberation, leading to a reduced speech quality and intelligibility. In this paper we focus on a blind speech dereverberation method based on multi-channel linear prediction (MCLP) in the short-time Fourier transform domain, which is typically performed in each frequency bin independently without taking into account the spectral structure of the speech signal. Since it is widely accepted that a speech spectrogram can be well approximated with a low-rank matrix, e.g., using a spectral dictionary, in this paper we propose to incorporate a low-rank matrix approximation of the speech spectrogram into the MCLP-based speech dereverberation. The low-rank approximation is obtained using nonnegative matrix factorization with Itakura-Saito divergence. Experimental results for several measured acoustic systems show that incorporating a low-rank approximation improves the dereverberation performance in terms of instrumental speech quality measures. Ante Jukic, Nasser Mohammadiha, Toon van Waterschoot, Timo Gerkmann, Simon Doclo |
ICASSP | 1 |
| 2015 | Multi-Channel Linear Prediction-Based Speech Dereverberation With Sparse PriorsabstractThe quality of speech signals recorded in an enclosure can be severely degraded by room reverberation. In this paper, we focus on a class of blind batch methods for speech dereverberation in a noiseless scenario with a single source, which are based on multi-channel linear prediction in the short-time Fourier transform domain. Dereverberation is performed by maximum-likelihood estimation of the model parameters that are subsequently used to recover the desired speech signal. Contrary to the conventional method, we propose to model the desired speech signal using a general sparse prior that can be represented in a convex form as a maximization over scaled complex Gaussian distributions. The proposed model can be interpreted as a generalization of the commonly used time-varying Gaussian model. Furthermore, we reformulate both the conventional and the proposed method as an optimization problem with an lp-norm cost function, emphasizing the role of sparsity in the considered speech dereverberation methods. Experimental evaluation in different acoustic scenarios show that the proposed approach results in an improved performance compared to the conventional approach in terms of instrumental measures for speech quality. Ante Jukic, Toon van Waterschoot, Timo Gerkmann, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Speech dereverberation using weighted prediction error with Laplacian model of the desired signalabstractReverberation has a considerable impact on the quality and intelligibility of captured speech signals. In this paper we present an approach for blind multi-microphone speech dereverberation based on the weighted prediction error method, where the reverberant observations are modeled using multi-channel linear prediction in the short-time Fourier transform domain. Instead of using the commonly employed Gaussian distribution for the desired speech signal, the proposed approach uses a Laplacian distribution which is known to be more accurate in modeling speech signals. Maximum-likelihood estimation is used for estimating the model parameters, leading to a linear programming optimization problem. Experimental results, obtained using measured impulse responses, indicate that the proposed approach could be used to improve the dereverberation performance compared to the classical technique. Ante Jukic, Simon Doclo |
ICASSP | 1 |
| 2013 | Supervised feature extraction for tensor objects based on maximization of mutual information
Ante Jukic, Marko Filipovic |
Pattern Recognit. Lett. | 1 |