VLDB 2026 Research / reviewers in the wild / expert
Srikanth Korse
dblp:164/9974
· DBLP profile ↗
8ranked-venue papers
4as first author
5since 2021 · last 2025
0009-0008-7564-9628ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Parametric Object Coding in IVAS: Efficient Coding of Multiple Audio Objects at Low Bit RatesabstractThe recently standardized 3GPP codec for Immersive Voice and Audio Services (IVAS) includes a parametric mode for efficiently coding multiple audio objects at low bit rates. In this mode, parametric side information is obtained from both the object metadata and the input audio objects. The side information comprises directional information, indices of two dominant objects, and the power ratio between these two dominant objects. It is transmitted to the decoder along with a stereo downmix. In IVAS, parametric object coding allows for transmitting three or four arbitrarily placed objects at bit rates of 24.4 or 32 kbit/s and faithfully reconstructing the spatial image of the original audio scene. Subjective listening tests confirm that IVAS provides a comparable immersive experience at lower bit rate and complexity compared to coding the audio objects independently using Enhanced Voice Services (EVS). Andrea Eichenseer, Srikanth Korse, Guillaume Fuchs, Markus Multrus |
ICASSP | 2 |
| 2024 | On Improving Error Resilience of Neural End-to-End Speech Codersabstract1755 Kishan Gupta, Nicola Pia, Srikanth Korse, Andreas Brendel, Guillaume Fuchs, Markus Multrus |
INTERSPEECH | 3 |
| 2022 | A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT DomainabstractFrequency domain processing, and in particular the use of Modified Discrete Cosine Transform (MDCT), is the most widespread approach to audio coding. However, at low bitrates, audio quality, especially for speech, degrades drastically due to the lack of available bits to directly code the transform coefficients. Traditionally, post-filtering has been used to mitigate artefacts in the coded speech by exploiting a-priori information of the source and extra transmitted parameters. Recently, datadriven post-filters have shown better results, but at the cost of significant additional complexity and delay. In this work, we propose a mask-based post-filter operating directly in MDCT domain of the codec, inducing no extra delay. The real-valued mask is applied to the quantized MDCT coefficients and is estimated from a relatively lightweight convolutional encoder-decoder network. Our solution is tested on the recently standardized low-delay, low-complexity codec (LC3) at lowest possible bitrate of 16 kbps. Objective and subjective assessments clearly show the advantage of this approach over the conventional post-filter, with an average improvement of 10 MUSHRA points over the LC3 coded speech. Kishan Gupta, Srikanth Korse, Bernd Edler, Guillaume Fuchs |
ICASSP | 2 |
| 2022 | PostGAN: A GAN-Based Post-Processor to Enhance the Quality of Coded SpeechabstractThe quality of speech coded by transform coding is affected by various artefacts especially when bitrates to quantize the frequency components become too low. In order to mitigate these coding artefacts and enhance the quality of coded speech, a post-processor that relies on a-priori information transmitted from the encoder is traditionally employed at the decoder side. In recent years, several data-driven post-postprocessors have been proposed which were shown to outperform traditional approaches. In this paper, we propose PostGAN, a GAN-based neural post-processor that operates in the sub-band domain and relies on the U-Net architecture and a learned affine transform. It has been tested on the recently standardized low-complexity, low-delay bluetooth codec (LC3) for wideband speech at the lowest bitrate (16 kbit/s). Subjective evaluations and objective scores show that the newly introduced post-processor surpasses previously published methods and can improve the quality of coded speech by around 20 MUSHRA points. Srikanth Korse, Nicola Pia, Kishan Gupta, Guillaume Fuchs |
ICASSP | 1 |
| 2022 | NESC: Robust Neural End-2-End Speech Coding with GANsabstract4212 Nicola Pia, Kishan Gupta, Srikanth Korse, Markus Multrus, Guillaume Fuchs |
INTERSPEECH | 3 |
| 2020 | Enhancement of Coded Speech Using a Mask-Based Post-FilterabstractThe quality of speech codecs deteriorates at low bitrates due to high quantization noise. A post-filter is generally employed to enhance the quality of the coded speech. In this paper, a data-driven post-filter relying on masking in the time-frequency domain is proposed. A fully connected neural network (FCNN), a convolutional encoder-decoder (CED) network and a long short-term memory (LSTM) network are implemeted to estimate a real-valued mask per time-frequency bin. The proposed models were tested on the five lowest operating modes (6.65 kbps-15.85 kbps) of the Adaptive Multi-Rate Wideband codec (AMR-WB). Both objective and subjective evaluations confirm the enhancement of the coded speech and also show the superiority of the mask-based neural network system over a conventional heuristic post-filter used in the standard like ITU-T G.718. Srikanth Korse, Kishan Gupta, Guillaume Fuchs |
ICASSP | 1 |
| 2018 | GMM-Based Iterative Entropy Coding for Spectral Envelopes of Speech and AudioabstractSpectral envelope modelling is a central part of speech and audio codecs and is traditionally based on either vector quantization or scalar quantization followed by entropy coding. To bridge the coding performance of vector quantization with the low complexity of the scalar case, we propose an iterative approach for entropy coding the spectral envelope parameters. For each parameter, a univariate probability distribution is derived from a Gaussian mixture model of the joint distribution and the previously quantized parameters used as a-priori information. Parameters are then iteratively and individually scalar quantized and entropy coded. Unlike vector quantization, the complexity of proposed method does not increase exponentially with dimension and bitrate. Moreover, the coding resolution and dimension can be adaptively modified without retraining the model. Experimental results show that these important advantages do not impair coding efficiency compared to a state-of-art vector quantization scheme. Srikanth Korse, Guillaume Fuchs, Tom Bäckström |
ICASSP | 1 |
| 2016 | Entropy Coding of Spectral Envelopes for Speech and Audio Coding Using Distribution QuantizationabstractS.2543-2547 Srikanth Korse, Tobias Jähnel, Tom Bäckström |
INTERSPEECH | 1 |