EDBT 2026 Demo / reviewers in the wild / expert
Guillaume Fuchs
dblp:40/8054
· DBLP profile ↗
22ranked-venue papers
8as first author
10since 2021 · last 2025
0009-0009-1045-6064ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Parametric Object Coding in IVAS: Efficient Coding of Multiple Audio Objects at Low Bit RatesabstractThe recently standardized 3GPP codec for Immersive Voice and Audio Services (IVAS) includes a parametric mode for efficiently coding multiple audio objects at low bit rates. In this mode, parametric side information is obtained from both the object metadata and the input audio objects. The side information comprises directional information, indices of two dominant objects, and the power ratio between these two dominant objects. It is transmitted to the decoder along with a stereo downmix. In IVAS, parametric object coding allows for transmitting three or four arbitrarily placed objects at bit rates of 24.4 or 32 kbit/s and faithfully reconstructing the spatial image of the original audio scene. Subjective listening tests confirm that IVAS provides a comparable immersive experience at lower bit rate and complexity compared to coding the audio objects independently using Enhanced Voice Services (EVS). Andrea Eichenseer, Srikanth Korse, Guillaume Fuchs, Markus Multrus |
ICASSP | 3 |
| 2025 | A first-order DirAC-based parametric Ambisonic coder for immersive communicationsabstractDirectional Audio Coding (DirAC) is a proven method for parametrically representing a 3D audio scene in B-format and is capable of reproducing it on arbitrary loudspeaker layouts. Although such a method seems well suited for low bitrate Ambisonic transmission, little work has been done on the feasibility of building a real system upon it. In this paper, we present a DirAC-based coding for Higher-Order Ambisonics (HOA), developed as part of a standardisation effort to extend the 3GPP EVS codec to immersive communications. Starting from the first-order DirAC model, we show how to reduce algorithmic delay, the bitrate required for the parameters and complexity by bringing the full synthesis in the spherical harmonic domain. The evaluation of the proposed technique for coding 3rdorder Ambisonics at bitrates from 32 to 128 kbps shows the relevance of the parametric approach compared with existing solutions. Guillaume Fuchs, Florin Ghido, Dominik Weckbecker, Oliver Thiergart |
ICASSP | 1 |
| 2025 | Hybrid predictive and parametric stereo coding for voice and audio communicationsabstractThe transmission of stereo audio in voice calls helps to improve immersion and user experience. However stereo coding has primarily been studied for broadcast or streaming applications. To extend the capabilities of 3GPP EVS, a hybrid stereo coding scheme combining predictive and parametric coding is proposed to meet the requirements of real-time communications. After the stereo channels are time-aligned, a sum-difference (M/S) decomposition is performed. The prediction of the side signal by the mid signal is used to capture the other spatial cues. The prediction residue can be parametrically modeled by a stereo filling or, at higher bitrates and for low frequencies, discretely coded. Listening test results confirm the stereo coding scheme’s great flexibility and robustness under different recording conditions, as well as the minimal quality impact of the low-bitrate and low-delay constraints. Guillaume Fuchs, Emmanuel Ravelli, Franz Reutelhuber, Eleni Fotopoulou |
ICASSP | 1 |
| 2025 | Ambisonics Coding in IVAS: A Hybrid SPAR and DirAC SystemabstractThe Ambisonics audio format represents a 3D sound field as a fixed set of audio channels, with a tradeoff between spatial detail and the required number of audio channels. Because the number of audio channels that must be coded increases quadratically with respect to Ambisonics order, high-quality coding of Ambisonics signals under practical bitrate and complexity constraints becomes a major challenge. To overcome this, the 3GPP IVAS codec employs a new hybrid parametric and residual coding scheme combining complementary Spatial Reconstruction (SPAR) and Directional Audio Coding (DirAC) techniques. In this paper, we describe the implementation of this hybrid coding system and demonstrate its advantages to discrete channel-coding approaches in terms of computational complexity and audio quality. Dominik Weckbecker, Stefanie Brown, Juan Torres 0001, Markus Multrus, Archit Tamarapu, Guillaume Fuchs |
ICASSP | 6 |
| 2025 | Benchmarking Neural Speech Codec Intelligibility with SIToolabstract5488 Anna Leschanowsky, Kishor Kayyar Lakshminarayana, Anjana Rajasekhar, Lyonel Behringer, Ibrahim Kilinc, Guillaume Fuchs, Emanuël A. P. Habets |
INTERSPEECH | 6 |
| 2024 | On Improving Error Resilience of Neural End-to-End Speech Codersabstract1755 Kishan Gupta, Nicola Pia, Srikanth Korse, Andreas Brendel, Guillaume Fuchs, Markus Multrus |
INTERSPEECH | 5 |
| 2022 | A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT DomainabstractFrequency domain processing, and in particular the use of Modified Discrete Cosine Transform (MDCT), is the most widespread approach to audio coding. However, at low bitrates, audio quality, especially for speech, degrades drastically due to the lack of available bits to directly code the transform coefficients. Traditionally, post-filtering has been used to mitigate artefacts in the coded speech by exploiting a-priori information of the source and extra transmitted parameters. Recently, datadriven post-filters have shown better results, but at the cost of significant additional complexity and delay. In this work, we propose a mask-based post-filter operating directly in MDCT domain of the codec, inducing no extra delay. The real-valued mask is applied to the quantized MDCT coefficients and is estimated from a relatively lightweight convolutional encoder-decoder network. Our solution is tested on the recently standardized low-delay, low-complexity codec (LC3) at lowest possible bitrate of 16 kbps. Objective and subjective assessments clearly show the advantage of this approach over the conventional post-filter, with an average improvement of 10 MUSHRA points over the LC3 coded speech. Kishan Gupta, Srikanth Korse, Bernd Edler, Guillaume Fuchs |
ICASSP | 4 |
| 2022 | PostGAN: A GAN-Based Post-Processor to Enhance the Quality of Coded SpeechabstractThe quality of speech coded by transform coding is affected by various artefacts especially when bitrates to quantize the frequency components become too low. In order to mitigate these coding artefacts and enhance the quality of coded speech, a post-processor that relies on a-priori information transmitted from the encoder is traditionally employed at the decoder side. In recent years, several data-driven post-postprocessors have been proposed which were shown to outperform traditional approaches. In this paper, we propose PostGAN, a GAN-based neural post-processor that operates in the sub-band domain and relies on the U-Net architecture and a learned affine transform. It has been tested on the recently standardized low-complexity, low-delay bluetooth codec (LC3) for wideband speech at the lowest bitrate (16 kbit/s). Subjective evaluations and objective scores show that the newly introduced post-processor surpasses previously published methods and can improve the quality of coded speech by around 20 MUSHRA points. Srikanth Korse, Nicola Pia, Kishan Gupta, Guillaume Fuchs |
ICASSP | 4 |
| 2022 | NESC: Robust Neural End-2-End Speech Coding with GANsabstract4212 Nicola Pia, Kishan Gupta, Srikanth Korse, Markus Multrus, Guillaume Fuchs |
INTERSPEECH | 5 |
| 2021 | StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive NormalizationabstractIn recent years, neural vocoders have surpassed classical speech generation approaches in naturalness and perceptual quality of the synthesized speech. Computationally heavy models like WaveNet and WaveGlow achieve best results, while lightweight GAN models, e.g. MelGAN and Parallel WaveGAN, remain inferior in terms of perceptual quality. We therefore propose StyleMelGAN, a lightweight neural vocoder allowing synthesis of high-fidelity speech with low computational complexity. StyleMelGAN employs temporal adaptive normalization to style a low-dimensional noise vector with the acoustic features of the target speech. For efficient training, multiple random-window discriminators adversarially evaluate the speech signal analyzed by a filter bank, with regularization provided by a multi-scale spectral reconstruction loss. The highly parallelizable speech generation is several times faster than real-time on CPUs and GPUs. MUSHRA and P.800 listening tests show that StyleMelGAN outperforms prior neural vocoders in copy-synthesis and Text-to-Speech scenarios. Ahmed Mustafa, Nicola Pia, Guillaume Fuchs |
ICASSP | 3 |
| 2020 | Enhancement of Coded Speech Using a Mask-Based Post-FilterabstractThe quality of speech codecs deteriorates at low bitrates due to high quantization noise. A post-filter is generally employed to enhance the quality of the coded speech. In this paper, a data-driven post-filter relying on masking in the time-frequency domain is proposed. A fully connected neural network (FCNN), a convolutional encoder-decoder (CED) network and a long short-term memory (LSTM) network are implemeted to estimate a real-valued mask per time-frequency bin. The proposed models were tested on the five lowest operating modes (6.65 kbps-15.85 kbps) of the Adaptive Multi-Rate Wideband codec (AMR-WB). Both objective and subjective evaluations confirm the enhancement of the coded speech and also show the superiority of the mask-based neural network system over a conventional heuristic post-filter used in the standard like ITU-T G.718. Srikanth Korse, Kishan Gupta, Guillaume Fuchs |
ICASSP | 3 |
| 2020 | Fundamental Frequency Model for Postfiltering at Low Bitrates in a Transform-Domain Speech and Audio CodecabstractS.2837-2841 Sneha Das, Tom Bäckström, Guillaume Fuchs |
INTERSPEECH | 3 |
| 2019 | Super-Wideband Spectral Envelope Modeling for Speech CodingabstractSignificant improvements in the quality of speech coders have been achieved by widening the coded frequency range from narrowband to wideband. However, existing speech coders still employ a limited band source-filter model extended by parametric coding of the higher band. In the present work, a superwideband source-filter model running at 32 kHz is considered and especially its spectral magnitude envelope modeling. To match super-wideband operating mode, we adapted and compared two methods; Linear Predictive Coding (LPC) and Distribution Quantization (DQ). LPC uses autoregressive modeling, while DQ quantifies the energy ratios between different parts of the spectrum. Parameters of both methods were quantized with a multi-stage vector quantization. Objective and subjective evaluations indicate that both methods used in a super-wideband source-filter coding scheme offer the same quality range, making them an attractive alternative to conventional speech coders that require additional bandwidth extension. Guillaume Fuchs, Chamran Ashour, Tom Bäckström |
INTERSPEECH | 1 |
| 2018 | GMM-Based Iterative Entropy Coding for Spectral Envelopes of Speech and AudioabstractSpectral envelope modelling is a central part of speech and audio codecs and is traditionally based on either vector quantization or scalar quantization followed by entropy coding. To bridge the coding performance of vector quantization with the low complexity of the scalar case, we propose an iterative approach for entropy coding the spectral envelope parameters. For each parameter, a univariate probability distribution is derived from a Gaussian mixture model of the joint distribution and the previously quantized parameters used as a-priori information. Parameters are then iteratively and individually scalar quantized and entropy coded. Unlike vector quantization, the complexity of proposed method does not increase exponentially with dimension and bitrate. Moreover, the coding resolution and dimension can be adaptively modified without retraining the model. Experimental results show that these important advantages do not impair coding efficiency compared to a state-of-art vector quantization scheme. Srikanth Korse, Guillaume Fuchs, Tom Bäckström |
ICASSP | 2 |
| 2015 | Low delay LPC and MDCT-based audio coding in the EVS codecabstractSpeech coders operating in time domain can be extended with a frequency domain mode to improve encoding of music, even though this is challenging at low delay. In such a scenario, the short analysis window limits the benefit of the transform coder, while a delayless switch between the two coders constrains the system further. The paper presents an LPC and MDCT-based audio coder part of the new 3GPP codec for Enhanced Voice Services, which aims to solve the issues. Several advanced coding tools are introduced to alleviate the constraints: transient handling is improved, harmonic structures are better preserved, and the modeling of the zero-quantized frequencies is enhanced. Test results show that the obtained low-delay switched coder brings a clear improvement over a speech coder and is competitive even in comparison to audio coders with higher delay. Guillaume Fuchs, Christian R. Helmrich, Goran Markovic, Matthias Neusinger, Emmanuel Ravelli, Takehiro Moriya |
ICASSP | 1 |
| 2015 | Frequency-domain Comfort Noise Generation for Discontinuous Transmission in EVSabstractDiscontinuous Transmission (DTX) is an efficient way to drastically reduce the transmission rate of a communication codec in the absence of voice input. In this mode, most frames that are determined to consist of background noise only are dropped from transmission and replaced by some Comfort Noise Generation (CNG) in the decoder. In this paper, we propose a novel CNG approach combining information gained about the actual background noise at both encoder and decoder side. It is able to better reproduce background noise types showing a pronounced spectral tilt, which is difficult for traditional schemes based on a linear prediction model. The proposed technique operates in the frequency domain. It is part of the Enhanced Voice Services (EVS) codec, where it is known as FD-CNG. Listening tests show the superior quality of FD-CNG over existing approaches for certain background noise such as car noise. Anthony Lombard, Stephan Wilde, Emmanuel Ravelli, Stefan Döhla, Guillaume Fuchs, Martin Dietz |
ICASSP | 5 |
| 2015 | Low-complexity and robust coding mode decision in the EVS coderabstractSeveral state-of-the-art switched audio codecs employ the closed-loop mode decision to select the best coding mode at every frame. The closed-loop mode selection is known to have good performance but also high complexity. The new approach we propose in this paper is a low-complexity version of the closed-loop approach, based on similar decisions which compute the coding distortion of each mode and select the one with the lowest distortion. Our approach differs mainly in the way the coding distortions are calculated. We are able to notably reduce the complexity by only estimating the distortions without encoding and decoding the input for each mode. The new approach was implemented in the EVS codec standard and evaluated both objectively and subjectively. Compared to the closed-loop approach, it yields similar performance and lower complexity. Emmanuel Ravelli, Christian R. Helmrich, Guillaume Fuchs, Markus Multrus |
ICASSP | 3 |
| 2013 | Embedded Voronoi codes for successive refinement lattice vector quantizationabstractLattice Vector Quantization (LVQ) is an interesting tool in source coding which can take advantage of a higher dimension than the scalar case while overcoming complexity limitations of conventional vector quantization. However, the high dimension and the relatively complex indexing of the codebooks make LVQ often unsuitable for getting a successive refinement of the source. For addressing this problem, the paper proposes a new class of LVQ called the embedded Voronoi codes. The new codes can gradually describe the source with a granularity of 1 bit/dimension by properly combining differently scaled Voronoi codes. A rate-distortion evaluation for a Gaussian source shows that the embedding of the codes comes at a minimal cost at low bit-rates while preserving LVQ advantages over scalar quantization. Guillaume Fuchs |
ICASSP | 1 |
| 2011 | Efficient context adaptive entropy coding for real-time applicationsabstractContext based entropy coding has the potential to provide higher gain over memoryless entropy coding. However serious difficulties arise regarding the practical implementation in real-time applications due to its very high memory requirements. This paper presents an efficient method for designing context adaptive entropy coding while fulfilling low memory requirements. From a study of coding gain scalability as a function of context size, new context design and validation procedures are derived. Further, supervised clustering and mapping optimization are introduced to model efficiently the context. The resulting context modelling associated with an arithmetic coder was successfully implemented in a transform-based audio coder for real-time processing. It shows significant improvement over the entropy coding used in MPEG-4 AAC. Guillaume Fuchs, Vignesh Subbaraman, Markus Multrus |
ICASSP | 1 |
| 2009 | Unified speech and audio coding scheme for high quality at low bitratesabstractTraditionally, speech coding and audio coding were separate worlds. Based on different technical approaches and different assumptions about the source signal, neither of the two coding schemes could efficiently represent both speech and music at low bitrates. This paper presents a unified speech and audio codec, which efficiently combines techniques from both worlds. This results in a codec that exhibits consistently high quality for speech, music and mixed audio content. The paper gives an overview of the codec architecture and presents results of formal listening tests comparing this new codec with HE-AAC(v2) and AMR-WB+. This new codec forms the basis of the reference model in the ongoing MPEG standardization activity for Unified Speech and Audio Coding. Max Neuendorf, Philippe Gournay, Markus Multrus, Jérémie Lecomte, Bruno Bessette, Ralf Geiger, Stefan Bayer, Guillaume Fuchs, Johannes Hilpert, Nikolaus Rettelbach, Redwan Salami, Gerald Schuller, Roch Lefebvre, Bernhard Grill |
ICASSP | 8 |
| 2006 | A New Post-Filtering for Artificially Replicated High-Band in Speech CodersabstractPure estimation in bandwidth extension of bandlimited speech does not provide transparent quality in the extended band. Combining coding and estimation of the missing band can achieve higher quality by transmitting additional information. However, substantial artifacts can remain in the extended band. In this paper, we introduce a new postfiltering designed to reduce these artifacts. Based on a source-filter model, the adopted bandwidth extension transmits spectral envelope parameters, while estimates the excitation at the receiver. The perceptual quality is then enhanced by processing the estimated excitation by a short-term and a long-term postfilter. Objective and subjective assessments show substantial gains when using the new postfiltering. It confirms that low bit-rate wideband coding based on bandwidth extension can get closer to the performances of the single band wideband coders Guillaume Fuchs, Roch Lefebvre |
ICASSP (1) | 1 |
| 2005 | A speech coder post-processor controlled by side-informationabstractSpeech coders provide high speech quality at low rates. However they perform poorly when encoding non-speech signals. This paper proposes a new enhancement algorithm requiring minimum side information to reduce the effect of this shortcoming. The enhancement algorithm consists of post-processing the speech decoder output in the spectral domain. Specifically, some frequency components are reduced or forced to zero when the corresponding frequency content is poorly described by the speech coder. The choice of modifying spectral components is determined at the encoder, thus requiring us to transmit the decision information. Experiments combining the AMR-WB speech codec and the proposed audio enhancement show that the quality for music signals is improved significantly while not affecting the quality for speech inputs. Guillaume Fuchs, Roch Lefebvre |
ICASSP (4) | 1 |