Lars F. Villemoes

dblp:38/4186 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 100%
Theoretical computer science
1 paper
Mathematical optimization · 67% Information theory · 33%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
audio coding
0.112008
A Backward-Compatible Multichannel Audio Codec · IEEE Trans. Speech Audio Process. 2008
Audio and music processing › audio coding
multichannel audio coding
0.112008
A Backward-Compatible Multichannel Audio Codec · IEEE Trans. Speech Audio Process. 2008
Mathematical optimization › sparse optimization
basis selection
0.012002
Reduction of correlated noise using a library of orthonormal bases · IEEE Trans. Inf. Theory 2002
Information theory › signal processing
denoising
0.012002
Reduction of correlated noise using a library of orthonormal bases · IEEE Trans. Inf. Theory 2002
Mathematical optimization
statistical estimation
0.012002
Reduction of correlated noise using a library of orthonormal bases · IEEE Trans. Inf. Theory 2002

Methods — techniques the papers use, named apart from their topics

perceptual redundancy exploitation · 0.1parametric coding · 0.1down-mixing · 0.1wavelet packet · 0.0thresholding · 0.0reverse holder inequality · 0.0
YearPublicationVenuePosition
2025 Audio Decoding by Inverse Problem Solving
abstract
We consider audio decoding as an inverse problem and solve it through diffusion posterior sampling. Explicit conditioning functions are developed for input signal measurements provided by an example of a transform domain perceptual audio codec. Viability is demonstrated by evaluating arbitrary pairings of a set of bitrates and task-agnostic prior models. For instance, we observe significant improvements on piano while maintaining speech performance when a speech model is replaced by a joint model trained on both speech and piano. With a more general music model, improved decoding compared to legacy methods is obtained for a broad range of content types and bitrates. The noisy mean model, underlying the proposed derivation of conditioning, enables a significant reduction of gradient evaluations for diffusion posterior sampling, compared to methods based on Tweedie’s mean. Combining Tweedie’s mean with our conditioning functions improves the objective performance. An audio demo is available at https://dpscodec-demo.github.io/
Pedro J. Villasana T., Lars F. Villemoes, Janusz Klejsa, Per Hedelin
ICASSP2
2023 High Quality Audio Coding with Mdctnet
abstract
We propose a neural audio generative model, MDCTNet, operating in the perceptually weighted domain of an adaptive modified discrete cosine transform (MDCT). The architecture of the model captures correlations in both time and frequency directions with recurrent layers (RNNs). An audio coding system is obtained by training MDCTNet on a diverse set of fullband monophonic audio signals at 48 kHz sampling, conditioned by a perceptual audio encoder. In a subjective listening test with ten excerpts chosen to be balanced across content types, yet critical for both codecs, the mean performance of the proposed system for 24 kb/s variable bitrate (VBR) is similar to that of Opus at twice the bitrate.
Grant A. Davidson, Mark Vinton, Per Ekstrand, Lars F. Villemoes, Lie Lu
ICASSP5
2020 Source Coding of Audio Signals with a Generative Model
abstract
We consider source coding of audio signals with the help of a generative model. We use a construction where a waveform is first quantized, yielding a finite bitrate representation. The waveform is then reconstructed by random sampling from a model conditioned on the quantized waveform. The proposed coding scheme is theoretically analyzed. Using SampleRNN as the generative model, we demonstrate that the proposed coding structure provides performance competitive with state-of-the-art source coding tools for specific categories of audio signals.
Roy Fejgin, Janusz Klejsa, Lars F. Villemoes
ICASSP3
2019 High-quality Speech Coding with Sample RNN
abstract
We provide a speech coding scheme employing a generative model based on SampleRNN that, while operating at significantly lower bitrates, matches or surpasses the perceptual quality of state-of-the-art classic wide-band codecs. Moreover, it is demonstrated that the proposed scheme can provide a meaningful rate-distortion trade-off without retraining. We evaluate the proposed scheme in a series of listening tests and discuss limitations of the approach.
Janusz Klejsa, Per Hedelin, Roy Fejgin, Lars F. Villemoes
ICASSP5
2018 Temporal Noise Shaping with Companding
Arijit Biswas, Per Hedelin, Lars F. Villemoes, Vinay Melkote
INTERSPEECH3
2017 Decorrelation for audio object coding
abstract
Object-based representations of audio content are increasingly used in entertainment systems to deliver immersive and personalized experiences. Efficient storage and transmission of such content can be achieved by joint object coding algorithms that convey a reduced number of downmix signals together with parametric side information that enables object reconstruction in the decoder. This paper presents an approach to improve the performance of joint object coding by adding one or more decorrelators to the decoding process. Listening test results illustrate the performance as a function of the number of decorrelators. The method is adopted as part of the Dolby AC-4 system standardized by ETSI.
Lars F. Villemoes, Toni Hirvonen, Heiko Purnhagen
ICASSP1
2011 Efficient transform coding of two-channel audio signals by means of complex-valued stereo prediction
abstract
Traditional MDCT-based perceptual audio coding schemes employ mid/side and intensity stereo techniques to allow efficient joint coding of the two channels of a stereophonic signal. These techniques, however, provide only little coding gain for critical stereo signals characterized by spectral components with a distinct level or phase difference between the channels. To overcome this deficiency, we propose an extension to the mid/side coding paradigm that utilizes complex-valued inter-channel linear prediction in the MDCT spectral domain. The required imaginary spectrum (MDST) is calculated in a computationally efficient manner without additional algorithmic delay. A formal listening test conducted in the course of the ISO/MPEG standardization of the unified speech and audio codec USAC illustrates that the proposed stereo prediction approach pro vides significant improvements in coding efficiency and shows that at 96 kb/s, excellent quality can be obtained even for critical signals.
Christian R. Helmrich, Pontus Carlsson, Sascha Disch, Bernd Edler, Johannes Hilpert, Matthias Neusinger, Heiko Purnhagen, Nikolaus Rettelbach, Julien Robilliard, Lars F. Villemoes
ICASSP10
2008 A Backward-Compatible Multichannel Audio Codec
abstract
We propose in this paper a backward-compatible multichannel audio codec. This codec represents a multichannel audio input signal by a down mix and parametric data. In order to enable backward compatibility, it is necessary to have the possibility of exerting control over the down-mixing procedure. At the same time, in order to achieve a high coding efficiency, both signal and perceptual redundancies should be exploited. In this paper, we describe a codec that unifies the above-mentioned conditions: backward compatibility and exploitation of both signal and perceptual redundancies. The codec combines a high audio quality and a low parameter bit rate. Moreover, its design is flexible, examples of which are the scalability of the audio quality to (in principle) transparency and the possibility to preserve the correlation structure of the original input signals by using synthetic signals. A stereo backward compatible version of the proposed codec is used as a component of the recently standardized MPEG Surround multichannel audio codec.
Gerard Hotho, Lars F. Villemoes, Jeroen Breebaart
IEEE Trans. Speech Audio Process.2
2002 Reduction of correlated noise using a library of orthonormal bases
abstract
We study the application of a library of orthonormal bases to the reduction of correlated Gaussian noise. A joint condition on the library and the noise covariance is derived which ensures that simple thresholding in an adaptively chosen basis yields an estimation error within a logarithmic factor of the ideal risk. In the model example of a wavelet packet library and stationary noise the condition can be translated into a reverse Holder inequality on the power spectrum.
Lars F. Villemoes
IEEE Trans. Inf. Theory1
1991 Optimizing time-frequency resolution of orthonormal wavelets
abstract
The authors consider the time and the frequency resolutions associated with the wavelet transform since these resolution properties are essential to evaluate a time-frequency analysis tool. A method is introduced to measure the resolution properties of the orthonormal wavelets. The authors exploit the degree of freedom in the construction of compactly supported wavelets in order to find the ones which yield the highest time-frequency concentration. It is shown that some of these orthonormal wavelets get close to the Heisenberg limit, which makes the corresponding wavelet transform particularly appropriate to signal analysis.>
Christian Dorize, Lars F. Villemoes
ICASSP2