EDBT 2026 Demo / reviewers in the wild / expert
Sebastian J. Schlecht
dblp:129/7864
· DBLP profile ↗
17ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0001-8858-4642ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FLAMO: An Open-Source Library for Frequency-Domain Differentiable Audio ProcessingabstractWe present FLAMO, a Frequency-sampling Library for Audio-Module Optimization designed to implement and optimize differentiable linear time-invariant audio systems. The library is open-source and built on the frequency-sampling filter design method, allowing for the creation of differentiable modules that can be used stand-alone or within the computation graph of neural networks, simplifying the development of differentiable audio systems. It includes predefined filtering modules and auxiliary classes for constructing, training, and logging the optimized systems, all accessible through an intuitive interface. Practical application of these modules is demonstrated through two case studies: the optimization of an artificial reverberator and an active acoustics system for improved response coloration. Gloria Dal Santo, Gian Marco De Bortoli, Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki |
ICASSP | 4 |
| 2025 | Multi-shelf graphic equalizerabstractA graphic equalizer (GEQ) is a standard tool in audio production and effect design. Adjustable gain control frequencies are fixed along the logarithmic frequency axis, and an automatic design method matches the magnitude response to them whenever target gains are changed. Most commonly, the GEQ comprises a set of peak filters centered an octave apart, possibly with a shelving filter at the bottom and top of the frequency range. While accurate designs were proposed, the dynamic range is typically limited to 24 dB. In this paper, we propose two innovations. First, we introduce a GEQ based on shelving filters only, which can cover an extensive dynamic range of over 60 dB. Secondly, we introduce an order-switching technique that combines shelf filters of different order. We demonstrate the performance and advantages of the proposed filter with design examples. The proposed shelf-filter-based GEQ offers a wider dynamic range and a smoother magnitude response than traditional peak-filter-based GEQ designs. Sebastian J. Schlecht, Tantep Sinjanakhom, Vesa Välimäki |
Signal Process. | 1 |
| 2024 | KLANN: Linearising Long-Term Dynamics in Nonlinear Audio Effects Using Koopman NetworksabstractIn recent years, neural network-based black-box modeling of nonlinear audio effects has improved considerably. Present convolutional and recurrent models can model audio effects with long-term dynamics, but the models require many parameters, thus increasing the processing time. In this letter, we propose KLANN, a Koopman-Linearised Audio Neural Network structure that lifts a one-dimensional signal (mono audio) into a high-dimensional approximately linear state-space representation with nonlinear mapping, and then uses differentiable biquad filters to predict linearly within the lifted state-space. Results show that the proposed models match the high performance of the state-of-the-art neural models while having a more compact architecture, reducing the number of parameters by tenfold, and having interpretable components. Ville Huhtala, Lauri Juvela, Sebastian J. Schlecht |
IEEE Signal Process. Lett. | 3 |
| 2024 | Modal Excitation in Feedback Delay NetworksabstractFeedback delay networks (FDNs) are used in audio processing and synthesis. The modal shapes of the system describe the modal excitation by input and output signals. Previously, the Ehrlich-Aberth method was used to find modes in large FDNs. Here, the method is extended to the corresponding eigenvectors indicating the modal shape. In particular, the computational complexity of the proposed analysis method does not depend on the delay-line lengths and is thus suitable for large FDNs, such as artificial reverberators. We show the relation between the compact generalized eigenvectors in the delay state space and the spatially extended modal shapes in the state space. We illustrate this method with an example FDN in which the suggested modal excitation control does not increase the computational cost. The modal shapes can help optimize input and output gains. This letter teaches how selecting the input and output points along the delay lines of an FDN adjusts the spectral shape of the system output. Sebastian J. Schlecht, Matteo Scerbo, Enzo De Sena, Vesa Välimäki |
IEEE Signal Process. Lett. | 1 |
| 2024 | Two-Stage Attenuation Filter for Artificial ReverberationabstractDelay networks are a common parametric method to synthesize the late part of the room reverberation. A delay network consists of several feedback loops, each containing a delay line and an attenuation filter, which approximates the same decay rate by appropriately setting the frequency-dependent loop gain. A remaining challenge is the design of the attenuation filters on a wide frequency range based on a measured room impulse response. This letter proposes a novel two-stage attenuation filter structure, sharpening the design. The first stage is a low-order pre-filter approximating the overall shape and determining the decay at the two ends of the frequency range, namely at the dc and the Nyquist limit. The second filter, an equalizer, fine-tunes the gain at different frequencies, such as on one-third-octave bands. It is shown that the proposed design is more accurate and robust than previous methods. A design example applying the proposed method to an interleaved velvet-noise reverberator is also exhibited. The proposed two-stage attenuation filter is a step toward a realistic parametric simulation of measured room impulse responses. Vesa Välimäki, Karolina Prawda, Sebastian J. Schlecht |
IEEE Signal Process. Lett. | 3 |
| 2023 | Interpolation of Spatial Room Impulse Responses Using Partial Optimal TransportabstractInterpolation between spatial room impulse responses (SRIRs) is necessary for dynamic acoustic rendering in which a listener can move with six degrees-of-freedom. The early part of the SRIR consists of sparse direct and reflected sound events, whose arrival time, direction and level vary with receiver position. Interpolation of the spatio-temporal structure necessitates the non-trivial task of mapping corresponding sound events. Instead of finding an exact map, we propose using partial optimal transport to find a coupling between reflections requiring neither estimation of the room geometry nor explicit knowledge of the source-receiver configuration. Each SRIR is first decomposed into a virtual source space. Then, the interpolated impulse response is calculated based on a partial optimal transport coupling obtained with linear programming. We compare the method against two baseline interpolation methods using simulated SRIR data, and show that it best preserves the temporal fine structure of the omnidirectional response. Aaron Geldert, Nils Meyer-Kahlen, Sebastian J. Schlecht |
ICASSP | 3 |
| 2023 | Grouped Feedback Delay Networks With Frequency-Dependent CouplingabstractFeedback Delay Networks are one of the most popular and efficient means of generating artificial reverberation. Recently, we proposed the Grouped Feedback Delay Network (GFDN), which couples multiple FDNs while maintaining system stability. The GFDN can be used to model reverberation in coupled spaces that exhibit multi-stage decay. The block feedback matrix determines the inter- and intra-group coupling. In this article, we expand on the design of the block feedback matrix to include frequency-dependent coupling among the various FDN groups. We show how paraunitary feedback matrices can be designed to emulate diffraction at the aperture connecting rooms. Several methods for the construction of nearly paraunitary matrices are investigated. The proposed method supports the efficient rendering of virtual acoustics for complex room topologies in games and XR applications. Orchisama Das, Sebastian J. Schlecht, Enzo De Sena |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Common-Slope Modeling of Late ReverberationabstractThe decaying sound field in rooms is typically described by energy decay functions (EDFs). Late reverberation can deviate considerably from the ideal diffuse field, for example, in multiple connected rooms or non-uniform absorption material distributions. This paper proposes the common-slope model of late reverberation. The model describes spatial and directional late reverberation as linear combinations of exponential decays called common slopes. Its fundamental idea is that common slopes have decay times that are invariant across space and direction, while their amplitudes vary across both. We explore different approaches for determining the common slopes for large EDF sets describing different source-receiver configurations of the same environment. Among the presented approaches, the k-means clustering of decay times is the most general. Our evaluation shows that the common-slope model introduces only a small error between the modeled and the true EDF, while being considerably more compact than the traditional multi-exponential model. The amplitude variations of the common slopes yield interpretable room acoustic analyses. The common-slope model has potential applications in all fields relying on late reverberation models, such as source separation, dereverberation, echo cancellation, and parametric spatial audio rendering. Georg Götz, Sebastian J. Schlecht, Ville Pulkki |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Decorrelation in Feedback Delay NetworksabstractThe feedback delay network (FDN) is a popular filter structure to generate artificial spatial reverberation. A common requirement for multichannel late reverberation is that the output signals are well decorrelated, as too high a correlation can lead to poor reproduction of source image and uncontrolled coloration. This article presents the analysis of multichannel correlation induced by FDNs. It is shown that the correlation depends primarily on the feedforward paths, while the long reverberation tail produced by the recursive path does not contribute to the inter-channel correlation. The impact of the feedback matrix type, size, and delays on the inter-channel correlation is demonstrated. The results show that small FDNs with a few feedback channels tend to have a high inter-channel correlation, and that the use of a filter feedback matrix significantly improves the decorrelation, often leading to the lowest inter-channel correlation among the tested cases. The learnings of this work support the practical design of multichannel artificial reverberators for immersive audio applications. Sebastian J. Schlecht, Jon Fagerström, Vesa Välimäki |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Audio Peak Reduction Using a Synced allpass FilterabstractPeak reduction is a common step used in audio playback chains to increase the loudness of a sound. The distortion introduced by a conventional nonlinear compressor can be avoided with the use of an allpass filter, which provides peak reduction by acting on the signal phase. This way, the signal energy around a waveform peak can be smeared while maintaining the total energy of the signal. In this paper, a new technique for linear peak amplitude reduction is proposed based on a Schroeder allpass filter, whose delay line and gain parameters are synced to match peaks of the signal’s auto-correlation function. The proposed method is compared with a previous search method and is shown to be often superior. An evaluation conducted over a variety of test signals indicates that the achieved peak reduction spans from 0 to 5 dB depending on the input waveform. The proposed method is widely applicable to real-time sound reproduction with a minimal computational processing budget. Sebastian J. Schlecht, Leonardo Fierro, Vesa Välimäki, Juha Backman |
ICASSP | 1 |
| 2021 | One-to-Many Conversion for Percussive SamplesabstractA filtering algorithm for generating subtle random variations in sampled sounds is proposed. Using only one recording for impact sound effects or drum machine sounds results in unrealistic repetitiveness during consecutive playback. This paper studies spectral variations in repeated knocking sounds and in three drum sounds: a hihat, a snare, and a tomtom. The proposed method uses a short pseudo-random velvet-noise filter and a low-shelf filter to produce timbral variations targeted at appropriate spectral regions, yielding potentially an endless number of new realistic versions of a single percussive sampled sound. The realism of the resulting processed sounds is studied in a listening test. The results show that the sound quality obtained with the proposed algorithm is at least as good as that of a previous method while using 77% fewer computational operations. The algorithm is widely applicable to computer-generated music and game audio. Jon Fagerström, Sebastian J. Schlecht, Vesa Välimäki |
DAFx | 2 |
| 2021 | The Role of Modal Excitation in Colorless ReverberationabstractA perceptual study revealing a novel connection between modal properties of feedback delay networks (FDNs) and colorless reverberation is presented. The coloration of the reverberation tail is quantified by the modal excitation distribution derived from the modal decomposition of the FDN. A homogeneously decaying allpass FDN is designed to be colorless such that the corresponding narrow modal excitation distribution leads to a high perceived modal density. Synthetic modal excitation distributions are generated to match modal excitations of FDNs. Three listening tests were conducted to demonstrate the correlation between the modal excitation distribution and the perceived degree of coloration. A fourth test shows a significant reduction of coloration by the colorless FDN compared to other FDN designs. The novel connection of modal excitation, allpass FDNs, and perceived coloration presents a beneficial design criterion for colorless artificial reverberation. Janis Heldmann, Sebastian J. Schlecht |
DAFx | 2 |
| 2021 | Acoustic Analysis and Dataset of Transitions Between Coupled RoomsabstractThe measurement of room acoustics plays a wide role in audio research, from physical acoustics modelling and virtual reality applications to speech enhancement. While vast literature exists on position-dependent room acoustics and coupling of rooms, little has explored the transition from one room to its neighbour. This paper presents the measurement and analysis of a dataset of spatial room impulse responses for the transition between four coupled room pairs. Each transition consists of 101 impulse responses recorded using a fourth-order spherical microphone array in 5 cm intervals, both with and without a continuous line-of-sight between the source and microphone. A numerical analysis of the room transitions is then presented, including direct-to-reverberant ratio and direction of arrival estimations, along with potential applications and uses of the dataset. Thomas McKenzie, Sebastian J. Schlecht, Ville Pulkki |
ICASSP | 2 |
| 2020 | Scattering in Feedback Delay NetworksabstractFeedback delay networks (FDNs) are recursive filters, which are widely used for artificial reverberation and decorrelation. One central challenge in the design of FDNs is the generation of sufficient echo density in the impulse response without compromising the computational efficiency. In a previous contribution, we have demonstrated that the echo density of an FDN can be increased by introducing so-called delay feedback matrices where each matrix entry is a scalar gain and a delay. In this contribution, we generalize the feedback matrix to arbitrary lossless filter feedback matrices (FFMs). As a special case, we propose the velvet feedback matrix, which can create dense impulse responses at a minimal computational cost. Further, FFMs can be used to emulate the scattering effects of non-specular reflections. We demonstrate the effectiveness of FFMs in terms of echo density and modal distribution. Sebastian J. Schlecht, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Perceptual Study of Near-Field Binaural Audio Rendering in Six-Degrees-of-Freedom Virtual RealityabstractAuditory localization cues in the near-field are significantly different than in the far-field. The near-field region is within an arm's length of the listener allowing to integrate proprioceptive cues to determine the location of an object in space. This perceptual study compares three non-individualized methods to apply head-related transfer functions (HRTFs) in six-degrees-of-freedom near-field audio rendering, namely, far-field measured HRTFs, multi-distance measured HRTFs, and spherical-model-based HRTFs with near-field extrapolation. To set our findings in context, we provide a real-world hand-held audio source for comparison along with a distance-invariant condition. Two modes of interaction are compared in an audio-visual virtual reality: one allowing the participant to move the audio object dynamically and the other with a stationary audio object but a freely moving listener. Olli S. Hurnmukainen, Sebastian J. Schlecht, Thomas Robotham, Axel Plinge, Emanuël A. P. Habets |
VR | 2 |
| 2017 | Evaluation of binaural reproduction systems from behavioral patterns in a six-degrees-of-freedom wayfinding taskabstractThis paper proposes a new method for evaluating real-time binaural reproduction systems by means of a wayfinding task in six degrees of freedom. Participants physically walk to sound objects in a virtual reality created by a head-mounted display and binaural audio. We show how the localization accuracy of spatial audio rendering is reflected by objective measures of the participants' behavior. The method allows for comparative evaluation of different rendering systems as well as the subjective assessment of the quality of experience. Olli Rummukainen, Sebastian J. Schlecht, Axel Plinge, Emanuël A. P. Habets |
QoMEX | 2 |
| 2017 | Feedback Delay Networks: Echo Density and Mixing TimeabstractFeedback delay networks (FDNs) are frequently used to generate artificial reverberation. This paper discusses the temporal features of impulse responses produced by FDNs, i.e., the number of echoes per time unit and its evolution over time. This so-called echo density is related to known measures of mixing time and their psychoacoustic correlates such as auditive perception of the room size. It is shown that the echo density of FDNs follows a polynomial function, whereby the polynomial coefficients can be derived from the lengths of the delays for which an explicit method is given. The mixing time of impulse responses can be predicted from the echo density, and conversely, a desired mixing time can be achieved by a derived mean delay length. A Monte Carlo simulation confirms the accuracy of the derived relation of mixing time and delay lengths. Sebastian J. Schlecht, Emanuël A. P. Habets |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |