Vesa Välimäki

dblp:94/3435 · DBLP profile ↗
← Back
95ranked-venue papers
19as first author
24since 2021 · last 2026
0000-0002-7869-292XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 68 · 11 first-author · 16 since 2021Artificial intelligence and machine learning · 25 · 7 first-author · 7 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Suppression of Nyquist Ringing in FFT-Based Sample Rate Conversion
abstract
Sample rate conversion, a common task in audio signal processing, can be performed with high quality using the fast Fourier transform (FFT) on the whole audio file. Before returning to the time domain using the inverse FFT, the sample rate of the signal is changed by either truncating or zero-padding the frequency-domain buffer. This operation leaves a discontinuity in the spectrum, which causes time-domain ringing at that frequency. The ringing can be suppressed by tapering the highest frequency bins. This letter introduces the double Dolph-Chebyshev window, a frequency-domain tapering function with a configurable level of ringing outside its main lobe in the transform domain. In comparison to basic cosine tapering, the proposed method provides, for example, a 150-dB suppression 91% faster. This letter improves the accuracy of FFT-based sample rate conversion, making it a practical tool for signal processing.
Roope Salmi, Vesa Välimäki
IEEE Signal Process. Lett.2
2025 FLAMO: An Open-Source Library for Frequency-Domain Differentiable Audio Processing
abstract
We present FLAMO, a Frequency-sampling Library for Audio-Module Optimization designed to implement and optimize differentiable linear time-invariant audio systems. The library is open-source and built on the frequency-sampling filter design method, allowing for the creation of differentiable modules that can be used stand-alone or within the computation graph of neural networks, simplifying the development of differentiable audio systems. It includes predefined filtering modules and auxiliary classes for constructing, training, and logging the optimized systems, all accessible through an intuitive interface. Practical application of these modules is demonstrated through two case studies: the optimization of an artificial reverberator and an active acoustics system for improved response coloration.
Gloria Dal Santo, Gian Marco De Bortoli, Karolina Prawda, Sebastian J. Schlecht, Vesa Välimäki
ICASSP5
2025 HRTF Estimation using a Score-based Prior
abstract
We present a head-related transfer function (HRTF) estimation method which relies on a data-driven prior given by a score-based diffusion model. The HRTF is estimated in reverberant environments using natural excitation signals, e.g. human speech. The impulse response of the room is estimated along with the HRTF by optimizing a parametric model of reverberation based on the statistical behaviour of room acoustics. The posterior distribution of HRTF given the reverberant measurement and excitation signal is modelled using the score-based HRTF prior and a log-likelihood approximation. We show that the resulting method outperforms several baselines, including an oracle recommender system that assigns the optimal HRTF in our training set based on the smallest distance to the true HRTF at the given direction of arrival. In particular, we show that the diffusion prior can account for the large variability of high-frequency content in HRTFs.
Etienne Thuillier, Jean-Marie Lemercier, Eloi Moliner, Timo Gerkmann, Vesa Välimäki
ICASSP5
2025 Multi-shelf graphic equalizer
abstract
A graphic equalizer (GEQ) is a standard tool in audio production and effect design. Adjustable gain control frequencies are fixed along the logarithmic frequency axis, and an automatic design method matches the magnitude response to them whenever target gains are changed. Most commonly, the GEQ comprises a set of peak filters centered an octave apart, possibly with a shelving filter at the bottom and top of the frequency range. While accurate designs were proposed, the dynamic range is typically limited to 24 dB. In this paper, we propose two innovations. First, we introduce a GEQ based on shelving filters only, which can cover an extensive dynamic range of over 60 dB. Secondly, we introduce an order-switching technique that combines shelf filters of different order. We demonstrate the performance and advantages of the proposed filter with design examples. The proposed shelf-filter-based GEQ offers a wider dynamic range and a smoother magnitude response than traditional peak-filter-based GEQ designs.
Sebastian J. Schlecht, Tantep Sinjanakhom, Vesa Välimäki
Signal Process.3
2024 Noise Morphing for Audio Time Stretching
abstract
This letter introduces an innovative method to enhance the quality of audio time stretching by precisely decomposing a sound into sines, transients, and noise and by improving the processing of the latter component. While there are established methods for time-stretching sines and transients with high quality, the manipulation of noise or residual components has lacked robust solutions in prior research. The proposed method combines sound decomposition with previous techniques for audio spectral resynthesis. The time-stretched noise component is achieved by morphing its time-interpolated spectral magnitude with a white-noise excitation signal. This method stands out for its simplicity, efficiency, and audio quality. The results of a subjective experiment affirm the superiority of this approach over current state-of-the-art methods across all evaluated stretch factors. The proposed technique notably excels in extreme stretching scenarios, signifying a substantial elevation in performance. The proposed method holds promise for a wide range of applications in slow-motion media content, such as music or sports video production.
Eloi Moliner, Leonardo Fierro, Alec Wright, Matti S. Hämäläinen, Vesa Välimäki
IEEE Signal Process. Lett.5
2024 Modal Excitation in Feedback Delay Networks
abstract
Feedback delay networks (FDNs) are used in audio processing and synthesis. The modal shapes of the system describe the modal excitation by input and output signals. Previously, the Ehrlich-Aberth method was used to find modes in large FDNs. Here, the method is extended to the corresponding eigenvectors indicating the modal shape. In particular, the computational complexity of the proposed analysis method does not depend on the delay-line lengths and is thus suitable for large FDNs, such as artificial reverberators. We show the relation between the compact generalized eigenvectors in the delay state space and the spatially extended modal shapes in the state space. We illustrate this method with an example FDN in which the suggested modal excitation control does not increase the computational cost. The modal shapes can help optimize input and output gains. This letter teaches how selecting the input and output points along the delay lines of an FDN adjusts the spectral shape of the system output.
Sebastian J. Schlecht, Matteo Scerbo, Enzo De Sena, Vesa Välimäki
IEEE Signal Process. Lett.4
2024 Two-Stage Attenuation Filter for Artificial Reverberation
abstract
Delay networks are a common parametric method to synthesize the late part of the room reverberation. A delay network consists of several feedback loops, each containing a delay line and an attenuation filter, which approximates the same decay rate by appropriately setting the frequency-dependent loop gain. A remaining challenge is the design of the attenuation filters on a wide frequency range based on a measured room impulse response. This letter proposes a novel two-stage attenuation filter structure, sharpening the design. The first stage is a low-order pre-filter approximating the overall shape and determining the decay at the two ends of the frequency range, namely at the dc and the Nyquist limit. The second filter, an equalizer, fine-tunes the gain at different frequencies, such as on one-third-octave bands. It is shown that the proposed design is more accurate and robust than previous methods. A design example applying the proposed method to an interleaved velvet-noise reverberator is also exhibited. The proposed two-stage attenuation filter is a step toward a realistic parametric simulation of measured room impulse responses.
Vesa Välimäki, Karolina Prawda, Sebastian J. Schlecht
IEEE Signal Process. Lett.1
2024 Blind Audio Bandwidth Extension: A Diffusion-Based Zero-Shot Approach
abstract
Audio bandwidth extension involves the realistic reconstruction of high-frequency spectra from bandlimited observations. In cases where the lowpass degradation is unknown, such as in restoring historical audio recordings, this becomes a blind problem. This paper introduces a novel method called BABE (Blind Audio Bandwidth Extension) that addresses the blind problem in a zero-shot setting, leveraging the generative priors of a pre-trained unconditional diffusion model. During the inference process, BABE utilizes a generalized version of diffusion posterior sampling, where the degradation operator is unknown but parametrized and inferred iteratively. The performance of the proposed method is evaluated using objective and subjective metrics, and the results show that BABE surpasses state-of-the-art blind bandwidth extension baselines and achieves competitive performance compared to informed methods when tested with synthetic data. Moreover, BABE exhibits robust generalization capabilities when enhancing real historical recordings, effectively reconstructing the missing high-frequency content while maintaining coherence with the original recording. Subjective preference tests confirm that BABE significantly improves the audio quality of historical music recordings. Examples of historical recordings restored with the proposed method are available on the companion webpage:http://research.spa.aalto.fi/publications/papers/ieee-taslp-babe/
Eloi Moliner, Filip Elvander, Vesa Välimäki
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 HRTF Interpolation Using a Spherical Neural Process Meta-Learner
abstract
Several individualization methods have recently been proposed to estimate a subject's Head-Related Transfer Function (HRTF) using convenient input modalities such as anthropometric measurements or pinnae photographs. There exists a need for adaptively correcting the estimation error committed by such methods using a few data point samples from the subject's HRTF, acquired using acoustic measurements or perceptual feedback. To facilitate this, we introduce a Convolutional Conditional Neural Process meta-learner specialized in HRTF error interpolation. In particular, the model includes a Spherical Convolutional Neural Network component to accommodate the spherical geometry of HRTF data. It also exploits potential symmetries between the HRTF's left and right channels about the median plane. In this work, we evaluate the proposed model's performance purely on time-aligned spectrum interpolation grounds under a simplified setup where a generic population-mean HRTF forms the initial estimates prior to corrections instead of individualized ones. The trained model achieves up to 3 dB relative error reduction compared to state-of-the-art interpolation methods despite being trained using only 85 subjects. This improvement translates up to nearly a halving of the data point count required to achieve comparable accuracy, in particular from 50 to 28 points to reach an average of -20 dB relative error per interpolated feature. Moreover, we show that the trained model provides well-calibrated uncertainty estimates. Accordingly, such estimates could inform the sequential decision problem of acquiring as few correcting HRTF data points as needed to meet a desired level of HRTF individualization accuracy.
Etienne Thuillier, Craig T. Jin, Vesa Välimäki
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Extreme Audio Time Stretching Using Neural Synthesis
abstract
A deep neural network solution for time-scale modification (TSM) focused on large stretching factors is proposed, targeting environmental sounds. Traditional TSM artifacts such as transient smearing, loss of presence, and phasiness are heavily accentuated and cause poor audio quality when the TSM factor is four or larger. The weakness of established TSM methods, often based on a phase vocoder structure, lies in the poor description and scaling of the transient and noise components, or nuances, of a sound. Our novel solution combines a sines-transients-noise decomposition with an independent WaveNet synthesizer to provide a better description of the noise component and an improve sound quality for large stretching factors. Results of a subjective listening test against four other TSM algorithms are reported, showing the proposed method to be often superior. The proposed method is stereo compatible and has a wide range of applications related to the slow motion of media content.
Leonardo Fierro, Alec Wright, Vesa Välimäki, Matti S. Hämäläinen
ICASSP3
2023 Solving Audio Inverse Problems with a Diffusion Model
abstract
This paper presents CQT-Diff, a data-driven generative audio model that can, once trained, be used for solving various different audio inverse problems in a problem-agnostic setting. CQT-Diff is a neural diffusion model with an architecture that is carefully constructed to exploit pitch-equivariant symmetries in music. This is achieved by preconditioning the model with an invertible Constant-Q Transform (CQT), whose logarithmically-spaced frequency axis represents pitch equivariance as translation equivariance. The proposed method is evaluated with solo piano music, using objective and subjective metrics in three different and varied tasks: audio bandwidth extension, inpainting, and declipping. The results show that CQT-Diff outperforms the compared baselines and ablations in audio bandwidth extension and, without retraining, delivers competitive performance against modern baselines in audio inpainting and declipping. This work represents the first diffusion-based general framework for solving inverse problems in audio processing.
Eloi Moliner, Jaakko Lehtinen, Vesa Välimäki
ICASSP3
2023 Adversarial Guitar Amplifier Modelling with Unpaired Data
abstract
We propose an audio effects processing framework that learns to emulate a target electric guitar tone from a recording. We train a deep neural network using an adversarial approach, with the goal of trans-forming the timbre of a guitar, into the timbre of another guitar after audio effects processing has been applied, for example, by a guitar amplifier. The model training requires no paired data, and the resulting model emulates the target timbre well whilst being capable of real-time processing on a modern personal computer. To verify our approach we present two experiments, one which carries out un-paired training using paired data, allowing us to monitor training via objective metrics, and another that uses fully unpaired data, corresponding to a realistic scenario where a user wants to emulate a guitar timbre only using audio data from a recording. Our listening test results confirm that the models are perceptually convincing.
Alec Wright, Vesa Välimäki, Lauri Juvela
ICASSP2
2023 BEHM-GAN: Bandwidth Extension of Historical Music Using Generative Adversarial Networks
abstract
Audio bandwidth extension aims to expand the spectrum of bandlimited audio signals. Although this topic has been broadly studied during recent years, the particular problem of extending the bandwidth of historical music recordings remains an open challenge. This paper proposes a method for the bandwidth extension of historical music using generative adversarial networks (BEHM-GAN) as a practical solution to this problem. The proposed method works with the complex spectrogram representation of audio and, thanks to a dedicated regularization strategy, can effectively extend the bandwidth of out-of-distribution real historical recordings. The BEHM-GAN is designed to be applied as a second step after denoising the recording to suppress any additive disturbances, such as clicks and background noise. We train and evaluate the method using solo piano classical music. The proposed method outperforms the compared baselines in both objective and subjective experiments. The results of a formal blind listening test show that BEHM-GAN significantly increases the perceptual sound quality in early-20th-century gramophone recordings. For several items, there is a substantial improvement in the mean opinion score after enhancing historical recordings with the proposed bandwidth-extension algorithm. This study represents a relevant step toward data-driven music restoration in real-world scenarios.
Eloi Moliner, Vesa Välimäki
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Decorrelation in Feedback Delay Networks
abstract
The feedback delay network (FDN) is a popular filter structure to generate artificial spatial reverberation. A common requirement for multichannel late reverberation is that the output signals are well decorrelated, as too high a correlation can lead to poor reproduction of source image and uncontrolled coloration. This article presents the analysis of multichannel correlation induced by FDNs. It is shown that the correlation depends primarily on the feedforward paths, while the long reverberation tail produced by the recursive path does not contribute to the inter-channel correlation. The impact of the feedback matrix type, size, and delays on the inter-channel correlation is demonstrated. The results show that small FDNs with a few feedback channels tend to have a high inter-channel correlation, and that the use of a filter feedback matrix significantly improves the decorrelation, often leading to the lowest inter-channel correlation among the tested cases. The learnings of this work support the practical design of multichannel artificial reverberators for immersive audio applications.
Sebastian J. Schlecht, Jon Fagerström, Vesa Välimäki
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Pruning Deep Neural Network Models of Guitar Distortion Effects
abstract
Deep neural networks have been successfully used in the task of black-box modeling of analog audio effects such as distortion. Improving the processing speed and memory requirements of the inference step is desirable to allow such models to be used on a wide range of hardware and concurrently with other software. In this paper, we propose a new application of recent advancements in neural network pruning methods to recurrent black-box models of distortion effects using a Long Short-Term Memory architecture. We compare the efficacy of the method on four different datasets; one distortion pedal and three vacuum tube amplifiers. Iterative magnitude pruning allows us to remove over 99% of parameters from some models without a loss of accuracy. We evaluate the real-time performance of the pruned models and find that a 3x-4x speedup can be achieved, compared to an unpruned baseline. We show that training a larger model and then pruning it outperforms an unpruned model of equivalent hidden size. A listening test confirms that pruning does not degrade the perceived sound quality, but may even slightly improve it. The proposed techniques can be used to design computationally efficient deep neural networks for processing the sound of the electric guitar in real time.
David Südholt, Alec Wright, Cumhur Erkut, Vesa Välimäki
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 A Two-Stage U-Net for High-Fidelity Denoising of Historical Recordings
abstract
Enhancing the sound quality of historical music recordings is a long-standing problem. This paper presents a novel denoising method based on a fully-convolutional deep neural network. A two-stage U-Net model architecture is designed to model and suppress the degradations with high fidelity. The method processes the time-frequency representation of audio, and is trained using realistic noisy data to jointly remove hiss, clicks, thumps, and other common additive disturbances from old analog discs. The proposed model outperforms previous methods in both objective and subjective metrics. The results of a formal blind listening test show that real gramophone recordings denoised with this method have significantly better quality than the baseline methods. This study shows the importance of realistic training data and the power of deep learning in audio restoration.
Eloi Moliner, Vesa Välimäki
ICASSP2
2022 Audio Peak Reduction Using a Synced allpass Filter
abstract
Peak reduction is a common step used in audio playback chains to increase the loudness of a sound. The distortion introduced by a conventional nonlinear compressor can be avoided with the use of an allpass filter, which provides peak reduction by acting on the signal phase. This way, the signal energy around a waveform peak can be smeared while maintaining the total energy of the signal. In this paper, a new technique for linear peak amplitude reduction is proposed based on a Schroeder allpass filter, whose delay line and gain parameters are synced to match peaks of the signal’s auto-correlation function. The proposed method is compared with a previous search method and is shown to be often superior. An evaluation conducted over a variety of test signals indicates that the achieved peak reduction spans from 0 to 5 dB depending on the input waveform. The proposed method is widely applicable to real-time sound reproduction with a minimal computational processing budget.
Sebastian J. Schlecht, Leonardo Fierro, Vesa Välimäki, Juha Backman
ICASSP3
2022 Sparse Graphic Equalizer Design
abstract
A typical graphic equalizer frequency resolution is one-third octave comprising 31 bands. A previous design based on a least-squares optimization of the band-filter gains with a single second-order section per band has an accuracy of 1 dB. However, the design always uses all the band filters even when a small number of gains is adjusted. This letter proposes a sparse design of a one-third-octave graphic equalizer, where the number of active bands is minimized using orthogonal matching pursuit and linear programming before the least-squares gain optimization. In addition, the bandwidths of the high-frequency band filters are gain dependent in order to optimize their shape relative to an analog prototype. The optimal bandwidth is obtained with linear interpolation during the filter design. The proposed design achieves approximately the same or better accuracy in comparison to the state-of-the-art non-sparse design and can be used to automatically reduce the computational load of equalization when only some band filters are active.
Mario Antonelli, Juho Liski, Vesa Välimäki
IEEE Signal Process. Lett.3
2022 Multicore implementation of a multichannel parallel graphic equalizer
abstract
Abstract Numerous signal processing applications are emerging on mobile computing systems. These applications are subject to responsiveness constraints for user interactivity and, at the same time, must be optimized for energy efficiency. Many current embedded devices are composed of low-power multicore processors that offer a good trade-off between computational capacity and low power consumption. In this context, equalizers are widely used in multiple mobile-based applications such as “Music streaming” to adjust the levels of bass and treble in sound reproduction. In this study, we evaluate a graphic equalizer from audio, computational capacity, and energy efficiency perspectives, as well as the execution of multiple real-time equalizers running on an embedded quad-core processor of a mobile device. To this end, we experiment with the working frequencies as well as the parallelism that can be extracted from a quad-core ARM Cortex-A57. Results show that using high CPU frequencies and three or four cores, our parallel algorithm is able to equalize more than five channels per watt in real time with an audio buffer of 4096 samples, which implies a latency of 92.8 ms at the standard sample rate of 44.1 kHz.
Jose A. Belloch, José M. Badía, German Leon, Balázs Bank, Vesa Välimäki
J. Supercomput.5
2021 One-to-Many Conversion for Percussive Samples
abstract
A filtering algorithm for generating subtle random variations in sampled sounds is proposed. Using only one recording for impact sound effects or drum machine sounds results in unrealistic repetitiveness during consecutive playback. This paper studies spectral variations in repeated knocking sounds and in three drum sounds: a hihat, a snare, and a tomtom. The proposed method uses a short pseudo-random velvet-noise filter and a low-shelf filter to produce timbral variations targeted at appropriate spectral regions, yielding potentially an endless number of new realistic versions of a single percussive sampled sound. The realism of the resulting processed sounds is studied in a listening test. The results show that the sound quality obtained with the proposed algorithm is at least as good as that of a previous method while using 77% fewer computational operations. The algorithm is widely applicable to computer-generated music and game audio.
Jon Fagerström, Sebastian J. Schlecht, Vesa Välimäki
DAFx3
2021 Sitrano: A Matlab App for Sines-Transients-Noise Decomposition of Audio Signals
abstract
Decomposition of sounds into their sinusoidal, transient, and noise components is an active research topic and a widely-used tool in audio processing. Multiple solutions have been proposed in recent years, using time-frequency representations to identify either horizontal and vertical structures or orientations and anisotropy in the spectrogram of the sound. In this paper, we present SiTraNo: an easy-to-use MATLAB application with a graphic user interface for audio decomposition that enables visualization and access to the sinusoidal, transient, and noise classes, individually. This application allows the user to choose between different well-known separation methods to analyze an input sound file, to instantaneously control and remix its spectral components, and to visually check the quality of the separation, before producing the desired output file. The visualization of common artifacts, such as birdies and dropouts, is demonstrated. This application promotes experimenting with the sound decomposition process by observing the effect of variations for each spectral component on the original sound and by comparing different methods against each other, evaluating the separation quality both audibly and visually. SiTraNo and its source code are available on a companion website and repository.
Leonardo Fierro, Vesa Välimäki
DAFx2
2021 Exposure Bias and State Matching in Recurrent Neural Network Virtual Analog Models
abstract
Virtual analog (VA) modeling using neural networks (NNs) has great potential for rapidly producing high-fidelity models. Recurrent neural networks (RNNs) are especially appealing for VA due to their connection with discrete nodal analysis. Furthermore, VA models based on NNs can be trained efficiently by directly exposing them to the circuit states in a gray-box fashion. However, exposure to ground truth information during training can leave the models susceptible to error accumulation in a free-running mode, also known as “exposure bias” in machine learning literature. This paper presents a unified framework for treating the previously proposed state trajectory network (STN) and gated recurrent unit (GRU) networks as special cases of discrete nodal analysis. We propose a novel circuit state-matching mechanism for the GRU and experimentally compare the previously mentioned networks for their performance in state matching, during training, and in ex-posure bias, during inference. Experimental results from modeling a diode clipper show that all the tested models exhibit some exposure bias, which can be mitigated by truncated backpropagation through time. Furthermore, the proposed state matching mechanism improves the GRU modeling performance of an overdrive pedal and a phaser pedal, especially in the presence of external modulation, apparent in a phaser circuit.
Aleksi Peussa, Eero-Pekka Damskägg, Thomas Sherson, Stylianos I. Mimilakis, Lauri Juvela, Athanasios Gotsopoulos, Vesa Välimäki
DAFx7
2021 Audibility of Group-Delay Equalization
abstract
This paper discusses the audibility of group-delay variations. Previous research has found limits of audibility as a function of frequency for different test signals, but extracting the tolerance for group delay to help audio reproduction system designers is hard. This study considers four critical test signals, three synthetic and one recorded, modified with digital allpass filters. The signals are filtered to produce a positive or negative group-delay peak covering the most sensitive frequency range from 500 Hz to 4 kHz, without changing the delay at other frequencies. ABX listening tests using headphones reveal the audibility thresholds for each signal. The perception is highly dependent on the signal, and the unit impulse and pink impulse are the most critical test signals. Negative group-delay variations are more easily audible than positive ones. The smallest mean threshold for the negative group delay was $-$0.56 ms and 0.64 ms for the positive group delay, obtained with a pink impulse. The thresholds are smaller than those obtained in previous studies. A synthetic hi-hat sound decaying 60 dB in 80 ms hides a positive group-delay variation. The variation is more difficult to hear in a recorded castanet sound than in the most critical synthetic signals. This work demonstrates how the group-delay response of headphones and loudspeakers can be perceptually tested, and leads to a better understanding of how audio systems should be equalized to avoid audible group-delay distortion.
Juho Liski, Aki Mäkivirta, Vesa Välimäki
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Late-Reverberation Synthesis Using Interleaved Velvet-Noise Sequences
abstract
This paper proposes a novel algorithm for simulating the late part of room reverberation. A well-known fact is that a room impulse response sounds similar to exponentially decaying filtered noise some time after the beginning. The algorithm proposed here employs several velvet-noise sequences in parallel and combines them so that their non-zero samples never occur at the same time. Each velvet-noise sequence is driven by the same input signal but is filtered with its own feedback filter which has the same delay-line length as the velvet-noise sequence. The resulting response is sparse and consists of filtered noise that decays approximately exponentially with a given frequency-dependent reverberation time profile. We show via a formal listening test that four interleaved branches are sufficient to produce a smooth high-quality response. The outputs of the branches connected in different combinations produce decorrelated output signals for multichannel reproduction. The proposed method is compared with a state-of-the-art delay-based reverberation method and its advantages are pointed out. The computational load of the method is 60% smaller than that of a comparable existing method, the feedback delay network. The proposed method is well suited to the synthesis of diffuse late reverberation in audio and music production.
Vesa Välimäki, Karolina Prawda
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Perceptual loss function for neural modeling of audio systems
abstract
This work investigates alternate pre-emphasis filters used as part of the loss function during neural network training for nonlinear audio processing. In our previous work, the error-to-signal ratio loss function was used during network training, with a first-order high-pass pre-emphasis filter applied to both the target signal and neural network output. This work considers more perceptually relevant pre-emphasis filters, which include low-pass filtering at high frequencies. We conducted listening tests to determine whether they offer an improvement to the quality of a neural network model of a guitar tube amplifier. Listening test results indicate that the use of an A-weighting pre-emphasis filter offers the best improvement among the tested filters. The proposed perceptual loss function improves the sound quality of neural network models in audio processing without affecting the computational cost.
Alec Wright, Vesa Välimäki
ICASSP2
2019 Deep Learning for Tube Amplifier Emulation
abstract
Analog audio effects and synthesizers often owe their distinct sound to circuit nonlinearities. Faithfully modeling such significant aspect of the original sound in virtual analog software can prove challenging. The current work proposes a generic data-driven approach to virtual analog modeling and applies it to the Fender Bassman 56F-A vacuum-tube amplifier. Specifically, a feedforward variant of the WaveNet deep neural network is trained to carry out a regression on audio waveform samples from input to output of a SPICE model of the tube amplifier. The output signals are pre-emphasized to assist the model at learning the high-frequency content. The results of a listening test suggest that the proposed model accurately emulates the reference device. In particular, the model responds to user control changes, and faithfully restitutes the range of sonic characteristics found across the configurations of the original device.
Eero-Pekka Damskägg, Lauri Juvela, Etienne Thuillier, Vesa Välimäki
ICASSP4
2019 Graphic Delay Equalizer
abstract
A graphic delay equalizer based on a high-order nonparametric allpass filter design is proposed. Command points at the centers of octave frequency bands are connected with polynomial interpolation to form a continuous target group-delay curve as function of frequency. The required number of all-pass sections depends on the area under the target curve. The group-delay area at low audio frequencies is small due to the linear frequency scale, so only a few allpass sections can be assigned there, which reduces the accuracy. Design accuracy can be improved by adding a constant delay to the target curve, which increases the area. Two use cases are presented: the group-delay equalization of a multi-way loudspeaker and the linearization of the phase response of a regular magnitude-only graphic equalizer. The graphic delay equalizer can be used to enhance the perceptual quality of audio systems or to produce audio effects.
Jussi Rämö, Vesa Välimäki
ICASSP2
2019 Neurally Controlled Graphic Equalizer
abstract
This paper describes a neural network based method to simplify the design of a graphic equalizer without sacrificing the accuracy of approximation. The key idea is to train a neural network to predict the mapping from target gains to the optimized band filter gains at specified center frequencies. The prediction is implemented with a feedforward neural network having a hidden layer with 20 neurons in the case of the ten-octave graphic equalizer. The band filter coefficients can then be quickly and easily computed using closed-form formulas. This work turns, for the first time, the accurate graphic equalization design into a feedforward calculation without matrix inversion or iterations. The filter gain control using the neural network reduces the computing time by 99.6% in comparison to the least-squares design method it is imitating and contributes an approximation error of less than 0.1 dB. The resulting neurally controlled graphic equalizer will be highly useful in various audio and music processing applications, which require time-varying equalization.
Vesa Välimäki, Jussi Rämö
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Antiderivative Antialiasing for Memoryless Nonlinearities
abstract
Aliasing is a commonly encountered problem in audio signal processing, particularly when memoryless nonlinearities are simulated in discrete time. A conventional remedy is to operate at an oversampled rate. A new aliasing reduction method is proposed here for discrete-time memoryless nonlinearities, which is suitable for operation at reduced oversampling rates. The method employs higher order antiderivatives of the nonlinear function used. The first-order form of the new method is equivalent to a technique proposed recently by Parker et al. Higher order extensions offer considerable improvement over the first antiderivative method, in terms of the signal-to-noise ratio. The proposed methods can be implemented with fewer operations than oversampling and are applicable to discrete-time modeling of a wide range of nonlinear analog systems.
Stefan Bilbao, Fabian Esqueda, Julian Parker, Vesa Välimäki
IEEE Signal Process. Lett.4
2017 Accurate Cascade Graphic Equalizer
abstract
A graphic equalizer is a high-order filter controlling the gain of several frequency bands. For good accuracy, graphic equalizers consisting of cascaded IIR filters have been of very high order. A previously proposed parallel graphic equalizer entailing twice as many second-order filter sections as there are bands can have a maximum approximation error of less than 1 dB, but its design is complicated. This letter proposes a cascade graphic equalizer having an accuracy comparable to the best parallel graphic equalizer, although only one second-order section is assigned per command gain. A key idea is to use band filters whose interaction with the two neighboring filters at their center frequency is exactly controlled. The filter gains are obtained using the least-squares method with one iteration step, which involves linear interpolation of the target gain vector, inversion of a square matrix, and a few matrix multiplications. The proposed method is compared with previous designs and is shown to be the most accurate one. The new graphic equalizer is widely useful for audio and music processing.
Vesa Välimäki, Juho Liski
IEEE Signal Process. Lett.1
2017 GPU-Based Dynamic Wave Field Synthesis Using Fractional Delay Filters and Room Compensation
abstract
Wave field synthesis (WFS) is a multichannel audio reproduction method, of a considerable computational cost that renders an accurate spatial sound field using a large number of loudspeakers to emulate virtual sound sources. The moving of sound source locations can be improved by using fractional delay filters, and room reflections can be compensated by using an inverse filter bank that corrects the room effects at selected points within the listening area. However, both the fractional delay filters and the room compensation filters further increase the computational requirements of the WFS system. This paper analyzes the performance of a WFS system composed of 96 loudspeakers which integrates both strategies. In order to deal with the large computational complexity, we explore the use of a graphics processing unit (GPU) as a massive signal co-processor to increase the capabilities of the WFS system. The performance of the method as well as the benefits of the GPU acceleration are demonstrated by considering different sizes of room compensation filters and fractional delay filters of order 9. The results show that a 96-speaker WFS system that is efficiently implemented on a state-of-art GPU can synthesize the movements of 94 sound sources in real time and, at the same time, can manage 9216 room compensation filters having more than 4000 coefficients each.
Jose A. Belloch, Alberto González 0001, Enrique S. Quintana-Ortí, Miguel Ferrer 0001, Vesa Välimäki
IEEE ACM Trans. Audio Speech Lang. Process.5
2017 Modeling Sparsely Reflecting Outdoor Acoustic Scenes Using the Waveguide Web
abstract
Computer games and virtual reality require digital reverberation algorithms, which can simulate a broad range of acoustic spaces, including locations in the open air. Additionally, the detailed simulation of environmental sound is an area of significant interest due to the propagation of noise pollution over distances and its related impact on well-being, particularly in urban spaces. This paper introduces the waveguide web digital reverberator design for modeling the acoustics of sparsely reflecting outdoor environments; a design that is, in part, an extension of the scattering delay network reverberator. The design of the algorithm is based on a set of digital waveguides connected by scattering junctions at nodes that represent the reflection points of the environment under study. The structure of the proposed reverberator allows for accurate reproduction of reflections between discrete reflection points. Approximation errors are caused when the assumption of point-like nodes does not hold true. Three example cases are presented comparing waveguide web simulated impulse responses for a traditional shoebox room, a forest scenario, and an urban courtyard, with impulse responses created using other simulation methods or from real-world measurements. The waveguide web algorithm can better enable the acoustic simulation of outdoor spaces and so contribute toward sound design for virtual reality applications, gaming, and auralization, with a particular focus on acoustic design for the urban environment.
Francis Stevens, Damian T. Murphy, Lauri Savioja, Vesa Välimäki
IEEE ACM Trans. Audio Speech Lang. Process.4
2016 Efficient target-response interpolation for a graphic equalizer
abstract
A graphic equalizer is an adjustable filter in which the command gain of each frequency band is practically independent of the gains of other bands. Designing a graphic equalizer with a high precision requires evaluating a target response that interpolates the magnitude response at several frequency points between the command gains. Good accuracy has been previously achieved by using polynomial interpolation methods such as cubic Hermite or spline interpolation. However, these methods require large computational resources, which is a limitation in real-time applications. This paper proposes an efficient way of computing the target response without sacrificing the approximation accuracy. This new approach called Linear Interpolation with Constant Segments (LICS) reduces the computing time of the target response by 55% and has an intrinsic parallel structure. Performance of the LICS method is assessed on an ARM Cortex-A7 core, which is commonly used in embedded systems.
Jose A. Belloch, Vesa Välimäki
ICASSP2
2014 Perceptual Linear Filters: Low-Order ARMA Approximation for Sound Synthesis
Rémi Mignot, Vesa Välimäki
DAFx2
2014 Examining the Oscillator Waveform Animation Effect
Joseph Timoney, Victor Lazzarini, Jari Kleimola, Vesa Välimäki
DAFx4
2014 Multi-channel IIR filtering of audio signals using a GPU
abstract
In the audio signal processing field, multiple IIR filters are required in many applications. As an example, equalizing a Wave Field Synthesis system requires massive filter processing in real time. Graphics Processing Units (GPUs) are well known for their potential in highly parallel data processing. Up to now, the use of the GPUs for implementing IIR filters has not been clearly tackled in audio processing because of its feedback loop that prevents its total parallelization. However, using the Parallel form of IIR filters, this feedback is reduced, since every single sample is computed in a parallel way. This paper analyzes the performance of multiple IIR filters using GPUs and compares it with a powerful multi-core computer. The proposed GPU implementation can run up to 1256 concurrent IIR filters of order 256th in real time, which means 321,536 total filter order, with a latency time of 0.72 ms (sampling frequency of 44.1 kHz). This demonstrates that GPUs are well suited for computing massive IIR filtering.
Jose A. Belloch, Balázs Bank, Lauri Savioja, Alberto González 0001, Vesa Välimäki
ICASSP5
2014 A nonlinear second-order digital oscillator for Virtual Acoustic Feedback
abstract
The guitar feedback effect, or howling, is well known to the general public and identified with many rock music genres and it is the only case of acoustic feedback employed for musical purposes. Virtual Acoustic Feedback (VAF), is regarded as the extension of this phenomenon to any instrument or sound source by means of virtual acoustics and is meant to enrich the sound palette of a musician. The study of the acoustic feedback as a musical tool and computational techniques for its emulation have been scarcely addressed in literature. In this paper a nonlinear feedback oscillator is proposed and its properties derived. The oscillator does not necessarily need to be connected to a virtual instrument, thus enables to process any kind of pitched real-time input.
Leonardo Gabrielli, M. Giobbi, Stefano Squartini, Vesa Välimäki
ICASSP4
2014 True discrete cepstrum: An accurate and smooth spectral envelope estimation for music processing
abstract
In the tradition of the spectral envelope estimation of periodic sounds, we propose a new accurate method, called True Discrete Cepstrum. Solving a constrained optimization problem, it provides a smooth envelope which fits exactly the given peak values. Moreover, based on the auditory masking, we propose a release of the constraint which improves the smoothness, without perceptual change. Contrarily to some other methods, the parametrization of this method is easy, and it gives a complete control of the maximal deviation of the spectral envelope from the peak values. The benefit of this new method is illustrated and an evaluation procedure validates it with a comparison with some other methods.
Rémi Mignot, Vesa Välimäki
ICASSP2
2014 Optimizing a High-Order Graphic Equalizer for Audio Processing
abstract
A high-order graphic equalizer has the advantage that the gain in one band is highly independent of the gains in the adjacent bands. However, all practical filters have transition bands, which interact with the adjacent bands and create errors in the desired magnitude response. This letter proposes a filter optimization algorithm for a high-order graphic equalizer, which minimizes the errors in the transition bands by iteratively optimizing the orders of adjacent band filters. The optimization of the filter order affects the shape of the transition band, thus enabling the search for the optimum shape relative to the adjacent filter. The optimization is done offline, and during filtering only the gains of the band filters are altered. In an example case, the proposed method was able to meet the given peak-error limitations of ±2 dB, when the total order of the graphical equalizer was 328, whereas the non-optimized filter could not meet the requirements even when the total order was raised to 672. Optimized high-order graphical equalizers can be widely used in audio signal processing applications.
Jussi Rämö, Vesa Välimäki
IEEE Signal Process. Lett.2
2014 Generalized Moog ladder filter: part I-linear analysis and parameterization
abstract
The Moog ladder filter, which consists of four cascaded first-order ladder stages in a feedback loop, falls within the class of devices that have attracted greatest interest in virtual analog research. On one hand, this work confirms that the presence of exactly four stages in the original analog circuit is motivated by specific filter control issues and, on the other, that such a limitation can be overcome in the digital domain with relative ease. First, a continuous-time large-signal model is defined for a version of the circuit that is generalized to an arbitrary number of ladder stages. Then, the linear behavior of the filter around its natural operating point and the effect of control parameters on the resulting frequency response are studied in depth, to obtain exact analytical expressions for the position of poles in the transfer function and for the dc gain of the filter, as well as a parameterization strategy that is consistent for any number of ladder stages. A previously-introduced linear digital model of the device suggested by Smith is eventually generalized based on these general results, which remain, however, relevant and similarly applicable to other discretizations of the filter. The proposed model faithfully reproduces the linear behavior of the generalized device while providing sensible parametric control for any number of ladder stages.
Stefano D'Angelo, Vesa Välimäki
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 Generalized Moog ladder filter: part II-explicit nonlinear model through a novel delay-free loop implementation method
abstract
One of the most critical aspects of virtual analog simulation of circuits for music production consists in accurate reproduction of their nonlinear behavior, yet this goal is in many cases difficult to achieve due to the presence of implicit differential equations in circuit models, since they naturally map to delay-free loops in the digital domain. This paper presents a novel and general method for non-iteratively implementing these loops in such a way that the linear response around a chosen operating point is preserved, the topology is minimally affected, and transformation of nonlinearities is not required. This technique is then applied to a generalized model of the Moog ladder filter, resulting in an implementation that outperforms its predecessors with only a modest computational load penalty. This digital version of the filter is shown to offer strong stability guarantees w.r.t. parameter variation, allows the extraction of different frequency response modes by simple mixing of individual ladder stage outputs, and is suitable for real-time sound synthesis and audio effects processing.
Stefano D'Angelo, Vesa Välimäki
IEEE ACM Trans. Audio Speech Lang. Process.2
2014 High-precision parallel graphic equalizer
abstract
This paper proposes a high-precision graphic equalizer based on second-order parallel filters. Previous graphic equalizers suffer from interaction between adjacent band filters, especially at high gain values, which can lead to substantial errors in the magnitude response. The fixed-pole design of the proposed parallel graphic equalizer avoids this problem, since the parallel second-order filters are optimized jointly. When the number of pole frequencies is twice the number of command points of the graphic equalizer, the proposed non-iterative design matches the target curve with high precision. In the three example cases presented in this paper, the proposed parallel equalizer clearly outperforms other non-iterative graphic equalizer designs, and its maximum global error is as low as 0.00-0.75 dB when compared to the target curve. While the proposed design has superior accuracy, the number of operations in the filter structure is increased only by 23% when compared to the second-order Regalia-Mitra structure. The parallel structure also enables the utilization of parallel computing hardware, which can nowadays easily outperform the traditional serial processing. The proposed graphic equalizer can be widely used in audio signal processing applications.
Jussi Rämö, Vesa Välimäki, Balázs Bank
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 An improved virtual analog model of the Moog ladder filter
abstract
The Moog ladder structure is a well known filter used in musical sound synthesizers and in music production. Previously several digital models have attempted to imitate its nonlinear and self-oscillating characteristics. In this paper we derive a novel circuit-based model for the Moog filter and discretize it using the bilinear transform. The proposed nonlinear digital filter compares favorably against Huovilainen's model, which is the best previous white-box model for the Moog filter. The harmonic distortion characteristics of the proposed model match closely with those of a SPICE simulation. Furthermore, the novel model realistically enters the self-oscillation mode and maintains it. The proposed model requires only 12 more basic operations per output sample than Huovilainen's model, but includes the same number of nonlinear functions, which dominate the computational load. The novel nonlinear digital filter is applicable in virtual analog music synthesis and in musical audio effects processing.
Stefano D'Angelo, Vesa Välimäki
ICASSP2
2013 An ideal integrator for higher-order integrated wavetable synthesis
abstract
Higher-order integrated wavetable synthesis (HOIWS) is an efficient technique to reduce aliasing in wavetable and sampling synthesis. A periodic audio signal is integrated repeatedly before it is stored in a wavetable. During playback, the pitch of the audio signal can be changed using interpolation techniques and the resulting signal is differentiated as many times as the wavetable has been integrated. Previous discrete-time integrators approximate ideal integration, which leads to magnitude and phase errors. This paper proposes an ideal integration method, which is applied in the frequency domain with the help of the FFT. Its remarkable advantage is that both the magnitude and the phase errors are completely avoided in the special case of periodic signals. The proposed ideal integrator shows a superior performance over previous digital integration methods. It improves the sound quality of the HOIWS algorithm and helps it to maintain the original waveform after interpolation and differentiation stages.
Andreas Franck, Vesa Välimäki
ICASSP2
2013 A directional diffuse reverberation model for excavated tunnels in rock
abstract
Acoustic impulse responses of an excavated tunnel were measured. Analysis of the impulse responses shows that they are very diffuse from the start. A reverberator suitable for reproducing this type of response is proposed. The input signal is first comb-filtered and then convolved with a sparse noise sequence of the same length as the filter's delay line. An IIR loop filter inside the comb filter determines the decay rate of the response and is derived from the Yule-Walker approximation of the measured frequency-dependent reverberation time. The particular sparse noise sequence proposed in this work combines three velvet noise sequences, two of which have time-varying weights. To simulate the directional soundfield in a tunnel, the use of multiple such reverberators, each associated with a virtual source distributed evenly around the listener, is suggested. The proposed tunnel acoustics simulation can be employed in gaming, in film sound, or in working machine simulators.
Sami Oksanen, Julian Parker, Archontis Politis, Vesa Välimäki
ICASSP4
2013 Perceptual headphone equalization for mitigation of ambient noise
abstract
An adaptive perceptual equalizer for headphones is introduced. It estimates the effect of auditory masking while considering the characteristics of the headphones, ambient noise, and music. The system utilizes a psychoacoustic masking model to estimate the level to which the music should be raised to have the same perceived tonal balance in noise as it has in a quiet environment. Prototype testing showed that the most important task is to make the music audible in each Bark band. The compensation of the partial masking further improves the perceived sound quality. The system uses a microphone of a headset to capture the ambient noise. The equalization is implemented using a high-order graphical equalizer that does not require subband decomposition of the music signal. The proposed equalizer also retains reasonable SPL levels: in an example case, the maximum gain in one Bark band was 11 dB while the overall SPL increase was only 2.5 dB.
Jussi Rämö, Vesa Välimäki, Miikka Tikander
ICASSP2
2013 Linear Dynamic Range Reduction of Musical Audio Using an Allpass Filter Chain
abstract
The reduction of signal dynamic range through limiting of peak amplitude is an important process in modern audio signal processing, mainly for loudness maximisation. Traditional processes are non-linear, and can produce significant distortion of the processed signal. In this paper we present a new linear technique that reduces the peak amplitude of transient signals using golden ratio allpass filters. The system is applied to test signals consisting of both isolated musical sounds and mixed musical audio. The average reduction of the peak amplitude of the musical passages considered is 2.5 dB. The system can be applied alongside non-linear methods, to reduce the distortion associated with a particular reduction in peak amplitude.
Julian Parker, Vesa Välimäki
IEEE Signal Process. Lett.2
2013 New Family of Wave-Digital Triode Models
abstract
A new family of wave-digital vacuum tube triode models is presented. These models are inspired by the triode model by Cardarilli , which provides realistic simulation of the triode's transconductance behavior, and hence high accuracy in saturation conditions. The triode is modeled as a single memoryless nonlinear three-port wave digital filter element in which the outgoing wave variables are computed by locally applying the monodimensional secant method to one or two port voltages, depending on whether the grid current effect is taken into account. The proposed algorithms were found to produce a richer static harmonic response, introducing comparable or less aliasing and requiring approximately 50% less CPU time than previous models. The proposed models are suitable for real-time virtual analog circuit simulation.
Stefano D'Angelo, Jyri Tapani Pakarinen, Vesa Välimäki
IEEE Trans. Speech Audio Process.3
2013 A Perceptual Study on Velvet Noise and Its Variants at Different Pulse Densities
abstract
This paper investigates sparse noise sequences, including the previously proposed velvet noise and its novel variants defined here. All sequences consist of sample values minus one, zero, and plus one only, and the location and the sign of each impulse is randomly chosen. Two of the proposed algorithms are direct variants of the original velvet noise requiring two random number sequences for determining the impulse locations and signs. In one of the proposed algorithms the impulse locations and signs are drawn from the same random number sequence, which is advantageous in terms of implementation. Moreover, two of the new sequences include known regions of zeros. The perceived smoothness of the proposed sequences was studied with a listening test in which test subjects compared the noise sequences against a reference signal that was a Gaussian white noise. The results show that the original velvet noise sounds smoother than the reference at 2000 impulses per second. At 4000 impulses per second, also three of the proposed algorithms are perceived smoother than the Gaussian noise sequence. These observations can be exploited in the synthesis of noisy sounds and in artificial reverberation.
Vesa Välimäki, Heidi-Maria Lehtonen, Marko Takanen
IEEE Trans. Speech Audio Process.1
2012 Wave-digital polarity and current inverters and their application to virtual analog audio processing
abstract
Wave digital filters (WDFs) allow for efficient real-time simulation of classic analog circuitry by DSP. This paper introduces two new nonenergic two-port WDF adaptors that allow mixing wave digital subnetworks adopting different polarity and sign conventions and extends the definitions of absorbed instantaneous and steady-state pseudopower to the case in which the active sign convention is used. This new knowledge is applied to a WDF triode tube amplifier model and it is shown to result in a more faithful reproduction of the simulated system than the previous model.
Stefano D'Angelo, Vesa Välimäki
ICASSP2
2012 Optimized Polynomial Spline Basis Function Design for Quasi-Bandlimited Classical Waveform Synthesis
abstract
Classical geometric waveforms used in virtual analog synthesis suffer from aliasing distortion when simple sampling is used. An efficient antialiasing technique is based on expressing the waveforms as a filtered sum of time-shifted approximately bandlimited polynomial-spline basis functions. It is shown that by optimizing the coefficients of the basis function so that the aliasing distortion is perceptually minimized, the alias-free bandwidth of classical waveforms can be expanded. With the best of the case examples given here, the generated impulse-train and sawtooth waveform are alias-free up to fundamental frequencies over 10 kHz when the sampling rate is 44.1 kHz.
Jussi Pekonen, Juhan Nam, Julius O. Smith III, Vesa Välimäki
IEEE Signal Process. Lett.4
2012 Fifty Years of Artificial Reverberation
abstract
The first artificial reverberation algorithms were proposed in the early 1960s, and new, improved algorithms are published regularly. These algorithms have been widely used in music production since the 1970s, and now find applications in new fields, such as game audio. This overview article provides a unified review of the various approaches to digital artificial reverberation. The three main categories have been delay networks, convolution-based algorithms, and physical room models. Delay-network and convolution techniques have been competing in popularity in the music technology field, and are often employed to produce a desired perceptual or artistic effect. In applications including virtual reality, predictive acoustic modeling, and computer-aided design of acoustic spaces, accuracy is desired, and physical models have been mainly used, although, due to their computational complexity, they are currently mainly used for simplified geometries or to generate reverberation impulse responses for use with a convolution method. With the increase of computing power, all these approaches will be available in real time. A recent trend in audio technology is the emulation of analog artificial reverberation units, such as spring reverberators, using signal processing algorithms. As a case study we present an improved parametric model for a spring reverberation unit.
Vesa Välimäki, Julian Parker, Lauri Savioja, Julius O. Smith III, Jonathan S. Abel
IEEE Trans. Speech Audio Process.1
2010 Robust, Efficient Design of Allpass Filters for Dispersive String Sound Synthesis
abstract
An efficient allpass filter design method is introduced to match the dispersion characteristics of vibrating stiff strings. The proposed method designs an allpass filter in cascaded biquad form directly from the target group delay, placing the poles at frequencies at which the group delay area function achieves odd integer multiples of ¿, and fixing the pole radii according to a smoothness parameter. The pole frequencies are seen to be roots of quartic polynomials, and an efficient approximation to the desired roots is provided. Design examples show the method to outperform a previous closed-form design. Furthermore, the proposed method can achieve an arbitrarily wide bandwidth of good approximation by increasing the filter order, as the method is numerically robust and yields stable allpass filters.
Jonathan S. Abel, Vesa Välimäki, Julius O. Smith III
IEEE Signal Process. Lett.2
2010 Analysis and Synthesis of Coupled Vibrating Strings Using a Hybrid Modal-Waveguide Synthesis Model
abstract
The linear coupling of two strings or of a single string vibrating in two orthogonal polarizations leads to two observable phenomena: two-stage decay and beating. In this paper, we present methods for accurately measuring and modeling the lower partials of a recorded guitar tone, where coupling effects are most audible. These estimated parameters are then used for accurate resynthesis in a hybrid modal/waveguide model. We make use of the fact that two-stage decay occurs in analyzed tones to allow direct measurement of sinusoidal decay rates. A traditional iterative optimization algorithm is explored and found to be most effective in the special case when only beating occurs. Sound examples are provided on the Web.
Nelson Lee 0001, Julius O. Smith III, Vesa Välimäki
IEEE Trans. Speech Audio Process.3
2010 Efficient Antialiasing Oscillator Algorithms Using Low-Order Fractional Delay Filters
abstract
One of the challenges in virtual analog synthesis is avoiding aliasing when generating classic waveforms such as sawtooth and square wave which have theoretically infinite bandwidth in their ideal forms. The human auditory system renders a certain amount of aliasing inaudible, which allows room for finding cost-effective algorithms. This paper suggests efficient algorithms to reduce the aliasing using low-order fractional delay filters in the framework of bandlimited impulse train (BLIT) synthesis. Examining Lagrange, B-spline interpolators and allpass fractional delay filters, optimized methods will be discussed for generating classic waveforms (sawtooth, square, and triangle). Techniques for generating more complicated harmonics such as pulse width modulation, hard-sync, and super-saw are also presented. The perceptual evaluation is performed by comparing the threshold of hearing and masking curve of oscillators with their aliasing levels. The result shows that the BLIT using the computationally efficient third-order B-spline generates waveforms that are perceptually free of aliasing within practically used fundamental frequencies.
Juhan Nam, Vesa Välimäki, Jonathan S. Abel, Julius O. Smith III
IEEE Trans. Speech Audio Process.2
2010 Introduction to the Special Issue on Virtual Analog Audio Effects and Musical Instruments
abstract
The 16 papers in this special issue focus on virtual audio effects and musical instruments.
Vesa Välimäki, Federico Fontana, Julius O. Smith III, Udo Zölzer
IEEE Trans. Speech Audio Process.1
2010 Alias-Suppressed Oscillators Based on Differentiated Polynomial Waveforms
abstract
An efficient approach to the generation of classical synthesizer waveforms with reduced aliasing is proposed. This paper introduces two new classes of polynomial waveforms that can be differentiated one or more times to obtain an improved version of the sampled sawtooth and triangular signals. The differentiated polynomial waveforms (DPW) extend the previous differentiated parabolic wave method to higher polynomial orders, providing improved alias-suppression. Suitable polynomials of order higher than two can be derived either by analytically integrating a previous lower order polynomial or by solving the polynomial coefficients directly from a set of equations based on constraints. We also show how rectangular waveforms can be easily produced by differentiating a triangular signal. Bandlimited impulse trains can be obtained by differentiating the sawtooth or the rectangular signal. An objective evaluation using masking and hearing threshold models shows that a fourth-order DPW method is perceptually alias-free over the whole register of the grand piano. The proposed methods are applicable in digital implementations of subtractive sound synthesis.
Vesa Välimäki, Juhan Nam, Julius O. Smith III, Jonathan S. Abel
IEEE Trans. Speech Audio Process.1
2009 Spectrally rich phase distortion sound synthesis using an allpass filter
abstract
This paper examines a recently introduced technique for sound synthesis that uses a coefficient modulated allpass filter to cause phase modifications to its input signal. The intention in this work is to outline some of the properties of the coefficient modulated allpass filter and then to establish a connection between this new method and the older technique of phase distortion. Results are presented to demonstrate how the allpass technique provides a spectrally richer output signal.
Joseph Timoney, Victor Lazzarini, Jussi Pekonen, Vesa Välimäki
ICASSP4
2008 A computationally efficient coefficient update technique for Lagrange fractional delay filters
abstract
A new algorithm for coefficient update of the Lagrange fractional delay FIR filter is proposed, which reduces the computational complexity dramatically. It is based on rearranging the polynomial terms of the Lagrange interpolation formula and computing the common product terms only once. Reordering the Lagrange interpolation formula yields two other methods for updating the coefficients, the direct and the division- based methods. The division-based method uses only one division per coefficient. The two latter methods reduce the computational load, although they are not as efficient as the new algorithm. Finally, the superiority of the direct form FIR implementation of the Lagrange fractional delay filter and the new coefficient update method over other existing methods is demonstrated in an audio signal processing application.
Azadeh Haghparast, Vesa Välimäki
ICASSP2
2008 Filter-based alias reduction for digital classical waveform synthesis
abstract
The classical waveforms used in the subtractive sound synthesis have rich spectral content, which causes their sampled digital implementations to suffer from aliasing distortion. Several antialiasing waveform synthesis algorithms have been suggested, and they either remove the aliasing completely or reduce it greatly. A new approach to alias reduction is proposed where the remaining aliased components are suppressed by applying digital highpass and/or comb filtering to the output of an antialiasing algorithm. Applicable filter designs for this novel postprocessing approach are discussed and evaluated with respect to the alias reduction performance using noise-to-mask ratio (NMR). The NMR can be reduced by 10 dB at high fundamental frequencies with a computationally efficient highpass filter. The NMR can be further reduced by using a combination of an IIR comb filter and a DC blocking filter, which provides the best alias reduction performance at high fundamental frequencies.
Jussi Pekonen, Vesa Välimäki
ICASSP2
2007 Fractional Delay Filter Design Based on Truncated Lagrange Interpolation
abstract
A new design method for fractional delay filters based on truncating the impulse response of the Lagrange interpolation filter is presented. The truncated Lagrange fractional delay filter introduces a wider approximation bandwidth than the Lagrange filter. However, because of truncation, a ripple caused by the Gibbs phenomenon appears in the filter's frequency response. Proper choices of filter order and prototype filter order allow adjusting the overshoot to a desired level and simultaneously reducing the overall frequency-response error. The design of the proposed filter is computationally efficient, because it is based on polynomial formulas, which have common terms for all coefficients.
Vesa Välimäki, Azadeh Haghparast
IEEE Signal Process. Lett.1
2007 Synthesis of Hand Clapping Sounds
abstract
We present two physics-based analysis, synthesis, and control systems for synthesizing hand clapping sounds. They both rely on the separation of the sound synthesis and event generation, and both are capable of producing individual hand-claps, or mimicking the asynchronous/synchronized applause of a group of clappers. The synthesis models consist of resonator filters, whose coefficients are derived from experimental measurements. The difference between these systems is mainly in the statistical event generation. While the first system allows an efficient parametric synthesis of large audiences, as well as flocking and synchronization by simple rules, the second one provides parametric extensions for synthesis of various clapping styles and enhanced control strategies. The synthesis and the control models of both systems are implemented as software running in real time at the audio sample rate, and they are available for download at at http://ccrma-www.stanford.edu/software/stk and http://www.acoustics.hut.fi/go/clapd
Leevi Peltola, Cumhur Erkut, Perry R. Cook, Vesa Välimäki
IEEE Trans. Speech Audio Process.4
2006 Simulation of Room Acoustics using 2-D Digital Waveguide Meshes
abstract
A novel method for simulation of acoustic spaces, such as concert halls or listening rooms, using several 2-D digital waveguide mesh simulations is discussed. The advantages of this approach include reduced computational load, reduced memory usage, and simplified model structure in comparison to a 3-D waveguide mesh simulation. In approximating the modal frequencies of rooms, all the most important lowest modes get modeled, but some higher modes are missing. The proposed method is useful for finding low-frequency modes and for detecting changes in modal distribution when a sound source is moved, for example. As an acoustic visualization tool, the method is superior over a 3-D simulation in that it can isolate a certain layer of the acoustic wave field and it automatically hides waves that propagate in other directions and thus confuse the visualization
Antti Kelloniemi, Vesa Välimäki, Lauri Savioja
ICASSP (5)2
2006 Parametric Excitation Model for Waveguide Piano Synthesis
abstract
In this paper, a method providing an excitation signal for the waveguide piano synthesis is presented. The waveguide synthesis string model needs an excitation signal, which stimulates the model to resonate at the partial frequencies. This signal simulates the force pulse, which occurs in the piano when the hammer hits the string. In the proposed method, the excitation signal is produced by using additive synthesis with matching partial amplitudes and frequencies, and by adding bandlimited white noise into the signal. The excitation model takes into account the velocity at which the piano key is pressed, using bandstop and lowpass filtering. The proposed method is suitable for real-time piano synthesis, as it is controllable and computationally efficient
Jukka Rauhala, Vesa Välimäki
ICASSP (5)2
2006 Tunable dispersion filter design for piano synthesis
abstract
The tunable dispersion filter is a new design approach presented in this letter to provide dispersion modeling for digital waveguide synthesis of musical instruments, which do not produce extremely inharmonic sounds, such as the piano. We propose to use a cascade of second-order allpass filters for modeling dispersion. The filter coefficients can be calculated by using simple formulae based on the Thiran allpass filter design, which is usually used for fractional delay approximation. Unlike the previous allpass filter approximations, this filter design is easily scalable to produce various inharmonicity values for a wide range of fundamental frequencies.
Jukka Rauhala, Vesa Välimäki
IEEE Signal Process. Lett.2
2005 Energy behavior in time-varying fractional delay filters for physical modeling synthesis of musical instruments
abstract
Time-varying fractional delays are applied, for example, in physics based modeling of musical instruments, particularly for string and wind instruments. While Lagrange interpolation and allpass filters are used routinely in such sound synthesis models, they are found somewhat problematic, for example, in plucked string simulation when the length of the string is varied due to glissando or vibrato. There can be problems with signal energy levels and aliasing. We study two variable delay filter designs that have a physically realistic energetic behavior and keep undesirable side effects, such as aliasing, in control. The first one is sliding termination point simulation with energy correction and the second one is based on controllable wave digital filter delay lines.
Jyri Tapani Pakarinen, Matti Karjalainen, Vesa Välimäki, Stefan Bilbao
ICASSP (3)3
2005 Acoustic guitar plucking point estimation in real time
abstract
The algorithm estimates the plucking point of guitar tones obtained with an undersaddle pickup. This problem is approached in the time domain by applying autocorrelation estimation. This work extends a recently developed algorithm and brings it to a practical and sufficiently robust level. Improvements have been made to the onset detection part of the algorithm. The algorithm also enables a new way to control, for example, audio effect parameters in real time by simply changing the plucking point. The paper discusses these issues and the real-time implementation in the Pd environment. In tests, the real-time implementation achieved a 96% hit rate while the estimation error remains smaller than one centimeter, except for a few outliers. Audio samples and a Pd implementation of the algorithm are available on-line at www.acoustics.hut.fi/demos/plucking-point/.
Henri Penttinen, Jaakko Siiskonen, Vesa Välimäki
ICASSP (3)3
2005 Spatial filter-based absorbing boundary for the 2-D digital waveguide mesh
abstract
The digital waveguide mesh is a method for simulating wave propagation, for example, in an acoustic system. Research on the boundary conditions has been going on for years, but adequate solutions for absorbing boundaries have not yet been presented for the digital waveguide mesh. In this work, a new method for constructing absorbing boundaries for a two-dimensional (2-D) rectangular mesh is introduced. With the use of the proposed numerically optimized spatial filtering with an interpolated mesh structure, the reflection was diminished to under -25 dB at incidence angles |/spl theta/|/spl les/79.26/spl deg/ on a frequency band limited only at the very lowest and highest ends.
Antti Kelloniemi, Lauri Savioja, Vesa Välimäki
IEEE Signal Process. Lett.3
2005 Discrete-time synthesis of the sawtooth waveform with reduced aliasing
abstract
An efficient signal processing algorithm for generating a sawtooth waveform is proposed. The algorithm improves the trivial waveform sampling method, which suffers from low sound quality due to aliasing. The basic version of the new algorithm differentiates a piecewise parabolic waveform. Another version of the algorithm oversamples and decimates the parabolic wave with a simple filter prior to differentiation. The two algorithm variants improve the signal-to-noise ratio (SNR) over the trivial method by 10 and 15 dB, respectively. A perceptually weighted SNR suggests a larger subjective improvement. The proposed methods are applicable in a digital signal processor (DSP) implementation of subtractive sound synthesis.
Vesa Välimäki
IEEE Signal Process. Lett.1
2004 Boundary conditions in a multi-dimensional digital waveguide mesh
abstract
The digital waveguide mesh is a modeling technique suitable for simulation of wave propagation in an acoustic system. Artificial boundary conditions are constructed for the digital waveguide mesh. Absorbing boundary conditions are evaluated and a new method for adjusting the reflection coefficient at values 0/spl les/r/spl les/1 is introduced. The frequency dependent error level of this method is minimized by the use of a second-order FIR filter.
Antti Kelloniemi, Damian T. Murphy, Lauri Savioja, Vesa Välimäki
ICASSP (4)4
2003 Robust loss filter design for digital waveguide synthesis of string tones
abstract
A robust loss filter design method is presented for digital waveguide string models, which can be used with high filter orders. The method aims at minimizing the decay time error in partials of the synthetic tone. This is achieved by a new weighting function based on the first-order Taylor series approximation of the decay time errors. Smoothing of decay time data and requiring the design to be minimum-phase are also proposed to facilitate the stability of the design. The new method is applicable to analysis-based sound synthesis of piano and guitar tones, for example.
Balázs Bank, Vesa Välimäki
IEEE Signal Process. Lett.2
2003 Interpolated rectangular 3-D digital waveguide mesh algorithms with frequency warping
abstract
Various interpolated three-dimensional (3-D) digital waveguide mesh algorithms are elaborated. We introduce an optimized technique that improves a formerly proposed trilinearly interpolated 3-D mesh and renders the mesh more homogeneous in different directions. Furthermore, various sparse versions of the interpolated mesh algorithm are investigated, which reduce the computational complexity at the expense of accuracy. Frequency-warping techniques are used to shift the frequencies of the output signal of the mesh in order to cancel the effect of dispersion error. The extensions improve the accuracy of 3-D digital waveguide mesh simulations enough so that in the future it can be used for acoustical simulations needed in the design of listening rooms, for example.
Lauri Savioja, Vesa Välimäki
IEEE Trans. Speech Audio Process.2
2001 Interpolated 3-D digital waveguide mesh with frequency warping
abstract
An interpolated 3-D digital waveguide mesh algorithm is elaborated. We introduce an optimized technique that improves a formerly proposed interpolated 3-D mesh and renders the 3-D mesh more homogeneous in different directions. Frequency-warping techniques are used to shift the frequencies of the output signal of the mesh in order to cancel the effect of dispersion error. The extensions improve the accuracy of 3-D digital waveguide mesh simulations enough so that in the future it can be used for acoustical simulations needed in the design of listening rooms, for example.
Lauri Savioja, Vesa Välimäki
ICASSP2
2001 Multiwarping for enhancing the frequency accuracy of digital waveguide mesh simulations
abstract
A multiwarping technique is introduced. The method is based on cascading frequency warping procedures and sample rate conversions. As a practical example, we show that with this new approach, the maximal frequency error of digital waveguide mesh simulations can be reduced by about 50%. This technique gives more degrees of freedom to attain a desired amount of nonlinear frequency shift still applying first-order allpass filters for processing finite-length digital impulse responses.
Lauri Savioja, Vesa Välimäki
IEEE Signal Process. Lett.2
2000 Model-based sound synthesis of tanbur, a Turkish long-necked lute
abstract
Physics-based simulation and sound synthesis of the tanbur, a traditional Turkish long-necked lute is tackled by two computational models based on digital waveguides. A linear generic dual-polarization simulation including sympathetic coupling is calibrated according to recorded sound examples. The other approach utilizes a specific model that incorporates a nonlinear tension modulation mechanism that is pronounced in the tanbur. The nonlinear implementation can accurately reproduce features such as variation of the fundamental frequency, nonlinear coupling of the harmonic partials, and tension modulation driving force coupling to the body. The synthesis results of both models are justified by audio demonstrations.
Cumhur Erkut, Vesa Välimäki
ICASSP2
2000 Acoustic sound from the electric guitar using DSP techniques
abstract
The electric guitar has been developed to withstand electric amplification and utilize (mostly analog) signal processing in order to create a multitude of timbres and sound types. Sometimes it would be desirable to play the same electric guitar, yet with a sound that resembles a good acoustic guitar or some other member of the plucked string instrument family. In this study we have investigated DSP techniques that can be used to shape the magnetic pickup output of the electric guitar to simulate acoustic instruments. This includes linear filtering for body simulation, time-varying modulation to generate beating of the harmonic components, and techniques to simulate the general temporal envelope of the plucked notes.
Matti Karjalainen, Henri Penttinen, Vesa Välimäki
ICASSP3
2000 Principles of fractional delay filters
abstract
In numerous applications, such as communications, audio and music technology, speech coding and synthesis, antenna and transducer arrays, and time delay estimation, not only the sampling frequency but the actual sampling instants are of crucial importance. Digital fractional delay (FD) filters provide a useful building block that can be used for fine-tuning the sampling instants, i.e., implement the required bandlimited interpolation. In this paper an overview of design techniques and applications is given.
Vesa Välimäki, Timo I. Laakso
ICASSP1
2000 Polynomial filtering approach to reconstruction and noise reduction of nonuniformly sampled signals
Timo I. Laakso, Andrzej Tarczynski, N. Paul Murphy, Vesa Välimäki
Signal Process.4
2000 Reducing the dispersion error in the digital waveguide mesh using interpolation and frequency-warping techniques
abstract
The digital waveguide mesh is an extension of the one-dimensional (1-D) digital waveguide technique. The mesh can be used for simulation of two- and three-dimensional (3-D) wave propagation in musical instruments and acoustic spaces. The original rectangular digital waveguide mesh algorithm suffers from direction-dependent dispersion. Alternative geometries, such as the triangular mesh, have been proposed previously to improve the performance of the mesh. In this paper, we show that the dispersion problem may be reduced using various other techniques. These methods include multidimensional interpolation, optimization of the point-spreading function, and frequency warping. We compare the accuracy and computational complexity of these techniques in the two-dimensional (2-D) case and conduct numerical simulations of a membrane. A rectangular mesh using second-order Lagrange interpolation can be implemented without multiplications, but its accuracy is worse than that of other enhanced structures. The most accurate technique in terms of the relative frequency error is the warped triangular mesh whose maximum error is 0.6%. The warped rectangular mesh with optimized weighting coefficients is not as exact, but still offers a 1.2% accuracy.
Lauri Savioja, Vesa Välimäki
IEEE Trans. Speech Audio Process.2
2000 Modeling of tension modulation nonlinearity in plucked strings
abstract
A nonlinear discrete-time model that simulates a vibrating string exhibiting tension modulation nonlinearity is developed. The tension modulation phenomenon is caused by string elongation during transversal vibration. Fundamental frequency variation and coupling of harmonic modes are among the perceptually most important effects of this nonlinearity. The proposed model extends the linear bidirectional digital waveguide model of a string. It is also formulated as a computationally more efficient single-delay-loop structure. A method of reducing the computational load of the string elongation approximation is described, and a technique of obtaining the tension modulation parameter from recorded plucked string instrument tones is presented. The performance of the model is demonstrated with analysis/synthesis experiments and with examples of synthetic tones.
Tero Tolonen, Vesa Välimäki, Matti Karjalainen
IEEE Trans. Speech Audio Process.2
1999 Reduction of the dispersion error in the interpolated digital waveguide mesh using frequency warping
abstract
The digital waveguide mesh is an extension of the one-dimensional digital waveguide technique. The mesh is used for simulation of two- and three-dimensional wave propagation in musical instruments and acoustic spaces. The rectangular digital waveguide mesh algorithm suffers from direction-dependent dispersion. By using the interpolated mesh, nearly uniform wave propagation characteristics are obtained in all directions. In this paper we show how the dispersion error of the interpolated mesh can be reduced by frequency warping. By using this technique the bandwidth where the frequency accuracy is within 1% tolerance is more than doubled.
Lauri Savioja, Vesa Välimäki
ICASSP2
1999 Plucked-string synthesis algorithms with tension modulation nonlinearity
abstract
Digital waveguide modeling of a nonlinear vibrating string is investigated when the nonlinearity is essentially caused by tension modulation. We derive synthesis models where the nonlinearity is implemented with a time-varying fractional delay filter. Also, conversion from a dual-delay-line physical model into a single-delay-loop model is explained. Realistic synthetic tones with nonlinear effects are obtained by introducing minor amendments to a linear string synthesis algorithm. It is shown how synthetic plucked-string tones are modified as a consequence of tension modulation.
Vesa Välimäki, Tero Tolonen, Matti Karjalainen
ICASSP1
1999 Reduction of the dispersion error in the triangular digital waveguide mesh using frequency warping
abstract
The digital waveguide mesh has been successfully used for simulation of two-dimensional (2-D) and three-dimensional (3-D) wave propagation in musical instruments and acoustic spaces. Nevertheless, digital waveguide mesh algorithms suffer from dispersion which increases with frequency. In this letter, we show how the dispersion error of the triangular digital waveguide mesh can be reduced by frequency warping. By using this technique, the worst-case dispersion error of 0.6% is obtained, whereas in the original triangular mesh it is about 6.5%.
Lauri Savioja, Vesa Välimäki
IEEE Signal Process. Lett.2
1998 Energy-based effective length of the impulse response of a recursive filter
abstract
A measure for the effective length of the impulse response of a stable recursive digital filter based on accumulated energy is proposed. A general definition and a simple algorithm for its evaluation are introduced, and closed-form expressions are derived for first-order IIR filters. The effect of zeros on the effective length is analyzed. An upper bound for the effective length of higher-order filters is derived using results for low-order filters. The new measure finds applications in several fields of digital signal processing, including estimation of the extent of attack transients for filters with dynamically varying inputs, elimination of transients in variable recursive filters, and design and implementation of linear-phase IIR systems.
Timo I. Laakso, Vesa Välimäki
ICASSP2
1998 Suppression of transients in time-varying recursive filters for audio signals
abstract
A new method for suppressing transients in time-varying recursive filters is proposed. The technique is based on modifying the state variables when the filter coefficients are changed so that the filter enters a new state smoothly without transient attacks, as originally proposed by Zetterberg and Zhang (1988). We modify the Zetterberg-Zhang algorithm to render it feasible for efficient implementation. We explain how to determine an optimal transient suppresser to cancel the transients down to a desired level with the minimum complexity of implementation. The application of the method to time-varying all-pole and direct form IIR filter structures is studied. The algorithm may be generalized for any recursive filter structure. The transient suppression technique finds applications in audio signal processing where the characteristics of a recursive filter needs to be changed in real time, such as in music synthesis, auralization, and equalization.
Vesa Välimäki, Timo I. Laakso
ICASSP1
1997 Improved discrete-time modeling of multi-dimensional wave propagation using the interpolated digital waveguide mesh
abstract
The digital waveguide mesh is an extension of the one-dimensional digital waveguide technique. Waveguide meshes are used for simulation of two- and three-dimensional wave propagation in musical instruments and acoustic spaces. The original waveguide mesh algorithm suffers from direction-dependent dispersion. In this paper we show that this problem may be reduced by using an interpolated rectilinear mesh. In the analysis part we show the analytical solution for the wave propagation speed and numerical simulations of the magnitude response and phase speed in both the original and the interpolated two-dimensional waveguide mesh algorithms. We demonstrate by simulation that the wave propagation characteristics of the proposed interpolated waveguide mesh are independent of direction and thus the remaining errors caused by dispersion may be corrected with a postprocessor.
Lauri Savioja, Vesa Välimäki
ICASSP2
1997 FIR filtering of nonuniformly sampled signals
abstract
Filtering signals sampled on a grid which is nonuniformly distributed in the time domain is not a simple task since the filter's coefficients have to be time varying. They must be updated at each sampling instant. The filtering becomes even more complicated when it has to be optimal (or at least suboptimal) in the sense of a certain design criterion. In this paper we present an effective algorithm for FIR filtering aiming at minimisation of the energy of the filtering error signal. The approach provides a solution which resembles the weighted least squares design method for FIR filters of uniformly sampled signals.
Andrzej Tarczynski, Vesa Välimäki, Gerald D. Cain
ICASSP2
1997 Low-order modeling of head-related transfer functions using balanced model truncation
abstract
We propose a novel technique for the design of low-order infinite impulse response (IIR) filter models of head-related transfer functions (HRTFs) that uses balanced model truncation. We design tenth-order IIR filters, suitable for efficient real-time implementation, from preprocessed HRTF impulse responses that are of significantly superior quality to current IIR models derived with the Prony and Yule-Walker methods.
Jonathan P. Mackenzie, Jyri Huopaniemi, Vesa Välimäki, Izzet Kale
IEEE Signal Process. Lett.3
1995 Tunable downsampling using fractional delay filters with applications to digital TV transmission
abstract
An efficient technique for sampling rate conversion for arbitrary (incommensurate) ratios is proposed. The technique is based on fractional delay filters that are efficient to implement and that can be controlled with a small number of arithmetic operations per output sample. The authors consider an application in digital television (DTV) transmission where, according to present standard proposals, conversions between several incommensurate sampling rates must be possible. Rather than trying to design separate rate fixed filters for each possible conversion, the authors outline a system which may be tuned for any possible downsampling ratio. A sampling rate conversion system based on the straightforward and simple Lagrange interpolation technique is illustrated with a level and highly efficient implementation structure. Various error sources involved are analyzed and a mean-square-error (MSE) type cost function is defined to aid in the system design.
Timo I. Laakso, Vesa Välimäki, Jukka Henriksson
ICASSP2
1995 Implementation of fractional delay waveguide models using allpass filters
abstract
This paper discusses a discrete-time modeling technique where the length of time delays can be arbitrarily adjusted. The new system is called a fractional delay waveguide model (FDWM). Formerly, FDWMs have only been implemented with FIR-type fractional delay filters. We show how an FDWM can be implemented using allpass filters. We use low-order allpass filters that are maximally-flat approximations of the ideal delay. The advantages of the allpass approach are computational efficiency and reduced approximation error. The proposed structure can be applied to discrete-time modeling of acoustic tubes, such as the human vocal tract or resonators of musical instruments.
Vesa Välimäki, Matti Karjalainen
ICASSP1
1995 A New Filter Implementation Strategy for Lagrange Interpolation
abstract
Fractional delays (FD) are used in DSP applications where accurate time delays are required. This paper introduces a new implementation technique for a maximally flat FIR FD filter, or Lagrange interpolator. The number of multiplications is reduced when compared with the direct-form FIR implementation plus coefficient update using Nth-order polynomials. The new structure is well-suited to implementation of a time-varying fractional delay.
Vesa Välimäki
ISCAS1
1994 Articulatory speech synthesis based on fractional delay waveguide filters
abstract
An extension to the traditional Kelly-Lochbaum (1962) vocal tract model is introduced. In the new model not only the diameter but also the length of each tube section can be continuously adjusted. This is achieved by using fractional delay filter techniques such as interpolation and deinterpolation. The filter structure consisting of bidirectional delay lines (digital waveguides) and interpolated ports that connect two or more waveguide sections together is called a fractional delay waveguide filter (FDWF). The interpolated version of the two-port scattering junction is presented and a technique for analyzing the degradation due to approximation errors in interpolation and deinterpolation is described. It is shown that when an FDWF structure with Lagrange interpolation is used a vocal tract model needs to be implemented using oversampling. For example, a sampling rate of 22 kHz is adequate for producing high-quality synthetic sounds at a 5 kHz bandwidth.>
Vesa Välimäki, Matti Karjalainen, Timo Kuisma
ICASSP (1)1
1994 Improving the kelly-lochbaum vocal tract model using conical tube sections and fractional delay filtering techniques
abstract
An articulatory model of speech production is usually constructed by approximating the profile of the vocal tract using cylindrical tube sections. This is implemented by a digital ladder filter that is called the Kelly--Lochbaum model. In this paper we propose an extended approach, where the tube sections approximating the profile of the tract are conical instead of cylindrical. Furthermore, the length of each tube section in our model can be accurately controlled using a novel fractional delay filtering scheme. These refinements result in an accurate and intuitively controllable vocal tract model that is well suited for articulatory speech synthesis. I. INTRODUCTION An articulatory model for the human vocal tract is traditionally constructed by approximating the profile of the vocal tract using cylindrical tube sections. The resulting tube system is then modeled by a digital ladder filter. This approach was first used in [1] and is called the Kelly-- Lochbaum (KL) model. There are tw...
Vesa Välimäki, Matti Karjalainen
ICSLP1
1993 Fractional delay digital filters
Vesa Välimäki, Matti Karjalainen, Timo I. Laakso
ISCAS1
1992 A real-time DSP implementation of a flute model
abstract
A computationally efficient semiphysical time-domain model of the flute is presented. The traditional method of modeling the one-dimensional wave propagation as a transmission-line, where all the losses as well as reflection and dispersion effects have been lumped to linear filters at the ends of the line, is used. The model is excited by white noise, and the nonlinear interaction between the excitation and the tube resonator has been modeled by a static memoryless nonlinearity. The model is simple enough to be computationally efficient, but has retained many important features of a real instrument. The real-time implementation has been done on a Texas Instruments TMS320C30 floating-point signal processor using a sampling rate of 44.1 kHz.>
Vesa Välimäki, Matti Karjalainen, Zoltán Jánosy, Unto K. Laine
ICASSP1