Julian Parker

dblp:25/8308 · also Julian D. Parker · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0003-7261-4271ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Stable Audio Open
abstract
Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model’s performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.
Zach Evans, Julian Parker, CJ Carr, Zack Zukowski, Josiah Taylor, Jordi Pons
ICASSP2
2025 Scaling Transformers for Low-Bitrate High-Quality Speech Coding
abstract
The tokenization of audio with neural audio codec models is a vital part of modern AI pipelines for the generation or understanding of speech, alone or in a multimodal context. Traditionally such tokenization models have concentrated on low parameter-count architectures using only components with strong inductive biases. In this work we show that by applying a transformer architecture with large parameter count to this problem, and applying a flexible Finite Scalar Quantization (FSQ) based bottleneck, it is possible to reach state-of-the-art speech quality at extremely low bit-rates of $400$ or $700$ bits-per-second. The trained models strongly out-perform existing baselines in both objective and subjective tests.
Julian Parker, Anton Smirnov, Jordi Pons, CJ Carr, Zack Zukowski, Zach Evans
ICLR1
2024 STEMGEN: A Music Generation Model That Listens
abstract
End-to-end generation of musical audio using deep learning techniques has seen an explosion of activity recently. However, most models concentrate on generating fully mixed music in response to abstract conditioning information. In this work, we present an alternative paradigm for producing music generation models that can listen and respond to musical context. We describe how such a model can be constructed using a non-autoregressive, transformer-based model architecture and present a number of novel architectural and sampling improvements. We train the described architecture on both an open-source and a proprietary dataset. We evaluate the produced models using standard quality metrics and a new approach based on music information retrieval descriptors. The resulting model reaches the audio quality of state-of-the-art text-conditioned models, as well as exhibiting strong musical coherence with its context.
Julian Parker, Janne Spijkervet, Katerina Kosta, Furkan Yesiler, Boris Kuznetsov, Ju-Chiang Wang, Matt Avent, Jitong Chen
ICASSP1
2021 Differentiable White-Box Virtual Analog Modeling
abstract
Component-wise circuit modeling, also known as “white-box” modeling, is a well established and much discussed technique in virtual analog modeling. This approach is generally limited in accuracy by lack of access to the exact component values present in a real example of the circuit. In this paper we show how this problem can be addressed by implementing the white-box model in a differentiable form, and allowing approximate component values to be learned from raw input-output audio measured from a real device.
Fabian Esqueda, Boris Kuznetsov, Julian Parker
DAFx3
2021 Combining Zeroth and First-Order Analysis with Lagrange Polynomials to Reduce Artefacts in Live Concatenative Granulation
abstract
This paper presents a technique addressing signal discontinuity and concatenation artefacts in real-time granular processing with rectangular windowing. By combining zero-crossing synchronicity, first-order derivative analysis, and Lagrange polynomials, we can generate streams of uncorrelated and non-overlapping sonic fragments with minimal low-order derivatives discontinuities. The resulting open-source algorithm, implemented in the Faust language, provides a versatile real-time software for dynamical looping, wavetable oscillation, and granulation with reduced artefacts due to rectangular windowing and no artefacts from overlap-add-to-one techniques commonly deployed in granular processing.
Dario Sanfilippo, Julian Parker
DAFx2
2021 On the Equivalence of Integrator- and Differentiator-Based Continuous- and Discrete-Time Systems
abstract
The article performs a generic comparison of integrator- and differentiator based continuous-time systems as well as their discrete-time models, aiming to answer the reoccurring question in the music DSP community of whether there are any benefits in using differentiators instead of conventionally employed integrators. It is found that both kinds of models are practically equivalent, but there are certain reservations about differentiator based models.
Vadim Zavalishin, Julian Parker
DAFx2
2017 Antiderivative Antialiasing for Memoryless Nonlinearities
abstract
Aliasing is a commonly encountered problem in audio signal processing, particularly when memoryless nonlinearities are simulated in discrete time. A conventional remedy is to operate at an oversampled rate. A new aliasing reduction method is proposed here for discrete-time memoryless nonlinearities, which is suitable for operation at reduced oversampling rates. The method employs higher order antiderivatives of the nonlinear function used. The first-order form of the new method is equivalent to a technique proposed recently by Parker et al. Higher order extensions offer considerable improvement over the first antiderivative method, in terms of the signal-to-noise ratio. The proposed methods can be implemented with fewer operations than oversampling and are applicable to discrete-time modeling of a wide range of nonlinear analog systems.
Stefan Bilbao, Fabian Esqueda, Julian Parker, Vesa Välimäki
IEEE Signal Process. Lett.3
2015 Designing microgames for assessment: a case study in rapid prototype iteration
abstract
In contrast to the trend toward large scale, immersive games that aspire toward the polish and experience of conventional commercial games, the authors offer a design case study for the potential of microgames and assessment. Microgames are designed to be small, pointed experiences more analogous to a single question than an entire exam. Instead of offering diverse mechanics, microgames are small punctuated play experiences. Microgames are rapidly developed games, targeting a relatively narrow set of skills.
Lindsay D. Grace, G. Tanner Jackson, Christopher Totten, Julian Parker, Joyce Rice
Advances in Computer Entertainment4
2013 A directional diffuse reverberation model for excavated tunnels in rock
abstract
Acoustic impulse responses of an excavated tunnel were measured. Analysis of the impulse responses shows that they are very diffuse from the start. A reverberator suitable for reproducing this type of response is proposed. The input signal is first comb-filtered and then convolved with a sparse noise sequence of the same length as the filter's delay line. An IIR loop filter inside the comb filter determines the decay rate of the response and is derived from the Yule-Walker approximation of the measured frequency-dependent reverberation time. The particular sparse noise sequence proposed in this work combines three velvet noise sequences, two of which have time-varying weights. To simulate the directional soundfield in a tunnel, the use of multiple such reverberators, each associated with a virtual source distributed evenly around the listener, is suggested. The proposed tunnel acoustics simulation can be employed in gaming, in film sound, or in working machine simulators.
Sami Oksanen, Julian Parker, Archontis Politis, Vesa Välimäki
ICASSP2
2013 Linear Dynamic Range Reduction of Musical Audio Using an Allpass Filter Chain
abstract
The reduction of signal dynamic range through limiting of peak amplitude is an important process in modern audio signal processing, mainly for loudness maximisation. Traditional processes are non-linear, and can produce significant distortion of the processed signal. In this paper we present a new linear technique that reduces the peak amplitude of transient signals using golden ratio allpass filters. The system is applied to test signals consisting of both isolated musical sounds and mixed musical audio. The average reduction of the peak amplitude of the musical passages considered is 2.5 dB. The system can be applied alongside non-linear methods, to reduce the distortion associated with a particular reduction in peak amplitude.
Julian Parker, Vesa Välimäki
IEEE Signal Process. Lett.1
2012 Fifty Years of Artificial Reverberation
abstract
The first artificial reverberation algorithms were proposed in the early 1960s, and new, improved algorithms are published regularly. These algorithms have been widely used in music production since the 1970s, and now find applications in new fields, such as game audio. This overview article provides a unified review of the various approaches to digital artificial reverberation. The three main categories have been delay networks, convolution-based algorithms, and physical room models. Delay-network and convolution techniques have been competing in popularity in the music technology field, and are often employed to produce a desired perceptual or artistic effect. In applications including virtual reality, predictive acoustic modeling, and computer-aided design of acoustic spaces, accuracy is desired, and physical models have been mainly used, although, due to their computational complexity, they are currently mainly used for simplified geometries or to generate reverberation impulse responses for use with a convolution method. With the increase of computing power, all these approaches will be available in real time. A recent trend in audio technology is the emulation of analog artificial reverberation units, such as spring reverberators, using signal processing algorithms. As a case study we present an improved parametric model for a spring reverberation unit.
Vesa Välimäki, Julian Parker, Lauri Savioja, Julius O. Smith III, Jonathan S. Abel
IEEE Trans. Speech Audio Process.2
2010 A Virtual Model of Spring Reverberation
abstract
The digital emulation of analog audio effects and synthesis components, through the simulation of lumped circuit components has seen a large amount of activity in recent years; electromechanical effects have seen rather less, primarily because they employ distributed mechanical components, which are not easily dealt with in a rigorous manner using typical audio processing constructs such as delay lines and digital filters. Spring reverberation is an example of such a system-a spring exhibits complex, highly dispersive behavior, including coupling between different types of wave propagation (longitudinal and transverse). Standard numerical techniques, such as finite difference schemes are a good match to such a problem, but require specialized design and analysis techniques in the context of audio processing. A model of helical spring vibration is introduced, along with a family of finite difference schemes suitable for time domain simulation. Various topics are covered, including numerical stability conditions, tuning of the scheme to the response of the model system, numerical boundary conditions and connection to an excitation and readout, implementation details, as well as computational requirements. Simulation results are presented, and full energy-based stability analysis appears in an Appendix.
Stefan Bilbao, Julian Parker
IEEE Trans. Speech Audio Process.2