VLDB 2026 Research / reviewers in the wild / expert
Joshua D. Reiss
dblp:05/342 · also Josh Reiss
· DBLP profile ↗
27ranked-venue papers
1as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Syncfusion: Multimodal Onset-Synchronized Video-to-Audio Foley SynthesisabstractSound design involves creatively selecting, recording, and editing sound effects for various media like cinema, video games, and virtual/augmented reality. One of the most time-consuming steps when designing sound is synchronizing audio with video. In some cases, environmental recordings from video shoots are available, which can aid in the process. However, in video games and animations, no reference audio exists, requiring manual annotation of event timings from the video. We propose a system to extract repetitive actions onsets from a video, which are then used - in conjunction with audio or textual embeddings - to condition a diffusion model trained to generate a new synchronized sound effects audio track. In this way, we leave complete creative control to the sound designer while removing the burden of synchronization with video. Furthermore, editing the onset track or changing the conditioning embedding requires much less effort than editing the audio track itself, simplifying the sonification process. We provide sound examples, source code, and pretrained models to faciliate reproducibility1. Marco Comunità, Riccardo F. Gramaccioni, Emilian Postolache, Emanuele Rodolà, Danilo Comminiello, Joshua D. Reiss |
ICASSP | 6 |
| 2024 | Enhanced Speech Emotion Recognition Incorporating Speaker-Sensitive Interactions in ConversationsabstractAccurately detecting emotions in conversation is a necessary yet challenging task due to the complexity of emotions and dynamics in dialogues. The emotional state of a speaker can be influenced by many different factors, such as interlocutor stimulus, dialogue scene, and topic. In this work, we propose a conversational speech emotion recognition method to deal with capturing attentive contextual dependency and speaker-sensitive interactions. First, we use a pretrained WavLM model to extract frame-based audio representation in individual utterances. Second, an attentive bi-directional gated recurrent unit (GRU) models contextual-sensitive information and explores listener dependency and speaker influence jointly in a simple, fast, parameter-efficient way. The experiments conducted on the standard conversational dataset MELD demonstrate the effectiveness of the proposed method when compared against state-of the-art methods. Jiachen Luo, Huy Phan, Lin Wang 0009, Joshua D. Reiss |
ICME | 4 |
| 2023 | Modelling Black-Box Audio Effects with Time-Varying Feature ModulationabstractDeep learning approaches for black-box modelling of audio effects have shown promise, however, the majority of existing work focuses on nonlinear effects with behaviour on relatively short time-scales, such as guitar amplifiers and distortion. While recurrent and convolutional architectures can theoretically be extended to capture behaviour at longer time scales, we show that simply scaling the width, depth, or dilation factor of existing architectures does not result in satisfactory performance when modelling audio effects such as fuzz and dynamic range compression. To address this, we propose the integration of time-varying feature-wise linear modulation into existing temporal convolutional backbones, an approach that enables learnable adaptation of the intermediate activations. We demonstrate that our approach more accurately captures long-range dependencies for a range of fuzz and compressor implementations across both time and frequency domain metrics. We provide sound examples, source code, and pretrained models to faciliate reproducibility1. Marco Comunità, Christian J. Steinmetz, Huy Phan, Joshua D. Reiss |
ICASSP | 4 |
| 2023 | Cross-Modal Fusion Techniques for Utterance-Level Emotion Recognition from Text and SpeechabstractMultimodal emotion recognition (MER) is a fundamental complex research problem due to the uncertainty of human emotional expression and the heterogeneity gap between different modalities. Audio and text modalities are particularly important for a human participant in understanding emotions. Although many successful attempts have been designed multimodal representations for MER, there still exist multiple challenges to be addressed: 1) bridging the heterogeneity gap between multimodal features and model inter- and intramodal interactions of multiple modalities; 2) effectively and efficiently modeling the contextual dynamics in the conversation sequence. In this paper, we propose Cross-Modal RoBERTa (CM-RoBERTa) model for emotion detection from spoken audio and corresponding transcripts. As the core unit of the CM-RoBERTa, parallel self- and cross- attention is designed to dynamically capture inter- and intra-modal interactions of audio and text. Specially, the mid-level fusion and residual module are employed to model longterm contextual dependencies and learn modality-specific patterns. We evaluate the approach on the MELD dataset and the experimental results show the proposed approach achieves the state-of-art performance on the dataset. Jiachen Luo, Huy Phan, Joshua D. Reiss |
ICASSP | 3 |
| 2023 | Fine-tuned RoBERTa Model with a CNN-LSTM Network for Conversational Emotion RecognitionabstractTextual emotion recognition in conversations has gained increasing attention in recent years for the growing amount of applications it can serve, e.g., human-robot interactions, recommended systems. However, most existing approaches are either based on BERT-based model which fail to exploit crucial information about the long-text context, or resort to complex entanglement of neural network architectures resulting in less stable training procedures and slower inference time. To bridge this gap, we first propose a fast, compact and parameter-efficient framework based on fine-tuned pre-trained RoBERTa model with a CNN-LSTM network for textual emotion recognition in conversations. First, we fine-tune the pre-tranined RoBERTa model to effectively learn long-term emotion-relevant context information. Second, convolutional neural network coupled with the bidirectional long short-term memory and joint reinforced blocks are utilized to recognize emotion in conversations. Extensive experiments are conducted on benchmark emotion MELD dataset, and the results show that our model outperforms a wide range of strong baselines and achieves competitive results with the state-of-art approaches. Jiachen Luo, Huy Phan, Joshua D. Reiss |
INTERSPEECH | 3 |
| 2022 | Direct Design of Biquad Filter Cascades with Deep Learning by Sampling Random PolynomialsabstractDesigning infinite impulse response filters to match an arbitrary magnitude response requires specialized techniques. Methods like modified Yule-Walker are relatively efficient, but may not be sufficiently accurate in matching high order responses. On the other hand, iterative optimization techniques often enable superior performance, but come at the cost of longer run-times and are sensitive to initial conditions, requiring manual tuning. In this work, we address some of these limitations by learning a direct mapping from the target magnitude response to the filter coefficient space with a neural network trained on millions of random filters. We demonstrate our approach enables both fast and accurate estimation of filter coefficients given a desired response. We investigate training with different families of random filters, and find training with a variety of filter families enables better generalization when estimating real-world filters, using head-related transfer functions and guitar cabinets as case studies. We compare our method against existing methods including modified Yule-Walker and gradient descent and show our approach is, on average, both faster and more accurate. Joseph T. Colonel, Christian J. Steinmetz, Marcus Michelen, Joshua D. Reiss |
ICASSP | 4 |
| 2020 | Modeling Plate and Spring Reverberation Using A DSP-Informed Deep Neural NetworkabstractPlate and spring reverberators are electromechanical systems first used and researched as means to substitute real room reverberation. Currently, they are often used in music production for aesthetic reasons due to their particular sonic characteristics. The modeling of these audio processors and their perceptual qualities is difficult since they use mechanical elements together with analog electronics resulting in an extremely complex response. Based on digital reverberators that use sparse FIR filters, we propose a signal processing-informed deep learning architecture for the modeling of artificial reverberators. We explore the capabilities of deep neural networks to learn such highly nonlinear electromechanical responses and we perform modeling of plate and spring reverberators. In order to measure the performance of the model, we conduct a perceptual evaluation experiment and we also analyze how the given task is accomplished and what the model is actually learning. Marco A. Martínez Ramírez, Emmanouil Benetos, Joshua D. Reiss |
ICASSP | 3 |
| 2019 | Modeling Nonlinear Audio Effects with End-to-end Deep Neural NetworksabstractIn the context of music production, distortion effects are mainly used for aesthetic reasons and are usually applied to electric musical instruments. Most existing methods for nonlinear modeling are often either simplified or optimized to a very specific circuit. In this work, we investigate deep learning architectures for audio processing and we aim to find a general purpose end-to-end deep neural network to perform modeling of nonlinear audio effects. We show the network modeling various nonlinearities and we discuss the generalization capabilities among different instruments. Marco A. Martínez Ramírez, Joshua D. Reiss |
ICASSP | 2 |
| 2019 | Unifying Probabilistic Models for Time-frequency AnalysisabstractIn audio signal processing, probabilistic time-frequency models have many benefits over their non-probabilistic counterparts. They adapt to the incoming signal, quantify uncertainty, and measure correlation between the signal’s amplitude and phase information, making time domain resynthesis straightforward. However, these models are still not widely used since they come at a high computational cost, and because they are formulated in such a way that it can be difficult to interpret all the modelling assumptions. By showing their equivalence to Spectral Mixture Gaussian processes, we illuminate the underlying model assumptions and provide a general framework for constructing more complex models that better approximate real-world signals. Our interpretation makes it intuitive to inspect, compare, and alter the models since all prior knowledge is encoded in the Gaussian process kernel functions. We utilise a state space representation to perform efficient inference via Kalman smoothing, and we demonstrate how our interpretation allows for efficient parameter learning in the frequency domain. William J. Wilkinson, Michael Riis Andersen, Joshua D. Reiss, Dan Stowell, Arno Solin |
ICASSP | 3 |
| 2019 | End-to-End Probabilistic Inference for Nonstationary Audio AnalysisabstractA typical audio signal processing pipeline includes multiple disjoint analysis stages, including calculation of a time-frequency representation followed by spectrogram-based feature analysis. We show how time-frequency analysis and nonnegative matrix factorisation can be jointly formulated as a spectral mixture Gaussian process model with nonstationary priors over the amplitude variance parameters. Further, we formulate this nonlinear model’s state space representation, making it amenable to infinite-horizon Gaussian process regression with approximate inference via expectation propagation, which scales linearly in the number of time steps and quadratically in the state dimensionality. By doing so, we are able to process audio signals with hundreds of thousands of data points. We demonstrate, on various tasks with empirical data, how this inference scheme outperforms more standard techniques that rely on extended Kalman filtering. William J. Wilkinson, Michael Riis Andersen, Joshua D. Reiss, Dan Stowell, Arno Solin |
ICML | 3 |
| 2018 | Perceptual Evaluation of Synthesized Sound EffectsabstractSound synthesis is the process of generating artificial sounds through some form of simulation or modelling. This article aims to identify which sound synthesis methods achieve the goal of producing a believable audio sample that may replace a recorded sound sample. A perceptual evaluation experiment of five different sound synthesis techniques was undertaken. Additive synthesis, statistical modelling synthesis with two different feature sets, physically inspired synthesis, concatenative synthesis, and sinusoidal modelling synthesis were all compared. Evaluation using eight different sound class stimuli and 66 different samples was undertaken. The additive synthesizer is the only synthesis method not considered significantly different from the reference sample across all sounds classes. The results demonstrate that sound synthesis can be considered as realistic as a recorded sample and makes recommendations for use of synthesis methods, given different sound class contexts. David Moffat, Joshua D. Reiss |
ACM Trans. Appl. Percept. | 2 |
| 2017 | Performance Evaluation of a New Flexible Time Division Multiplexing Protocol on Mixed Traffic TypesabstractThe broadcasting industry has recently begun to adopt statistical multiplexing based network platform in their workflow to support professional live audio/video (AV) transmission instead of the Time Division Multiplexing (TDM) based system. These audio-over-packet switched systems require a carefully designed and managed network to ensure key quality measures of the real-time (RT) media, such as low jitter and low latency. Often the best effort traffic or different types of media are still physically or logically segregated from these dedicated systems, or require large redundant links. The proposed Flexilink architecture is an alternative that combines both circuit switched and best effort features. However, there is no research evaluation that shows the actual performance of this proposed architecture. In this paper, we give a simulation based study and critical evaluation of the performance of the Flexilink network. The simulation results show that Flexilink has a better and more stable RT performance when compared with both Ethernet and priority queueing networks, especially when given a burst of traffic and/or multiple RT traffic sources. In addition, Unlike other networking protocols, jitter in Flexilink is below the audible threshold. Yangyang Song, Yonghao Wang, Peter Bull, Joshua D. Reiss |
AINA | 4 |
| 2017 | Stem audio mixing as a content-basec transformation of audio featuresabstractMultitrack audio mixing is an essential part of music production and one of the first steps consist on processing individual stems from raw recordings. In this paper, we investigate this stage as a content-based transformation. We explore which audio features are relevant to interpret this specific process and which set of features gets modified by the mixing of stems in the most consistent way. We show that the number of features can be reduced with a procedure based on the permutation importance method of random forest classifiers. Thus, the selected audio features are used to train various classification models and we analyse which set of features lead to a better classification accuracy. We conclude that the underlying characteristics of manipulating raw recordings into individual stems can be described by this selected set of features. Marco A. Martínez Ramírez, Joshua D. Reiss |
MMSP | 2 |
| 2016 | Semantic Description of Timbral Transformations in Music ProductionabstractIn music production, descriptive terminology is used to define perceived sound transformations. By understanding the underlying statistical features associated with these descriptions, we can aid the retrieval of contextually relevant processing parameters using natural language, and create intelligent systems capable of assisting in audio engineering. In this study, we present an analysis of a dataset containing descriptive terms gathered using a series of processing modules, embedded within a Digital Audio Workstation. By applying hierarchical clustering to the audio feature space, we show that similarity in term representations exists within and between transformation classes. Furthermore, the organisation of terms in low-dimensional timbre space can be explained using perceptual concepts such as size and dissonance. We conclude by performing Latent Semantic Indexing to show that similar groupings exist based on term frequency. Ryan Stables, Brecht De Man, Sean Enderby, Joshua D. Reiss, György Fazekas, Thomas Wilmering |
ACM Multimedia | 4 |
| 2016 | An Iterative Approach to Source Counting and Localization Using Two Distant MicrophonesabstractWe propose a time difference of arrival (TDOA) estimation framework based on time-frequency inter-channel phase difference (IPD) to count and localize multiple acoustic sources in a reverberant environment using two distant microphones. The time-frequency (T-F) processing enables exploitation of the nonstationarity and sparsity of audio signals, increasing robustness to multiple sources and ambient noise. For inter-channel phase difference estimation, we use a cost function, which is equivalent to the generalized cross correlation with phase transform (GCC) algorithm and which is robust to spatial aliasing caused by large inter-microphone distances. To estimate the number of sources, we further propose an iterative contribution removal (ICR) algorithm to count and locate the sources using the peaks of the GCC function. In each iteration, we first use IPD to calculate the GCC function, whose highest peak is detected as the location of a sound source; then we detect the T-F bins that are associated with this source and remove them from the IPD set. The proposed ICR algorithm successfully solves the GCC peak ambiguities between multiple sources and multiple reverberant paths. Lin Wang 0009, Tsz-Kin Hon, Joshua D. Reiss, Andrea Cavallaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Over-Determined Source Separation and Localization Using Distributed MicrophonesabstractWe propose an overdetermined source separation and localization method for a set of M microphones distributed around an unknown number, N <; M, of sources. We reformulate the overdetermined acoustic mixing procedure with a new determined mixing model and apply a determined M × M independent component analysis ('CA) in each frequency bin directly. The reformulated 'CA operates without knowing N and also leads to better separation in reverberant scenarios. To solve the challenging permutation ambiguity problem, we first employ a time activity-based clustering approach to cluster the separated frequency components into M channels. We then propose a remixing procedure to detect and merge channels from the same source. The detection is done by analyzing time and frequency activities, spectral likeliness, and spatial location. To estimate the spatial location, we propose a time-frequency masking-based steered response power algorithm. Simulated and real-data experiments in a very challenging reverberant scenario confirm the effectiveness of the proposed method in obtaining the number of sources, the separated signals, and the location and spatial likelihood of each source. Lin Wang 0009, Joshua D. Reiss, Andrea Cavallaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Audio Fingerprinting for Multi-Device Self-LocalizationabstractWe investigate the self-localization problem of an ad-hoc network of randomly distributed and independent devices in an open-space environment with low reverberation but heavy noise (e.g. smartphones recording videos of an outdoor event). Assuming a sufficient number of sound sources, we estimate the distance between a pair of devices from the extreme (minimum and maximum) time difference of arrivals (TDOAs) from the sources to the pair of devices without knowing the time offset. The obtained inter-device distances are then exploited to derive the geometrical configuration of the network. In particular, we propose a robust audio fingerprinting algorithm for noisy recordings and perform landmark matching to construct a histogram of the TDOAs of multiple sources. The extreme TDOAs can be estimated from this histogram. By using audio fingerprinting features, the proposed algorithm works robustly in very noisy environments. Experiments with free-field simulation and open-space recordings prove the effectiveness of the proposed algorithm. Tsz-Kin Hon, Lin Wang 0009, Joshua D. Reiss, Andrea Cavallaro |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | A Cross-Adaptive Dynamic Spectral Panning Technique
Pedro D. Pestana, Joshua D. Reiss |
DAFx | 2 |
| 2014 | MIXPLORATION: rethinking the audio mixer interfaceabstractA typical audio mixer interface consists of faders and knobs that control the amplitude level as well as processing (e.g. equalization, compression and reverberation) parameters of individual tracks. This interface, while widely used and effective for optimizing a mix, may not be the best interface to facilitate exploration of different mixing options. In this work, we rethink the mixer interface, describing an alternative interface for exploring the space of possible mixes of four audio tracks. In a user study with 24 participants, we compared the effectiveness of this interface to the traditional paradigm for exploring alternative mixes. In the study, users responded that the proposed alternative interface facilitated exploration and that they considered the process of rating mixes to be beneficial. Mark Cartwright, Bryan Pardo, Joshua D. Reiss |
IUI | 3 |
| 2013 | Model-Based Inversion of Dynamic Range CompressionabstractIn this work it is shown how a dynamic nonlinear time-variant operator, such as a dynamic range compressor, can be inverted using an explicit signal model. By knowing the model parameters that were used for compression one is able to recover the original uncompressed signal from a “broadcast” signal with high numerical accuracy and very low computational complexity. A compressor-decompressor scheme is worked out and described in detail. The approach is evaluated on real-world audio material with great success. Stanislaw Gorlow, Joshua D. Reiss |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | A Wiener Filter Approach to Microphone Leakage Reduction in Close-Microphone ApplicationsabstractMicrophone leakage is one of the most prevalent problems in audio applications involving multiple instruments and multiple microphones. Currently, sound engineers have limited solutions available to them. In this paper, the applicability of two widely used signal enhancement methods to this problem is discussed, namely blind source separation and noise suppression. By extending previous work, it is shown that the noise suppression framework is a valid choice and can effectively address the problem of microphone leakage. Here, an extended form of the single channel Wiener filter is used which takes into account the individual audio sources to derive a multichannel noise term. A novel power spectral density (PSD) estimation method is also proposed based on the identification of dominant frequency bins by examining the microphone and output signal PSDs. The performance of the method is examined for simulated environments with various source-microphone setups and it is shown that the proposed approach efficiently suppresses leakage. Elias K. Kokkinis, Joshua D. Reiss, John Mourjopoulos |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Design of Audio Parametric Equalizer Filters Directly in the Digital DomainabstractMost design procedures for a digital parametric equalizer begin with analog design techniques, followed by applying the bilinear transform to an analog prototype. As an alternative, an approximation to the parametric equalizer is sometimes designed using pole-zero placement techniques. In this paper, we present an exact derivation of the parametric equalizer without resorting to an analog prototype. We show that there are many solutions to the parametric equalizer design constraints as usually stated, but only one of which consistently yields stable, minimum phase behavior with the upper and lower cutoff frequencies positioned around the center frequency. The conditions for complex conjugate poles and zeros are found and the resultant pole zero placements are examined. Joshua D. Reiss |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Evaluation of Distance Based Amplitude panning for spatial audio
Dimitar Kostadinov, Joshua D. Reiss, Valeri M. Mladenov |
ICASSP | 2 |
| 2010 | A Real-Time Framework for Video Time and Pitch Scale ModificationabstractA framework is presented which addresses the issues related to the real-time implementation of synchronized video and audio time-scale and pitch-scale modification algorithms. It allows for seamless real-time transition between continually varying, independent time-scale and pitch-scale parameters arising as a result of manual or automatic intervention. We illuminate the problems which arise in a real-time context as well as provide novel solutions to prevent artifacts, minimize latency, and improve synchronization. The time and pitch scaling approach is based on a modified phase vocoder with optional phase locking and an integrated transient detector which enables high-quality transient preservation in real-time. A novel method for audio/visual synchronization was implemented in order to ensure no perceptible latency between audio and video while real-time time scaling and pitch shifting is applied. Evaluation results are reported which demonstrate both high audio quality and minimal synchronization error. Ivan Damnjanovic 0001, Dan Barry, David Dorran, Joshua D. Reiss |
IEEE Trans. Multim. | 4 |
| 2008 | Enabling access to sound archives through integration, enrichment and retrievalabstractMany digital sound archives still suffer from tremendous problems concerning access. Materials are often in different formats, with related media in separate collections, and with non-standard, specialist, incomplete or even erroneous metadata. Thus, the end user is unable to discover the full value of the archived material. EASAIER addresses these issues with the development of an innovative remote access system which extends beyond standard content management and retrieval systems. The EASAIER system has been designed with sound archives, libraries, museums, broadcast archives, and music schools in mind. However, the tools may be used by anyone interested in accessing archived material; amateur or professional, regardless of the material involved. Furthermore, it enriches the access experience enabling the user to experiment with the materials in exciting new ways. The system features; enhanced cross media retrieval functionality, multi-media synchronisation, audio and video processing, analysis and visualisation tools, all combined within in a single user configurable interface. Ivan Damnjanovic 0001, Joshua D. Reiss, Dan Barry |
ICME | 2 |
| 2006 | Noise Analysis of Modulated Quantizer based on Oversampled SignalsabstractIn this paper, a noise analysis of a modulated quantizer is performed. If input signals are oversampled, then the quantization error could be reduced by modulating both the input and the output of the quantizer. The working principle is based on the fact that convolutions of bandpass signals would spread wider in the frequency spectrum than that of lowpass signals. Hence, by filtering the high frequency components, the signal-to-noise ratio (SNR) could be increased. Numerical simulation results show that the modulated quantization scheme could achieve an average of 13.0960dB to 21.4700dB improvements on SNR over the conventional scheme, depends on the types of bandlimited input signals. Charlotte Yuk-Fan Ho, Bingo Wing-Kuen Ling, Joshua D. Reiss |
ICASSP (3) | 3 |
| 2005 | Nonlinear behaviors of bandpass sigma delta modulators with stable system matricesabstractIt has been established that a class of bandpass sigma delta modulators (SDMs) may exhibit state space dynamics which are represented by elliptical or fractal patterns confined within trapezoidal regions when the system matrices are marginally stable. It is found that fractal patterns may also be exhibited in the phase plane when the system matrices are strictly stable. This occurs when the sets of initial conditions corresponding to convergent or limit cycle behavior do not cover the whole phase plane. Based on the derived analytical results, some interesting results are found. If the bandpass SDM exhibits periodic output, then the period of the symbolic sequence must equal the limiting period of the state space variables. Second, if the state vector converges to some fixed points on the phase portrait, these fixed points do not depend directly on the initial conditions. Bingo Wing-Kuen Ling, Charlotte Yuk-Fan Ho, Joshua D. Reiss, Xinghuo Yu 0001 |
ICASSP (4) | 3 |