Michael M. Goodwin

dblp:82/7145 · also Michael Mark Goodwin, Mike Goodwin · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 NOLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping
abstract
Speech codec enhancement methods are designed to remove distortions added by speech codecs. While classical methods are very low in complexity and add zero delay, their effectiveness is rather limited. Compared to that, DNN-based methods deliver higher quality but they are typically high in complexity and/or require delay. The recently proposed Linear Adaptive Coding Enhancer (LACE) addresses this problem by combining DNNs with classical long-term/short-term post-filtering resulting in a causal low-complexity model. A short-coming of the LACE model is, however, that quality quickly saturates when the model size is scaled up. To mitigate this problem, we propose a novel adatpive temporal shaping module that adds high temporal resolution to the LACE model resulting in the Non-Linear Adaptive Coding Enhancer (NoLACE). We adapt NoLACE to enhance the Opus codec and show that NoLACE significantly outperforms both the Opus baseline and an enlarged LACE model at 6, 9 and 12 kb/s. We also show that LACE and NoLACE are well-behaved when used with an ASR system.
Jan Büthe, Ahmed Mustafa, Jean-Marc Valin, Karim Helwani, Michael M. Goodwin
ICASSP5
2024 Noise-Robust DSP-Assisted Neural Pitch Estimation With Very Low Complexity
abstract
Pitch estimation is an essential step of many speech processing algorithms, including speech coding, synthesis, and enhancement. Recently, pitch estimators based on deep neural networks (DNNs) have been outperforming well-established DSP-based techniques. Unfortunately, these new estimators can be impractical to deploy in real-time systems, both because of their relatively high complexity, and the fact that some require significant lookahead. We show that a hybrid estimator using a small deep neural network (DNN) with traditional DSP-based features can match or exceed the performance of pure DNN-based models, with a complexity and algorithmic delay comparable to traditional DSP-based algorithms. We further demonstrate that this hybrid approach can provide benefits for a neural vocoding task.
Krishna Subramani, Jean-Marc Valin, Jan Büthe, Paris Smaragdis, Michael M. Goodwin
ICASSP5
2024 Real-Time Stereo Speech Enhancement with Spatial-Cue Preservation Based on Dual-Path Structure
abstract
We introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source has unique spatial information, our method utilizes a dual-path structure, ensuring the spatial cues remain unaffected during enhancement by applying source-specific common-band gain. This method also seamlessly integrates pretrained monaural speech enhancement, eliminating the need for retraining on stereo inputs. Source separation from stereo mixtures is achieved via spatial beamforming, with the steering vector for each source being adaptively updated using post-enhancement output signal. This ensures accurate tracking of the spatial information. The final stereo output is derived by merging the spatial images of the enhanced sources, with its efficacy not heavily reliant on the separation performance of the beamforming. The algorithm runs in real-time on 10-ms frames with a 40 ms of look-ahead. Evaluations reveal its effectiveness in enhancing speech and preserving spatial cues in both fully and sparsely overlapped mixtures.
Masahito Togami, Jean-Marc Valin, Karim Helwani, Ritwik Giri, Umut Isik, Michael M. Goodwin
ICASSP6
2023 Generative Modeling Based Manifold Learning for Adaptive Filtering Guidance
abstract
In most practical adaptive filtering problems, estimated filters are not arbitrary, but instead lie on a manifold that encapsulates characteristics of the problem at hand. Consequently, it is desirable to steer adaptation towards filters that lie on that manifold. In this paper, we propose a novel approach to learn the manifold of a set of impulse responses and subsequently employ that learned manifold in an adaptation algorithm for system identification. The presented approach is a practical adaptive filtering recipe for enforcing a data-driven search domain constraint, instead of using conventional constrained optimization methods.
Karim Helwani, Paris Smaragdis, Michael M. Goodwin
ICASSP3
2023 Framewise Wavegan: High Speed Adversarial Vocoder In Time Domain With Very Low Computational Complexity
abstract
GAN vocoders are currently one of the state-of-the-art methods for building high-quality neural waveform generative models. However, most of their architectures require dozens of billion floating-point operations per second (GFLOPS) to generate speech waveforms in samplewise manner. This makes GAN vocoders still challenging to run on normal CPUs without accelerators or parallel computers. In this work, we propose a new architecture for GAN vocoders that mainly depends on recurrent and fully-connected networks to directly generate the time domain signal in framewise manner. This results in considerable reduction of the computational cost and enables very fast generation on both GPUs and low-complexity CPUs. Experimental results show that our Framewise WaveGAN vocoder achieves significantly higher quality than auto-regressive maximum-likelihood vocoders such as LPCNet at a very low complexity of 1.2GFLOPS. This makes GAN vocoders more practical on edge and low-power devices.
Ahmed Mustafa, Jean-Marc Valin, Jan Büthe, Paris Smaragdis, Michael M. Goodwin
ICASSP5
2023 A Framework for Unified Real-Time Personalized and Non-Personalized Speech Enhancement
abstract
In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning input that specifies the type of enhancement output. To improve the quality of the enhanced output and mitigate oversuppression, we experiment with re-weighting frames by the presence or absence of speech activity and applying augmentations to speaker embeddings. By training under a multi-task learning setting, we empirically show that the proposed unified model obtains promising results on both personalized and non-personalized speech enhancement benchmarks and reaches similar performance to models that are trained specialized for either task. The strong performance of the proposed method demonstrates that the unified model is a more economical alternative compared to keeping separate task-specific models during inference.
Zhepei Wang, Ritwik Giri, Devansh Shah, Jean-Marc Valin, Michael M. Goodwin, Paris Smaragdis
ICASSP5
2022 Improved Singing Voice Separation with Chromagram-Based Pitch-Aware Remixing
abstract
Singing voice separation aims to separate music into vocals and accompaniment components. One of the major constraints for the task is the limited amount of training data with separated vocals. Data augmentation techniques such as random source mixing have been shown to make better use of existing data and mildly improve model performance. We propose a novel data augmentation technique, chromagram-based pitch-aware remixing, where music segments with high pitch alignment are mixed. By performing controlled experiments in both supervised and semi-supervised settings, we demonstrate that training models with pitch-aware remixing significantly improves the test signal-to-distortion ratio (SDR).
Siyuan Yuan, Zhepei Wang, Umut Isik, Ritwik Giri, Jean-Marc Valin, Michael M. Goodwin, Arvindh Krishnaswamy
ICASSP6
2022 Clock Skew Robust Acoustic Echo Cancellation
Karim Helwani, Erfan Soltanmohammadi, Michael M. Goodwin, Arvindh Krishnaswamy
INTERSPEECH3
2008 Geometric signal decompositions for spatial audio enhancement
abstract
Decomposition of audio signals into primary and ambient components is useful for realizing spatial enhancements such as upmix and stereo widening. In this paper, we present several methods for primary-ambient decomposition of two-channel audio signals based on signal-space geometry. We discuss the performance of the various methods with respect to target orthogonality conditions on the estimated primary and ambient components, which cannot all be satisfied due to the need to constrain the model components to the signal subspace in light of limitations on implementation complexity.
Michael M. Goodwin
ICASSP1
2007 Primary-Ambient Signal Decomposition and Vector-Based Localization for Spatial Audio Coding and Enhancement
abstract
Spatial audio coding and enhancement address the growing commercial need to store and distribute multichannel audio and to render content optimally on arbitrary reproduction systems. In this paper, we discuss a spatial analysis-synthesis scheme which applies principal component analysis to an STFT-domain representation of the original audio to separate it into primary and ambient components, which are then respectively analyzed for cues that describe the spatial percept of the audio scene on a per-tile basis; these cues are used by the synthesis to render the audio appropriately on the available playback system. The proposed framework can be tailored for robust spatial audio coding, or it can be applied directly to enhancement scenarios where there are no rate constraints on the intermediate spatial data and audio representation.
Michael M. Goodwin, Jean-Marc Jot
ICASSP (1)1
2006 Efficient Designs for Broadband Linear Arrays
abstract
Linear electroacoustic arrays are useful for various applications in audio acquisition and reproduction. For robust broadband performance, a filter network is typically incorporated to achieve roughly frequency-invariant beamforming. In this paper, we are concerned with the design of cost-effective broadband arrays of limited size for consumer audio applications. Filter-based broadband beamformers are suboptimal for such applications due to high cost and the large number of elements generally needed to achieve invariance. Here, three alternative design methods are presented which lead to effective far-field broadband performance of small arrays at low cost: allpass weighting, optimal nonuniform spacing, and optimal delay dispersion. Performance improvements are demonstrated for each method and various design tradeoffs are considered
Michael M. Goodwin
ICASSP (5)1
2005 Predicting and preventing unmasking incurred in coded audio post-processing
abstract
In modern audio compression algorithms, the masking properties of the auditory system are exploited to improve the coding gain, namely, quantization noise is introduced in the signal in time-frequency regions where it will be masked. However, since signal modifications change the characteristics of the masking regions, degradations may result if a decoded audio signal is modified. In this paper, we explain and demonstrate how modifying an audio signal can result in the unmasking of signal components that were imperceptible in the unmodified signal. We consider both pitch-shifting and linear filtering modifications; synthetic and natural audio examples are provided to verify the unmasking phenomenon. We discuss how modification of decoded audio may lead to unmasking of quantization noise, describe conditions for which such unmasking may occur, and propose a method for adjusting the masking threshold employed in the audio coder to make the decoded signal robust to quantization noise unmasking for a given set of signal modifications.
Michael M. Goodwin, A. J. Hipple, B. Link
IEEE Trans. Speech Audio Process.1
2004 A dynamic programming approach to audio segmentation and speech/music discrimination
abstract
We consider the problem of segmenting an audio signal into characteristic regions based on feature-set similarities. In the proposed approach, a feature-space representation of the signal is generated; sequences of these feature-space samples are then aggregated into clusters corresponding to distinct signal regions. The algorithm consists of using linear discriminant analysis (LDA) to condition the feature space and dynamic programming (DP) to identify data clusters. We consider the design of the dynamic program cost functions; we are able to derive effective cost functions without relying on significant prior information about the structure of the expected data clusters. We demonstrate the application of the LDA-DP segmentation algorithm to speech/music discrimination. Experimental results are given and discussed.
Michael M. Goodwin, Jean Laroche
ICASSP (4)1
1998 Multiresolution sinusoidal modeling using adaptive segmentation
abstract
The sinusoidal model has proven useful for representation and modification of speech and audio. One drawback, however, is that a sinusoidal signal model is typically derived using a fixed frame size, which corresponds to a rigid signal segmentation. For nonstationary signals, the resolution limitations that result from this rigidity lead to reconstruction artifacts. It is shown in this paper that such artifacts can be significantly reduced by using a signal-adaptive segmentation derived by a dynamic program. An atomic interpretation of the sinusoidal model is given; this perspective suggests that algorithms for adaptive segmentation can be viewed as methods for adapting the time scales of the constituent atoms so as to improve the model by employing appropriate time-frequency tradeoffs.
Michael M. Goodwin
ICASSP1
1997 Matching pursuit with damped sinusoids
abstract
The matching pursuit algorithm derives an expansion of a signal in terms of the elements of a large dictionary of time-frequency atoms. This paper considers the use of matching pursuit for computing signal expansions in terms of damped sinusoids. First, expansion based on complex damped sinusoids is explored; it is shown that the expansion can be efficiently derived using the FFT and simple recursive filterbanks. Then, the approach is extended to provide decompositions in terms of real damped sinusoids. This extension relies on generalizing the matching pursuit algorithm to derive expansions with respect to dictionary subspaces; of specific interest is the subspace spanned by a complex atom and its conjugate. Developing this particular case leads to a framework for deriving real-valued expansions of real signals using complex atoms. Applications of the damped sinusoidal decomposition include system identification, spectral estimation, and signal modeling for coding and analysis-modification-synthesis.
Michael M. Goodwin
ICASSP1
1997 Optimal time segmentation for signal modeling and compression
abstract
The idea of optimal joint time segmentation and resource allocation for signal modeling is explored with respect to arbitrary segmentations and arbitrary representation schemes. When the chosen signal modeling techniques can be quantified in terms of a cost function which is additive over distinct segments, a dynamic programming approach guarantees the global optimality of the scheme while keeping the computational requirements of the algorithm sufficiently low. Two immediate applications of the algorithm to LPC speech coding and to sinusoidal modeling of musical signals are presented.
Paolo Prandoni, Michael M. Goodwin, Martin Vetterli
ICASSP2
1996 Residual modeling in music analysis-synthesis
abstract
In analysis-synthesis of musical sounds based on a sinusoidal model, the difference between the original signal and the synthesized signal, termed the residual, is typically a broadband noise process. It contains such musical phenomena as flute breath noise or violin bow noise. Synthesis without such "noise" tends to sound artificial; it is desirable to improve the synthesis realism by modeling the residual in such a way that it can be reinjected in the synthesized signal. This paper deals with a model of noise perception based on the equivalent rectangular bands (ERBs) of the auditory system. Since a broadband noise is perceptually well-represented by the time-varying energy in each of these frequency bands, the residual is parametrized in terms of these energies in the proposed model. An application of the model to music synthesis based on the inverse fast Fourier transform (FFT) is described in detail.
Michael M. Goodwin
ICASSP1
1993 Beam dithering: Acoustic feedback control using a modulated-directivity loudspeaker array
Gary W. Elko, Michael M. Goodwin
ICASSP (1)2
1993 Constant beamwidth beamforming
Michael M. Goodwin, Gary W. Elko
ICASSP (1)1