Haici Yang

dblp:259/3072 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0001-5503-1259ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Keeping the Balance: Anomaly Score Calculation for Domain Generalization
abstract
Emitted sounds may drastically change when using different microphones, when properties of the sound sources change, or when recording in different acoustic environments. Ideally, anomalous sound detection (ASD) systems should be able to generalize well to unseen target domains by only providing a few target domain samples to define how normal data samples sound like, without needing to re-train or modify the system. In contrast with the source domain, for which many normal training samples are available, accurately estimating the underlying distribution of normal data after a domain shift based on very few samples is challenging. This usually leads to a mismatch between the corresponding anomaly scores of source and target domains and significantly reduces performance. In this work, we propose a framework for re-scaling anomaly scores based on the ratio between the cosine distance of a test sample to a normal reference sample and the distances to this sample’s next-closest neighbors in the reference set. In experimental evaluations, it is shown that the re-scaled anomaly scores reduce the domain mismatch for multiple domains. As a result, we obtain new state-of-the-art performances on the DCASE2020 and DCASE2023 ASD datasets.
Kevin Wilkinghoff, Haici Yang, Janek Ebbers, François G. Germain, Gordon Wichern, Jonathan Le Roux
ICASSP2
2025 Investigating continuous autoregressive generative speech enhancement
Haici Yang, Gordon Wichern, Ryo Aihara, Yoshiki Masuyama, Sameer Khurana, François G. Germain, Jonathan Le Roux
INTERSPEECH1
2024 Personalized Neural Speech Codec
abstract
In this paper, we propose a personalized neural speech codec, envisioning that personalization can reduce the model complexity or improve perceptual speech quality. Despite the common usage of speech codecs where only a single talker is involved on each side of the communication, personalizing a codec for the specific user has rarely been explored in the literature. First, we assume speakers can be grouped into smaller subsets based on their perceptual similarity. Then, we also postulate that a group-specific codec can focus on the group’s speech characteristics to improve its perceptual quality and computational efficiency. To this end, we first develop a Siamese network that learns the speaker embeddings from the LibriSpeech dataset, which are then grouped into underlying speaker clusters. Finally, we retrain the LPCNet-based speech codec baselines on each of the speaker clusters. Subjective listening tests show that the proposed personalization scheme introduces model compression while maintaining speech quality. In other words, with the same model complexity, personalized codecs produce better speech quality.
Inseon Jang, Haici Yang, Wootaek Lim, Seungkwon Beack
ICASSP2
2024 Generative De-Quantization for Neural Speech Codec Via Latent Diffusion
abstract
End-to-end speech coding models achieve high coding gains by learning compact yet expressive features and a powerful decoder in a single network. A challenging problem as such results in unwelcome complexity increase and inferior speech quality. In this paper, we propose to separate the representation learning and information reconstruction tasks. We leverage an end-to-end codec for learning low-dimensional discrete tokens. Instead of using its decoder, we employ a latent diffusion model to de-quantize coded features into a high-dimensional continuous space, relieving the decoder’s burden of de-quantizing and upsampling. To mitigate the issue of over-smooth generation, we introduce midway-infilling with less noise reduction and stronger conditioning. We investigate the hyperparameters for midway-infilling and latent diffusion space with different dimensions in ablation studies. Subjective listening tests show that our model outperforms the state-of-the-art at two low bitrates, 1.5 and 3 kbps. We open-source the project for reproducibility1.
Haici Yang, Inseon Jang
ICASSP1
2024 Genhancer: High-Fidelity Speech Enhancement via Generative Modeling on Discrete Codec Tokens
Haici Yang, Jiaqi Su, Zeyu Jin
INTERSPEECH1
2023 Neural Feature Predictor and Discriminative Residual Coding for Low-Bitrate Speech Coding
abstract
Low and ultra-low-bitrate neural speech codecs achieved unprecedented coding gain by generating speech signals from compact features. This paper introduces additional coding efficiency in speech coding by reducing the temporal redundancy existing in the frame-level feature sequence via a feature predictor. This predictor produces low-entropy residual representations, and we discriminatively code them based on their contribution to the signal reconstruction. Combining feature prediction and discriminative coding optimizes bitrate efficiency by assigning more bits to hard-to-predict events. We demonstrate the advantage of the proposed methods using the LPCNet as a neural vocoder, resulting in a scalable, lightweight, low-latency, and low-bitrate neural speech coding system. While our approach guarantees strict causality in the frame-level prediction, the subjective tests and feature space analysis show that our model achieves superior coding efficiency compared to the loosely-causal LPCNet and Lyra V2 in the very low bitrates.
Haici Yang, Wootaek Lim, Minje Kim 0001
ICASSP1
2022 Don't Separate, Learn To Remix: End-To-End Neural Remixing With Joint Optimization
abstract
The task of manipulating the level and/or effects of individual instruments to recompose a mixture of recordings, or remixing, is common across a variety of applications such as music production, audio-visual post-production, podcasts, and more. This process, however, traditionally requires access to individual source recordings, restricting the creative process. To work around this, source separation algorithms can separate a mixture into its respective components. Then, a user can adjust their levels and mix them back together. This two-step approach, however, still suffers from audible artifacts and motivates further work. In this work, we re-purpose Conv-TasNet, a well-known source separation model, into two neural remixing architectures that learn to remix directly rather than just to separate sources. We use an explicit loss term that directly measures remix quality and jointly optimize it with a separation loss. We evaluate our methods using the Slakh and MUSDB18 datasets and report remixing performance as well as the impact on source separation as a byproduct. Our results suggest that learning-to-remix significantly outperforms a strong separation baseline and is particularly useful for small volume changes.
Haici Yang, Shivani Firodiya, Nicholas J. Bryan
ICASSP1
2022 Upmixing Via Style Transfer: A Variational Autoencoder for Disentangling Spatial Images And Musical Content
abstract
In the stereo-to-multichannel upmixing problem for music, one of the main tasks is to set the directionality of the instrument sources in the multichannel rendering results. In this paper, we propose a modified variational autoencoder model that learns a latent space to describe the spatial images in multichannel music. We seek to disentangle the spatial images and music content, so the learned latent variables are invariant to the music. At test time, we use the latent variables to control the panning of sources. We propose two upmixing use cases: transferring the spatial images from one song to another and blind panning based on the generative model. We report objective and subjective evaluation results to empirically show that our model captures spatial images separately from music content and achieves transfer-based interactive panning.
Haici Yang, Sanna Wager, Spencer Russell, Mike Luo, Wontak Kim
ICASSP1
2021 Source-Aware Neural Speech Coding for Noisy Speech Compression
abstract
This paper introduces a novel neural network-based speech coding system that can process noisy speech effectively. The proposed source-aware neural audio coding (SANAC) system harmonizes a deep autoencoder-based source separation model and a neural coding system, so that it can explicitly perform source separation and coding in the latent space. An added benefit of this system is that the codec can allocate a different amount of bits to the underlying sources, so that the more important source sounds better in the decoded signal. We target a new use case where the user on the receiver side cares about the quality of the non-speech components in the speech communication, while the speech source still carries the most important information. Both objective and subjective evaluation tests show that SANAC can recover the original noisy speech better than the baseline neural audio coding system, which is with no source-aware coding mechanism, and two conventional codecs.
Haici Yang, Kai Zhen, Seungkwon Beack
ICASSP1
2020 Boosted Locality Sensitive Hashing: Discriminative Binary Codes for Source Separation
abstract
Speech enhancement tasks have seen significant improvements with the advance of deep learning technology, but with the cost of increased computational complexity. In this study, we propose an adaptive boosting approach to learning locality sensitive hash codes, which represent audio spectra efficiently. We use the learned hash codes for single-channel speech denoising tasks as an alternative to a complex machine learning model, particularly to address the resource-constrained environments. Our adaptive boosting algorithm learns simple logistic regressors as the weak learners. Once trained, their binary classification results transform each spectrum of test noisy speech into a bit string. Simple bitwise operations calculate Hamming distance to find the K-nearest matching frames in the dictionary of training noisy speech spectra, whose associated ideal binary masks are averaged to estimate the denoising mask for that test mixture. Our proposed learning algorithm differs from AdaBoost in the sense that the projections are trained to minimize the distances between the self-similarity matrix of the hash codes and that of the original spectra, rather than the misclassification rate. We evaluate our discriminative hash codes on the TIMIT corpus with various noise types, and show comparative performance to deep learning methods in terms of denoising performance and complexity.
Sunwoo Kim 0003, Haici Yang, Minje Kim 0001
ICASSP2