Chouchang Yang

dblp:119/3821 · also Chouchang (Jack) Yang · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
12since 2021 · last 2025
0009-0007-3506-6121ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorComputer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 Better Exploiting Spatial Separability in Multichannel Speech Enhancement with an Align-and-Filter Network
abstract
Multichannel speech enhancement (SE) techniques combine multiple microphone signals to extract clean speech from noisy mixtures based on spatial filtering. As the target speech may come from arbitrary, unknown directions, current deep learning-based SE systems could suffer from performance bottleneck in denoising speech within one stage. In contrast, conventional signal processing algorithms often feature a two-stage design, where the first stage focuses on spatially aligning the received signals with respect to the speech source, followed by the second stage to filter out noise. In this paper, we introduce Align-and-Filter network (AFnet) for deep learning-based SE that decouples the primal denoising problem into two sub-problems, which imitates the alignment-followed-by-filtering wisdom from signal processing. The key is to leverage the relative transfer functions (RTFs) that encode meaningful spatial information via a tactically designed alignment strategy. Experimental results show that by leveraging the proposed RTF-based spatial alignment supervision, AFnet learns interpretable directional features to better exploit spatial separability of sound sources for improved SE performance.
Ching Hua Lee, Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Yilin Shen, Hongxia Jin
ICASSP2
2025 MIB: Mixed Information Bottleneck for Out-of-Distribution Keyword Spotting
abstract
Deep Keyword Spotting (KWS) systems continuously process audio streams to detect keywords. However, performance of deep neural networks degrade when the input data diverges from the training data; referred to as Out-of-Distribution (OOD) data problem. In this paper, we show performance degradation of existing State-of-the-Art (SOTA) keyword spotting models on OOD data w.r.t. in-domain testing data, and propose a training mechanism to improve performance on OOD data. Specifically, we propose a novel combination of Mixup and Information Bottleneck, called MIB, to achieve SOTA performance on OOD data. Considering on-device applications, we show across multiple models ranging from sizes of 12.5K parameters to 350K parameters, that MIB achieves as much as 2.5% (absolute) improvement in performance over OOD data. Further, in the more realistic case where OOD keywords are uttered in the presence of OOD noise, MIB achieves as much as 10% (absolute) performance improvement over SOTA models. The proposed MIB is model-agnostic, i.e., it can be applied to enhance the training of any deep keyword spotting model.
Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP5
2025 RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned Prior
abstract
Denoising diffusion probabilistic models (DDPMs) can be utilized to recover a clean signal from its degraded observation(s) by conditioning the model on the degraded signal. The degraded signals are themselves contaminated versions of the clean signals; due to this correlation, they may encompass certain useful information about the target clean data distribution. However, existing adoption of the standard Gaussian as the prior distribution in turn discards such information when shaping the prior, resulting in sub-optimal performance. In this paper, we propose to improve conditional DDPMs for signal restoration by leveraging a more informative prior that is jointly learned with the diffusion model. The proposed framework, called RestoreGrad, seamlessly integrates DDPMs into the variational autoencoder (VAE) framework, taking advantage of the correlation between the degraded and clean signals to encode a better diffusion prior. On speech and image restoration tasks, we show that RestoreGrad demonstrates faster convergence (5-10 times fewer training steps) to achieve better quality of restored signals over existing DDPM baselines and improved robustness to using fewer sampling steps in inference time (2-2.5 times fewer), advocating the advantages of leveraging jointly learned prior for efficiency improvements in the diffusion process.
Ching Hua Lee, Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Yilin Shen, Hongxia Jin
ICML2
2024 End-To-End Personalized Cuff-Less Blood Pressure Monitoring Using ECG and PPG Signals
abstract
Cuffless blood pressure (BP) monitoring offers the potential for continuous, non-invasive healthcare but has been limited in adoption by existing models relying on handcrafted features from ECG and PPG signals. To overcome this, researchers have looked to deep learning. Along these lines, in this paper, we introduce a novel end-to-end model based on transformers. Further, we also introduce a novel contrastive loss-based loss function for robust training. To study the limits of performance for our proposed ideas, we first study personalized models trained on large subject-specific datasets, and achieve an average mean absolute error of 1.08/0.68 mmHg for systolic (SBP) and diastolic BP (DBP) across all subjects while achieving a best case of 0.29/0.19 mmHg. Further, in the case where subject-specific data is scarce, we leverage transfer learning using multi-subject data, and show that our model outperforms State-of-the-Art (SOTA) methods across varying amounts of subject-specific data.
Suhas BN, Rakshith Sharma Srinivasa, Yashas Malur Saidutta, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP6
2024 Zero-Shot Intent Classification Using a Semantic Similarity Aware Contrastive Loss and Large Language Model
abstract
Zero-shot systems can reduce the cost of collecting data and training in a new domain since they can work directly with the test data without further training. In this paper, we build zero-shot systems for intent classification, based on Semantic Similarity-aware Contrastive Loss (SSCL) that addresses an issue in the original CL which treats non-corresponding pairs indiscriminately. We confirm that SSCL outperforms CL through experiments. Then, we explore how including text or speech in-domain data during the SSCL training affects the out-of-domain intent classification.During the zero-shot classification, embeddings for a set of classes in the new domain are generated to calculate the similarities between each class embedding and an input utterance embedding, after which the most similar class is predicted for the utterance’s intent. Although manually-collected text sentences per class can be used to generate the class embedding, the data collection can be costly. Thus, we explore how to generate better class embeddings without human-collected text data in the target domain. The best proposed method employing an instruction-tuned Llama2, a public large language model, shows the performance comparable to the case where the human-collected text data was used, implying the importance of accurate class embedding generation.
Rakshith Sharma Srinivasa, Ching Hua Lee, Yashas Malur Saidutta, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP5
2024 An MVDR-Embedded U-Net Beamformer for Effective and Robust Multichannel Speech Enhancement
abstract
In multichannel speech enhancement (SE) systems, deep neural networks (DNNs) are often utilized to directly estimate the clean speech for effective beamforming. This approach, however, may not generalize adequately to new acoustic or noise conditions. Alternatively, DNNs can indirectly perform SE by predicting the time-frequency masks of speech and noise patterns to assist classic statistical beamformers. Despite being robust, its effectiveness is constrained by the later statistical component relying on certain modeling assumptions, e.g., covariance-based modeling in the minimum-variance-distortionless-response (MVDR) beamformer. In this paper, we propose a novel integration of the two types of methodology, by introducing an intra-MVDR module embedded in the U-Net beamformer, that encompasses the merits of both, i.e., effectiveness and robustness. Experiments show that intra-MVDR leads to improvements that are not achievable by simply enlarging the baseline SE network.
Ching Hua Lee, Kashyap Patel, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP3
2024 Leveraging Self-Supervised Speech Representations for Domain Adaptation in Speech Enhancement
abstract
Deep learning based speech enhancement (SE) approaches could suffer from performance degradation due to mismatch between training and testing environments. A realistic situation is that an SE model trained on parallel noisy-clean utterances from one environment, the source domain, may fail to perform adequately in another environment, the target (new) domain of unseen acoustic or noise conditions. Even though we can improve the target domain performance by leveraging paired data in that domain, in reality, noisy data is more straightforward to collect. Therefore, it is worth studying unsupervised domain adaptation techniques for SE that utilize only noisy data from the target domain, together with exploiting the knowledge available from the source domain paired data, for improved SE in the new domain. In this paper, we present a novel adaptation framework for SE by leveraging self-supervised learning (SSL) based speech models. SSL models are pre-trained with large amount of raw speech data to extract representations rich in phonetic and acoustics information. We explore the potential of leveraging SSL representations for effective SE adaptation to new domains. To our knowledge, it is the first attempt to apply SSL models for domain adaptation in SE.
Ching Hua Lee, Chouchang Yang, Rakshith Sharma Srinivasa, Yashas Malur Saidutta, Yilin Shen, Hongxia Jin
ICASSP2
2024 CIFD: Controlled Information Flow to Enhance Knowledge Distillation
abstract
Knowledge Distillation is the mechanism by which the insights gained from a larger teacher model are transferred to a smaller student model. However, the transfer suffers when the teacher model is significantly larger than the student. To overcome this, prior works have proposed training intermediately sized models, Teacher Assistants (TAs) to help the transfer process. However, training TAs is expensive, as training these models is a knowledge transfer task in itself. Further, these TAs are larger than the student model and training them especially in large data settings can be computationally intensive. In this paper, we propose a novel framework called Controlled Information Flow for Knowledge Distillation (CIFD) consisting of two components. First, we propose a significantly smaller alternatives to TAs, the Rate-Distortion Module (RDM) which uses the teacher's penultimate layer embedding and a information rate-constrained bottleneck layer to replace the Teacher Assistant model. RDMs are smaller and easier to train than TAs, especially in large data regimes, since they operate on the teacher embeddings and do not need to relearn low level input feature extractors. Also, by varying the information rate across the bottleneck, RDMs can replace TAs of different sizes. Secondly, we propose the use of Information Bottleneck Module in the student model, which is crucial for regularization in the presence of a large number of RDMs. We show comprehensive state-of-the-art results of the proposed method over large datasets like Imagenet. Further, we show the significant improvement in distilling CLIP like models over a huge 12M image-text dataset. It outperforms CLIP specialized distillation methods across five zero-shot classification datasets and two zero-shot image-text retrieval datasets.
Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin
NeurIPS5
2023 Improved Mask-Based Neural Beamforming for Multichannel Speech Enhancement by Snapshot Matching Masking
abstract
In multichannel speech enhancement (SE), time-frequency (T-F) mask-based neural beamforming algorithms take advantage of deep neural networks to predict T-F masks that represent speech and noise dominance. The predicted masks are subsequently leveraged to estimate the speech and noise power spectral density (PSD) matrices for computing the beamformer filter weights based on signal statistics. However, in the literature most networks are trained to estimate some pre-defined masks, e.g., the ideal binary mask (IBM) and ideal ratio mask (IRM) that lack direct connection to the PSD estimation. In this paper, we propose a new masking strategy to predict the Snapshot Matching Mask (SMM) that aims to minimize the distance between the predicted and the true signal snapshots, thereby estimating the PSD matrices in a more systematic way. Performance of SMM compared with existing IBM- and IRM-based PSD estimation for mask-based neural beamforming is presented on several datasets to demonstrate its effectiveness for the SE task.
Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP2
2023 To Wake-Up or Not to Wake-Up: Reducing Keyword False Alarm by Successive Refinement
abstract
Keyword spotting systems continuously process audio streams to detect keywords. One of the most challenging tasks in designing such systems is to reduce False Alarm (FA) which happens when the system falsely registers a keyword despite the keyword not being uttered. In this paper, we propose a simple yet elegant solution to this problem that follows from the law of total probability. We show that existing deep keyword spotting mechanisms can be improved by Successive Refinement, where the system first classifies whether the input audio is speech or not, followed by whether the input is keyword-like or not, and finally classifies which keyword was uttered. We show across multiple models with size ranging from 13K parameters to 2.41M parameters, the successive refinement technique reduces FA by up to a factor of 8 on in-domain held-out FA data, and up to a factor of 7 on out-of-domain (OOD) FA data. Further, our proposed approach is "plug-and-play" and can be applied to any deep keyword spotting model.
Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin
ICASSP4
2023 Robust Keyword Spotting for Noisy Environments by Leveraging Speech Enhancement and Speech Presence Probability
Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Yilin Shen, Hongxia Jin
INTERSPEECH1
2023 CWCL: Cross-Modal Transfer with Continuously Weighted Contrastive Loss
abstract
This paper considers contrastive training for cross-modal 0-shot transfer wherein a pre-trained model in one modality is used for representation learning in another domain using pairwise data. The learnt models in the latter domain can then be used for a diverse set of tasks in a 0-shot way, similar to Contrastive Language-Image Pre-training (CLIP) and Locked-image Tuning (LiT) that have recently gained considerable attention. Classical contrastive training employs sets of positive and negative examples to align similar and repel dissimilar training data samples. However, similarity amongst training examples has a more continuous nature, thus calling for a more `non-binary' treatment. To address this, we propose a new contrastive loss function called Continuously Weighted Contrastive Loss (CWCL) that employs a continuous measure of similarity. With CWCL, we seek to transfer the structure of the embedding space from one modality to another. Owing to the continuous nature of similarity in the proposed loss function, these models outperform existing methods for 0-shot transfer across multiple models, datasets and modalities. By using publicly available datasets, we achieve 5-8% (absolute) improvement over previous state-of-the-art methods in 0-shot image classification and 20-30% (absolute) improvement in 0-shot speech-to-intent classification and keyword classification.
Rakshith Sharma Srinivasa, Chouchang Yang, Yashas Malur Saidutta, Ching Hua Lee, Yilin Shen, Hongxia Jin
NeurIPS3
2018 Wall++: Room-Scale Interactive and Context-Aware Sensing
abstract
Human environments are typified by walls, homes, offices, schools, museums, hospitals and pretty much every indoor context one can imagine has walls. In many cases, they make up a majority of readily accessible indoor surface area, and yet they are static their primary function is to be a wall, separating spaces and hiding infrastructure. We present Wall++, a low-cost sensing approach that allows walls to become a smart infrastructure. Instead of merely separating spaces, walls can now enhance rooms with sensing and interactivity. Our wall treatment and sensing hardware can track users' touch and gestures, as well as estimate body pose if they are close. By capturing airborne electromagnetic noise, we can also detect what appliances are active and where they are located. Through a series of evaluations, we demonstrate Wall++ can enable robust room-scale interactive and context-aware applications.
Yang Zhang 0041, Chouchang Yang, Scott E. Hudson, Chris Harrison 0001, Alanson P. Sample
CHI2
2017 Riding the airways: Ultra-wideband ambient backscatter via commercial broadcast systems
abstract
Communication costs dominate the energy consumption, and ultimately limit the utility, of low power devices and sensor nodes. Backscatter communication based on deliberate and ambient sources has the potential to radically alter this paradigm by offering two to three orders of magnitude better communication efficiency (in terms of nJ/Bit) then conventional radio architectures. Initial work on ambient backscatter shows promising results but has focused on narrow band operation in well controlled laboratory settings. The goal of this work is to enable the ubiquitous deployment of ultra-low power nodes that communicate via ambient backscatter to wired Universal Backscatter Readers, in real-world environments. This is accomplished through ultra-wideband backscatter techniques that leverage the breath of commercial broadcast signals in the 80 MHz to 900 MHz range from FM radios, digital TVs, and cellular networks. Additionally the use of powered Universal Backscatter Readers allows a network of ultra-low power nodes to operate on ambient carriers as low as -80 dBm, which is typical for indoor home and office environments. For the first time we demonstrate the simultaneous use of 17 ambient signal sources to achieve node-to-reader communication distances of 50 meters, with data rates up to 1 kbps.
Chouchang Yang, Jeremy Gummeson, Alanson P. Sample
INFOCOM1
2015 EM-Sense: Touch Recognition of Uninstrumented, Electrical and Electromechanical Objects
abstract
Most everyday electrical and electromechanical objects emit small amounts of electromagnetic (EM) noise during regular operation. When a user makes physical contact with such an object, this EM signal propagates through the user, owing to the conductivity of the human body. By modifying a small, low-cost, software-defined radio, we can detect and classify these signals in real-time, enabling robust on-touch object detection. Unlike prior work, our approach requires no instrumentation of objects or the environment; our sensor is self-contained and can be worn unobtrusively on the body. We call our technique EM-Sense and built a proof-of-concept smartwatch implementation. Our studies show that discrimination between dozens of objects is feasible, independent of wearer, time and local environment.
Gierad Laput, Chouchang Yang, Robert Xiao, Alanson P. Sample, Chris Harrison 0001
UIST2
2013 Optimized relay-route assignment for anonymity in wireless networks
abstract
Anonymous wireless networks use covert relays to prevent unauthorized entities from determining communicating parties through traffic timing analysis. In a multipath anonymous network, the choice of which relay nodes should be covert, as well as the route selection by the network nodes, affect both the anonymity and network performance. Although assigning relays as covert and selecting routes composed of covert relays can provide higher anonymity, the selection of these two parameters will increase the packet dropping rate of the network. In this paper, we introduce an analytical framework for joint relay assignment and route selection in multi-path anonymous wireless networks. The main contributions of this work are two-fold. First, we show that joint relay assignment and route selection can be formulated as a convex optimization problem which guarantees global optimum solution. Second, as special cases of our formulation, we derive solutions for the problem of route selection to maximize anonymity when the relay configuration is given, as well as the problem of relay configuration for a given route selection.
Chouchang Yang, Basel Alomair, Radha Poovendran
ISIT1
2012 Optimized flow allocation for anonymous communication in multipath wireless networks
abstract
In anonymous networks, a subset of nodes is chosen to act as covert relays to hide timing information from unauthorized observers. While such covert relays increase anonymity, they cause performance degradation by delaying or dropping packets. In this paper, we propose flow allocation methods that maximize anonymity for multipath wireless networks with predetermined covert relay nodes, while taking into account packet-loss as a constraint. Using a rate-distortion framework, we show how to assign probabilities which split the flows from source to destination among all possible routes and show that selecting routes according to the assigned probabilities achieves maximum anonymity given the packet-loss constraint.
Chouchang Yang, Basel Alomair, Radha Poovendran
ISIT1