EDBT 2026 Demo / reviewers in the wild / expert
Ching Hua Lee
dblp:120/7651
· DBLP profile ↗
19ranked-venue papers
10as first author
15since 2021 · last 2025
0000-0001-6116-3653ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Better Exploiting Spatial Separability in Multichannel Speech Enhancement with an Align-and-Filter NetworkabstractMultichannel speech enhancement (SE) techniques combine multiple microphone signals to extract clean speech from noisy mixtures based on spatial filtering. As the target speech may come from arbitrary, unknown directions, current deep learning-based SE systems could suffer from performance bottleneck in denoising speech within one stage. In contrast, conventional signal processing algorithms often feature a two-stage design, where the first stage focuses on spatially aligning the received signals with respect to the speech source, followed by the second stage to filter out noise. In this paper, we introduce Align-and-Filter network (AFnet) for deep learning-based SE that decouples the primal denoising problem into two sub-problems, which imitates the alignment-followed-by-filtering wisdom from signal processing. The key is to leverage the relative transfer functions (RTFs) that encode meaningful spatial information via a tactically designed alignment strategy. Experimental results show that by leveraging the proposed RTF-based spatial alignment supervision, AFnet learns interpretable directional features to better exploit spatial separability of sound sources for improved SE performance. Ching Hua Lee, Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Yilin Shen, Hongxia Jin |
ICASSP | 1 |
| 2025 | MIB: Mixed Information Bottleneck for Out-of-Distribution Keyword SpottingabstractDeep Keyword Spotting (KWS) systems continuously process audio streams to detect keywords. However, performance of deep neural networks degrade when the input data diverges from the training data; referred to as Out-of-Distribution (OOD) data problem. In this paper, we show performance degradation of existing State-of-the-Art (SOTA) keyword spotting models on OOD data w.r.t. in-domain testing data, and propose a training mechanism to improve performance on OOD data. Specifically, we propose a novel combination of Mixup and Information Bottleneck, called MIB, to achieve SOTA performance on OOD data. Considering on-device applications, we show across multiple models ranging from sizes of 12.5K parameters to 350K parameters, that MIB achieves as much as 2.5% (absolute) improvement in performance over OOD data. Further, in the more realistic case where OOD keywords are uttered in the presence of OOD noise, MIB achieves as much as 10% (absolute) performance improvement over SOTA models. The proposed MIB is model-agnostic, i.e., it can be applied to enhance the training of any deep keyword spotting model. Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 4 |
| 2025 | RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned PriorabstractDenoising diffusion probabilistic models (DDPMs) can be utilized to recover a clean signal from its degraded observation(s) by conditioning the model on the degraded signal. The degraded signals are themselves contaminated versions of the clean signals; due to this correlation, they may encompass certain useful information about the target clean data distribution. However, existing adoption of the standard Gaussian as the prior distribution in turn discards such information when shaping the prior, resulting in sub-optimal performance. In this paper, we propose to improve conditional DDPMs for signal restoration by leveraging a more informative prior that is jointly learned with the diffusion model. The proposed framework, called RestoreGrad, seamlessly integrates DDPMs into the variational autoencoder (VAE) framework, taking advantage of the correlation between the degraded and clean signals to encode a better diffusion prior. On speech and image restoration tasks, we show that RestoreGrad demonstrates faster convergence (5-10 times fewer training steps) to achieve better quality of restored signals over existing DDPM baselines and improved robustness to using fewer sampling steps in inference time (2-2.5 times fewer), advocating the advantages of leveraging jointly learned prior for efficiency improvements in the diffusion process. Ching Hua Lee, Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Yilin Shen, Hongxia Jin |
ICML | 1 |
| 2024 | End-To-End Personalized Cuff-Less Blood Pressure Monitoring Using ECG and PPG SignalsabstractCuffless blood pressure (BP) monitoring offers the potential for continuous, non-invasive healthcare but has been limited in adoption by existing models relying on handcrafted features from ECG and PPG signals. To overcome this, researchers have looked to deep learning. Along these lines, in this paper, we introduce a novel end-to-end model based on transformers. Further, we also introduce a novel contrastive loss-based loss function for robust training. To study the limits of performance for our proposed ideas, we first study personalized models trained on large subject-specific datasets, and achieve an average mean absolute error of 1.08/0.68 mmHg for systolic (SBP) and diastolic BP (DBP) across all subjects while achieving a best case of 0.29/0.19 mmHg. Further, in the case where subject-specific data is scarce, we leverage transfer learning using multi-subject data, and show that our model outperforms State-of-the-Art (SOTA) methods across varying amounts of subject-specific data. Suhas BN, Rakshith Sharma Srinivasa, Yashas Malur Saidutta, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 5 |
| 2024 | Zero-Shot Intent Classification Using a Semantic Similarity Aware Contrastive Loss and Large Language ModelabstractZero-shot systems can reduce the cost of collecting data and training in a new domain since they can work directly with the test data without further training. In this paper, we build zero-shot systems for intent classification, based on Semantic Similarity-aware Contrastive Loss (SSCL) that addresses an issue in the original CL which treats non-corresponding pairs indiscriminately. We confirm that SSCL outperforms CL through experiments. Then, we explore how including text or speech in-domain data during the SSCL training affects the out-of-domain intent classification.During the zero-shot classification, embeddings for a set of classes in the new domain are generated to calculate the similarities between each class embedding and an input utterance embedding, after which the most similar class is predicted for the utterance’s intent. Although manually-collected text sentences per class can be used to generate the class embedding, the data collection can be costly. Thus, we explore how to generate better class embeddings without human-collected text data in the target domain. The best proposed method employing an instruction-tuned Llama2, a public large language model, shows the performance comparable to the case where the human-collected text data was used, implying the importance of accurate class embedding generation. Rakshith Sharma Srinivasa, Ching Hua Lee, Yashas Malur Saidutta, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 3 |
| 2024 | An MVDR-Embedded U-Net Beamformer for Effective and Robust Multichannel Speech EnhancementabstractIn multichannel speech enhancement (SE) systems, deep neural networks (DNNs) are often utilized to directly estimate the clean speech for effective beamforming. This approach, however, may not generalize adequately to new acoustic or noise conditions. Alternatively, DNNs can indirectly perform SE by predicting the time-frequency masks of speech and noise patterns to assist classic statistical beamformers. Despite being robust, its effectiveness is constrained by the later statistical component relying on certain modeling assumptions, e.g., covariance-based modeling in the minimum-variance-distortionless-response (MVDR) beamformer. In this paper, we propose a novel integration of the two types of methodology, by introducing an intra-MVDR module embedded in the U-Net beamformer, that encompasses the merits of both, i.e., effectiveness and robustness. Experiments show that intra-MVDR leads to improvements that are not achievable by simply enlarging the baseline SE network. Ching Hua Lee, Kashyap Patel, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 1 |
| 2024 | Leveraging Self-Supervised Speech Representations for Domain Adaptation in Speech EnhancementabstractDeep learning based speech enhancement (SE) approaches could suffer from performance degradation due to mismatch between training and testing environments. A realistic situation is that an SE model trained on parallel noisy-clean utterances from one environment, the source domain, may fail to perform adequately in another environment, the target (new) domain of unseen acoustic or noise conditions. Even though we can improve the target domain performance by leveraging paired data in that domain, in reality, noisy data is more straightforward to collect. Therefore, it is worth studying unsupervised domain adaptation techniques for SE that utilize only noisy data from the target domain, together with exploiting the knowledge available from the source domain paired data, for improved SE in the new domain. In this paper, we present a novel adaptation framework for SE by leveraging self-supervised learning (SSL) based speech models. SSL models are pre-trained with large amount of raw speech data to extract representations rich in phonetic and acoustics information. We explore the potential of leveraging SSL representations for effective SE adaptation to new domains. To our knowledge, it is the first attempt to apply SSL models for domain adaptation in SE. Ching Hua Lee, Chouchang Yang, Rakshith Sharma Srinivasa, Yashas Malur Saidutta, Yilin Shen, Hongxia Jin |
ICASSP | 1 |
| 2024 | CIFD: Controlled Information Flow to Enhance Knowledge DistillationabstractKnowledge Distillation is the mechanism by which the insights gained from a larger teacher model are transferred to a smaller student model. However, the transfer suffers when the teacher model is significantly larger than the student. To overcome this, prior works have proposed training intermediately sized models, Teacher Assistants (TAs) to help the transfer process. However, training TAs is expensive, as training these models is a knowledge transfer task in itself. Further, these TAs are larger than the student model and training them especially in large data settings can be computationally intensive. In this paper, we propose a novel framework called Controlled Information Flow for Knowledge Distillation (CIFD) consisting of two components. First, we propose a significantly smaller alternatives to TAs, the Rate-Distortion Module (RDM) which uses the teacher's penultimate layer embedding and a information rate-constrained bottleneck layer to replace the Teacher Assistant model. RDMs are smaller and easier to train than TAs, especially in large data regimes, since they operate on the teacher embeddings and do not need to relearn low level input feature extractors. Also, by varying the information rate across the bottleneck, RDMs can replace TAs of different sizes. Secondly, we propose the use of Information Bottleneck Module in the student model, which is crucial for regularization in the presence of a large number of RDMs. We show comprehensive state-of-the-art results of the proposed method over large datasets like Imagenet. Further, we show the significant improvement in distilling CLIP like models over a huge 12M image-text dataset. It outperforms CLIP specialized distillation methods across five zero-shot classification datasets and two zero-shot image-text retrieval datasets. Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
NeurIPS | 4 |
| 2023 | A DNN Based Normalized Time-Frequency Weighted Criterion for Robust Wideband DoA EstimationabstractDeep neural networks (DNNs) have greatly benefited direction of arrival (DoA) estimation methods for speech source localization in noisy environments. However, their localization accuracy is still far from satisfactory due to the vulnerability to nonspeech interference. To improve the robustness against interference, we propose a DNN based normalized time-frequency (T-F) weighted criterion which minimizes the distance between the candidate steering vectors and the filtered snapshots in the T-F domain. Our method requires no eigendecomposition and uses a simple normalization to prevent the optimization objective from being misled by noisy filtered snapshots. We also study different designs of T-F weights guided by a DNN. We find that duplicating the Hadamard product of speech ratio masks is highly effective and better than other techniques such as direct masking and taking the mean in the proposed approach. However, the best-performing design of T-F weights is criterion-dependent in general. Experiments show that the proposed method outperforms popular DNN based DoA estimation methods including widely used subspace methods in noisy and reverberant environments. Kuan-Lin Chen 0002, Ching Hua Lee, Bhaskar D. Rao, Harinath Garudadri |
ICASSP | 2 |
| 2023 | Improved Mask-Based Neural Beamforming for Multichannel Speech Enhancement by Snapshot Matching MaskingabstractIn multichannel speech enhancement (SE), time-frequency (T-F) mask-based neural beamforming algorithms take advantage of deep neural networks to predict T-F masks that represent speech and noise dominance. The predicted masks are subsequently leveraged to estimate the speech and noise power spectral density (PSD) matrices for computing the beamformer filter weights based on signal statistics. However, in the literature most networks are trained to estimate some pre-defined masks, e.g., the ideal binary mask (IBM) and ideal ratio mask (IRM) that lack direct connection to the PSD estimation. In this paper, we propose a new masking strategy to predict the Snapshot Matching Mask (SMM) that aims to minimize the distance between the predicted and the true signal snapshots, thereby estimating the PSD matrices in a more systematic way. Performance of SMM compared with existing IBM- and IRM-based PSD estimation for mask-based neural beamforming is presented on several datasets to demonstrate its effectiveness for the SE task. Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 1 |
| 2023 | To Wake-Up or Not to Wake-Up: Reducing Keyword False Alarm by Successive RefinementabstractKeyword spotting systems continuously process audio streams to detect keywords. One of the most challenging tasks in designing such systems is to reduce False Alarm (FA) which happens when the system falsely registers a keyword despite the keyword not being uttered. In this paper, we propose a simple yet elegant solution to this problem that follows from the law of total probability. We show that existing deep keyword spotting mechanisms can be improved by Successive Refinement, where the system first classifies whether the input audio is speech or not, followed by whether the input is keyword-like or not, and finally classifies which keyword was uttered. We show across multiple models with size ranging from 13K parameters to 2.41M parameters, the successive refinement technique reduces FA by up to a factor of 8 on in-domain held-out FA data, and up to a factor of 7 on out-of-domain (OOD) FA data. Further, our proposed approach is "plug-and-play" and can be applied to any deep keyword spotting model. Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 3 |
| 2023 | Robust Keyword Spotting for Noisy Environments by Leveraging Speech Enhancement and Speech Presence Probability
Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Yilin Shen, Hongxia Jin |
INTERSPEECH | 4 |
| 2023 | CWCL: Cross-Modal Transfer with Continuously Weighted Contrastive LossabstractThis paper considers contrastive training for cross-modal 0-shot transfer wherein a pre-trained model in one modality is used for representation learning in another domain using pairwise data. The learnt models in the latter domain can then be used for a diverse set of tasks in a 0-shot way, similar to Contrastive Language-Image Pre-training (CLIP) and Locked-image Tuning (LiT) that have recently gained considerable attention. Classical contrastive training employs sets of positive and negative examples to align similar and repel dissimilar training data samples. However, similarity amongst training examples has a more continuous nature, thus calling for a more `non-binary' treatment. To address this, we propose a new contrastive loss function called Continuously Weighted Contrastive Loss (CWCL) that employs a continuous measure of similarity. With CWCL, we seek to transfer the structure of the embedding space from one modality to another. Owing to the continuous nature of similarity in the proposed loss function, these models outperform existing methods for 0-shot transfer across multiple models, datasets and modalities. By using publicly available datasets, we achieve 5-8% (absolute) improvement over previous state-of-the-art methods in 0-shot image classification and 20-30% (absolute) improvement in 0-shot speech-to-intent classification and keyword classification. Rakshith Sharma Srinivasa, Chouchang Yang, Yashas Malur Saidutta, Ching Hua Lee, Yilin Shen, Hongxia Jin |
NeurIPS | 5 |
| 2021 | ResNEsts and DenseNEsts: Block-based DNN Models with Improved Representation GuaranteesabstractModels recently used in the literature proving residual networks (ResNets) are better than linear predictors are actually different from standard ResNets that have been widely used in computer vision. In addition to the assumptions such as scalar-valued output or single residual block, the models fundamentally considered in the literature have no nonlinearities at the final residual representation that feeds into the final affine layer. To codify such a difference in nonlinearities and reveal a linear estimation property, we define ResNEsts, i.e., Residual Nonlinear Estimators, by simply dropping nonlinearities at the last residual representation from standard ResNets. We show that wide ResNEsts with bottleneck blocks can always guarantee a very desirable training property that standard ResNets aim to achieve, i.e., adding more blocks does not decrease performance given the same set of basis elements. To prove that, we first recognize ResNEsts are basis function models that are limited by a coupling problem in basis learning and linear prediction. Then, to decouple prediction weights from basis learning, we construct a special architecture termed augmented ResNEst (A-ResNEst) that always guarantees no worse performance with the addition of a block. As a result, such an A-ResNEst establishes empirical risk lower bounds for a ResNEst using corresponding bases. Our results demonstrate ResNEsts indeed have a problem of diminishing feature reuse; however, it can be avoided by sufficiently expanding or widening the input space, leading to the above-mentioned desirable property. Inspired by the densely connected networks (DenseNets) that have been shown to outperform ResNets, we also propose a corresponding new model called Densely connected Nonlinear Estimator (DenseNEst). We show that any DenseNEst can be represented as a wide ResNEst with bottleneck blocks. Unlike ResNEsts, DenseNEsts exhibit the desirable property without any special architectural re-design. Kuan-Lin Chen 0002, Ching Hua Lee, Harinath Garudadri, Bhaskar D. Rao |
NeurIPS | 2 |
| 2021 | Proportionate Adaptive Filtering Algorithms Derived Using an Iterative Reweighting FrameworkabstractSSR methods that majorize the regularized objective function during the optimization process. We show that introducing the majorizers leads to the same algorithm as simply using the gradient update of the regularized objective function, as is done in existing approaches. Different from the past works, the reweighting formulation naturally leads to an affine scaling transformation (AST) strategy, which effectively introduces a diagonal weighting on the gradient, giving rise to new algorithms that demonstrate improved convergence properties. Interestingly, setting the regularization coefficient to zero in the proposed AST-based framework leads to the Sparsity-promoting LMS (SLMS) and Sparsity-promoting Normalized LMS (SNLMS) algorithms, which exploit but do not strictly enforce the sparsity of the system response if it already exists. The SLMS and SNLMS realize proportionate adaptation for convergence speedup should sparsity be present in the underlying system response. In this manner, we develop a new way for rigorously deriving a large class of proportionate algorithms, and also explain why they are useful in applications where the underlying systems admit certain sparsity, e.g., in acoustic echo and feedback cancellation. Ching Hua Lee, Bhaskar D. Rao, Harinath Garudadri |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | SSGD: Sparsity-Promoting Stochastic Gradient Descent Algorithm for Unbiased Dnn PruningabstractWhile deep neural networks (DNNs) have achieved state-of-the-art results in many fields, they are typically over-parameterized. Parameter redundancy, in turn, leads to inefficiency. Sparse signal recovery (SSR) techniques, on the other hand, find compact solutions to overcomplete linear problems. Therefore, a logical step is to draw the connection between SSR and DNNs. In this paper, we explore the application of iterative reweighting methods popular in SSR to learning efficient DNNs. By efficient, we mean sparse networks that require less computation and storage than the original, dense network. We propose a reweighting framework to learn sparse connections within a given architecture without biasing the optimization process, by utilizing the affine scaling transformation strategy. The resulting algorithm, referred to as Sparsity-promoting Stochastic Gradient Descent (SSGD), has simple gradient-based updates which can be easily implemented in existing deep learning libraries. We demonstrate the sparsification ability of SSGD on image classification tasks and show that it outperforms existing methods on the MNIST and CIFAR-10 datasets. Ching Hua Lee, Igor Fedorov, Bhaskar D. Rao, Harinath Garudadri |
ICASSP | 1 |
| 2020 | A Sparse Conjugate Gradient Adaptive FilterabstractIn this letter, we propose a novel conjugate gradient (CG) adaptive filtering algorithm for online estimation of system responses that admit sparsity. Specifically, the Sparsity-promoting Conjugate Gradient (SCG) algorithm is developed based on iterative reweighting methods popular in the sparse signal recovery area. We propose an affine scaling transformation strategy within the reweighting framework, leading to an algorithm that allows the usage of a zero sparsity regularization coefficient. This enables SCG to leverage the sparsity of the system response if it already exists, while not compromising the optimization process. Simulation results show that SCG demonstrates improved convergence and steady-state properties over existing methods. Ching Hua Lee, Bhaskar D. Rao, Harinath Garudadri |
IEEE Signal Process. Lett. | 1 |
| 2019 | On Mitigating Acoustic Feedback in Hearing Aids with Frequency Warping by All-Pass NetworksabstractAcoustic feedback control continues to be a challenging problem due to the emerging form factors in advanced hearing aids (HAs) and hearables. In this paper, we present a novel use of well-known all-pass filters in a network to perform frequency warping that we call "freping." Freping helps in breaking the Nyquist stability criterion and improves adaptive feedback cancellation (AFC). Based on informal subjective assessments, distortions due to freping are fairly benign. While common objective metrics like the perceptual evaluation of speech quality (PESQ) and the hearing-aid speech quality index (HASQI) may not adequately capture distortions due to freping and acoustic feedback artifacts from a perceptual perspective, they are still instructive in assessing the proposed method. We demonstrate quality improvements with freping for a basic AFC (PESQ: 2.56 to 3.52 and HASQI: 0.65 to 0.78) at a gain setting of 20; and an advanced AFC (PESQ: 2.75 to 3.17 and HASQI: 0.66 to 0.73) for a gain of 30. From our investigations, freping provides larger improvement for basic AFC, but still improves overall system performance for many AFC approaches. Ching Hua Lee, Kuan-Lin Chen 0002, Fredric J. Harris, Bhaskar D. Rao, Harinath Garudadri |
INTERSPEECH | 1 |
| 2018 | Bone-Conduction Sensor Assisted Noise Estimation for Improved Speech EnhancementabstractState-of-the-art noise power spectral density (PSD) estimation techniques for speech enhancement utilize the so-called speech presence probability (SPP). However, in highly non-stationary environments, SPP-based techniques could still suffer from inaccurate estimation, leading to significant amount of residual noise or speech distortion. In this paper, we propose to improve speech enhancement by deploying the bone-conduction (BC) sensor, which is known to be relatively insensitive to the environmental noise compared to the regular air-conduction (AC) microphone. A strategy is suggested to utilized the BC sensor characteristics for assisting the AC microphone in better SPP-based noise estimation. To our knowledge, no previous work has incorporated the BC sensor in this noise estimation aspect. Consequently, the proposed strategy can possibly be combined with other BC sensor assisted speech enhancement techniques. We show the feasibility and potential of the proposed method for improving the enhanced speech quality by both objective and subjective tests. Ching Hua Lee, Bhaskar D. Rao, Harinath Garudadri |
INTERSPEECH | 1 |