VLDB 2026 Research / reviewers in the wild / expert
Karim Helwani
dblp:75/8759
· DBLP profile ↗
19ranked-venue papers
8as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | NOLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal ShapingabstractSpeech codec enhancement methods are designed to remove distortions added by speech codecs. While classical methods are very low in complexity and add zero delay, their effectiveness is rather limited. Compared to that, DNN-based methods deliver higher quality but they are typically high in complexity and/or require delay. The recently proposed Linear Adaptive Coding Enhancer (LACE) addresses this problem by combining DNNs with classical long-term/short-term post-filtering resulting in a causal low-complexity model. A short-coming of the LACE model is, however, that quality quickly saturates when the model size is scaled up. To mitigate this problem, we propose a novel adatpive temporal shaping module that adds high temporal resolution to the LACE model resulting in the Non-Linear Adaptive Coding Enhancer (NoLACE). We adapt NoLACE to enhance the Opus codec and show that NoLACE significantly outperforms both the Opus baseline and an enlarged LACE model at 6, 9 and 12 kb/s. We also show that LACE and NoLACE are well-behaved when used with an ASR system. Jan Büthe, Ahmed Mustafa, Jean-Marc Valin, Karim Helwani, Michael M. Goodwin |
ICASSP | 4 |
| 2024 | Real-Time Stereo Speech Enhancement with Spatial-Cue Preservation Based on Dual-Path StructureabstractWe introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source has unique spatial information, our method utilizes a dual-path structure, ensuring the spatial cues remain unaffected during enhancement by applying source-specific common-band gain. This method also seamlessly integrates pretrained monaural speech enhancement, eliminating the need for retraining on stereo inputs. Source separation from stereo mixtures is achieved via spatial beamforming, with the steering vector for each source being adaptively updated using post-enhancement output signal. This ensures accurate tracking of the spatial information. The final stereo output is derived by merging the spatial images of the enhanced sources, with its efficacy not heavily reliant on the separation performance of the beamforming. The algorithm runs in real-time on 10-ms frames with a 40 ms of look-ahead. Evaluations reveal its effectiveness in enhancing speech and preserving spatial cues in both fully and sparsely overlapped mixtures. Masahito Togami, Jean-Marc Valin, Karim Helwani, Ritwik Giri, Umut Isik, Michael M. Goodwin |
ICASSP | 3 |
| 2023 | Generative Modeling Based Manifold Learning for Adaptive Filtering GuidanceabstractIn most practical adaptive filtering problems, estimated filters are not arbitrary, but instead lie on a manifold that encapsulates characteristics of the problem at hand. Consequently, it is desirable to steer adaptation towards filters that lie on that manifold. In this paper, we propose a novel approach to learn the manifold of a set of impulse responses and subsequently employ that learned manifold in an adaptation algorithm for system identification. The presented approach is a practical adaptive filtering recipe for enforcing a data-driven search domain constraint, instead of using conventional constrained optimization methods. Karim Helwani, Paris Smaragdis, Michael M. Goodwin |
ICASSP | 1 |
| 2022 | Clock Skew Robust Acoustic Echo Cancellation
Karim Helwani, Erfan Soltanmohammadi, Michael M. Goodwin, Arvindh Krishnaswamy |
INTERSPEECH | 1 |
| 2021 | Enhancing Audio Augmentation Methods with Consistency LearningabstractData augmentation is an inexpensive way to increase training data diversity, and is commonly achieved via transformations of existing data. For tasks such as classification, there is a good case for learning representations of the data that are invariant to such transformations, yet this is not explicitly enforced by classification losses such as the cross-entropy loss. This paper investigates the use of training objectives that explicitly impose this consistency constraint, and how it can impact downstream audio classification tasks. In the context of deep convolutional neural networks in the supervised setting, we show empirically that certain measures of consistency are not implicitly captured by the cross-entropy loss, and that incorporating such measures into the loss function can improve the performance of tasks such as audio tagging. Put another way, we demonstrate how existing augmentation methods can further improve learning by enforcing consistency. Turab Iqbal, Karim Helwani, Arvindh Krishnaswamy, Wenwu Wang 0001 |
ICASSP | 2 |
| 2021 | Low-Complexity, Real-Time Joint Neural Echo Control and Speech Enhancement Based On PercepnetabstractSpeech enhancement algorithms based on deep learning have greatly surpassed their traditional counterparts and are now being considered for the task of removing acoustic echo from hands-free communication systems. This is a challenging problem due to both real-world constraints like loudspeaker non-linearities, and to limited compute capabilities in some communication systems. In this work, we propose a system combining a traditional acoustic echo canceller, and a low-complexity joint residual echo and noise suppressor based on a hybrid signal processing/deep neural network (DSP/DNN) approach. We show that the proposed system outperforms both traditional and other neural approaches, while requiring only 5.5% CPU for real-time operation. We further show that the system can scale to even lower complexity levels. Jean-Marc Valin, Srikanth V. Tenneti, Karim Helwani, Umut Isik, Arvindh Krishnaswamy |
ICASSP | 3 |
| 2020 | PoCoNet: Better Speech Enhancement with Frequency-Positional Embeddings, Semi-Supervised Conversational Data, and Biased LossabstractNeural network applications generally benefit from larger-sized models, but for current speech enhancement models, larger scale networks often suffer from decreased robustness to the variety of real-world use cases beyond what is encountered in training data. We introduce several innovations that lead to better large neural networks for speech enhancement. The novel PoCoNet architecture is a convolutional neural network that, with the use of frequency-positional embeddings, is able to more efficiently build frequency-dependent features in the early layers. A semi-supervised method helps increase the amount of conversational training data by pre-enhancing noisy datasets, improving performance on real recordings. A new loss function biased towards preserving speech quality helps the optimization better match human perceptual opinions on speech quality. Ablation experiments and objective and human opinion metrics show the benefits of the proposed improvements. Umut Isik, Ritwik Giri, Neerad Phansalkar, Jean-Marc Valin, Karim Helwani, Arvindh Krishnaswamy |
INTERSPEECH | 5 |
| 2020 | A Perceptually-Motivated Approach for Low-Complexity, Real-Time Enhancement of Fullband SpeechabstractOver the past few years, speech enhancement methods based on deep learning have greatly surpassed traditional methods based on spectral subtraction and spectral estimation. Many of these new techniques operate directly in the the short-time Fourier transform (STFT) domain, resulting in a high computational complexity. In this work, we propose PercepNet, an efficient approach that relies on human perception of speech by focusing on the spectral envelope and on the periodicity of the speech. We demonstrate high-quality, real-time enhancement of fullband (48 kHz) speech with less than 5% of a CPU core. Jean-Marc Valin, Umut Isik, Neerad Phansalkar, Ritwik Giri, Karim Helwani, Arvindh Krishnaswamy |
INTERSPEECH | 5 |
| 2019 | Unsupervised Bayesian Estimation and Tracking of Time-Varying Convolutive Multichannel Systems
Herbert Buchner, Karim Helwani, Simon J. Godsill |
FUSION | 2 |
| 2019 | Blind Signal Processing for Time-varying Convolutive Mixing Systems Based on Sequence Estimation on Partly Smooth ManifoldsabstractIn this paper we focus on Bayesian blind and semi-blind adaptive signal processing based on a broadband MIMO FIR model (e.g., for blind source separation (BSS) and blind system identification (BSI)). Specifically, we study in this paper a framework allowing us to systematically incorporate various types of prior knowledge: (1) source signal statistics, (2) deterministic knowledge on the mixing system, and (3) stochastic knowledge on the mixing system. In order to exploit all possible types of source signal statistics (1), our considerations are based on TRINICON, a previously introduced generic framework for broadband blind (and semi-blind) adaptive MIMO signal processing. The motivation for this paper is threefold: (a) the extension of TRINICON to Bayesian point estimation to address (3) in addition to (1), and (b) more specifically to unify system-based blind adaptive MIMO signal processing with the tracking of time-varying scenarios, and finally (c) to show how the Bayesian TRINICON-based tracking can be formulated as a sequence estimation approach on arbitrary partly smooth manifolds. As we will see in this paper, the Bayesian approach to incorporate stochastic priors and the manifold learning approach to exploit deterministic system knowledge (2) complement one another very efficiently in the context of TRINICON. Herbert Buchner, Karim Helwani, Simon J. Godsill |
ICASSP | 2 |
| 2017 | Efficient adaptive filtering in compressive domains for sparse systems and relation to transform-domain adaptive filteringabstractIn this paper we introduce a novel class of efficient multichannel adaptive filtering algorithms for sparse FIR systems. By suitably integrating ideas from compressed sensing and adaptive filter theory, this class of algorithms allows to significantly reduce the actual number of adaptive coefficients in an efficient way. These algorithms, termed compressive-domain adaptive filters, can be interpreted as a novel type of transform-domain techniques. They can also be seen as adaptive approach in an efficiently self-learning manifold based on the prior knowledge of sparseness of the system. An important property of this concept is that it does not place additional restrictions on the input signal characteristics. Based on the well-known RLS algorithm as a reference, the simulation results confirm that the proposed algorithm converges at acceptable rates, even for strongly colored signals such as speech and audio. Herbert Buchner, Karim Helwani, Bashar I. Ahmad, Simon J. Godsill |
ICASSP | 2 |
| 2016 | Localization of sound sources with known statistics in the presence of interferersabstractMethods are available for simultaneous localization of multiple (unknown) audio sources using microphone arrays. Typical algorithms aim at localizing all active sources. They moreover require that the number of sources is known and is less than or equal the number of microphones. This constraint cannot be satisfied in many reallife situations and noisy environments. We present an algorithm for localizing an audio source with known statistics in a multi-source environment. The proposed method circumvents the mentioned problems by using a phase-preserving signal extraction method on the input signal. A binary mask is estimated and used to retain only the information of the target source in the original microphone signals. The masked signals are fed to a modified version of a conventional localization algorithm, which now localizes only the target source. Experimental results obtained from real recordings show that the proposed method can successfully detect and localize the target source. Kainan Chen, Jürgen T. Geiger, Karim Helwani, Mohammad Javad Taghizadeh |
ICASSP | 3 |
| 2015 | Low complexity unsupervised multi-camera color calibration with application to panoramic video capturingabstractIn this study, a novel approach to perform unsupervised color correction on images captured with multiple cameras is presented. The proposed algorithm requires a minimal common reference area for each pair of cameras in the setup, which makes it well suited for high resolution panoramic images. Moreover, due to its simplicity it allows to cope with realtime constraints in an online scenario. Karim Helwani, Lukasz Kondrad, Nicola Piotto |
ICIP | 1 |
| 2013 | On the eigenspace estimation for supervised multichannel system identificationabstractRecently developed multichannel adaptive filtering algorithms aim at spatio-temporal decoupling of the signals by suitably chosen transformations. In this paper we establish the relation between the techniques of transform-domain adaptive filtering with application to multichannel acoustic cancellation. The link between the recently introduced source-domain and eigenspace adaptive filtering algorithms is shown by means of a generic spatial transform-domain adaptive filtering algorithm. We discuss the difference between regularizing the identification problem in the source domain and in the system eigenspace. Further, we study the estimation of the multiple-input multiple-output (MIMO) system eigenspace without modifying the highly cross correlated input signals or requiring prior knowledge of the system and highlight the validity of the estimated eigenspace due to system changes. Finally, we give simulation results proving our concept. Karim Helwani, Herbert Buchner |
ICASSP | 1 |
| 2013 | Multichannel acoustic echo suppressionabstractAcoustic echo suppression (AES) provides an attractive alternative to acoustic echo cancellation (AEC) techniques for full-duplex communication in low-complexity systems. However, so far AES techniques are commonly known to introduce significant distortions to the desired signal. Moreover, most traditional echo control techniques typically require accurately detecting the contribution of the near-end speaker to the microphone signal (“double talk”). The extension of AES techniques to the multichannel case usually assumes a symmetric system design which is often not fulfilled by typical scenarios. In this paper we propose a novel approach to multichannel acoustic echo suppression, which aims at extracting the near-end signal using a constraint for a distortionless output, without requiring a double-talk detector, or a symmetric system design. In addition to the above mentioned properties, the multichannel AES is also shown to overcome the known challenges in conventional multichannel acoustic echo control setups. Karim Helwani, Herbert Buchner, Jacob Benesty, Jingdong Chen |
ICASSP | 1 |
| 2013 | A study of the MVDR filter for acoustic echo suppressionabstractThis paper studies an echo suppression approach to reducing the undesired echoes that result from the acoustic coupling between a loudspeaker and a microphone in duplex voice communication. The approach consists of four basic steps. First, both the loudspeaker and microphone signals are partitioned into small overlapping frames. Second, each frame is transformed into the short-time Fourier transform (STFT) domain. Third, a minimum variance distortionless response (MVDR) filter is designed in each subband by explicitly using the interframe signal correlation. This MVDR filter is then used to estimate the echo signal and the obtained estimate is subsequently subtracted from the microphone signal. Finally, the time-domain processed signal is constructed using the overlap-add technique with the inverse STFT. Experiments are performed and the results demonstrate that this proposed method can achieve significant amount of echo suppression in practical room environments. Jacob Benesty, Jingdong Chen, Karim Helwani, Herbert Buchner |
ICASSP | 4 |
| 2013 | A Single-Channel MVDR Filter for Acoustic Echo SuppressionabstractAcoustic echo suppression techniques for full-duplex communication in low-complexity systems are commonly known to introduce distortion to the desired signal (i.e., near-end speech). Moreover, most traditional echo control techniques typically require accurately detecting the contribution of the near-end speaker to the microphone signal (“double talk”). In this letter, we propose a novel approach to acoustic echo suppression, which aims at extracting the near-end signal using a constraint for minimizing the distortion, and without requiring a double-talk detector. Karim Helwani, Herbert Buchner, Jacob Benesty, Jingdong Chen |
IEEE Signal Process. Lett. | 1 |
| 2011 | Spatio-temporal signal preprocessing for multichannel acoustic echo cancellationabstractHands-free full-duplex communication systems require acoustic echo cancelers to reduce echoes. Multichannel sound reproduction enhances realism in virtual reality and multimedia communication systems. However, in the case of multichannel systems the acoustic echo cancellation problem is challenging because of the non-uniqueness of the solution in the least squares sense. There fore, preprocessing techniques are required. Known preprocessing approaches act in the temporal domain. In this paper we propose an acoustic echo cancellation preprocessing stage which bases on the spatial diversity offered by massive multichannel reproduction systems. Karim Helwani, Sascha Spors, Herbert Buchner |
ICASSP | 1 |
| 2010 | Source-domain adaptive filtering for MIMO systems with application to acoustic echo cancellationabstractCombining the channels of a multiple input/multiple output (MIMO) system into suitably chosen modes by a domain transformation offers great improvements of adaptive filtering algorithms. In this paper we present an algorithm for adaptive MIMO filtering, called source-domain adaptive filtering (SDAF), with application to multichannel acoustic echo cancellation operating in an optimally adjusted transform domain without requiring a-priori knowledge about the system. Experimental results show a significant performance improvement compared to fixed transformation bases. Karim Helwani, Herbert Buchner, Sascha Spors |
ICASSP | 1 |