EDBT 2026 Demo / reviewers in the wild / expert
Andy W. H. Khong
dblp:66/1858 · also Andy Wai Hoong Khong
· DBLP profile ↗
92ranked-venue papers
11as first author
32since 2021 · last 2026
0000-0002-0708-4791ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 9 first-author · 16 since 2021Artificial intelligence and machine learning · 25 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoachLah: A Singlish-English Parallel Corpus of Health Coaching Conversations with Behavior Goal Annotations
Iva Bojic, Mathieu Ravaut, Stephanie Hilary Xinyi Ma, Doreen Tan, Andy Hau Yan Ho, Andy W. H. Khong |
LREC | 6 |
| 2026 | Singlish to English Translation with Precision: A Dataset and Language Detection-Driven Masked Modeling for Singlish to English Translation
Sujit Kumar, Gerome Kusuma Ang, Stephanie Hilary Xinyi Ma, Andy Hau Yan Ho, Andy W. H. Khong |
LREC | 5 |
| 2026 | Joint Enhancement and Bandwidth Extension for Radar Through-Barrier Speech AcquisitionabstractJoint speech enhancement and bandwidth extension are challenging for through-barrier radar vibrometry. Deep learning approaches achieve reasonable success but lack temporal context and bandwidth extension capability for significantly band-limited speech. In this work, we propose a joint speech enhancement and bandwidth extension module consisting of a newly formulated features fusion transformer that leverages temporal context to facilitate the consistency of spectral features of the early extended subbands by large dilation deconvolution operations with the enhancement operation. The proposed EBEnet was evaluated under a through-barrier setup and demonstrated higher speech intelligibility than encoder-decoder-based CNNs. Zhi-Wei Tan, V. G. Reju, Ritesh Chandra Tewari, Ruotong Ding, Andy W. H. Khong |
IEEE Signal Process. Lett. | 5 |
| 2025 | ViKIENet: Towards Efficient 3D Object Detection with Virtual Key Instance Enhanced NetworkabstractThe sparsity of point clouds and inadequacy of semantic information pose challenges to current LiDAR-only 3D object detection methods. Recent methods alleviate these challenges by converting RGB images into virtual points via depth completion to be fused with LiDAR points. Although these methods have shown outstanding results, they often introduce significant computation overhead due to the high density of virtual points and noise due to inaccurate depth completion. Besides, they do not thoroughly leverage semantic information from images. In this work, we propose the virtual key instance enhanced network (ViKIENet), a highly efficient and effective multi-modal feature fusion framework that fuses the features of virtual key instances (VKIs) and LiDAR points through multiple stages. Our contributions include three main components: semantic key instance selection (SKIS), virtual-instance-focused fusion (VIFF), and virtual-instance-to-real attention (VIRA). We also propose the extended version ViKIENet-R with VIFF-R which includes rotationally equivariant features. Experiment results show that ViKIENet and ViKIENet-R achieve significant improvements in detection performance on the KITTI, JRDB, and nuScenes datasets compared to existing works. On the KITTI dataset, ViKIENet and ViKIENet-R operate at 22.7 and 15.0 FPS, respectively. As of CVPR submission (Nov. 15th, 2024), ViKIENet ranks first on the car detection and orientation estimation leaderboard, while ViKIENet-R ranks second (compared with officially published papers) on the 3D car detection leaderboard. Zhuochen Yu, Bijie Qiu, Andy W. H. Khong |
CVPR | 3 |
| 2025 | Improving Course Recommendation Systems with Explainable AI: LLM-Based Frameworks and Evaluations
Qianru Lyu, Andy W. H. Khong |
EDM | 4 |
| 2025 | Contactless Vital Sign Monitoring for Multiple People Using a Millimeter-wave MIMO RadarabstractRadar technology offers much appeal for contactless vital sign monitoring. While most radar-based approaches achieve reasonable performance for single-person scenarios, they suffer from inaccurate vital sign estimates for multiple individuals, especially when the subjects occupy the same range bin. In addition, complex interferences due to environmental clutter and random body movements (RBMs) further exacerbate the issue of poor performance. We propose a resonance-based sparse separation (RBSS) algorithm for contactless vital sign measurement using millimeter wave (mmWave) multiple-input multiple-output (MIMO) radar. The proposed algorithm utilizes the resonance-based signal decomposition and can achieve reliable reconstruction of cardiopulmonary signals and precise estimates of respiration rate (RR) and heart rate (HR) for multiple individuals located within the same range bin. Experiment results demonstrate the efficacy of the proposed method, even under conditions of heavy clutter and moderate RBMs. Yuan Liu 0007, Xuemei Fu, Ritesh Chandra Tewari, Xingze Wang, Andy W. H. Khong |
ICASSP | 5 |
| 2025 | Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal RecommendationabstractIncorporating multi-modal features as side information has recently become a trend in recommender systems. To elucidate user-item preferences, recent studies focus on fusing modalities via concatenation, element-wise sum, or attention mechanisms. Despite having notable success, existing approaches do not account for the modality-specific noise encapsulated within each modality. As a result, direct fusion of modalities will lead to the amplification of cross-modality noise. Moreover, the variation of noise that is unique within each modality results in noise alleviation and fusion being more challenging. In this work, we propose a new Spectrum-based Modality Representation (SMORE) fusion graph recommender that aims to capture both uni-modal and fusion preferences while simultaneously suppressing modality noise. Specifically, SMORE projects the multi-modal features into the frequency domain and leverages the spectral space for fusion. To reduce dynamic contamination that is unique to each modality, we introduce a filter to attenuate and suppress the modality noise adaptively while capturing the universal modality patterns effectively. Furthermore, we explore the item latent structures by designing a new multi-modal graph learning module to capture associative semantic correlations and universal fusion patterns among similar items. Finally, we formulate a new modality-aware preference module, which infuses behavioral features and balances the uni- and multi-modal features for precise preference modeling. This empowers SMORE with the ability to infer both user modality-specific and fusion preferences more accurately. Experiments on three real-world datasets show the efficacy of our proposed model. The source code for this work has been made publicly available at https://github.com/kennethorq/SMORE. Rongqing Kenneth Ong, Andy W. H. Khong |
WSDM | 2 |
| 2025 | Collaborative Multiobjective Decisions for Cyber-Physical Production Systems Under Time-Varying DemandsabstractThe advent of cyber-physical production systems (CPPSs) has greatly improved production responsiveness. However, effective control and decision-making in CPPSs remain challenging due to the dynamic nature of both internal operations and external environments. We present a multiobjective optimization approach for managing operation, maintenance, and support decisions in CPPSs under time-varying demands. Specifically, a decision-making framework is developed to enable collaborative control, incorporating reliability-based risk assessment and multiobjective optimization techniques. To facilitate continuous decision-making in response to uncertainties, a biobjective optimization model is formulated using a receding horizon control architecture, addressing conflicting objectives simultaneously. An enhanced multiobjective pigeon-inspired optimization algorithm is proposed to generate Pareto-optimal solutions by co-minimizing the production risks and costs. Experimental validations are carried out through both numerical simulations and real-world experiments on a subsea production system in the South China Sea, involving two support sites, six production sites, thirty-six machines, and 288 components. Meng Liu 0019, Qiang Feng 0003, Xingshuo Hai, Qianming Zhang, Changyun Wen, Andy W. H. Khong |
IEEE Trans. Cybern. | 6 |
| 2025 | Capability-Oriented Decision-Making in Multi-UAV Deployment and Task Allocation: A Hierarchical Game-Based FrameworkabstractHigh-level decision-making for multiple uncrewed aerial vehicles (multi-UAV) mission planning is crucial, especially with the rising demand for long-term services in geo-distributed environments. However, the interrelated issues of multi-UAV deployment and task allocation are often addressed separately. This article integrates these two problems and introduces a hierarchical framework for effective decision-making. This is achieved by proposing balanced capability (BC), a customized metric tailored for long-term multi-UAV missions with geographically dispersed targets. By considering the global objective and self-organized coordination, a joint optimization model is established from a game-theoretical perspective. Additionally, a novel tangent and cotangent search algorithm (TCSA) is proposed to steer cooperative players toward the global objective in the upper layer, while in the lower layer, a modified distributed task allocation algorithm (MDT2A) incentivizes each autonomous player to efficiently maximize their individual benefits. Simulations validate the effectiveness of the proposed method, with comparative results highlighting the superiority of the algorithms. Xingshuo Hai, Qiang Feng 0003, Weike Chen, Changyun Wen, Andy W. H. Khong |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2024 | Enhancing Code-Switching Speech Recognition With Interactive Language BiasesabstractLanguages usually switch within a multilingual speech signal, especially in a bilingual society. This phenomenon is referred to as code-switching (CS), making automatic speech recognition (ASR) challenging under a multilingual scenario. We propose to improve CS-ASR by biasing the hybrid CTC/attention ASR model with multi-level language information comprising frame-and token-level language posteriors. The interaction between various resolutions of language biases is subsequently explored in this work. We conducted experiments on datasets from the ASRU 2019 code-switching challenge. Compared to the baseline, the proposed interactive language biases (ILB) method achieves higher performance and ablation studies highlight the effects of different language biases and their interactions. In addition, the results presented indicate that language bias implicitly enhances internal language modeling, leading to performance degradation after employing an external language model. Hexin Liu, L. Paola García-Perera, Xiangyu Zhang 0005, Andy W. H. Khong, Sanjeev Khudanpur |
ICASSP | 4 |
| 2024 | Enhanced Student-graph Representation for At-risk Student DetectionabstractPredicting examination grades is essential to facilitate early interventions and to enhance student retention rates in an academic institution. We propose a predictive model based solely on historical academic performance made available before the beginning of each semester. The proposed model employs singular value decomposition to distill the underlying student-course graph structure, resulting in a student representation vector that holistically captures a student’s academic ability across courses in relation to their cohort. This representation vector is then fused with the student’s historical academic records for grade prediction. Data for training the proposed prediction model was sourced from approximately five thousand Electrical and Electronic Engineering students across seventeen core courses, including Circuit Analysis and Analog Electronics taught in the sophomore year. Students identified as at risk of failing a course at the beginning of each semester may be offered targeted (academic) support such as peer tutoring programs. Andy W. H. Khong, Fun Siong Lim |
ISCAS | 2 |
| 2024 | Quad-Faceted Feature-Based Graph Network for Domain-Agnostic Text Classification to Enhance Learning EffectivenessabstractEnhancing learning effectiveness requires one to define suitable learning outcomes and align assessment constructs with these outcomes. We present a quad-faceted feature-based graph network to classify assessment texts into domain-agnostic class labels more accurately. The proposed model incorporates four complementary graphs (syntactic, semantic, sequential, and topical) with observable and latent node types and unique edge weight computations that are dependent on node properties to extract unique features from a given text. The purpose of incorporating syntactic information is to consider the dependency parsing between word nodes, while the semantic information is to provide the algorithm with contextual similarity between phrase nodes that are more effective than words in encapsulating the meaning of a text. The sequential graph is applied to regular expression nodes that contribute to a domain-agnostic class label, while the topical graph identifies topics that are convergent to each other based on their distributions. As opposed to existing techniques that construct graphs solely based on word nodes, the proposed model exploits the benefits of term weighting, nested phrases, regular expressions, and topic modeling to develop a diverse heterogeneous architecture for text classification. We evaluate the classification performance on questions with different class labels such as cognitive complexities, reasoning capabilities, and question types, as well as longer documents. Experiment results show that the proposed model outperforms in terms of macroaverage F1 score when compared with existing deep learning techniques. We also demonstrate the application of the classification model to understand learners’ attitudes via an empirical study in a workplace-learning environment. S. Supraja, Andy W. H. Khong |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2023 | Quad-Tier Entity Fusion Contrastive Representation Learning for Knowledge Aware Recommendation SystemabstractKnowledge graph (KG) has recently emerged as a powerful source of auxiliary information in the realm of knowledge-aware recommendation (KGR) systems. However, due to the lack of supervision signals caused by the sparse nature of user-item interactions, existing supervised graph neural network (GNN) models suffer from performance degradation. Moreover, the over-smoothing issue further limits the number of GNN layers or hops required to propagate messages - these models ignore the non-local information concealed deep within the knowledge graph. We propose the Quad-Tier Entity Fusion Contrastive Representation Learning (QTEF-CRL) knowledge-aware framework to achieve learning of deep user preferences from four perspectives: the collaborative, semantic, preference, and structural view. Unlike existing methods, the proposed tri-local and single-global quad-tier architecture exploits the knowledge graph holistically to achieve effective self-supervised representation learning. The newly-introduced preference view constructed from the collaborative knowledge graph (CKG) comprises a preference graph and preference-guided GNN that are specifically designed to capture non-local information explicitly. Experiments conducted on three datasets highlight the efficacy of our proposed model. Rongqing Kenneth Ong, Andy W. H. Khong |
CIKM | 3 |
| 2023 | Reducing Language Confusion for Code-Switching Speech Recognition with Token-Level Language DiarizationabstractCode-switching (CS) occurs when languages switch within a speech signal and leads to language confusion for automatic speech recognition (ASR). We address the problem of language confusion for improving CS-ASR from two perspectives: incorporating and disentangling language information. We incorporate language information within the CS-ASR model by dynamically biasing the model with token-level language posteriors corresponding to outputs of a sequence-to-sequence auxiliary language diarization (LD) module. In contrast, the disentangling process reduces the difference between languages via adversarial training so as to normalize two languages. We conduct experiments on the SEAME dataset. Compared to the baseline model, both the joint optimization with LD and the language posterior bias achieve performance improvement. Comparison of the proposed methods indicates that incorporating language information is more effective than disentangling for reducing language confusion in CS speech. Hexin Liu, Haihua Xu 0001, L. Paola García-Perera, Andy W. H. Khong, Sanjeev Khudanpur |
ICASSP | 4 |
| 2023 | Improving Performance of Real-Time Full-Band Blind Packet-Loss Concealment with Predictive NetworkabstractPacket loss concealment (PLC) is a tool for enhancing speech degradation caused by poor network conditions or underflow/overflow in audio processing pipelines. We propose a real-time recurrent method that leverages previous outputs to mitigate artefact of lost packets without the prior knowledge of loss mask. The proposed full-band recurrent network (FRN) model operates at 48 kHz, which is suitable for high-quality telecommunication applications. Experiment results highlight the superiority of FRN over an offline non-causal baseline and a top performer in a recent PLC challenge. Anh H. T. Nguyen, Andy W. H. Khong |
ICASSP | 3 |
| 2023 | MERLIon CCS Challenge: A English-Mandarin code-switching child-directed speech corpus for language identification and diarizationabstractTo enhance the reliability and robustness of language identification (LID) and language diarization (LD) systems for heterogeneous populations and scenarios, there is a need for speech processing models to be trained on datasets that feature diverse language registers and speech patterns. We present the MERLIon CCS challenge, featuring a first-of-its-kind Zoom video call dataset of parent-child shared book reading, of over 30 hours with over 300 recordings, annotated by multilingual transcribers using a high-fidelity linguistic transcription protocol. The audio corpus features spontaneous and in-the-wild English-Mandarin code-switching, child-directed speech in non-standard accents with diverse language-mixing patterns recorded in a variety of home environments. This report describes the corpus, as well as LID and LD results for our baseline and several systems submitted to the MERLIon CCS challenge using the corpus. Yi Han Victoria Chua, Hexin Liu, L. Paola García-Perera, Fei Ting Woon, Jinyi Wong, Xiangyu Zhang 0005, Sanjeev Khudanpur, Andy W. H. Khong, Justin Dauwels, Suzy J. Styles |
INTERSPEECH | 8 |
| 2023 | Investigating model performance in language identification: beyond simple error statisticsabstractLanguage development experts need tools that can automatically identify languages from fluent, conversational speech and provide reliable estimates of usage rates at the level of an individual recording. However, LID systems are typically evaluated on metrics such as equal error rate and balanced accuracy, applied at the level of an entire speech corpus. These overview metrics do not provide information about model performance at the level of individual speakers, recordings, or units of speech with different linguistic characteristics. Overview statistics may mask systematic errors in model performance for some subsets of the data, and consequently, have worse performance on data derived from some subsets of human speakers, creating a kind of algorithmic bias. Here, we investigate how well a number of LID systems perform on individual recordings and speech units with different linguistic properties in the MERLIon CCS Challenge featuring accented code-switched child-directed speech. Suzy J. Styles, Yi Han Victoria Chua, Fei Ting Woon, Hexin Liu, L. Paola García-Perera, Sanjeev Khudanpur, Andy W. H. Khong, Justin Dauwels |
INTERSPEECH | 7 |
| 2023 | Gridless DOA Estimation Using Complex-Valued Convolutional Neural Network With Phasor NormalizationabstractWe propose a complex LeDIM-net (C-LeDIM-net) convolutional neural network (CNN) that employs a newly-formulated complex phasor normalization for gridless direction-of-arrival (DOA) estimation. Unlike existing deep learning (DL) approaches, C-LeDIM-net extracts explicit phase information in its intermediate complex-valued feature maps to estimate unknown source DOAs. Given its explicit phase representation, the proposed complex phasor normalization leverages the phase-to-sensor relationship of the feature maps which, as a consequence, improves the robustness of C-LeDIM-net to array imperfections when operating with limited number of snapshots. Simulation results show that the proposed method outperforms the existing methods, including the subspace-based and DL-based methods. Zhi-Wei Tan, Yuan Liu 0007, Andy W. H. Khong, Anh H. T. Nguyen |
IEEE Signal Process. Lett. | 3 |
| 2022 | Grade Prediction via Prior Grades and Text Mining on Course Descriptions: Course Outlines and Intended Learning Outcomes
S. Supraja, Andy W. H. Khong |
EDM | 4 |
| 2022 | Toward Better Grade Prediction via A2GP - An Academic Achievement Inspired Predictive Model
S. Supraja, Andy W. H. Khong |
EDM | 3 |
| 2022 | Factors Impacting Students' Creativity-related Self-efficacy in an Undergraduate Makerspace-based CourseabstractThe need to cultivate creativity in engineering education calls for opportunities for students to exercise freedom in proposing and pursuing projects aligned with their interests. This paper presents insights into an undergraduate makerspace-based course in terms of factors affecting students’ creativity-related self-efficacy. We conducted a survey on students who come from an engineering and science background to gain their opinions about the impact of this course on enhancing their creativity. To establish if there is a significant difference in the students’ creativity, we performed the non-parametric Wilcoxon signed-rank test comparing the first and second survey, with results showing that there is a statistically significant increase in students’ creativity-related self-efficacy. There was a general increase for all the items, especially students’ perceptions toward the relevance of the course, the conduciveness of the learning environment, opportunities to make and learn from mistakes, and the resourcefulness of their team. Results obtained via quantitative statistical analysis was backed up by qualitative analysis that employed text mining techniques such as automatic key phrase extraction and sentiment analysis on the open-ended responses and the reason(s) for the Likert-scale answer choice. In addition, we used the Spearman’s rho to report correlations between Likert-scale items and determine the variables that are significantly and positively correlated with the creativity-related self-efficacy construct. A multivariate regression model was then constructed to observe the extent to which each highly correlated variable impacts creativity-related self-efficacy; of which, a sense of relevance appears to have the largest effect. Through gaining insights into the factors that may impact students’ creativity-related self-efficacy, this study contributed to a deeper understanding on how this important attribute could be developed through a makerspace-based university course. S. Supraja, Fun Siong Lim, Sophia Tan, Shen Yong Ho, Beng Koon Ng, Andy W. H. Khong |
EDUCON | 6 |
| 2022 | Freshmen Orientation Program Using Minecraft: Designed by Students for Students during the Covid-19 PandemicabstractThis Innovative Practice Full Paper presents experiences in designing a student-led virtual freshmen orientation program that uses a Minecraft environment. We describe the planning process, roles of the organizing committee members, and how the game was constructed for participants to learn and interact with one another. The student organizers not only created a virtual environment that scales the college map where more than a hundred freshmen (participants) could have an immersive experience of the campus, but also ensured the branding and marketing, logistics, and safety/well-being aspects of the event. In this paper, we present students’ experience of this program from both the designers’ as well as the participants’ perspectives. We conducted surveys with the organizing committee members and interviewed the participants to gain insights on their perception of this event. Our analysis showed that student organizers had the autonomy to brainstorm, suggest creative ideas, develop novel games, and procure materials. They also felt that they developed authentic programming and leadership skills. On the other hand, participants felt engaged as the event was well-organized, had clear delivery of information, introduced them to new technology, made them more familiar with the campus, provided a conducive environment to hone their soft skills such as communication and teamwork even before they officially enrolled as undergraduate students in an engineering program, and helped them establish social networks to support them throughout their undergraduate education journey. S. Supraja, Sophia Tan, Fun Siong Lim, Beng Koon Ng, Shen Yong Ho, Andy W. H. Khong |
FIE | 6 |
| 2022 | Joint Source Localization and Association Through Overcomplete Representation Under Multipath Propagation EnvironmentabstractThis work addresses the source localization and association problem in a multipath propagation environment. By focusing on the limitation of the prior information in practical applications, we propose a target localization and association method based on iterative optimization with semi-unitary constraint and eigen-decomposition techniques. In contrast to the previous works, the proposed method can localize spatial sources and associate the incident paths to each source without prior knowledge pertaining to the propagation environment. Moreover, the proposed approach can be applied to an arbitrary array geometry without reducing the effective array aperture. Both simulations and real data experiments validate the effectiveness and robustness of the proposed method. Yuan Liu 0007, Zhi-Wei Tan, Andy W. H. Khong, Hongwei Liu 0001 |
ICASSP | 3 |
| 2022 | Tunet: A Block-Online Bandwidth Extension Model Based On Transformers And Self-Supervised PretrainingabstractWe introduce a block-online variant of the temporal feature-wise linear modulation (TFiLM) model to achieve bandwidth extension. The proposed architecture simplifies the UNet backbone of the TFiLM to reduce inference time and employs an efficient transformer at the bottleneck to alleviate performance degradation. We also utilize self-supervised pretraining and data augmentation to enhance the quality of bandwidth extended signals and reduce the sensitivity with respect to downsampling methods. Experiment results on the VCTK dataset show that the proposed method outperforms several recent baselines in both intrusive and non-intrusive metrics. Pretraining and filter augmentation also help stabilize and enhance the overall performance. Anh H. T. Nguyen, Andy W. H. Khong |
ICASSP | 3 |
| 2022 | Multichannel Noise Reduction Using Dilated Multichannel U-Net and Pre-Trained Single-Channel NetworkabstractPre-trained single-channel neural networks have become more prevalent for noise reduction in recent years. However, unlike their multichannel counterparts, these monoaural approaches do not exploit spatial information during the optimization process. Furthermore, while multichannel neural networks exploit spatial information, they are optimized for a specific microphone array configuration; extensive data collection and training are required if a new array configuration is deployed. We propose a transfer learning approach that leverages existing pre-trained single-channel neural networks for the optimization of multichannel neural networks. Simulation results on the CHiME-3 dataset show that the proposed method outperforms the state-of-the-art multichannel neural network and neural beamformer. Zhi-Wei Tan, Anh H. T. Nguyen, Yuan Liu 0007, Andy W. H. Khong |
ICASSP | 4 |
| 2022 | PHO-LID: A Unified Model Incorporating Acoustic-Phonetic and Phonotactic Information for Language IdentificationabstractWe propose a novel model to hierarchically incorporate phoneme and phonotactic information for language identification (LID) without requiring phoneme annotations for training.In this model, named PHO-LID, a self-supervised phoneme segmentation task and a LID task share a convolutional neural network (CNN) module, which encodes both language identity and sequential phonemic information in the input speech to generate an intermediate sequence of "phonotactic" embeddings.These embeddings are then fed into transformer encoder layers for utterance-level LID.We call this architecture CNN-Trans.We evaluate it on AP17-OLR data and the MLS14 set of NIST LRE 2017, and show that the PHO-LID model with multitask optimization exhibits the highest LID performance among all models, achieving over 40% relative improvement in terms of average cost on AP17-OLR data compared to a CNN-Trans model optimized only for LID.The visualized confusion matrices imply that our proposed method achieves higher performance on languages of the same cluster in NIST LRE 2017 data than the CNN-Trans model.A comparison between predicted phoneme boundaries and corresponding audio spectrograms illustrates the leveraging of phoneme information for LID. Hexin Liu, L. Paola García-Perera, Andy W. H. Khong, Suzy J. Styles, Sanjeev Khudanpur |
INTERSPEECH | 3 |
| 2022 | Iterative Implementation Method for Robust Target Localization in a Mixed Interference EnvironmentabstractFor the problem of target localization under the multipath propagation environment, the existing methods are mainly restricted to the limited prior information of complex reflections, especially when the target is embedded in a mixed interference environment. They may suffer from performance degradation due to the shortage of target classification ability. To address this problem, we propose a target localization method based on iterative implementation with semiunitary constraint and eigen-decomposition technique, where a practical propagation scenario based on the spherical Earth model is considered. Compared to the previous works, the proposed method can automatically distinguish a real target from the mixed interference environment with improved localization accuracy. Neither additional decorrelation preprocessing nor prior information of the dynamic scenario is required. Both simulations and real data experiments validate the effectiveness and robustness of the proposed method. Yuan Liu 0007, Xiang-Gen Xia 0001, Hongwei Liu 0001, Anh H. T. Nguyen, Andy W. H. Khong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Hybrid Method Based on Random Convolution Nodes for Short-Term Wind Speed ForecastingabstractDespite having a plethora of works, wind speed time-series forecasting capabilities are prone to errors due to their intermittent and nonstationary nature as well as the limited generalization capabilities of forecasting methods for non-Gaussian distributed data. In this article, a hybrid method that consists of elastic variational mode decomposition (eVMD) andforecastingrandom convolution nodes (fRCN) is proposed to forecast the Gaussian heteroscedastic wind speed time-series. The proposed eVMD algorithm gauges the nonstationary characteristics (complexity) of the wind speed signal and thereafter decomposes the signal into its intrinsic components (ICs) accordingly. The fRCN method rely on local receptive fields to extract features that contribute to the local variations and the global trend in each IC. These features are subsequently learned using extreme learning machines theories. An ensemble unit is employed to learn appropriate weightages for each forecasted IC before yielding the final forecasting values. Suitability of the proposed hybrid method for wind speed forecasting is evaluated via an actual wind speed dataset and comparing against various existing hybrid methods. Yubo Wang 0001, Andy W. H. Khong |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | An Adaptive Non-Linear Process for Under-Determined Virtual Microphone BeamformingabstractVirtual microphone beamforming techniques are attractive for devices limited by space constraints. These techniques synthesize virtual microphone signals via interpolation algorithms. We propose to extend existing virtual microphone signal interpolation by employing an adaptive non-linear (ANL) process for acoustic beamforming. The proposed ANL based interpolation utilizes a target-presence probability criteria to determine the degree of non-linearity. The beamformer output is then derived using a combination between interpolations during target inactive zones and target active zones. Such combination offers a trade-off between reducing interference and target signal distortion. We apply the proposed ANL-based interpolator to the maximum signal-to-noise ratio (MSNR) beamformer and compare its performance against conventional beamforming and virtual microphone based beamforming methods in under-determined situations. Mehdi Bekrani, Anh H. T. Nguyen, Andy W. H. Khong |
ICASSP | 3 |
| 2021 | Directional Sparse Filtering Using Weighted Lehmer Mean for Blind Separation of Unbalanced Speech MixturesabstractIn blind source separation of speech signals, the inherent imbalance in the source spectrum poses a challenge for methods that rely on single-source dominance for the estimation of the mixing matrix. We propose an algorithm based on the directional sparse filtering (DSF) framework that utilizes the Lehmer mean with learnable weights to adaptively account for source imbalance. Performance evaluation in multiple real acoustic environments show improvements in source separation compared to the baseline methods. Karn Watcharasupat, Anh H. T. Nguyen, Ching-Hui Ooi, Andy W. H. Khong |
ICASSP | 4 |
| 2021 | End-to-End Language Diarization for Bilingual Code-Switching SpeechabstractWe propose two end-to-end neural configurations for language diarization on bilingual code-switching speech. The first, a BLSTM-E2E architecture, includes a set of stacked bidirectional LSTMs to compute embeddings and incorporates the deep clustering loss to enforce grouping of languages belonging to the same class. The second, an XSA-E2E architecture, is based on an x-vector model followed by a self-attention encoder. The former encodes frame-level features into segmentlevel embeddings while the latter considers all those embeddings to generate a sequence of segment-level language labels. We evaluated the proposed methods on the dataset obtained from the shared task B in WSTCSMC 2020 and our handcrafted simulated data from the SEAME dataset. Experimental results show that our proposed XSA-E2E architecture achieved a relative improvement of 12.1% in equal error rate and a 7.4% relative improvement on accuracy compared with the baseline algorithm in the WSTCSMC 2020 dataset. Our proposed XSA-E2E architecture achieved an accuracy of 89.84% with a baseline of 85.60% on the simulated data derived from the SEAME dataset. Hexin Liu, L. Paola García-Perera, Justin Dauwels, Andy W. H. Khong, Sanjeev Khudanpur, Suzy J. Styles |
Interspeech | 5 |
| 2021 | Regularized Phrase-Based Topic Model for Automatic Question Classification With Domain-Agnostic Class LabelsabstractClassification of questions according to domain-agnostic class labels relies on a suitable feature extraction process. We propose the use of phrases that are more effective than words to represent questions. The proposed phrase-based topic modeling technique employs asymmetric priors that are scaled with a new C-value for nested regular expressions. In addition, to suppress high-frequency words in phrases, we deploy term weightages computed using the modified distinguishing feature selector. The proposed approach also incorporates a new topic regularization mechanism to facilitate efficient mapping of questions to class labels. We validate the performance of the above approach via four datasets across different domain-agnostic class labels comprising question types, reasoning capabilities, and cognitive complexities. Results obtained highlight that the proposed technique outperforms existing methods in terms of macro-average F1 score. S. Supraja, Andy W. H. Khong |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Controllable Question Generation via Sequence-to-Sequence Neural Model with Auxiliary InformationabstractAutomatic question generation (QG) has found applications in the education sector and to enhance human-machine interactions in chatbots. Existing neural QG models can be categorized into answer-unaware and answer-aware models. One of the main challenges faced by existing neural QG models is the degradation in performance due to the issue of one-to-many mapping, where, given a passage, both answer (query interest/question intent) and auxiliary information (context information present in the question) can result in different questions being generated. We propose a controllable question generation model (CQG) that employs an attentive sequence-to-sequence (seq2seq) based generative model with copying mechanism. The proposed CQG also incorporates query interest and auxiliary information as controllers to address the one-to-many mapping problem in QG. Two variants of embedding strategies are designed for CQG to achieve good performance. To verify its performance, an automatic labeling scheme for harvesting auxiliary information is first developed. A QG dataset is also annotated with auxiliary information from a reading comprehension dataset. Performance evaluation shows that the proposed model not only outperforms existing QG models, it also has the potential to generate multiple questions that are relevant given a single passage. Andy W. H. Khong |
IJCNN | 3 |
| 2020 | Distilling Essence of a Question: A Hierarchical Architecture for Question Quality in Community Question Answering SitesabstractCommunity question answering (CQA) sites have grown to be useful platforms where users search for highly specific information to resolve a problem. However, the significant increase in the number of user-generated content with high variance in quality on these sites not only presents challenges for user navigation but also outgrow the community's peer reviewing capacity. This necessitates ways to automatically assess the quality of new questions so as to maintain quality of content served to its users. While existing methods commonly employ social network indicators as features, our model aims to avoid social influence biases arising from these indicators by predicting the quality from semantic evaluation of the question text. Formulation of the proposed model is non-trivial as it requires the extraction of meaningful features from the noisy question text at different granularities while filtering redundant information. In this work, a neural architecture is proposed to address this problem by aggregating the textual features extracted at word- and sentence-level in a hierarchical manner. In addition, a unique attention mechanism that focuses on sentence segments for interpreting a question is developed. This new mechanism employs the global topical information from common problem contexts. The proposed approach is verified on the Stack Overflow question dataset and is shown to outperform existing neural models. Mun Kit Ho, Andy W. H. Khong |
IJCNN | 3 |
| 2019 | A Method Based on L-bfgs to Solve Constrained Complex-valued IcaabstractComplex-valued independent component analysis (ICA) is a celebrated method in blind separation of complex-valued signals. In this paper, we propose to transform the constrained optimization problems of complex-valued ICA into unconstrained optimization problems which can be solved by limited-memory Broyden-Fletcher-Goldfarb-Shanno update (L-BFGS). As opposed to previous approaches, the proposed method does not apply any restriction on the Hessian matrix of ICA cost function. It can separate mixed sub-Gaussian, super-Gaussian, circular, and non-circular sources. Simulations show promising results. Anh H. T. Nguyen, V. G. Reju, Andy W. H. Khong |
ICASSP | 3 |
| 2018 | Online Education Evaluation for Signal Processing Course Through Student Learning PathwaysabstractImpact of online learning sequences to forecast course outcomes for an undergraduate digital signal processing (DSP) course is studied in this work. A multi-modal learning schema based on deep-learning techniques with learning sequences, psychometric measures, and personality traits as input features is developed in this work. The aim is to identify any underlying patterns in the learning sequences and subsequently forecast the learning outcomes. Experiments are conducted on the data acquired for the DSP course taught over 13 teaching weeks to underpin the forecasting efficacy of various deep-learning models. Results showed that the proposed multi-modal schema yields better forecasting performance compared to existing frequency-based methods in existing literature. It is further observed that the psychometric measures incorporated in the proposed multimodal schema enhance the ability of distinguishing nuances in the input sequences when the forecasting task is highly dependent on human behavior. Kelvin H. R. Ng, Andy W. H. Khong |
ICASSP | 3 |
| 2018 | Automatically Linking Digital Signal Processing Assessment Questions to Key Engineering Learning OutcomesabstractTo deliver on the potential outcome-based teaching and learning holds for engineering education, it is important for engineering courses to provide students with different types of deliberate practice opportunities that align to the program's learning outcomes. Working from these requirements, we increased the design and measurement intentionality of a digital signal processing (DSP) course. To align the course's learning outcomes more constructively with its assessment measures, we automated the process of classifying DSP questions according to learning outcomes by introducing a model that integrates topic modeling and machine learning. In this work, we explored the effect of pre-processing procedures in terms of stopword selection and word co-occurrence redundancy issue in question classification inferences. In this work, we proposed a customized variant of the Word Network Topic Model, q-WNTM, which is able to use its pre-classified DSP questions to reliably classify new questions according to the course's learning outcomes. S. Supraja, Kevin Hartman, Andy W. H. Khong |
ICASSP | 4 |
| 2018 | Multisource DOA Estimation in a Reverberant Environment Using a Single Acoustic Vector SensorabstractWe address the problem of direction-of-arrival (DOA) estimation for multiple speech sources in an enclosed environment using a single acoustic vector sensor. The challenges in such scenario include reverberation and overlapping of the source signals. In this work, we exploit low-reverberant-single-source (LRSS) points in the time-frequency (TF) domain, where a particular source is dominant with high signal-to-reverberation ratio. Unlike conventional algorithms having limitation that such potential points need to be detected at “TF-zone” level, the proposed algorithm performs LRSS detection at “TF-point” level. Therefore, for the proposed algorithm, the potential LRSS points need not be neighbors of each other within a TF zone to be detected, resulting an increased number of detected LRSS points. The detected LRSS points are further screened by an outlier removal step such that only reliable LRSS points will be used for DOA estimation. Simulations and experiments were conducted to demonstrate the effectiveness of the proposed algorithm in multisource reverberant environments. Kai Wu 0005, V. G. Reju, Andy W. H. Khong |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Impact Localization on Rigid Surfaces Using Hermitian Angle Distribution for Human-Computer Interface ApplicationsabstractWe propose an algorithm to localize impacts on rigid surfaces using induced vibration signals. This allows for the conversion of daily objects, such as tabletops and glass panels, into human–computer touch interfaces using low-cost piezoelectric sensors. Impact localization is achieved by estimating the time-of-arrivals and subsequently time-difference-of-arrivals of the sensor-received signals. Time-of-arrival estimation is highly challenging with increasing source–sensor distance due to the occurrence of a gradual noise-to-signal transition at the sensor output. We address this problem by first converting the signal into Hermitian angle distributions. The time-varying probability contributions of the background noise and vibration signal in each of the distributions are subsequently monitored to identify the instant when the signal begins to dominate the noise, signifying the signals arrival. The proposed framework also allows simultaneous time-of-arrival estimation across all the sensors to minimize errors in the resultant time-difference-of-arrival estimates. Experimental results show that the proposed algorithm outperforms existing techniques for source localization on solid surfaces of different materials. Nguyen Quang Hanh, V. G. Reju, Andy W. H. Khong |
IEEE Trans. Multim. | 3 |
| 2017 | Toward the Automatic Labeling of Course Questions for Ensuring their Alignment with Learning Outcomes
S. Supraja, Kevin Hartman, Andy W. H. Khong |
EDM | 4 |
| 2017 | On TOA estimation of vibration signals for localizing impacts on solid surfacesabstractWe propose a TDOA-based algorithm for source localization on rigid surfaces. This allows the conversion of readily available large surfaces into touch interfaces using surface-mounted vibration sensors. To achieve this, we characterize the arrival of each sensor-received signal by the arrival times of its frequency components. To estimate the arrival time of each frequency component, we first model each component as a harmonic random process. A kurtosis sequence, which exhibits a sharp rising edge when the signal begins to deviate from the background noise, can then be obtained for each component. By accurately estimating the starting point of the rising edge, our algorithm can avoid the uncertainty due to gradual noise-to-signal transition. Experiment results show that the proposed algorithm achieves better localization performance than existing techniques on large surfaces. Nguyen Quang Hanh, V. G. Reju, Andy W. H. Khong |
ICASSP | 3 |
| 2017 | Learning complex-valued latent filters with absolute cosine similarityabstractWe propose a new sparse coding technique based on the power mean of phase-invariant cosine distances. Our approach is a generalization of sparse filtering and K-hyperlines clustering. It offers a better sparsity enforcer than the L1/L2norm ratio that is typically used in sparse filtering. At the same time, the proposed approach scales better than the clustering counterparts for high-dimensional input. Our algorithm fully exploits the prior information obtained by preprocessing the observed data with whitening via an efficient row-wise decoupling scheme. In our simulating experiments, the algorithm produces better estimates than previous approaches do. It yields better separation of live recorded speech mixtures as well. Anh H. T. Nguyen, V. G. Reju, Andy W. H. Khong, Ing Yann Soon |
ICASSP | 3 |
| 2017 | Swarm Intelligence Based Particle Filter for Alternating Talker Localization and Tracking Using Microphone ArraysabstractWe address the problem of localizing and tracking alternating (moving or stationary) talkers using microphone arrays in a room environment. One of the main challenges is the frequent (and possibly abrupt) change of talker positions, which requires the algorithm to capture the active talker rapidly. In addition, the presence of interference, background noise, and room reverberation degrades the tracking performance. We propose a new algorithm that jointly exploits the advantages of the particle filter (PF) and particle swarm intelligence. The PF is used as a general tracking framework, which incorporates a proposed alternating source-dynamic model for recursive estimation of talker position. Unlike the conventional PF, where particles operate independently in the particle sampling stage, the use of swarm intelligence allows particles to interact with each other, thereby improving convergence toward the active talker location. In addition, the memory mechanism in swarm intelligence allows particles to remain at their previous best-fit state estimate when signals are corrupted by interference, noise, and/or reverberation. Simulations and experiments were conducted to demonstrate the effectiveness of the proposed algorithm. Kai Wu 0005, V. G. Reju, Andy W. H. Khong, Shu Ting Goh |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Modelling the way: Using action sequence archetypes to differentiate learning pathways from learning outcomes
Kelvin H. R. Ng, Kevin Hartman, Andy W. H. Khong |
EDM | 4 |
| 2016 | Source localization on solids utilizing logistic modeling of energy transition in vibration signalsabstractWe propose a new algorithm for source localization on rigid surfaces, which allows one to convert daily objects into human-computer touch interfaces using surface-mounted vibration sensors. This is achieved via estimating the time-difference-of-arrivals (TDOA) of the signals across the sensors. In this work, we employ a smooth parametrized function to model the gradual noise-to-signal energy transition at each sensor. Specifically, the noise-to-signal transition is modeled by a four-parameter logistic function. The TDOA is then estimated as the difference in time shifts of the functions fitted to the sensor data. Experiment results show that the proposed algorithm significantly outperforms existing techniques which adopt the abrupt change model for time-of-arrival estimation. Nguyen Quang Hanh, V. G. Reju, Andy W. H. Khong |
ICASSP | 3 |
| 2015 | Direction-of-arrival estimation of speech sources under aliasing conditionsabstractDue to practical considerations the microphone spacing is increased to achieve improved resolution by violating the spatial Nyquist criterion. Accompanied aliasing components adversely affect the identifiability of the source direction peaks. We investigate the effect of aliasing on the spatial spectrum of the steered minimum variance distortionless response (STMV) method and propose a novel multi-stage scheme assisted by subband decomposition for suppressing aliasing components. The performance of the proposed technique, evaluated with simulations and recorded room responses, reflects the improvement in the identifiability of accurate source directions under aliasing conditions. Vinod V. Reddy, Andy W. H. Khong |
ICASSP | 2 |
| 2015 | Multi-source direction-of-arrival estimation in a reverberant environment using single acoustic vector sensorabstractWe address the problem of estimating direction-of-arrivals (DOAs) for multiple sound sources using a single acoustic vector sensor (AVS) in an enclosed room environment. It is well-known that multi-source DOA estimation in an enclosed environment is challenging due to room reverberation, environmental noise and overlapping of the source spectra. In this work, we propose a multi-source DOA estimation algorithm which exploits co-location of the sensor elements in AVS. We identify time-frequency (TF) zones of the received signals in which only one source is dominant with a high signal-to-reverberation ratio. DOA estimation is then achieved via the use of clustering of the Hermitian angle feature. Simulation results show that the proposed DOA estimation algorithm is robust to both reverberation and environmental noise. Kai Wu 0005, V. G. Reju, Andy W. H. Khong |
ICASSP | 3 |
| 2015 | Single-channel speech enhancement in a transient noise environment by exploiting speech harmonicityabstractThis paper focuses on the problem of single-channel noise reduction in a transient noise environment for speech enhancement application. A typical speech enhancement algorithm requires an estimate of the noise statistics. However, the problem of noise estimation is challenging when the statistics of the noise vary significantly with time. By exploiting the fact that for speech signal most of the energy is concentrated on the harmonic bands in voiced frames, we propose an algorithm for the estimation of speech presence probability in the time-frequency domain. The estimated speech presence probability is then used for noise estimation for speech enhancement application. Evaluations are conducted to compare the speech enhancement performance between the proposed algorithm and the existing algorithm for various types of transient noise. Kai Wu 0005, V. G. Reju, Andy W. H. Khong |
ICASSP | 3 |
| 2014 | Misalignment analysis and insights into the performance of clipped-input LMS with correlated Gaussian dataabstractThe three-level clipped input least-mean-square (CLMS) adaptive algorithm is known to have low complexity that is suitable for the identification of long finite impulse response of unknown systems. In this paper we analyze the performance of CLMS which allows one to gain insights into its convergence property and the amount of steady-state misalignment error for both time-invariant and time-varying systems perturbed by correlated Gaussian input. Arising from our analysis, we derive the optimal step-size for CLMS to achieve the minimum possible steady-state misalignment and compare its results with the performance of LMS adaptive algorithm. The accuracy of our derivations is evaluated with simulation results. Mehdi Bekrani, Andy W. H. Khong |
ICASSP | 2 |
| 2014 | Source localization on solids utilizing time-frequency analysis of parameterized warped signalsabstractWe propose a new approach for source localization on solids with applications to human-computer interface. We analyze the wave propagation of flexural modes of vibration, generated by an impact on a solid surface, to characterize the dispersive linear time-varying system having non-linear phase response. We show that a difference in dispersion between two signals propagating through solids can be mapped directly to the relative propagation distance if the signals are appropriately time-warped. We then exploit this important property for source localization by computing the similarity of the warped signals in the time-frequency domain. As the proposed source localization algorithm jointly estimates warping-based polynomial parameters and source location, the method does not require pre-calibration. Arun R. Kattukandy, V. G. Reju, Andy W. H. Khong |
ICASSP | 3 |
| 2014 | Adaptive multichannel equalization applied to room acoustics exploiting the sparsity of target responseabstractThe adaptive multiple-input/output inverse theorem (A-MINT) multichannel equalization algorithm was proposed to address the computational complexity of the MINT algorithm. However, similar to MINT, A-MINT also assumes an arbitrary modeling delay for the target response which, if inappropriately chosen, results in misconvergence or poor performance. We present a new adaptive algorithm that exploits the sparsity of a target response to estimate equalization filters. The proposed algorithm automatically selects a suitable modeling delay for the equalized response so as to achieve a minimum-norm solution. Rajan S. Rashobh, Andy W. H. Khong |
ICASSP | 2 |
| 2014 | A GMM Post-Filter for Residual Crosstalk Suppression in Blind Source SeparationabstractExisting algorithms employ the Wiener filter to suppress residual crosstalk in the outputs of blind source separation algorithms. We show that, in the context of BSS, the Wiener filter is optimal in the maximum likelihood (ML) sense only for normally-distributed signals. We then propose to model the distribution of speech signals using the Gaussian mixture model (GMM) and then derive a post-filter in the ML sense using the expectation-maximization algorithm. We show that the GMM introduces a probabilistic sample weight that is able to emphasize speech segments that are free of crosstalk components in the BSS output and this results in a better estimate of the post-filter. Simulation results show that the proposed post-filter achieves better crosstalk suppression than the Wiener filter for BSS. Benxu Liu, V. G. Reju, Andy W. H. Khong, Vinod V. Reddy |
IEEE Signal Process. Lett. | 3 |
| 2014 | Multichannel Equalization in the KLT and Frequency Domains With Application to Speech DereverberationabstractEqualization of acoustic channels usually involves inversion of acoustic impulse responses (AIRs), and generally employs multichannel techniques. In this paper, we propose three equalization algorithms, one in the Karhunen-Loève transform (KLT) domain and the other two in the frequency domain. Our proposed algorithm in the KLT domain provides a platform to achieve equalization in conjunction with denoising. Existing multiple-input/output inverse theorem (MINT)-based non-adaptive algorithms require the inversion of a matrix with dimension that is proportional to the AIR length, and is computationally expensive. To overcome this limitation, we propose the frequency-domain algorithm which is computationally very efficient and thus can be employed for the equalization of high-order AIRs in practical applications. In addition, the frequency-domain method is more robust to AIR estimation errors. To achieve further reduction in the complexity without significant performance degradation, we then propose a modified version of the frequency-domain algorithm. Rajan S. Rashobh, Andy W. H. Khong |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Unambiguous speech DOA estimation under spatial aliasing conditionsabstractWith the bandwidth of speech signals extending over several octaves, the spatial Nyquist criterion constrains the microphone array design. Violating this criterion by increasing microphone spacing in order to achieve high resolution introduces ambiguity in identifying the source directions due to the aliasing components. In this work, we investigate the effect of spatial aliasing on the direction-of-arrival (DOA) spectrum due to wideband sources. Noting that the extent of aliasing is frequency dependent, we propose a multi-stage scheme for speech DOA estimation following a subband decomposition. To observe the advantage of this scheme, we verify it with the steered minimum variance distortionless response (STMV) and approximate kernel density estimators. The performance is evaluated with simulations and recorded room impulse responses. Vinod V. Reddy, Andy W. H. Khong, Boon Poh Ng |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Underdetermined instantaneous blind source separation of sparse signals with temporal structure using the state-space modelabstractIn this work, we exploit, in addition to sparseness, the temporal structure of the source signals to address the problem of underdetermined blind source separation. To achieve good separation performance and reduction of artifacts, a two-stage algorithm is proposed. In the first stage, the auto-regressive (AR) coefficients of the source signals are estimated using partially separated sources that have been derived from conventional sparseness-based algorithm. In the second stage, the AR model is combined with the mixing equation to form a state-space model. This model is subsequently solved using the Kalman filter in order to obtain the refined source estimate. Simulation results show the effectiveness of proposed sparseness-based AR-Kalman (SPARK) algorithm compared to the conventional sparseness-based algorithms. Benxu Liu, V. G. Reju, Andy W. H. Khong |
ICASSP | 3 |
| 2013 | A multichannel time-domain subspace approach exploiting multiple time-delays for acoustic channel equalizationabstractIt is well-known that the performance of acoustic multichannel equalization (MCEQ) algorithms depend on the modeling delay that has to be pre-defined. In this work, we propose a MCEQ algorithm which achieves full equalization of acoustic room impulse responses in the presence of blind-system identification error. We achieve the above by modeling the inverse filters using multiple delays via columns of the identity matrix which serve as basis vectors for this time-domain sub-space approach. We further show that the proposed algorithm allows one to determine an optimal set of inverse filters corresponding to an appropriate modeling delay. Rajan S. Rashobh, Andy W. H. Khong |
ICASSP | 2 |
| 2013 | On the use of the quaternion generalized Gaussian distribution for footstep detectionabstractWe propose a method to detect human footsteps from a vector-quaternion signal acquired by a tri-axial geophone. The quaternion generalized Gaussian distribution (QGGD) is derived to parameterize variations in the vector-quaternion signal using a shape parameter, quantifying non-Gaussianity and quaternion augmented covariance matrix, quantifying inter-channel correlation. The detection of footsteps is then formulated as binary hypotheses tests in terms of the parameters of the QGGD. The effectiveness of the proposed metrics is evaluated on recorded seismic data. Divya Venkatraman, Vinod V. Reddy, Andy W. H. Khong |
ICASSP | 3 |
| 2013 | Speaker localization and tracking in the presence of sound interference by exploiting speech harmonicityabstractThe performance of conventional acoustic source localization and tracking system reduces significantly when reverberation, noise, and acoustic interference are present. In this paper, a robust speaker tracking algorithm for an enclosed environment in the presence of interference and noise is proposed. We exploit the harmonic structure which is a distinctive feature in speech to enhance the robustness against acoustic interference. In order to extract the speech harmonic information, a beamformer is employed to enhance the signal from a prior estimated source location. A new particle weight update is then computed based on the steered response power function given the estimated speech harmonic information. Simulation results show that the proposed method achieves robustness in localization and tracking of a speech source in the presence of interference, noise and reverberation. Kai Wu 0005, Shu Ting Goh, Andy W. H. Khong |
ICASSP | 3 |
| 2013 | Convergence Analysis of Narrowband Feedback Active Noise Control System With Imperfect Secondary Path EstimationabstractIn many practical active noise control (ANC) applications, feedback structure using estimated secondary path to synthesize reference signal is preferred under various conditions. This paper analyzes the convergence behavior of the narrowband feedback ANC systems with imperfect secondary path estimation. Existing approaches do not include the analysis of the reference signal synthesis errors due to its interrelated feedback nature. In this paper, the reconstruction error is modeled using the secondary path estimation error. Using this model, the effects of estimation errors on the convergence of the feedback ANC system is investigated. To further examine the effects of error in the filtered- x and filtered- y signal paths, these two paths are analyze separately to isolate the effects caused by these paths. Computer simulations are conducted to verify the theoretical analysis presented in the paper. Liang Wang 0007, Woon-Seng Gan, Andy W. H. Khong, Sen M. Kuo |
IEEE Trans. Speech Audio Process. | 3 |
| 2013 | Localization of Taps on Solid Surfaces for Human-Computer Touch InterfacesabstractLocalization of impacts on solid surfaces is a challenging task due to dispersion where the velocity of wave propagation is frequency dependent. In this work, we develop a source localization algorithm on solids with applications to human-computer interface. We employ surface-mounted piezoelectric shock sensors that, in turn, allow us to convert existing flat surfaces to a low-cost touch interface. The algorithm estimates the time-differences-of-arrival between the signals via onset detection in the time-frequency domain. The proposed algorithm is suitable for vibration signals generated by a metal stylus and a finger. The validity of the algorithm is then verified on an aluminium and a glass plate surface. V. G. Reju, Andy W. H. Khong, Amir Bin Sulaiman |
IEEE Trans. Multim. | 2 |
| 2012 | A variable step-size multichannel equalization algorithm exploiting sparseness measure for room acousticsabstractNon-adaptive multichannel equalization (MCEQ) algorithms based on multiple input/output inverse theorem (MINT) is computationally expensive as MINT involves the inversion of a convolution matrix with dimension that is proportional to the length of the acoustic impulse responses. To address this, we propose a MINT-based algorithm that estimates inverse filters by minimizing a cost function iteratively. To further enhance the convergence rate, we formulate an algorithm that employs an adaptive step-size that is derived as a function of the sparseness measure. The proposed algorithm is then applied to existing MINT-based equalization algorithms such as A-MINT and the currently proposed MCEQ-based algorithms. Rajan S. Rashobh, Andy W. H. Khong |
ISCAS | 2 |
| 2012 | DOA estimation of wideband sources without estimating the number of sources
Vinod V. Reddy, Boon Poh Ng, Andy W. H. Khong |
Signal Process. | 4 |
| 2012 | A Fast Frequency-Domain Algorithm for Equalizing Acoustic Impulse ResponsesabstractThe multiple-input/output inverse theorem (MINT) algorithm for multichannel equalization is computationally demanding. Although adaptive MINT reduces the computational complexity, it suffers from slow convergence. In this letter, we propose a low-complexity fast-converging adaptive algorithm for multichannel equalization. The novelty of the approach lies in the adaptive equalization for each frequency bin and its ability to achieve fast convergence in a single step. The proposed algorithm can achieve better equalization of high-order acoustic impulse responses with significant reduction in complexity. Rajan S. Rashobh, Andy W. H. Khong |
IEEE Signal Process. Lett. | 2 |
| 2012 | A Forced Spectral Diversity Algorithm for Speech Dereverberation in the Presence of Near-Common ZerosabstractBlind identification of single-input multiple-output (SIMO) systems is not normally possible if common zeros exist in the channels. Studies of measured acoustic SIMO systems show that near-common zeros occur in such systems as encountered in the speech dereverberation task. We therefore introduce a method to add additional diversity to the SIMO system to be identified which we term forced spectral diversity (FSD) and we show that its use leads to an identification-equalization approach that gives improved dereverberation. As part of this work, we show the link between channel diversity and the effect of common zeros. We also define and discuss in more detail the concept and impact of near-common zeros. The proposed algorithm is presented specifically for a two-channel system where such near-common zeros exist. Andy W. H. Khong, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Equalization of multichannel acoustic system using sub-systems for speech dereverberationabstractAn auto-relation aided multiple-input/output inverse filtering algorithm (A-RAM) is proposed for the inverse filtering of room acoustics in speech dereverberation. In A-RAM, we exploit the relationship between the received reverberant signals and the acoustic impulse responses as well as constraining the reconstructed signals of the two sub-systems. This results in an auto-relation which is then used as a constraint for the adaptive multiple input/output inversion theorem (A-MINT) algorithm. Simulation results, using both synthetic and recorded room impulse responses, show that the proposed A-RAM achieves a fast convergence compared to A-MINT. Andy W. H. Khong |
ICASSP | 2 |
| 2011 | A wavenumber-fitting extrapolation method for FFT-based near-field acoustic holography using microphone arrayabstractNear-field acoustic holography (NAH) is an important acoustic visualization technique which can be implemented efficiently through the fast Fourier transform (FFT). However, when the FFT is applied on a small measurement aperture via a microphone array, significant spectral leakage occurs. Windowing is often employed to reduce this leakage but at the expense of degradation in reconstruction performance due to the corruption of measured data, especially when the source is located around the edge of the sensor array. A wavenumber-fitting method is proposed to extrapolate the measurement aperture effectively. This is achieved by exploiting the wave equation using measurement aperture as the boundary condition. The quality of NAH source reconstruction is shown to be improved significantly using the proposed approach. Benxu Liu, Ramachandran Bremananth, Andy W. H. Khong |
ICASSP | 3 |
| 2011 | Adaptive Channel Equalization of Room Acoustics Exploiting Sparseness ConstraintabstractThe adaptive multiple input/output inverse theorem (A-MINT) algorithm has been applied to equalize acoustic channels. To further increase its convergence rate, we suppress any undesirable non-zero coefficients in the estimated Kronecker delta function at each iteration. We achieve this by introducing the concept of sparseness measure of the estimated Kronecker delta function and using this as an additional constraint to A-MINT. Simulation and experimental results illustrate that the proposed algorithm can achieve faster convergence than A-MINT. Andy W. H. Khong |
IEEE Signal Process. Lett. | 2 |
| 2011 | A Linear Neural Network-Based Approach to Stereophonic Acoustic Echo CancellationabstractWe propose a new adaptive filtering algorithm for stereophonic acoustic echo cancellation. This algorithm uses a linear single-layer feedforward neural network to efficiently decorrelate the tap-input vectors. It achieves an improvement in the misalignment convergence by means of applying the resulted decorrelated tap-input vectors to the coefficient update of the adaptive filters. The advantage of our approach as compared with existing techniques is that our algorithm, in use with the nonlinear preprocessor, can achieve a high rate of misalignment convergence without significantly degrading the quality and stereophonic image of the transmitted signals since our neural network operates on the tap-input vectors as opposed to the transmitted audio signals. We then show that we can achieve an efficient implementation for the proposed decorrelation method by considering the structure of the joint-input covariance matrix of the stereophonic signals. Mehdi Bekrani, Andy W. H. Khong, Mojtaba Lotfizad |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | A Clipping-Based Selective-Tap Adaptive Filtering Approach to Stereophonic Acoustic Echo CancellationabstractStereophonic acoustic echo cancellation remains one of the challenging areas for tele/video-conferencing applications. However, the existence of high interchannel coherence between the two input signals for such systems leads to considerable degradation in misalignment convergence of the adaptive filters. We propose a new algorithm for improving the convergence performance and steady-state misalignment by considering robustness to the source position in the transmission room. We achieve this by exploiting the inherent decorrelating properties of selective-tap adaptive filtering as well as employing a variable clipping threshold for the unselected taps. Simulation results using colored noise and speech signals show an improvement over existing algorithms both in terms of convergence rate as well as steady-state normalized misalignment. Mehdi Bekrani, Andy W. H. Khong, Mojtaba Lotfizad |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Time-Reversal Approach to the Stereophonic Acoustic Echo Cancellation ProblemabstractStereophonic acoustic echo cancellation (SAEC) plays an important role in delivering realistic teleconferencing experience. The fundamental problem of SAEC system is that stereophonic channels are linearly related and this results in slow convergence of the adaptive filters. In this paper, we present a novel algorithm by employing a selective time-reversal block to solve the SAEC problem which results in a significant increase in the convergence performance of adaptive filters such that the stereophonic image as well as quality are preserved. The proposed algorithm employs time-reversal operation on selective blocks of input data samples for one of the two channels to decorrelate stereophonic channels in the SAEC system. To achieve good stereophonic perception, time-reversal operation is only applied to the selective blocks whose magnitudes fall below a pre-determined threshold. Theoretical and numerical simulation results are also studied and investigated to show that the proposed algorithm achieves faster convergence in terms of normalized misalignment and better stereophonic perception with less audio distortion compared to the well-known nonlinear transformation algorithm for the SAEC system. Dinh-Quy Nguyen, Woon-Seng Gan, Andy W. H. Khong |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | A Touch Interface Exploiting Time-Frequency Classification Using Zak Transform for Source Localization on SolidsabstractWe propose a new approach to the development of a touch interface using surface-mounted sensors which allows one to convert a hard surface into a touch pad. This is achieved by using location template matching (LTM), a source localization algorithm that is robust to dispersion and multipath. In this interdisciplinary research, we employ mechanical vibration theories that model wave propagation of the flexural modes of vibration generated by an impact on the surface. We then verify that the amplitude variance across time for each propagating mode frequency is unique to each location on a surface. We show that the Zak transform allows us to faithfully track these amplitude variations and we exploit the uniqueness of this variance as a time-frequency classifier which in turn allows us to localize a finger tap in the context of a human-computer interface. The performance of the proposed algorithm is compared with existing LTM approaches on real surfaces. Kattukandy Rajan Arun, XueXin Yap, Andy W. H. Khong |
IEEE Trans. Multim. | 3 |
| 2010 | Improvement of Speech Source Localization in Noisy Environment Using Overcomplete Rational-Dilation Wavelet TransformsabstractThe generalized cross-correlation using the phase transform prefilter remains popular for the estimation of time-differences-of-arrival. However it is not robust to noise and as a consequence, the performance of direction-of-arrival algorithms is often degraded under low signal-to-noise condition. We propose to address this problem through the use of a wavelet-based speech enhancement technique since the wavelet transform can achieve good denoising performance. The over complete rational-dilation wavelet transform is then exploited to effectively process speech signals due to its higher frequency resolution. In addition, we exploit the joint distribution of the speech in the wavelet domain and develop a novel local noise variance estimator based on the bivariate shrinkage function. As will be shown, our proposed algorithm achieves good direction-of-arrival performance in the presence of noise. Andy W. H. Khong |
CW | 2 |
| 2010 | Source Localization in the Presence of Dispersion for Next Generation Touch InterfaceabstractWe propose a new paradigm of touch interface that allows one to convert daily objects to a touch pad through the use of surface mounted sensors. To achieve a successful touch interface, localization of the finger tap is important. We present an inter-disciplinary approach to improve source localization on solids by means of a mathematical model. It utilizes mechanical vibration theories to simulate the output signals derived from sensors mounted on a physical surface. Utilizing this model, we provide an insight into how phase is distorted in vibrational waves within an aluminium plate which in turn serves as a motivation for our work. We then propose a source localization algorithm based on the phase information of the received signals. We verify the performance of our algorithm using both simulated and recorded data. Amir Bin Sulaiman, Kirill Poletkin, Andy W. H. Khong |
CW | 3 |
| 2010 | A clipping-based adaptive filtering approach for stereophonic acoustic echo cancellationabstractThe use of partial-updating algorithm for reducing interchannel coherence in stereophonic acoustic echo cancellation has been proposed recently. In this work, we show that this algorithm suffers from the lack of robustness against source positions in the transmission room. To address this, we present an insight into this problem and propose a center clipping algorithm that improves the joint optimization between reducing interchannel coherence and increasing energies of selected taps for different source positions. Simulation results using both colored and speech inputs verify the robustness of the proposed algorithm against source positions. Mehdi Bekrani, Andy W. H. Khong, Mojtaba Lotfizad |
ICASSP | 2 |
| 2010 | Localization of acoustic source on solids: A linear predictive coding based algorithm for location template matchingabstractLocation template matching (LTM) is a source localization technique in solids that is robust to dispersion and multipath. This is possible since LTM compares the input with a database of signals made at known locations. With this in place, it is possible to employ LTM in situations where the surface of interest takes an irregular shape. However, one of the existing LTM approaches uses crosscorrelation to compare the input and the database. It should be noted that if any two of the known locations stored in the database are too close, the cross-correlation method may have difficulties differentiating between signals generated from the neighboring points. To address this, we propose an algorithm which employs the linear predictive coding (LPC) that takes into account the dominant frequencies of a received signal. Using this approach, we show that the proposed algorithm is able to improve LTM's source localization accuracy under a real environment in the context of source localization for a touch interface. XueXin Yap, Andy W. H. Khong, Woon-Seng Gan |
ICASSP | 2 |
| 2010 | Neural network based adaptive echo cancellation for stereophonic teleconferencing applicationabstractAcoustic transmission for conferencing systems have progressed from the use of single channel to one that employs stereophonic channels. One of the most important challenges for such stereophonic system is the problem of stereophonic acoustic echo cancellation (SAEC) where a pair of echo cancellers are deployed to estimate the acoustic impulse responses of the receiving room. We propose, in this paper, a neural network based adaptive filtering approach for SAEC. The neural network is employed to decorrelate the input vectors for efficient filter updating, resulting in a high convergence rate of the adaptive filters for this multi-channel acoustic application. To further enhance the efficiency of the proposed algorithm, we then utilize the joint-input correlation matrix of the stereophonic signals so as to simplify the proposed neural network. Simulation results show the improvement in performance of the proposed adaptive SAEC approach over the state-of-the-art algorithms. Mehdi Bekrani, Andy W. H. Khong, Mojtaba Lotfizad |
ICME | 2 |
| 2010 | A touch interface exploiting the use of vibration theories and infinite impulse response filter modeling based localization algorithmabstractResearch into human-machine computer interface (HMI) has been very active in recent years due to the proliferation and advances in software applications. Such devices are aimed at providing a more natural interface for which humans and machines interact. In this multi-disciplinary research, we propose a new approach to the development of a touch interface through the use of surface mounted sensors which allow one to convert hard surfaces into touch pads. We first develop, using mechanical vibration theories, a mathematical model that simulates the output signals derived from sensors mounted on a physical surface such. Utilizing this model, we show that the profile of the output signals is unique not only in time but also in the frequency domain. We then exploit this important property to localize finger taps by developing a source localization algorithm based on infinite impulse response filter model for location template matching. The performance of the proposed algorithm is compared with existing approaches and verified both in a synthetic as well as a real environment for the localization of a finger tap on a touch interface. Kirill Poletkin, XueXin Yap, Andy W. H. Khong |
ICME | 3 |
| 2009 | Blind system identification for speech dereverberation with Forced Spectral DiversityabstractThe common zeros problem for blind system identification (BSI) is well known. It degrades the performance of classic BSI algorithms and therefore imposes the limit on the performance of subsequent speech dereverberation. The effect of near-common zeros has recently been studied in terms of channel diversity and the degradation in performance of BSI and multichannel equalization algorithms has been shown. We now introduce a novel approach to improve channel diversity which we refer to as Forced Spectral Diversity (FSD). The FSD concept uses a combination of spectral shaping filters and effective channel undermodelling. Simulation results show that the proposed approach achieves improved performance with reduced complexity for multichannel BSI in a room acoustics example. Andy W. H. Khong, Patrick A. Naylor |
ICASSP | 2 |
| 2009 | A Class of Sparseness-Controlled Algorithms for Echo CancellationabstractIn the context of acoustic echo cancellation (AEC), it is shown that the level of sparseness in acoustic impulse responses can vary greatly in a mobile environment. When the response is strongly sparse, convergence of conventional approaches is poor. Drawing on techniques originally developed for network echo cancellation (NEC), we propose a class of AEC algorithms that can not only work well in both sparse and dispersive circumstances, but also adapt dynamically to the level of sparseness using a new sparseness-controlled approach. Simulation results, using white Gaussian noise (WGN) and speech input signals, show improved performance over existing methods. The proposed algorithms achieve these improvement with only a modest increase in computational complexity. Pradeep Loganathan, Andy W. H. Khong, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | A lowcomplexity fast converging partial update adaptive algorithm employing variable step-size for acoustic echo cancellationabstractPartial update adaptive algorithms have been proposed as a means of reducing complexity for adaptive filtering. The MMax tap-selection is one of the most popular tap-selection algorithms. It is well known that the performance of such partial update algorithm reduces with reducing number of filter coefficients selected for adaptation. We propose a low complexity and fast converging adaptive algorithm that exploits the MMax tap-selection. We achieve fast convergence with low complexity by deriving a variable step-size for the MMax normalized least-mean-square (MMax-NLMS) algorithm using its mean square deviation. Simulation results verify that the proposed algorithm achieves higher rate of convergence with lower computational complexity compared to the NLMS algorithm. Andy W. H. Khong, Woon-Seng Gan, Patrick A. Naylor, Mike Brookes |
ICASSP | 1 |
| 2008 | Frequency domain selective tap adaptive algorithms for sparse system identificationabstractWe propose a new low complexity and fast converging frequency-domain adaptive algorithm for sparse system identification. This is achieved by exploiting the MMax and SP tap-selection criteria for complexity reduction and fast convergence respectively. We incorporate these tap-selection techniques into the multi-delay filtering (MDF) algorithm in order to reduce the delay inherent in frequency-domain algorithms. We illustrate two such approaches and discuss the tradeoff between convergence performance and computational complexity for these approaches. Simulation results show an improvement in convergence rate for the proposed algorithm over MDF with reduced complexity. The proposed algorithm achieves a convergence performance close to that of the recently proposed but substantially more complex improved proportionate MDF algorithm. Andy W. H. Khong, Milos Doroslovacki, Patrick A. Naylor |
ICASSP | 1 |
| 2008 | Algorithms for identifying clusters of near-common zeros in multichannel blind system identification and equalizationabstractBlind system identification (BSI) and equalization algorithms have been applied to multichannel systems with high order such as found in acoustic impulse responses. Studies on the performance of such algorithms in the presence of near-common zeros have been limited to low order systems. In this work, we propose two high order clustering algorithms which efficiently extract clusters of near-common zeros within a specified pairwise distance in the z-plane. Using these algorithms, we then quantify the number of common zeros that exist in acoustic systems. In addition, we show how these algorithms can be applied to study of BSI and equalization algorithms in the presence of near-common zeros for such acoustic systems. Andy W. H. Khong, Patrick A. Naylor |
ICASSP | 1 |
| 2007 | The Effect of Calibration Errors on Source Localization with Microphone ArraysabstractSource localization employing time-differences-of-arrival has been employed for many applications. The accuracy of source localization is limited by the errors in the time differences of arrival estimation as well as microphone position calibration errors. Because a microphone position error will affect multiple time differences of arrival, correlation between these quantities will be introduced. This work presents a new mathematical framework in which we quantify the localization performance of a microphone array in which the microphone positions are subject to such errors. Andy W. H. Khong, Mike Brookes |
ICASSP (1) | 1 |
| 2007 | Misalignment Performance of Selective Tap Adaptive Algorithms for System Identification of Time-Varying Unknown SystemsabstractSelective tap algorithms have been proposed as a means of reducing complexity for adaptive filtering. MMax tap selection has been employed in many algorithms due to its straightforward implementation. This paper formulates the analysis of two MMax-based algorithms under time-varying unknown system conditions as are often found in practical applications. The steady-state misalignment for the MMax normalized least mean square and the MMax recursive least squares algorithms are derived and their performance is compared to that of their respective full-update algorithms. The tradeoff between computational complexity and misalignment performance is also shown for the MMax normalized least mean square case. Patrick A. Naylor, Andy W. H. Khong, Mike Brookes |
ICASSP (1) | 2 |
| 2007 | Selective-Tap Adaptive Filtering With Performance Analysis for Identification of Time-Varying SystemsabstractSelective-tap algorithms employing the MMax tap selection criterion were originally proposed for low-complexity adaptive filtering. The concept has recently been extended to multichannel adaptive filtering and applied to stereophonic acoustic echo cancellation. This paper first briefly reviews least mean square versions of MMax selective-tap adaptive filtering and then introduces new recursive least squares and affine projection MMax algorithms. We subsequently formulate an analysis of the MMax algorithms for time-varying system identification by modeling the unknown system using a modified Markov process. Analytical results are derived for the tracking performance of MMax selective tap algorithms for normalized least mean square, recursive least squares, and affine projection algorithms. Simulation results are shown to verify the analysis. Andy W. H. Khong, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Proportionate Frequency Domain Adaptive Algorithms for Blind Channel IdentificationabstractWe present fast-converging adaptive blind channel identification algorithms for acoustic room impulse responses. These new algorithms exploit the fast-convergence of the improved proportionate normalized least-mean-square (IPNLMS) algorithm and address the problem of delay inherent in frequency domain algorithms by employing the multi-delay filter (MDF) structure. Simulation results for both speech and white Gaussian noise show that the proposed algorithms outperform current frequency domain blind channel estimation algorithms Rehan Ahmad, Andy W. H. Khong, Patrick A. Naylor |
ICASSP (5) | 2 |
| 2006 | Effect of Interchannel Coherence on Conditioning and Misalignment Performance for Stereo Acoustic ECHO CancellationabstractIt is well known that the performance in terms of misalignment of adaptive algorithms, in general, is dependent on the conditioning of the input signal covariance matrix. For two-channel (stereophonic) adaptive algorithms, this performance is further degraded by the high interchannel coherence between the two input signals. In this paper, we establish the relationship between interchannel coherence of the two input signals and condition of the corre- corresonding covariance matrix for stereo acoustic echo cancellation application. We further show how this relationship affects the misalignment performance of a two-channel frequency-domain adaptive algorithm. We provide simulation results for both WGN and speech input to verify our mathematical analysis. Andy W. H. Khong, Jacob Benesty, Patrick A. Naylor |
ICASSP (5) | 1 |
| 2006 | Stereophonic acoustic echo cancellation: analysis of the misalignment in the frequency domainabstractThe performance in terms of misalignment of adaptive algorithms, in general, is dependent on the conditioning of the input signal covariance matrix. The performance of two-channel adaptive algorithms is further degraded by the high interchannel coherence between the two input signals. In this letter, we establish the relationship between interchannel coherence of the two input signals and condition of the corresponding covariance matrix for stereo acoustic echo cancellation application. We show how this relationship affects the misalignment of a frequency-domain adaptive algorithm. We provide simulation results for both white Gaussian noise and speech input to verify our mathematical analysis. Andy W. H. Khong, Jacob Benesty, Patrick A. Naylor |
IEEE Signal Process. Lett. | 1 |
| 2006 | Stereophonic acoustic echo cancellation employing selective-tap adaptive algorithmsabstractStereophonic acoustic echo cancellation has generated much interest in recent years due to the nonuniqueness and misalignment problems that are caused by the strong interchannel signal coherence. In this paper, we introduce a novel adaptive filtering approach to reduce interchannel coherence which is based on a selective-tap updating procedure. This tap-selection technique is then applied to the normalized least-mean-square, affine projection and recursive least squares algorithms for stereophonic acoustic echo cancellation. Simulation results for the proposed algorithms have shown a significant improvement in convergence rate compared with existing techniques. Andy W. H. Khong, Patrick A. Naylor |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | A family of selective-tap algorithms for stereo acoustic echo cancellationabstractThe use of adaptive filters employing tap-selection for stereophonic acoustic echo cancellation (SAEC) is investigated. We propose to employ subsampling of the tap-input vector, that is intrinsic to partial update schemes, to improve the conditioning of the tap-input autocorrelation matrix hence improving convergence. We investigate the effect of MMax tap-selection on the convergence rate for the single channel case by proposing a new measure which is then used as an optimization parameter in the development of our tap-selection scheme in the two channel case. The resultant exclusive maximum tap-selection is then applied to two channel NLMS, AP and RLS algorithms. Although our main motivation is not the reduction of complexity of SAEC, the proposed tap-selection nevertheless brings significant computation savings in additional to an improved rate of convergence over algorithms using only a nonlinear preprocessor. Andy W. H. Khong, Patrick A. Naylor |
ICASSP (3) | 1 |
| 2005 | Selective-tap adaptive algorithms in the solution of the nonuniqueness problem for stereophonic acoustic echo cancellationabstractWe investigate stereophonic acoustic echo cancellation in which solutions for the system can be nonunique and propose the use of selective-tap adaptive filters to address this problem. The main concept is to employ tap selection to optimize jointly for minimum interchannel coherence and maximum L/sub 2/-norm of the subselected tap-input vectors. The exclusive maximum (XM) tap-selection approach is proposed and applied to normalized least-mean squares (NLMS) and recursive least-squares (RLS) algorithm. We propose an approach for solving the nonuniqueness problem employing XM tap selection in combination with a nonlinear preprocessor. Simulation results show a significant improvement in convergence rate compared with existing techniques. Andy W. H. Khong, Patrick A. Naylor |
IEEE Signal Process. Lett. | 1 |
| 2005 | Corrections to "Selective-Tap Adaptive Algorithms in the Solution of the Nonuniqueness Problem for Stereophonic Acoustic Echo Cancellation"
Andy W. H. Khong, Patrick A. Naylor |
IEEE Signal Process. Lett. | 1 |