VLDB 2026 Research / reviewers in the wild / expert
Woon-Seng Gan
dblp:g/WoonSengGan · also Woon S. Gan
· DBLP profile ↗
121ranked-venue papers
8as first author
48since 2021 · last 2026
0000-0002-7143-1823ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 85 · 5 first-author · 36 since 2021Artificial intelligence and machine learning · 31 · 3 first-author · 11 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reinforcement learning-based selective fixed-filter active noise control (RL-SFANC): From theory to real-time headphone implementation
Zhengding Luo, Haozhe Ma, Dongyuan Shi, Woon-Seng Gan |
Signal Process. | 5 |
| 2026 | Spatial-frequency cued generative fixed-filter active noise control based on deep learning in reverberant environments
Dongyuan Shi, Junwei Ji, Zhengding Luo, Woon-Seng Gan |
Signal Process. | 7 |
| 2026 | A Practical Data-Driven Step-Size Selection Method for Adaptive Active Noise Control Based on Modified Meta-LearningabstractActive noise control (ANC) is widely recognized as an effective and efficient solution for attenuating urban noise. Least mean square (LMS)-based adaptive algorithms, particularly the filtered-reference LMS (FxLMS) algorithm, play a central role in adaptive ANC systems due to their computational efficiency and optional steady-state performance. However, their effectiveness heavily depends on appropriate step-size selection. An unsuitable step size can severely degrade convergence speed and stability. Traditional step-size strategies, such as variable step-size approaches, often involve high computational complexity and are limited to specific noise types. To address this, this letter proposes a data-driven step-size selection method for the FxLMS algorithm based on modified model-agnostic meta-learning (MAML), incorporating a forgetting factor to mitigate the filter's initial zero effect. Compared to conventional methods, the proposed approach can determine an optimal step size across various noise types without requiring additional computations during control, making it highly suitable for practical deployment. Numerical simulations using real-world paths and noise further verify its effectiveness. Luyuan Li, Xiruo Su, Dongyuan Shi, Jie Chen 0022, Woon-Seng Gan |
IEEE Signal Process. Lett. | 5 |
| 2026 | WINTER: Wrapped Interval Normalization for Elevation Representation in Stereo 3-D Sound Event Localization and DetectionabstractStereo 3-D Sound Event Localization and Detection (SELD) faces a fundamental geometric conflict: the simultaneous regression of unitless angular direction and unbounded physical distance. This heterogeneity destabilizes training, as large-magnitude distance errors often overwhelm fine-grained directional gradients in the loss landscape. To resolve this, we propose WINTER (Wrapped Interval NormalizaTion for Elevation Representation), a novel embedding that projects the heterogeneous target space onto a unified, homogeneous manifold. By mapping scalar distance to the latent elevation axis of stereo signals, WINTER transforms the problem into a single-head regression on a pseudo 3-D unit sphere. This geometric unification restores the validity of angular metrics and enables the introduction of Stereo-DICE, a differentiable proxy for the polyphonic F1-score that optimizes joint localization and detection. Validated on the DCASE 2025 Task 3 benchmark, our framework achieves a relative reduction of up to 8.19% in overall SELD error compared to standard regression baselines, significantly improving far-field stability. Code is available at:https://github.com/itsjunwei/WINTER-SELD Junwei Yeow, Ee-Leng Tan, Santi Peksi, Woon-Seng Gan |
IEEE Signal Process. Lett. | 4 |
| 2025 | Preventing output saturation in active noise control: An output-constrained Kalman filter approachabstractThe Kalman filter (KF)-based active noise control (ANC) system demonstrates superior tracking and faster convergence compared to the least mean square (LMS) method, particularly in dynamic noise cancellation scenarios. However, in environments with extremely high noise levels, the power of the control signal can exceed the system’s rated output power due to hardware limitations, leading to output saturation and subsequent non-linearity. To mitigate this issue, a modified KF with an output constraint is proposed. In this approach, the disturbance treated as an measurement is re-scaled by a constraint factor, which is determined by the system’s rated power, the secondary path gain, and the disturbance power. As a result, the output power of the system, i.e. the control signal, is indirectly constrained within the maximum output of the system, ensuring stability. Simulation results indicate that the proposed algorithm not only achieves rapid suppression of dynamic noise but also effectively prevents non-linearity due to output saturation, highlighting its practical significance. Junwei Ji, Dongyuan Shi, Xiaoyi Shen, Zhengding Luo, Woon-Seng Gan |
ICASSP | 6 |
| 2025 | Transferable Selective Virtual Sensing Active Noise Control Technique Based on Metric LearningabstractVirtual sensing (VS) technology enables active noise control (ANC) systems to attenuate noise at virtual locations distant from the physical error microphones. Appropriate auxiliary filters (AF) can significantly enhance the effectiveness of VS approaches. The selection of appropriate AF for various types of noise can be automatically achieved using convolutional neural networks (CNNs). However, training the CNN model for different ANC systems is often labour-intensive and timeconsuming. To tackle this problem, we propose a novel method, Transferable Selective VS, by integrating metric-learning technology into CNN-based VS approaches. The Transferable Selective VS method allows a pre-trained CNN to be applied directly to new ANC systems without requiring retraining, and it can handle unseen noise types. Numerical simulations demonstrate the effectiveness of the proposed method in attenuating suddenvarying broadband noises and real-world noises. Dongyuan Shi, Zhengding Luo, Xiaoyi Shen, Junwei Ji, Woon-Seng Gan |
ICASSP | 6 |
| 2025 | Multi-granularity acoustic information fusion for sound event detection
Han Yin, Jisheng Bai, Mou Wang, Susanto Rahardja, Dongyuan Shi, Woon-Seng Gan |
Signal Process. | 7 |
| 2025 | Self-Boosted Weight-Constrained FxLMS: A Robustness Distributed Active Noise Control Algorithm Without Internode CommunicationabstractCompared to the conventional centralized multichannel active noise control (MCANC) algorithm, which requires substantial computational resources, decentralized approaches exhibit higher computational efficiency but typically result in inferior noise reduction performance. To enhance performance, distributed ANC methods have been introduced, enabling information exchange among ANC nodes; however, the resulting communication latency often compromises system stability. To overcome these limitations, we propose a self-boosted weight-constrained filtered-reference least mean square (SB-WCFxLMS) algorithm for the distributed MCANC system without internode communication. The WCFxLMS algorithm is specifically designed to mitigate divergence issues caused by the internode cross-talk effect. The self-boosted strategy lets each ANC node independently adapt its constraint parameters based on its local noise reduction performance, thus ensuring effective noise cancellation without the need for inter-node communication. With the assistance of this mechanism, this approach significantly reduces both computational complexity and communication overhead. Numerical simulations employing real acoustic paths and compressor noise validate the effectiveness and robustness of the proposed system. The results demonstrate that our proposed method achieves satisfactory noise cancellation performance with minimal resource requirements. Junwei Ji, Dongyuan Shi, Zhengding Luo, Woon-Seng Gan |
IEEE Signal Process. Lett. | 5 |
| 2025 | Data-Driven Method to Accelerate Convergence of Adaptive Hybrid Active Noise Control: Two-Stage Model-Agnostic Meta-LearningabstractHybrid active noise control (ANC) is widely employed in portable commercial products to attenuate both broadband and narrowband noise. Although the adaptive hybrid ANC, updated by the filtered reference least mean square (FxLMS) algorithm, can achieve optimal noise control even with uncorrelated noise, its slow convergence speed significantly decreases dynamic noise reduction performance. To address this challenge, we propose a two-stage model-agnostic meta-learning (MAML) approach to compute the optimal initial coefficients of the control filters for the hybrid ANC, effectively reducing the convergence time of adaptive algorithms. Different from conventional variable step-size strategies, this data-driven method determines the optimal initial coefficients based on the statistical characteristics of the noise, ensuring system stability. Furthermore, numerical simulations demonstrate that two-stage MAML initialization of adaptive hybrid ANC significantly accelerates convergence speed for attenuating broadband and real-world noise. Xiaoyi Shen, Dongyuan Shi, Woon-Seng Gan |
IEEE Signal Process. Lett. | 3 |
| 2025 | Corrigendum to "ARAUS: A Large-Scale Dataset and Baseline Models of Affective Responses to Augmented Urban Soundscapes"abstractPresents corrections to the article “ARAUS: A Large-Scale Dataset and Baseline Models of Affective Responses to Augmented Urban Soundscapes”. Kenneth Ooi, Zhen-Ting Ong, Karn Watcharasupat, Bhan Lam, Joo Young Hong, Woon-Seng Gan |
IEEE Trans. Affect. Comput. | 6 |
| 2024 | Unsupervised Learning Based End-to-End Delayless Generative Fixed-Filter Active Noise ControlabstractDelayless noise control is achieved by our earlier generative fixed-filter active noise control (GFANC) framework through efficient coordination between the co-processor and real-time controller. However, the one-dimensional convolutional neural network (1D CNN) in the co-processor requires initial training using labelled noise datasets. Labelling noise data can be resource-intensive and may introduce some biases. In this paper, we propose an unsupervised-GFANC approach to simplify the 1D CNN training process and enhance its practicality. During training, the co-processor and real-time controller are integrated into an end-to-end differentiable ANC system. This enables us to use the accumulated squared error signal as the loss for training the 1D CNN. With this unsupervised learning paradigm, the unsupervised-GFANC method not only omits the labelling process but also exhibits better noise reduction performance compared to the supervised GFANC method in real noise experiments. Zhengding Luo, Dongyuan Shi, Xiaoyi Shen, Woon-Seng Gan |
ICASSP | 4 |
| 2024 | Audiolog: LLMs-Powered Long Audio Logging with Hybrid Token-Semantic Contrastive LearningabstractPrevious studies in automated audio captioning have faced difficulties in accurately capturing the complete temporal details of acoustic scenes and events within long audio sequences. This paper presents AudioLog, a large language models (LLMs)-powered audio logging system with hybrid token-semantic contrastive learning. Specifically, we propose to fine-tune the pre-trained hierarchical token-semantic audio Transformer by incorporating contrastive learning between hybrid acoustic representations. We then leverage LLMs to generate audio logs that summarize textual descriptions of the acoustic environment. Finally, we evaluate the AudioLog system on two datasets with both scene and event annotations. Experiments show that the proposed system achieves exceptional performance in acoustic scene classification and sound event detection, surpassing existing methods in the field. Further analysis of the prompts to LLMs demonstrates that AudioLog can effectively summarize long audio sequences1. To the best of our knowledge, this approach is the first attempt to leverage LLMs for summarizing long audio sequences. Jisheng Bai, Han Yin, Mou Wang, Dongyuan Shi, Woon-Seng Gan, Susanto Rahardja |
ICME | 5 |
| 2024 | GFANC-RL: Reinforcement Learning-based Generative Fixed-filter Active Noise Control
Zhengding Luo, Haozhe Ma, Dongyuan Shi, Woon-Seng Gan |
Neural Networks | 4 |
| 2024 | What is behind the meta-learning initialization of adaptive filter? - A naive method for accelerating convergence of adaptive multichannel active noise control
Dongyuan Shi, Woon-Seng Gan, Xiaoyi Shen, Zhengding Luo, Junwei Ji |
Neural Networks | 2 |
| 2024 | A survey on adaptive active noise control algorithms overcoming the output saturation effect
Dongyuan Shi, Xiaoyi Shen, Junwei Ji, Woon-Seng Gan |
Signal Process. | 5 |
| 2024 | GFANC-Kalman: Generative Fixed-Filter Active Noise Control With CNN-Kalman FilteringabstractSelective Fixed-filter Active Noise Control (SFANC) is limited by its selection of a single candidate from pre-trained control filters. In contrast, Generative Fixed-filter Active Noise Control (GFANC) addresses this limitation by employing an adaptive combination of sub control filters to generate more suitable control filters for different primary noises. However, GFANC solely relies on the information from the current noise frame to generate its control filter, resulting in potential inaccuracies when dealing with dynamic noises. Therefore, we propose a GFANC-Kalman approach that integrates an efficient one-dimensional convolutional neural network (1D CNN) with a Kalman filter to further improve the performance of GFANC. Specifically, the weight vector used to combine sub control filters is predicted by the 1D CNN for each noise frame, and then processed by the Kalman filter with minimal complexity. By considering the correlation between adjacent noise frames, the Kalman filter can enhance the accuracy and robustness of weight vector prediction. Hence, GFANC-Kalman is more able to adapt to changes in noise distribution, particularly for dynamic noises. Numerical simulations validate the efficacy of the proposed GFANC-Kalman approach in dealing with real-world dynamic noises. Zhengding Luo, Dongyuan Shi, Xiaoyi Shen, Junwei Ji, Woon-Seng Gan |
IEEE Signal Process. Lett. | 5 |
| 2024 | Spatial-Frequency-Based Selective Fixed-Filter Algorithm for Multichannel Active Noise ControlabstractThe multichannel active noise control (MCANC) approach is widely regarded as an effective solution to achieve a large noise cancellation zone in a complicated acoustic environment. However, the sluggish convergence and massive computation of traditional adaptive multichannel active control algorithms typically impede the MCANC system's practical applications. The recently developed selective fixed-filter method offers a way to decrease the computational load in real-time scenarios and enhance the reaction time. Nevertheless, this method is specifically designed for the single-channel ANC system and only considers the frequency information of the noise. This inevitably impacts the effectiveness of reducing noise from various directions, particularly in the MCANC system. Therefore, we proposed a spatial-frequency-based selective fixed-filter ANC technique that adopts the Bhattacharyya Distance Matching (SFANC-BdM). In our work, the BdM is a one-step spectra and is designed by calculating similarity of different data distribution. According to the most similar case, the corresponding control filter is then selected. By avoiding separately extracting the direction and frequency information, the proposed method significantly increases the algorithm's efficiency. Compared to the conventional SFANC method, it enables a more accurate filter choice and achieves better noise reduction. Xiruo Su, Dongyuan Shi, Zhijuan Zhu, Woon-Seng Gan, Lingyun Ye |
IEEE Signal Process. Lett. | 4 |
| 2024 | ARAUS: A Large-Scale Dataset and Baseline Models of Affective Responses to Augmented Urban SoundscapesabstractChoosing optimal maskers for existing soundscapes to effect a desired perceptual change via soundscape augmentation is non-trivial due to extensive varieties of maskers and a dearth of benchmark datasets with which to compare and develop soundscape augmentation models. To address this problem, we make publicly available the ARAUS (Affective Responses to Augmented Urban Soundscapes) dataset, which comprises a five-fold cross-validation set and independent test set totaling 25,440 unique subjective perceptual responses to augmented soundscapes presented as audio-visual stimuli. Each augmented soundscape is made by digitally adding “maskers” (bird, water, wind, traffic, construction, or silence) to urban soundscape recordings at fixed soundscape-to-masker ratios. Responses were then collected by asking participants to rate how pleasant, annoying, eventful, uneventful, vibrant, monotonous, chaotic, calm, and appropriate each augmented soundscape was, in accordance with ISO/TS 12913-2:2018. Participants also provided relevant demographic information and completed standard psychological questionnaires. We perform exploratory and statistical analysis of the responses obtained to verify internal consistency and agreement with known results in the literature. Finally, we demonstrate the benchmarking capability of the dataset by training and comparing four baseline models for urban soundscape pleasantness: a low-parameter regression model, a high-parameter convolutional neural network, and two attention-based networks in the literature. Kenneth Ooi, Zhen-Ting Ong, Karn Watcharasupat, Bhan Lam, Joo Young Hong, Woon-Seng Gan |
IEEE Trans. Affect. Comput. | 6 |
| 2024 | Delayless Generative Fixed-Filter Active Noise Control Based on Deep Learning and Bayesian FilterabstractThe selective fixed-filter active noise control (SFANC) method can select suitable pre-trained control filters to attenuate incoming noises. However, the limited number of pre-trained filters is insufficient to effectively control various forms of noise, especially when the incoming noise differs much from the filter-training noises. To address this limitation and generate more appropriate control filters, a generative fixed-filter active noise control approach based on Bayesian filter (GFANC-Bayes) is proposed in this paper. The GFANC-Bayes method can automatically generate suitable control filters by combining sub control filters. The combination weights of sub control filters are predicted via a one-dimensional convolutional neural network (1D CNN). Based on prior information and predicted information, Bayesian filtering technique is applied to decide the combination weights. By considering the correlation between adjacent noise frames, the Bayesian filter can enhance the accuracy and robustness of predicting combination weights. Simulations on real-world noises indicate that the GFANC-Bayes method achieves superior noise reduction performance than SFANC and a faster response time than FxLMS. Moreover, experiments on different acoustic paths demonstrate its robustness and transferability. Zhengding Luo, Dongyuan Shi, Woon-Seng Gan |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | A Practical Distributed Active Noise Control Algorithm Overcoming Communication RestrictionsabstractBy assigning the massive computing tasks of the traditional multichannel active noise control (MCANC) system to several distributed control nodes, distributed multichannel active noise control (DM-CANC) techniques have become effective global noise reduction solutions with low computational costs. However, existing DMCANC algorithms simply complete the distribution of traditional centralized algorithms by combining neighbour nodes’ information but rarely consider the degraded control performance and system stability of distributed units caused by delays and interruptions in communication. Hence, this paper develops a novel DMCANC algorithm that utilizes the compensation filters and neighbour nodes’ information to counterbalance the cross-talk effect between channels while maintaining independent weight updating. Since the neighbours’ information required barely affects the local control filter updating in each node, this approach can tolerate communication delay and interruption to some extent. Numerical simulations demonstrate that the proposed algorithm can achieve satisfactory noise reduction performance and high robustness to real-world communication challenges. Junwei Ji, Dongyuan Shi, Zhengding Luo, Xiaoyi Shen, Woon-Seng Gan |
ICASSP | 5 |
| 2023 | Real-Time Modelling of Observation Filter in the Remote Microphone Technique for an Active Noise Control ApplicationabstractThe remote microphone technique (RMT) is often used in active noise control (ANC) applications to overcome design constraints in microphone placements by estimating the acoustic pressure at inconvenient locations using a pre-calibrated observation filter (OF), albeit limited to stationary primary acoustic fields. While the OF estimation in varying primary fields can be significantly improved through the recently proposed source decomposition technique, it requires knowledge of the relative source strengths between incoherent primary noise sources. This paper proposes a method for combining the RMT with a new source-localization technique to estimate the source ratio parameter. Unlike traditional source-localization techniques, the proposed method is capable of being implemented in a real-time RMT application. Simulations with measured responses from an open-aperture ANC application showed a good estimation of the source ratio parameter, which allows the observation filter to be modelled in real-time. Chung Kwan Lai, Bhan Lam, Dongyuan Shi, Woon-Seng Gan |
ICASSP | 4 |
| 2023 | Deep Generative Fixed-Filter Active Noise ControlabstractDue to the slow convergence and poor tracking ability, conventional LMS-based adaptive algorithms are less capable of handling dynamic noises. Selective fixed-filter active noise control (SFANC) can significantly reduce response time by selecting appropriate pre-trained control filters for different noises. Nonetheless, the limited number of pre-trained control filters may affect noise reduction performance, especially when the incoming noise differs much from the initial noises during pre-training. Therefore, a generative fixed-filter active noise control (GFANC) method is proposed in this paper to overcome the limitation. Based on deep learning and a perfect-reconstruction filter bank, the GFANC method only requires a few prior data (one pre-trained broadband control filter) to automatically generate suitable control filters for various noises. The efficacy of the GFANC method is demonstrated by numerical simulations on real-recorded noises. Zhengding Luo, Dongyuan Shi, Xiaoyi Shen, Junwei Ji, Woon-Seng Gan |
ICASSP | 5 |
| 2023 | Autonomous Soundscape Augmentation with Multimodal Fusion of Visual and Participant-Linked InputsabstractAutonomous soundscape augmentation systems typically use trained models to pick optimal maskers to effect a desired perceptual change. While acoustic information is paramount to such systems, contextual information, including participant demographics and the visual environment, also influences acoustic perception. Hence, we propose modular modifications to an existing attention-based deep neural network, to allow early, mid-level, and late feature fusion of participant-linked, visual, and acoustic features. Ablation studies on module configurations and corresponding fusion methods using the ARAUS dataset show that contextual features improve the model performance in a statistically significant manner on the normalized ISO Pleasantness, to a mean squared error of 0.1194±0.0012 for the best-performing all-modality model, against 0.1217±0.0009 for the audio-only model. Soundscape augmentation systems can thereby leverage multimodal inputs for improved performance. We also investigate the impact of individual participant-linked factors using trained models to illustrate improvements in model explainability. Kenneth Ooi, Karn Watcharasupat, Bhan Lam, Zhen-Ting Ong, Woon-Seng Gan |
ICASSP | 5 |
| 2023 | A Momentum Two-Gradient Direction Algorithm with Variable Step Size Applied to Solve Practical Output Constraint Issue for Active Noise ControlabstractActive noise control (ANC) has been widely utilized to reduce unwanted environmental noise. The primary objective of ANC is to generate an anti-noise with the same amplitude but the opposite phase of the primary noise using the secondary source. However, the effectiveness of the ANC application is impacted by the speaker’s output saturation. This paper proposes a two-gradient direction ANC algorithm with a momentum factor to solve the saturation with faster convergence. In order to make it implemented in real-time, a computation-effective variable step size approach is applied to further reduce the steady-state error brought on by the changing gradient directions. The time constant and step size bound for the momentum two-gradient direction algorithm is analyzed. Simulation results show that the proposed algorithm performs effectively in the time-unvaried and time-varied environment. Xiaoyi Shen, Dongyuan Shi, Zhengding Luo, Junwei Ji, Woon-Seng Gan |
ICASSP | 5 |
| 2023 | Implementing Continuous HRTF Measurement in Near-FieldabstractHead-related transfer function (HRTF) is an essential component to create an immersive listening experience over headphones for virtual reality (VR) and augmented reality (AR) applications. Metaverse combines VR and AR to create immersive digital experiences, and users are very likely to interact with virtual objects in the near-field (NF). The HRTFs of such objects are highly individualized and dependent on directions and distances. Hence, a significant number of HRTF measurements at different distances in the NF would be needed. Using conventional static stop-and-go HRTF measurement methods to acquire these measurements would be time-consuming and tedious for human listeners. In this paper, we propose a continuous measurement system targeted for the NF, and efficiently capturing HRTFs in the horizontal plane within 45 secs. Comparative experiments are performed on head and torso similar (HATS) and human listeners to evaluate system consistency and robustness. Ee-Leng Tan, Santi Peksi, Woon-Seng Gan |
ICASSP | 3 |
| 2023 | Partially Randomizing Transformer Weights for Dialogue Response Diversity
Jing Yang Lee, Kong-Aik Lee, Woon-Seng Gan |
PACLIC | 3 |
| 2023 | Multichannel two-gradient direction filtered reference least mean square algorithm for output-constrained multichannel active noise control
Dongyuan Shi, Bhan Lam, Xiaoyi Shen, Woon-Seng Gan |
Signal Process. | 4 |
| 2023 | MOV-Modified-FxLMS Algorithm With Variable Penalty Factor in a Practical Power Output Constrained Active Control SystemabstractPractical Active Noise Control (ANC) systems typically require a restriction in their maximum output power, to prevent overdriving the loudspeaker and causing system instability. Recently, the minimum output variance filtered-reference least mean square (MOV-FxLMS) algorithm was shown to have optimal control under output constraint with an analytically formulated penalty factor, but it needs offline knowledge of disturbance power and secondary path gain. The constant penalty factor in MOV-FxLMS is also susceptible to variations in disturbance power that could cause output power constraint violations. This paper presents a new variable penalty factor that utilizes the estimated disturbance in the established Modified-FxLMS (MFxLMS) algorithm, resulting in a computationally efficient MOV-MFxLMS algorithm that can adapt to changes in disturbance levels in real-time. Numerical simulation with real noise and plant response showed that the variable penalty factor always manages to meet its maximum power output constraint despite sudden changes in disturbance power, whereas the fixed penalty factor has suffered from a constraint mismatch. Chung Kwan Lai, Dongyuan Shi, Bhan Lam, Woon-Seng Gan |
IEEE Signal Process. Lett. | 4 |
| 2023 | Transferable Latent of CNN-Based Selective Fixed-Filter Active Noise ControlabstractPractical active noise control (ANC) systems, like the active noise cancellation headphone, usually adopt a control filter with preset coefficients to achieve satisfactory noise reduction performance for dynamic noise and higher robustness. In this strategy, selecting the appropriate control filter for different types of noise is critical to the noise cancellation performance, and this selection mechanism is typically determined by trial and error. Hence, this article proposes a computation-efficient one-dimensional convolutional neural network capable of selecting the most suitable pre-trained control filter for each distinct primary noise. Applying the similarity matching method allows the proposed model to have a better generalization and can even deal with zero-shot noise, whose class does not exist in the training set. The Large-margin softmax (L-softmax) is also investigated to improve the proposed model's performance. Furthermore, when dealing with the N-shot learning problem, where there are few known real-world noise samples for the ANC system, an additional fine-tuning strategy is used to improve control filter selection accuracy. Numerical simulations on measured primary and secondary paths validate the proposed method's efficacy. Dongyuan Shi, Woon-Seng Gan, Bhan Lam, Zhengding Luo, Xiaoyi Shen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | A Frequency-Domain Output-Constrained Active Noise Control Algorithm Based on an Intuitive Circulant Convolutional Penalty FactorabstractDue to their computational efficiency, least mean square (LMS)–based algorithms are still widely utilized to achieve optimal control in active noise control (ANC) applications. Real-world implementation of advanced ANC functionalities, such as selective cancellation of frequencies, is nonetheless hampered by complexity trade-offs, especially with computationally-expensive frequency-domain approaches. Prevailing time-domain adaptive algorithms – proposed to alleviate complexities from transformation – continue to incur increased complexities while constraining the magnitude of frequency bins in the time-domain filters. To address existing complexities in time-domain approaches, this paper proposes a circulant convolutional penalty factor that assists the extended leaky filtered-reference LMS (FxLMS) algorithm in achieving frequency constraint without any frequency-domain transform. This circulant convolutional penalty factor is readily determined by methods for designing finite response filters, such as frequency sampling. Additionally, the coordinate descent method is adopted to further reduce the proposed algorithm's computations, significantly increasing its feasibility for implementation on conventional real-time processors. Finally, the numerical simulations performed on the measured primary and secondary paths demonstrate the effectiveness of the proposed algorithm. Dongyuan Shi, Woon-Seng Gan, Bhan Lam, Xiaoyi Shen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | DLVGen: A Dual Latent Variable Approach to Personalized Dialogue GenerationabstractThe generation of personalized dialogue is vital to natural and human-like conversation. Typically, personalized dialogue generation models involve conditioning the generated response on the dialogue history and a representation of the persona/personality of the interlocutor. As it is impractical to obtain the persona/personality representations for every interlocutor, recent works have explored the possibility of generating personalized dialogue by finetuning the model with dialogue examples corresponding to a given persona instead. However, in real-world implementations, a sufficient number of corresponding dialogue examples are also rarely available. Hence, in this paper, we propose a Dual Latent Variable Generator (DLVGen) capable of generating personalized dialogue in the absence of any persona/personality information or any corresponding dialogue examples. Unlike prior work, DLVGen models the latent distribution over potential responses as well as the latent distribution over the agent's potential persona. During inference, latent variables are sampled from both distributions and fed into the decoder. Empirical results show that DLVGen is capable of generating diverse responses which accurately incorporate the agent's persona. Jing Yang Lee, Kong-Aik Lee, Woon-Seng Gan |
ICAART (2) | 3 |
| 2022 | Improving Contextual Coherence in Variational Personalized and Empathetic Dialogue AgentsabstractIn recent years, latent variable models, such as the Conditional Variational Auto Encoder (CVAE), have been applied to both personalized and empathetic dialogue generation. Prior work have largely focused on generating diverse dialogue responses that exhibit persona consistency and empathy. However, when it comes to the contextual coherence of the generated responses, there is still room for improvement. Hence, to improve the contextual coherence, we propose a novel Uncertainty Aware CVAE (UA-CVAE) framework. The UA-CVAE framework involves approximating and incorporating the aleatoric uncertainty during response generation. We apply our framework to both personalized and empathetic dialogue generation. Empirical results show that our framework significantly improves the contextual coherence of the generated response. Additionally, we introduce a novel automatic metric for measuring contextual coherence, which was found to correlate positively with human judgement. Jing Yang Lee, Kong-Aik Lee, Woon-Seng Gan |
ICASSP | 3 |
| 2022 | SALSA-Lite: A Fast and Effective Feature for Polyphonic Sound Event Localization and Detection with Microphone ArraysabstractPolyphonic sound event localization and detection (SELD) has many practical applications in acoustic sensing and monitoring. However, the development of real-time SELD has been limited by the demanding computational requirement of most recent SELD systems. In this work, we introduce SALSA-Lite, a fast and effective feature for polyphonic SELD using microphone array inputs. SALSA-Lite is a lightweight variation of a previously proposed SALSA feature for polyphonic SELD. SALSA, which stands for Spatial Cue-Augmented Log-Spectrogram, consists of multichannel log-spectrograms stacked channel-wise with the normalized principal eigenvectors of the spectrotemporally corresponding spatial covariance matrices. In contrast to SALSA, which uses eigenvector-based spatial features, SALSA-Lite uses normalized inter-channel phase differences as spatial features, allowing a 30-fold speedup compared to the original SALSA feature. Experimental results on the TAU-NIGENS Spatial Sound Events 2021 dataset showed that the SALSA-Lite feature achieved competitive performance compared to the full SALSA feature, and significantly outperformed the traditional feature set of multichannel log-mel spectrograms with generalized cross-correlation spectra. Specifically, using SALSA-Lite features increased localization-dependent F1 score and class-dependent localization recall by 15% and 5%, respectively, compared to using multichannel log-mel spectrograms with generalized cross-correlation spectra. Thi Ngoc Tho Nguyen, Douglas L. Jones, Karn Watcharasupat, Huy Phan, Woon-Seng Gan |
ICASSP | 5 |
| 2022 | Probably Pleasant? A Neural-Probabilistic Approach to Automatic Masker Selection for Urban Soundscape AugmentationabstractSoundscape augmentation, which involves the addition of sounds known as "maskers" to a given soundscape, is a human-centric urban noise mitigation measure aimed at improving the overall sound-scape quality. However, the choice of maskers is often predicated on laborious processes and is inflexible to the time-varying nature of real-world soundscapes. Owing to the perceptual uniqueness of each soundscape and the inherent subjectiveness of human perception, we propose a probabilistic perceptual attribute predictor (PPAP) that predicts parameters of random distributions as outputs instead of a single deterministic value. Using the PPAP, we developed a novel automatic masker selection system (AMSS), which selects optimal masker candidates based on the predicted distribution of the ISO 12913-3 Pleasantness score for a given soundscape. Via a largescale listening test with 300 participants, we collected 12600 subjective responses, each to a unique augmented soundscape, to train the PPAP models in a 5-fold cross-validation scheme. Using a convolutional recurrent neural network backbone and experimenting with several variants of the attention mechanism for the PPAP, we evaluated the proposed system on a blind test set with 48 unseen augmented soundscapes to assess the effectiveness of the probabilistic output scheme over traditional deterministic systems. Kenneth Ooi, Karn Watcharasupat, Bhan Lam, Zhen-Ting Ong, Woon-Seng Gan |
ICASSP | 5 |
| 2022 | A Hybrid Approach to Combine Wireless and Earcup Microphones for ANC Headphones with Error Separation ModuleabstractActive noise control (ANC) technology is commonly used to cancel acoustic noise in daily life. The conventional ANC headphone, being one of the mature commercial products that implement this approach, utilizes microphones on its earcup to pick up the reference signal. However, in a multi-noise source situation, the reference signal mixed with uncorrelated interference usually results in poor noise reduction performance of the ANC system. Hence, we proposed a novel hybrid approach that employs wireless microphones to acquire high signal-to-noise-ratio reference signals from far-end noise sources, increasing coherence and thus improving noise reduction performance. Additionally, an error separation model is applied in the proposed structure to enhance the coherence between the error signal and each adaptive filter. As a result, the proposed hybrid approach to combine wireless and earcup microphones for ANC headphone significantly improves its noise reduction performance when dealing with multi-noise sources. Furthermore, numerical simulation and real-time experiments of the proposed structure have shown that it improves noise reduction performance by 4 − 6 dB when compared to a conventional ANC headphone. Xiaoyi Shen, Dongyuan Shi, Woon-Seng Gan |
ICASSP | 3 |
| 2022 | End-to-End Complex-Valued Multidilated Convolutional Neural Network for Joint Acoustic Echo Cancellation and Noise SuppressionabstractEcho and noise suppression is an integral part of a full-duplex communication system. Many recent acoustic echo cancellation (AEC) systems rely on a separate adaptive filtering module for linear echo suppression and a neural module for residual echo suppression. However, in practice, adaptive filtering modules require time to converge and remain susceptible to changes in the acoustic environment. This introduces unnecessary delays to AEC systems using this two-stage framework, despite neural modules already having the capability to suppress both linear and nonlinear echo components. In this paper, we exploit the offset-compensating property of complex time-frequency masks and propose an end-to-end complex-valued neural network architecture. The building block of the proposed model is a pseudocomplex extension of the densely-connected multidilated DenseNet (D3Net), resulting in a very small network of only 354K parameters. The architecture utilized the multi-resolution nature of the D3Net to eliminate the need for pooling, allowing feature extraction using large receptive fields without any loss of output resolution. We also propose a dual-mask technique for joint echo and noise suppression with simultaneous speech enhancement. Evaluation on both synthetic and real test sets demonstrated promising results across multiple energy-based metrics and perceptual proxies. Karn Watcharasupat, Thi Ngoc Tho Nguyen, Woon-Seng Gan, Shengkui Zhao, Bin Ma 0001 |
ICASSP | 3 |
| 2022 | FRCRN: Boosting Feature Representation Using Frequency Recurrence for Monaural Speech EnhancementabstractConvolutional recurrent networks (CRN) integrating a convolutional encoder-decoder (CED) structure and a recurrent structure have achieved promising performance for monaural speech enhancement. However, feature representation across frequency context is highly constrained due to limited receptive fields in the convolutions of CED. In this paper, we propose a convolutional recurrent encoder-decoder (CRED) structure to boost feature representation along the frequency axis. The CRED applies frequency recurrence on 3D convolutional feature maps along the frequency axis following each convolution, therefore, it is capable of catching long-range frequency correlations and enhancing feature representations of speech inputs. The proposed frequency recurrence is realized efficiently using a feedforward sequential memory network (FSMN). Besides the CRED, we insert two stacked FSMN layers between the encoder and the decoder to model further temporal dynamics. We name the proposed framework as Frequency Recurrent CRN (FRCRN). We design FRCRN to predict complex Ideal Ratio Mask (cIRM) in complex-valued domain and optimize FRCRN using both time-frequency-domain and time-domain losses. Our proposed approach achieved state-of-the-art performance on wideband bench-mark datasets and achieved 2nd place for the real-time fullband track in terms of Mean Opinion Score (MOS) and Word Accuracy (WAcc) in the ICASSP 2022 Deep Noise Suppression (DNS) challenge. Shengkui Zhao, Bin Ma 0001, Karn Watcharasupat, Woon-Seng Gan |
ICASSP | 4 |
| 2022 | Selective fixed-filter active noise control based on convolutional neural network
Dongyuan Shi, Bhan Lam, Kenneth Ooi, Xiaoyi Shen, Woon-Seng Gan |
Signal Process. | 5 |
| 2022 | A Hybrid SFANC-FxNLMS Algorithm for Active Noise Control Based on Deep LearningabstractThe selective fixed-filter active noise control (SFANC) method selecting the best pre-trained control filters for various types of noise can achieve a fast response time. However, it may lead to large steady-state errors due to inaccurate filter selection and the lack of adaptability. In comparison, the filtered-X normalized least-mean-square (FxNLMS) algorithm can obtain lower steady-state errors through adaptive optimization. Nonetheless, its slow convergence has a detrimental effect on dynamic noise attenuation. Therefore, this paper proposes a hybrid SFANC-FxNLMS approach to overcome the adaptive algorithm’s slow convergence and provide a better noise reduction level than the SFANC method. A lightweight one-dimensional convolutional neural network (1D CNN) is designed to automatically select the most suitable pre-trained control filter for each frame of the primary noise. Meanwhile, the FxNLMS algorithm continues to update the coefficients of the chosen pre-trained control filter at the sampling rate. Owing to the effective combination of the two algorithms, experimental results show that the hybrid SFANC-FxNLMS algorithm can achieve a rapid response time, a low noise reduction error, and a high degree of robustness. Zhengding Luo, Dongyuan Shi, Woon-Seng Gan |
IEEE Signal Process. Lett. | 3 |
| 2022 | Optimal Penalty Factor for the MOV-FxLMS Algorithm in Active Noise Control SystemabstractThe minimum output variance filtered reference least mean square (MOV-FxLMS) algorithm is a effective algorithm that utilizes the penalty mechanism to help the active noise control (ANC) system achieve noise cancellation with constrained output variance or power. As it can constrain output power, the MOV-FxLMS algorithm can freely determine the ANC system’s control effort, avoiding output saturation, and improving system stability. However, its performance is determined by a penalty factor, which is normally chosen by trial and error. Hence, this work proposes an optimal penalty factor and its feasible estimation that does not require any assumptions of Gaussian reference signal or input independence. This factor assists the MOV-FxLMS in achieving the optimal solution under the target output-variance constraint. Numerical simulations on measured paths demonstrate its effectiveness for various types of noise. Dongyuan Shi, Woon-Seng Gan, Bhan Lam, Xiaoyi Shen |
IEEE Signal Process. Lett. | 2 |
| 2022 | Autonomous In-Situ Soundscape Augmentation via Joint Selection of Masker and GainabstractThe selection of maskers and playback gain levels in an in-situ soundscape augmentation system is crucial to its effectiveness in improving the overall acoustic comfort of a given environment. Traditionally, the selection of appropriate maskers and gain levels has been informed by expert opinion, which may not be representative of the target population, or by listening tests, which can be time- and labor-intensive. Furthermore, the resulting static choices of masker and gain are often inflexible to dynamic real-world soundscapes. In this work, we utilized a deep learning model to perform joint selection of the optimal masker and its gain level for a given soundscape. The proposed model was designed with highly modular building blocks, allowing for an optimized inference process that can quickly search through a large number of masker-gain combinations. In addition, we introduced the use of feature-domain soundscape augmentation conditioned on the digital gain level, eliminating the computationally expensive waveform-domain mixing process during inference, as well as the tedious gain adjustment process required for new maskers. The proposed system was evaluated on a large-scale dataset of subjective responses to augmented soundscapes with 442 participants, with the best model achieving a mean squared error of${0.122}\mathbf {\pm }{0.005}$on pleasantness score, validating the ability of the model to predict combined effect of the masker and its gain level on the perceptual pleasantness level. The proposed system thus allows in-situ or mixed-reality soundscape augmentation to be performed autonomously with near real-time latency while continuously accounting for changes in acoustic environments. Karn Watcharasupat, Kenneth Ooi, Bhan Lam, Trevor Wong, Zhen-Ting Ong, Woon-Seng Gan |
IEEE Signal Process. Lett. | 6 |
| 2022 | SALSA: Spatial Cue-Augmented Log-Spectrogram Features for Polyphonic Sound Event Localization and DetectionabstractSound event localization and detection (SELD) consists of two subtasks, which are sound event detection and direction-of-arrival estimation. While sound event detection mainly relies on time-frequency patterns to distinguish different sound classes, direction-of-arrival estimation uses amplitude and/or phase differences between microphones to estimate source directions. As a result, it is often difficult to jointly optimize these two subtasks. We propose a novel feature calledSpatial cue-Augmented Log-SpectrogrAm(SALSA) with exact time-frequency mapping between the signal power and the source directional cues, which is crucial for resolving overlapping sound sources. The SALSA feature consists of multichannel log-spectrograms stacked along with the normalized principal eigenvector of the spatial covariance matrix at each corresponding time-frequency bin. Depending on the microphone array format, the principal eigenvector can be normalized differently to extract amplitude and/or phase differences between the microphones. As a result, SALSA features are applicable for different microphone array formats such as first-order ambisonics (FOA) and multichannel microphone array (MIC). Experimental results on the TAU-NIGENS Spatial Sound Events 2021 dataset with directional interferences showed that SALSA features outperformed other state-of-the-art features. Specifically, the use of SALSA features in the FOA format increased the F1 score and localization recall by$6 \,\%$each, compared to the multichannel log-mel spectrograms with intensity vectors. For the MIC format, using SALSA features increased F1 score and localization recall by$16 \,\%$and$7 \,\%$, respectively, compared to using multichannel log-mel spectrograms with generalized cross-correlation spectra. Thi Ngoc Tho Nguyen, Karn Watcharasupat, Ngoc Khanh Nguyen 0003, Douglas L. Jones, Woon-Seng Gan |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | A General Network Architecture for Sound Event Localization and Detection Using Transfer Learning and Recurrent Neural NetworkabstractPolyphonic sound event detection and localization (SELD) task is challenging because it is difficult to jointly optimize sound event detection (SED) and direction-of-arrival (DOA) estimation in the same network. We propose a general network architecture for SELD in which the SELD network comprises sub-networks that are pre-trained to solve SED and DOA estimation independently, and a recurrent layer that combines the SED and DOA estimation outputs into SELD outputs. The recurrent layer does the alignment between the sound classes and DOAs of sound events while being unaware of how these outputs are produced by the upstream SED and DOA estimation algorithms. This simple network architecture is compatible with different existing SED and DOA estimation algorithms. It is highly practical since the sub-networks can be improved independently. The experimental results using the DCASE 2020 SELD dataset show that the performances of our proposed network architecture using different SED and DOA estimation algorithms and different audio formats are competitive with other state-of-the-art SELD algorithms. The source code for the proposed SELD network architecture is available at Github1. Thi Ngoc Tho Nguyen, Ngoc Khanh Nguyen 0003, Huy Phan, Lam Pham, Kenneth Ooi, Douglas L. Jones, Woon-Seng Gan |
ICASSP | 7 |
| 2021 | A Wireless Reference Active Noise Control Headphone Using Coherence Based Selection TechniqueabstractFeedforward active noise control (ANC) is widely utilized to attenuate the broadband noise picked up by the reference microphone. However, in some situations, it is impractical to obtain a clean reference signal when the noise source is far away from the controller. Hence, we adopt a wireless reference microphone to pick up the reference signals around the noise sources. Furthermore, a coherence-based-selection algorithm is proposed to select the reference signals with high coherence. The proposed method improves the quality of the reference signals and the noise reduction performance of the ANC system. Numerical simulations and real-time experiments are conducted to validate the effectiveness of the proposed algorithm. Xiaoyi Shen, Dongyuan Shi, Woon-Seng Gan |
ICASSP | 3 |
| 2021 | Extracting Urban Sound Information for Residential Areas in Smart Cities Using an End-to-End IoT SystemabstractWith rapid urbanization comes the increase of community, construction, and transportation noise in residential areas. The conventional approach of solely relying on sound pressure-level information to decide on the noise environment and to plan out noise control and mitigation strategies is inadequate. This article presents an end-to-end Internet-of-Things (IoT) system that extracts real-time urban sound metadata using edge devices, providing information on the sound type, location and duration, rate of occurrence, loudness, and azimuth of a dominant noise in nine residential areas. The collected metadata on environmental sound is transmitted to and aggregated in a cloud-based platform to produce detailed descriptive analytics and visualization. Our approach in integrating different building blocks, namely, hardware, software, cloud technologies, and signal processing algorithms to form our real-time IoT system is outlined. We demonstrate how some of the sound metadata extracted by our system are used to provide insights into the noise in residential areas. A scalable workflow to collect and prepare audio recordings from nine residential areas to construct our urban sound data set for training and evaluating a location-agnostic model is discussed. Some practical challenges of managing and maintaining a sensor network deployed at numerous locations are also addressed. Ee-Leng Tan, Furi Andi Karnapi, Linus Junjia Ng, Kenneth Ooi, Woon-Seng Gan |
IEEE Internet Things J. | 5 |
| 2021 | Comb-partitioned frequency-domain constraint adaptive algorithm for active noise control
Dongyuan Shi, Woon-Seng Gan, Bhan Lam, Xiaoyi Shen |
Signal Process. | 2 |
| 2021 | Fast Adaptive Active Noise Control Based on Modified Model-Agnostic Meta-Learning AlgorithmabstractWith the advent of efficient low-cost processors and electroacoustic components, there is renewed interest in the practical implementation of active noise control (ANC). However, the slow convergence of conventional adaptive algorithms deployed in ANC restricts its handling of typical amplitude-varying noise. Hence, we proposed a modified model-agnostic, meta-learning (MAML) strategy to obtain an initial control filter, which accelerates an adaptive algorithm's convergence when dealing with different types of amplitude-varying low-frequency noise. Numerical simulations with measured paths and real noise sources demonstrate its convergence acceleration efficacy in practical scenarios. Dongyuan Shi, Woon-Seng Gan, Bhan Lam, Kenneth Ooi |
IEEE Signal Process. Lett. | 2 |
| 2021 | Optimal Output-Constrained Active Noise Control Based on Inverse Adaptive Modeling Leak Factor EstimateabstractOutput saturation, mainly caused by the power amplifier, is a critical issue influencing the performance and stability of an adaptive system, such as in active noise control. In this paper, a quadratically constrained quadratic program (QCQP) is defined to achieve optimal control under the averaging-output-power constraint, which ensures the output of the system operates linearly and hence, avoids the output saturation. To solve this QCQP problem recursively in practice, this paper utilizes one of the leaky-based filtered-x least mean square algorithm with an optimal leak factor. However, this method only can be applied when the statistical feature of the control signal with maximum output-power is known, which is difficult to obtain in practice. Hence, by incorporating the adaptive inverse modeling technique, we can derive a practical estimation of the optimal leaky factor, which is applicable to different noise types. Furthermore, as the optimal output-constraint control forces the output to operate linearly, the nonlinear amplifier model is not required for the leak factor estimate. The simulation of the proposed algorithm is carried out on measured nonlinear paths to validate its efficacy. Dongyuan Shi, Woon-Seng Gan, Bhan Lam, Shulin Wen, Xiaoyi Shen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | A Sequence Matching Network for Polyphonic Sound Event Localization and DetectionabstractPolyphonic sound event detection and direction-of-arrival estimation require different input features from audio signals. While sound event detection mainly relies on time-frequency patterns, direction-of-arrival estimation relies on magnitude or phase differences between microphones. Previous approaches use the same input features for sound event detection and direction-of-arrival estimation, and train the two tasks jointly or in a two-stage transfer-learning manner. We propose a two-step approach that decouples the learning of the sound event detection and directional-of-arrival estimation systems. In the first step, we detect the sound events and estimate the directions-of-arrival separately to optimize the performance of each system. In the second step, we train a deep neural network to match the two output sequences of the event detector and the direction-of-arrival estimator. This modular and hierarchical approach allows the flexibility in the system design, and increase the performance of the whole sound event localization and detection system. The experimental results using the DCASE 2019 sound event localization and detection dataset show an improved performance compared to the previous state-of-the-art solutions. Thi Ngoc Tho Nguyen, Douglas L. Jones, Woon-Seng Gan |
ICASSP | 3 |
| 2020 | Multichannel Active Noise Control with Spatial Derivative Constraints to Enlarge the Quiet ZoneabstractActive noise control is an efficient approach in dealing with unwanted acoustic disturbances. However, most of the active noise control algorithms aim to control the signal of the error sensor leading to local noise attenuation only around the error microphones. One of the approaches to enlarge the quiet zone is by restraining the derivative of the sound field around the error microphone to zero. It achieves noise cancellation not only at the error microphone but also in its vicinity. This paper proposes a time-domain adaptive algorithm to implement the spatial derivative constraint, which avoids the complex analytic acoustic calculations. Furthermore, the proposed algorithm does not require extra microphones to acquire the sound field information during control. Numerical simulations are carried out to validate the effectiveness of the proposed method. Dongyuan Shi, Bhan Lam, Shulin Wen, Woon-Seng Gan |
ICASSP | 4 |
| 2020 | An Improved Selective Active Noise Control Algorithm Based on Empirical Wavelet TransformabstractThe gradual adaptation and possibility of divergence have been the two main obstacles in the efficient implementation of conventional adaptive active noise control (ANC) to a wider range of applications. Selective ANC (SANC) has been proposed to rapidly reduce noise by selecting a pre-trained control filter for different primary noise detected, and improve the robustness of the system. For stationary noise, considerable noise reduction performance and system stability are obtained by SANC. However, for non-stationary noise, in order to track the variability of the signal, frequency-band-match and selection have to be conducted constantly, which results in high computational burden. To confront this problem, empirical wavelet transform (EWT) is introduced to simplify the matching and selection of SANC in this paper. This EWT based SANC (SANC_EWT) algorithm extracts the first mode of random noises, and attenuates the noise immediately by picking the optimal pre-trained control filter labeled by the first boundary. Therefore, computational complexity is reduced drastically. Simulation results show that convergence could be reached rapidly. Better noise reduction performance is achieved by SANC_EWT compared to conventional FxLMS and SANC algorithms. Shulin Wen, Woon-Seng Gan, Dongyuan Shi |
ICASSP | 2 |
| 2020 | Robust Source Counting and DOA Estimation Using Spatial Pseudo-Spectrum and Convolutional Neural NetworkabstractMany signal processing-based methods for sound source direction-of-arrival estimation produce a spatial pseudo-spectrum of which the local maxima strongly indicate the source directions. Due to different levels of noise, reverberation and different number of overlapping sources, the spatial pseudo-spectra are noisy even after smoothing. In addition, the number of sources is often unknown. As a result, selecting the peaks from these spectra is susceptible to error. Convolutional neural network has been successfully applied to many image processing problems in general and direction-of-arrival estimation in particular. In addition, deep learning-based methods for direction-of-arrival estimation show good generalization to different environments. We propose to use a 2D convolutional neural network with multi-task learning to robustly estimate the number of sources and the directions-of-arrival from short-time spatial pseudo-spectra, which have useful directional information from audio input signals. This approach reduces the tendency of the neural network to learn unwanted association between sound classes and directional information, and helps the network generalize to unseen sound classes. The simulation and experimental results show that the proposed methods outperform other directional-of-arrival estimation methods in different levels of noise and reverberation, and different number of sources. Thi Ngoc Tho Nguyen, Woon-Seng Gan, Rishabh Ranjan, Douglas L. Jones |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Feedforward Selective Fixed-Filter Active Noise Control: Algorithm and ImplementationabstractConventional real-time active noise control (ANC) usually employs the adaptive filtered-x least mean square (FxLMS) algorithm to approach optimum coefficients for the control filter. However, lengthy training is usually required, and the perceived noise reduction is not immediately realized. Motivated by the practical implementation, we propose a selective fixed-filter active noise control (SFANC) algorithm, which selects a pretrained control filter to attenuate the detected primary noise rapidly. On top of improved robustness, the complexity analysis reveals that SFANC appears to be more efficient. The SFANC algorithm chooses the most suitable control filter based on the frequency-band-match approach implemented in a partitioned frequency-domain filter. Through simulations, SFANC is shown to exhibit a satisfactory response time and steady-state noise reduction performance, even for time-varying noise and real non-stationary disturbance. Dongyuan Shi, Woon-Seng Gan, Bhan Lam, Shulin Wen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Practical Implementation of Multichannel Filtered-x Least Mean Square Algorithm Based on the Multiple-Parallel-Branch With Folding Architecture for Large-Scale Active Noise ControlabstractMultichannel active noise control (MCANC) is widely recognized as an effective and efficient solution for acoustic noise and vibration cancellation, such as in high-dimensional ventilation ducts, open windows, and mechanical structures. The feedforward multichannel filtered-x least mean square (FFMCFxLMS) algorithm is commonly used to dynamically adjust the transfer function of the multichannel controllers for different noise environments. The computational load incurred by the FFMCFxLMS algorithm, however, increases exponentially with increasing channel count, thus requiring high-end field-programmable gate array (FPGA) processors. Nevertheless, such processors still need specific configurations to cope with soaring computing loads as the channel count increases. To achieve a high-efficiency implementation of the FFMCFxLMS algorithm with floating-point arithmetic, a novel architecture based on multiple-parallel-branch with folding (MPBF) technique is proposed. This architecture parallelizes the branches and reuses the multiplier and adder in each folded branch so that the tradeoff between throughput and the usage of the hardware resources is balanced. The proposed architecture is validated in an experimental setup that implements the FFMCFxLMS algorithm for the MCANC system with 24 reference sensors, 24 secondary sources, and 24 error sensors, at a sampling and throughput rates of 25 kHz and 260 Mb/s, respectively. Dongyuan Shi, Woon-Seng Gan, Jianjun He 0001, Bhan Lam |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | Parametric Hear through Equalization for Augmented Reality AudioabstractAugmented Reality (AR) audio applications require headphones to be acoustically transparent so that real sounds can pass through unaltered for natural fusion with virtual sounds. In this paper, we consider a multiple source scenario for hear through (HT) equalization (EQ) using closed-back circumaural headsets. AR headset prototype (described in our previous study) is used to capture real sounds from external microphones and compute the directional HT filters using adaptive filtering. This method is best suited for single source scenarios as one best filter corresponding to the estimated source direction is optimally used for HT filtering. In this paper, we propose parametric HT EQ for multiple-source scenarios in time-frequency domain by estimating a sub-band Direction of Arrival (DOA) using neural networks (NN) and selecting the corresponding HT filters from a pre-computed database. Objective analysis using spectral difference (SD) is used to evaluate the performance of different HT EQ filters with open ear scenario used as a reference. Using dummy head measurements with bandlimited pink noise and real source signals, it was found that the proposed integrated system significantly improves the performance over the conventional HT system in multiple source scenarios. Rishabh Ranjan, Jianjun He 0001, Woon-Seng Gan |
ICASSP | 4 |
| 2019 | Analysis of Multichannel Virtual Sensing Active Noise Control to Overcome Spatial Correlation and Causality ConstraintsabstractThis paper revisits the virtual sensing active noise control (VS-ANC) technique and extends it to a general multichannel ANC (MCANC) implementation. A frequency domain analysis shows that the multichannel virtual sensing ANC (VS-MCANC) technique arrives at an optimal control filter to cancel the noise disturbance at the virtual locations and overcomes the spatial correlation and causality constraints between the physical microphone and the virtual microphone. A real-time control of broadband noise with a 4-channel VS-MCANC implemented in a test chamber validates its theoretical analysis and demonstrates its active control effectiveness. Dongyuan Shi, Bhan Lam, Woon-Seng Gan |
ICASSP | 3 |
| 2019 | Optimal Leak Factor Selection for the Output-Constrained Leaky Filtered-Input Least Mean Square AlgorithmabstractThe leaky filtered-input least mean square (LFxLMS) algorithm is widely used in active noise control applications to minimize the degradation of attenuation performance due to output saturation distortion. However, the leak factor, which is critical in determining the steady-state error and robustness of the algorithm, is usually selected through trial and error. This letter proposes a leak factor selection approach, which ensures the LFxLMS algorithm converges to its optimal solution under the average-output-power constraint and can be readily derived in practice. Both broadband and narrowband cases are considered in the derivation without the independence assumption, and the simulations are conducted based on real primary and secondary paths to verify its effectiveness. Dongyuan Shi, Bhan Lam, Woon-Seng Gan, Shulin Wen |
IEEE Signal Process. Lett. | 3 |
| 2018 | A Novel Selective Active Noise Control Algorithm to Overcome Practical Implementation IssueabstractSelective active noise control (SANC) is a method to select a pre-trained control filter for different primary noises, instead of using conventional real-time computation of the control filter coefficients. SANC has the advantage of improving the robustness of control filter while reducing the computational complexity. This paper presents a practical strategy in choosing a suitable control filter based on the frequency-band-match mechanism implemented in a partitioned frequency domain filter structure. Both simulation and real-time experiment are carried out validate the noise reduction performance of the SANC compared to the conventional FxLMS algorithm. Dongyuan Shi, Bhan Lam, Woon-Seng Gan |
ICASSP | 3 |
| 2017 | Fast HRFT measurement system with unconstrained head movements for 3D audio in virtual and augmented reality applicationsabstractBinaural audio plays an indispensable role in virtual reality (VR) and augmented reality (AR). Binaural audio recreates the sensation of the three dimensional auditory experience using Head- Related Transfer Functions (HRTFs). HRTFs are as unique as our fingerprint. To achieve an immersive audio experience, HRTFs measured from every particular user is required. Nowadays, the conventional methods for HRTF measurements requires a wellcontrolled environment, hardly any movement of the user, and projecting to the user a high level of unpleasant sound in a rather long duration. Such difficulties have greatly limited the use of individually measurement HRTFs and hinder the authenticity of immersive audio. To solve these problems, we proposed a fast and convenient HRTF measurement system that is an order of magnitude faster and more importantly, it does not place any constraints on the user's movement. With the help of a head-tracker and advanced adaptive signal processing algorithms, this system is able to achieve satisfactory HRTF measurement accuracy. In this demonstration, we will present a fast real-time HRTF acquisition system and show how the individualized HRTFs improve the audio experience in VR/AR applications. Nguyen Duy Hai, Nitesh Kumar Chaudhary, Santi Peksi, Rishabh Ranjan, Jianjun He 0001, Woon-Seng Gan |
ICASSP | 6 |
| 2017 | Multiple parallel branch with folding architecture for multichannel filtered-x least mean square algorithmabstractMultichannel active noise control (MCANC) systems are commonly used in acoustic noise or vibration control, such as large-dimension ventilation ducts, open windows and mechanical structures. However, its computational load far exceeds the capabilities of digital signal processors (DSPs) and microcontrollers. Even the field programmable gate array (FPGA) cannot straightforwardly cope with the exponential increase in the computation load of MCANC systems. A novel architecture, called the multiple parallel branch with folding, is proposed for the J × J × M (J reference microphones, J secondary sources and Merror microphones) MCANC implementation with the floating-point arithmetic. This architecture addresses the tradeoff between throughput and hardware resource consumption by using parallel execution and folding. The proposed architecture is validated in an experimental setup that carries out a 4 × 4 × 4 multichannel filtered-x least mean square (FxLMS) algorithm achieving the sampling rate and throughput of 24 KHz and 18.4 Mbps, respectively. Dongyuan Shi, Jianjun He 0001, Chuang Shi, Tatsuya Murao, Woon-Seng Gan |
ICASSP | 5 |
| 2016 | Fast continuous HRTF acquisition with unconstrained movements of human subjectsabstractHead related transfer function (HRTF) is widely used in 3D audio reproduction, especially over headphones. Conventionally, HRTF database is acquired at discrete directions and the acquisition process is time-consuming. Recent works have been proposed to improve HRTF acquisition efficiency via continuous acquisition. However, these HRTF acquisition techniques still require subject to sit still (with limited head movement) in a rotating chair. In this paper, we further relax the head movement constraint during acquisition by using a head tracker. The proposed continuous HRTF acquisition technique relies on the activation based normalized least-mean-square (ANLMS) algorithm to extract HRTF on the fly. Experimental results validated the accuracy of the proposed technique, when compared with the standard static acquisition technique. Jianjun He 0001, Rishabh Ranjan, Woon-Seng Gan |
ICASSP | 3 |
| 2016 | Comparison of different development kits and its suitability in signal processing educationabstractWith the availability of many low-cost programmable development kits in the market, real-time signal processing projects can now be readily introduced into today's signal processing and embedded system course curriculum. In this paper, we group these popular development kits in terms of their cost, hardware architecture, development methodology and software resource. We further illustrate the programming efforts in implementing a real-time digital signal processing algorithm using different types of programmable development kits. Dongyuan Shi, Woon-Seng Gan |
ICASSP | 2 |
| 2015 | Multi-shift principal component analysis based primary component extraction for spatial audio reproductionabstractIn spatial audio analysis-synthesis, one of the key issues is to decompose a signal into primary and ambient components based on their spatial features. Principal component analysis (PCA) has been widely employed in primary component extraction, and shifted PCA (SPCA) is employed to enhance the primary extraction for input signals involving the inter-channel time difference. However, SPCA generally requires the primary components to come from one direction and cannot produce good results in the case of multiple directions. To solve this problem, we propose multi-shift PCA (MSPCA) by extending SPCA to multiple shifts. Two structures of MSPCA with different weighting methods are discussed. From the results of our simulations and listening tests, the proposed consecutive MSPCA with proper weighting is found to be superior to the conventional PCA and SPCA based primary extraction methods. Jianjun He 0001, Woon-Seng Gan |
ICASSP | 2 |
| 2015 | On the preprocessing and postprocessing of HRTF individualization based on sparse representation of anthropometric featuresabstractIndividualization of head-related transfer functions (HRTFs) can be realized using the person's anthropometry with a pretrained model. This model usually establishes a direct linear or non-linear mapping from anthropometry to HRTFs in the training database. Due to the complex relation between anthropometry and HRTFs, the accuracy of this model depends heavily on the correct selection of the anthropometric features. To alleviate this problem and improve the accuracy of HRTF individualization, an indirect HRTF individualization framework was proposed recently, where HRTFs are synthesized using a sparse representation trained from the anthropometric features. In this paper, we extend their study on this framework by investigating the effects of different preprocessing and postprocessing methods on HRTF individualization. Our experimental results showed that preprocessing and postprocessing methods are crucial for achieving accurate HRTF individualization. Jianjun He 0001, Woon-Seng Gan, Ee-Leng Tan |
ICASSP | 2 |
| 2015 | A virtual bass system with improved overflow controlabstractThe virtual bass system (VBS) can enhance the bass performance of small or flat loudspeakers by tricking the human brain to perceive the fundamental frequency from its higher harmonics. However, additional harmonics may lead to arithmetic overflow and cause distortion due to clipping, especially during high-level transient components. Past research pay little attention on this problem, and manual control of VBS gain settings is required to prevent overflow. Users need to manually adjust the gain settings for different sound tracks, which can be very troublesome. In this paper, we propose a VBS that can efficiently prevent the overflow problem by automatically controlling the gain settings for additional harmonics. This new approach pre-computes the gain limitation for additional harmonics and can be adopted for real-time audio implementation. Objective measurements are carried out to compare the proposed method with the commonly used limiter method. Hao Mu, Woon-Seng Gan |
ICASSP | 2 |
| 2015 | A hybrid speaker array-headphone system for immersive 3D audio reproductionabstractSpatial sound systems aim at rendering realistic sound experience to the listeners with uniform sound fields in the entire listening area. Today with the advancement of multichannel surround sound techniques, such systems are being practically realized, especially, at theatres, lecture halls, auditoriums, etc. Current practices, which are most widely used as home theatre systems, are based on multichannel stereophony, like 5.1, 10.2 and higher surround channel system. These systems require multiple loudspeakers to be placed in fixed configuration but often constrained by the room size. Sound reproduction systems like wave field synthesis (WFS) based on principle of natural propagation of sound waves, can create replica of true sound field uniformly over an extended listening area. However, WFS based systems too require hundreds of densely spaced loudspeakers enclosing the listener area and thus, difficult to realize in homes. In this paper, we introduce a new hybrid system by combining the WFS and binaural synthesis over headphones (based on active noise control techniques) to reduce the need of installing loudspeakers everywhere in a living room. Rishabh Ranjan, Woon-Seng Gan |
ICASSP | 2 |
| 2015 | Primary-Ambient Extraction Using Ambient Phase Estimation with a Sparsity ConstraintabstractSpatial audio reproduction addresses the growing commercial need to recreate an immersive listening experience of digital media content, such as movies and games. Primary-ambient extraction (PAE) is one of the key approaches to facilitate flexible and optimal rendering in spatial audio reproduction. Existing approaches, such as principal component analysis and time-frequency masking, often suffer from severe extraction error. This problem is more evident when the sound scene contains a relatively strong ambient component, which is frequently encountered in digital media. In this Letter, we propose a novel PAE approach by estimating the ambient phase with a sparsity constraint (APES). This approach exploits the equal magnitude of the uncorrelated ambient components in the two channels of a stereo signal and reformulates the PAE problem as an ambient phase estimation problem, which is then solved using the criterion that the primary component is sparse. Our experimental results demonstrate that the proposed approach significantly outperforms existing approaches, especially when the ambient component is relatively strong. Jianjun He 0001, Woon-Seng Gan, Ee-Leng Tan |
IEEE Signal Process. Lett. | 2 |
| 2015 | Primary-Ambient Extraction Using Ambient Spectrum Estimation for Immersive Spatial Audio ReproductionabstractThe diversity of today's playback systems requires a flexible, efficient, and immersive reproduction of sound scenes in digital media. Spatial audio reproduction based on primary-ambient extraction (PAE) fulfills this objective, where accurate extraction of primary and ambient components from sound mixtures in channel-based audio is crucial. Severe extraction error was found in existing PAE approaches when dealing with sound mixtures that contain a relatively strong ambient component, a commonly encountered case in the sound scenes of digital media. In this paper, we propose a novel ambient spectrum estimation (ASE) framework to improve the performance of PAE. The ASE framework exploits the equal magnitude of the uncorrelated ambient components in two channels of a stereo signal, and reformulates the PAE problem into the problem of estimating either ambient phase or magnitude. In particular, we take advantage of the sparse characteristic of the primary components to derive sparse solutions for ASE based PAE, together with an approximate solution that can significantly reduce the computational cost. Our objective and subjective experimental results demonstrate that the proposed ASE approaches significantly outperform existing approaches, especially when the ambient component is relatively strong. Jianjun He 0001, Woon-Seng Gan, Ee-Leng Tan |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Time-Shifting Based Primary-Ambient Extraction for Spatial Audio ReproductionabstractOne of the key issues in spatial audio analysis and reproduction is to decompose a signal into primary and ambient components based on their directional and diffuse spatial features, respectively. Existing approaches employed in primary-ambient extraction (PAE), such as principal component analysis (PCA), are mainly based on a basic stereo signal model. The performance of these PAE approaches has not been well studied for the input signals that do not satisfy all the assumptions of the stereo signal model. In practice, one such case commonly encountered is that the primary components of the stereo signal are partially correlated at zero lag, referred to as the primary-complex case. In this paper, we take PCA as a representative of existing PAE approaches and investigate the performance degradation of PAE with respect to the correlation of the primary components in the primary-complex case. A time-shifting technique is proposed in PAE to alleviate the performance degradation due to the low correlation of the primary components in such stereo signals. This technique involves time-shifting the input signal according to the estimated inter-channel time difference of the primary component prior to the signal decomposition using conventional PAE approaches. To avoid the switching artifacts caused by the varied time-shifting in successive time frames, overlapped output mapping is suggested. Based on the results from our experiments, PAE approaches with the proposed time-shifting technique are found to be superior to the conventional PAE approaches in terms of extraction accuracy and spatial accuracy. Jianjun He 0001, Woon-Seng Gan, Ee-Leng Tan |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | An Objective Analysis Method for Perceptual Quality of a Virtual Bass SystemabstractDue to the physical size and frequency response constraints of miniaturized and flat panel loudspeakers, low frequency reproduction from these loudspeakers is generally limited and unsatisfactory. The virtual bass system (VBS) enhances the bass performance of such loudspeakers by tricking the human brain to perceive the fundamental frequency from its higher harmonics. Problematically, the additional harmonics from VBS can also introduce perceptual distortion and reduce the audio quality. Therefore, a reliable method to assess the quality of VBS-enhanced signals is necessary in the designing of VBS. Since subjective experiments are often time-consuming and may be inconsistent, it is desirable to develop an objective assessment method for VBS. Earlier studies only utilized some simple objective metrics, which generally do not consider the human auditory model and are unable to accurately predict the perceptual quality of VBS. In this paper, we introduce a perceptual quality-assessment method for VBS based on the model output variables (MOVs) of the ITU Recommendation ITU-R BS.1387. Suitable combinations of MOVs are selected to derive perceptual quality metrics that correlate closely to the subjective quality. A verification experiment is presented to justify the accuracy of the metrics. Hao Mu, Woon-Seng Gan, Ee-Leng Tan |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Natural Listening over Headphones in Augmented Reality Using Adaptive Filtering TechniquesabstractAugmented reality (AR), which composes of virtual and real world environments, is becoming one of the major topics of research interest due to the advent of wearable devices. Today, AR is commonly used as assistive display to enhance the perception of reality in education, gaming, navigation, sports, entertainment, simulators, etc. However, most of the past works have mainly concentrated on the visual aspects of AR. Auditory events are one of the essential components in human perceptions in daily life but the augmented reality solutions have been lacking in this regard till now compared to visual aspects. Therefore, there is a need of natural listening in AR systems to give a holistic experience to the user. A new headphones configuration is presented in this work with two pairs of binaural microphones attached to headphones (one internal and one external microphone on each side). This paper focuses on enabling natural listening using open headphones employing adaptive filtering techniques to equalize the headset such that virtual sources are perceived as close as possible to sounds emanating from the physical sources. This would also require a superposition of virtual sources with the physical sound sources, as well as ambience. Modified versions of the filtered-x normalized least mean square algorithm (FxNLMS) are proposed in the paper to converge faster to the optimum solution as compared to the conventional FxNLMS. Measurements are carried out with open structure type headphones to evaluate their performance. Subjective test was conducted using individualized binaural room impulse responses (BRIRs) to evaluate the perceptual similarity between real and virtual sounds. Rishabh Ranjan, Woon-Seng Gan |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | A study on the frequency-domain Primary-ambient extraction for stereo audio signalsabstractPrimary-ambient extraction (PAE) has been playing an important role in spatial audio analysis-synthesis. Based on the spatial features, PAE decomposes a signal into primary and ambient components, which are then rendered separately. PAE is performed in subband domain for complex input signals having multiple point-like sound sources. However, the performance of PAE approaches and their key influences for such signals have not been well-studied so far. In this paper, we conducted a study on frequency-domain PAE using principal component analysis (PCA) in the case of multiple sources. We found that the partitioning of the frequency bins is very critical in PAE. Simulation results reveal that the proposed top-down adaptive partitioning method achieves superior performance as compared to the conventional partitioning methods. Jianjun He 0001, Woon-Seng Gan, Ee-Leng Tan |
ICASSP | 2 |
| 2014 | Fast and efficient real-time GPU based implementation of wave field synthesisabstractWave Field Synthesis (WFS) aims to replicate true sound field in an extended listening area with the help of loudspeaker arrays. WFS practical setups are heavily computational, as they need to drive many loudspeakers to accurately render multiple virtual sources. Thus, performance bottleneck occurs due to the sequential implementation on PCs with few cores. In addition, real-time spatial audio reproduction systems like WFS are subjected to hard real-time constraints, limiting system throughput and require cascading of several PCs to improve performance. In this paper, a fast and efficient graphics processing unit (GPU) based implementation of WFS is proposed to enhance the system throughput by extracting maximum data parallelism in the algorithm. The proposed method, implemented on NVidia C2075 GPU, uses block based partitioning approach to achieve peak system throughput of 1,400 Msamples per second, while rendering up to 200 real-time sound sources. Rishabh Ranjan, Woon-Seng Gan |
ICASSP | 2 |
| 2014 | New feedback active noise control system with improved performanceabstractIn many practical active noise control (ANC) applications, the undesired primary noise consists of multiple harmonic-related tones, which can be effectively reduced by internal model control (IMC) feedback ANC systems. In this paper, a new feedback ANC system is proposed with improved convergence rate and noise reduction when the frequency separation between two adjacent tones is small. In the proposed system, adaptive notch filter (ANF) is used to estimate the frequencies of the multi-tone noise, based on which the reference signals are internally generated and regrouped to increase the frequency separation in each channel. Compared with the conventional IMC feedback ANC system, the proposed system converges faster, has a better tracking capability when the frequencies of the primary noise vary, and is also less sensitive to impulsive noise. Computer simulations are conducted for both synthesized tonal noise and motorcycle engine noise to show the improved performance of the proposed system. Tongwei Wang, Woon-Seng Gan, Sen M. Kuo |
ICASSP | 2 |
| 2014 | Stochastic analysis of FXLMS-based internal model control feedback active noise control systems
Tongwei Wang, Woon-Seng Gan |
Signal Process. | 2 |
| 2014 | Linear Estimation Based Primary-Ambient Extraction for Stereo Audio SignalsabstractAudio signals for moving pictures and video games are often linear combinations of primary and ambient components. In spatial audio analysis-synthesis, these mixed signals are usually decomposed into primary and ambient components to facilitate flexible spatial rendering and enhancement. Existing approaches such as principal component analysis (PCA) and least squares (LS) are widely used to perform this decomposition from stereo signals. However, the performance of these approaches in primary-ambient extraction (PAE) has not been well studied and no comparative analysis among the existing approaches has been carried out so far. In this paper, we generalize the existing approaches into a linear estimation framework. Under this framework, we propose a series of performance measures to identify the components that contribute to the extraction error. Based on the generalized linear estimation framework and our proposed performance measures, a comparative study and experimental testing of the linear estimation based PAE approaches including existing PCA, LS, and three proposed variant LS approaches are presented. Jianjun He 0001, Ee-Leng Tan, Woon-Seng Gan |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Time-shifted principal component analysis based cue extraction for stereo audio signalsabstractIn spatial audio analysis-synthesis, one of the key issues is to decompose a signal into cue and ambient components based on their spatial features. Principal component analysis (PCA) has been widely employed in cue extraction. However, the performance of PCA based cue extraction is highly dependent on the assumptions of the input signal model. One of these assumptions is the input signal contains highly correlated cue at zero lag. However, this assumption is often unmet. To overcome this problem, time shifted PCA is proposed in this paper, which involves time-shifting the input signal according to the estimated inter-channel time difference (ITD) of the input signal before cue extraction. From our simulation and listening tests results, the proposed method is found to be superior to the conventional PCA based cue extraction method. Jianjun He 0001, Ee-Leng Tan, Woon-Seng Gan |
ICASSP | 3 |
| 2013 | A timbre matching approach to enhance audio quality of psychoacoustic bass enhancement systemabstractSmall and flat loudspeakers usually result in poor low-frequency (or bass) responses. Conventional gain equalization does not help significantly and may even result in overdriving and distortion. A psychoacoustic approach has been found to be suitable in tricking the human ear to perceive the fundamental frequency from its higher harmonics. Past research efforts have generally focused on weighting the harmonics based on the loudness matching method, but no work on timbre matching has been carried out so far. In this paper, we propose a new timbre matching technique, which can improve the sound quality of the psychoacoustically enhanced bass. This approach adjusts the amplitude of harmonics to produce similar timbre as the original audio content. Objective and subjective tests are carried out to compare the audio quality of the psychoacoustic bass enhanced signal using the equal-loudness weighting and the timbre matching methods. Hao Mu, Woon-Seng Gan, Ee-Leng Tan |
ICASSP | 2 |
| 2013 | A psychoacoustical preprocessing technique for virtual bass enhancement of the parametric loudspeakerabstractThe parametric loudspeaker is a novel type of loudspeaker that can project a directional sound beam. It is commonly used in creating personal sound zone and projecting private messages to a targeted audience. However, the parametric loudspeaker possesses a very poor bass (or low-frequency) response due inherently to the nonlinear acoustic principle generating sound from ultrasound in air. A psychoacoustic signal processing method known as “virtual bass” has been successfully implemented in some consumer electronics with miniature or flat loudspeaker unit, aiming to enhance their bass performances. In this paper, we adapt this “virtual bass“ approach for parametric loudspeakers. Unlike conventional loudspeakers, the parametric loudspeaker brings in an added degree of complexity in “virtual bass” enhancement due to its inherent nonlinear acoustic property. Accordingly, a new preprocessing technique is proposed for the parametric loudspeaker to psychoacoustically reproduce the low-frequency components within an octave below its cut-off frequency. Chuang Shi, Hao Mu, Woon-Seng Gan |
ICASSP | 3 |
| 2013 | Convergence Analysis of Narrowband Feedback Active Noise Control System With Imperfect Secondary Path EstimationabstractIn many practical active noise control (ANC) applications, feedback structure using estimated secondary path to synthesize reference signal is preferred under various conditions. This paper analyzes the convergence behavior of the narrowband feedback ANC systems with imperfect secondary path estimation. Existing approaches do not include the analysis of the reference signal synthesis errors due to its interrelated feedback nature. In this paper, the reconstruction error is modeled using the secondary path estimation error. Using this model, the effects of estimation errors on the convergence of the feedback ANC system is investigated. To further examine the effects of error in the filtered- x and filtered- y signal paths, these two paths are analyze separately to isolate the effects caused by these paths. Computer simulations are conducted to verify the theoretical analysis presented in the paper. Liang Wang 0007, Woon-Seng Gan, Andy W. H. Khong, Sen M. Kuo |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | A psychoacoustic bass enhancement system with improved transient and steady-state performanceabstractPsychoacoustic bass (low frequency) enhancement approach has received strong interests from consumer electronics manufacturers who demand good bass quality from small or flat panel speakers. Due to physical limitations and cost constraints, good bass performance is lacking from such speakers. A typical solution is based on the psychoacoustic phenomenon known as the “missing fundamental”, whereby human auditory system can perceive the fundamental frequency from its higher harmonics. This psychoacoustic bass enhancement system generally uses nonlinear devices (NLD) or phase vocoders (PV) to generate harmonics that enhance bass virtually. However, both approaches have their strengths and weaknesses. This paper presents a hybrid system, which combines these two approaches, to overcome their drawbacks and achieve good bass performance. The new approach first separates musical signals into transient and steady-state components, and applies NLD and PV on the separated signals. MUSHRA subjective test is used to evaluate the bass effect and audio quality of the proposed hybrid system against the NLD and PV methods. Hao Mu, Woon-Seng Gan, Ee-Leng Tan |
ICASSP | 2 |
| 2012 | Fixed-point square rootsabstractSquare root (SQRT) is a common arithmetic operation used in many DSP algorithms. In this paper, we evaluate square rooting methods suitable for implementation on fixed-point (FxP) DSP processors with a fast multiplying unit. The finite wordlength effect on the square rooting methods is highlighted, and it is shown that the theoretically derived convergence rate for the Newton-Raphson (NR) based square rooting methods are not suitable for FxP processor. Also, the most efficient methods for 8-bit and 16-bit FxP processors are identified. Abhishek Seth, Woon-Seng Gan |
ICASSP | 2 |
| 2012 | Psychoacoustic hybrid active noise control systemabstractIn conventional active noise control (ANC) system, the primary noise is attenuated over the frequency band of interest based simply on the error signal, which does not take into account the human perception of noise. Hence, researchers have developed psychoacoustic ANC system to improve its noise reduction performance from psychoacoustic point of view. However, in this psychoacoustic ANC system, there may be a disturbance that is uncorrelated with the primary noise at the error sensor, which can severely degrade the system's noise reduction performance. Hence, in this paper, a psychoacoustic hybrid ANC system is proposed, which can simultaneously control both the correlated primary noise and uncorrelated disturbance from psychoacoustic point of view. Loudness is used as the psychoacoustic criterion for evaluating the system's noise reduction performance. Simulation results show the effectiveness of the proposed psychoacoustic hybrid ANC system. Tongwei Wang, Woon-Seng Gan, Yong Kim Chong |
ICASSP | 2 |
| 2012 | Model-Free Iterative Learning Control for Repetitive Impulsive Noise Using FFT
Yali Zhou, Yixin Yin, Qizhi Zhang 0002, Woon-Seng Gan |
ISNN (2) | 4 |
| 2011 | Versatile and portable DSP platform for learning embedded signal processingabstractThis paper presents a versatile and portable digital signal processing (DSP) platform that is highly suitable for learning embedded signal processing anywhere and anytime. This DSP platform is based on the Texas Instruments VC5505 eZDSP USB Stick. We outline some of the important features in this development tool, such as the internal fast Fourier transform (FFT) hardware accelerator and the programmable high-speed codec that can be used in learning real-time embedded signal processing. This portable and easy-to-use platform extends the real-time DSP activities beyond the traditional laboratory environment. We highlight some project examples that used this platform. Woon-Seng Gan, Abhishek Seth, Sen M. Kuo |
ICASSP | 1 |
| 2011 | A Comparative Analysis of Preprocessing Methods for the Parametric Loudspeaker Based on the Khokhlov-Zabolotskaya-Kuznetsov Equation for Speech ReproductionabstractBased on the Berktay's farfield solution, various preprocessing methods were proposed to reduce the distortion of the highly directional audible signal in the parametric loudspeaker. However, the Berktay's farfield solution is an approximated model of nonlinear acoustic propagation. To determine the effectiveness of these methods, we analyze various preprocessing methods theoretically for directional speech reproduction using the Khokhlov-Zabolotskaya-Kuznetsov (KZK) equation, which provides a more accurate model of nonlinear acoustic propagation. In order to reduce the distortion effectively in the parametric loudspeaker with these preprocessing methods, the initial sound pressure level of the carrier frequency is found to be less than 132 dB according to the KZK equation. Unlike the Berktay' farfield solution that results in a +12 dB/octave gain slope, different gain slopes are derived using the KZK equation and appropriate equalizers are proposed to improve the frequency response of the parametric loudspeaker. The optimal preprocessing method for directional speech reproduction is established based on the KZK equation, which has a relatively flat frequency response of the desired speech signal and the best total harmonic distortion performance. Peifeng Ji, Ee-Leng Tan, Woon-Seng Gan, Jun Yang 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2011 | Time-Reversal Approach to the Stereophonic Acoustic Echo Cancellation ProblemabstractStereophonic acoustic echo cancellation (SAEC) plays an important role in delivering realistic teleconferencing experience. The fundamental problem of SAEC system is that stereophonic channels are linearly related and this results in slow convergence of the adaptive filters. In this paper, we present a novel algorithm by employing a selective time-reversal block to solve the SAEC problem which results in a significant increase in the convergence performance of adaptive filters such that the stereophonic image as well as quality are preserved. The proposed algorithm employs time-reversal operation on selective blocks of input data samples for one of the two channels to decorrelate stereophonic channels in the SAEC system. To achieve good stereophonic perception, time-reversal operation is only applied to the selective blocks whose magnitudes fall below a pre-determined threshold. Theoretical and numerical simulation results are also studied and investigated to show that the proposed algorithm achieves faster convergence in terms of normalized misalignment and better stereophonic perception with less audio distortion compared to the well-known nonlinear transformation algorithm for the SAEC system. Dinh-Quy Nguyen, Woon-Seng Gan, Andy W. H. Khong |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Generalized harmonic analysis of Arc-Tangent Square Root (ATSR) nonlinear device for virtual bass systemabstractNowadays, portable devices demand small-sized and low-end loudspeaker, however, the physical acoustic bass (or low-frequency) reproduction is usually poor. Bass enhancement by equalization is not feasible, and may even overload or damage the loudspeaker. Bass enhancement for such low-end loudspeaker can be psychoacoustically accomplished by exploiting the missing fundamental phenomenon. These systems are known as virtual bass systems (VBS), and generally use nonlinear device (NLD) as the main processing block to generate harmonics for virtual pitch perception. One of the recently developed NLD is the Arc-Tangent Square Root (ATSR) function, which can be used to control the pitch perception by a set of parameters. Mathematical relation between the input and output of the NLD is derived under a single-tone analysis framework. A detailed study on how the parameters of this NLD affect the harmonic's decay pattern and intensity is presented in this paper. Nay Oo, Woon-Seng Gan, Wee-Tong Lim |
ICASSP | 2 |
| 2010 | Localization of acoustic source on solids: A linear predictive coding based algorithm for location template matchingabstractLocation template matching (LTM) is a source localization technique in solids that is robust to dispersion and multipath. This is possible since LTM compares the input with a database of signals made at known locations. With this in place, it is possible to employ LTM in situations where the surface of interest takes an irregular shape. However, one of the existing LTM approaches uses crosscorrelation to compare the input and the database. It should be noted that if any two of the known locations stored in the database are too close, the cross-correlation method may have difficulties differentiating between signals generated from the neighboring points. To address this, we propose an algorithm which employs the linear predictive coding (LPC) that takes into account the dominant frequencies of a received signal. Using this approach, we show that the proposed algorithm is able to improve LTM's source localization accuracy under a real environment in the context of source localization for a touch interface. XueXin Yap, Andy W. H. Khong, Woon-Seng Gan |
ICASSP | 3 |
| 2010 | Performance analysis on recursive single-sideband amplitude modulation for parametric loudspeakersabstractA highly directional speech signal can be generated using parametric loudspeaker. The generation of highly directional sound beam is due to the nonlinear interaction of amplitude-modulated ultrasound waves in air. However, severe distortion is also generated during the reproduction of directional speech and several preprocessing techniques based on the Berktay's farfield model have been proposed by researchers to reduce the distortion. In this paper, we carried out a thorough investigation on the analytical performance of the recursive single-sideband amplitude modulation (RSSB-AM) technique, which has been found to perform well for directional speech reproduction. Several important characteristics of the performance of the RSSB-AM are observed and optimal parameters of the RSSB-AM are also presented. Peifeng Ji, Woon-Seng Gan, Ee-Leng Tan, Jun Yang 0004 |
ICME | 2 |
| 2009 | Synthesis of Polynomial-based Nonlinear Device and Harmonic Shifting Technique for Virtual Bass SystemabstractLow frequency bandwidth limitation is a common problem faced by miniature speakers and highly-directional speakers that have relatively high cut-off frequency. As such, low frequencies cannot be effectively reproduced. One of the current methods in addressing this problem is to employ psychoacoustic signal processing based on the “missing fundamental phenomenon”. Nonlinear device is generally used to induce virtual pitch that enhances the low frequency perception. However, some of the difficulties in using nonlinear device include the need of precise adjustment of harmonics' magnitudes, harmonic order and its decay rate to achieve good perceived bass. In this paper, we propose two techniques on synthesis of polynomial-based nonlinear device and harmonic shifting by modulation in an attempt to overcome these difficulties. Real-Time implementation and listening tests were conducted to verify the effectiveness of the proposed algorithm. Wee-Tong Lim, Nay Oo, Woon-Seng Gan |
ISCAS | 3 |
| 2009 | Model-Free Control of Nonlinear Noise Processes Based on C-FLAN
Yali Zhou, Qizhi Zhang 0002, Woon-Seng Gan |
ISNN (4) | 4 |
| 2009 | Convergence Analysis of Narrowband Active Noise Equalizer System Under Imperfect Secondary Path EstimationabstractActive noise equalizer systems are used to adjust the noise level in an environment, based on the preference of retaining noise information. Several researches have been carried out to determine the maximum step size bond of narrowband active noise control system with perfect secondary path estimation without gain factor consideration. However, in practical environment, secondary path estimation error of the system exists. In this paper, a stochastic approach analysis is applied to determine the maximum step size of the system under imperfect secondary path estimation. Simulation results are conducted to verify the analysis. Results show that the gain factor, sampling frequency, and secondary path estimation errors are all major factors governing the maximum step size of the narrowband active noise equalizer system under imperfect secondary path estimation. Liang Wang 0007, Woon-Seng Gan |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | A lowcomplexity fast converging partial update adaptive algorithm employing variable step-size for acoustic echo cancellationabstractPartial update adaptive algorithms have been proposed as a means of reducing complexity for adaptive filtering. The MMax tap-selection is one of the most popular tap-selection algorithms. It is well known that the performance of such partial update algorithm reduces with reducing number of filter coefficients selected for adaptation. We propose a low complexity and fast converging adaptive algorithm that exploits the MMax tap-selection. We achieve fast convergence with low complexity by deriving a variable step-size for the MMax normalized least-mean-square (MMax-NLMS) algorithm using its mean square deviation. Simulation results verify that the proposed algorithm achieves higher rate of convergence with lower computational complexity compared to the NLMS algorithm. Andy W. H. Khong, Woon-Seng Gan, Patrick A. Naylor, Mike Brookes |
ICASSP | 2 |
| 2008 | Novel DORT Method in Non-Well-Resolved Scatterer CaseabstractThe decomposition of time reversal operator (DORT) method is an efficient technique to select focusing signal on the target in well-resolved scatterer case. The DORT method requires the measurement of the inter-element impulse responses and diagonalizes these responses to find eigenstructure of the medium. However, the DORT method is not able to produce satisfactory result for non-well-resolved scatterer case (i.e., two scatterers that are placed closely about 0.2lambda). Therefore, in this letter, a weighted least squares (WLS) algorithm is introduced in the DORT method to perform selective focusing in non-well-resolved scatterer case. This technique called the WLS-DORT method performs better spatial focusing resolution than the conventional DORT method and results in simpler computational complexity than the minimum variance beamforming method under practical implementation. Dinh-Quy Nguyen, Woon-Seng Gan |
IEEE Signal Process. Lett. | 2 |
| 2007 | On Delayless Architecture for the Normalized Subband Adaptive FilterabstractDelayless architecture for the recently proposed normalized subband adaptive filter (NSAF) is described and analyzed in this paper. The NSAF has a unique weight-control mechanism whereby error signals estimated in subbands are used to adapt a fullband filter. In the delayless architecture, we implement the subband weight adaptation in an auxiliary loop, and place only the fullband filter along the input signal path. By so doing, delay due to the filter banks is moved to the auxiliary loop out of the signal path, thereby making the algorithm attractive for applications where excessive signal path delay is intolerable. Simulation results demonstrate that the proposed delayless NSAF outperforms other delayless approaches in terms of convergence rate. Kong-Aik Lee, Woon-Seng Gan |
ICME | 2 |
| 2007 | Fast Arbitrary Resizing of Images in DCT DomainabstractLimited storage and bandwidth have long necessitated the use of images and videos in compressed formats. Resizing of images in compressed domains such as discrete cosine transform (DCT) are developed primarily to avoid the computation associated to decompression and compression for spatial techniques. An important application of this work is transcoding of image and video content to achieve reduction of spatial resolution. In this work, arbitrary resizing of images by factors of P/Q times R/S, where P, Q, R, and S are positive integers is considered. An efficient approach is derived by performing up-sampling and down-sampling simultaneously. All computations are carried out with N times N DCT blocks to simplify integration with block based coders such as JPEG and MPEG. Simulation results revealed that our proposed approach achieves significant reduction up to 50% of the computational cost and visually sharper images compared to existing methods. Ee-Leng Tan, Woon-Seng Gan, Meng-Tong Wong |
ICME | 2 |
| 2007 | New Equalizing Scheme of Active Noise Equalization System in Automobile CabinabstractIn this paper, two adapting schemes for gain factors of active noise equalizer in application of bass enhancement and engine noise cancellation are presented. The novel system can adaptively equalize the engine noise into a proper enhancement of the bass audio production in automobile cabin. Computer simulations are performed to verify the proposed schemes. In noise cancellation mode, a maximum of 3 dB bass enhancement can be achieved with 6 dB noise suppression. More than 6 dB of bass enhancement can be achieved in the bass enhancement mode. The results showed that the proposed system could be a promising approach for solving the bass audio reproduction, together with the noise control problems in automobile. Liang Wang 0007, Woon-Seng Gan, Yong Kim Chong |
ICME | 2 |
| 2007 | A Nonlinear ANC System with a SPSA-Based Recurrent Fuzzy Neural Network Controller
Qizhi Zhang 0002, Yali Zhou, Xiaohe Liu, Woon-Seng Gan |
ISNN (1) | 5 |
| 2006 | A Novel Approach to Bass Enhancement in Automobile CabinabstractAn audio enhancement system is presented to improve the sound quality, especially the bass reproduction in the automobile cabin. High quality audio reproduction in cabin can be difficult due to a number of factors, including the noise in the cabin, performance, size of the loudspeakers in the car. The proposed system uses digital signal processing techniques with the built-in car audio system to tune and make use of the engine noise. The main problems are the tracking of the engine noise and tune the noise to match the audio signal. In this paper, we propose a solution based on frequency sampling filters to track and extract the audio signal, and a multi-frequency active noise equalizer to tune the engine noise. The results showed that the proposed system could be a promising approach for solving the bass audio reproduction, together with the noise control problems in automobile Liang Wang 0007, Woon-Seng Gan, Yong Kim Chong, Sen M. Kuo |
ICASSP (3) | 2 |
| 2006 | Convergence Analysis of Narrowband Active Noise ControlabstractA narrowband feedforward active noise control (ANC) system consists of an adaptive filter excited by the sum of multiple sinusoids corresponding to the harmonic "frequencies of the primary noise. The convergence of this direct-form ANC system is dependent on the frequency separation between two adjacent sinusoids in the reference signal. The analysis also proves that the length of the adaptive filter required for achieving the same convergence speed decreases with an increase in frequency separation. Computer simulations are conducted to verify the analysis presented in the paper Sen M. Kuo, Ajay B. Puvvala, Woon-Seng Gan |
ICASSP (5) | 3 |
| 2006 | Efficient VLSI Architecture of Lifting-Based Wavelet Packet Transform for Audio and Speech ApplicationsabstractThis paper presents a novel VLSI architecture for discrete wavelet packet transform (DWPT). By exploiting the in-place nature of the DWPT algorithm, this architecture has an efficient pipeline structure to implement high-throughput processing. Folded architecture for lifting-based wavelet filters is proposed to compute wavelet butterflies in different groups simultaneously, at each decomposition level. Internal pipelining and by-pass mode are employed on each processing element to increase computation throughput and provide easy configuration for arbitrary decomposition, respectively. According to the comparison results, our proposed VLSI architecture is more efficient than previous proposed architectures in terms of arithmetic operations, storage requirement, and throughput Woon-Seng Gan |
ICASSP (3) | 2 |
| 2006 | Model-Free Control of a Nonlinear ANC System with a SPSA-Based Neural Network Controller
Yali Zhou, Qizhi Zhang 0002, Woon-Seng Gan |
ISNN (2) | 4 |
| 2006 | A digital beamsteerer for difference frequency in a parametric arrayabstractA steerable audio system can be realized using parametric array. However, the available steerable angle is often limited by the sampling interval used in the digital system. As such, the smallest steerable angle is large (/spl sim/26/spl deg/) for several hundred kilohertz of sampling frequency. Although there are some fractional delay or frequency domain algorithms can be used to improve the steering angle, most of the algorithms are either computational intensive or introduce error during the process. In this paper, an algorithm is proposed to rectify this problem by applying separate delays to the carrier and sideband frequencies. Different weighting functions also added to the carrier and sideband frequencies to control the difference frequency's beamwidth and sidelobe. Most importantly, the proposed system can steer the difference frequency to a small angle with minimal computation. Woon-Seng Gan, Jun Yang 0004, Khim Sia Tan, Meng Hwa Er |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Bandwidth-efficient recursive pth-order equalization for correcting baseband distortion in parametric loudspeakersabstractA bandwidth-efficient recursive implementation of pth-order equalization is developed in order to correct the inherent baseband distortion in parametric loudspeakers. Assuming that the nonlinear system is largely quadratic, a substitution of it can be made by using an analytic pair of all-pass sections cascaded with a squared magnitude block. The all-pass sections comprise of a delay block and an approximate linear-phase infinite impulse response (IIR) Hilbert transformer to give a Hilbert transform pair. In this way, the distortion terms produced are due to the difference frequencies only and therefore within the bandwidth of the original input signal. When used together with a single-sideband (SSB) amplitude modulation (SSB-AM) scheme, this new method allows the equalized SSB output signal to be faithfully reproduced to correct the distortion. Simulation results in this paper show that the new method is able to suppress residual in-band distortion components by -70 dB or lower. Kelvin Chee-Mun Lee, Woon-Seng Gan |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Reconfigurable Context-Sensitive Bio-Bridge Middleware for Smart Bio-LaboratoriesabstractRecent advances in genomic research and biotechnology have already led to an increased level of technological sophistication in today's bio-laboratories ("wet" labs), where complex experiments are now routinely carried out with various high-tech lab equipment. With the current proliferation of wireless technologies and the emergence of advanced embedded computing technologies, we envision that the wet lab can benefit greatly from smart technologies implemented under a scalable smart unwired access infrastructure. In this paper, we show that with our custom-designed Bio-Bridge (BB) device - a smart middleware that possesses: (a) a common interface platform for communication between multi-vendor bio-instruments and mobile devices; (b) next-generation middleware characteristics such as transparent computing and awareness functionalities; and (c) low-cost scalability and easy reconfiguration, we can provide a smart lab platform for operation in the common IEEE 802.1 Ib wireless network environment. We also propose novel solutions based on our BB devices to address two of the challenging research issues in the Smart Lab, namely: reconfiguration and context-sensitive computing. Wei-Khing For, Xiaoming Bao, Woon-Seng Gan, See-Kiong Ng |
AINA | 3 |
| 2005 | Experimental Investigation of Active Vibration Control Using a Filtered-Error Neural Network and Piezoelectric Actuators
Yali Zhou, Qizhi Zhang 0002, Woon-Seng Gan |
ISNN (3) | 4 |
| 2005 | Nonlinear least-square solution to flat-top pattern synthesis using arbitrary linear array
Yuan Wen, Woon-Seng Gan, Jun Yang 0004 |
Signal Process. | 2 |
| 2004 | An efficient digital beamsteering system for difference frequency in parametric arrayabstractFor a digital beamsteering system, the smallest time delay available is equal to the sampling period of the digital signal processing (DSP) board. As most of the time the sampling frequency is not high enough, the smallest steering angle available is large, which is undesirable. This limitation also occurs when performing beamsteering in a parametric array digitally. Although partial delay or frequency domain algorithms can be used to improve the steering angle, most of the algorithms are either computational intensive or introduce error during the process. In this paper, an algorithm is proposed to beamsteer the difference frequency in parametric array. The proposed system can be used to steer the difference frequency to a small angle, without the need to increase the sampling frequency or implement partial delay. Khim Sia Tan, Woon-Seng Gan, Jun Yang 0004, Meng Hwa Er |
ICASSP (2) | 2 |
| 2004 | Improving convergence of the NLMS algorithm using constrained subband updatesabstractWe propose a new design criterion for subband adaptive filters (SAFs). The proposed multiple-constraint optimization criterion is based on the principle of minimal disturbance, where the multiple constraints are imposed on the updated subband filter outputs. Compared to the classical fullband least-mean-square (LMS) algorithm, the subband adaptive filtering algorithm derived from the proposed criterion exhibits faster convergence under colored excitation. Furthermore, the recursive tap-weight adaptation can be expressed in a simple form comparable to that of the normalized LMS (NLMS) algorithm. We also show that the proposed multiple-constraint optimization criterion is related to another known weighted criterion. The efficacy of the proposed criterion and algorithm are examined and validated via mathematical analysis and simulation. Kong-Aik Lee, Woon-Seng Gan |
IEEE Signal Process. Lett. | 2 |
| 2003 | Constant beamwidth beamformer for difference frequency in parametric arrayabstractSound reproduction in air by using a parametric acoustic array has been investigated for a few decades. Two inaudible ultrasonic frequencies are produced from the parametric array. Due to the nonlinearity of air, it is possible to produce an audible frequency with its frequency equal to the difference in the two ultrasonic frequencies. However, there is not much work done in controlling the beam pattern of the difference frequency generated by the primary waves. In this paper, an algorithm is proposed to control the sidelobe level of the difference frequency directivity. By making use of array signal processing techniques, the algorithm is also capable of producing a constant beamwidth for broadband difference frequency. Khim Sia Tan, Woon-Seng Gan, Jun Yang 0004, Meng Hwa Er |
ICASSP (5) | 2 |
| 2003 | Constant beamwidth beamformer for difference frequency in parametric arrayabstractThe sound reproduction in air by using a parametric acoustic array [P.J. Westervelt, 1963] has been reported for a few decades. Two inaudible ultrasonic frequencies are produced from the parametric array. Due to the nonlinearity of air, it is possible to produce an audible frequency with its frequency equal to the difference in the two ultrasonic frequencies. However, there is not much work done in controlling the beam pattern of the difference frequency generated by the primary waves. In this paper, an algorithm is proposed to control the sidelobe level of the difference frequency directivity. By making use of array signal processing techniques, the algorithm is also capable of producing a constant beamwidth for broadband difference frequency. Khim Sia Tan, Woon-Seng Gan, Jun Yang 0004, Meng Hwa Er |
ICME | 2 |
| 1998 | Broadband active noise compressorabstractA broadband active noise compressor (ANCP) is presented to adapt the active noise equalization (ANE) technique suitable for practical usage. Compared to the existing broadband ANE system, the novel ANCP not only has the ability to shape the residual noise spectrum, but can also automatically adjust the residual noise power. This algorithm is analyzed in steady state and verified by computer simulations. Jin Wei Feng, Woon-Seng Gan |
IEEE Signal Process. Lett. | 2 |
| 1997 | A broadband self-tuning active noise equaliser
Jin Wei Feng, Woon-Seng Gan |
Signal Process. | 2 |
| 1996 | Fuzzy step-size adjustment for the LMS algorithm
Woon-Seng Gan |
Signal Process. | 1 |
| 1995 | Neural networks and multivariate currency forecastingabstractA neural network approach to multivariate currency forecasting is presented. The performance of this model is compared with a univariate currency model for the major currencies, the Swiss Franc; Deutschemark and the Yen. The multivariate currency model outperforms the univariate model in prediction for all three currencies for single-step and multi-step forecasting. Kah Hwa Ng, Woon-Seng Gan |
CIFEr | 2 |
| 1994 | Functional-link models for adaptive channel equaliserabstractThis paper presents a study of a new class of adaptive channel equaliser, known as the functional-link (FL) equaliser which utilises functional expansion of the equaliser's input data. By carefully selecting the appropriate functions of the input, significant performance improvement can be obtained. These studies will lead to significant reduction in the exponentially increased expansion of the polynomial-perceptron equaliser. Numerical simulation results are presented to highlight the better bit error rate performance of the FL based equaliser compared to other non-linear, neural network based equalisers.> Woon-Seng Gan, John J. Soraghan, Tariq S. Durrani |
ICASSP (3) | 1 |
| 1994 | Acoustical Chaotic Fractal Images for Medical Imaging
Woon-Seng Gan |
IPMU | 1 |
| 1992 | Development of fuzzy neural tool for medical signal processing and imagingabstractThe author proposes the use of fuzzy neural networks to improve the resolution of medical images and the segmentation of medical images. The backpropagation neural network is used to obtain an optimized membership function. The author works out the algorithms to implement the fuzzy neural networks for both types of application. Preliminary results are given. An advantage of using fuzzy neural networks compared with conventional neural networks is the reduction of the number of elements in each neural network layer. Thus, computation time can be reduced. Another advantage of using neural networks is the solution of the ill-posed problem in the universe scattering problem such as the divergence problem.> Woon-Seng Gan |
CBMS | 1 |
| 1992 | Stability analysis of the noncanonical LMS (NCLMS) algorithmabstractThe stability of the noncanonical least mean square (NCLMS) algorithm is investigated. The NCLMS effectively uses a different step size for each tap coefficient position during adaptation. The classical LMS step size bound cannot be directly applied to the NCLMS. The weight error vector is modeled as a first-order difference equation and a stability bound for the NCLMS is derived. Simulation results are presented to back up the analysis.> Woon-Seng Gan, John J. Soraghan, Robert W. Stewart, Tariq S. Durrani |
ICASSP | 1 |
| 1991 | The non-canonical LMS algorithm (NCLMS): characteristics and analysisabstractThe authors present analysis and simulations of an LMS (least mean square) based adaptive filtering algorithm called the NCLMS (non-canonical LMS). Rather than using the standard FIR (finite impulse response) filter as for the LMS algorithm, a modified structure called the NCFIR (non-canonical FIR) is used. The NCFIR allows a faster VLSI implementation than the conventional FIR. A comparison of the performances of the NCLMS and conventional LMS algorithm is presented for an inverse system modeling application. Simulation results are given which show a reduced EMSE (excess mean square error) level and an improved performance in an impulsive noise environment for the NCLMS over the LMS algorithm.> Woon-Seng Gan, John J. Soraghan, Robert W. Stewart, Tariq S. Durrani |
ICASSP | 1 |