Boon Poh Ng

dblp:82/6317 · DBLP profile ↗
← Back
35ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0001-7394-8503ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge
abstract
Large-scale Video Foundation Models (VFMs) have significantly advanced various video-related tasks, either through task-specific models or Multi-modal Large Language Models (MLLMs). However, the open accessibility of VFMs also introduces critical security risks, as adversaries can exploit full knowledge of the VFMs to launch potent attacks. This paper investigates a novel and practical adversarial threat scenario: attacking downstream models or MLLMs fine-tuned from open-source VFMs, without requiring access to the victim task, training data, model query, and architecture. In contrast to conventional transfer-based attacks that rely on task-aligned surrogate models, we demonstrate that adversarial vulnerabilities can be exploited directly from the VFMs. To this end, we propose the Transferable Video Attack (TVA), a temporal-aware adversarial attack method that leverages the temporal representation dynamics of VFMs to craft effective perturbations. TVA integrates a bidirectional contrastive learning mechanism to maximize the discrepancy between the clean and adversarial features, and introduces a temporal consistency loss that exploits motion cues to enhance the sequential impact of perturbations. TVA avoids the need to train expensive surrogate models or access to domain-specific data, thereby offering a more practical and efficient attack strategy. Extensive experiments across 24 video-related tasks demonstrate the efficacy of TVA against downstream models and MLLMs, revealing a previously underexplored security vulnerability in the deployment of video models.
Yi Yu 0011, Song Xia, Deepu Rajan, Boon Poh Ng, Alex Chichung Kot, Xudong Jiang 0001
AAAI6
2025 Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the Wild
Peijun Bao, Chenqi Kong, Siyuan Yang 0001, Zihao Shao, Xinghao Jiang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot
ICCV6
2024 Omnipotent Distillation with LLMs for Weakly-Supervised Natural Language Video Localization: When Divergence Meets Consistency
abstract
Natural language video localization plays a pivotal role in video understanding, and leveraging weakly-labeled data is considered a promising approach to circumvent the laborintensive process of manual annotations. However, this approach encounters two significant challenges: 1) limited input distribution, namely that the limited writing styles of the language query, annotated by human annotators, hinder the model’s generalization to real-world scenarios with diverse vocabularies and sentence structures; 2) the incomplete ground truth, whose supervision guidance is insufficient. To overcome these challenges, we propose an omnipotent distillation algorithm with large language models (LLM). The distribution of the input sample is enriched to obtain diverse multi-view versions while a consistency then comes to regularize the consistency of their results for distillation. Specifically, we first train our teacher model with the proposed intra-model agreement, where multiple sub-models are supervised by each other. Then, we leverage the LLM to paraphrase the language query and distill the teacher model to a lightweight student model by enforcing the consistency between the localization results of the paraphrased sentence and the original one. In addition, to assess the generalization of the model across different dimensions of language variation, we create extensive datasets by building upon existing datasets. Our experiments demonstrate substantial performance improvements adaptively to diverse kinds of language queries.
Peijun Bao, Zihao Shao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot
AAAI4
2024 Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video Grounding
abstract
This paper for the first time leverages multi-modal videos for weakly-supervised temporal video grounding. As labeling the video moment is labor-intensive and subjective, the weakly-supervised approaches have gained increasing attention in recent years. However, these approaches could inherently compromise performance due to inadequate supervision. Therefore, to tackle this challenge, we for the first time pay attention to exploiting complementary information extracted from multi-modal videos (e.g., RGB frames, optical flows), where richer supervision is naturally introduced in the weaklysupervised context. Our motivation is that by integrating different modalities of the videos, the model is learned from synergic supervision and thereby can attain superior generalization capability. However, addressing multiple modalities† would also inevitably introduce additional computational overhead, and might become inapplicable if a particular modality is inaccessible. To solve this issue, we adopt a novel route: building a multi-modal distillation algorithm to capitalize on the multi-modal knowledge as supervision for model training, while still being able to work with only the single modal input during inference. As such, we can utilize the benefits brought by the supplementary nature of multiple modalities, without compromising the applicability in practical scenarios. Specifically, we first propose a cross-modal mutual learning framework and train a sophisticated teacher model to learn collaboratively from the multi-modal videos. Then we identify two sorts of knowledge from the teacher model, i.e., temporal boundaries and semantic activation map. And we devise a local-global distillation algorithm to transfer this knowledge to a student model of single-modal input at both local and global levels. Extensive experiments on large-scale datasets demonstrate that our method achieves state-of-the-art performance with/without multi-modal inputs.
Peijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot
AAAI4
2024 E3M: Zero-Shot Spatio-Temporal Video Grounding with Expectation-Maximization Multimodal Modulation
Peijun Bao, Zihao Shao, Wenhan Yang, Boon Poh Ng, Alex Chichung Kot
ECCV (83)4
2023 Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event Localization
abstract
This paper for the first time explores audio-visual event localization in an unsupervised manner. Previous methods tackle this problem in a supervised setting and require segment-level or video-level event category ground-truth to train the model. However, building large-scale multi-modality datasets with category annotations is human-intensive and thus not scalable to real-world applications. To this end, we propose cross-modal label contrastive learning to exploit multi-modal information among unlabeled audio and visual streams as self-supervision signals. At the feature representation level, multi-modal representations are collaboratively learned from audio and visual components by using self-supervised representation learning. At the label level, we propose a novel self-supervised pretext task i.e. label contrasting to self-annotate videos with pseudo-labels for localization model training. Note that irrelevant background would hinder the acquisition of high-quality pseudo-labels and thus lead to an inferior localization model. To address this issue, we then propose an expectation-maximization algorithm that optimizes the pseudo-label acquisition and localization model in a coarse-to-fine manner. Extensive experiments demonstrate that our unsupervised approach performs reasonably well compared to the state-of-the-art supervised methods.
Peijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot
AAAI3
2023 A Hybrid Differential Detection Scheme for the Ultra-Wideband Orientational Beamforming System
abstract
Differential space-time coding methods have been investigated for ultra-wideband communication systems to avoid channel estimation. However, their performance in the line-of-sight (LOS) environment is worse than the recently proposed orientational beamforming (OBF) system under low signal-to-noise ratios (SNRs). On the other hand, the OBF system cannot work well in the non-LOS (NLOS) environment. To address these issues, a hybrid differential detection (HDD) scheme is proposed in this paper, which combines the OBF system with a proposed differential OBF (DOBF) system. First, the DOBF system with a fully differential detection scheme is proposed, whose performance is almost the same as the differential space-time block coding method. However, its detection complexity only linearly increases with the number of transmitting antennas. Then, the HDD scheme is proposed, with its decision statistic being a combination of a modified OBF decision statistic and the DOBF decision statistic. In the NLOS environment, the HDD scheme is reduced to the DOBF system. In the LOS environment, a combination coefficient is obtained by mapping an SNR-related factor through a sigmoid function. The optimal parameters for the sigmoid function are determined through simulations, with which the proposed HDD scheme can achieve overall better performance than the OBF and DOBF systems.
Jiangyan Han, Boon Poh Ng, Meng Hwa Er
IEEE Trans. Wirel. Commun.2
2022 An Adaptive Orientational Beamforming Technique for Narrowband Interference Rejection
abstract
In this paper, we investigate and extend the linearly constrained minimum variance (LCMV) algorithm for conventional wideband beamforming system to the recently proposed orientational beamforming (OBF) system. An orientational LCMV (O-LCMV) algorithm is proposed. It is constructed on the orientation dimension, and an orientational constraint instead of directional constraint is used to guarantee an orientational gain for the desired signal. This perfectly solves the problem of rejecting an interference arriving from the same direction as the desired signal, which current adaptive directional beamforming algorithms cannot handle. Numerous simulations show that the O-LCMV algorithm for the OBF system works effectively regardless of the number of narrowband interferences (NBIs) or their DOAs if the NBIs have the same center frequency. As the number of different center frequencies of the NBIs increases, the performance degrades slightly.
Jiangyan Han, Boon Poh Ng, Meng Hwa Er
ICASSP2
2022 Adaptive orientational beamforming techniques for narrowband interference rejection
Jiangyan Han, Boon Poh Ng, Meng Hwa Er
Signal Process.2
2021 Learning Sparsifying Transforms for Image Reconstruction in Electrical Impedance Tomography
abstract
Electrical Impedance Tomography (EIT) is a fast and non-invasive imaging technology that reconstructs the internal electrical properties of a subject. However, its functionality is limited by low spatial resolution arising from an ill-posed and ill-conditioned inverse problem. Several sparsity-promoting regularization methods have been applied to improve the quality of EIT image reconstruction, including various ℓ0and ℓ1-based analytical models (TV, TwIST, etc.), and a patch-based sparse representation via a learned dictionary (using the K-SVD algorithm), dubbed CS-EIT. To further exploit the potential of compressed sensing in Electrical Impedance Tomography, this paper incorporates the recent novel method of transform learning for EIT image reconstruction. We propose a blind compressed sensing algorithm, dubbed TL-EIT, which simultaneously optimizes the sparsifying transform and updates the reconstructed image. We demonstrate using both synthetic and in vivo data that the proposed TL-EIT is more effective than other sparsity-based algorithms for reconstructing high-quality EIT images. In addition, TL-EIT also accelerates the reconstruction process in comparison to other learning-based algorithms like CS-EIT.
Kaiyi Yang, Narong Borijindargoon, Boon Poh Ng, Saiprasad Ravishankar, Bihan Wen
ICASSP3
2020 A flexible model for edge representation in digital imagery
Dhimas Arief Dharmawan, Boon Poh Ng
Signal Process.2
2019 Residual U-Net for Retinal Vessel Segmentation
abstract
In recent years, the influence of deep learning on retinal vessel segmentation has grown rapidly. Most of the available deep learning based methods use relatively shallow structures. However, due to the limited representative capacity, shallow networks will restrain deep learning models to segment both vessel and non-vessel pixels accurately. In this paper, we propose a residual U-Net for retinal vessel segmentation. Our network has several advantages. First, the network uses a new residual block structure. In the new structure, batch normalization layers are placed before the activation unit to achieve better performance and accelerate the convergence. Also, a dropout layer is utilized in the structure to alleviate over-fitting problems. Second, the depth of the network is increased by adding more residual blocks and strong dropouts which then allow the network to extract features better. Fundus images from the publicly available DRIVE and STARE datasets are used to evaluate the proposed network. Experimental result shows that the proposed modified residual U-Net has better performance than existing state-of-the-art algorithms.
Di Li 0006, Dhimas Arief Dharmawan, Boon Poh Ng, Susanto Rahardja
ICIP3
2019 Matrix Ill-Condition Analysis in Spectrum Recovery of DPCA HRWS SAR Imaging
abstract
Conventional single-platform spaceborne synthetic aperture radar has a conflict between wide imaging swath width and fine-cross-range resolution. Displaced phase center antennas is one method to overcome this limitation. For spectrum recovery of nonuniform sampling, a detailed analysis on the relationship between coefficient matrix singularity and the number of effective phase centers (EPCs) passed by the platform during one pulse repetition time (PRT) is given. When the distance passed by the platform during one PRT is multiple of the distance between two EPCs plus a small value, the coefficient matrix is singularity. An explicit expression of the coefficient matrix is also derived.
Changzheng Ma, Xianyang Hu, Tat Soon Yeo, Boon Poh Ng, Qun Zhang 0001
IEEE Geosci. Remote. Sens. Lett.4
2019 Aperiodic geometry design for DOA estimation of broadband sources using compressive sensing
Zeeshan Asghar Sayed, Boon Poh Ng
Signal Process.2
2019 Design of Optimal Adaptive Filters for Two-Dimensional Filamentary Structures Segmentation
abstract
The problems of filamentary structures segmentation encompass retinal vessel detection in fundus images, reconstruction and tracing of neurons in microscopic images and segmentation of human vasculatures in two-dimensional digital angiography. In this letter, we focus on retinal vessel segmentation which is an important and challenging problem among others. Many retinal vessel segmentation algorithms have been developed, where most of them rely on the filtering technique. However, the available filters suffer from undesirable responses at some vessel and non-vessel structures. Thus, we propose a new framework based on optimal adaptive filters for retinal vessel segmentation. The filter coefficients are obtained by solving multiple inverse problems under a regularization framework. The performance of the proposed framework is evaluated on fundus images from the DRIVE and STARE datasets. In the experimental section, we show that the proposed framework outperforms all of the compared methods in four indicators namely sensitivity, F1-score, G-mean and Mathews correlation coefficient.
Dhimas Arief Dharmawan, Boon Poh Ng, Narong Borijindargoon
IEEE Signal Process. Lett.2
2017 Orientational Beamforming For UWB Signals
abstract
In this paper, we propose a new beamforming scheme for pulse-based signals, particularly, for ultra wideband pulses, by employing a random array geometry and an orientation matching of the arrays. Unlike conventional directional beamforming, the introduced scheme forms an array beam, being independent of a signal impinging direction for both near-fields and far-fields. As a consequence, it is advantageous to suppress unwanted interferences having same frequency contents and coexisting in the same arrival direction with reference to the desired signal. The proposed beamforming concept is analyzed together with the definitions of the -3dB beamwidth and restlobe level. Besides, the orientational beampatterns are compared with the directional beampatterns, and observed that the array responses are different according to the respective spatial processing domains.
Aye Su Yee, Boon Poh Ng
IEEE Trans. Commun.2
2015 Reduced-Complexity Super-Resolution DOA Estimation with Unknown Number of Sources
abstract
When the number of sources is inaccurately estimated, it is well-known that the conventional subspace-based super-resolution direction-of-arrival (DOA) estimation techniques provide inconsistent spatial spectrum, and hence the DOA estimates. In this work, we present a novel technique which provides resolution capability comparable with that of the super-resolution techniques. While the working principle of the proposed technique is similar to that of the minimum-norm algorithm, the algorithm is insensitive to the estimated model-order. Simulation studies show that the proposed technique is advantageous over the use of subspace-based techniques with the number of sources estimated by well-known model order estimation techniques.
Vinod V. Reddy, Mohamed Mubeen, Boon Poh Ng
IEEE Signal Process. Lett.3
2014 Filter-and-forward relay beamforming using output power minimization
abstract
In this paper, we consider designing the Filter-and-forward (FF) relay beamforming in frequency-selective channels using a new approach. The proposed approach aims to minimize the output power at the destination side while keeping the response to the desired signal at a constant level, and the problem is subject to both total and individual relay transmit power constraints. It is shown that the proposed beamforming design scheme is equivalent to the SINR maximization formulation, in terms of achieving the same output SINR. Despite the equivalence in performance, the proposed approach requires significantly lower computational load for solving the problem.
Meng Hwa Er, Boon Poh Ng
ICASSP3
2014 Unambiguous speech DOA estimation under spatial aliasing conditions
abstract
With the bandwidth of speech signals extending over several octaves, the spatial Nyquist criterion constrains the microphone array design. Violating this criterion by increasing microphone spacing in order to achieve high resolution introduces ambiguity in identifying the source directions due to the aliasing components. In this work, we investigate the effect of spatial aliasing on the direction-of-arrival (DOA) spectrum due to wideband sources. Noting that the extent of aliasing is frequency dependent, we propose a multi-stage scheme for speech DOA estimation following a subband decomposition. To observe the advantage of this scheme, we verify it with the steered minimum variance distortionless response (STMV) and approximate kernel density estimators. The performance is evaluated with simulations and recorded room impulse responses.
Vinod V. Reddy, Andy W. H. Khong, Boon Poh Ng
IEEE ACM Trans. Audio Speech Lang. Process.3
2012 Indoor contaminant source estimation using a multiple model unscented Kalman filter
Rong Yang 0002, Pek Hui Foo, Peng Yen Tan, Elaine Mei Eng See, Gee Wah Ng, Boon Poh Ng
FUSION6
2012 DOA estimation of amplitude modulated signals with less array sensors than sources
abstract
This paper addresses the Direction-of-Arrival (DOA) estimation problem for amplitude modulated signals whose number is more than that of the array sensors. The proposed method is based on an idea of virtual array. For source signals with amplitude modulation, such as binary phase shift keying (BPSK) and M-ary amplitude shift keying (M-ASK), we show that introducing in virtual array actually gives rise to processing the fourth-order moments of array output, which is related to higher-order statistics (HOS) techniques. While traditional HOS methods in array processing mainly exploit higher-order cumulants of the received data, we propose a DOA estimation method based on the fourth-order moments, which is of lower computational load than the fourth-order cumulants. Simulation results demonstrate the effectiveness of the proposed method for estimating DOAs of more source signals than array elements.
Boon Poh Ng, Meng Hwa Er
ICASSP2
2012 DOA estimation of wideband sources without estimating the number of sources
Vinod V. Reddy, Boon Poh Ng, Andy W. H. Khong
Signal Process.2
2010 Tracking an accelerated target with a nonlinear constant heading model
Rong Yang 0002, Gee Wah Ng, Boon Poh Ng
FUSION3
2010 Natural-ordered complex Hadamard transform
Aye Aung, Boon Poh Ng
Signal Process.2
2010 WTDM-Based M3H Filter for Target Tracking in the Presence of Outliers
abstract
This letter proposes a waiting-time-dependent semi-Markov (WTDM) switching based multiple model multiple hypothesis (M3H) filter to track a target in the presence of outliers. Two models, namely, a normal noise model and an outlier model, are constructed in the semi-Markov system. The adaptive transition probabilities are derived as the functions of the waiting time of outliers (or the interval of outlier occurrences), and this waiting time is treated as a discrete random variable governed by an exponential pmf. Performance of the WTDM-based M3H filter is demonstrated through simulation experiments. The proposed WTDM-based M3H filter outperforms the existing interacting multiple model (IMM) filter, which was proposed for the same purpose.
Rong Yang 0002, Boon Poh Ng, Gee Wah Ng
IEEE Signal Process. Lett.2
2009 Robust Adaptive Trimming for High-Resolution Direction Finding
abstract
The presence of impulsive noise can severely degrade the accuracy performance of conventional direction of arrival (DOA) estimation algorithms, such as MUSIC and ESPIRIT. We propose a two-stage robust adaptive trimming approach. We first apply Shapiro-Wilk's goodness-of-fitWtest for Gaussianity, as a preprocessing stage. We then robustly estimate the covariance matrix in order to minimize the impact of impulsive noise on conventional DOA estimation algorithms. Numerical simulations are presented to illustrate the efficacy of the proposed approach for high resolution direction finding in highly-impulsive environments.
Chin-Heng Lim, Chong Meng Samson See, Abdelhak M. Zoubir, Boon Poh Ng
IEEE Signal Process. Lett.4
2009 Multiple Model Multiple Hypothesis Filter With Sojourn-Time-Dependent Semi-Markov Switching
abstract
This letter suggests a maneuvering target tracking algorithm using sojourn-time-dependent semi-Markov (STDM) model switching system on the basis of multiple model multiple hypothesis (M3H) filter. In the M3H filter, the target model sequences are constructed by a normal Markov switching system. A set of fixed model transition probabilities is used throughout the whole Markov process. In this letter, we propose the STDM-based M3H filter, which adapts the transition probability to the system sojourn time in the current model. This adaptation makes the target model sequence closer to the target natural behavior, and leads to the better tracking performance. Simulation results are presented to demonstrate the performance improvement after the STDM being introduced in the M3H filter.
Rong Yang 0002, Boon Poh Ng, Gee Wah Ng
IEEE Signal Process. Lett.2
2008 Applications of the SRV constraint in broadband pattern synthesis
Huiping Duan, Boon Poh Ng, Chong Meng Samson See, Jun Fang 0001
Signal Process.2
2008 Pipelined Hardware Structure for Sequency-Ordered Complex Hadamard Transform
abstract
This letter presents a fast algorithm for the sequency-ordered complex Hadamard transform (SCHT) based on the decomposition method of decimation-in-sequency. To support high-speed real-time applications, a pipelined hardware structure is also proposed to deal with sequentially presented input/output data streams. This structure achieves a full hardware utilization and requires only complex adder/subtracters and complex data stores for an N-point SCHT.
Guoan Bi, Aye Aung, Boon Poh Ng
IEEE Signal Process. Lett.3
2007 Spatial Resolutions of the Broadband Nonredundant and Minimum Redundancy Arrays
abstract
Approximate formulations for the 3-dB beamwidth are derived in this letter to measure the spatial resolution of the broadband nonredundant array (NRA) and minimum redundancy array (MRA), which assume the ideal continuous-time, infinite-length filters with the frequency responses obtained by the linearly constrained minimum variance (LCMV) optimization. By these formulations, the beamwidths of NRA and MRA are compared with that of the uniform linear array (ULA). Moreover, the accuracy of the derived formulations is assessed by numerical studies.
Huiping Duan, Boon Poh Ng, Chong Meng Samson See, Jun Fang 0001
IEEE Signal Process. Lett.2
2006 A Stereo to Mono Dowmixing Scheme for MPEG-4 Parametric Stereo Encoder
abstract
In this paper, a signal-adaptive, stereo-to-mono downmixing scheme associated with MPEG-4 Parametric Stereo (PS) encoding is presented. The proposed scheme minimizes signal cancellation and coloration due to inter-channel phase misalignment. By using the inter-channel phase difference information which is calculated as a PS spatial parameter, the phase of the stereo signals are aligned. Subsequently, a simple averaging is carried out to mix the signals. The need to perform power equalization to preserve the overall power of the stereo signals in the downmix signal is eliminated. This leads to a significant saving of computational power. The scheme allows overall phase difference (OPD), one of the spatial parameter sent in PS bitstream, to be coded with minimum bit consumption. As a result, the entropy is reduced from 1.31 bits/symbol to 1 bit/symbol. This scheme is useful especially for stereo audio with a significant amount of side signal component.
Samsudin Ng, Evelyn Kurniawati, Boon Poh Ng, Farook Sattar, Sapna George
ICASSP (5)3
2006 A CFAR based model order selection criterion for complex sinusoids
Changzheng Ma, Boon Poh Ng
Signal Process.2
2005 A new broadband beamformer using IIR filters
abstract
In this letter, a new broadband beamformer using infinite impulse response (IIR) filters is proposed. The unique feature of our letter is that by replacing all the delay elements with tap-to-tap IIR filters under some structural restrictions, the Frost processor is naturally extended from the finite impulse response (FIR) beamformer to the IIR beamformer. In the new beamformer, the feedforward and feedback weights are computed with the constrained and unconstrained least-mean-squares algorithms, respectively. The stability of the IIR filters in the proposed beamformer could be monitored easily. Simulation results demonstrate the improved performance of the new IIR beamformer relative to the Frost beamformer.
Huiping Duan, Boon Poh Ng, Chong Meng Samson See
IEEE Signal Process. Lett.2
1997 Sensor array calibration using measured steering vectors of uncertain location
abstract
We present a maximum likelihood approach for calibrating sensor arrays in the presence of mutual coupling, channel gain and phase mismatch and array geometry uncertainties using measured steering vectors of uncertain locations. The estimated perturbation parameters is used to calibrate the array manifold, hence enabling many high resolution array processing algorithms to attain their potential advantages. We present two methods for optimizing the highly nonlinear and multimodal ML cost function. The first method is a linearized local gradient search algorithm. The second method is derived from combining the fast local search of gradient methods with the nonlinear global search ability of the genetic algorithm. The resulting hybrid optimizer is both fast and globally converging. Simulation results are presented to illustrate the usefulness of the proposed approach.
Chong Meng Samson See, Boon Poh Ng, Colin Cowan
ICASSP2
1994 A MUSIC approach for estimation of directions of arrival of multiple narrowband and broadband sources
Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot
Signal Process.1