Abdesselam Bouzerdoum

dblp:00/5421 · also Salim Bouzerdoum · DBLP profile ↗
← Back
132ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0002-9163-0045ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 74 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 57 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 NaviFormer: Multimodal scene segmentation for assistive navigation
abstract
Safe and independent navigation is an essential part of daily life for individuals with visual impairments. Recent advances in assistive navigation have leveraged deep learning-based semantic scene segmentation with promising results by using RGB images. In comparison, RGB-D images provide richer geometric and appearance information, but their integration into deep learning frameworks for assistive navigation could be further explored to enhance segmentation accuracy. In modeling cross-modal interactions, the existing RGB-D segmentation methods tend to lose fine-grained spatial details or discard useful information, which hinders performance on dense prediction tasks. To address these gaps, we propose NaviFormer, a novel RGB-D semantic segmentation architecture tailored for assistive navigation. NaviFormer features a dual-stream transformer encoder with shared weights to efficiently extract latent features from RGB and depth modalities. It also incorporates a Local-Global Cross-Modal Fusion module, which facilitates effective information exchange between the two modalities across both local and global feature levels. In training NaviFormer, we further employ a pixel-wise contrastive loss to enhance the separability of pixel-level embeddings in the RGB-D feature space. Extensive experiments on TrueSight and Cityscapes datasets indicate that NaviFormer achieves superior performance compared to existing RGB-D segmentation methods. Our findings highlight the importance of leveraging RGB-D data for enhancing semantic understanding in assistive navigation systems, and establish NaviFormer as a solid baseline for future research in this domain.
Ly Bui, Son Lam Phung, Yang Di, Soan Thi Minh Duong, Abdesselam Bouzerdoum
Comput. Vis. Image Underst.5
2026 CSN: A compact semantic segmentation network for visual scene perception in assistive navigation
abstract
Accuracy and efficiency are essential in assistive navigation algorithms to ensure accessibility and reliability for visually impaired individuals. However, existing deep learning models often require substantial computational resources to achieve high accuracy, making them impractical for deployment on mobile devices. To address this problem, we introduce CSN, a compact semantic segmentation network designed for assistive navigation, optimizing performance in resource-constrained environments. With CSN, we introduce two innovative modules, the cascaded atrous multi-scale enhancement (CAME) layer and the dual-path residual bottleneck (DPRB) block. The CAME layer efficiently enhances multi-scale representation through feature resampling, while the DPRB block improves feature refinement with minimal computational cost. These modules enable CSN to achieve robust and reliable segmentation across diverse and complex pedestrian environments. The proposed approach achieves the best result on the challenging TrueSight dataset, demonstrating superior prediction accuracy and computational efficiency compared to state-of-the-art lightweight models. CSN achieves a mean intersection over union of 60.99%, while maintaining a low computational cost of 83.46 giga floating-point operations, a compact model size of 8.41 million parameters, and a real-time inference speed of 52.59 frames per second.
Yunjia Lei, Son Lam Phung, Yang Di, Abdesselam Bouzerdoum
Comput. Vis. Image Underst.4
2026 A Lightweight Hybrid Gabor Deep Learning Approach and its Application to Medical Image Classification
abstract
Abstract Deep learning has revolutionized image analysis, but its applications are limited by the need for large datasets and high computational resources. Hybrid approaches that combine domain-specific, universal feature extractor with learnable neural networks offer a promising balance of efficiency and accuracy. This paper presents a hybrid model integrating a Gabor filter bank front-end with compact neural networks for efficient feature extraction and classification. Gabor filters, inherently bandpass, extract early-stage features with spatially shifted filters covering the frequency plane to balance spatial and spectral localization. We introduce separate channels capturing low- and high-frequency components to enhance feature representation while maintaining efficiency. The approach reduces trainable parameters and training time while preserving accuracy, making it suitable for resource-constrained environments. Compared to MobileNetV2 and EfficientNetB0, our model trains approximately 4–6 × faster on average while using fewer parameters and FLOPs. We compare it to pretrained networks used as feature extractors, lightweight fine-tuned models, and classical descriptors (HOG, LBP). It achieves competitive results with faster training and reduced computation. The hybrid model uses only around 0.60 GFLOPs and 0.34 M parameters, and we apply statistical significance testing (ANOVA, paired t-tests) to validate performance gains. Inference takes 0.01–0.02 s per image, up to 15 × faster than EfficientNetB0 and 8 × faster than MobileNetV2. Grad-CAM visualizations confirm localized attention on relevant regions. This work highlights integrating traditional features with deep learning to improve efficiency for resource-limited applications. Future work will address color fusion, robustness to noise, and automated filter optimization.
Rayyan Ahmed, Hamza Baali, Abdesselam Bouzerdoum
Int. J. Comput. Vis.3
2025 Improving perceptual quality in spatiotemporal timeseries forecasting
Shehel Yoosuf, Hamza Baali, Abdesselam Bouzerdoum
Eng. Appl. Artif. Intell.3
2025 AMT-Net: Attention-based multi-task network for scene depth and semantics prediction in assistive navigation
abstract
Traveling safely and independently in unfamiliar environments remains a significant challenge for people with visual impairments. Conventional assistive navigation systems, while aiming to enhance spatial awareness, typically handle crucial tasks like semantic segmentation and depth estimation separately, resulting in high computational overhead and reduced inference speed. To address this limitation, we introduce AMT-Net, a novel multi-task deep neural network designed for joint semantic segmentation and monocular depth estimation. The AMT-Net is designed with a single unified decoder, which boosts not only the model’s efficiency but also its scalability on portable devices with limited computational resources. We propose two self-attention-based modules, CSAPP and RSAB, to leverage the strengths of convolutional neural networks for extracting robust local features and Transformers for capturing essential long-range dependencies. This design enhances the ability of our model to interpret complex scenes effectively. Furthermore, AMT-Net has low computational complexity and achieves real-time performance, making it suitable for assistive navigation applications. Extensive experiments on the public NYUD-v2 dataset and the TrueSight dataset demonstrated our model’s state-of-the-art performance and the effectiveness of the proposed components.
Yunjia Lei, Joshua Luke Thompson, Son Lam Phung, Abdesselam Bouzerdoum, Hoang Thanh Le 0001
Neurocomputing4
2024 MSSP: A Multi-view Benchmark for Street Scene Perception in Assistive Navigation
abstract
Finding safe paths in autonomous or assistive navigation systems is a challenging task. In this paper, we introduce MSSP, a novel multi-perspective street scene perception benchmark dataset for navigation assistance. Compared to single-view perception, multi-view provides more comprehensive information about the surroundings to the pedestrians for enhancing safety. The MSSP dataset includes 2,044 samples from complex scenes in both first-person and third-person perspectives, covering four categories: Sidewalk, Traffic lane, Verge, and Lawn. This dataset primarily focuses on pedestrian navigation assistance and is the first pedestrian multi-perspective dataset. Furthermore, in comparison to single-perspective datasets in other studies, MSSP contains more categories. The dataset supports pixel-level semantic segmentation approaches. We also provide experimental results of recent semantic segmentation methods on this dataset for evaluation. Specifically, the transformer-based SegFormer method outperforms others on MSSP. Our dataset is accessible through: https://github.com/yangdi-cv/MSSP.
Yang Di, Son Lam Phung, Abdesselam Bouzerdoum
IJCNN3
2024 Bio-DETR: A Transformer-based Network for Pest and Seed Detection with Hyperspectral Images
abstract
Exotic pests and seeds pose a serious threat to agricultural production and ecosystems, leading to significant labour and economic losses. Therefore, automated pest and seed detection systems are crucial for biosecurity and agriculture. Biosecurity detection systems need to identify pests and seeds of various sizes and types in a variable and complex environment, making high-precision automated detection a challenging task. To address this, we propose Bio-DETR, a lightweight transformer-based architecture for accurate pest and seed detection. We introduce two self-attention-based modules, Hybrid Scale Attention and Dynamic Bilateral Attention, for enhanced feature extraction and multiscale information fusion. The effectiveness of these modules is validated experimentally. We also propose HSI-Bio, a large-scale dataset with 8,000 images across 23 categories, collected using a hyperspectral camera on diverse backgrounds. Compared to RGB images, hyperspectral images (HSI) offer rich channel information. The representative spectra are selected from HSI for experiments. Bio-DETR achieves an AP50of 87.4% and an AP of 62.2% on HSI-Bio, outperforming other state-of-the-art methods and achieving real-time detection of 52 FPS. Our code is available at: https://github.com/yangdi-cv/Bio-DETR.
Yang Di, Son Lam Phung, Julian Van Den Berg, Jason Clissold, Ly Bui, Hoang Thanh Le 0001, Abdesselam Bouzerdoum
IJCNN7
2024 UOW-Vessel: A Benchmark Dataset of High-Resolution Optical Satellite Images for Vessel Detection and Segmentation
abstract
In this paper, we introduce UOW-Vessel, a benchmark dataset of high-resolution optical satellite images for vessel detection and segmentation. Our dataset consists of 3,500 images, collected from 14 countries across 4 continents. With a total of 35,598 instances in 10 vessel categories, UOW-Vessel is to date the largest satellite image dataset for vessel recognition. Furthermore, compared to the existing public datasets that only provide bounding box ground-truth, our new dataset offers more accurate polygon annotations of vessel objects. This dataset is expected to support instance segmentation-based approaches, which is a less investigated area in vessel surveillance. We also report extensive evaluations of the recent algorithms for instance segmentation on the new benchmark dataset.
Ly Bui, Son Lam Phung, Yang Di, Hoang Thanh Le 0001, Tran Thanh Phong Nguyen, Sandy Burden, Abdesselam Bouzerdoum
WACV7
2024 Oversampling techniques for imbalanced data in regression
abstract
Our study addresses the challenge of imbalanced regression data in Machine Learning (ML) by introducing tailored methods for different data structures. We adapt K-Nearest Neighbor Oversampling-Regression (KNNOR-Reg), originally for imbalanced classification, to address imbalanced regression in low population datasets, evolving to KNNOR-Deep Regression (KNNOR-DeepReg) for high-population datasets. For tabular data, we also present the Auto-Inflater neural network, utilizing an exponential loss function for Autoencoders. For image datasets, we employ Multi-Level Autoencoders, consisting of Convolutional and Fully Connected Autoencoders. For such high-dimension data our approach outperforms the Synthetic Minority Oversampling Technique for Regression (SMOTER) algorithm for the IMDB-WIKI and AgeDB image datasets. For tabular data we conducted a comprehensive experiment using various models trained on both augmented and non-augmented datasets, followed by performance comparisons on test data. The outcomes revealed a positive impact of data augmentation, with a success rate of 83.75% for Light Gradient Boosting Method (LightGBM) and 71.57% for the 18 other regressors employed in the study. This success rate is determined by the frequency of instances where models performed better when augmented data was used compared to instances with no augmentation. Access to the comparative code can be found in GitHub.
Samir Brahim Belhaouari, Ashhadul Islam, Khelil Kassoul, Ala I. Al-Fuqaha, Abdesselam Bouzerdoum
Expert Syst. Appl.5
2024 DAP: A dataset-agnostic predictor of neural network performance
abstract
Training a deep neural network on a large dataset to convergence is a time-demanding task. This task often must be repeated many times, especially when developing a new deep learning algorithm or performing a neural architecture search. This problem can be mitigated if a deep neural network’s performance can be estimated without actually training it. In this work, we investigate the feasibility of two tasks: (i) predicting a deep neural network’s performance accurately given only its architectural descriptor, and (ii) generalizing the predictor across different datasets without re-training. To this end, we propose a dataset-agnostic regression framework that uses a novel dual-LSTM model and a new dataset difficulty feature. The experimental results show that both tasks above are indeed feasible, and the proposed method outperforms the existing techniques in all experimental cases. Additionally, we also demonstrate several practical use-cases of the proposed predictor.
Sui Paul Ang, Soan Thi Minh Duong, Son Lam Phung, Abdesselam Bouzerdoum
Neurocomputing4
2024 AdaptorNAS: A New Perturbation-Based Neural Architecture Search for Hyperspectral Image Segmentation
abstract
Hyperspectral image segmentation is an emerging area with numerous applications, including agriculture, forestry, environment monitoring, and remote sensing. This paper proposes a new neural architecture search algorithm, named AdaptorNAS, for hyperspectral image segmentation. AdaptorNAS aims to design the optimum decoder for any given encoder. In our approach, the search space of AdaptorNAS is a large deep neural network (DNN), and the optimal decoder is derived by pruning the large DNN via a perturbation-based pruning strategy. Verified on three popular encoders, i.e., ResNet-34, MobileNet-V2, and EfficientNet-B2, AdaptorNAS can design high-speed decoders that are significantly better than six common hand-crafted decoders. Additionally, with the EfficientNet-B2 encoder, AdaptorNAS (mIoU of 92.47% and mDice of 95.15%) outperforms the state-of-the-art NAS algorithms and hand-crafted network architectures on the hyperspectral image segmentation task. We also introduce a new hyperspectral image dataset of 4,625 images for objective evaluation in hyperspectral image segmentation research.
Sui Paul Ang, Son Lam Phung, Ly Bui, Abdesselam Bouzerdoum
IEEE Trans. Circuits Syst. Video Technol.4
2023 TP-YOLO: A Lightweight Attention-Based Architecture for Tiny Pest Detection
abstract
Automatic detection of agricultural pests is a challenging problem that is of great interest in biosecurity and precision agriculture. The detection model must cope well with the dense distribution of small-sized pests in complex backgrounds. This paper proposes a lightweight attention-based network, called TP-YOLO, for tiny pest detection. We introduce two attention-based components, namely Contextual Transformer and Omni-Dimensional Dynamic Convolution modules, to enhance feature extraction. The proposed modules are integrated into the YOLOv8 backbone, a state-of-the-art baseline for object detection. This paper also introduces a new benchmark dataset consisting of 1,600 images of Khapra beetles for objective evaluation of pest detection algorithms. Extensive experiments on two datasets indicate that TP-YOLO achieves competitive detection accuracy while having a significantly smaller model size and fast prediction time. We have made the code available to the public at: https://github.com/yangdi-cv/TP-YOLO.
Yang Di, Son Lam Phung, Julian Van Den Berg, Jason Clissold, Abdesselam Bouzerdoum
ICIP5
2023 MSD-NAS: multi-scale dense neural architecture search for real-time pedestrian lane detection
abstract
Abstract Accurate detection of pedestrian lanes is a crucial criterion for vision-impaired people to navigate freely and safely. The current deep learning methods have achieved reasonable accuracy at this task. However, they lack practicality for real-time pedestrian lane detection due to non-optimal accuracy, speed, and model size trade-off. Hence, an optimized deep neural network (DNN) for pedestrian lane detection is required. Designing a DNN from scratch is a laborious task that requires significant experience and time. This paper proposes a novel neural architecture search (NAS) algorithm, named MSD-NAS, to automate this laborious task. The proposed method designs an optimized deep network with multi-scale input branches, allowing the derived network to utilize local and global contexts for predictions. The search is also performed in a large and generic space that includes many existing hand-designed network architectures as candidates. To further boost performance, we propose a Short-term Visual Memory mechanism to improve information facilitation within the derived networks. Evaluated on the PLVP3 dataset of 10,000 images, the DNN designed by MSD-NAS achieves state-of-the-art accuracy (0.9781) and mIoU (0.9542), while being 20.16 times faster and 2.56 times smaller than the current best deep learning model.
Sui Paul Ang, Son Lam Phung, Soan Thi Minh Duong, Abdesselam Bouzerdoum
Appl. Intell.4
2022 Class Similarity Weighted Knowledge Distillation for Continual Semantic Segmentation
abstract
Deep learning models are known to suffer from the problem of catastrophic forgetting when they incrementally learn new classes. Continual learning for semantic segmentation (CSS) is an emerging field in computer vision. We identify a problem in CSS: A model tends to be confused between old and new classes that are visually similar, which makes it forget the old ones. To address this gap, we propose REMINDER - a new CSS framework and a novel class similarity knowledge distillation (CSW-KD) method. Our CSW-KD method distills the knowledge of a previous model on old classes that are similar to the new one. This provides two main benefits: (i) selectively revising old classes that are more likely to be forgotten, and (ii) better learning new classes by relating them with the previously seen classes. Extensive experiments on Pascal-Voc 2012 and ADE20k datasets show that our approach outperforms state-of-the-art methods on standard CSS settings by up to 7.07% and 8.49%, respectively.
Minh-Hieu Phan, The-Anh Ta, Son Lam Phung, Long Tran-Thanh, Abdesselam Bouzerdoum
CVPR5
2022 Efficient hyperspectral image segmentation for biosecurity scanning using knowledge distillation from multi-head teacher
Minh-Hieu Phan, Son Lam Phung, Khoa Luu, Abdesselam Bouzerdoum
Neurocomputing4
2022 Bayesian Gabor Network With Uncertainty Estimation for Pedestrian Lane Detection in Assistive Navigation
abstract
Automatic pedestrian lane detection is a challenging problem that is of great interest in assistive navigation and autonomous driving. Such a detection system must cope well with variations in lane surfaces and illumination conditions so that a vision-impaired user can navigate safely in unknown environments. This paper proposes a new lightweight Bayesian Gabor Network (BGN) for camera-based detection of pedestrian lanes in unstructured scenes. In our approach, each Gabor parameter is represented as a learnable Gaussian distribution using variational Bayesian inference. For the safety of vision-impaired users, in addition to an output segmentation map, the network provides two full-resolution maps of aleatoric uncertainty and epistemic uncertainty as well-calibrated confidence measures. Our Gabor-based method has fewer weights than the standard CNNs, therefore it is less prone to overfitting and requires fewer operations to compute. Compared to the state-of-the-art semantic segmentation methods, the BGN maintains a competitive segmentation performance while achieving a significantly compact model size (from$1.8\times $to$237.6\times $reduction), a fast prediction time (from$1.2\times $to$67.5\times $faster), and a well-calibrated uncertainty measure. We also introduce a new lane dataset of 10,000 images for objective evaluation in pedestrian lane detection research.
Hoang Thanh Le 0001, Son Lam Phung, Abdesselam Bouzerdoum
IEEE Trans. Circuits Syst. Video Technol.3
2022 Classification of Improvised Explosive Devices Using Multilevel Projective Dictionary Learning With Low-Rank Prior
abstract
Improvised explosive devices (IEDs) pose a significant threat to defense forces and humanitarian demining personnel. They are weapons of modern times, made from nonconventional military materials, rendering them difficult to identify when buried in the ground. Numerous studies focus on detecting these explosive threats and reducing the false alarm rate. However, there are few attempts to identify the detected explosive devices to take proper countermeasures. This article presents a multilevel projective dictionary learning (DL) method to classify ground-penetrating radar signals from IEDs. The proposed dictionary learning method solves three different tasks simultaneously: suppressing background clutter, learning a set of discriminative features for classification, and training a classifier. The suppression of ground clutter is formulated as a low-rank (LR) optimization problem with sparse constraints, where a low-rank subspace is learned from background clutter signals. Dictionary learning is used to transform the target signals into discriminative feature vectors, which are in turn used by the classifier to predict the target class. Experiments were conducted on real radar data. The results showed that the proposed method is more effective than the existing dictionary models and machine learning methods.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum, Canicious Abeynayake
IEEE Trans. Geosci. Remote. Sens.2
2021 Sparsity And Nonnegativity Constrained Krylov Approach For Direction Of Arrival Estimation
abstract
The conventional delay-and-sum beamforming technique results in blurred source maps due to its poor spatial resolution and high side-lobe levels. To overcome these limitations, the deconvolution approach for the mapping of acoustic sources (DAMAS) has been proposed as a postprocessing stage for image enhancement. DAMAS solves an inverse problem in the form of a system of linear equations. However, this is computationally intensive. This paper presents an approach that imposes two additional constraints to the inverse problem, namely sparsity and nonnegativity of the solution. The resulting constrained problem is solved within the Krylov projection framework. Moreover, the mapping of the sparsity penalty into the Krylov subspace is approximated by a sequence of ℓ2-norm problems via the iteratively reweighted norm (IRN) approach. Experimental results are presented which demonstrate the merits of the proposed method compared to several state-of-the-art approaches in terms of reconstruction accuracy and computation time.
Hamza Baali, Abdesselam Bouzerdoum, Abdelkrim Khelif
ICASSP2
2021 A Trajectory-Based Method for Dynamic Scene Recognition
abstract
Existing methods for dynamic scene recognition mostly use global features extracted from the entire video frame or a video segment. In this paper, a trajectory-based dynamic scene recognition method is proposed. A trajectory is formed by a pixel moving across consecutive frames of a video segment. The local regions surrounding the trajectory provide useful appearance and motion information about a portion of the video segment. The proposed method works at several stages. First, dense and evenly distributed trajectories are extracted from a video segment. Then, the fully-connected-layer features are extracted from each trajectory using a pre-trained Convolutional Neural Networks (CNNs) model, forming a feature sequence. Next, these feature sequences are fed into a Long-Short-Term-Memory (LSTM) network to learn their temporal behavior. Finally, by aggregating the information of the trajectories, a global representation of the video segment can be obtained for classification purposes. The LSTM is trained using synthetic trajectory feature sequences instead of real ones. The synthetic feature sequences are generated with a series of generative adversarial networks (GANs). In addition to classification, category-specific discriminative trajectories are located in a video segment, which help reveal what portions of a video segment are more important than others. This is achieved by formulating an optimization problem to learn discriminative part detectors for all categories simultaneously. Experimental results on two benchmark dynamic scene datasets show that the proposed method is very competitive with six other methods.
Xiaoming Peng, Abdesselam Bouzerdoum, Son Lam Phung
Int. J. Pattern Recognit. Artif. Intell.2
2021 A part-based spatial and temporal aggregation method for dynamic scene recognition
Xiaoming Peng, Abdesselam Bouzerdoum, Son Lam Phung
Neural Comput. Appl.2
2021 Variational Bayesian Compressive Multipolarization Indoor Radar Imaging
abstract
This article introduces a probabilistic Bayesian model for addressing the problem of compressive multipolarization through-wall radar imaging (TWRI). The proposed approach formulates the task of wall-clutter mitigation and multipolarization image reconstruction as a Bayesian inference problem for a joint distribution between observed radar measurements and latent wall-clutter matrix and indoor target images. The joint probability distribution incorporates three prior beliefs: low-dimensional structure of the wall reflections, group sparsity structure of the target images, and joint sparsity among the polarization images. These signal attributes are modeled through hierarchical priors, whose parameters and hyperparameters are treated with a full Bayesian formulation. Furthermore, this article presents a variational Bayesian inference algorithm that estimates wall-clutter and multipolarization images as posterior distributions and optimizes the model parameters and hyperparameters simultaneously. Experimental results on simulated and real radar data show that the proposed model is very effective at removing wall clutter and enhancing target localization even when the radar measurements are significantly reduced.
Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung
IEEE Trans. Geosci. Remote. Sens.2
2021 Clutter Removal in Through-the-Wall Radar Imaging Using Sparse Autoencoder With Low-Rank Projection
abstract
Through-the-wall radar imaging is a sensing technology that can be used by first responders to see through obscure barriers during search-and-rescue missions or deployed by law enforcement and military personnel to maintain situational awareness during tactical operations. However, the strong reflections from the front wall and other obstacles render the detection of stationary targets very difficult. In this article, a learning-based approach is proposed to mitigate the effect of the wall and background clutter. A sparse autoencoder with a low-rank projection is developed to mitigate the wall clutter and recover the target signal. The weights of the proposed autoencoder are determined by solving an augmented Lagrange multiplier optimization problem, and the regularization parameters are estimated using the Bayesian optimization technique. Experiments using real data from a stepped-frequency radar were conducted to illustrate its effectiveness for wall clutter removal. The results show that the proposed method achieves superior performance compared with the existing approaches.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
IEEE Trans. Geosci. Remote. Sens.2
2021 Toward Moving Target Detection in Through-the-Wall Radar Imaging
abstract
With the advances in radar technology, through-the-wall radar imaging (TWRI) has become a viable sensing modality that can allow fire-and-rescue personnel, police, and military forces to detect, localize, and identify targets behind opaque obstacles. Many of the existing TWRI approaches detect either stationary or moving targets but not both of them simultaneously. In this article, a method is proposed to detect both stationary and moving targets from a sequence of radar signals. The proposed method decomposes the 3-D radar data, i.e., frequency, space, and time data into a low-rank tensor and two sets of sparse images. One set of images comprises the stationary targets, and the other set of images contains the moving targets. Wall clutter removal and target detection are formulated into an optimization problem regularized by tensor low-rank, joint sparsity, and total variation constraints. Then, an alternating direction technique is developed to reconstruct the sets of stationary and moving target images. Experiments using simulated and real radar signals are conducted. The experimental results illustrate the effectiveness of the proposed method to detect and separate the stationary and moving targets into a pair of sparse images.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
IEEE Trans. Geosci. Remote. Sens.2
2020 Multichannel Signal Classification Using Vector Autoregression
abstract
The analysis of multichannel signals (MCS) has received a great deal of attention in the past few years. Modeling MCS requires depicting not only the temporal correlations within each single-channel signal (SCS) but also the interdependencies between marginal signals. The vector autoregressive (VAR) model is well adapted to providing insights to these ubiquitous dependencies, which is why it has been widely adopted for forecasting and analyzing impulse responses. Despite that, only a few studies have employed the VAR model for classification. To further explore this area, we propose a simple yet effective approach based on modeling MCS with a VAR process. To demonstrate the performance of our approach, we test it on real EEG recordings to discriminate between control and alcoholic subjects. Experimental results show that the proposed VAR approach can be very effective in MCS classification; it achieves competitive results on the benchmark dataset compared to existing state-of-the-art techniques.
Amine Haboub, Hamza Baali, Abdesselam Bouzerdoum
ICASSP3
2020 A Variational Bayesian Approach for Multichannel Through-Wall Radar Imaging with Low-Rank and Sparse Priors
abstract
This paper considers the problem of multichannel through-wall radar (TWR) imaging from a probabilistic Bayesian perspective. Given the observed radar signals, a joint distribution of the observed data and latent variables is formulated by incorporating two important beliefs: low-dimensional structure of wall reflections and joint sparsity among channel images. These priors are modeled through probabilistic distributions whose hyperparameters are treated with a full Bayesian formulation. Furthermore, the paper presents a variational Bayesian inference algorithm that captures wall clutter and provides channel images as full posterior distributions. Experimental results on real data show that the proposed model is very effective at removing wall clutter and enhancing target localization.
Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung
ICASSP2
2020 Infer The Input To The Generator Of Auxiliary Classifier Generative Adversarial Networks
abstract
Generative Adversarial Networks (GANs) are deep-learning-based generative models. This paper presents three methods to infer the input to the generator of auxiliary classifier generative adversarial networks (ACGANs), which are a type of conditional GANs. The first two methods, named i-ACGAN- r and i-ACGAN-d, are “inverting” methods, which obtain an inverse mapping from an image to the class label and the latent sample. By contrast, the third method, referred to as i-ACGAN-e, directly infers both the class label and the latent sample by introducing an encoder into an ACGAN. The three methods were evaluated on two natural scene datasets, using two performance measures: the class recovery accuracy and the image reconstruction error. Experimental results show that i-ACGAN-e outperforms the other two methods in terms of the class recovery accuracy. However, the images generated by the other two methods have smaller image reconstruction errors. The source code is publicly available from https://github.com/XMPeng/Infer-Input-ACGAN.
Xiaoming Peng, Abdesselam Bouzerdoum, Son Lam Phung
ICIP2
2020 Real-time Pedestrian Lane Detection for Assistive Navigation using Neural Architecture Search
abstract
Pedestrian lane detection is a core component in many assistive and autonomous navigation systems. These systems are usually deployed in environments that require realtime processing. Many state-of-the-art deep neural networks only focus on detection accuracy but not inference speed. Without further modifications, they are not suitable for real-time applications. Furthermore, the task of designing a high-performing deep neural network is time-consuming and requires experience. To tackle these issues, we propose a neural architecture search algorithm that can find the best deep network for pedestrian lane detection automatically. The proposed method searches in a network-level space using the gradient descent algorithm. Evaluated on a dataset of 5,000 images, the deep network found by the proposed algorithm achieves comparable segmentation accuracy, while being significantly faster than other state-of-the-art methods. The proposed method has been successfully implemented as a real-time pedestrian lane detection tool.
Sui Paul Ang, Son Lam Phung, Abdesselam Bouzerdoum, Thi Nhat Anh Nguyen, Soan Thi Minh Duong, Mark M. Schira
ICPR3
2020 Ordinal Depth Classification Using Region-based Self-attention
abstract
Depth perception is essential for scene understanding, autonomous navigation and augmented reality. Depth estimation from a single 2D image is challenging due to the lack of reliable cues, e.g. stereo correspondences and motions. Modern approaches exploit multi-scale feature extraction to provide more powerful representations for deep networks. However, these studies only use simple addition or concatenation to combine the extracted multi-scale features. This paper proposes a novel region-based self-attention (rSA) unit for effective feature fusions. The rSA recalibrates the multi-scale responses by explicitly modelling the dependency between channels in separate image regions. We discretize continuous depths to formulate an ordinal depth classification problem in which the relative order between categories is preserved. The experiments are performed on a dataset of 4410 RGB-D images, captured in outdoor environments at the University of Wollongong's campus. The proposed module improves the models on small-sized datasets by 22% to 40%.
Minh-Hieu Phan, Son Lam Phung, Abdesselam Bouzerdoum
ICPR3
2020 Compressive Radar Imaging of Stationary Indoor Targets With Low-Rank Plus Jointly Sparse and Total Variation Regularizations
abstract
This paper addresses the problem of wall clutter mitigation and image reconstruction for through-wall radar imaging (TWRI) of stationary targets by seeking a model that incorporates low-rank (LR), joint sparsity (JS), and total variation (TV) regularizers. The motivation of the proposed model is that LR regularizer captures the low-dimensional structure of wall clutter; JS guarantees a small fraction of target occupancy and the similarity of sparsity profile among channel images; TV regularizer promotes the spatial continuity of target regions and mitigates background noise. The task of wall clutter mitigation and target image reconstruction is formulated as an optimization problem comprising LR, JS, and TV regularization terms. To handle this problem efficiently, an iterative algorithm based on the forward-backward proximal gradient splitting technique is introduced, which captures wall clutter and yields target images simultaneously. Extensive experiments are conducted on real radar data under compressive sensing scenarios. The results show that the proposed model enhances target localization and clutter mitigation even when radar measurements are significantly reduced.
Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung
IEEE Trans. Image Process.2
2020 Hybrid Deep Learning-Gaussian Process Network for Pedestrian Lane Detection in Unstructured Scenes
abstract
Pedestrian lane detection is an important task in many assistive and autonomous navigation systems. This article presents a new approach for pedestrian lane detection in unstructured environments, where the pedestrian lanes can have arbitrary surfaces with no painted markers. In this approach, a hybrid deep learning-Gaussian process (DL-GP) network is proposed to segment a scene image into lane and background regions. The network combines a compact convolutional encoder-decoder net and a powerful nonparametric hierarchical GP classifier. The resulting network with a smaller number of trainable parameters helps mitigate the overfitting problem while maintaining the modeling power. In addition to the segmentation output for each test image, the network also generates a map of uncertainty-a measure that is negatively correlated with the confidence level with which we can trust the segmentation. This measure is important for pedestrian lane-detection applications, since its prediction affects the safety of its users. We also introduce a new data set of 5000 images for training and evaluating the pedestrian lane-detection algorithms. This data set is expected to facilitate research in pedestrian lane detection, especially the application of DL in this area. Evaluated on this data set, the proposed network shows significant performance improvements compared with several existing methods.
Thi Nhat Anh Nguyen, Son Lam Phung, Abdesselam Bouzerdoum
IEEE Trans. Neural Networks Learn. Syst.3
2019 Radar Stationary and Moving Indoor Target Localization with Low-rank and Sparse Regularizations
abstract
This paper proposes a low-rank and sparse regularized optimization model to address the problem of wall clutter mitigation, stationary, and moving target indications using through-wall radar. The task of wall clutter suppression and target image reconstruction is formulated as a nuclear and ℓ1penalized least squares optimization problem in which the nuclear-norm term enforces for a low-rank wall clutter matrix and the ℓ1-norm term promotes the sparsity of the target images. An iterative algorithm based on the proximal gradient technique is introduced to solve the optimization problem. The solution comprises the wall clutter and images of stationary and moving targets. Experiments are conducted on real radar data under compressive sensing scenarios. The results show that the proposed model is very effective at removing unwanted wall clutter, reconstructing stationary targets, and capturing moving targets.
Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung
ICASSP2
2019 A Super Descriptor Tensor Decomposition for Dynamic Scene Recognition
abstract
This paper presents a new approach for dynamic scene recognition based on a super descriptor tensor decomposition. Recently, local feature extraction based on dense trajectories has been used for modeling motion. However, dense trajectories usually include a large number of unnecessary trajectories, which increase noise, add complexity, and limit the recognition accuracy. Another problem is that the traditional bag-of-words techniques encode and concatenate the local features extracted from multiple descriptors to form a single large vector for classification. This concatenation not only destroys the spatio-temporal structure among the features but also yields high dimensionality. To address these problems, first, we propose to refine the dense trajectories by selecting only salient trajectories in a region of interest containing motion. Visual descriptors consisting of oriented gradient and motion boundary histograms are then computed along the refined dense trajectories. In case of camera motion, a short-window video stabilization is integrated to compensate for global motion. Second, the extracted features from multiple descriptors are encoded using a super descriptor tensor model. To this end, the TUCKER-3 tensor decomposition is employed to obtain a compact set of salient features, followed by feature selection via Fisher ranking. Experiments are conducted using two benchmark dynamic scene recognition datasets: Maryland “in-the-wild” and YUPPEN dynamic scenes. Experimental results show that the proposed approach outperforms several existing methods in terms of recognition accuracy and achieves a performance comparable with the state-of-the-art deep learning methods. The proposed approach achieves classification rates of 89.2% for Maryland and 98.1% for YUPPEN datasets.
Muhammad Rizwan Khokher, Abdesselam Bouzerdoum, Son Lam Phung
IEEE Trans. Circuits Syst. Video Technol.2
2019 Through-the-Wall Target Separation Using Low-Rank and Variational Mode Decomposition
abstract
Through-the-wall radar imaging (TWRI) is a sensing technology that can be used for detecting, locating, and identifying targets inside enclosed building structures. Many of the existing target classification approaches focus on single-target scene. For multi-target classification, the radar signal has to be segregated into different target components. However, target separation in TWRI is a challenging problem since the radar signals consist of both strong wall reflections and weak target echoes. Furthermore, the target signals are attenuated and distorted when propagating through the wall. In this paper, a variational model with low-rank constraint is proposed for decomposing the radar signal into target components and removing the wall returns. Experimental results show that the proposed method can effectively separate the radar signal into different target components.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
IEEE Trans. Geosci. Remote. Sens.2
2019 GPR Target Detection by Joint Sparse and Low-Rank Matrix Decomposition
abstract
Ground penetrating radar (GPR) uses electromagnetic waves to image, locate, and identify changes in electric and magnetic properties in the ground. The received signal comprises not only the target echoes but also strong reflections from the rough, uneven ground surface, which impair subsurface inspections and visualization of buried objects. In this paper, a background clutter mitigation and target detection method using low-rank and sparse priors is proposed for GPR data. The radar signal is decomposed into the sum of a low-rank component and a sparse component, plus noise. The low-rank component captures the ground surface reflections and background clutter, whereas the sparse component contains the target reflections. The effectiveness of the proposed method is evaluated on real radar signals collected from buried landmines and improvised explosive devices. The experimental results show that the proposed method successfully removes the background clutter and estimates the target signals.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum, Canicious Abeynayake
IEEE Trans. Geosci. Remote. Sens.2
2019 Short-Term Forecasting of Electricity Spot Prices Containing Random Spikes Using a Time-Varying Autoregressive Model Combined With Kernel Regression
abstract
Forecasting spot prices of electricity is challenging because it not only contains seasonal variations, but also random, abrupt spikes, which depend on market conditions and network contingencies. In this paper, a hybrid model has been developed to forecast the spot prices of electricity in two main stages. In the first stage, the prices are forecasted using autoregressive time varying (ARXTV) model with exogenous variables. To improve the forecasting ability of the ARXTV model, the price variations in the training process have been smoothened using the wavelet technique. In the second stage, a kernel regression is used to estimate the price spikes, which are detected using support vector machine based model. In addition, mutual information technique is employed to select appropriate input variables for the model. A case study is carried out with the aid of price data obtained from the Australian energy market operator. It is demonstrated that the proposed hybrid method can accurately forecast electricity prices containing spikes.
Dao H. Vu, Kashem M. Muttaqi, Ashish P. Agalgaonkar, Abdesselam Bouzerdoum
IEEE Trans. Ind. Informatics4
2018 Pooling-Based Feature Extraction and Coarse-to-fine Patch Matching for Optical Flow Estimation
Xiaolin Tang, Son Lam Phung, Abdesselam Bouzerdoum, Van Ha Tang
ACCV (4)3
2018 Anatomy-Guided Inverse-Gradient Susceptibility Artifact Correction Method for High-Resolution FMRI
abstract
Functional Magnetic Resonance Imaging (fMRI) is a widely used and non-invasive technique for recording changes in brain activity. However, susceptibility artifacts are ubiquitous distortions in fMRI, especially strong in high-resolution images, causing the misrepresentation of brain function and structure in the affected regions. Here, we present a novel method for correcting these distortions in high-resolution fMRI images based on the hyper-elastic susceptibility artifact correction (HySCO) method. The novelty of the proposed method is the utilization of the easily-acquired T1-weighted (T1w) anatomy image as a ground-truth measurement to regularize deformations, thereby obtaining meaningful corrections. The performance of the new method is compared to that of HySCO. Results from high-resolution (1mm) EPI data are presented demonstrating the robustness of the new method for image correction and its suitability for subsequent fMRI analysis.
Soan Thi Minh Duong, Mark M. Schira, Son Lam Phung, Abdesselam Bouzerdoum, H. G. B. Taylor
ICASSP4
2018 Human Motion Classification with Micro-Doppler Radar and Bayesian-Optimized Convolutional Neural Networks
abstract
In recent years, Doppler radar has emerged as an alternative sensing modality for human gait classification since it measures not only the target speed, but also the local dynamics of the moving body parts, thereby creating a unique spectral signature. This paper presents a learning-based method for classifying human motions from micro-Doppler signals. Inspired by the applications of deep learning, the proposed method extracts features from the time-frequency representation of the radar signal using a cascaded of convolutional network layers. To design a optimal network architecture, the Bayesian optimization with Gaussian process priors is employed. Experimental results on real data are presented, which show a significant improvement compared to three existing approaches.
Hoang Thanh Le 0001, Son Lam Phung, Abdesselam Bouzerdoum, Fok Hing Chi Tivive
ICASSP3
2018 A Matrix Completion Approach for Wall-Clutter Mitigation in Compressive Radar Imaging of Indoor Targets
abstract
This paper presents a low-rank matrix completion approach to tackle the problem of wall clutter mitigation for through-wall radar imaging in the compressive sensing context. In particular, the task of wall clutter removal is reformulated as a matrix completion problem in which a low-rank matrix containing wall clutter is reconstructed from compressive measurements. The proposed model regularizes the low-rank prior of the wall-clutter matrix via the nuclear norm, casting the wall-clutter mitigation task as a nuclear-norm penalized least squares problem. To solve this optimization problem, an iterative algorithm based on the proximal gradient technique is introduced. Experiments on simulated full-wave electromagnetic data are conducted under compressive sensing scenarios. The results show that the proposed matrix completion approach is very effective at suppressing unwanted wall clutter and enhancing the targets.
Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung
ICASSP2
2018 Human Gait Recognition with Micro-Doppler Radar and Deep Autoencoder
abstract
The micro-Doppler signals from moving objects contain useful information about their motions. This paper introduces a novel approach for human gait recognition based on backscattered signals from a micro-Doppler radar. Three different signal techniques are utilized for the extraction of micro-Doppler features via time-frequency and time-scale representations. To classify the human motions into various types, this paper presents a deep autoencoder with the use of local patches extracted along the spectrogram and scalogram. The network configuration and the learning parameters of the deep autoencoder, which are considered as hyperparameters, are optimized by a Bayesian optimization algorithm. Experimental results produced by the proposed technique on real radar data show a significant improvement compared to several existing approaches.
Hoang Thanh Le 0001, Son Lam Phung, Abdesselam Bouzerdoum
ICPR3
2018 Convolutional neural network acceleration with hardware/software co-design
Andrew Tzer-Yeu Chen, Morteza Biglari-Abhari, Kevin I-Kai Wang, Abdesselam Bouzerdoum, Fok Hing Chi Tivive
Appl. Intell.4
2018 Stochastic variational hierarchical mixture of sparse Gaussian processes for regression
Thi Nhat Anh Nguyen, Abdesselam Bouzerdoum, Son Lam Phung
Mach. Learn.2
2018 Multipolarization Through-Wall Radar Imaging Using Low-Rank and Jointly-Sparse Representations
abstract
Compressed sensing techniques have been applied to through-the-wall radar imaging (TWRI) and multipolarization TWRI for fast data acquisition and enhanced target localization. The studies so far in this area have either assumed effective wall clutter removal prior to image formation or performed signal estimation, wall clutter mitigation, and image formation independently. This paper proposes a low-rank and sparse imaging model for jointly addressing the problem of wall clutter mitigation and image formation in multichannel TWRI. The proposed model exploits two important structures of through-wall radar signals: low-rank structure of the wall reflections and jointly-sparse structure among the different polarization images. The task of removing wall clutter and reconstructing multichannel images of the same scene behind-the-wall is formulated as a regularized least squares problem, where low-rank regularization is enforced for the wall components, and joint-sparsity penalty is imposed on channel images. To solve the optimization problem, an iterative algorithm based on the proximal gradient technique is introduced, which simultaneously estimates the wall interferences and yields multichannel images of the indoor targets. Experiments on real and simulated radar data are conducted under full measurements and compressive sensing scenarios. The results show that the proposed model is very effective at removing unwanted wall clutter and enhancing the stationary targets, even under considerable reduction in measurements.
Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung
IEEE Trans. Image Process.2
2017 Through-the-wall radar signal classification using discriminative dictionary learning
abstract
Through-the-wall radar imaging is an electromagnetic wave sensing technology capable of detecting targets behind walls, doors, and opaque obstacles. Identification of stationary targets is often achieved by first forming an image of the scene, and then segmenting and classifying the targets of interest. In order to provide prompt and reliable situational awareness, this paper proposes a radar signal classification approach that does not rely on image formation. Here, a dictionary learning based method is employed to classify targets behind a wall using the signals received from individual antennas. The cepstrum coefficients of the high resolution range profile are first extracted as features. Then, the latent consistent K-SVD algorithm is used to learn a discriminative dictionary and a linear classifier simultaneously. Experimental results show that the proposed method can classify individual radar signals with high accuracy, without having recourse to image formation.
Abdesselam Bouzerdoum, Fok Hing Chi Tivive, Jia Fei
ICASSP1
2017 Human interaction recognition using low-rank matrix approximation and super descriptor tensor decomposition
abstract
Audio-visual recognition systems rely on efficient feature extraction. Many spatio-temporal interest point detectors for visual feature extraction are either too sparse, leading to loss of information, or too dense resulting in noisy and redundant information. Furthermore, interest point detectors designed for a controlled environment can be affected by camera motion. In this paper, a salient spatio-temporal interest point detector is proposed based on a low-rank and group-sparse matrix approximation. The detector handles the camera motion through a short-window video stabilization. The multimodal audio-visual features from multiple descriptors are represented by a super descriptor, from which a compact set of features is extracted through a tensor decomposition and feature selection. This tensor decomposition retains the spatiotemporal structure among features obtained from multiple descriptors. Experimental validation is conducted using two benchmark human interaction recognition datasets: TVHID and Parliament. Experimental results are presented which show that the proposed approach outperforms many state-ofthe-art methods, achieving classification rates of 74.7% and 88.5% on the TVHID and Parliament datasets, respectively.
Muhammad Rizwan Khokher, Abdesselam Bouzerdoum, Son Lam Phung
ICASSP2
2017 Enhanced pixel-wise voting for image vanishing point detection in road scenes
abstract
Vanishing point estimation is a crucial task in vision-based road detection. This paper presents a new texture-based voting scheme, which enhances both accuracy and speed of vanishing point estimation. In the proposed method, color tensors analysis is adopted to calculate local orientations and color edges. The search space is reduced by optimizing the set of vanishing point candidates and voters. A new strategy based on Bayesian classifier is proposed to select a suitable voting function. The proposed method is evaluated on a benchmark dataset of 4000 images of pedestrian lanes with annotated vanishing points. The experimental results show that it offers an improved accuracy and significantly faster processing time compared with other state-of-the-art methods.
L. Nguyen, Son Lam Phung, Abdesselam Bouzerdoum
ICASSP3
2017 An efficient local method for stereo matching using daisy features
abstract
In this paper, a local method is proposed to estimate the visibility and disparity of pixels from a stereo pair using the DAISY feature. The problem is formulated as a joint optimization over disparity and visibility of individual pixels. The constraints on the range of disparities and the binary visibility variables are enforced by incorporating penalty terms into the cost function. Finally, the unconstrained optimization problem is solved using a Newton scheme with appropriate approximations to the Hessian matrices and gradients. The computation time of the proposed optimization method is around one minute to run for 768 × 512 stereo pairs using the DAISY feature descriptor in a C++ implementation.
Xiaoming Peng, Abdesselam Bouzerdoum, Son Lam Phung
ICIP2
2016 Variational inference for infinite mixtures of sparse Gaussian processes through KL-correction
abstract
We propose a new approximation method for Gaussian process (GP) regression based on the mixture of experts structure and variational inference. Our model is essentially an infinite mixture model in which each component is composed of a Gaussian distribution over the input space, and a Gaussian process expert over the output space. Each expert is a sparse GP model augmented with its own set of inducing points. Variational inference is made feasible by assuming that the training outputs are independent given the inducing points. In previous works on variational mixture of GP experts, the inducing points are selected through a greedy selection algorithm, which is computationally expensive. In our method, both the inducing points and hyperparameters of the experts are learned through maximizing an improved lower bound of the marginal likelihood. Experiments on benchmark datasets show the advantages of the proposed method.
Thi Nhat Anh Nguyen, Abdesselam Bouzerdoum, Son Lam Phung
ICASSP2
2016 Radar imaging of stationary indoor targets using joint low-rank and sparsity constraints
abstract
This paper introduces a joint low-rank and sparsity-based model to address the problem of wall-clutter mitigation in compressed through-the-wall radar imaging. The proposed model is motivated by two observations that wall reflections reside in a low-rank subspace, and target signals tend to be sparse. In the proposed approach, the task of segregating target returns from wall reflections is formulated as a joint low-rank and sparsity constrained optimization problem. Here, the low rank constraint is imposed on the wall component and the sparsity constraint is used to model the target component. An iterative soft thresholding algorithm is developed to estimate a low-rank matrix of wall clutter and a sparse matrix of target reflections from a reduced measurement set. Once the wall and target components are estimated, the target signals are used for scene reconstruction. Experimental evaluation was conducted using real radar data. The results show that the proposed model is very effective at removing wall clutter and reconstructing the image of behind-the-wall targets from reduced measurements.
Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung, Fok Hing Chi Tivive
ICASSP2
2016 Hardware/Software Co-design for a Gender Recognition Embedded System
Andrew Tzer-Yeu Chen, Morteza Biglari-Abhari, Kevin I-Kai Wang, Abdesselam Bouzerdoum, Fok Hing Chi Tivive
IEA/AIE4
2016 Pedestrian lane detection in unstructured scenes for assistive navigation
abstract
Automatic detection of the pedestrian lane in a scene is an important task in assistive and autonomous navigation. This paper presents a vision-based algorithm for pedestrian lane detection in unstructured scenes, where lanes vary significantly in color, texture, and shape and are not indicated by any painted markers. In the proposed method, a lane appearance model is constructed adaptively from a sample image region, which is identified automatically from the image vanishing point. This paper also introduces a fast and robust vanishing point estimation method based on the color tensor and dominant orientations of color edge pixels. The proposed pedestrian lane detection method is evaluated on a new benchmark dataset that contains images from various indoor and outdoor scenes with different types of unmarked lanes. Experimental results are presented which demonstrate its efficiency and robustness in comparison with several existing methods.
Son Lam Phung, Manh Cuong Le, Abdesselam Bouzerdoum
Comput. Vis. Image Underst.3
2015 Multi-view indoor scene reconstruction from compressed through-wall radar measurements using a joint bayesian sparse representation
abstract
This paper addresses the problem of scene reconstruction, incorporating wall-clutter mitigation, for compressed multi-view through-the-wall radar imaging. We consider the problem where the scene is sensed using different reduced sets of frequencies at different antennas. A joint Bayesian sparse recovery framework is first employed to estimate the antenna signal coefficients simultaneously, by exploiting the sparsity and correlations between antenna signals. Following joint signal coefficient estimation, a subspace projection technique is applied to segregate the target coefficients from the wall contributions. Furthermore, a multitask linear model is developed to relate the target coefficients to the scene, and a composite scene image is reconstructed by a joint Bayesian sparse framework, taking into account the inter-view dependencies. Experimental results show that the proposed approach improves reconstruction accuracy and produces a composite scene image in which the targets are enhanced and the background clutter is attenuated.
Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung, Fok Hing Chi Tivive
ICASSP2
2015 A tensor-based subspace wall clutter mitigation method for through-the-wall radar imaging
abstract
In through-the-wall radar imaging, targets behind the wall reflect weak electromagnetic signals that are obscured by the strong returns from exterior wall, rendering the detection and classification of indoor stationary targets very difficult. In this paper, a tensor-based subspace method is proposed for wall clutter mitigation. The radar signals received from the antenna array are transformed into a data tensor. Higher-order singular value decomposition is used to segregate the wall reflections from the target returns. Simulation and experimental results show that the proposed method is effective in removing reflections backscattered from both homogeneous and heterogeneous walls.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
ICASSP2
2015 Depth image super-resolution using internal and external information
abstract
The fast development of 3-D imaging techniques has increased demands for high-resolution depth images. Conventional depth super-resolution methods reconstruct the high-resolution image by accessing high frequency information, either internally from a high-resolution intensity image or externally from a high-resolution image database. In this paper, a new depth super-resolution method based on joint regularization is proposed, which exploits both internal and external high frequency information. Specifically, a joint regularization problem with different constraints is formulated, which allows us to solve for the high-resolution image and a sparse code simultaneously. These constraints are constructed by utilizing information from both internal and external high-frequency sources. Experimental evaluation suggests that the proposed method provides improved results over existing approaches, in terms of both visual appearance and objective image quality.
Haoheng Zheng, Abdesselam Bouzerdoum, Son Lam Phung
ICASSP2
2015 Invariant image recognition under projective deformations: An image normalization approach
abstract
Robustness in image recognition refers to the ability to perceive an image pattern regardless of factors including camera views and locations. This paper proposes an image normalization algorithm that allows an image with arbitrary projective distortions to be recognized efficiently. The normalization algorithm calculates the required projective transformation matrix using image moments. For an input image, a set of 8 output images that are independent of projective deformations are generated. The proposed algorithm is evaluated on three benchmark data sets. The experimental results show that the proposed normalization is significantly more accurate than the existing rank minimization and affine normalization methods.
Son Lam Phung, Abdesselam Bouzerdoum, Amine Bermak
VCIP3
2015 A Subspace Projection Approach for Wall Clutter Mitigation in Through-the-Wall Radar Imaging
abstract
One of the main challenges in through-the-wall radar imaging (TWRI) is the strong exterior wall returns, which tend to obscure indoor stationary targets, rendering target detection and classification difficult, if not impossible. In this paper, an effective wall clutter mitigation approach is proposed for TWRI that does not require knowledge of the background scene nor does it rely on accurate modeling and estimation of wall parameters. The proposed approach is based on the relative strength of the exterior wall returns compared to behind-wall targets. It applies singular value decomposition to the data matrix constructed from the space-frequency measurements to identify the wall subspace. Orthogonal subspace projection is performed to remove the wall electromagnetic signature from the radar signals. Furthermore, this paper provides an analysis of the wall and target subspace characteristics, demonstrating that both wall and target subspaces can be multidimensional. While the wall subspace depends on the wall type and building material, the target subspace depends on the location of the target, the number of targets in the scene, and the size of the target. Experimental results using simulated and real data demonstrate the effectiveness of the subspace projection method in mitigating wall clutter while preserving the target image. It is shown that the performance of the proposed approach, in terms of the improvement factor of the target-to-clutter ratio, is better than existing approaches and is comparable to that of background subtraction, which requires knowledge of a reference background scene.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum, Moeness G. Amin
IEEE Trans. Geosci. Remote. Sens.2
2014 Lane Detection in Unstructured Environments for Autonomous Navigation Systems
Manh Cuong Le, Son Lam Phung, Abdesselam Bouzerdoum
ACCV (1)3
2014 Enhanced wall clutter mitigation for compressed through-the-wall radar imaging using joint Bayesian sparse signal recovery
abstract
This paper addresses the problem of wall clutter mitigation in compressed sensing through-the-wall radar imaging, where a different set of frequencies is sensed at different antenna locations. A joint Bayesian sparse approximation framework is first employed to reconstruct all the signals simultaneously by exploiting signal sparsity and correlations between antenna signals. This is in contrast to previous approaches where the signal at each antenna location is reconstructed independently. Furthermore, to promote sparsity and improve reconstruction accuracy, a sparsifying wavelet dictionary is employed in the sparse signal recovery. Following signal reconstruction, a subspace projection technique is applied to remove wall clutter, prior to image formation. Experimental results on real data show that the proposed approach produces significantly higher reconstruction accuracy and requires far fewer measurements for forming high-quality images, compared to the single-signal compressed sensing model, where each antenna signal is reconstructed independently.
Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung, Fok Hing Chi Tivive
ICASSP2
2014 Affine-invariant scene categorization
abstract
This paper presents a scene categorization method that is invariant to affine transformations. We propose a new moment-based normalization algorithm to generate an output image that is independent of the position, rotation, shear, and scale of the input image. In the proposed approach, an affine transform matrix is determined subject to the normalized image satisfying a set of moment constraints. After image normalization, a dense set of local features is extracted using scattering transform, and the global features are then formed via a sparse coding method. We evaluate the proposed method and other state-of-the-art algorithms on a benchmark dataset. The experimental results show that for images distorted with affine transformations, the proposed normalization increases the classification rate by about 28%, compared with the scene categorization approach that uses no normalization.
Son Lam Phung, Abdesselam Bouzerdoum
ICIP3
2014 Object segmentation and classification using 3-D range camera
Son Lam Phung, Abdesselam Bouzerdoum
J. Vis. Commun. Image Represent.3
2013 A compressed sensing method for complex-valued signals with application to through-the-wall radar imaging
abstract
In this paper, we present a compressed sensing method for complex-valued signals based on multiple measurement vector compressed sensing model. The proposed method constrains the real and imaginary parts of the recovered signal to have the same sparsity profile. It is applied to a compressed sensing through-the-wall radar imaging problem. Experiments based on synthetic data shows that the proposed method achieves lower reconstruction error than the existing CS method.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
ICASSP2
2013 Depth image super-resolution using multi-dictionary sparse representation
abstract
In this paper, we propose a new depth super-resolution technique based on multiple dictionary learning. A novel dictionary selection method using basis pursuit is proposed to generate multiple dictionaries adaptively. A sparse representation of each low-resolution input patch is derived based on the learned dictionaries, and then used to reconstruct the corresponding high-resolution patch. Experimental results are presented which show that the proposed multi-dictionary scheme outperforms existing depth super-resolution methods.
Abdesselam Bouzerdoum, Son Lam Phung
ICIP2
2013 Two-Stage Fuzzy Fusion With Applications to Through-the-Wall Radar Imaging
abstract
A two-stage fuzzy image fusion approach, which combines multiple radar images of the same scene, is proposed to produce a more informative image. In this approach, two different image fusion methods are first applied. Then, a fuzzy logic fusion method is applied to the outputs of the first fusion stage. The performance of the proposed approach is evaluated on through-the-wall radar images obtained using different polarizations. Experimental results show that the proposed approach enhances image quality by producing outputs with high target intensity values and low clutter.
Cher Hau Seng, Abdesselam Bouzerdoum, Moeness G. Amin, Son Lam Phung
IEEE Geosci. Remote. Sens. Lett.2
2013 Biologically inspired approaches for visual information processing and analysis
Azeddine Beghdadi, Abdesselam Bouzerdoum, Khan M. Iftekharuddin, Mohamed-Chaker Larabi
Signal Process. Image Commun.2
2013 A survey of perceptual image processing methods
Azeddine Beghdadi, Mohamed-Chaker Larabi, Abdesselam Bouzerdoum, Khan M. Iftekharuddin
Signal Process. Image Commun.3
2013 Sparse Representation of GPR Traces With Application to Signal Classification
abstract
Sparse representation (SR) models a signal with a small number of elementary waves using an overcomplete dictionary. It has been employed for a wide range of signal and image processing applications, including denoising, deblurring, and compression. In this paper, we present an adaptive SR method for modeling and classifying ground penetrating radar (GPR) signals. The proposed method decomposes each GPR trace into elementary waves using an adaptive Gabor dictionary. The sparse decomposition is used to extract salient features for SR and classification of GPR signals. Experimental results on real-world data show that the proposed sparse decomposition achieves efficient signal representation and yields discriminative features for pattern classification.
Wenbin Shao, Abdesselam Bouzerdoum, Son Lam Phung
IEEE Trans. Geosci. Remote. Sens.2
2013 Probabilistic Fuzzy Image Fusion Approach for Radar Through Wall Sensing
abstract
This paper addresses the problem of combining multiple radar images of the same scene to produce a more informative composite image. The proposed approach for probabilistic fuzzy logic-based image fusion automatically forms fuzzy membership functions using the Gaussian-Rayleigh mixture distribution. It fuses the input pixel values directly without requiring fuzzification and defuzzification, thereby removing the subjective nature of the existing fuzzy logic methods. In this paper, the proposed approach is applied to through-the-wall radar imaging in urban sensing and evaluated on real multi-view and polarimetric data. Experimental results show that the proposed approach yields improved image contrast and enhances target detection.
Cher Hau Seng, Abdesselam Bouzerdoum, Moeness G. Amin, Son Lam Phung
IEEE Trans. Image Process.2
2012 Adaptive Autoregressive Logarithmic Search for 3D Human Tracking
abstract
Human tracking is an important vision task in video surveillance and perceptual human-computer interfaces. This paper presents a novel algorithm for region-based human tracking using color and depth features. We propose an adaptive autoregressive logarithmic search (ARLS) to estimate the target position, and use depth information to further reduce the false alarm rate. The new ARLS algorithm is evaluated on a color and depth (RGBD) video dataset acquired with the Kinect sensor. The dataset contains various real-world scenarios with illumination and speed variations, and partial occlusion. The experimental results show that the ARLS algorithm is able to handle difficult tracking scenarios, it achieves a tracking accuracy of 91.26% on the test dataset. The proposed algorithm is compared with two tracking algorithms, namely the particle filtering and a modified logarithmic search algorithm.
Peiyao Li, Abdesselam Bouzerdoum, Son Lam Phung
AVSS2
2012 A Gaussian-Rayleigh mixture modeling approach for through-the-wall radar image segmentation
abstract
In this paper, we propose a Gausssian-Rayleigh mixture modeling approach to segment indoor radar images in urban sensing applications. The performance of the proposed method is evaluated on real 2D polarimetric data. Experimental results show that the proposed method enhances image quality by distinguishing between target and clutter regions. The proposed method is also compared to an existing Neyman-Pearson (NP) target detector that has been recently devised for through-the-wall radar imaging. Performance evaluation of both methods shows that the proposed method outperforms the NP detector in enhancing the input images.
Cher Hau Seng, Abdesselam Bouzerdoum, Moeness G. Amin, Fauzia Ahmad
ICASSP2
2012 Compressed sensing-based frequency selection for classification of ground penetrating radar signals
abstract
In this paper we present an automatic classification system for ground penetrating radar (GPR) signals. The system extracts the magnitude spectra at resonant frequencies and classifies them using support vector machines. To locate the resonant frequencies, we propose an approach based on compressed sensing and orthogonal matching pursuit. The performance of the system is evaluated by classifying GPR traces from different ballast fouling conditions. The experimental results show that the proposed approach, compared to the approach of using frequencies at local maxima, represents the GPR signal more efficiently using a small number of coefficients, and obtains higher classification accuracy.
Wenbin Shao, Abdesselam Bouzerdoum, Son Lam Phung
ICASSP2
2012 Scene Segmentation and Pedestrian Classification from 3-D Range and Intensity Images
abstract
This paper proposes a new approach to classify obstacles using a time-of-flight camera, for applications in assistive navigation of the visually impaired. Combining range and intensity images enables fast and accurate object segmentation, and provides useful navigation cues such as distances to the nearby obstacles and obstacle types. In the proposed approach, a 3-D range image is first segmented using histogram thresholding and mean-shift grouping. Then Fourier and GIST descriptors are applied on each segmented object to extract shape and texture features. Finally, support vector machines are used to recognize the obstacles. This paper focuses on classifying pedestrian and non-pedestrian obstacles. Evaluated on an image data set acquired using a time-of-flight camera, the proposed approach achieves a classification rate of 99.5%.
Son Lam Phung, Abdesselam Bouzerdoum
ICME3
2012 Pedestrian lane detection for assistive navigation of blind people
Manh Cuong Le, Son Lam Phung, Abdesselam Bouzerdoum
ICPR3
2012 A short length window-based method for islanding detection in distributed generation
abstract
Distributed generation (DG) has recently drawn the interest to meet the increased load demand with minimum investment. But cohesive operation of these DG sources, in a grid-connected environment, gives rise to several issues during abnormal conditions of the utility system. This paper addresses the detection method of one such crucial event which is “islanding”. A short length window based Mahalanobis Distance method has been proposed in this paper to detect islanding. A trade-off between computational time and accuracy has been maintained to make it reliable and acceptable. In this method, network parameters such as rate of change of frequency (ROCOF), rate of change of voltage (ROCOV), rate of change of real power (ROCOP) and rate of change of reactive power (ROCOQ) have been extracted from the voltage and current signal. Standard Deviations of the network features have been used as parameters for islanding and non-islanding events. These parameters have been classified with the proposed Mahalanobis Distance method incorporating short length window. The proposed method has been simulated in a test distribution system and it has been compared with Support Vector Machine (SVM), and Feed-forward Multi-layer Neural Network (FFML NN) classifiers to show its reliability and acceptability.
Mollah R. Alam, Kashem M. Muttaqi, Abdesselam Bouzerdoum
IJCNN3
2012 Automatic classification of human motions using Doppler radar
abstract
This paper presents a new approach to classify human motions using a Doppler radar for applications in security and surveillance. Traditionally, the Doppler radar is an effective tool for detecting the position and velocity of a moving target, even in adverse weather conditions and from a long range. In this paper, we are interested in using the Doppler radar to recognize the micro-motions exhibited by people. In the proposed approach, a frequency modulated continuous wave radar is applied to scan the target, and the short-time Fourier transform is used to convert the radar signal into spectrogram. Then, the new two-directional, two-dimensional principal component analysis and linear discriminant analysis are performed to obtain the feature vectors. This approach is more computationally efficient than the traditional principal component analysis. Finally, support vector machines are applied to classify feature vectors into different human motions. Evaluated on a radar data set with three types of motions, the proposed approach has a classification rate of 91.9%.
Jingli Li, Son Lam Phung, Fok Hing Chi Tivive, Abdesselam Bouzerdoum
IJCNN4
2011 Multiple-Measurement Vector model and its application to Through-the-Wall Radar Imaging
abstract
This paper addresses the problem of Through-the-Wall Radar Imaging (TWRI) using the Multiple-Measurement Vector (MMV) compressive sensing model. TWR image formation is reformulated as a compressed sensing (CS) problem, seeking a sparse representation in the spatial domain. In traditional CS-based through-the-wall radar imaging (TWRI) methods, the measurement matrix is vectorized so that a single measurement vector (SMV) model is applied to generate a sparse solution, which represents a scene comprising point-like targets. For multiple measurement TWRI problems, the SMV model may produce a sub-optimum sparse solution. On the other hand, the proposed MMV model for TWRI generates a more sparse scene by processing all the measurements simultaneously. To evaluate the effectiveness of the proposed method, it is applied to fuse multiple polarization data to form the radar image. Based on simulated data with different number of measurements and noise levels, the proposed MMV-based TWRI method produces better TWR images in terms of image quality and detection accuracy.
Jie Yang 0009, Abdesselam Bouzerdoum, Fok Hing Chi Tivive, Moeness G. Amin
ICASSP2
2011 Adaptive regularization for multiple image restoration using an extended Total Variations approach
abstract
In this paper a Variational Inequality method for multiple in- put, multiple output image restoration is presented using an extended Total Variations (TV) regularizer. This approach calculates an adaptive regularization parameter for each image based on their respective degradations. The proposed ex- tended Total Variations regularizer combines both intra-image and inter-image pixel information for improved restoration performance. Hyperparameters for controlling this new TV measure are calculated using a Bayesian joint maximum a posteriori approach.
Matthew Andrew Kitchener, Abdesselam Bouzerdoum, Son Lam Phung
ICIP2
2011 Optical flow estimation using sparse gradient representation
abstract
This paper introduces a sparsity based optical flow estimation method in digital video sequences. The method stems from the key observation that the gradient field of optical flow, in digital video sequences, is usually structured and sparse in spatial domain, provided there is a small number of multiple motions in the scene. The gradient field of motion vectors is formed by the pixels forming the edges of moving objects. We utilize this fact and formulate the optical flow estimation problem in sparse representation framework. We then use a minimization algorithm over ℓ1norm of the gradient flow field to find the solution to this problem. The proposed algorithm has been evaluated on Middlebury's benchmark video sequence database.
Muhammad Wasim Nawaz, Abdesselam Bouzerdoum, Son Lam Phung
ICIP2
2011 Automatic Classification of Ground-Penetrating-Radar Signals for Railway-Ballast Assessment
abstract
The ground-penetrating radar (GPR) has been widely used in many applications. However, the processing and interpretation of the acquired signals remain challenging tasks since an experienced user is required to manage the entire operation. In this paper, we present an automatic classification system to assess railway-ballast conditions. It is based on the extraction of magnitude spectra at salient frequencies and their classification using support vector machines. The system is evaluated on real-world railway GPR data. The experimental results show that the proposed method efficiently represents the GPR signal using a small number of coefficients and achieves a high classification rate when distinguishing GPR signals reflected by ballasts of different conditions.
Wenbin Shao, Abdesselam Bouzerdoum, Son Lam Phung, Lijun Su, Buddhima Indraratna, Cholachat Rujikiatkamjorn
IEEE Trans. Geosci. Remote. Sens.2
2010 A particle swarm optimization algorithm based on orthogonal design
abstract
The last decade has witnessed a great interest in using evolutionary algorithms, such as genetic algorithms, evolutionary strategies and particle swarm optimization (PSO), for multivariate optimization. This paper presents a hybrid algorithm for searching a complex domain space, by combining the PSO and orthogonal design. In the standard PSO, each particle focuses only on the error propagated back from the best particle, without “communicating” with other particles. In our approach, this limitation of the standard PSO is overcome by using a novel crossover operator based on orthogonal design. Furthermore, instead of the “generating-and-updating” model in the standard PSO, the elitism preservation strategy is applied to determine the possible movements of the candidate particles in the subsequent iterations. Experimental results demonstrate that our algorithm has a better performance compared to existing methods, including five PSO algorithms and three evolutionary algorithms.
Jie Yang 0009, Abdesselam Bouzerdoum, Son Lam Phung
IEEE Congress on Evolutionary Computation2
2010 Multi-resolution Mean-Shift Algorithm for Vector Quantization
abstract
Here we propose to apply the mean-shift algorithm to the four image subbands generated by a DWT, namely the LL, LH, HL and HH subbands. The simulated annealing technique traditionally used to identify the modes of the distribution for different resolutions can be performed by exploration of the multi-scale DWT pyramid, avoiding the costly estimation by a full mean-shift at each level.
Philippe Loic Marie Bouttefroy, Abdesselam Bouzerdoum, Azeddine Beghdadi, Son Lam Phung
DCC2
2010 On the analysis of background subtraction techniques using Gaussian Mixture Models
abstract
In this paper, we conduct an investigation into background subtraction techniques using Gaussian Mixture Models (GMM) in the presence of large illumination changes and background variations. We show that the techniques used to date suffer from the trade-off imposed by the use of a common learning rate to update both the mean and variance of the component densities, which leads to a degeneracy of the variance and creates “saturated pixels”. To address this problem, we propose a simple yet effective technique that differentiates between the two learning rates, and imposes a constraint on the variance so as to avoid the degeneracy problem. Experimental results are presented which show that, compared to existing techniques, the proposed algorithm provides more robust segmentation in the presence of illumination variations and abrupt changes in background distribution.
Philippe Loic Marie Bouttefroy, Abdesselam Bouzerdoum, Son Lam Phung, Azeddine Beghdadi
ICASSP2
2010 A training algorithm for sparse LS-SVM using Compressive Sampling
abstract
Least Squares Support Vector Machine (LS-SVM) has become a fundamental tool in pattern recognition and machine learning. However, the main disadvantage is lack of sparseness of solutions. In this article Compressive Sampling (CS), which addresses the sparse signal representation, is employed to find the support vectors of LS-SVM. The main difference between our work and the existing techniques is that the proposed method can locate the sparse topology while training. In contrast, most of the traditional methods need to train the model before finding the sparse support vectors. An experimental comparison with the standard LS-SVM and existing algorithms is given for function approximation and classification problems. The results show that the proposed method achieves comparable performance with typically a much sparser model.
Jie Yang 0009, Abdesselam Bouzerdoum, Son Lam Phung
ICASSP2
2010 Adaptive regularization for image restoration using a variational inequality approach
abstract
In this paper, a generalized image restoration method is formulated as a variational inequality problem, whose solution is obtained using a dynamic system approach. In this method, the restored image and the regularization parameter are obtained simultaneously. In particular, the optimum regularization parameter is determined adaptively, depending on noise and image content. The restoration problem is presented in a generalized form so that it maybe be implemented using different norms; only L1and L2norms have been implemented in this paper. A comparison based on experimental results shows that the proposed method achieves comparable if not better performance as some of the existing state-of-the-art techniques.
Matthew Andrew Kitchener, Abdesselam Bouzerdoum, Son Lam Phung
ICIP2
2010 Wavelet based nonlocal-means super-resolution for video sequences
abstract
Video sequence resolution enhancement became a popular research area during the last two decades. Although traditional super-resolution techniques have been successful in dealing with image sequences, many constraints such as global translation between frames, have to be imposed to obtain good performance. In this paper, we present a new wavelet-based nonlocal-means (WNLM) framework to bypass the motion estimation stage. It can handle complex motion changes between frames. Compared with the nonlocal-means (NLM) super-resolution framework, the proposed method provides better result in terms of PSNR and faster processing.
Haoheng Zheng, Abdesselam Bouzerdoum, Son Lam Phung
ICIP2
2010 Improved Facial Expression Recognition with Trainable 2-D Filters and Support Vector Machines
abstract
Facial expression is one way humans convey their emotional states. Accurate recognition of facial expressions is essential in perceptual human-computer interface, robotics and mimetic games. This paper presents a novel approach to facial expression recognition from static images that combines fixed and adaptive 2-D filters in a hierarchical structure. The fixed filters are used to extract primitive features. They are followed by the adaptive filters that are trained to extract more complex facial features. Both types of filters are non-linear and are based on the biological mechanism of shunting inhibition. The features are finally classified by a support vector machine. The proposed approach is evaluated on the JAFFE database with seven types of facial expressions: anger, disgust, fear, happiness, neutral, sadness and surprise. It achieves a classification rate of 96.7%, which compares favorably with several existing techniques for facial expression recognition tested on the same database.
Peiyao Li, Son Lam Phung, Abdesselam Bouzerdoum, Fok Hing Chi Tivive
ICPR3
2010 Efficient SVM training with reduced weighted samples
abstract
This paper presents an efficient training approach for support vector machines that will improve their ability to learn from a large or imbalanced data set. Given an original training set, the proposed approach applies unsupervised learning to extract a smaller set of salient training exemplars, which are represented by weighted cluster centers and the target outputs. In subsequent supervised learning, the objective function is modified by introducing a weight for each new training sample and the corresponding penalty term. In this paper, we investigate two methods of defining the weight based on cluster vectors. The proposed SVM training is implemented and tested on two problems: (i) gender classification of facial images using the FERET data set; (ii) income prediction using the UCI Adult Census data set. Experiment results show that compared to standard SVM training, the proposed approach leads to much faster SVM training, produces a more compact classifier while maintaining generalization ability.
Giang Hoang Nguyen, Son Lam Phung, Abdesselam Bouzerdoum
IJCNN3
2010 Improved learning in grid-to-grid neural network via clustering
abstract
The maze traversal problem involves finding the shortest distance to the goal from any position in a maze. Such maze solving problems have been an interesting challenge in computational intelligence. Previous work has shown that grid-to-grid neural networks such as the cellular simultaneous recurrent neural network (CSRN) can effectively solve simple maze traversing problems better than other iterative algorithms such as the feedforward multi layer perceptron (MLP). In this work, we investigate improved learning for the CSRN maze solving problem by exploiting relevant information about the maze. We cluster parts of the maze using relevant state information and show an improvement in learning performance. We also study the effect of the number of clusters on the learning rate for the maze solving problem. Furthermore, we investigate a few code optimization techniques to improve the run time efficiency. The outcome of this research may have direct implication in rapid search and recovery, disaster planning and autonomous navigation among others.
William E. White, Khan M. Iftekharuddin, Abdesselam Bouzerdoum
IJCNN3
2010 Dimensionality reduction using compressed sensing and its application to a large-scale visual recognition task
abstract
This paper presents a novel algorithm for the dimensionality reduction which employs compressed sensing (CS) to improve the generalization capability of a classifier, especially for large-scale data. Compared to traditional dimensionality reduction methods, the proposed algorithm makes no use of the problem-dependent parameters, nor does it require additional computation for the eigenvalue decomposition like PCA or LDA. Mathematically, the derived algorithm regards the input features as the dictionary in CS, and selects the features that minimize the residual output error iteratively, thus the resulting features have a direct correspondence to the performance requirements of the given problem. Furthermore, the proposed algorithm can be regarded as a sparse classifier, which selects discriminative features and classifies the training data simultaneously. Experimentally, the CS-based algorithm is tested with a hierarchical visual pattern recognition architecture. The simulation results show that not only does the proposed method utilize only 25% of full features while achieving the test accuracy of the original full architecture, but also its performance is competitive when compared to existing dimensionality reduction methods.
Jie Yang 0009, Abdesselam Bouzerdoum, Fok Hing Chi Tivive, Son Lam Phung
IJCNN2
2009 A New Approach to Sparse Image Representation Using MMV and K-SVD
Jie Yang 0009, Abdesselam Bouzerdoum, Son Lam Phung
ACIVS2
2009 Vehicle Tracking Using Projective Particle Filter
abstract
This article introduces a new particle filtering approach for object tracking in video sequences. The projective particle filter uses a linear fractional transformation, which projects the trajectory of an object from the real world onto the camera plane, thus providing a better estimate of the object position. In the proposed particle filter, samples are drawn from an importance density integrating the linear fractional transformation. This provides a better coverage of the feature space and yields a finer estimate of the posterior density. Experiments conducted on traffic video surveillance sequences show that the variance of the estimated trajectory is reduced, resulting in more robust tracking.
Philippe Loic Marie Bouttefroy, Abdesselam Bouzerdoum, Son Lam Phung, Azeddine Beghdadi
AVSS2
2009 A car detection system based on hierarchical visual features
abstract
In this paper, we address the problem of detecting and localizing cars in still images. The proposed car detection system is based on a hierarchical feature detector in which the processing units are shunting inhibitory neurons. To reduce the training time and complexity of the network, the shunting inhibitory neurons in the first layer are implemented as directional nonlinear filters, whereas the neurons in the second layer have trainable parameters. A multi-resolution processing scheme is implemented so as to detect cars of different sizes, and to reduce the number of false positives during the detection stage, an adaptive thresholding strategy is developed. Tested on the UIUC car database, the proposed method achieves better classification results than some of the existing car detection approaches.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
CIMSIVP2
2009 A Neural Network pruning approach based on Compressive Sampling
abstract
The balance between computational complexity and the architecture bottlenecks the development of neural networks (NNs). An architecture that is too large or too small will influence the performance to a large extent in terms of generalization and computational cost. In the past, saliency analysis has been employed to determine the most suitable structure, however, it is time-consuming and the performance is not robust. In this paper, a family of new algorithms for pruning elements (weighs and hidden neurons) in neural networks is presented based on compressive sampling (CS) theory. The proposed framework makes it possible to locate the significant elements, and hence find a sparse structure, without computing their saliency. Experiment results are presented which demonstrate the effectiveness of the proposed approach.
Jie Yang 0009, Abdesselam Bouzerdoum, Son Lam Phung
IJCNN2
2008 A supervised learning approach for imbalanced data sets
abstract
This paper presents a new learning approach for pattern classification applications involving imbalanced data sets. In this approach, a clustering technique is employed to resample the original training set into a smaller set of representative training exemplars, represented by weighted cluster centers and their target outputs. Based on the proposed learning approach, four training algorithms are derived for feed-forward neural networks. These algorithms are implemented and tested on three benchmark data sets. Experimental results show that with the proposed learning approach, it is possible to design networks to tackle the class imbalance problem, without compromising the overall classification performance.
Giang Hoang Nguyen, Abdesselam Bouzerdoum, Son Lam Phung
ICPR2
2008 Efficient supervised learning with reduced training exemplars
abstract
In this article, we propose a new supervised learning approach for pattern classification applications involving large or imbalanced data sets. In this approach, a clustering technique is employed to reduce the original training set into a smaller set of representative training exemplars, represented by weighted cluster centers and their target outputs. Based on the proposed learning approach, two training algorithms are derived for feed-forward neural networks. These algorithms are implemented and tested on two pattern classification applications - skin detection and image classification. Experimental results show that with the proposed learning approach, it is possible to design networks in a fraction of time taken by the standard learning approach, without compromising the generalization ability and overall classification performance.
Giang Hoang Nguyen, Abdesselam Bouzerdoum, Son Lam Phung
IJCNN2
2008 A biologically inspired visual pedestrian detection system
abstract
In this paper, we present a biologically inspired method for detecting pedestrians in images. The method is based on a convolutional neural network architecture, which combines feature extraction and classification. The proposed network architecture is much simpler and easier to train than earlier versions. It differs from its predecessors in that the first processing layer consists of a set of pre-defined nonlinear derivative filters for computing gradient information. The subsequent processing layer has trainable shunting inhibitory feature detectors, which are used as inputs to a pattern classifier. The proposed pedestrian detection system is evaluated on the DaimlerChrysler pedestrian classification benchmark database and its performance is compared to the performance of support vector machines and Adaboost classifiers.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
IJCNN2
2008 A hierarchical learning network for face detection with in-plane rotation
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
Neurocomputing2
2007 Detecting People in Images: An Edge Density Approach
abstract
In this paper, we present a new method for detecting visual objects in digital images and video. The novelty of the proposed method is that it differentiates objects from non-objects using image edge characteristics. Our approach is based on a fast object detection method developed by Viola and Jones. While Viola and Jones use Harr-like features, we propose a new image feature - the edge density - that can be computed more efficiently. When applied to the problem of detecting people and pedestrians in images, the new feature shows a very good discriminative capability compared to the Harr-like features.
Son Lam Phung, Abdesselam Bouzerdoum
ICASSP (1)2
2007 A Nonlinear Feature Extractor for Texture Segmentation
abstract
This article presents a feed-forward network architecture that can be used as a nonlinear feature extractor for texture segmentation. It comprises two layers of feature extraction units; each layer is arranged into several planes, called feature maps. The features extracted from the second layer are used as the final texture features. The feature maps are characterised by a set of masks (or weights), which are shared among all the units of a single feature map. Combining the nonlinear feature extractor with a classifier, we have developed a texture segmentation system that does not rely on pre-defined filters for feature extraction; the weights of the feature maps are found during a supervised learning stage. Tested on the Brodatz texture images, the proposed texture segmentation system achieves better classification accuracy than some of the most popular texture segmentation approaches.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
ICIP (2)2
2007 A Pyramidal Neural Network For Visual Pattern Recognition
abstract
In this paper, we propose a new neural architecture for classification of visual patterns that is motivated by the two concepts of image pyramids and local receptive fields. The new architecture, called pyramidal neural network (PyraNet), has a hierarchical structure with two types of processing layers: Pyramidal layers and one-dimensional (1-D) layers. In the new network, nonlinear two-dimensional (2-D) neurons are trained to perform both image feature extraction and dimensionality reduction. We present and analyze five training methods for PyraNet [gradient descent (GD), gradient descent with momentum, resilient back-propagation (RPROP), Polak-Ribiere conjugate gradient (CG), and Levenberg-Marquadrt (LM)] and two choices of error functions [mean-square-error (mse) and cross-entropy (CE)]. In this paper, we apply PyraNet to determine gender from a facial image, and compare its performance on the standard facial recognition technology (FERET) database with three classifiers: The convolutional neural network (NN), the k-nearest neighbor (k-NN), and the support vector machine (SVM).
Son Lam Phung, Abdesselam Bouzerdoum
IEEE Trans. Neural Networks2
2006 Gender Classification Using a New Pyramidal Neural Network
Son Lam Phung, Abdesselam Bouzerdoum
ICONIP (2)2
2006 Rotation Invariant Face Detection Using Convolutional Neural Networks
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
ICONIP (2)2
2006 A Gender Recognition System using Shunting Inhibitory Convolutional Neural Networks
abstract
In this paper, we employ shunting inhibitory convolutional neural networks to develop an automatic gender recognition system. The system comprises two modules: a face detector and a gender classifier. The human faces are first detected and localized in the input image. Each detected face is then passed to the gender classifier to determine whether it is a male or female. Both the face detection and gender classification modules employ the same neural network architecture; however, the two modules are trained separately to extract different features for face detection and gender classification. Tested on two different databases, Web and BioID database, the face detector has an average detection accuracy of 97.9%. The gender classifier, on the other hand, achieves 97.2% classification accuracy on the FERET database. The combined system achieves a recognition rate of 85.7% when tested on a large set of digital images collected from the Web and BioID face databases.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
IJCNN2
2006 Application of SiCONnets to Handwritten Digit Recognition
abstract
In this paper, we apply a new neural network model, namely shunting inhibitory convolutional neural networks, or SICoNNets for short, to the problem of handwritten digit recognition. This type of networks has a generic and flexible architecture, where the processing is based on the physiologically plausible mechanism of shunting inhibition. A hybrid first-order training method, called QRProp, is developed based on the three training algorithms Rprop, Quickprop, and SuperSAB. The MNIST database is used to train and evaluate the performance of SICoNNets in handwritten digit recognition. A network with 24 feature maps and 2722 free parameters achieves a recognition accuracy of 97.3%.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
Int. J. Comput. Intell. Appl.2
2005 Skin Segmentation Using Color Pixel Classification: Analysis and Comparison
abstract
This paper presents a study of three important issues of the color pixel classification approach to skin segmentation: color representation, color quantization, and classification algorithm. Our analysis of several representative color spaces using the Bayesian classifier with the histogram technique shows that skin segmentation based on color pixel classification is largely unaffected by the choice of the color space. However, segmentation performance degrades when only chrominance channels are used in classification. Furthermore, we find that color quantization can be as low as 64 bins per channel, although higher histogram sizes give better segmentation performance. The Bayesian classifier with the histogram technique and the multilayer perceptron classifier are found to perform better compared to other tested classifiers, including three piecewise linear classifiers, three unimodal Gaussian classifiers, and a Gaussian mixture classifier.
Son Lam Phung, Abdesselam Bouzerdoum, Douglas Chai
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Efficient training algorithms for a class of shunting inhibitory convolutional neural networks
abstract
This article presents some efficient training algorithms, based on first-order, second-order, and conjugate gradient optimization methods, for a class of convolutional neural networks (CoNNs), known as shunting inhibitory convolution neural networks. Furthermore, a new hybrid method is proposed, which is derived from the principles of Quickprop, Rprop, SuperSAB, and least squares (LS). Experimental results show that the new hybrid method can perform as well as the Levenberg-Marquardt (LM) algorithm, but at a much lower computational cost and less memory storage. For comparison sake, the visual pattern recognition task of face/nonface discrimination is chosen as a classification problem to evaluate the performance of the training algorithms. Sixteen training algorithms are implemented for the three different variants of the proposed CoNN architecture: binary-, Toeplitz- and fully connected architectures. All implemented algorithms can train the three network architectures successfully, but their convergence speed vary markedly. In particular, the combination of LS with the new hybrid method and LS with the LM method achieve the best convergence rates in terms of number of training epochs. In addition, the classification accuracies of all three architectures are assessed using ten-fold cross validation. The results show that the binary- and Toeplitz-connected architectures outperform slightly the fully connected architecture: the lowest error rates across all training algorithms are 1.95% for Toeplitz-connected, 2.10% for the binary-connected, and 2.20% for the fully connected network. In general, the modified Broyden-Fletcher-Goldfarb-Shanno (BFGS) methods, the three variants of LM algorithm, and the new hybrid/LS method perform consistently well, achieving error rates of less than 3% averaged across all three architectures.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
IEEE Trans. Neural Networks2
2004 Naive bayes face/nonface classifier: a study of preprocessing and feature extraction techniques
abstract
This paper presents a classifier of face and nonface patterns that is based on the naive Bayes model. Using this classifier as a tool. We analyze the effects on classification performance of preprocessing, feature extraction and classifier combination techniques. Our analysis shows that image normalization techniques that reduce the effects of different lighting conditions improve face-nonface classification significantly. In addition, techniques such as background masking and combining classifiers that use different feature vectors are shown to enhance classification performance. Over a test set of 12,000 patterns, the combined classifier using four feature vectors has correct detection rates (CDRs) of 96.2% and 99.2% at false detection rates (FDRs) of 1% and 5%, respectively.
Son Lam Phung, Abdesselam Bouzerdoum, Douglas Chai, Anthony Watson
ICIP2
2004 A face detection system using shunting inhibitory convolutional neural networks
abstract
We present a face detection system based on a class of convolutional neural networks, namely shunting inhibitory convolutional neural networks (SICoNNets). The topology of these networks is a flexible feedforward architecture with three different connections schemes: fully-connected, toeplitz-connected and binary-connected. SICoNNets were trained, using a hybrid method based on Rprop, Quickprop and least squares, to discriminate between face and non-face patterns. All three connection schemes achieve 99% detection accuracy at 5% false alarm rate, based on a test set of 7000 face and non-face patterns. Furthermore, toeplitz-connected network was trained on a larger training set and has achieved a 99% correct classification rate with only 1% false alarm rate based on the same test set. A face detection system is built based on the trained convolutional neural networks. The system accepts an input image of arbitrary size and localizes the face patterns in the image. To localize faces of different sizes, the convolutional neural network is applied as a face detection filter at different scales. The detection scores from different scales are aggregated together to form the final decision.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
IJCNN2
2003 A Generalized Feedforward Neural Network Architecture and Its Training Using Two Stochastic Search Methods
Abdesselam Bouzerdoum, Rainer Mueller
GECCO1
2003 Adaptive skin segmentation in color images
abstract
A new skin segmentation technique for color images is proposed. The proposed technique uses a human skin color model that is based on the Bayesian decision theory and developed using a large training set of skin colors and nonskin colors. The proposed technique is novel and unique in that texture characteristics of the human skin are used to select appropriate skin color thresholds for skin segmentation. Two homogeneity measures for skin regions that take into account both global and local image features are also proposed. Experimental results showed that the proposed technique can achieve good skin segmentation performance (false detection rate of 4.5% and false rejection rate of 4.0%).
Son Lam Phung, Douglas Chai, Abdesselam Bouzerdoum
ICASSP (3)3
2003 Adaptive skin segmentation in color images
abstract
A new skin segmentation technique for color images is proposed. The proposed technique uses a human skin color model that is based on the Bayesian decision theory and developed using a large training set of skin colors and nonskin colors. The proposed technique is novel and unique in that texture characteristics of the human skin are used to select appropriate skin color thresholds for skin segmentation. Two homogeneity measures for skin regions that take into account both global and local image features are also proposed. Experimental results showed that the proposed technique can achieve good skin segmentation performance (false detection rate of 4.5% and false rejection rate of 4.0%).
Son Lam Phung, Douglas Chai, Abdesselam Bouzerdoum
ICME3
2003 A generalized feedforward neural network classifier
abstract
In this article a new generalized feedforward neural network (GFNN) architecture for pattern classification is proposed. The GFNNs are an expansion of shunting inhibitory artificial neural networks (SIANNs), proposed previously for classification and function approximations. The GFNN architecture uses as its basic computing unit the generalized shunting neuron (GSN), which includes as special cases the perceptron and the shunting inhibitory neuron. Generalized shunting neurons are capable of forming complex, nonlinear decision boundaries. This allows the GFNN architecture to learn complex pattern classification problems using few neurons. In this article, GFNNs are applied to several benchmark classification problems, and their performance compared to the performance of SIANNs and multilayer perceptrons.
Ganesh Arulampalam, Abdesselam Bouzerdoum
IJCNN2
2003 A new class of convolutional neural networks (SICoNNets) and their application of face detection
abstract
Artificial neural networks (ANNs), evolved from biological insights, have equipped computers with the capacity to actually learn from examples using real world data. With this remarkable ability, ANNs are able to extract patterns and detect trends that are too complex to be noticed or perceived by either humans or classical computer techniques. Nevertheless, as the amount of data to be processed increases significantly there is a demand for developing other types of artificial neural networks to perform complex pattern recognition tasks. In this article, a new class of convolutional neural networks, namely shunting inhibitory convolutional neural networks (SICoNNets), is introduced, and a training algorithm is developed using supervised learning based on resilient backpropagation with momentum. Three different network topologies, ranging from fully-connected to partially-connected, are implemented and trained to discriminate between face and non-face patterns. All three architectures achieve more than 96% correct face classification; the best architecture achieves 97.6% correct face classification at a false alarm rate of 3.4%.
Fok Hing Chi Tivive, Abdesselam Bouzerdoum
IJCNN2
2003 A generalized feedforward neural network architecture for classification and regression
Ganesh Arulampalam, Abdesselam Bouzerdoum
Neural Networks2
2002 A novel skin color model in YCbCr color space and its application to human face detection
abstract
This paper presents a new human skin color model in YCbCr color space and its application to human face detection. Skin colors are modeled by a set of three Gaussian clusters, each of which is characterized by a centroid and a covariance matrix. The centroids and covariance matrices are estimated from large set of training samples after a k-means clustering process. Pixels in a color input image can be classified into skin or non-skin based on the Mahalanobis distances to the three clusters. Efficient post-processing techniques namely noise removal, shape criteria, elliptic curve fitting and face/non-face classification are proposed in order to further refine skin segmentation results for the purpose of face detection.
Son Lam Phung, Abdesselam Bouzerdoum, Douglas Chai
ICIP (1)2
2000 Foreground/Background Bit Allocation for Region-of-Interest Coding
abstract
Two bit allocation strategies, namely, maximum bit transfer (MET) and joint bit assignment (JBA) are proposed for region-of-interest coding. The MBT strategy uses a pair of quantizers to facilitate maximum bit transfer from the background to foreground image region. It assigns the highest quantization parameter to the background quantizer, and then determines the finest value of foreground quantizer that can be used without increasing the bit rate. In this approach, the background is always encoded with the coarsest quantization level, but this is not always desirable. Therefore, the JBA strategy can be used instead. It has two modes: user defined and automatic. In the user defined mode, the user can adjust bit consumption using a scale that ranges from "not coding foreground" to "coding foreground only". In the automatic mode, bit allocation is based on the characteristics of the image regions, which include size, motion and priority. The coding results showed that improved image quality was achieved at the same bit rate.
Douglas Chai, King Ngi Ngan, Abdesselam Bouzerdoum
ICIP3
2000 Supervised Texture Segmentation using DWT and a Modified K-NN Classifier
abstract
We present a texture segmentation scheme based on the discrete wavelet transform (DWT). The DWT is a non-redundant representation which can reduce computational complexity in the processing. The texture segmentation scheme presented here consists of three steps: feature extraction, conditioning, and clustering. For feature conditioning, a number of smoothing windows have been tested. Clustering is performed with a modified k-nearest neighbour clustering algorithm. The proposed scheme consistently achieves error rates of less than 10% with the best average error of 5.62%.
Brian W. Ng, Abdesselam Bouzerdoum
ICPR2
2000 Classification and Function Approximation Using Feed-Forward Shunting Inhibitory Artificial Neural Networks
abstract
In this article we propose a new class of artificial neural networks for classification and function approximation. These networks are referred to as shunting inhibitory artificial neural networks (SIANN). A SIANN consists of one or more hidden layers comprised of shunting neurons, the outputs of which are combined linearly to form the desired output. The basic synaptic interaction of the hidden units is shunting inhibition. Due to the inherent nonlinearity mediated by shunting inhibition, SIANN networks are capable of constructing a large repertoire of decision surfaces, ranging from simple hyperplanes to very complex nonlinear hypersurfaces. Therefore, developing efficient training algorithms for these networks should simplify the design of very powerful classifiers and function approximators. In this paper some examples of complex decision regions formed by SIANN are illustrated. Furthermore, a method for training feedforward SIANN is developed based on the error backpropagation algorithm. Finally, simulation results which illustrate the performance of SIANN in function approximation and classification tasks are presented and compared with results obtained from multilayer perceptron networks.
Abdesselam Bouzerdoum
IJCNN (6)1
2000 A high fill-factor native logarithmic pixel: Simulation, design and layout optimization
abstract
In this paper we investigate important issues in the design of the logarithmic CMOS pixel. In particular, much attention is paid to the optimization of pixel performance in terms of output gain, dynamic range, and fill-factor. In order to increase the gain-bandwidth product, we propose to use the native transistor as source follower. The performance of such a pixel is compared with that of conventional logarithmic pixels. It is shown that the native source follower yields a significant increase in the gain-bandwidth product. In addition, we propose a layout floor-planning strategy which allows us to achieve a 46% fill-factor. In order to compare the performance of the proposed pixel with the conventional NMOS and PMOS logarithmic pixels, a VLSI prototype has been realized using 0.7 /spl mu/m CMOS technology.
Amine Bermak, Abdesselam Bouzerdoum, Jason Kamran Eshraghian
ISCAS2
2000 CMOS circuit for high-speed flexible read-out of CMOS imagers
Amine Bermak, Abdesselam Bouzerdoum, Jason Kamran Eshraghian, Jean L. Noullet
VCIP2
2000 Digital implementation of shunting-inhibitory cellular neural network
Tarik Hammadou, Abdesselam Bouzerdoum, Amine Bermak
VCIP2
2000 Supervised texture segmentation using DT-CWT and a modified k-NN classifier
Brian W. Ng, Abdesselam Bouzerdoum
VCIP2
1999 A Novel Fingerprint Image Compression Technique Using the Wavelet Transform and Piecewise-Uniform Pyramid Lattice Vector Quantisation
abstract
A novel compression algorithm for fingerprint images is introduced. Using wavelet packets and lattice vector quantisation, a new vector quantisation scheme based on an accurate model for the distribution of the wavelet coefficients is presented. In the new algorithm, no assumptions are made about the lattice parameters and no training and multi-quantising are required. The proposed algorithms achieve the best rate-distortion performance by adapting to the statistical characteristics of the source image in each sub-image. Compared to other available image compression algorithms, the proposed algorithms result in higher quality reconstructed images for identical bit rates.
Mohamed Deriche 0001, Shohreh Kasaei, Abdesselam Bouzerdoum
ICIP (3)3
1998 From Receptive Field to Connection Strengths
Nicolangelo Iannella, Abdesselam Bouzerdoum, Shigeru Tanaka
ICONIP2
1998 A combined quadratic optimization/median filtering technique for image restoration
abstract
Recently, we have introduced a new iterative technique for signal and image restoration. The technique solves a bound-constrained quadratic optimization problem with preconditioning. Although this technique yields a good solution in just few iterations, the error starts increasing after it reaches a minimum. This is a common problem in iterative techniques where the first few iterations restore the low frequency components of the signal and as the number of iterations increases the algorithm attempts to restore the high frequency components, which are dominated by noise. In this article we combine the proposed new iterative technique with median filtering. Median filtering helps maintain a low error by preserving the edge information while reducing the high frequency noise.
Abdesselam Bouzerdoum
SMC1
1998 Illumination invariant face recognition
abstract
Few of the face recognition methods reported in the literature are capable of recognising faces under varying illumination conditions. The paper discusses a method which can achieve a higher recognition rate than those obtained for existing methods. The novelty of this new method is the use of an embossing technique to process a face image before presenting it to a standard face recognition system. Using a large database of face images, the performance of the proposed method is evaluated by comparing it against the performances of three existing methods. The experimental results demonstrate the successfulness of the proposed method.
Abbas Z. Kouzani, Fangpo He, Karl Sammut, Abdesselam Bouzerdoum
SMC4
1996 Evaluation of biologically inspired motion detection systems as a basis for local motion processing systems
abstract
The mechanisms employed to detect motion by a variety of biological systems have been investigated for many years and a number of models that explain different aspects of these systems have been developed. Most of the models have been developed to explain aspects of wide field operation of motion detector arrays. This paper investigates suitability of these model as a front end to a system that requires local motion information. Results of simulations of different motion detection models using real scenes captured using a video camera are also presented.
Richard Beare, Abdesselam Bouzerdoum
ICIP (3)2
1995 Neural Betwork Region-Based Inexact Relational Medical Image Matching
Steven M. Vajdic, Abdesselam Bouzerdoum, Michael J. Brooks, A. R. Downing, H. E. Katz
IEA/AIE2
1995 Neural Network Architectures: Medical Image Segmentation Approach as a Step Towards Medical Image Fusion
Steven M. Vajdic, Abdesselam Bouzerdoum, Michael J. Brooks, A. R. Downing, H. E. Katz
IEA/AIE2
1994 A neural architecture for hierarchical clustering
abstract
An hierarchical neural network structure for clustering problems is presented and a statistical analysis of its performance is conducted. This neural network architecture aims to find, through competition and cooperation, maximally related objects in a scene. The architecture was first introduced by Maren and Ali (1983), and was named the hierarchical scene structure (HSS). We propose an enhancement of the original HSS and demonstrate that this leads to an improved performance. It is also shown that further improvement in performance can be achieved by cascading two enhanced HSS networks.>
Abdesselam Bouzerdoum, Michael L. Southcott, Jihan Zhu, Robert E. Bogner
ICASSP (2)1
1994 Dual-Purpose Interpretation of Sensory Information
abstract
Fully autonomous mobile robots often rely on an array of sensors to provide them with an adequate picture of their environment. Furthermore, these systems tend to have strict limitations in terms of available processing capability. Hence the so-called "smart sensing" approach is particularly appropriate as it combines a small size with a reduced requirement for interpretation of sensory input. This paper describes how a visual micro-sensor implemented in VLSI can be used for obstacle avoidance as well as localised path planning or navigation.>
Andre Yakovleff, X. Thong Nguyen, Abdesselam Bouzerdoum, Alireza Moini, Robert E. Bogner, Jason Kamran Eshraghian
ICRA3
1993 Neural network for quadratic optimization with bound constraints
abstract
A recurrent neural network is presented which performs quadratic optimization subject to bound constraints on each of the optimization variables. The network is shown to be globally convergent, and conditions on the quadratic problem and the network parameters are established under which exponential asymptotic stability is achieved. Through suitable choice of the network parameters, the system of differential equations governing the network activations is preconditioned in order to reduce its sensitivity to noise and to roundoff errors. The optimization method employed by the neural network is shown to fall into the general class of gradient methods for constrained nonlinear optimization and, in contrast with penalty function methods, is guaranteed to yield only feasible solutions.
Abdesselam Bouzerdoum, Tim R. Pattison
IEEE Trans. Neural Networks1
1990 A shunting inhibitory motion detector that can account for the functional characteristics of fly motion-sensitive interneurons
abstract
The multiplicative inhibitory motion detector (MIMD) has been proposed by the authors (Visual Communications and Image Processing IV, Proceedings of SPIE. 1199, 1229-1240, 1989). The applicability of this motion detector to the activity of motion-sensitive interneurons of the lobula plate, the posterior pat of the third visual ganglion in the fly's optic lobe, is investigated. In particular, it is demonstrated that an array of MIMDs can simulate the characteristics of transient and steady-state response and of contrast sensitivity functions for these neurons
Abdesselam Bouzerdoum, R. B. Pinter
IJCNN1