EDBT 2026 Demo / reviewers in the wild / expert
Son Lam Phung
dblp:93/737
· DBLP profile ↗
87ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0002-3076-0540ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 59 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 37 · 4 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NaviFormer: Multimodal scene segmentation for assistive navigationabstractSafe and independent navigation is an essential part of daily life for individuals with visual impairments. Recent advances in assistive navigation have leveraged deep learning-based semantic scene segmentation with promising results by using RGB images. In comparison, RGB-D images provide richer geometric and appearance information, but their integration into deep learning frameworks for assistive navigation could be further explored to enhance segmentation accuracy. In modeling cross-modal interactions, the existing RGB-D segmentation methods tend to lose fine-grained spatial details or discard useful information, which hinders performance on dense prediction tasks. To address these gaps, we propose NaviFormer, a novel RGB-D semantic segmentation architecture tailored for assistive navigation. NaviFormer features a dual-stream transformer encoder with shared weights to efficiently extract latent features from RGB and depth modalities. It also incorporates a Local-Global Cross-Modal Fusion module, which facilitates effective information exchange between the two modalities across both local and global feature levels. In training NaviFormer, we further employ a pixel-wise contrastive loss to enhance the separability of pixel-level embeddings in the RGB-D feature space. Extensive experiments on TrueSight and Cityscapes datasets indicate that NaviFormer achieves superior performance compared to existing RGB-D segmentation methods. Our findings highlight the importance of leveraging RGB-D data for enhancing semantic understanding in assistive navigation systems, and establish NaviFormer as a solid baseline for future research in this domain. Ly Bui, Son Lam Phung, Yang Di, Soan Thi Minh Duong, Abdesselam Bouzerdoum |
Comput. Vis. Image Underst. | 2 |
| 2026 | CSN: A compact semantic segmentation network for visual scene perception in assistive navigationabstractAccuracy and efficiency are essential in assistive navigation algorithms to ensure accessibility and reliability for visually impaired individuals. However, existing deep learning models often require substantial computational resources to achieve high accuracy, making them impractical for deployment on mobile devices. To address this problem, we introduce CSN, a compact semantic segmentation network designed for assistive navigation, optimizing performance in resource-constrained environments. With CSN, we introduce two innovative modules, the cascaded atrous multi-scale enhancement (CAME) layer and the dual-path residual bottleneck (DPRB) block. The CAME layer efficiently enhances multi-scale representation through feature resampling, while the DPRB block improves feature refinement with minimal computational cost. These modules enable CSN to achieve robust and reliable segmentation across diverse and complex pedestrian environments. The proposed approach achieves the best result on the challenging TrueSight dataset, demonstrating superior prediction accuracy and computational efficiency compared to state-of-the-art lightweight models. CSN achieves a mean intersection over union of 60.99%, while maintaining a low computational cost of 83.46 giga floating-point operations, a compact model size of 8.41 million parameters, and a real-time inference speed of 52.59 frames per second. Yunjia Lei, Son Lam Phung, Yang Di, Abdesselam Bouzerdoum |
Comput. Vis. Image Underst. | 2 |
| 2025 | Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual GroundingabstractVisual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features corresponding to textual descriptions, improving model transparency and trustworthiness for wider adoption of deep learning models in clinical practice. Current models struggle to associate textual descriptions with disease regions due to inefficient attention mechanisms and a lack of fine-grained token representations. In this paper, we empirically demonstrate two key observations. First, current VLMs assign high norms to background tokens, diverting the model's attention from regions of disease. Second, the global tokens used for cross-modal learning are not representative of local disease tokens. This hampers identifying correlations between the text and disease tokens. To address this, we introduce simple, yet effective Disease-Aware Prompting (DAP) process, which uses the explainability map of a VLM to identify the appropriate image features. This simple strategy amplifies disease-relevant regions while suppressing background interference. Without any additional pixel-level annotations, DAP improves visual grounding accuracy by 20.74% compared to state-of-the-art methods across three major chest X-ray datasets. Ta Duc Huy, Duy Anh Huynh, Yutong Xie 0001, Yuankai Qi, Qi Chen 0014, Phi-Le Nguyen, Sen Kim Tran, Son Lam Phung, Anton van den Hengel, Zhibin Liao, Minh-Son To, Johan Verjans, Vu Minh Hieu Phan |
ICCV | 8 |
| 2025 | Localizing Before Answering: A Benchmark for Grounded Medical Visual Question AnsweringabstractMedical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate localization reasoning. This work reveals a critical limitation in current medical LMMs: instead of analyzing relevant pathological regions, they often rely on linguistic patterns or attend to irrelevant image areas when responding to disease-related queries. To address this, we introduce HEAL-MedVQA (Hallucination Evaluation via Localization MedVQA), a comprehensive benchmark designed to evaluate LMMs' localization abilities and hallucination robustness. HEAL-MedVQA features (i) two innovative evaluation protocols to assess visual and textual shortcut learning, and (ii) a dataset of 67K VQA pairs, with doctor-annotated anatomical segmentation masks for pathological regions. To improve visual reasoning, we propose the Localize-before-Answer (LobA) framework, which trains LMMs to localize target regions of interest and self-prompt to emphasize segmented pathological areas, generating grounded and reliable answers. Experimental results demonstrate that our approach significantly outperforms state-of-the-art biomedical LMMs on the challenging HEAL-MedVQA benchmark, advancing robustness in medical VQA. Minh Khoi Ho, Ta Duc Huy, Thanh Tam Nguyen, Qi Chen 0014, Kumar Rav, Quy Duong Dang, Satwik Ramchandre, Son Lam Phung, Zhibin Liao, Minh-Son To, Johan Verjans, Phi-Le Nguyen, Vu Minh Hieu Phan |
IJCAI | 9 |
| 2025 | AMT-Net: Attention-based multi-task network for scene depth and semantics prediction in assistive navigationabstractTraveling safely and independently in unfamiliar environments remains a significant challenge for people with visual impairments. Conventional assistive navigation systems, while aiming to enhance spatial awareness, typically handle crucial tasks like semantic segmentation and depth estimation separately, resulting in high computational overhead and reduced inference speed. To address this limitation, we introduce AMT-Net, a novel multi-task deep neural network designed for joint semantic segmentation and monocular depth estimation. The AMT-Net is designed with a single unified decoder, which boosts not only the model’s efficiency but also its scalability on portable devices with limited computational resources. We propose two self-attention-based modules, CSAPP and RSAB, to leverage the strengths of convolutional neural networks for extracting robust local features and Transformers for capturing essential long-range dependencies. This design enhances the ability of our model to interpret complex scenes effectively. Furthermore, AMT-Net has low computational complexity and achieves real-time performance, making it suitable for assistive navigation applications. Extensive experiments on the public NYUD-v2 dataset and the TrueSight dataset demonstrated our model’s state-of-the-art performance and the effectiveness of the proposed components. Yunjia Lei, Joshua Luke Thompson, Son Lam Phung, Abdesselam Bouzerdoum, Hoang Thanh Le 0001 |
Neurocomputing | 3 |
| 2024 | MSSP: A Multi-view Benchmark for Street Scene Perception in Assistive NavigationabstractFinding safe paths in autonomous or assistive navigation systems is a challenging task. In this paper, we introduce MSSP, a novel multi-perspective street scene perception benchmark dataset for navigation assistance. Compared to single-view perception, multi-view provides more comprehensive information about the surroundings to the pedestrians for enhancing safety. The MSSP dataset includes 2,044 samples from complex scenes in both first-person and third-person perspectives, covering four categories: Sidewalk, Traffic lane, Verge, and Lawn. This dataset primarily focuses on pedestrian navigation assistance and is the first pedestrian multi-perspective dataset. Furthermore, in comparison to single-perspective datasets in other studies, MSSP contains more categories. The dataset supports pixel-level semantic segmentation approaches. We also provide experimental results of recent semantic segmentation methods on this dataset for evaluation. Specifically, the transformer-based SegFormer method outperforms others on MSSP. Our dataset is accessible through: https://github.com/yangdi-cv/MSSP. Yang Di, Son Lam Phung, Abdesselam Bouzerdoum |
IJCNN | 2 |
| 2024 | Bio-DETR: A Transformer-based Network for Pest and Seed Detection with Hyperspectral ImagesabstractExotic pests and seeds pose a serious threat to agricultural production and ecosystems, leading to significant labour and economic losses. Therefore, automated pest and seed detection systems are crucial for biosecurity and agriculture. Biosecurity detection systems need to identify pests and seeds of various sizes and types in a variable and complex environment, making high-precision automated detection a challenging task. To address this, we propose Bio-DETR, a lightweight transformer-based architecture for accurate pest and seed detection. We introduce two self-attention-based modules, Hybrid Scale Attention and Dynamic Bilateral Attention, for enhanced feature extraction and multiscale information fusion. The effectiveness of these modules is validated experimentally. We also propose HSI-Bio, a large-scale dataset with 8,000 images across 23 categories, collected using a hyperspectral camera on diverse backgrounds. Compared to RGB images, hyperspectral images (HSI) offer rich channel information. The representative spectra are selected from HSI for experiments. Bio-DETR achieves an AP50of 87.4% and an AP of 62.2% on HSI-Bio, outperforming other state-of-the-art methods and achieving real-time detection of 52 FPS. Our code is available at: https://github.com/yangdi-cv/Bio-DETR. Yang Di, Son Lam Phung, Julian Van Den Berg, Jason Clissold, Ly Bui, Hoang Thanh Le 0001, Abdesselam Bouzerdoum |
IJCNN | 2 |
| 2024 | Structural Attention: Rethinking Transformer for Unpaired Medical Image Synthesis
Vu Minh Hieu Phan, Yutong Xie 0001, Bowen Zhang 0009, Yuankai Qi, Zhibin Liao, Antonios Perperidis, Son Lam Phung, Johan Verjans, Minh-Son To |
MICCAI (7) | 7 |
| 2024 | UOW-Vessel: A Benchmark Dataset of High-Resolution Optical Satellite Images for Vessel Detection and SegmentationabstractIn this paper, we introduce UOW-Vessel, a benchmark dataset of high-resolution optical satellite images for vessel detection and segmentation. Our dataset consists of 3,500 images, collected from 14 countries across 4 continents. With a total of 35,598 instances in 10 vessel categories, UOW-Vessel is to date the largest satellite image dataset for vessel recognition. Furthermore, compared to the existing public datasets that only provide bounding box ground-truth, our new dataset offers more accurate polygon annotations of vessel objects. This dataset is expected to support instance segmentation-based approaches, which is a less investigated area in vessel surveillance. We also report extensive evaluations of the recent algorithms for instance segmentation on the new benchmark dataset. Ly Bui, Son Lam Phung, Yang Di, Hoang Thanh Le 0001, Tran Thanh Phong Nguyen, Sandy Burden, Abdesselam Bouzerdoum |
WACV | 2 |
| 2024 | DAP: A dataset-agnostic predictor of neural network performanceabstractTraining a deep neural network on a large dataset to convergence is a time-demanding task. This task often must be repeated many times, especially when developing a new deep learning algorithm or performing a neural architecture search. This problem can be mitigated if a deep neural network’s performance can be estimated without actually training it. In this work, we investigate the feasibility of two tasks: (i) predicting a deep neural network’s performance accurately given only its architectural descriptor, and (ii) generalizing the predictor across different datasets without re-training. To this end, we propose a dataset-agnostic regression framework that uses a novel dual-LSTM model and a new dataset difficulty feature. The experimental results show that both tasks above are indeed feasible, and the proposed method outperforms the existing techniques in all experimental cases. Additionally, we also demonstrate several practical use-cases of the proposed predictor. Sui Paul Ang, Soan Thi Minh Duong, Son Lam Phung, Abdesselam Bouzerdoum |
Neurocomputing | 3 |
| 2024 | Multi-camera multi-object tracking on the move via single-stage global association approach
Pha A. Nguyen, Kha Gia Quach, Chi Nhan Duong, Son Lam Phung, T. Hoang Ngan Le, Khoa Luu |
Pattern Recognit. | 4 |
| 2024 | AdaptorNAS: A New Perturbation-Based Neural Architecture Search for Hyperspectral Image SegmentationabstractHyperspectral image segmentation is an emerging area with numerous applications, including agriculture, forestry, environment monitoring, and remote sensing. This paper proposes a new neural architecture search algorithm, named AdaptorNAS, for hyperspectral image segmentation. AdaptorNAS aims to design the optimum decoder for any given encoder. In our approach, the search space of AdaptorNAS is a large deep neural network (DNN), and the optimal decoder is derived by pruning the large DNN via a perturbation-based pruning strategy. Verified on three popular encoders, i.e., ResNet-34, MobileNet-V2, and EfficientNet-B2, AdaptorNAS can design high-speed decoders that are significantly better than six common hand-crafted decoders. Additionally, with the EfficientNet-B2 encoder, AdaptorNAS (mIoU of 92.47% and mDice of 95.15%) outperforms the state-of-the-art NAS algorithms and hand-crafted network architectures on the hyperspectral image segmentation task. We also introduce a new hyperspectral image dataset of 4,625 images for objective evaluation in hyperspectral image segmentation research. Sui Paul Ang, Son Lam Phung, Ly Bui, Abdesselam Bouzerdoum |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Bilateral Coarse-to-Fine Network for Point Cloud CompletionabstractPoint cloud completion aims to accurately estimate complete point clouds from partial observations. Existing methods of-ten directly infer the missing points from the partial shape, but they suffer from limited structural information. To address this, we propose the Bilateral Coarse-to-Fine Network (BCF-Net), which leverages 2D images as guidance to compensate for structural information loss. Our method introduces a multi-level codeword skip-connection to estimate structural details. Experimental results show that BCF-Net outperforms state-of-the-art point cloud completion networks on synthetic and real-world datasets. Tran Thanh Phong Nguyen, Son Lam Phung, Vinod Gopaldasani, Jane Whitelaw |
ICASSP | 2 |
| 2023 | TP-YOLO: A Lightweight Attention-Based Architecture for Tiny Pest DetectionabstractAutomatic detection of agricultural pests is a challenging problem that is of great interest in biosecurity and precision agriculture. The detection model must cope well with the dense distribution of small-sized pests in complex backgrounds. This paper proposes a lightweight attention-based network, called TP-YOLO, for tiny pest detection. We introduce two attention-based components, namely Contextual Transformer and Omni-Dimensional Dynamic Convolution modules, to enhance feature extraction. The proposed modules are integrated into the YOLOv8 backbone, a state-of-the-art baseline for object detection. This paper also introduces a new benchmark dataset consisting of 1,600 images of Khapra beetles for objective evaluation of pest detection algorithms. Extensive experiments on two datasets indicate that TP-YOLO achieves competitive detection accuracy while having a significantly smaller model size and fast prediction time. We have made the code available to the public at: https://github.com/yangdi-cv/TP-YOLO. Yang Di, Son Lam Phung, Julian Van Den Berg, Jason Clissold, Abdesselam Bouzerdoum |
ICIP | 2 |
| 2023 | MSD-NAS: multi-scale dense neural architecture search for real-time pedestrian lane detectionabstractAbstract Accurate detection of pedestrian lanes is a crucial criterion for vision-impaired people to navigate freely and safely. The current deep learning methods have achieved reasonable accuracy at this task. However, they lack practicality for real-time pedestrian lane detection due to non-optimal accuracy, speed, and model size trade-off. Hence, an optimized deep neural network (DNN) for pedestrian lane detection is required. Designing a DNN from scratch is a laborious task that requires significant experience and time. This paper proposes a novel neural architecture search (NAS) algorithm, named MSD-NAS, to automate this laborious task. The proposed method designs an optimized deep network with multi-scale input branches, allowing the derived network to utilize local and global contexts for predictions. The search is also performed in a large and generic space that includes many existing hand-designed network architectures as candidates. To further boost performance, we propose a Short-term Visual Memory mechanism to improve information facilitation within the derived networks. Evaluated on the PLVP3 dataset of 10,000 images, the DNN designed by MSD-NAS achieves state-of-the-art accuracy (0.9781) and mIoU (0.9542), while being 20.16 times faster and 2.56 times smaller than the current best deep learning model. Sui Paul Ang, Son Lam Phung, Soan Thi Minh Duong, Abdesselam Bouzerdoum |
Appl. Intell. | 2 |
| 2022 | Class Similarity Weighted Knowledge Distillation for Continual Semantic SegmentationabstractDeep learning models are known to suffer from the problem of catastrophic forgetting when they incrementally learn new classes. Continual learning for semantic segmentation (CSS) is an emerging field in computer vision. We identify a problem in CSS: A model tends to be confused between old and new classes that are visually similar, which makes it forget the old ones. To address this gap, we propose REMINDER - a new CSS framework and a novel class similarity knowledge distillation (CSW-KD) method. Our CSW-KD method distills the knowledge of a previous model on old classes that are similar to the new one. This provides two main benefits: (i) selectively revising old classes that are more likely to be forgotten, and (ii) better learning new classes by relating them with the previously seen classes. Extensive experiments on Pascal-Voc 2012 and ADE20k datasets show that our approach outperforms state-of-the-art methods on standard CSS settings by up to 7.07% and 8.49%, respectively. Minh-Hieu Phan, The-Anh Ta, Son Lam Phung, Long Tran-Thanh, Abdesselam Bouzerdoum |
CVPR | 3 |
| 2022 | DirecFormer: A Directed Attention in Transformer Approach to Robust Action RecognitionabstractHuman action recognition has recently become one of the popular research topics in the computer vision community. Various 3D-CNN based methods have been presented to tackle both the spatial and temporal dimensions in the task of video action recognition with competitive results. However, these methods have suffered some fundamental limitations such as lack of robustness and generalization, e.g., how does the temporal ordering of video frames affect the recognition results? This work presents a novel end-to-end Transformer-based Directed Attention (Direc-Former) framework11The implementation of DirecFormer is available at https://github.com/uark-cviu/DirecFormer for robust action recognition. The method takes a simple but novel perspective of Transformer-based approach to understand the right order of sequence actions. Therefore, the contributions of this work are three-fold. Firstly, we introduce the problem of ordered temporal learning issues to the action recognition problem. Secondly, a new Directed Attention mechanism is introduced to understand and provide attentions to human actions in the right order. Thirdly, we introduce the conditional dependency in action sequence modeling that includes orders and classes. The proposed approach consistently achieves the state-of-the-art (SOTA) results compared with the recent action recognition methods [4, 18, 72, 74]. on three standard large-scale benchmarks, i.e. Jester, Kinetics-400 and Something-Something-V2. Thanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo, Son Lam Phung, Xin Li 0005, Khoa Luu |
CVPR | 5 |
| 2022 | Efficient hyperspectral image segmentation for biosecurity scanning using knowledge distillation from multi-head teacher
Minh-Hieu Phan, Son Lam Phung, Khoa Luu, Abdesselam Bouzerdoum |
Neurocomputing | 2 |
| 2022 | Bayesian Gabor Network With Uncertainty Estimation for Pedestrian Lane Detection in Assistive NavigationabstractAutomatic pedestrian lane detection is a challenging problem that is of great interest in assistive navigation and autonomous driving. Such a detection system must cope well with variations in lane surfaces and illumination conditions so that a vision-impaired user can navigate safely in unknown environments. This paper proposes a new lightweight Bayesian Gabor Network (BGN) for camera-based detection of pedestrian lanes in unstructured scenes. In our approach, each Gabor parameter is represented as a learnable Gaussian distribution using variational Bayesian inference. For the safety of vision-impaired users, in addition to an output segmentation map, the network provides two full-resolution maps of aleatoric uncertainty and epistemic uncertainty as well-calibrated confidence measures. Our Gabor-based method has fewer weights than the standard CNNs, therefore it is less prone to overfitting and requires fewer operations to compute. Compared to the state-of-the-art semantic segmentation methods, the BGN maintains a competitive segmentation performance while achieving a significantly compact model size (from$1.8\times $to$237.6\times $reduction), a fast prediction time (from$1.2\times $to$67.5\times $faster), and a well-calibrated uncertainty measure. We also introduce a new lane dataset of 10,000 images for objective evaluation in pedestrian lane detection research. Hoang Thanh Le 0001, Son Lam Phung, Abdesselam Bouzerdoum |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | BiMaL: Bijective Maximum Likelihood Approach to Domain Adaptation in Semantic Scene SegmentationabstractSemantic segmentation aims to predict pixel-level labels. It has become a popular task in various computer vision applications. While fully supervised segmentation methods have achieved high accuracy on large-scale vision datasets, they are unable to generalize on a new test environment or a new domain well. In this work, we first introduce a new Unaligned Domain Score to measure the efficiency of a learned model on a new target domain in unsupervised manner. Then, we present the new Bijective Maximum Likelihood1(BiMaL) loss that is a generalized form of the Adversarial Entropy Minimization without any assumption about pixel independence. We have evaluated the proposed BiMaL on two domains. The proposed BiMaL approach consistently outperforms the SOTA methods on empirical experiments on "SYNTHIA to Cityscapes", "GTA5 to Cityscapes", and "SYNTHIA to Vistas". Thanh-Dat Truong, Chi Nhan Duong, T. Hoang Ngan Le, Son Lam Phung, Chase Rainwater, Khoa Luu |
ICCV | 4 |
| 2021 | A Trajectory-Based Method for Dynamic Scene RecognitionabstractExisting methods for dynamic scene recognition mostly use global features extracted from the entire video frame or a video segment. In this paper, a trajectory-based dynamic scene recognition method is proposed. A trajectory is formed by a pixel moving across consecutive frames of a video segment. The local regions surrounding the trajectory provide useful appearance and motion information about a portion of the video segment. The proposed method works at several stages. First, dense and evenly distributed trajectories are extracted from a video segment. Then, the fully-connected-layer features are extracted from each trajectory using a pre-trained Convolutional Neural Networks (CNNs) model, forming a feature sequence. Next, these feature sequences are fed into a Long-Short-Term-Memory (LSTM) network to learn their temporal behavior. Finally, by aggregating the information of the trajectories, a global representation of the video segment can be obtained for classification purposes. The LSTM is trained using synthetic trajectory feature sequences instead of real ones. The synthetic feature sequences are generated with a series of generative adversarial networks (GANs). In addition to classification, category-specific discriminative trajectories are located in a video segment, which help reveal what portions of a video segment are more important than others. This is achieved by formulating an optimization problem to learn discriminative part detectors for all categories simultaneously. Experimental results on two benchmark dynamic scene datasets show that the proposed method is very competitive with six other methods. Xiaoming Peng, Abdesselam Bouzerdoum, Son Lam Phung |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2021 | A part-based spatial and temporal aggregation method for dynamic scene recognition
Xiaoming Peng, Abdesselam Bouzerdoum, Son Lam Phung |
Neural Comput. Appl. | 3 |
| 2021 | Variational Bayesian Compressive Multipolarization Indoor Radar ImagingabstractThis article introduces a probabilistic Bayesian model for addressing the problem of compressive multipolarization through-wall radar imaging (TWRI). The proposed approach formulates the task of wall-clutter mitigation and multipolarization image reconstruction as a Bayesian inference problem for a joint distribution between observed radar measurements and latent wall-clutter matrix and indoor target images. The joint probability distribution incorporates three prior beliefs: low-dimensional structure of the wall reflections, group sparsity structure of the target images, and joint sparsity among the polarization images. These signal attributes are modeled through hierarchical priors, whose parameters and hyperparameters are treated with a full Bayesian formulation. Furthermore, this article presents a variational Bayesian inference algorithm that estimates wall-clutter and multipolarization images as posterior distributions and optimizes the model parameters and hyperparameters simultaneously. Experimental results on simulated and real radar data show that the proposed model is very effective at removing wall clutter and enhancing target localization even when the radar measurements are significantly reduced. Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | A Variational Bayesian Approach for Multichannel Through-Wall Radar Imaging with Low-Rank and Sparse PriorsabstractThis paper considers the problem of multichannel through-wall radar (TWR) imaging from a probabilistic Bayesian perspective. Given the observed radar signals, a joint distribution of the observed data and latent variables is formulated by incorporating two important beliefs: low-dimensional structure of wall reflections and joint sparsity among channel images. These priors are modeled through probabilistic distributions whose hyperparameters are treated with a full Bayesian formulation. Furthermore, the paper presents a variational Bayesian inference algorithm that captures wall clutter and provides channel images as full posterior distributions. Experimental results on real data show that the proposed model is very effective at removing wall clutter and enhancing target localization. Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung |
ICASSP | 3 |
| 2020 | Infer The Input To The Generator Of Auxiliary Classifier Generative Adversarial NetworksabstractGenerative Adversarial Networks (GANs) are deep-learning-based generative models. This paper presents three methods to infer the input to the generator of auxiliary classifier generative adversarial networks (ACGANs), which are a type of conditional GANs. The first two methods, named i-ACGAN- r and i-ACGAN-d, are “inverting” methods, which obtain an inverse mapping from an image to the class label and the latent sample. By contrast, the third method, referred to as i-ACGAN-e, directly infers both the class label and the latent sample by introducing an encoder into an ACGAN. The three methods were evaluated on two natural scene datasets, using two performance measures: the class recovery accuracy and the image reconstruction error. Experimental results show that i-ACGAN-e outperforms the other two methods in terms of the class recovery accuracy. However, the images generated by the other two methods have smaller image reconstruction errors. The source code is publicly available from https://github.com/XMPeng/Infer-Input-ACGAN. Xiaoming Peng, Abdesselam Bouzerdoum, Son Lam Phung |
ICIP | 3 |
| 2020 | Real-time Pedestrian Lane Detection for Assistive Navigation using Neural Architecture SearchabstractPedestrian lane detection is a core component in many assistive and autonomous navigation systems. These systems are usually deployed in environments that require realtime processing. Many state-of-the-art deep neural networks only focus on detection accuracy but not inference speed. Without further modifications, they are not suitable for real-time applications. Furthermore, the task of designing a high-performing deep neural network is time-consuming and requires experience. To tackle these issues, we propose a neural architecture search algorithm that can find the best deep network for pedestrian lane detection automatically. The proposed method searches in a network-level space using the gradient descent algorithm. Evaluated on a dataset of 5,000 images, the deep network found by the proposed algorithm achieves comparable segmentation accuracy, while being significantly faster than other state-of-the-art methods. The proposed method has been successfully implemented as a real-time pedestrian lane detection tool. Sui Paul Ang, Son Lam Phung, Abdesselam Bouzerdoum, Thi Nhat Anh Nguyen, Soan Thi Minh Duong, Mark M. Schira |
ICPR | 2 |
| 2020 | Ordinal Depth Classification Using Region-based Self-attentionabstractDepth perception is essential for scene understanding, autonomous navigation and augmented reality. Depth estimation from a single 2D image is challenging due to the lack of reliable cues, e.g. stereo correspondences and motions. Modern approaches exploit multi-scale feature extraction to provide more powerful representations for deep networks. However, these studies only use simple addition or concatenation to combine the extracted multi-scale features. This paper proposes a novel region-based self-attention (rSA) unit for effective feature fusions. The rSA recalibrates the multi-scale responses by explicitly modelling the dependency between channels in separate image regions. We discretize continuous depths to formulate an ordinal depth classification problem in which the relative order between categories is preserved. The experiments are performed on a dataset of 4410 RGB-D images, captured in outdoor environments at the University of Wollongong's campus. The proposed module improves the models on small-sized datasets by 22% to 40%. Minh-Hieu Phan, Son Lam Phung, Abdesselam Bouzerdoum |
ICPR | 2 |
| 2020 | Compressive Radar Imaging of Stationary Indoor Targets With Low-Rank Plus Jointly Sparse and Total Variation RegularizationsabstractThis paper addresses the problem of wall clutter mitigation and image reconstruction for through-wall radar imaging (TWRI) of stationary targets by seeking a model that incorporates low-rank (LR), joint sparsity (JS), and total variation (TV) regularizers. The motivation of the proposed model is that LR regularizer captures the low-dimensional structure of wall clutter; JS guarantees a small fraction of target occupancy and the similarity of sparsity profile among channel images; TV regularizer promotes the spatial continuity of target regions and mitigates background noise. The task of wall clutter mitigation and target image reconstruction is formulated as an optimization problem comprising LR, JS, and TV regularization terms. To handle this problem efficiently, an iterative algorithm based on the forward-backward proximal gradient splitting technique is introduced, which captures wall clutter and yields target images simultaneously. Extensive experiments are conducted on real radar data under compressive sensing scenarios. The results show that the proposed model enhances target localization and clutter mitigation even when radar measurements are significantly reduced. Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung |
IEEE Trans. Image Process. | 3 |
| 2020 | Hybrid Deep Learning-Gaussian Process Network for Pedestrian Lane Detection in Unstructured ScenesabstractPedestrian lane detection is an important task in many assistive and autonomous navigation systems. This article presents a new approach for pedestrian lane detection in unstructured environments, where the pedestrian lanes can have arbitrary surfaces with no painted markers. In this approach, a hybrid deep learning-Gaussian process (DL-GP) network is proposed to segment a scene image into lane and background regions. The network combines a compact convolutional encoder-decoder net and a powerful nonparametric hierarchical GP classifier. The resulting network with a smaller number of trainable parameters helps mitigate the overfitting problem while maintaining the modeling power. In addition to the segmentation output for each test image, the network also generates a map of uncertainty-a measure that is negatively correlated with the confidence level with which we can trust the segmentation. This measure is important for pedestrian lane-detection applications, since its prediction affects the safety of its users. We also introduce a new data set of 5000 images for training and evaluating the pedestrian lane-detection algorithms. This data set is expected to facilitate research in pedestrian lane detection, especially the application of DL in this area. Evaluated on this data set, the proposed network shows significant performance improvements compared with several existing methods. Thi Nhat Anh Nguyen, Son Lam Phung, Abdesselam Bouzerdoum |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Radar Stationary and Moving Indoor Target Localization with Low-rank and Sparse RegularizationsabstractThis paper proposes a low-rank and sparse regularized optimization model to address the problem of wall clutter mitigation, stationary, and moving target indications using through-wall radar. The task of wall clutter suppression and target image reconstruction is formulated as a nuclear and ℓ1penalized least squares optimization problem in which the nuclear-norm term enforces for a low-rank wall clutter matrix and the ℓ1-norm term promotes the sparsity of the target images. An iterative algorithm based on the proximal gradient technique is introduced to solve the optimization problem. The solution comprises the wall clutter and images of stationary and moving targets. Experiments are conducted on real radar data under compressive sensing scenarios. The results show that the proposed model is very effective at removing unwanted wall clutter, reconstructing stationary targets, and capturing moving targets. Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung |
ICASSP | 3 |
| 2019 | A Super Descriptor Tensor Decomposition for Dynamic Scene RecognitionabstractThis paper presents a new approach for dynamic scene recognition based on a super descriptor tensor decomposition. Recently, local feature extraction based on dense trajectories has been used for modeling motion. However, dense trajectories usually include a large number of unnecessary trajectories, which increase noise, add complexity, and limit the recognition accuracy. Another problem is that the traditional bag-of-words techniques encode and concatenate the local features extracted from multiple descriptors to form a single large vector for classification. This concatenation not only destroys the spatio-temporal structure among the features but also yields high dimensionality. To address these problems, first, we propose to refine the dense trajectories by selecting only salient trajectories in a region of interest containing motion. Visual descriptors consisting of oriented gradient and motion boundary histograms are then computed along the refined dense trajectories. In case of camera motion, a short-window video stabilization is integrated to compensate for global motion. Second, the extracted features from multiple descriptors are encoded using a super descriptor tensor model. To this end, the TUCKER-3 tensor decomposition is employed to obtain a compact set of salient features, followed by feature selection via Fisher ranking. Experiments are conducted using two benchmark dynamic scene recognition datasets: Maryland “in-the-wild” and YUPPEN dynamic scenes. Experimental results show that the proposed approach outperforms several existing methods in terms of recognition accuracy and achieves a performance comparable with the state-of-the-art deep learning methods. The proposed approach achieves classification rates of 89.2% for Maryland and 98.1% for YUPPEN datasets. Muhammad Rizwan Khokher, Abdesselam Bouzerdoum, Son Lam Phung |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Pooling-Based Feature Extraction and Coarse-to-fine Patch Matching for Optical Flow Estimation
Xiaolin Tang, Son Lam Phung, Abdesselam Bouzerdoum, Van Ha Tang |
ACCV (4) | 2 |
| 2018 | Anatomy-Guided Inverse-Gradient Susceptibility Artifact Correction Method for High-Resolution FMRIabstractFunctional Magnetic Resonance Imaging (fMRI) is a widely used and non-invasive technique for recording changes in brain activity. However, susceptibility artifacts are ubiquitous distortions in fMRI, especially strong in high-resolution images, causing the misrepresentation of brain function and structure in the affected regions. Here, we present a novel method for correcting these distortions in high-resolution fMRI images based on the hyper-elastic susceptibility artifact correction (HySCO) method. The novelty of the proposed method is the utilization of the easily-acquired T1-weighted (T1w) anatomy image as a ground-truth measurement to regularize deformations, thereby obtaining meaningful corrections. The performance of the new method is compared to that of HySCO. Results from high-resolution (1mm) EPI data are presented demonstrating the robustness of the new method for image correction and its suitability for subsequent fMRI analysis. Soan Thi Minh Duong, Mark M. Schira, Son Lam Phung, Abdesselam Bouzerdoum, H. G. B. Taylor |
ICASSP | 3 |
| 2018 | Human Motion Classification with Micro-Doppler Radar and Bayesian-Optimized Convolutional Neural NetworksabstractIn recent years, Doppler radar has emerged as an alternative sensing modality for human gait classification since it measures not only the target speed, but also the local dynamics of the moving body parts, thereby creating a unique spectral signature. This paper presents a learning-based method for classifying human motions from micro-Doppler signals. Inspired by the applications of deep learning, the proposed method extracts features from the time-frequency representation of the radar signal using a cascaded of convolutional network layers. To design a optimal network architecture, the Bayesian optimization with Gaussian process priors is employed. Experimental results on real data are presented, which show a significant improvement compared to three existing approaches. Hoang Thanh Le 0001, Son Lam Phung, Abdesselam Bouzerdoum, Fok Hing Chi Tivive |
ICASSP | 2 |
| 2018 | Scalable Hierarchical Mixture of Gaussian Processes for Pattern ClassificationabstractThis paper introduces a novel Gaussian process (GP) classification method that combines advantages of global and local GP approximators through a two-layer hierarchical model. The upper layer consists of a global sparse GP to coarsely model the entire dataset. The lower layer is a mixture of GP experts which uses local information to learn a fine-grained model. A variational inference algorithm is developed for simultaneous learning of the global GP, the experts and the gating network. Stochastic optimization can be employed for large-scale problems. Experiments on benchmark binary classification datasets demonstrate the advantages of the method in terms of scalability and classification accuracy. Thi Nhat Anh Nguyen, Abdesselam Bouserdoum, Son Lam Phung |
ICASSP | 3 |
| 2018 | A Matrix Completion Approach for Wall-Clutter Mitigation in Compressive Radar Imaging of Indoor TargetsabstractThis paper presents a low-rank matrix completion approach to tackle the problem of wall clutter mitigation for through-wall radar imaging in the compressive sensing context. In particular, the task of wall clutter removal is reformulated as a matrix completion problem in which a low-rank matrix containing wall clutter is reconstructed from compressive measurements. The proposed model regularizes the low-rank prior of the wall-clutter matrix via the nuclear norm, casting the wall-clutter mitigation task as a nuclear-norm penalized least squares problem. To solve this optimization problem, an iterative algorithm based on the proximal gradient technique is introduced. Experiments on simulated full-wave electromagnetic data are conducted under compressive sensing scenarios. The results show that the proposed matrix completion approach is very effective at suppressing unwanted wall clutter and enhancing the targets. Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung |
ICASSP | 3 |
| 2018 | Human Gait Recognition with Micro-Doppler Radar and Deep AutoencoderabstractThe micro-Doppler signals from moving objects contain useful information about their motions. This paper introduces a novel approach for human gait recognition based on backscattered signals from a micro-Doppler radar. Three different signal techniques are utilized for the extraction of micro-Doppler features via time-frequency and time-scale representations. To classify the human motions into various types, this paper presents a deep autoencoder with the use of local patches extracted along the spectrogram and scalogram. The network configuration and the learning parameters of the deep autoencoder, which are considered as hyperparameters, are optimized by a Bayesian optimization algorithm. Experimental results produced by the proposed technique on real radar data show a significant improvement compared to several existing approaches. Hoang Thanh Le 0001, Son Lam Phung, Abdesselam Bouzerdoum |
ICPR | 2 |
| 2018 | Stochastic variational hierarchical mixture of sparse Gaussian processes for regression
Thi Nhat Anh Nguyen, Abdesselam Bouzerdoum, Son Lam Phung |
Mach. Learn. | 3 |
| 2018 | Multipolarization Through-Wall Radar Imaging Using Low-Rank and Jointly-Sparse RepresentationsabstractCompressed sensing techniques have been applied to through-the-wall radar imaging (TWRI) and multipolarization TWRI for fast data acquisition and enhanced target localization. The studies so far in this area have either assumed effective wall clutter removal prior to image formation or performed signal estimation, wall clutter mitigation, and image formation independently. This paper proposes a low-rank and sparse imaging model for jointly addressing the problem of wall clutter mitigation and image formation in multichannel TWRI. The proposed model exploits two important structures of through-wall radar signals: low-rank structure of the wall reflections and jointly-sparse structure among the different polarization images. The task of removing wall clutter and reconstructing multichannel images of the same scene behind-the-wall is formulated as a regularized least squares problem, where low-rank regularization is enforced for the wall components, and joint-sparsity penalty is imposed on channel images. To solve the optimization problem, an iterative algorithm based on the proximal gradient technique is introduced, which simultaneously estimates the wall interferences and yields multichannel images of the indoor targets. Experiments on real and simulated radar data are conducted under full measurements and compressive sensing scenarios. The results show that the proposed model is very effective at removing unwanted wall clutter and enhancing the stationary targets, even under considerable reduction in measurements. Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung |
IEEE Trans. Image Process. | 3 |
| 2017 | Human interaction recognition using low-rank matrix approximation and super descriptor tensor decompositionabstractAudio-visual recognition systems rely on efficient feature extraction. Many spatio-temporal interest point detectors for visual feature extraction are either too sparse, leading to loss of information, or too dense resulting in noisy and redundant information. Furthermore, interest point detectors designed for a controlled environment can be affected by camera motion. In this paper, a salient spatio-temporal interest point detector is proposed based on a low-rank and group-sparse matrix approximation. The detector handles the camera motion through a short-window video stabilization. The multimodal audio-visual features from multiple descriptors are represented by a super descriptor, from which a compact set of features is extracted through a tensor decomposition and feature selection. This tensor decomposition retains the spatiotemporal structure among features obtained from multiple descriptors. Experimental validation is conducted using two benchmark human interaction recognition datasets: TVHID and Parliament. Experimental results are presented which show that the proposed approach outperforms many state-ofthe-art methods, achieving classification rates of 74.7% and 88.5% on the TVHID and Parliament datasets, respectively. Muhammad Rizwan Khokher, Abdesselam Bouzerdoum, Son Lam Phung |
ICASSP | 3 |
| 2017 | Enhanced pixel-wise voting for image vanishing point detection in road scenesabstractVanishing point estimation is a crucial task in vision-based road detection. This paper presents a new texture-based voting scheme, which enhances both accuracy and speed of vanishing point estimation. In the proposed method, color tensors analysis is adopted to calculate local orientations and color edges. The search space is reduced by optimizing the set of vanishing point candidates and voters. A new strategy based on Bayesian classifier is proposed to select a suitable voting function. The proposed method is evaluated on a benchmark dataset of 4000 images of pedestrian lanes with annotated vanishing points. The experimental results show that it offers an improved accuracy and significantly faster processing time compared with other state-of-the-art methods. L. Nguyen, Son Lam Phung, Abdesselam Bouzerdoum |
ICASSP | 2 |
| 2017 | An efficient local method for stereo matching using daisy featuresabstractIn this paper, a local method is proposed to estimate the visibility and disparity of pixels from a stereo pair using the DAISY feature. The problem is formulated as a joint optimization over disparity and visibility of individual pixels. The constraints on the range of disparities and the binary visibility variables are enforced by incorporating penalty terms into the cost function. Finally, the unconstrained optimization problem is solved using a Newton scheme with appropriate approximations to the Hessian matrices and gradients. The computation time of the proposed optimization method is around one minute to run for 768 × 512 stereo pairs using the DAISY feature descriptor in a C++ implementation. Xiaoming Peng, Abdesselam Bouzerdoum, Son Lam Phung |
ICIP | 3 |
| 2017 | Breast cancer diagnosis in DCE-MRI using mixture ensemble of convolutional neural networks
Reza Rasti, Mohammad Teshnehlab, Son Lam Phung |
Pattern Recognit. | 3 |
| 2016 | Variational inference for infinite mixtures of sparse Gaussian processes through KL-correctionabstractWe propose a new approximation method for Gaussian process (GP) regression based on the mixture of experts structure and variational inference. Our model is essentially an infinite mixture model in which each component is composed of a Gaussian distribution over the input space, and a Gaussian process expert over the output space. Each expert is a sparse GP model augmented with its own set of inducing points. Variational inference is made feasible by assuming that the training outputs are independent given the inducing points. In previous works on variational mixture of GP experts, the inducing points are selected through a greedy selection algorithm, which is computationally expensive. In our method, both the inducing points and hyperparameters of the experts are learned through maximizing an improved lower bound of the marginal likelihood. Experiments on benchmark datasets show the advantages of the proposed method. Thi Nhat Anh Nguyen, Abdesselam Bouzerdoum, Son Lam Phung |
ICASSP | 3 |
| 2016 | Radar imaging of stationary indoor targets using joint low-rank and sparsity constraintsabstractThis paper introduces a joint low-rank and sparsity-based model to address the problem of wall-clutter mitigation in compressed through-the-wall radar imaging. The proposed model is motivated by two observations that wall reflections reside in a low-rank subspace, and target signals tend to be sparse. In the proposed approach, the task of segregating target returns from wall reflections is formulated as a joint low-rank and sparsity constrained optimization problem. Here, the low rank constraint is imposed on the wall component and the sparsity constraint is used to model the target component. An iterative soft thresholding algorithm is developed to estimate a low-rank matrix of wall clutter and a sparse matrix of target reflections from a reduced measurement set. Once the wall and target components are estimated, the target signals are used for scene reconstruction. Experimental evaluation was conducted using real radar data. The results show that the proposed model is very effective at removing wall clutter and reconstructing the image of behind-the-wall targets from reduced measurements. Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung, Fok Hing Chi Tivive |
ICASSP | 3 |
| 2016 | Pedestrian lane detection in unstructured scenes for assistive navigationabstractAutomatic detection of the pedestrian lane in a scene is an important task in assistive and autonomous navigation. This paper presents a vision-based algorithm for pedestrian lane detection in unstructured scenes, where lanes vary significantly in color, texture, and shape and are not indicated by any painted markers. In the proposed method, a lane appearance model is constructed adaptively from a sample image region, which is identified automatically from the image vanishing point. This paper also introduces a fast and robust vanishing point estimation method based on the color tensor and dominant orientations of color edge pixels. The proposed pedestrian lane detection method is evaluated on a new benchmark dataset that contains images from various indoor and outdoor scenes with different types of unmarked lanes. Experimental results are presented which demonstrate its efficiency and robustness in comparison with several existing methods. Son Lam Phung, Manh Cuong Le, Abdesselam Bouzerdoum |
Comput. Vis. Image Underst. | 1 |
| 2015 | Multi-view indoor scene reconstruction from compressed through-wall radar measurements using a joint bayesian sparse representationabstractThis paper addresses the problem of scene reconstruction, incorporating wall-clutter mitigation, for compressed multi-view through-the-wall radar imaging. We consider the problem where the scene is sensed using different reduced sets of frequencies at different antennas. A joint Bayesian sparse recovery framework is first employed to estimate the antenna signal coefficients simultaneously, by exploiting the sparsity and correlations between antenna signals. Following joint signal coefficient estimation, a subspace projection technique is applied to segregate the target coefficients from the wall contributions. Furthermore, a multitask linear model is developed to relate the target coefficients to the scene, and a composite scene image is reconstructed by a joint Bayesian sparse framework, taking into account the inter-view dependencies. Experimental results show that the proposed approach improves reconstruction accuracy and produces a composite scene image in which the targets are enhanced and the background clutter is attenuated. Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung, Fok Hing Chi Tivive |
ICASSP | 3 |
| 2015 | Depth image super-resolution using internal and external informationabstractThe fast development of 3-D imaging techniques has increased demands for high-resolution depth images. Conventional depth super-resolution methods reconstruct the high-resolution image by accessing high frequency information, either internally from a high-resolution intensity image or externally from a high-resolution image database. In this paper, a new depth super-resolution method based on joint regularization is proposed, which exploits both internal and external high frequency information. Specifically, a joint regularization problem with different constraints is formulated, which allows us to solve for the high-resolution image and a sparse code simultaneously. These constraints are constructed by utilizing information from both internal and external high-frequency sources. Experimental evaluation suggests that the proposed method provides improved results over existing approaches, in terms of both visual appearance and objective image quality. Haoheng Zheng, Abdesselam Bouzerdoum, Son Lam Phung |
ICASSP | 3 |
| 2015 | Invariant image recognition under projective deformations: An image normalization approachabstractRobustness in image recognition refers to the ability to perceive an image pattern regardless of factors including camera views and locations. This paper proposes an image normalization algorithm that allows an image with arbitrary projective distortions to be recognized efficiently. The normalization algorithm calculates the required projective transformation matrix using image moments. For an input image, a set of 8 output images that are independent of projective deformations are generated. The proposed algorithm is evaluated on three benchmark data sets. The experimental results show that the proposed normalization is significantly more accurate than the existing rank minimization and affine normalization methods. Son Lam Phung, Abdesselam Bouzerdoum, Amine Bermak |
VCIP | 2 |
| 2014 | Lane Detection in Unstructured Environments for Autonomous Navigation Systems
Manh Cuong Le, Son Lam Phung, Abdesselam Bouzerdoum |
ACCV (1) | 2 |
| 2014 | Enhanced wall clutter mitigation for compressed through-the-wall radar imaging using joint Bayesian sparse signal recoveryabstractThis paper addresses the problem of wall clutter mitigation in compressed sensing through-the-wall radar imaging, where a different set of frequencies is sensed at different antenna locations. A joint Bayesian sparse approximation framework is first employed to reconstruct all the signals simultaneously by exploiting signal sparsity and correlations between antenna signals. This is in contrast to previous approaches where the signal at each antenna location is reconstructed independently. Furthermore, to promote sparsity and improve reconstruction accuracy, a sparsifying wavelet dictionary is employed in the sparse signal recovery. Following signal reconstruction, a subspace projection technique is applied to remove wall clutter, prior to image formation. Experimental results on real data show that the proposed approach produces significantly higher reconstruction accuracy and requires far fewer measurements for forming high-quality images, compared to the single-signal compressed sensing model, where each antenna signal is reconstructed independently. Van Ha Tang, Abdesselam Bouzerdoum, Son Lam Phung, Fok Hing Chi Tivive |
ICASSP | 3 |
| 2014 | Affine-invariant scene categorizationabstractThis paper presents a scene categorization method that is invariant to affine transformations. We propose a new moment-based normalization algorithm to generate an output image that is independent of the position, rotation, shear, and scale of the input image. In the proposed approach, an affine transform matrix is determined subject to the normalized image satisfying a set of moment constraints. After image normalization, a dense set of local features is extracted using scattering transform, and the global features are then formed via a sparse coding method. We evaluate the proposed method and other state-of-the-art algorithms on a benchmark dataset. The experimental results show that for images distorted with affine transformations, the proposed normalization increases the classification rate by about 28%, compared with the scene categorization approach that uses no normalization. Son Lam Phung, Abdesselam Bouzerdoum |
ICIP | 2 |
| 2014 | Object segmentation and classification using 3-D range camera
Son Lam Phung, Abdesselam Bouzerdoum |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Depth image super-resolution using multi-dictionary sparse representationabstractIn this paper, we propose a new depth super-resolution technique based on multiple dictionary learning. A novel dictionary selection method using basis pursuit is proposed to generate multiple dictionaries adaptively. A sparse representation of each low-resolution input patch is derived based on the learned dictionaries, and then used to reconstruct the corresponding high-resolution patch. Experimental results are presented which show that the proposed multi-dictionary scheme outperforms existing depth super-resolution methods. Abdesselam Bouzerdoum, Son Lam Phung |
ICIP | 3 |
| 2013 | Two-Stage Fuzzy Fusion With Applications to Through-the-Wall Radar ImagingabstractA two-stage fuzzy image fusion approach, which combines multiple radar images of the same scene, is proposed to produce a more informative image. In this approach, two different image fusion methods are first applied. Then, a fuzzy logic fusion method is applied to the outputs of the first fusion stage. The performance of the proposed approach is evaluated on through-the-wall radar images obtained using different polarizations. Experimental results show that the proposed approach enhances image quality by producing outputs with high target intensity values and low clutter. Cher Hau Seng, Abdesselam Bouzerdoum, Moeness G. Amin, Son Lam Phung |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2013 | Sparse Representation of GPR Traces With Application to Signal ClassificationabstractSparse representation (SR) models a signal with a small number of elementary waves using an overcomplete dictionary. It has been employed for a wide range of signal and image processing applications, including denoising, deblurring, and compression. In this paper, we present an adaptive SR method for modeling and classifying ground penetrating radar (GPR) signals. The proposed method decomposes each GPR trace into elementary waves using an adaptive Gabor dictionary. The sparse decomposition is used to extract salient features for SR and classification of GPR signals. Experimental results on real-world data show that the proposed sparse decomposition achieves efficient signal representation and yields discriminative features for pattern classification. Wenbin Shao, Abdesselam Bouzerdoum, Son Lam Phung |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2013 | Probabilistic Fuzzy Image Fusion Approach for Radar Through Wall SensingabstractThis paper addresses the problem of combining multiple radar images of the same scene to produce a more informative composite image. The proposed approach for probabilistic fuzzy logic-based image fusion automatically forms fuzzy membership functions using the Gaussian-Rayleigh mixture distribution. It fuses the input pixel values directly without requiring fuzzification and defuzzification, thereby removing the subjective nature of the existing fuzzy logic methods. In this paper, the proposed approach is applied to through-the-wall radar imaging in urban sensing and evaluated on real multi-view and polarimetric data. Experimental results show that the proposed approach yields improved image contrast and enhances target detection. Cher Hau Seng, Abdesselam Bouzerdoum, Moeness G. Amin, Son Lam Phung |
IEEE Trans. Image Process. | 4 |
| 2012 | Adaptive Autoregressive Logarithmic Search for 3D Human TrackingabstractHuman tracking is an important vision task in video surveillance and perceptual human-computer interfaces. This paper presents a novel algorithm for region-based human tracking using color and depth features. We propose an adaptive autoregressive logarithmic search (ARLS) to estimate the target position, and use depth information to further reduce the false alarm rate. The new ARLS algorithm is evaluated on a color and depth (RGBD) video dataset acquired with the Kinect sensor. The dataset contains various real-world scenarios with illumination and speed variations, and partial occlusion. The experimental results show that the ARLS algorithm is able to handle difficult tracking scenarios, it achieves a tracking accuracy of 91.26% on the test dataset. The proposed algorithm is compared with two tracking algorithms, namely the particle filtering and a modified logarithmic search algorithm. Peiyao Li, Abdesselam Bouzerdoum, Son Lam Phung |
AVSS | 3 |
| 2012 | Compressed sensing-based frequency selection for classification of ground penetrating radar signalsabstractIn this paper we present an automatic classification system for ground penetrating radar (GPR) signals. The system extracts the magnitude spectra at resonant frequencies and classifies them using support vector machines. To locate the resonant frequencies, we propose an approach based on compressed sensing and orthogonal matching pursuit. The performance of the system is evaluated by classifying GPR traces from different ballast fouling conditions. The experimental results show that the proposed approach, compared to the approach of using frequencies at local maxima, represents the GPR signal more efficiently using a small number of coefficients, and obtains higher classification accuracy. Wenbin Shao, Abdesselam Bouzerdoum, Son Lam Phung |
ICASSP | 3 |
| 2012 | Scene Segmentation and Pedestrian Classification from 3-D Range and Intensity ImagesabstractThis paper proposes a new approach to classify obstacles using a time-of-flight camera, for applications in assistive navigation of the visually impaired. Combining range and intensity images enables fast and accurate object segmentation, and provides useful navigation cues such as distances to the nearby obstacles and obstacle types. In the proposed approach, a 3-D range image is first segmented using histogram thresholding and mean-shift grouping. Then Fourier and GIST descriptors are applied on each segmented object to extract shape and texture features. Finally, support vector machines are used to recognize the obstacles. This paper focuses on classifying pedestrian and non-pedestrian obstacles. Evaluated on an image data set acquired using a time-of-flight camera, the proposed approach achieves a classification rate of 99.5%. Son Lam Phung, Abdesselam Bouzerdoum |
ICME | 2 |
| 2012 | Pedestrian lane detection for assistive navigation of blind people
Manh Cuong Le, Son Lam Phung, Abdesselam Bouzerdoum |
ICPR | 2 |
| 2012 | Automatic classification of human motions using Doppler radarabstractThis paper presents a new approach to classify human motions using a Doppler radar for applications in security and surveillance. Traditionally, the Doppler radar is an effective tool for detecting the position and velocity of a moving target, even in adverse weather conditions and from a long range. In this paper, we are interested in using the Doppler radar to recognize the micro-motions exhibited by people. In the proposed approach, a frequency modulated continuous wave radar is applied to scan the target, and the short-time Fourier transform is used to convert the radar signal into spectrogram. Then, the new two-directional, two-dimensional principal component analysis and linear discriminant analysis are performed to obtain the feature vectors. This approach is more computationally efficient than the traditional principal component analysis. Finally, support vector machines are applied to classify feature vectors into different human motions. Evaluated on a radar data set with three types of motions, the proposed approach has a classification rate of 91.9%. Jingli Li, Son Lam Phung, Fok Hing Chi Tivive, Abdesselam Bouzerdoum |
IJCNN | 2 |
| 2011 | Adaptive regularization for multiple image restoration using an extended Total Variations approachabstractIn this paper a Variational Inequality method for multiple in- put, multiple output image restoration is presented using an extended Total Variations (TV) regularizer. This approach calculates an adaptive regularization parameter for each image based on their respective degradations. The proposed ex- tended Total Variations regularizer combines both intra-image and inter-image pixel information for improved restoration performance. Hyperparameters for controlling this new TV measure are calculated using a Bayesian joint maximum a posteriori approach. Matthew Andrew Kitchener, Abdesselam Bouzerdoum, Son Lam Phung |
ICIP | 3 |
| 2011 | Optical flow estimation using sparse gradient representationabstractThis paper introduces a sparsity based optical flow estimation method in digital video sequences. The method stems from the key observation that the gradient field of optical flow, in digital video sequences, is usually structured and sparse in spatial domain, provided there is a small number of multiple motions in the scene. The gradient field of motion vectors is formed by the pixels forming the edges of moving objects. We utilize this fact and formulate the optical flow estimation problem in sparse representation framework. We then use a minimization algorithm over ℓ1norm of the gradient flow field to find the solution to this problem. The proposed algorithm has been evaluated on Middlebury's benchmark video sequence database. Muhammad Wasim Nawaz, Abdesselam Bouzerdoum, Son Lam Phung |
ICIP | 3 |
| 2011 | Automatic Classification of Ground-Penetrating-Radar Signals for Railway-Ballast AssessmentabstractThe ground-penetrating radar (GPR) has been widely used in many applications. However, the processing and interpretation of the acquired signals remain challenging tasks since an experienced user is required to manage the entire operation. In this paper, we present an automatic classification system to assess railway-ballast conditions. It is based on the extraction of magnitude spectra at salient frequencies and their classification using support vector machines. The system is evaluated on real-world railway GPR data. The experimental results show that the proposed method efficiently represents the GPR signal using a small number of coefficients and achieves a high classification rate when distinguishing GPR signals reflected by ballasts of different conditions. Wenbin Shao, Abdesselam Bouzerdoum, Son Lam Phung, Lijun Su, Buddhima Indraratna, Cholachat Rujikiatkamjorn |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2010 | A particle swarm optimization algorithm based on orthogonal designabstractThe last decade has witnessed a great interest in using evolutionary algorithms, such as genetic algorithms, evolutionary strategies and particle swarm optimization (PSO), for multivariate optimization. This paper presents a hybrid algorithm for searching a complex domain space, by combining the PSO and orthogonal design. In the standard PSO, each particle focuses only on the error propagated back from the best particle, without “communicating” with other particles. In our approach, this limitation of the standard PSO is overcome by using a novel crossover operator based on orthogonal design. Furthermore, instead of the “generating-and-updating” model in the standard PSO, the elitism preservation strategy is applied to determine the possible movements of the candidate particles in the subsequent iterations. Experimental results demonstrate that our algorithm has a better performance compared to existing methods, including five PSO algorithms and three evolutionary algorithms. Jie Yang 0009, Abdesselam Bouzerdoum, Son Lam Phung |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | Multi-resolution Mean-Shift Algorithm for Vector QuantizationabstractHere we propose to apply the mean-shift algorithm to the four image subbands generated by a DWT, namely the LL, LH, HL and HH subbands. The simulated annealing technique traditionally used to identify the modes of the distribution for different resolutions can be performed by exploration of the multi-scale DWT pyramid, avoiding the costly estimation by a full mean-shift at each level. Philippe Loic Marie Bouttefroy, Abdesselam Bouzerdoum, Azeddine Beghdadi, Son Lam Phung |
DCC | 4 |
| 2010 | On the analysis of background subtraction techniques using Gaussian Mixture ModelsabstractIn this paper, we conduct an investigation into background subtraction techniques using Gaussian Mixture Models (GMM) in the presence of large illumination changes and background variations. We show that the techniques used to date suffer from the trade-off imposed by the use of a common learning rate to update both the mean and variance of the component densities, which leads to a degeneracy of the variance and creates “saturated pixels”. To address this problem, we propose a simple yet effective technique that differentiates between the two learning rates, and imposes a constraint on the variance so as to avoid the degeneracy problem. Experimental results are presented which show that, compared to existing techniques, the proposed algorithm provides more robust segmentation in the presence of illumination variations and abrupt changes in background distribution. Philippe Loic Marie Bouttefroy, Abdesselam Bouzerdoum, Son Lam Phung, Azeddine Beghdadi |
ICASSP | 3 |
| 2010 | A training algorithm for sparse LS-SVM using Compressive SamplingabstractLeast Squares Support Vector Machine (LS-SVM) has become a fundamental tool in pattern recognition and machine learning. However, the main disadvantage is lack of sparseness of solutions. In this article Compressive Sampling (CS), which addresses the sparse signal representation, is employed to find the support vectors of LS-SVM. The main difference between our work and the existing techniques is that the proposed method can locate the sparse topology while training. In contrast, most of the traditional methods need to train the model before finding the sparse support vectors. An experimental comparison with the standard LS-SVM and existing algorithms is given for function approximation and classification problems. The results show that the proposed method achieves comparable performance with typically a much sparser model. Jie Yang 0009, Abdesselam Bouzerdoum, Son Lam Phung |
ICASSP | 3 |
| 2010 | Adaptive regularization for image restoration using a variational inequality approachabstractIn this paper, a generalized image restoration method is formulated as a variational inequality problem, whose solution is obtained using a dynamic system approach. In this method, the restored image and the regularization parameter are obtained simultaneously. In particular, the optimum regularization parameter is determined adaptively, depending on noise and image content. The restoration problem is presented in a generalized form so that it maybe be implemented using different norms; only L1and L2norms have been implemented in this paper. A comparison based on experimental results shows that the proposed method achieves comparable if not better performance as some of the existing state-of-the-art techniques. Matthew Andrew Kitchener, Abdesselam Bouzerdoum, Son Lam Phung |
ICIP | 3 |
| 2010 | Wavelet based nonlocal-means super-resolution for video sequencesabstractVideo sequence resolution enhancement became a popular research area during the last two decades. Although traditional super-resolution techniques have been successful in dealing with image sequences, many constraints such as global translation between frames, have to be imposed to obtain good performance. In this paper, we present a new wavelet-based nonlocal-means (WNLM) framework to bypass the motion estimation stage. It can handle complex motion changes between frames. Compared with the nonlocal-means (NLM) super-resolution framework, the proposed method provides better result in terms of PSNR and faster processing. Haoheng Zheng, Abdesselam Bouzerdoum, Son Lam Phung |
ICIP | 3 |
| 2010 | Improved Facial Expression Recognition with Trainable 2-D Filters and Support Vector MachinesabstractFacial expression is one way humans convey their emotional states. Accurate recognition of facial expressions is essential in perceptual human-computer interface, robotics and mimetic games. This paper presents a novel approach to facial expression recognition from static images that combines fixed and adaptive 2-D filters in a hierarchical structure. The fixed filters are used to extract primitive features. They are followed by the adaptive filters that are trained to extract more complex facial features. Both types of filters are non-linear and are based on the biological mechanism of shunting inhibition. The features are finally classified by a support vector machine. The proposed approach is evaluated on the JAFFE database with seven types of facial expressions: anger, disgust, fear, happiness, neutral, sadness and surprise. It achieves a classification rate of 96.7%, which compares favorably with several existing techniques for facial expression recognition tested on the same database. Peiyao Li, Son Lam Phung, Abdesselam Bouzerdoum, Fok Hing Chi Tivive |
ICPR | 2 |
| 2010 | Efficient SVM training with reduced weighted samplesabstractThis paper presents an efficient training approach for support vector machines that will improve their ability to learn from a large or imbalanced data set. Given an original training set, the proposed approach applies unsupervised learning to extract a smaller set of salient training exemplars, which are represented by weighted cluster centers and the target outputs. In subsequent supervised learning, the objective function is modified by introducing a weight for each new training sample and the corresponding penalty term. In this paper, we investigate two methods of defining the weight based on cluster vectors. The proposed SVM training is implemented and tested on two problems: (i) gender classification of facial images using the FERET data set; (ii) income prediction using the UCI Adult Census data set. Experiment results show that compared to standard SVM training, the proposed approach leads to much faster SVM training, produces a more compact classifier while maintaining generalization ability. Giang Hoang Nguyen, Son Lam Phung, Abdesselam Bouzerdoum |
IJCNN | 2 |
| 2010 | Dimensionality reduction using compressed sensing and its application to a large-scale visual recognition taskabstractThis paper presents a novel algorithm for the dimensionality reduction which employs compressed sensing (CS) to improve the generalization capability of a classifier, especially for large-scale data. Compared to traditional dimensionality reduction methods, the proposed algorithm makes no use of the problem-dependent parameters, nor does it require additional computation for the eigenvalue decomposition like PCA or LDA. Mathematically, the derived algorithm regards the input features as the dictionary in CS, and selects the features that minimize the residual output error iteratively, thus the resulting features have a direct correspondence to the performance requirements of the given problem. Furthermore, the proposed algorithm can be regarded as a sparse classifier, which selects discriminative features and classifies the training data simultaneously. Experimentally, the CS-based algorithm is tested with a hierarchical visual pattern recognition architecture. The simulation results show that not only does the proposed method utilize only 25% of full features while achieving the test accuracy of the original full architecture, but also its performance is competitive when compared to existing dimensionality reduction methods. Jie Yang 0009, Abdesselam Bouzerdoum, Fok Hing Chi Tivive, Son Lam Phung |
IJCNN | 4 |
| 2009 | A New Approach to Sparse Image Representation Using MMV and K-SVD
Jie Yang 0009, Abdesselam Bouzerdoum, Son Lam Phung |
ACIVS | 3 |
| 2009 | Vehicle Tracking Using Projective Particle FilterabstractThis article introduces a new particle filtering approach for object tracking in video sequences. The projective particle filter uses a linear fractional transformation, which projects the trajectory of an object from the real world onto the camera plane, thus providing a better estimate of the object position. In the proposed particle filter, samples are drawn from an importance density integrating the linear fractional transformation. This provides a better coverage of the feature space and yields a finer estimate of the posterior density. Experiments conducted on traffic video surveillance sequences show that the variance of the estimated trajectory is reduced, resulting in more robust tracking. Philippe Loic Marie Bouttefroy, Abdesselam Bouzerdoum, Son Lam Phung, Azeddine Beghdadi |
AVSS | 3 |
| 2009 | A Neural Network pruning approach based on Compressive SamplingabstractThe balance between computational complexity and the architecture bottlenecks the development of neural networks (NNs). An architecture that is too large or too small will influence the performance to a large extent in terms of generalization and computational cost. In the past, saliency analysis has been employed to determine the most suitable structure, however, it is time-consuming and the performance is not robust. In this paper, a family of new algorithms for pruning elements (weighs and hidden neurons) in neural networks is presented based on compressive sampling (CS) theory. The proposed framework makes it possible to locate the significant elements, and hence find a sparse structure, without computing their saliency. Experiment results are presented which demonstrate the effectiveness of the proposed approach. Jie Yang 0009, Abdesselam Bouzerdoum, Son Lam Phung |
IJCNN | 3 |
| 2008 | A supervised learning approach for imbalanced data setsabstractThis paper presents a new learning approach for pattern classification applications involving imbalanced data sets. In this approach, a clustering technique is employed to resample the original training set into a smaller set of representative training exemplars, represented by weighted cluster centers and their target outputs. Based on the proposed learning approach, four training algorithms are derived for feed-forward neural networks. These algorithms are implemented and tested on three benchmark data sets. Experimental results show that with the proposed learning approach, it is possible to design networks to tackle the class imbalance problem, without compromising the overall classification performance. Giang Hoang Nguyen, Abdesselam Bouzerdoum, Son Lam Phung |
ICPR | 3 |
| 2008 | Efficient supervised learning with reduced training exemplarsabstractIn this article, we propose a new supervised learning approach for pattern classification applications involving large or imbalanced data sets. In this approach, a clustering technique is employed to reduce the original training set into a smaller set of representative training exemplars, represented by weighted cluster centers and their target outputs. Based on the proposed learning approach, two training algorithms are derived for feed-forward neural networks. These algorithms are implemented and tested on two pattern classification applications - skin detection and image classification. Experimental results show that with the proposed learning approach, it is possible to design networks in a fraction of time taken by the standard learning approach, without compromising the generalization ability and overall classification performance. Giang Hoang Nguyen, Abdesselam Bouzerdoum, Son Lam Phung |
IJCNN | 3 |
| 2007 | Detecting People in Images: An Edge Density ApproachabstractIn this paper, we present a new method for detecting visual objects in digital images and video. The novelty of the proposed method is that it differentiates objects from non-objects using image edge characteristics. Our approach is based on a fast object detection method developed by Viola and Jones. While Viola and Jones use Harr-like features, we propose a new image feature - the edge density - that can be computed more efficiently. When applied to the problem of detecting people and pedestrians in images, the new feature shows a very good discriminative capability compared to the Harr-like features. Son Lam Phung, Abdesselam Bouzerdoum |
ICASSP (1) | 1 |
| 2007 | A Pyramidal Neural Network For Visual Pattern RecognitionabstractIn this paper, we propose a new neural architecture for classification of visual patterns that is motivated by the two concepts of image pyramids and local receptive fields. The new architecture, called pyramidal neural network (PyraNet), has a hierarchical structure with two types of processing layers: Pyramidal layers and one-dimensional (1-D) layers. In the new network, nonlinear two-dimensional (2-D) neurons are trained to perform both image feature extraction and dimensionality reduction. We present and analyze five training methods for PyraNet [gradient descent (GD), gradient descent with momentum, resilient back-propagation (RPROP), Polak-Ribiere conjugate gradient (CG), and Levenberg-Marquadrt (LM)] and two choices of error functions [mean-square-error (mse) and cross-entropy (CE)]. In this paper, we apply PyraNet to determine gender from a facial image, and compare its performance on the standard facial recognition technology (FERET) database with three classifiers: The convolutional neural network (NN), the k-nearest neighbor (k-NN), and the support vector machine (SVM). Son Lam Phung, Abdesselam Bouzerdoum |
IEEE Trans. Neural Networks | 1 |
| 2006 | Gender Classification Using a New Pyramidal Neural Network
Son Lam Phung, Abdesselam Bouzerdoum |
ICONIP (2) | 1 |
| 2005 | Skin Segmentation Using Color Pixel Classification: Analysis and ComparisonabstractThis paper presents a study of three important issues of the color pixel classification approach to skin segmentation: color representation, color quantization, and classification algorithm. Our analysis of several representative color spaces using the Bayesian classifier with the histogram technique shows that skin segmentation based on color pixel classification is largely unaffected by the choice of the color space. However, segmentation performance degrades when only chrominance channels are used in classification. Furthermore, we find that color quantization can be as low as 64 bins per channel, although higher histogram sizes give better segmentation performance. The Bayesian classifier with the histogram technique and the multilayer perceptron classifier are found to perform better compared to other tested classifiers, including three piecewise linear classifiers, three unimodal Gaussian classifiers, and a Gaussian mixture classifier. Son Lam Phung, Abdesselam Bouzerdoum, Douglas Chai |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Naive bayes face/nonface classifier: a study of preprocessing and feature extraction techniquesabstractThis paper presents a classifier of face and nonface patterns that is based on the naive Bayes model. Using this classifier as a tool. We analyze the effects on classification performance of preprocessing, feature extraction and classifier combination techniques. Our analysis shows that image normalization techniques that reduce the effects of different lighting conditions improve face-nonface classification significantly. In addition, techniques such as background masking and combining classifiers that use different feature vectors are shown to enhance classification performance. Over a test set of 12,000 patterns, the combined classifier using four feature vectors has correct detection rates (CDRs) of 96.2% and 99.2% at false detection rates (FDRs) of 1% and 5%, respectively. Son Lam Phung, Abdesselam Bouzerdoum, Douglas Chai, Anthony Watson |
ICIP | 1 |
| 2003 | Adaptive skin segmentation in color imagesabstractA new skin segmentation technique for color images is proposed. The proposed technique uses a human skin color model that is based on the Bayesian decision theory and developed using a large training set of skin colors and nonskin colors. The proposed technique is novel and unique in that texture characteristics of the human skin are used to select appropriate skin color thresholds for skin segmentation. Two homogeneity measures for skin regions that take into account both global and local image features are also proposed. Experimental results showed that the proposed technique can achieve good skin segmentation performance (false detection rate of 4.5% and false rejection rate of 4.0%). Son Lam Phung, Douglas Chai, Abdesselam Bouzerdoum |
ICASSP (3) | 1 |
| 2003 | Adaptive skin segmentation in color imagesabstractA new skin segmentation technique for color images is proposed. The proposed technique uses a human skin color model that is based on the Bayesian decision theory and developed using a large training set of skin colors and nonskin colors. The proposed technique is novel and unique in that texture characteristics of the human skin are used to select appropriate skin color thresholds for skin segmentation. Two homogeneity measures for skin regions that take into account both global and local image features are also proposed. Experimental results showed that the proposed technique can achieve good skin segmentation performance (false detection rate of 4.5% and false rejection rate of 4.0%). Son Lam Phung, Douglas Chai, Abdesselam Bouzerdoum |
ICME | 1 |
| 2002 | A novel skin color model in YCbCr color space and its application to human face detectionabstractThis paper presents a new human skin color model in YCbCr color space and its application to human face detection. Skin colors are modeled by a set of three Gaussian clusters, each of which is characterized by a centroid and a covariance matrix. The centroids and covariance matrices are estimated from large set of training samples after a k-means clustering process. Pixels in a color input image can be classified into skin or non-skin based on the Mahalanobis distances to the three clusters. Efficient post-processing techniques namely noise removal, shape criteria, elliptic curve fitting and face/non-face classification are proposed in order to further refine skin segmentation results for the purpose of face detection. Son Lam Phung, Abdesselam Bouzerdoum, Douglas Chai |
ICIP (1) | 1 |