EDBT 2026 Demo / reviewers in the wild / expert
Alexander Wong
dblp:52/4401 · also Andy Wong
· DBLP profile ↗
131ranked-venue papers
29as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 18 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 42 · 9 first-author · 7 since 2021Artificial intelligence and machine learning · 32 · 5 first-author · 11 since 2021Computer networks · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCALEX: Scalable Concept and Latent Exploration for Diffusion ModelsabstractImage generation models frequently encode social biases, including stereotypes tied to gender, race, and profession. Existing methods for analyzing these biases in diffusion models either focus narrowly on predefined categories or depend on manual interpretation of latent directions. These constraints limit scalability and hinder the discovery of subtle or unanticipated patterns.We introduce SCALEX, a framework for scalable and automated exploration of diffusion model latent spaces. SCALEX extracts semantically meaningful directions from H-space using only natural language prompts, enabling zero-shot interpretation without retraining or labelling. This allows systematic comparison across arbitrary concepts and large-scale discovery of internal model associations. We show that SCALEX detects gender bias in profession prompts, ranks semantic alignment across identity descriptors, and reveals clustered conceptual structure without supervision. By linking prompts to latent directions directly, SCALEX makes bias analysis in diffusion models more scalable, interpretable, and extensible than prior approaches. E. Zhixuan Zeng, Yuhao Chen 0001, Alexander Wong |
WACV | 3 |
| 2025 | Facilitating Long Context Understanding via Supervised Chain-of-Thought ReasoningabstractRecent advances in Large Language Models (LLMs) have enabled them to process increasingly longer sequences, ranging from 2K to 2M tokens and even beyond.However, simply extending the input sequence length does not necessarily lead to effective long-context understanding.In this study, we integrate Chainof-Thought (CoT) reasoning into LLMs in a supervised manner to facilitate effective longcontext understanding.To achieve this, we introduce LongFinanceQA, a synthetic dataset in the financial domain designed to improve longcontext reasoning.Unlike existing long-context synthetic data, LongFinanceQA includes intermediate CoT reasoning before the final conclusion, which encourages LLMs to perform explicit reasoning, improving accuracy and interpretability in long-context understanding.To generate synthetic CoT reasoning, we propose Property-based Agentic Inference (PAI), an agentic framework that simulates humanlike reasoning steps, including property extraction, retrieval, and summarization.We evaluate PAI's reasoning capabilities by assessing GPT-4o-mini w/ PAI on the Loong benchmark, outperforming standard GPT-4o-mini by 20.0%.Furthermore, we fine-tune LLaMA-3.1-8B-Instruct on LongFinanceQA, achieving a 28.0% gain on Loong's financial subset. Alexander Wong, Shenghua He, Jiebo Luo 0001 |
EMNLP | 2 |
| 2025 | 3D Human Pose Estimation with MusclesabstractWe introduce MusclePose as an end-to-end learnable physics-infused 3D human pose estimator that incorporates muscle-dynamics modeling to infer human dynamics from monocular video. Current physics pose estimators aim to predict physically plausible poses by enforcing the underlying dynamics equations that govern motion. Since this is an underconstrained problem without force-annotated data, methods often estimate kinetics with external physics optimizers that may not be compatible with existing learning frameworks, or are too slow for real-time inference. While more recent methods use a regression-based approach to overcome these issues, the estimated kinetics can be seen as auxiliary predictions, and may not be physically plausible. To this end, we build on existing regression-based approaches, and aim to improve the biofidelity of kinetic inference with a multihypothesis approach --- by inferring joint torques via Lagrange’s equations and via muscle dynamics modeling with muscle torque generators. Furthermore, MusclePose predicts detailed human anthropometrics based on values from biomechanics studies, in contrast to existing physics pose estimators that construct their human models with shape primitives. We show that MusclePose is competitive with existing 3D pose estimators in positional accuracy, while also able to infer plausible human kinetics and muscle signals consistent with values from biomechanics studies, without requiring an external physics engine. Kevin Zhu, Ali Asghar Mohammadi Nasrabadi, Alexander Wong, John McPhee 0001 |
NeurIPS | 3 |
| 2025 | VAIRO: A Vision-Based Adaptive Impedance-Control Robotic FrameworkabstractIn this work, we present VAIRO, a Vision-based Adaptive Impedance-control RObotic framework for the purpose of manipulating soft materials, centered around the use case of rolling croissant dough for use in artisanal bakeries. Traditional automated processes for the industrial production of croissants consist of overly bulky equipment and fail to preserve the artisanal quality of hand-rolled croissants, with one of the major challenges being the high variability in the dough properties. VAIRO addresses these challenges by introducing a novel vision-based adaptive Cartesian impedance control strategy for collaborative robot arms to regulate rolling forces in real-time without the need for estimating the properties of the soft material. As such, VAIRO mimics the tactile adjustments made by human pastry chefs, ensuring consistent layer thickness and eliminating gaps. Using a Kinova Gen3 robotic arm and a custom-designed end-effector, we demonstrate that VAIRO can successfully manipulate various "doughs" without estimating any material properties. These results are promising and offer a cost-effective, small-scale alternative for local craft bakeries to leverage automation while maintaining high artisanal quality. Jeffrey Lee, Alexander Wong, Yue Hu 0001 |
RO-MAN | 2 |
| 2024 | Spot the Difference! Temporal Coarse to Fine to Finer Difference Spotting for Action Recognition in VideosabstractIn this paper, we present a novel difference-spotting strategy for video action recognition inspired by the cognitive challenges posed by the childhood puzzle game "Spot the Difference". Our approach aims to enhance the model’s capability to capture time-series variation and intricate details by gradually integrating distinctive information between action and non-action segments in a temporal "coarse-to-fine-to-finer" manner within a discriminative learning framework. To achieve this, we propose a model-agnostic discriminative learning mechanism that can be easily integrated into existing action recognition networks. Firstly, we incorporate coarse-level discriminative information of action and non-action segments across all videos in a corpus using novel booster nets. Secondly, we introduce a fine-level discrimination objective in the penultimate layer of the network through a novel contrastive learning approach, increasing the distinction between different segments within the same video. Lastly, we incorporate finer discrimination through a novel clip matching mechanism, enhancing the distinction of different consecutive clips within an action segment. Experimental results on multiple benchmark datasets (ActivityNet, HACS, FineAction) and backbone architectures (TSN, TSM, TANet, TPN, Timesformer, VideoSwin) demonstrate the effectiveness of our proposed mechanism. We consistently achieve significant improvements (0.33 - 4%) over the baselines, with competitive single crop results on ActivityNet (87.9%) and HACS (90.21%) datasets. Moreover, our technique achieves stateof-the-art classifier results (94.8%) in the ActivityNet 2022 challenge’s validation set. Yaoxin Li, Deepak Sridhar, Hanwen Liang, Alexander Wong |
ICME | 4 |
| 2024 | Empowering Tuberculosis Screening with Explainable Self-Supervised Deep Neural NetworksabstractTuberculosis remains a global health crisis, disproportionately affecting resource-limited populations and remote regions, with over 10 million new infections annually. Though curable, early detection is crucial. Chest X-rays are the primary screening tool, but their use requires skilled radiologists, often unavailable in underserved areas. This highlights the need for AI-powered systems to assist in rapid screening. However, training reliable AI models requires large-scale, high-quality data, which is costly and challenging to obtain. To address this, we introduce an explainable self-supervised learning network for tuberculosis screening, achieving 98.14% accuracy, with recall and precision rates of 95.72% and 99.44%, respectively. Alexander Wong, Ashkan Ebadi |
ICMLA | 2 |
| 2024 | Synthetic Local Data AugmentationabstractModern object segmentation models are crucial in sports analytics, particularly in dynamic sports like hockey where fast-paced action often results in blurred imagery, such as motion-blurred hockey sticks. Given the shortage of segmentation data for uncommon objects like hockey sticks, data augmentation emerges as a natural solution to enhance training datasets. However, traditional data augmentation methods, which apply transformations at the image level, can distort critical relational cues between objects and their surroundings, undermining a model's ability to accurately segment objects in such challenging conditions. To address this, we propose the Synthetic Local Data Augmentation (SLDA) technique, which selectively applies traditional DA transformations-like scaling, rotation, blurring, and motion blur-directly to individual target objects. This technique allows precise customization of transformations to specifically enhance model robustness against particular types of distortions, such as the motion blur frequently observed with fast-moving hockey sticks. Utilizing a segmented dataset of hockey sticks, SLDA introduces a greater variety of stick instances by inserting elements in the scene with different examples of the same category. This focused approach significantly enhances the model's ability to recognize hockey sticks across a range of visual conditions, thereby improving its generalization capabilities. SLDA detailed experiments in a case study on hockey stick seg-mentation, we demonstrate how SLDA surpasses existing object-level and traditional data augmentation methods in promoting model robustness and adaptive precision. Surpassing alternative by 2.1 %, i.e. from 85% to 87% in F1 Score on small model complexity, and by 5.8%, i.e. from 86% to 92% in mAP50 on large model complexity. Vasyl Chomko, Yuhao Chen 0001, David A. Clausi, Alexander Wong |
MMSP | 4 |
| 2024 | Integrating deep transformer and temporal convolutional networks for SMEs revenue and employment growth prediction
Dening Lu, Shimon Schwartz, Linlin Xu, Mohammad Javad Shafiee, Norman G. Vinson, Chris Czarnecki, Alexander Wong |
Expert Syst. Appl. | 7 |
| 2024 | Modeling the Role of Contour Integration in Visual InferenceabstractUnder difficult viewing conditions, the brain's visual system uses a variety of recurrent modulatory mechanisms to augment feedforward processing. One resulting phenomenon is contour integration, which occurs in the primary visual (V1) cortex and strengthens neural responses to edges if they belong to a larger smooth contour. Computational models have contributed to an understanding of the circuit mechanisms of contour integration, but less is known about its role in visual perception. To address this gap, we embedded a biologically grounded model of contour integration in a task-driven artificial neural network and trained it using a gradient-descent variant. We used this model to explore how brain-like contour integration may be optimized for high-level visual objectives as well as its potential roles in perception. When the model was trained to detect contours in a background of random edges, a task commonly used to examine contour integration in the brain, it closely mirrored the brain in terms of behavior, neural responses, and lateral connection patterns. When trained on natural images, the model enhanced weaker contours and distinguished whether two points lay on the same versus different contours. The model learned robust features that generalized well to out-of-training-distribution stimuli. Surprisingly, and in contrast with the synthetic task, a parameter-matched control network without recurrence performed the same as or better than the model on the natural-image tasks. Thus, a contour integration mechanism is not essential to perform these more naturalistic contour-related tasks. Finally, the best performance in all tasks was achieved by a modified contour integration model that did not distinguish between excitatory and inhibitory neurons. Alexander Wong, Bryan P. Tripp |
Neural Comput. | 2 |
| 2024 | MetaGraspNetV2: All-in-One Dataset Enabling Fast and Reliable Robotic Bin Picking via Object Relationship Reasoning and Dexterous GraspingabstractGrasping unknown objects in unstructured environments is one of the most challenging and demanding tasks for robotic bin picking systems. Developing a holistic approach is crucial to building such dexterous bin picking systems to meet practical requirements on speed, cost and reliability. Proposed datasets so far focus only on challenging sub-problems and are therefore limited in their ability to leverage the complementary relationship between individual tasks. In this paper, we tackle this holistic data challenge and design MetaGraspNetV2, an all-in-one bin picking dataset consisting of (i) a photo-realistic dataset with over 296k images, which has been created through physics-based metaverse synthesis; and (ii) a real-world test dataset with 3.2k images featuring task-specific difficulty levels. Both datasets provide full annotations for amodal panoptic segmentation, object relationship detection, occlusion reasoning, 6-DoF pose estimation, and grasp detection for a parallel-jaw as well as a vacuum gripper. Extensive experiments demonstrate that our dataset outperforms state-of-the-art datasets in object detection, instance segmentation, amodal detection, parallel-jaw grasping, and vacuum grasping. Furthermore, leveraging the potential of our data for building holistic perception systems, we propose a single-shot-multi-pick (SSMP) grasping policy for scene understanding accelerated fast picking in high clutter. SSMP reasons about suitable manipulation orders for blindly picking multiple items given a single image acquisition. Physical robot experiments demonstrate that SSMP effectively speeds up cycle times through reducing image acquisitions by more than 47% while providing better grasp performance compared to state-of-the-art bin picking methods.Note to Practitioners—In robotic bin picking, most proposed methods and datasets focus on solving only one aspect of the grasping task, such as grasp point detection, object detection, or relationship reasoning. They do not address practical aspects such as the widespread use of vacuum grasp technology or the need for short cycle times. In practice, however, efficient bin picking solutions often rely on multiple task-specific methods. Hence, having one dataset for a large variety of vision-related tasks in robotic picking reduces data redundancy and enables the development of holistic methods. While deep learning has been proven highly effective for bin picking vision systems, it demands large, high-quality training datasets. Collecting such datasets in the real-world, while assuring label quality and consistency, is prohibitively expensive and time-consuming. To overcome these challenges, we set up a photo-realistic metaverse data generation pipeline and create a large-scale synthetic training dataset. Furthermore, we design a comprehensive real-world dataset for testing. Unlike previously proposed datasets, our datasets provide difficulty levels and annotations in simulation and real-world for a comprehensive list of high-level tasks, including amodal object detection, scene layout reasoning, and grasp detection. In real-world applications, cycle time is a critical factor affecting the productivity and profitability of a robotic system. We tackle time-efficiency through scene understanding and demonstrate the capability of our data regarding holistic system development by proposing a single-shot-multi-pick (SSMP) policy. Our SSMP algorithm, trained exclusively on our synthetic data, distinguishes between uncovered and occluded items, and infers specific manipulation orders to perform multiple blind picks in a single shot. Physical robot experiments show that SSMP was able to reduce image acquisitions by more than 47% without compromising grasp performance. This clearly demonstrates that SSMP, together with our dataset, paves the way for application-oriented research in time-critical bin picking. Maximilian Gilles, Yuhao Chen 0001, E. Zhixuan Zeng, Yifan Wu 0004, Kai Furmans, Alexander Wong, Rania Rayyes |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2023 | SolderNet: Towards Trustworthy Visual Inspection of Solder Joints in Electronics Manufacturing Using Explainable Artificial IntelligenceabstractIn electronics manufacturing, solder joint defects are a common problem affecting a variety of printed circuit board components. To identify and correct solder joint defects, the solder joints on a circuit board are typically inspected manually by trained human inspectors, which is a very time-consuming and error-prone process. To improve both inspection efficiency and accuracy, in this work we describe an explainable deep learning-based visual quality inspection system tailored for visual inspection of solder joints in electronics manufacturing environments. At the core of this system is an explainable solder joint defect identification system called SolderNet which we design and implement with trust and transparency in mind. While several challenges remain before the full system can be developed and deployed, this study presents important progress towards trustworthy visual inspection of solder joints in electronics manufacturing. Hayden Gunraj, Paul Guerrier, Sheldon Fernandez, Alexander Wong |
AAAI | 4 |
| 2023 | High-Throughput, High-Performance Deep Learning-Driven Light Guide Plate Surface Visual Quality Inspection Tailored for Real-World Manufacturing EnvironmentsabstractLight guide plates are essential optical components widely used in a diverse range of applications ranging from medical lighting fixtures to back-lit TV displays. An essential step in the manufacturing of light guide plates is the quality inspection of defects such as scratches, bright/dark spots, and impurities. This is mainly done in industry through manual visual inspection for plate pattern irregularities, which is time-consuming and prone to human error and thus act as a significant barrier to high-throughput production. Advances in deep learning-driven computer vision has led to the exploration of automated visual quality inspection of light guide plates to improve inspection consistency, accuracy, and efficiency. However, given the computational constraints and high-throughput nature of real-world manufacturing environments, the widespread adoption of deep learning-driven visual inspection systems for inspecting light guide plates in real-world manufacturing environments has been greatly limited due to high computational requirements and integration challenges of existing deep learning approaches in research literature. In this work, we introduce a fully-integrated, high-throughput, high-performance deep learning-driven workflow for light guide plate surface visual quality inspection (VQI) tailored for real-world manufacturing environments. To enable automated VQI on the edge computing within the fully-integrated VQI system, a highly compact deep anti-aliased attention condenser neural network (which we name Light-DefectNet) tailored specifically for light guide plate surface defect detection in resource-constrained scenarios was created via machine-driven design exploration with computational and “best-practices” constraints as well as L1 paired classification discrepancy loss. Experiments show that Light-DetectNet achieves a detection accuracy of ∼98.2% on the LGPSDD benchmark while having just 770K parameters (∼33× and ∼6.9× lower than ResNet-50 and EfficientNet-B0, respectively) and ∼93M FLOPs (∼88× and ∼8.4× lower than ResNet-50 and EfficientNet-B0, respectively) and ∼8.8× faster inference speed than EfficientNet-B0 on an embedded ARM processor. As such, the proposed deep learning-driven workflow, integrated with the aforementioned LightDefectNet neural network, is highly suited for high-throughput, high-performance light plate surface VQI within real-world manufacturing environments. Carol Xu, Mahmoud Famouri, Gautam Bathla, Mohammad Javad Shafiee, Alexander Wong |
AAAI | 5 |
| 2023 | AI-Powered Noncontact In-Home Gait Monitoring and Activity Recognition System Based on mm-Wave FMCW Radar and Cloud ComputingabstractIn this work, we present a cloud-based system for non-contact, real-time recognition and monitoring of physical activities and walking periods within a domestic environment. The proposed system employs standalone Internet of Things (IoT)-based millimeter wave radar devices and deep learning models to enable autonomous, free-living activity recognition and gait analysis. To train deep learning models, we utilize range-Doppler maps generated from a dataset of real-life in-home activities. The performance of several deep learning models is evaluated based on accuracy and prediction time, with the gated recurrent network (GRU) model selected for real-time deployment due to its balance of speed and accuracy compared to 2D Convolutional Neural Network Long Short-Term Memory (2D-CNNLSTM) and Long Short-Term Memory (LSTM) models. The overall accuracy of the GRU model for classifying in-home physical activities of trained subjects is 93%, with 86% accuracy for a new subject. In addition to recognizing and differentiating various activities and walking periods, the system also records the subject’s activity level over time, washroom use frequency, sleep/sedentary/active/out-of-home durations, current state, and gait parameters. Importantly, the system maintains privacy by not requiring the subject to wear or carry any additional devices. Hajar Abedi, Ahmad Ansariyan, Plinio Pelegrini Morita, Alexander Wong, Jennifer Boger, George Shaker |
IEEE Internet Things J. | 4 |
| 2023 | Unsupervised Bayesian Subpixel Mapping Autoencoder Network for Hyperspectral ImagesabstractUnsupervised subpixel mapping (SPM) of hyperspectral image (HSI) is a challenging task due to the difficulties to integrate different prior information and model constraints into a coherent framework. This paper presents a Bayesian neural network for unsupervised HSI SPM, which has the following characteristics. First, the deep image prior (DIP) achieved by a fully convolutional neural network (FCNN) is used to model the spatial correlation efficiently and adaptively in the subpixel label domain. Second, a discrete spectral mixture model (DSMM) is designed to leverage the forward model for enhanced SPM. Third, an auto-encoder architecture is designed to integrate the FCNN and the DSMM to allow efficient unsupervised representational learning using both data and knowledge. Fourth, an expectation-maximization approach is designed to solve the resulting maximum a posteriori problem, where a purified means approach extracts endmembers, and the gradient descent approach updates FCNN parameters for subpixel label estimation. Comparative experiments on both real and simulated HSIs demonstrate that the proposed method outperforms other state-of-the-art methods in terms of both numerical accuracies and visual subpixel mapping results. Yuan Fang 0003, Yuxian Wang, Linlin Xu, Yujia Chen 0002, Alexander Wong, David A. Clausi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | LexSubCon: Integrating Knowledge from Lexical Resources into Contextual Embeddings for Lexical SubstitutionabstractLexical substitution is the task of generating meaningful substitutes for a word in a given textual context.Contextual word embedding models have achieved state-of-the-art results in the lexical substitution task by relying on contextual information extracted from the replaced word within the sentence.However, such models do not take into account structured knowledge that exists in external lexical databases.We introduce LexSubCon, an end-to-end lexical substitution framework based on contextual embedding models that can identify highly-accurate substitute candidates.This is achieved by combining contextual information with knowledge from structured lexical resources.Our approach involves: (i) introducing a novel mix-up embedding strategy to the target word's embedding through linearly interpolating the pair of the target input embedding and the average embedding of its probable synonyms; (ii) considering the similarity of the sentence-definition embeddings of the target word and its proposed candidates; and, (iii) calculating the effect of each substitution on the semantics of the sentence through a fine-tuned sentence similarity model.Our experiments show that LexSubCon outperforms previous state-of-the-art methods by at least 2% over all the official lexical substitution metrics on LS07 and CoInCo benchmark datasets that are widely used for lexical substitution tasks. George Michalopoulos, Ian McKillop, Alexander Wong, Helen H. Chen |
ACL (1) | 3 |
| 2022 | Rethinking Keypoint Representations: Modeling Keypoints and Poses as Objects for Multi-person Human Pose Estimation
William J. McNally, Kanav Vats, Alexander Wong, John McPhee 0001 |
ECCV (6) | 3 |
| 2022 | TAL: Topography-Aware Multi-Resolution Fusion Learning for Enhanced Building Footprint ExtractionabstractAutomatic building footprint extraction from remote sensing imagery is a challenging task with important applications in geomatics and environmental science. Significant advances have been made in this field as a result of the emergence of deep convolutional neural networks (CNNs) designed for semantic segmentation. Although CNNs have demonstrated state-of-the-art performance in coarse annotation and identification of buildings, the accuracy of extracted building footprints is still insufficient for high-precision applications such as mapping and navigation. We propose the topography-aware multi-resolution fusion learning strategy tailored to the problem of enhanced building footprint extraction. More specifically, we introduce a topography-aware loss (TAL) for enhancing a deep CNN’s ability to learn heterogeneous building features for better boundary preservation during segmentation. We then incorporate the proposed TAL loss within a multi-resolution fusion architecture to boost high-resolution segmentation performance. Finally, we introduce a novel metric named average thresholded contour accuracy (tCA) which specifically measures the accuracy of segmentation boundaries. The experimental results on the SpaceNet buildings dataset show significant improvements in boundary integrity of extracted building footprints when compared with previously proposed methods. Hence, this method enables accurate boundary annotation toward automatic production of building footprint maps for high-precision applications. Yifan Wu 0004, Linlin Xu, Yuhao Chen 0001, Alexander Wong, David A. Clausi |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | BCUN: Bayesian Fully Convolutional Neural Network for Hyperspectral Spectral UnmixingabstractSpectral unmixing (SU) plays a fundamental role in hyperspectral image (HSI) processing. Effective SU relies on the accurate and efficient characterization of the noise effect, the endmembers, and the spatial correlation effect in abundances, as well as efficient optimization techniques to estimate these effects. To address these issues, this article presents a Bayesian fully convolutional hyperspectral unmixing network (BCUN) with the following key characteristics. First, a fully convolutional neural network (FCNN)-based deep image prior (DIP) is designed for enhanced characterization and estimation of the spatial context information in abundance maps, leading to more efficient and accurate abundance modeling than the traditional nonnegative least squares (NNLS) approaches. Second, a multivariate Gaussian distribution with an anisotropic covariance matrix is designed to characterize the conditional distribution of the spectral observations, leading to a novel Mahalanobis distance-based loss for FCNN training that is better capable of addressing the noise heterogeneous effect in HSI than the Euclidean distance-based mean squared error (MSE) loss in traditional deep neural networks. Third, the designed conditional distribution of spectral observations also enables the incorporation of the spectral mixture model (SMM) into the FCNN training process for effectively leveraging the knowledge in the forward spectral model. Fourth, the endmembers are modeled and estimated by a “purified means” approach that is capable of better characterizing endmembers. Finally, the above key components are coherently integrated into a Bayesian framework, and the resulting maximuma posteriori(MAP) problem is solved by a designed expectation–maximization (EM) algorithm. Experimental results on both simulated and real HSIs demonstrate that the proposed BCUN approach outperforms the other classical and state-of-the-art methods on both endmember estimation and abundance estimation. Yuan Fang 0003, Yuxian Wang, Linlin Xu, Rongming Zhuo, Alexander Wong, David A. Clausi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | VidAF: A Motion-Robust Model for Atrial Fibrillation Screening From Facial VideosabstractAtrial fibrillation (AF) is the most common arrhythmia, but an estimated 30% of patients with AF are unaware of their conditions. The purpose of this work is to design a model for AF screening from facial videos, with a focus on addressing typical motion disturbances in our real life, such as head movements and expression changes. This model detects a pulse signal from the skin color changes in a facial video by a convolution neural network, incorporating a phase-driven attention mechanism to suppress motion signals in the space domain. It then encodes the pulse signal into discriminative features for AF classification by a coding neural network, using a de-noise coding strategy to improve the robustness of the features to motion signals in the time domain. The proposed model was tested on a dataset containing 1200 samples of 100 AF patients and 100 non-AF subjects. Experimental results demonstrated that VidAF had significant robustness to facial motions, predicting clean pulse signals with the mean absolute error of inter-pulse intervals less than 100 milliseconds. Besides, the model achieved promising performance in AF identification, showing an accuracy of more than 90% in multiple challenging scenarios. VidAF provides a more convenient and cost-effective approach for opportunistic AF screening in the community. Xuenan Liu, Xuezhi Yang, Dingliang Wang, Alexander Wong, Likun Ma |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | UmlsBERT: Clinical Domain Knowledge Augmentation of Contextual Embeddings Using the Unified Medical Language System MetathesaurusabstractGeorge Michalopoulos, Yuanxin Wang, Hussam Kaka, Helen Chen, Alexander Wong. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. George Michalopoulos, Yuanxin Wang 0001, Hussam Kaka, Helen H. Chen, Alexander Wong |
NAACL-HLT | 5 |
| 2021 | Unsupervised Bayesian Subpixel Mapping of Hyperspectral Imagery Based on Band-Weighted Discrete Spectral Mixture Model and Markov Random FieldabstractAlthough accurate training and initialization information is difficult to acquire, unsupervised hyperspectral subpixel mapping (SPM) without relying on this predefined information is an insufficiently addressed research issue. This letter presents a novel Bayesian approach for unsupervised SPM of hyperspectral imagery (HSI) based on the Markov random field (MRF) and a band-weighted discrete spectral mixture model (BDSMM), with the following key characteristics. First, this is an unsupervised approach that allows adjustment of abundance and endmember information adaptively for less relying on algorithm initialization. Second, this approach consists of the BDSMM for accommodating the noise heterogeneity and the hidden label field of subpixels in HSI. The BDSMM also integrates SPM into the spectral mixture analysis and allows enhanced SPM by fully exploring the endmember-abundance patterns in HSI. Third, the MRF and BDSMM are integrated into a Bayesian framework to use both the spatial and spectral information efficiently, and an expectation-maximization (EM) approach is designed to solve the model by iteratively estimating the endmembers and the label field. Experiments on both simulated and real HSI demonstrate that the proposed algorithm can yield better performance than traditional methods. Yujia Chen 0002, Linlin Xu, Yuan Fang 0003, Junhuan Peng, Wenfu Yang, Alexander Wong, David A. Clausi |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2020 | Investigating the Impact of Inclusion in Face Recognition Training Data on Individual Face IdentificationabstractModern face recognition systems leverage datasets containing images of hundreds of thousands of specific individuals' faces to train deep convolutional neural networks to learn an embedding space that maps an arbitrary individual's face to a vector representation of their identity. The performance of a face recognition system in face verification (1:1) and face identification (1:N) tasks is directly related to the ability of an embedding space to discriminate between identities. Recently, there has been significant public scrutiny into the source and privacy implications of large-scale face recognition training datasets such as MS-Celeb-1M and MegaFace, as many people are uncomfortable with their face being used to train dual-use technologies that can enable mass surveillance. However, the impact of an individual's inclusion in training data on a derived system's ability to recognize them has not previously been studied. In this work, we audit ArcFace, a state-of-the-art, open source face recognition system, in a large-scale face identification experiment with more than one million distractor images. We find a Rank-1 face identification accuracy of 79.71% for individuals present in the model's training data and an accuracy of 75.73% for those not present. This modest difference in accuracy demonstrates that face recognition systems using deep learning work better for individuals they are trained on, which has serious privacy implications when one considers all major open source face recognition training datasets do not obtain informed consent from individuals during their collection. Chris Dulhanty, Alexander Wong |
AIES | 2 |
| 2020 | Learn2Perturb: An End-to-End Feature Perturbation Learning to Improve Adversarial RobustnessabstractWhile deep neural networks have been achieving state-of-the-art performance across a wide variety of applications, their vulnerability to adversarial attacks limits their widespread deployment for safety-critical applications. Alongside other adversarial defense approaches being investigated, there has been a very recent interest in improving adversarial robustness in deep neural networks through the introduction of perturbations during the training process. However, such methods leverage fixed, pre-defined perturbations and require significant hyper-parameter tuning that makes them very difficult to leverage in a general fashion. In this study, we introduce Learn2Perturb, an end-to-end feature perturbation learning approach for improving the adversarial robustness of deep neural networks. More specifically, we introduce novel perturbation-injection modules that are incorporated at each layer to perturb the feature space and increase uncertainty in the network. This feature perturbation is performed at both the training and the inference stages. Furthermore, inspired by the Expectation-Maximization, an alternating back-propagation training algorithm is introduced to train the network and noise parameters consecutively. Experimental results on CIFAR-10 and CIFAR-100 datasets show that the proposed Learn2Perturb method can result in deep neural networks which are 4-7% more robust on l_inf FGSM and PDG adversarial attacks and significantly outperforms the state-of-the-art against l_2 C\&W attack and a wide range of well-known black-box attacks. Ahmadreza Jeddi, Mohammad Javad Shafiee, Michelle Karg, Christian Scharfenberger, Alexander Wong |
CVPR | 5 |
| 2020 | Squeeze-and-Attention Networks for Semantic SegmentationabstractThe recent integration of attention mechanisms into segmentation networks improves their representational capabilities through a great emphasis on more informative features. However, these attention mechanisms ignore an implicit sub-task of semantic segmentation and are constrained by the grid structure of convolution kernels. In this paper, we propose a novel squeeze-and-attention network (SANet) architecture that leverages an effective squeeze-and-attention (SA) module to account for two distinctive characteristics of segmentation: i) pixel-group attention, and ii) pixel-wise prediction. Specifically, the proposed SA modules impose pixel-group attention on conventional convolution by introducing an 'attention' convolutional channel, thus taking into account spatial-channel inter-dependencies in an efficient manner. The final segmentation results are produced by merging outputs from four hierarchical stages of a SANet to integrate multi-scale contexts for obtaining an enhanced pixel-wise prediction. Empirical experiments on two challenging public datasets validate the effectiveness of the proposed SANets, which achieves 83.2 % mIoU (without COCO pre-training) on PASCAL VOC and a state-of-the-art mIoU of 54.4 % on PASCAL Context. Zilong Zhong, Zhong Qiu Lin, Rene Bidart, Xiaodan Hu, Ibrahim Ben Daya, Wei-Shi Zheng 0001, Jonathan Li 0001, Alexander Wong |
CVPR | 9 |
| 2020 | Quantization in Relative Gradient Angle Domain For Building Polygon EstimationabstractBuilding footprint extraction in remote sensing data benefits many important applications, such as urban planning and population estimation. Recently, rapid development of convolutional neural networks (CNNs) and open-sourced high resolution satellite building image datasets have pushed the performance boundary further for automated building extractions. However, CNN approaches often generate imprecise building morphologies including noisy edges and round corners. In this paper, we leverage the performance of CNNs, and propose a module that uses prior knowledge of building corners to create angular and concise building polygons from CNN segmentation outputs. We describe a new transform, Relative Gradient Angle Transform (RGA Transform) that converts object contours from time vs. space to time vs. angle. We propose a new shape descriptor, Boundary Orientation Relation Set (BORS), to describe angle relationship between edges in RGA domain, such as orthogonality and parallelism. Finally, we develop an energy minimization framework that makes use of the angle relationship in BORS to straighten edges and reconstruct sharp corners, and the resulting corners create a polygon. Experimental results demonstrate that our method refines CNN output from a rounded approximation to a more clear-cut angular shape of the building footprint. Yuhao Chen 0001, Yifan Wu 0004, Linlin Xu, Alexander Wong |
ICPR | 4 |
| 2020 | Unsupervised Domain Adaptation in Person re-ID via k-Reciprocal Clustering and Large-Scale Heterogeneous Environment SynthesisabstractAn ongoing major challenge in computer vision is the task of person re-identification, where the goal is to match individuals across different, non-overlapping camera views. While recent success has been achieved via supervised learning using deep neural networks, such methods have limited widespread adoption due to the need for large-scale, customized data annotation. As such, there has been a recent focus on unsupervised learning approaches to mitigate the data annotation issue; however, current approaches in literature have limited performance compared to supervised learning approaches as well as limited applicability for adoption in new environments. In this paper, we address the aforementioned challenges faced in person re-identification for real-world, practical scenarios by introducing a novel, unsupervised domain adaptation approach for person re-identification. This is accomplished through the introduction of: i) k-reciprocal tracklet Clustering for Unsupervised Domain Adaptation (ktCUDA) (for pseudo-label generation on target domain), and ii) Synthesized Heterogeneous RE-id Domain (SHRED) composed of large-scale heterogeneous independent source environments (for improving robustness and adaptability to a wide diversity of target environments). Experimental results across four different image and video benchmark datasets show that the proposed ktCUDA and SHRED approach achieves an average improvement of +5.7 mAP in re-identification performance when compared to existing state-of-the-art methods, as well as demonstrate better adaptability to different types of environments. Devinder Kumar, Parthipan Siva, Paul Marchwica, Alexander Wong |
WACV | 4 |
| 2020 | Generative Adversarial Networks and Conditional Random Fields for Hyperspectral Image ClassificationabstractIn this paper, we address the hyperspectral image (HSI) classification task with a generative adversarial network and conditional random field (GAN-CRF)-based framework, which integrates a semisupervised deep learning and a probabilistic graphical model, and make three contributions. First, we design four types of convolutional and transposed convolutional layers that consider the characteristics of HSIs to help with extracting discriminative features from limited numbers of labeled HSI samples. Second, we construct semisupervised generative adversarial networks (GANs) to alleviate the shortage of training samples by adding labels to them and implicitly reconstructing real HSI data distribution through adversarial training. Third, we build dense conditional random fields (CRFs) on top of the random variables that are initialized to the softmax predictions of the trained GANs and are conditioned on HSIs to refine classification maps. This semisupervised framework leverages the merits of discriminative and generative models through a game-theoretical approach. Moreover, even though we used very small numbers of labeled training HSI samples from the two most challenging and extensively studied datasets, the experimental results demonstrated that spectral-spatial GAN-CRF (SS-GAN-CRF) models achieved top-ranking accuracy for semisupervised HSI classification. Zilong Zhong, Jonathan Li 0001, David A. Clausi, Alexander Wong |
IEEE Trans. Cybern. | 4 |
| 2020 | Nonlocal Band-Weighted Iterative Spectral Mixture Model for Hyperspectral Imagery DenoisingabstractAlthough efficient hyperspectral image (HSI) denoising relies on complete and accurate description and modeling the spatial-spectral signal in HSI, the current approaches do not fully account for key characteristics of HSI, i.e., the mixed spectra effect, the spatial nonstationarity effect, and noise variance heterogeneity effect. To address this issue, this article presents a linear spectral mixture model with nonlocal means constraint (LSMM-NLMC), with the following advantages. First, LSMM-NLMC can effectively learn the signal in mixed pixels in HSI by estimating clean endmembers and abundances for image restoration. Second, LSMM-NLMC can efficiently address nonstationary spatial correlation effect by imposing NLMC on the latent scene signal. Last, LSMM-NLMC provides accurate noise characterization by accounting for noise variance heterogeneity effect using a band-dependent noise model and a band-weighted Mahalanobis distance for similarity measurement. A novel optimization method based on the expectation-maximization (EM) algorithm and the purified means approach is used to efficiently solve the resulting maximum a posterior (MAP) problem. The experiments on both simulated and real HSI data sets demonstrate that the visual quality and denoising accuracy are significantly improved by the proposed LSMM-NLMC compared with previous methods. Longshan Yang, Linlin Xu, Junhuan Peng, Yongze Song, Alexander Wong, David A. Clausi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Real-Time Vehicle Make and Model Recognition Using Unsupervised Feature LearningabstractVehicle Make and Model Recognition (VMMR) systems provide a fully automatic framework to recognize and classify different vehicle models. Several approaches have been proposed to address this challenge; however, they can perform in restricted conditions. Here, in this paper, we formulate the VMMR as a fine-grained classification problem and propose a new configurable on-road VMMR framework. We benefit from the unsupervised feature learning methods, and in more details, we employ Locality-constraint Linear Coding (LLC) method as a fast feature encoder for encoding the input SIFT features. The proposed method can perform in real environments of different conditions. This framework can recognize 50 models of vehicles and has the advantage to classify every other vehicle not belonging to one of the specified 50 classes as an unknown vehicle. The proposed VMMR framework can be configured to become faster or more accurate based on the application domain. The proposed approach is examined on two datasets, including Iranian on-road vehicle (IORV) dataset and CompuCar dataset. The IORV dataset contains images of 50 models of vehicles captured in real situations by traffic-cameras in different weather and lighting conditions. The experimental results show the advantage of the real-time configuration of the proposed framework over the state-of-the-art methods on the IORV datatset and comparable results on CompuCar dataset with 97.5% and 98.4% accuracies, respectively and acceptable running time. Amir Nazemi, Zohreh Azimifar, Mohammad Javad Shafiee, Alexander Wong |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2019 | StressedNets: Efficient feature representations via stress-induced evolutionary synthesis of deep neural networks
Mohammad Javad Shafiee, Brendan Chwyl, Francis Li, Rongyan Chen, Michelle Karg, Christian Scharfenberger, Alexander Wong |
Neurocomputing | 7 |
| 2019 | A Bayesian Joint Decorrelation and Despeckling of SAR ImageryabstractDespeckling of synthetic aperture radar (SAR) is a known research challenge. A novel solution to this problem has been developed and evaluated via an iterative maximum a posterior estimation incorporating a Bayesian joint decorrelation and despeckling based on a correlation model. This model realistically explores the physical correlation process of SAR speckle noise and is determined automatically via Bayesian estimation in the log-Fourier domain. A patchwise computation is used to account for the spatial nonstationarity associated with SAR image data. The proposed approach is compared to the existing despeckling techniques using both simulated and real SAR data, and the experimental results demonstrate the improvement in preserving the structural details while suppressing speckle noise. Linlin Xu, David A. Clausi, Alexander Wong |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2019 | A Novel Motion Plane-Based Approach to Vehicle Speed EstimationabstractSpeed limit violation by vehicles is one of the most frequent reasons for road crashes, which take the lives of many people every year, resulting in an increasing demand for video-based vehicle speed estimation systems. One of the biggest challenges to achieve monocular video-based vehicle speed estimation is the projection displacement difference (PDD) problem, where there is a plane disparity between street-level indicators and above-plane feature points of vehicles, thus resulting in unreliable speed estimates. In this paper, a novel motion plane-based approach for vehicle speed estimation is proposed, which addresses the problem of PDD. In the proposed method, we consider the center of a vehicle license plate as the vehicle reference point and estimate the hypothetical plane (named motion plane), on which license plate moves. Subsequently, the plate position is mapped on the motion plane and the displacement is then calculated, thus mitigating the effects of PPD. To estimate the motion plane, a texture-based shape-from-template technique is used. Unlike existing methods, the proposed method needs neither to apply any indicator nor to use any information about extrinsic parameters of the camera. Furthermore, since all the license plates are approximately located on a flat plane, the motion plane can be estimated by using several extracted 3-D points. Experimental results show that the proposed method performs better than the state-of-the-art vehicle speed estimation methods, illustrating the efficacy of this approach for achieving reliable vehicle speed estimation. Mahmoud Famouri, Zohreh Azimifar, Alexander Wong |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Deep Learning with Darwin: Evolutionary Synthesis of Deep Neural Networks
Mohammad Javad Shafiee, Akshaya Kumar Mishra, Alexander Wong |
Neural Process. Lett. | 3 |
| 2018 | Scalable image segmentation via decoupled sub-graph compression
R. S. Medeiros, Alexander Wong, Jacob Scharcanski |
Pattern Recognit. | 2 |
| 2018 | ST-IRGS: A Region-Based Self-Training Algorithm Applied to Hyperspectral Image Classification and SegmentationabstractThe problem of limited labeled training samples is challenging for the classification of remote sensing imagery. We develop a joint classification and segmentation algorithm to address this problem. Our algorithm combines semisupervised learning and conditional random fields (CRFs) into a single framework. The multimodal Gaussian maximum-likelihood classifier is used to estimate the probabilities for the unary potentials of the CRF. Unlike traditional methods based on random fields, region merging is concatenated with the CRF inference to reduce the number of nodes iteratively. Moreover, a semisupervised technique called self-training is used, which iteratively enlarges the training sample set and retrains the classifier. The selection of training samples is based on the region information, so that the risk of assigning wrong labels is largely reduced. The proposed algorithm is applied to hyperspectral image classification, and results on benchmark data sets show that the proposed algorithm significantly improves classification performance after using self-training, and outperforms state-of-the-art spectral-spatial methods for limited labeled training samples. Fan Li 0005, David A. Clausi, Linlin Xu, Alexander Wong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Cognitive-Empowered Femtocells: An Intelligent Paradigm for Femtocell NetworksabstractDeploying femtocells has been taken as an effective solution for removing coverage holes and improving wireless service performance in 3G‐beyond wireless networks such as WiMAX and Long Term Evolution (LTE). This article investigates a novel framework of dynamic spectrum management for femtocell networks, called cognitive‐empowered femtocells (CEF), aiming at mitigating both cross‐tier and intratier interferences with minimum modifications required on the corresponding macrocell network. With the proposed framework, each CEF base station (BS) and the femtocell users can utilize spatiotemporally available radio resources for the access traffic. We conclude that the proposed CEF framework can effectively complement the existing femtocell design and serve as a value‐added feature to the state‐of‐the‐art femtocell technologies, while achieving high scalability and interoperability by minimizing the required modifications on the macrocell protocol design. Xiao-Yu Wang 0010, Pin-Han Ho, Alexander Wong, Limei Peng |
Wirel. Commun. Mob. Comput. | 3 |
| 2017 | Unsupervised Domain Adaptation with a Relaxed Covariate Shift AssumptionabstractDomain adaptation addresses learning tasks where training is performed on data from one domain whereas testing is performed on data belonging to a different but related domain. Assumptions about the relationship between the source and target domains should lead to tractable solutions on the one hand, and be realistic on the other hand. Here we propose a generative domain adaptation model that allows for modelling different assumptions about this relationship, among which is a newly introduced assumption that replaces covariate shift with a possibly more realistic assumption without losing tractability due to the efficient variational inference procedure developed. In addition to the ability to model less restrictive relationships between source and target, modelling can be performed without any target labeled data (unsupervised domain adaptation). We also provide a Rademacher complexity bound of the proposed algorithm. We evaluate the model on the Amazon reviews and the CVC pedestrian detection datasets. Tameem Adel, Han Zhao 0002, Alexander Wong |
AAAI | 3 |
| 2017 | Comparing methods to estimate consumptive use with remote sensing in California's Sacramento-San Joaquin DeltaabstractThis study compares estimates of crop evapotranspiration (ET) from the Sacramento San Joaquin Delta (the Delta) for the 2014-2015 and 2015-2017 water years using seven estimation methods. In addition, direct field measurements of ET from bare soil and three crops (corn, alfalfa and pasture) were taken at several fields using eddy covariance and surface renewal stations. The project also included land use surveys for the 2015 and 2016 irrigation seasons to serve as an input for some estimation methods. Josué Medellín-Azuara, U. Kyaw Tha Paw, Yufang Jin, Eric Kent, Jenae' Clay, Alexander Wong, Michelle M. Leinfelder-Miles, Jay R. Lund |
IGARSS | 6 |
| 2017 | An Enhanced Probabilistic Posterior Sampling Approach for Synthesizing SAR Imagery With Sea Ice and Oil SpillsabstractAlthough the synthesis of the synthetic aperture radar (SAR) imagery with both sea ice and oil spills can significantly benefit in improving the consistency and comprehensiveness of testing and evaluating algorithms that are designed for mapping cold ocean regions, creating such imagery is difficult due to the heterogeneity and complexity of the source images. This letter presents an enhanced region-based probabilistic posterior sampling approach to effectively synthesize SAR imagery with different ocean features. In the proposed approach, instead of relying entirely on the SAR intensity values, the posterior sampling is performed based on a number of quantitative factors, such as intensity, label field, and the prior class probability of sampling candidates, constituting a complete probabilistic framework that addresses key aspects in the synthesis of SAR imagery from heterogeneous sources. The experiments demonstrate that the proposed approach can better address the difficulties caused by the heterogeneity in the source images compared with the existing state-of-the-art ice synthesis method, and it will improve the consistency, comprehensiveness, and fairness of the evaluation of the remote sensing classification and segmentation algorithms. Linlin Xu, Alexander Wong, David A. Clausi |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Weakly Supervised Classification of Remotely Sensed Imagery Using Label Constraint and Edge PenaltyabstractThe classification of pixels in remotely sensed imagery (RSI) into land cover classes typically requires knowing the labels of some image pixels for model training. However, accurate pixel-level label information is usually difficult and expensive to acquire, which restricts the applicability of supervised image classification methods. In contrast, the region labels information that specifies which classes are contained in a region of the image that is easier to acquire and less susceptible to identification errors. To utilize the region label information for remotely sensed image classification, this paper presents a weakly supervised image classification approach using label constraint and edge penalty (ILCEP), which has the following key characteristics. First, the predefined region labels are used as constraints in ILCEP to guide the inference of pixel labels in the image. Second, the edges between neighboring pixels are used as penalties to address the spatial contextual information in the image. Third, the label constraint and edge penalty are incorporated into the conditional random field framework, and simultaneous model learning and label inference are achieved by solving the maximum a posteriori problem through an enhanced simulated annealing algorithm. Experiments on both simulated and real RSIs demonstrate that the proposed approach can achieve high classification accuracy by knowing only the region-level label information. Linlin Xu, David A. Clausi, Fan Li 0005, Alexander Wong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2017 | A Novel Bayesian Spatial-Temporal Random Field Model Applied to Cloud Detection From Remotely Sensed ImageryabstractWith the fast advancement of remote sensing platforms and sensors, remotely sensed imagery (RSI) is increasingly being characterized by both high spatial resolution and high temporal resolution. How to efficiently use the rich spatial and temporal information in RSI for highly accurate object detection and classification is an important research question. Nevertheless, there is still a lack of a probabilistic framework that is capable of fully accounting for the spatial-temporal information in RSI for improved applications. In this paper, we present a Bayesian spatial-temporal random field model that constitutes a complete probabilistic framework for fully explaining the spatial-temporal correlation in RSI, leading to an enhanced object detection approach that is used for cloud detection from RSI. Under the Bayesian theorem, the posterior distribution of a label field is decomposed into the label prior, the data likelihood, the temporal label likelihood, and the temporal data likelihood. To address the difficulties in modeling the complex spatial-temporal correlation effect in the temporal data likelihood, a stochastic sampling approach is presented. Based on the maximum a posteriori approach, the posterior distribution is seamlessly integrated into the graph-cut optimization framework, and, therefore, the model optimization can be efficiently solved. The proposed algorithm is tested for cloud detection on both simulated and real RSIs and the results demonstrate that the proposed algorithm can effectively exploit the spatial-temporal information for achieving higher detection accuracy. Linlin Xu, Alexander Wong, David A. Clausi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Saliency-guided projection geometric correction using a projector-camera systemabstractProjecting an image onto an arbitrary non-flat screen surface leads to undesired geometric distortions in the image's projection. Geometric correction pre-distorts the image being projected such that the image's projection appears geometrically correct. In this work, we propose a novel saliency-guided projection geometric correction (SPGC) method that leverages calibration parameters along with 3D surface geometry captured by the projector-camera system to compensate for the geometric distortions created by non-flat screen surfaces. The proposed SPGC method incorporates a novel sampling scheme that selects a small set of surface points for geometric correction estimation based on local surface saliency, which greatly reduces the computational complexity of geometric correction estimation process. Experimental results using a test non-flat screen surface with abrupt edges and curve showed that the proposed SPGC approach achieved superior distortion compensation performance both quantitatively and qualitatively when compared to an unguided projection geometric correction method, while requiring just 3% of the samples used by a conventional densely-sampled projection geometric correction method. Ameneh Boroomand, Hicham Sekkati, Mark Lamm, David A. Clausi, Alexander Wong |
ICIP | 5 |
| 2016 | Random feature maps via a Layered Random Projection (LARP) framework for object classificationabstractThe approximation of nonlinear kernels via linear feature maps has recently gained interest due to their applications in reducing the training and testing time of kernel-based learning algorithms. Current random projection methods avoid the curse of dimensionality by embedding the nonlinear feature space into a low dimensional Euclidean space to create nonlinear kernels. We introduce a Layered Random Projection (LaRP) framework, where we model the linear kernels and nonlinearity separately for increased training efficiency. The proposed LaRP framework was assessed using the MNIST hand-written digits database and the COIL-100 object database, and showed notable improvement in object classification performance relative to other state-of-the-art random projection methods. Audrey G. Chung, Mohammad Javad Shafiee, Alexander Wong |
ICIP | 3 |
| 2016 | SAPPHIRE: Stochastically acquired photoplethysmogram for heart rate inference in realistic environmentsabstractA novel method, Stochastically Acquired Photoplethysmo-gram for Heart rate Inference in Realistic Environments (SAPPHIRE), is proposed for robust remote heart rate measurement through broadband video. A set of stochastically sampled points from the cheek region is tracked and used to construct corresponding time series observations via skin erythema transforms. From these observations, a photo-plethysmogram (PPG) waveform is estimated via Bayesian minimization, with the required posterior probability inferred using a Monte Carlo approach. To mitigate the effects of noise, the contribution of each observation is weighted based on the observation's likelihood to contain relevant data. A bandpass filter is applied to the estimated PPG waveform to omit implausible heart rate frequencies, and the heart rate is estimated through frequency domain analysis. Experimental results acquired from a set of thirty videos indicate significantly improved performance in comparison to state-of-the-art methods. Brendan Chwyl, Audrey G. Chung, Robert Amelard, Jason Deglint, David A. Clausi, Alexander Wong |
ICIP | 6 |
| 2016 | Spatio-temporal saliency detection using abstracted fully-connected graphical modelsabstractA novel approach to spatio-temporal saliency detection in video is proposed. Saliency computation is considered as an optimization problem that maximizes the energy of a fully-connected graphical model based on spatio-temporal feature distinctiveness. Each pixel in a video is modeled by a node, and the spatio-temporal feature distinctiveness between pixels by edges connecting the nodes in the graph. The computational complexity is addressed by compressing the fully-connected graph into an abstracted, fully-connected graph with far fewer nodes, where each node in the new graph characterizes nodal groups. The saliency value of each pixel is then computed based on spatio-temporal feature distinctiveness and the energy representation of its nodal group given the constructed graphical model. Experimental results show that our approach outperforms existing approaches to spatio-temporal salient region detection. Mohammad Javad Shafiee, Christian Scharfenberger, I. BenDaya, Shahid A. Haider, N. Talukdar, David A. Clausi, Alexander Wong |
ICIP | 8 |
| 2016 | Bag of Bags: Nested Multi Instance Classification for Prostate Cancer DetectionabstractComputer-aided detection (CAD) algorithms have been proposed for auto-detection of different types of cancer. CAD algorithms rely on machine learning methods to classify regions of interest in images into cancerous and healthy regions. In cancer screening, the foremost problem to solve is whether a patient has cancer, regardless of the location of cancerous regions in the organ. This allows early detection of the disease leading to a right course of action in terms of treatment to be taken. In machine learning, this problem has been formulated as multi-instance learning (MIL) where bags of instances are classified rather than the individual instances. In this paper, we propose a bag of bags (BoB) nested MIL algorithm where high-level bags (or parent bags), each contains multiple smaller bags of instances. We applied the proposed BoB MIL algorithm to prostate cancer detection problem using magnetic resonance imaging data to first detect which patients have cancer and consequently, to detect which slices in the 3D volume imaging data of the detected patients contain cancerous regions. Experimental results obtained from the imaging data of 30 patients with ground-truth data based on biopsy results show that the proposed algorithm is not only capable of detecting prostate cancer at patient level, it is also able to detect the cancerous regions at slice level of imaging data with high accuracy. Farzad Khalvati, Junjie Zhang 0005, Alexander Wong, Masoom A. Haider |
ICMLA | 3 |
| 2016 | Image segmentation via multi-scale stochastic regional texture appearance models
R. S. Medeiros, Jacob Scharcanski, Alexander Wong |
Comput. Vis. Image Underst. | 3 |
| 2016 | NeRD: A Neural Response Divergence Approach to Visual Saliency DetectionabstractIn this letter, a novel approach to visual saliency detection via neural response divergence (NeRD) is proposed, where synaptic portions of deep neural networks, previously trained for complex object recognition, are leveraged to compute low-level cues that can be used to compute image region distinctiveness. Based on this concept, an efficient visual saliency detection framework is proposed using deep convolutional StochasticNets. Experimental results using complex scene saliency dataset and MSRA10k natural image datasets show that the proposed NeRD approach can achieve improved performance when compared to state-of-the-art image saliency approaches, while attaining low computational complexity necessary for near-real-time computer vision applications. Mohammad Javad Shafiee, Parthipan Siva, Christian Scharfenberger, Paul W. Fieguth, Alexander Wong |
IEEE Signal Process. Lett. | 5 |
| 2016 | Intrinsic Representation of Hyperspectral Imagery for Unsupervised Feature ExtractionabstractUnsupervised feature extraction from hyperspectral images (HSIs) relies on efficient data representation. However, classical data representation techniques, e.g., principal component analysis and independent component analysis, do not reflect the intrinsic characteristics of HSI, and as such, they are less efficient for producing discriminative features. To address this issue, we have developed an intrinsic representation (IR) approach to support HSI classification. Based on the linear spectral mixture model, the IR approach explains the underlying physical factors that are responsible for generating HSI. Moreover, it addresses other important characteristics of HSI, i.e., the noise variance heterogeneity effect in the spectral domain and the spatial correlation effect in image domain. The IR model is solved iteratively by alternating the estimation of IR coefficients given IR bases and the update of IR bases given the coefficients. The resulting IR coefficients are discriminative, compact, and noise resistant, thereby constituting powerful features for improved HSI classification. The experiments on both simulated and real HSI demonstrate that the features extracted by the IR model are more capable of boosting the classification performance than the other referenced techniques. Linlin Xu, Alexander Wong, Fan Li 0005, David A. Clausi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Noise-Compensated, Bias-Corrected Diffusion Weighted Endorectal Magnetic Resonance Imaging via a Stochastically Fully-Connected Joint Conditional Random Field ModelabstractDiffusion weighted magnetic resonance imaging (DW-MR) is a powerful tool in imaging-based prostate cancer screening and detection. Endorectal coils are commonly used in DW-MR imaging to improve the signal-to-noise ratio (SNR) of the acquisition, at the expense of significant intensity inhomogeneities (bias field) that worsens as we move away from the endorectal coil. The presence of bias field can have a significant negative impact on the accuracy of different image analysis tasks, as well as prostate tumor localization, thus leading to increased inter- and intra-observer variability. Retrospective bias correction approaches are introduced as a more efficient way of bias correction compared to the prospective methods such that they correct for both of the scanner and anatomy-related bias fields in MR imaging. Previously proposed retrospective bias field correction methods suffer from undesired noise amplification that can reduce the quality of bias-corrected DW-MR image. Here, we propose a unified data reconstruction approach that enables joint compensation of bias field as well as data noise in DW-MR imaging. The proposed noise-compensated, bias-corrected (NCBC) data reconstruction method takes advantage of a novel stochastically fully connected joint conditional random field (SFC-JCRF) model to mitigate the effects of data noise and bias field in the reconstructed MR data. The proposed NCBC reconstruction method was tested on synthetic DW-MR data, physical DW-phantom as well as real DW-MR data all acquired using endorectal MR coil. Both qualitative and quantitative analysis illustrated that the proposed NCBC method can achieve improved image quality when compared to other tested bias correction methods. As such, the proposed NCBC method may have potential as a useful retrospective approach for improving the consistency of image interpretations. Ameneh Boroomand, Mohammad Javad Shafiee, Farzad Khalvati, Masoom A. Haider, Alexander Wong |
IEEE Trans. Medical Imaging | 5 |
| 2015 | A Probabilistic Covariate Shift Assumption for Domain AdaptationabstractThe aim of domain adaptation algorithms is to establish a learner, trained on labeled data from a source domain, that can classify samples from a target domain, in which few or no labeled data are available for training. Covariate shift, a primary assumption in several works on domain adaptation, assumes that the labeling functions of source and target domains are identical. We present a domain adaptation algorithm that assumes a relaxed version of covariate shift where the assumption that the labeling functions of the source and target domains are identical holds with a certain probability. Assuming a source deterministic large margin binary classifier, the farther a target instance is from the source decision boundary, the higher the probability that covariate shift holds. In this context, given a target unlabeled sample and no target labeled data, we develop a domain adaptation algorithm that bases its labeling decisions both on the source learner and on the similarities between the target unlabeled instances. The source labeling function decisions associated with probabilistic covariate shift, along with the target similarities are concurrently expressed on a similarity graph. We evaluate our proposed algorithm on a benchmark sentiment analysis (and domain adaptation) dataset, where state-of-the-art adaptation results are achieved. We also derive a lower bound on the performance of the algorithm. Tameem Adel, Alexander Wong |
AAAI | 2 |
| 2015 | Tiger: A texture-illumination guided energy response model for illumination robust local saliencyabstractLocal saliency models are a cornerstone in image processing and computer vision, used in a wide variety of applications ranging from keypoint detection and feature extraction, to image matching and image representation. However, current models exhibit difficulties in achieving consistent results under varying, non-ideal illumination conditions. In this paper, a novel texture-illumination guided energy response (TIGER) model for illumination robust local saliency is proposed. In the TIGER model, local saliency is quantified by a modified Hessian energy response guided by a weighted aggregate of texture and illumination aspects from the image. A stochastic Bayesian disassociation approach via Monte Carlo sampling is employed to decompose the image into its texture and illumination aspects for the saliency computation. Experimental results demonstrate that higher correlation between local saliency maps constructed from the same scene under different illumination conditions can be achieved using the TIGER model when compared to common local saliency approach, i.e., Laplacian of Gaussian, Difference of Gaussians, and Hessian saliency models. Brendan Chwyl, Audrey G. Chung, F. Y. Li, Alexander Wong, David A. Clausi |
ICIP | 4 |
| 2015 | High dynamic range map estimation via fully connected random fields with stochastic cliquesabstractThe reconstruction of high dynamic range (HDR) images via conventional camera systems and low dynamic range (LDR) images is a growing field of research in image acquisition. The radiance map associated with the HDR image of a scene is typically computed using multiple images of the same scene captured at different exposures (i.e., bracketed LDR imzages). This approach, though inexpensive, is sensitive to noise under high camera ISO. Each bracketed image is associated with a different level of noise due to the change in exposure time, and the noise is further amplified when tone-mapping the HDR image for display. A new framework is proposed to address the associated noise in the context of random fields. The estimation of the HDR image from a set of LDR images is formulated as a stochastically fully connected conditional random field where the spatial information is incorporated to compute the HDR value in combination with the LDR image values. Experimental results show that the proposed framework compensated the non-stationary ISO noise while preserving the boundaries in the estimated HDR images. F. Y. Li, Mohammad Javad Shafiee, Audrey G. Chung, Brendan Chwyl, Farnoud Kazemzadeh, Alexander Wong, John S. Zelek |
ICIP | 6 |
| 2015 | DESIRe: Discontinuous energy seam carving for image retargeting via structural and textural energy functionalsabstractThis paper proposes DESIRe (Discontinuous Energy Seam-carving Image Retargeting), an improved seam carving approach for content-aware image retargeting. The proposed algorithm introduces a novel discontinuous seam carving optimization process that incorporates not only a structural energy cost functional, but also a texture energy cost functional to handle scenes with complex structural and textural characteristics. Experimental results using a variety of scenes with dense image detail characteristics show that the proposed DESIRe seam carving approach can provide improved visual quality when compared to a number of existing seam carving methods. These results illustrate the potential of the proposed DESIRe approach for improving content-aware image retargeting that reduces the loss of important image information without causing significant visual distortions or artifacts. Akshaya Kumar Mishra, Christian Scharfenberger, Parthipan Siva, Fan Li 0005, Alexander Wong, David A. Clausi |
ICIP | 5 |
| 2015 | DESIRe: Discontinuous energy seam carving for image retargeting via structural and textural energy functionalsabstractThis paper proposes DESIRe (Discontinuous Energy Seam-carving Image Retargeting), an improved seam carving approach for content-aware image retargeting. The proposed algorithm introduces a novel discontinuous seam carving optimization process that incorporates not only a structural energy cost functional, but also a texture energy cost functional to handle scenes with complex structural and textural characteristics. Experimental results using a variety of scenes with dense image detail characteristics show that the proposed DESIRe seam carving approach can provide improved visual quality when compared to a number of existing seam carving methods. These results illustrate the potential of the proposed DESIRe approach for improving content-aware image retargeting that reduces the loss of important image information without causing significant visual distortions or artifacts. Akshaya Kumar Mishra, Christian Scharfenberger, Parthipan Siva, Fan Li 0005, Alexander Wong, David A. Clausi |
ICIP | 5 |
| 2015 | Improved fine structure modeling via guided stochastic clique formation in fully connected conditional random fieldsabstractMarkov random fields (MRFs) and conditional random fields (CRFs) are influential tools in image modeling, particularly for applications such as image segmentation. Local MRFs and CRFs utilize local nodal interactions when modeling, leading to excessive smoothness on boundaries (i.e., the short-boundary bias problem). Recently, the concept of fully connected conditional random fields with stochastic cliques (SFCRF) was proposed to enable long-range nodal interactions while addressing the computational complexity associated with fully connected random fields. While SFCRF was shown to provide significant improvements in segmentation accuracy, there were still limitations with the preservation of fine structure boundaries. To address these limitations, we propose a new approach to stochastic clique formation for fully connected random fields (G-SFCRF) that is guided by the structural characteristics of different nodes within the random field. In particular, fine structures surrounding a node are modeled statistically by probability distributions, and stochastic cliques are formed by considering the statistical similarities between nodes within the random fields. Experimental results show that G-SFCRF outperforms existing fully connected CRF frameworks, SFCRF, and the principled deep random field framework for image segmentation. Mohammad Javad Shafiee, Audrey G. Chung, Alexander Wong, Paul W. Fieguth |
ICIP | 3 |
| 2015 | Return of grid seams: A superpixel algorithm using discontinuous multi-functional energy seam carvingabstractSuperpixels have been widely used for compact image representation in a large number of computer vision applications such as object recognition, segmentation, and depth estimation. Recently, a novel seam carving approach for superpixel generation called Grid Seams was introduced that significantly improved image structure preservation while maintaining global spatial constraints. While Grid Seams was able to achieve state-of-the-art superpixel accuracy, it only took advantage of rudimentary edge information in its energy function and as such did not account for other important image characteristics such as texture. Motivated by this, we present Return of Grid Seams (RGS), a novel extension of Grid Seams that takes advantage of not only structural variations, but also textural variations (in the form of texture distinctiveness) into a unified multi-functional energy along with global spatial constraints. Furthermore, RGS incorporates discontinuous seams in the optimization process to allow for greater flexibility in preserving structural information. Experimental results using the Berkeley Segmentation Dataset show that RGS is able to outperform Grid Seams as well as a number of other seam carving based superpixel methods. Parthipan Siva, Christian Scharfenberger, Ibrahim Ben Daya, Akshaya Kumar Mishra, Alexander Wong |
ICIP | 5 |
| 2015 | PIRM: Fast background subtraction under sudden, local illumination changes via probabilistic illumination range modellingabstractWe present an illumination-compensation method to enable fast and reliable background subtraction under sudden, local illumination changes in wide area surveillance videos. We use Probabilistic Illumination Range Modeling (PIRM) to model the conditional probability distribution of current frame intensity given background intensity. With this model, we can identify a continuous range of current frame intensities that map to the same background intensity, and scale all pixels within that range in the current frame appropriately to enable illumination-compensated background subtraction. Experimental results using a standard academic dataset as well as very challenging industry videos show that PIRM can achieve improvements in compensating for sudden, local illumination changes. Parthipan Siva, Mohammad Javad Shafiee, Francis Li, Alexander Wong |
ICIP | 4 |
| 2015 | Optimized sampling distribution based on nonparametric learning for improved compressive sensing performance
Shimon Schwartz, Alexander Wong, David A. Clausi |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | QMCTLS: Quasi Monte Carlo Texture Likelihood Sampling for Despeckling of Complex Polarimetric SAR ImagesabstractDespeckling of complex polarimetric synthetic aperture radar (SAR) images is more difficult than denoising of general images due to the low signal-to-noise ratio and the complex signals. A novel stochastic polarimetric SAR despeckling technique based on quasi Monte Carlo sampling (QMCS) and region-based probabilistic similarity likelihood has been developed. The despeckling of complex polarimetric SAR images is formulated as a Bayesian least squares optimization problem, where the posterior distribution is estimated by QMCS in a nonparametric manner. The QMCS approach allows the incorporation of the statistical description of local texture pattern similarity. Experiments on two benchmark quad-pol SAR images demonstrate that the proposed QMC texture likelihood sampling (QMCTLS) filter outperforms referenced methods in terms of both noise removal and detail preservation. Fan Li 0005, Linlin Xu, Alexander Wong, David A. Clausi |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2015 | Feature Extraction for Hyperspectral Imagery via Ensemble Localized Manifold LearningabstractA feature extraction approach for hyperspectral image classification has been developed. Multiple linear manifolds are learned to characterize the original data based on their locations in the feature space, and an ensemble of classifier is then trained using all these manifolds. Such manifolds are localized in the feature space (which we will refer to as “localized manifolds”) and can overcome the difficulty of learning a single global manifold due to the complexity and nonlinearity of hyperspectral data. Two state-of-the-art feature extraction methods are used to implement localized manifolds. Experimental results show that classification accuracy is improved using both localized manifold learning methods on standard hyperspectral data sets. Fan Li 0005, Linlin Xu, Alexander Wong, David A. Clausi |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2015 | Structure-Guided Statistical Textural Distinctiveness for Salient Region Detection in Natural ImagesabstractWe propose a simple yet effective structure-guided statistical textural distinctiveness approach to salient region detection. Our method uses a multilayer approach to analyze the structural and textural characteristics of natural images as important features for salient region detection from a scale point of view. To represent the structural characteristics, we abstract the image using structured image elements and extract rotational-invariant neighborhood-based textural representations to characterize each element by an individual texture pattern. We then learn a set of representative texture atoms for sparse texture modeling and construct a statistical textural distinctiveness matrix to determine the distinctiveness between all representative texture atom pairs in each layer. Finally, we determine saliency maps for each layer based on the occurrence probability of the texture atoms and their respective statistical textural distinctiveness and fuse them to compute a final saliency map. Experimental results using four public data sets and a variety of performance evaluation metrics show that our approach provides promising results when compared with existing salient region detection approaches. Christian Scharfenberger, Alexander Wong, David A. Clausi |
IEEE Trans. Image Process. | 2 |
| 2015 | Apparent Ultra-High b-Value Diffusion-Weighted Image Reconstruction via Hidden Conditional Random FieldsabstractA promising, recently explored, alternative to ultra-high b-value diffusion weighted imaging (UHB-DWI) is apparent ultra-high b-value diffusion-weighted image reconstruction (AUHB-DWR), where a computational model is used to assist in the reconstruction of apparent DW images at ultra-high b -values. Firstly, we present a novel approach to AUHB-DWR that aims to improve image quality. We formulate the reconstruction of an apparent DW image as a hidden conditional random field (HCRF) in which tissue model diffusion parameters act as hidden states in this random field. The second contribution of this paper is a new generation of fully connected conditional random fields, called the hidden stochastically fully connected conditional random fields (HSFCRF) that allows for efficient inference with significantly reduced computational complexity via stochastic clique structures. The proposed AUHB-DWR algorithms, HCRF and HSFCRF, are evaluated quantitatively in nine different patient cases using Fisher's criteria, probability of error, and coefficient of variation metrics to validate its effectiveness for the purpose of improving intensity delineation between expert identified suspected cancerous and healthy tissue within the prostate gland. The proposed methods are also examined using a prostate phantom, where the apparent ultra-high b-value DW images reconstructed using the tested AUHB-DWR methods are compared with real captured UHB-DWI. The results illustrate that the proposed AUHB-DWR methods has improved reconstruction quality and improved intensity delineation compared with existing AUHB-DWR approaches. Mohammad Javad Shafiee, Shahid A. Haider, Alexander Wong, Dorothy Lui, Andrew Cameron, Amen Modhafar, Paul W. Fieguth, Masoom A. Haider |
IEEE Trans. Medical Imaging | 3 |
| 2014 | Image saliency detection via multi-scale statistical non-redundancy modelingabstractA multi-scale statistical non-redundancy modeling approach is introduced for saliency detection in images. The statistical non-redundancy of pixels at different wavelet sub-bands is characterized using a multi-dimensional lattice of non-parametric statistical models, thus taking into account image saliency at multiple scales. This identifies saliency in image attributes at multiple scales, and makes saliency detection strongly robust against noisy input images. Results based on images from a public database show that the proposed approach outperforms existing single and multi-scale approaches, particularly when dealing with noisy images. Christian Scharfenberger, Aanchal Jain, Alexander Wong, Paul W. Fieguth |
ICIP | 3 |
| 2014 | Efficient Bayesian inference using fully connected conditional random fields with stochastic cliquesabstractConditional random fields (CRFs) are one of the most powerful frameworks in image modeling. However practical CRFs typically have edges only between nearby nodes; using more interactions and expressive relations among nodes make these methods impractical for large-scale applications, due to the high computational complexity. Recent work has shown that fully connected CRFs can be tractable by defining specific potential functions. In this paper, we present a novel framework to tackle the computational complexity of a fully connected graph without requiring specific potential functions. Instead, inspired by random graph theory and sampling methods, we propose a new clique structure called stochastic cliques. The stochastically fully connected CRF (SFCRF) is a marriage between random graphs and random fields, benefiting from the advantages of fully connected graphs while maintaining computational tractability. The effectiveness of SFCRF was examined by binary image labeling of highly noisy images. The results show that the proposed framework outperforms an adjacency CRF and a CRF with a large neighborhood size. Mohammad Javad Shafiee, Alexander Wong, Parthipan Siva, Paul W. Fieguth |
ICIP | 2 |
| 2014 | Comparison of unsupervised segmentation methods for surficial materials mapping in Nunavut, Canada using RADARSAT-2 polarimetric, Landsat-7, and DEM dataabstractIn this paper, unsupervised segmentation methods are investigated for surficial materials mapping in Nunavut, Canada. Different satellite data sources including RADARSAT-2 polarimetric image, LANDSAT-7 image, and DEM data are combined and three unsupervised segmentation methods are compared. Results show that IRGS has better performance than the other two methods. Fan Li 0005, Alexander Wong, David A. Clausi |
IGARSS | 2 |
| 2014 | Comparative study of feature space projection methods for hyperspectral image classificationabstractFeature space projection, or feature projection is an active research topic in machine learning. Some projection methods have been used in remote sensing for dimension reduction, especially for hyperspectral data due to high dimensionality. Projection methods can improve the performance of classifiers susceptible to the Hughes phenomenon. However, the effect of feature projection for more advanced classifiers has not been well-studied, and there are few studies comparing projection methods for hyperspectral image classification. A comprehensive study has been performed on the effect of feature projection for classification using both reduced and full dimensions. The performance of six feature projection methods (PCA, LLE, LDA, LFDA, LMNN, and SPCA) using three classifiers has been explored on three hyperspectral data sets. Results show that the performance of feature projection methods on different classifiers are mainly consistent for different data sets. LFDA achieves the best overall performance considering all data sets and all classifiers. Fan Li 0005, Alexander Wong, David A. Clausi |
IGARSS | 2 |
| 2014 | Combining rotation forests and adaboost for hyperspectral imagery classification using few labeled samplesabstractClassification of hyperspectral imagery using too few labeled samples is a challenging problem considering the high dimensionality of hyperspectral imagery. In this paper, an ensemble method combining rotation forests and AdaBoost is proposed to tackle this problem. By adaptive boosting, AdaBoost can significantly reduce classification error in an iteration compared to a single classifier, and the rotation matrix can increase diversity so that the ensemble performance can be further improved. Experimental resutls show that the final classification accuracy of the proposed algorithm consistently outperforms other state-of-the-art classification methods. Fan Li 0005, Alexander Wong, David A. Clausi |
IGARSS | 2 |
| 2014 | URC: Unsupervised regional clustering of remote sensing imageryabstractConditional random fields (CRF) have been used extensively for spatially coherent segmentation and classification of images. As a result, may techniques for finding the optimal inference of CRFs have been developed. However, CRFs have seen little use in clustering and segmenting of satellite imagery due to the large number of pixels in satellite images. In this paper we present a means of defining remote sensing imagery as a region based conditional random field (CRF). Unlike the pixel based CRF, our region based CRF can be optimally solved using the latest advancements in CRF inference techniques because the region based CRF is a computationally simpler problem to tackle. To reduce the loss of information when forming the region based CRF, the regions are defined using a superpixel algorithm which decreases the spatial resolution while preserving the original satellite image's structure. Furthermore, unlike previous approaches, we show that the optimal number of clusters in the satellite images can be automatically determined by formulating the number of clusters as part of the CRF cost function. This unsupervised region based approach is a non-parametric formulation which is validated using different imaging modalities: SAR and hyper-spectral imaging. Parthipan Siva, Alexander Wong |
IGARSS | 2 |
| 2014 | Hybrid structural and texture distinctiveness vector field convolution for region segmentation
Khalil Fergani, Dorothy Lui, Christian Scharfenberger, Alexander Wong, David A. Clausi |
Comput. Vis. Image Underst. | 4 |
| 2014 | K-P-Means: A Clustering Algorithm of K "Purified" Means for Hyperspectral Endmember EstimationabstractThis letter presents K-P-Means, a novel approach for hyperspectral endmember estimation. Spectral unmixing is formulated as a clustering problem, with the goal of K-P-Means to obtain a set of “purified” hyperspectral pixels to estimate endmembers. The K-P-Means algorithm alternates iteratively between two main steps (abundance estimation and endmember update) until convergence to yield final endmember estimates. Experiments using both simulated and real hyperspectral images show that the proposed K-P-Means method provides strong endmember and abundance estimation results compared with existing approaches. Linlin Xu, Jonathan Li 0001, Alexander Wong, Junhuan Peng |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2014 | Despeckling of Synthetic Aperture Radar Images Using Monte Carlo Texture Likelihood SamplingabstractSpeckle noise is found in synthetic aperture radar (SAR) images and can affect visualization and analysis. A novel stochastic texture-based algorithm is proposed to suppress speckle noise while preserving the underlying structural and texture detail. Based on a sorted local texture model and a Fisher-Tippett logarithmic-space speckle distribution model, a Monte Carlo texture likelihood sampling strategy is proposed to estimate the true signal. The algorithm is compared to six other classic and state-of-the-art despeckling techniques. The comparison is performed both on synthetic noisy images added and on actual SAR images. Using peak signal-to-noise ratio, contrast-to-noise ratio, and structural similarity index as image quality metrics, the proposed algorithm shows strong despeckling performance when compared to existing despeckling algorithms. Jeffrey Glaister, Alexander Wong, David A. Clausi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | Enhanced Decoupled Active Contour Using Structural and Textural Variation Energy FunctionalsabstractActive contours are a popular approach for object segmentation that uses an energy minimizing spline to extract an object's boundary. Nonparametric approaches can be computationally complex, whereas parametric approaches can be impacted by parameter sensitivity. A decoupled active contour (DAC) overcomes these problems by decoupling the external and internal energies and optimizing them separately. However a drawback of this approach is its reliance on the edge gradient as the external energy. This can lead to poor convergence toward the object boundary in the presence of weak object and strong background edges. To overcome these issues with convergence, a novel approach is proposed that takes advantage of a sparse texture model, which explicitly considers texture for boundary detection. The approach then defines the external energy as a weighted combination of textural and structural variation maps and feeds it into a multifunctional hidden Markov model for more robust object boundary detection. The enhanced DAC (EDAC) is qualitatively and visually analyzed on two natural image data sets as well as Brodatz images. The results demonstrate that EDAC effectively combines texture and structural information to extract the object boundary without impact on computation time and a reliance on color. Dorothy Lui, Christian Scharfenberger, Khalil Fergani, Alexander Wong, David A. Clausi |
IEEE Trans. Image Process. | 4 |
| 2013 | Statistical Textural Distinctiveness for Salient Region Detection in Natural ImagesabstractA novel statistical textural distinctiveness approach for robustly detecting salient regions in natural images is proposed. Rotational-invariant neighborhood-based textural representations are extracted and used to learn a set of representative texture atoms for defining a sparse texture model for the image. Based on the learnt sparse texture model, a weighted graphical model is constructed to characterize the statistical textural distinctiveness between all representative texture atom pairs. Finally, the saliency of each pixel in the image is computed based on the probability of occurrence of the representative texture atoms, their respective statistical textural distinctiveness based on the constructed graphical model, and general visual attentive constraints. Experimental results using a public natural image dataset and a variety of performance evaluation metrics show that the proposed approach provides interesting and promising results when compared to existing saliency detection methods. Christian Scharfenberger, Alexander Wong, Khalil Fergani, John S. Zelek, David A. Clausi |
CVPR | 2 |
| 2013 | Natural scene segmentation based on a stochastic texture region merging approachabstractThis paper presents an approach for segmenting natural scenes based on the underlying texture characteristics using a stochastic region merging strategy. Texture region models are constructed from patch-based stochastic texture features using a texton dictionary learning approach. Finally, a stochastic region merging strategy performs the image segmentation based on texture region likelihood. Compared with other state-of-the-art texture segmentation methods, our experimental results suggest that our approach potentially can handle better highly textured regions commonly found in natural scenes, and also can be more robust to color and illumination variations. R. S. Medeiros, Jacob Scharcanski, Alexander Wong |
ICASSP | 3 |
| 2013 | Unsupervised classification of agricultural land cover using polarimetric synthetic aperture radar via a sparse texture dictionary modelabstractA sparse texture dictionary learning method for unsupervised land cover classification is presented. The method takes the stance that land cover in remote sensing data is best analysed in texture patches rather than localized pixels. To this end, a feature vector is designed that describes local texture information in a spatially coherent manner. This texture model is extracted for each pixel in the scene. A sparse dictionary of global texture models is then learned to characterize the underlying texture distribution of the scene in a simplified manner. An unsupervised classifier is learned using these global texture models for grouping pixels exhibiting high similarity. Being an unsupervised classifier, the class labels that are learned are unbiased toward human interpretation of the scene, and rather are learned according to the texture information. The method is validated using polarimetric SAR data over a Flevoland, Netherlands agriculture scene, but may be generalized to any remote sensing data. Promising experimental results show how the proposed method retains the spatial coherence of crops, and attains higher accuracy than recent unsupervised and supervised classification methods using the same data. Robert Amelard, Alexander Wong, David A. Clausi |
IGARSS | 2 |
| 2013 | Unsupervised classification of sea-ice using synthetic aperture radar via an adaptive texture sparsifying transformabstractA texture sparsifying transform for use in unsupervised classification of sea-ice in polarimetric synthetic aperture radar (SAR) imagery is presented. The goal of the sparsifying transform is to compactly represent the underlying information of the SAR imagery to eliminate sources of unwanted noise and complexities (e.g., banding effect on RADARSAT-2) commonly found in SAR imagery. The proposed algorithm is designed to be simple to implement and discriminative in sea-ice scenes. Performing unsupervised classification on the sparsifying transform space using scenes captured with C-band HV polarization yields experimental results that are much more accurate than common pixel-based methods, and performs comparably to a recent more complex method. Robert Amelard, Alexander Wong, Fan Li 0005, David A. Clausi |
IGARSS | 2 |
| 2013 | Continuous sea ice thickness estimation using a joint MODIS and AMSR-E guided variational modelabstractEstimates of sea ice thickness are important for shipping and weather forecasting applications. Sea ice thickness can be estimated using data from the thermal channels on the Moderate Resolution Imaging Spectroradiometer (MODIS). However, using this data for studies of surface conditions is significantly hampered by cloud cover. This is particularly problematic for studies of the marginal ice zone, where atmospheric conditions often lead to persistent cloudy conditions. In this study a new method is proposed in which data from a passive microwave sensor is used to guide the estimation of surface temperature in cloud-covered regions. The impact of the method is verified by checking sea ice thickness values calculated using the guided surface temperature against values from operational sea ice charts. Alexander Wong, Katharine Andrea Scott, Edward Li, Robert Amelard |
IGARSS | 1 |
| 2013 | Saliency-guided compressive sensing approach to efficient laser range measurement
Shimon Schwartz, Alexander Wong, David A. Clausi |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | ETVOS: An Enhanced Total Variation Optimization Segmentation Approach for SAR Sea-Ice Image SegmentationabstractThis paper presents a novel enhanced total variation optimization segmentation (ETVOS) approach consisting of two phases to segmentation of various sea-ice types. In the total variation optimization phase, the Rudin-Osher-Fatemi total variation model was modified and implemented iteratively to estimate the piecewise constant state from a nonpiecewise constant state (the original noisy imagery) by minimizing the total variation constraints. In the finite mixture model classification phase, based on the pixel distribution, an expectation maximization method was performed to estimate the final class likelihood using a Gaussian mixture model. Then, a maximum likelihood classification technique was utilized to estimate the final class of each pixel that appeared in the product of the total variation optimization phase. The proposed method was tested on a synthetic image and various subsets of RADARSAT-2 imagery, and the results were compared with other well-established approaches. With the advantage of a short processing time, the visual inspection and quantitative analysis of segmentation results confirm the superiority of the proposed ETVOS method over other existing methods. Tae J. Kwon, Jonathan Li 0001, Alexander Wong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2013 | A Decoupled Approach to Illumination-Robust Optical Flow EstimationabstractDespite continuous improvements in optical flow in the last three decades, the ability for optical flow algorithms to handle illumination variation is still an unsolved challenge. To improve the ability to interpret apparent object motion in video containing illumination variation, an illumination-robust optical flow method is designed. This method decouples brightness into reflectance and illumination components using a stochastic technique; reflectance is given higher weight to ensure robustness against illumination, which is suppressed. Illumination experiments using the Middlebury and University of Oulu databases demonstrate the decoupled method's improvement when compared with state-of-the-art. In addition, a novel technique is implemented to visualize optical flow output, which is especially useful to compare different optical flow methods in the absence of the ground truth. Frederick Tung, Alexander Wong, David A. Clausi |
IEEE Trans. Image Process. | 3 |
| 2012 | Saliency detection via statistical non-redundancyabstractA novel algorithm based on statistical non-redundancy is proposed for saliency detection in natural images. By modeling site neighbourhoods as realizations of other site neighbourhoods under a Gaussian process, the saliency of any arbitrary site can be characterized by the statistical non-redundancy of its site neighbourhood with respect to other site neighbourhoods in a given image. Preliminary results using natural images show that the proposed method provides improved precision vs. recall characteristics over previous methods such as spectral residuals and spectral whitening. Aanchal Jain, Alexander Wong, Paul W. Fieguth |
ICIP | 2 |
| 2012 | Multi-scale tensor vector field active contourabstractTwo major challenges faced in active contours are poor capture range and high sensitivity towards noise. Recently, the concept of tensor vector convolution (TVF) was introduced and shown to be promising in handling these challenges. However, in the presence of high noise levels, TVF may have difficulty in converging to the desired object boundary, particularly if the distance is great between the initial contour and the object boundary. To tackle this challenge, the concept of a multi-scale tensor vector field (MTVF) active contour is introduced to further reduce noise sensitivity. Comparing the performance of MTVF with multi-scale gradient vector field and multi-scale vector field convolution demonstrates that MTVF is more resilient to high noise levels as well as significantly reducing computation time. Alexander Wong, Paul W. Fieguth, David A. Clausi |
ICIP | 2 |
| 2012 | Monte Carlo cluster refinement for noise robust image segmentation
Alexander Wong, Xiao-Yu Wang 0010 |
J. Vis. Commun. Image Represent. | 1 |
| 2012 | A Bayesian Theoretic Approach to Multiscale Complex-Phase-Order RepresentationsabstractThis paper explores a Bayesian theoretic approach to constructing multiscale complex-phase-order representations. We formulate the construction of complex-phase-order representations at different structural scales based on the scale-space theory. Linear and nonlinear deterministic approaches are explored, and a Bayesian theoretic approach is introduced for constructing representations in such a way that strong structure localization and noise resilience are achieved. Experiments illustrate its potential for constructing robust multiscale complex-phase-order representations with well-localized structures across all scales under high-noise situations. Illustrative examples of applications of the proposed approach is presented in the form of multimodal image registration and feature extraction. Alexander Wong |
IEEE Trans. Image Process. | 1 |
| 2011 | Tensor vector field based active contoursabstractAmong the main limitations of active contours are their high noise sensitivity and poor capture range from the target object. One of the most promising approaches for addressing these limitations is the concept of Vector Field Convolution (VFC). However, due to its isotropic vector field kernel, VFC does not take full advantage of the underlying image structural characteristics. By specifically addressing this idea, a novel local tensor vector field approach is developed to adaptively account for these structural characteristics. Experimental results demonstrate that the proposed adaptive method leads to more accurate segmentation. Alexander Wong, Akshaya Kumar Mishra, David A. Clausi, Paul W. Fieguth |
ICIP | 2 |
| 2011 | A structure-guided conditional sampling model for video resolution enhancementabstractIn this paper, a novel approach is introduced to video resolution enhancement based on a structure-guided conditional sampling model. The proposed conditional sampling model is composed of hybrid models to handle strongly structured and weakly structured regions separately. To accomplish this, a binary hidden field indicating image structure is introduced into the conditional reconstruction process for estimating the high resolution video frames from the low resolution video frames. Preliminary results show that the proposed approach has the potential to produce high resolution video frames with improved image quality when compared to existing structure-guided video resolution enhancement approaches. Ying Liu 0010, Alexander Wong, Paul W. Fieguth |
ICIP | 2 |
| 2011 | Multi-scale 3D representation via volumetric quasi-random scale spaceabstractA novel nonlinear volumetric scale-space framework is proposed for multi-scale volumetric data representation. The problem is formulated as a Bayesian least-squares estimator, and a quasi-random density estimation approach is introduced for estimating the posterior distribution between consecutive volumetric scale space realizations. Experimental results using both synthetic and real MR volumetric data demonstrate the effectiveness of the proposed scale-space framework for three-dimensional representation with significantly better structural separation and localization across all scales when compared to existing volumetric scale-space frameworks such as volumetric anisotropic diffusion and volumetric linear Gaussian scale-space, especially under scenarios with high noise levels. Akshaya Kumar Mishra, Alexander Wong, Paul W. Fieguth, David A. Clausi |
ICIP | 2 |
| 2011 | Shot Boundary Detection Using Genetic Algorithm OptimizationabstractThis paper presents a novel method for shot boundary detection via an optimization of traditional scoring based metrics using a genetic algorithm search heuristic. The advantage of this approach is that it allows for the detection of shots without requiring the direct use of thresholds. The methodology is described using the edge-change ratio metric and applied to several test video segments from the TREC 2002 video track and contemporary television shows. The shot boundary detection results are evaluated using recall, precision and F1 metrics, which demonstrate that the proposed approach provides superior overall performance when compared to the effective edge-change ratio method. In addition, the convergence of the genetic algorithm is examined to show that the proposed method is both efficient and stable. Calvin Chan, Alexander Wong |
ISM | 2 |
| 2011 | Hybrid Video Compression Using Selective Keyframe Identification and Patch-Based Super-ResolutionabstractThis paper details a novel video compression pipeline using selective key frame identification to encode video and patch-based super-resolution to decode for playback. Selective key frame identification uses shot boundary detection and frame differencing methods to identify representative frames which are subsequently kept in high resolution within the compressed container. All other non-key frames are downscaled for compression purposes. Patch-based super-resolution finds similar patches between an up scaled non-key frame and the associated, high-resolution key frame to regain lost detail via a super-resolution process. The algorithm was integrated into the H.264 video compression pipeline tested on web cam, cartoon and live-action video for both streaming and storage purposes. Experimental results show that the proposed hybrid video compression pipeline successfully achieved higher compression ratios than standard H.264, while achieving superior video quality than low resolution H.264 at similar compression ratios. Jeffrey Glaister, Calvin Chan, Michael Frankovich, Alexander Wong |
ISM | 5 |
| 2011 | Comprehensive Analysis on the Effects of Noise Estimation Strategies on Image Noise Artifact Suppression PerformanceabstractIn this paper, the effects of employing different noise estimation strategies on the performance of noise artifact suppression techniques in achieving high image quality has been investigated. Most literature on the subject tends to use the true noise level of the noisy image when performing noise artifact suppression. However, this approach does not reflect how such techniques would be used in practical situations where the true noise level is unknown, which is common in most image and video processing applications. Therefore, in practical situations, the noise level must first be estimated before a noise artifact suppression technique can be applied using the estimated noise level. Through a comprehensive analysis of different noise estimation strategies using empirical testing on a variety of images with different characteristics, the MAD wavelet noise estimation technique was found to be the overall preferred noise estimation technique for all popular noise artifact suppression techniques investigated (BM3D, bilateral, Neigh Shrink, BLS-GSM and non-local means). Furthermore, the BM3D noise artifact suppression technique, combined with the MAD wavelet noise estimation technique, was found to offer the best performance in achieving high image quality in situations where the noise level is unknown and must be estimated. The outcome of this research is clear recommendations that can be used in practise when suppressing noise artifacts exhibited in digital imagery and video. Angus Leigh, Alexander Wong, David A. Clausi, Paul W. Fieguth |
ISM | 2 |
| 2011 | A Novel Hierarchical Model-Based Frame Rate Up-Conversion via Spatio-temporal Conditional Random FieldsabstractIn this paper, a hierarchical model-based approach to frame rate-up conversion is presented. Given a sequence of consecutive video frames, a Spatio-Temporal Conditional Random Field (ST-CRF) is trained to capture both the motion and shape characteristics of objects within consecutive frames. A hierarchical tree is then constructed via hierarchical segmentation that sub-divides frames into regions based on color intensity and regional velocity. A hierarchical sampling approach is then introduced to construct new intermediate frames between adjacent video frames, where estimated intermediate frames are constructed at each level of a hierarchical tree constructed such that the probability of the ST-CRF is maximized. Preliminary results using videos with different motion characteristics show that the proposed approach has potential for producing intermediate frames with high visual quality. Mohammad Javad Shafiee, Zohreh Azimifar, Alexander Wong, Paul W. Fieguth |
ISM | 3 |
| 2011 | Stochastic Medium Access for Cognitive Radio Ad Hoc NetworksabstractIn ad hoc cognitive radio (CR) networks, medium access control (MAC) design has been raised as a major challenge due to its highly dynamic nature and strong user diversity, particularly in situations where a dedicated control channel is not reserved among the distributed CR nodes. In this paper, we propose a novel Stochastic Medium Access (SMA) scheme that takes interference constraints into account to improve spectrum sharing efficiency. Specifically, the proposed SMA scheme is developed to serve in a CR network without dedicated control channels, such that the probability of successful channel accesses can be maximized. The formulated optimization problem is then solved by using a dynamic Markov-Chain Monte-Carlo scheme. Moreover, the paper introduces a suite of mechanisms for implementation of the proposed SMA scheme, including segmentation of long packets and contention resolution, which is working on top of power controlled Request-to-Send (RTS) and Clear-to-Send (CTS) exchanges in a multichannel environment. An analytical model is developed on the proposed SMA scheme using an absorbing Markov chain model to evaluate throughput of the secondary user network. Extensive simulation is conducted to study the impact of some important factors on the proposed SMA scheme, such as channel conditions and secondary traffic loads. Xiao-Yu Wang 0010, Alexander Wong, Pin-Han Ho |
IEEE J. Sel. Areas Commun. | 2 |
| 2011 | Efficient Globally Optimal Registration of Remote Sensing Imagery via Quasi-Random Scale-Space Structural Correlation Energy FunctionalabstractA novel energy functional for automatic registration of remote sensing imagery based on quasi-random scale-space structural correlation is presented. The structural correlation energy functional takes advantage of the fact that, for many types of remote sensing imagery, there exist common structures at different scales even if the acquired images have very different intensity characteristics. The proposed energy functional also takes advantage of the noise robustness and feature localization properties of quasi-random scale-space theory. An efficient globally exhaustive optimization strategy in the frequency domain is developed for registering remote sensing imagery based on the proposed energy functional. Promising test results on interband, intraband, and intermodal remote sensing image sets show that the proposed method has the advantage of being robust to differing sensing conditions and large misalignments. Wen Zhang 0012, Alexander Wong, Akshaya Kumar Mishra, Paul W. Fieguth, David A. Clausi |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2011 | Stochastic image denoising based on Markov-chain Monte Carlo sampling
Alexander Wong, Akshaya Kumar Mishra, Wen Zhang 0012, Paul W. Fieguth, David A. Clausi |
Signal Process. | 1 |
| 2011 | Enhanced Seam Carving via Integration of Energy Gradient FunctionalsabstractThis letter proposes an improved seam carving approach for content-aware image retargeting. The proposed algorithm extends upon the backward and forward energy cost functionals used in previous seam carving methods by incorporating an energy gradient cost functional in the optimization process. This combined absolute energy cost functional penalizes seam candidates that cross areas of local extrema, which characterizes regions with high concentrations of “important” image detail. Experiments show superior visual quality results using the proposed absolute energy cost functional over existing methods on a variety of images characterized by high image detail concentration. Michael Frankovich, Alexander Wong |
IEEE Signal Process. Lett. | 2 |
| 2011 | Synthesis of Remote Sensing Label Fields Using a Tree-Structured Hierarchical ModelabstractThe systematic evaluation of synthetic aperture radar (SAR) data analysis tools, such as segmentation and classification algorithms for geographic information systems, is difficult given the unavailability of ground-truth data in most cases. Therefore, testing is typically limited to small sets of pseudoground-truth data collected manually by trained experts, or primitive synthetic sets composed of simple geometries. To address this issue, we investigate the potential of employing an alternative approach, which involves the synthesis of SAR data and corresponding label fields from real SAR data for use as a reliable evaluation testbed. Given the scale-dependent nonstationary nature of SAR data, a new modeling approach that combines a resolution-oriented hierarchical method with a region-oriented binary tree structure is introduced to synthesize such complex data in a realistic manner. Experimental results using operational RADARSAT SAR sea-ice data and SIR-C/X-SAR land-mass data show that the proposed hierarchical approach can better model complex nonstationary scale structures than local MRF approaches and existing nonparametric methods, thus making it well suited for synthesizing SAR data and the corresponding label fields for potential use in the systematic evaluation of SAR data analysis tools. Ying Liu 0010, Alexander Wong, Paul W. Fieguth |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2011 | Fisher-Tippett Region-Merging Approach to Transrectal Ultrasound Prostate Lesion SegmentationabstractIn this paper, a computerized approach to segmenting prostate lesions in transrectal ultrasound (TRUS) images is presented. The segmentation of prostate lesions from TRUS images is very challenging due to issues, such as poor contrast, low SNRs, and irregular shape variations. To address these issues, a novel approach is employed to segment the lesions from the surrounding prostate, where region merging is performed via a region-merging likelihood function based on regional statistics, as well as Fisher-Tippett statistics. Experimental results using TRUS prostate images demonstrate that the proposed Fisher-Tippett region-merging approach achieves more accurate segmentation of prostate lesions when compared to other segmentation methods. Alexander Wong, Jacob Scharcanski |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2011 | Automatic Skin Lesion Segmentation via Iterative Stochastic Region MergingabstractAn automatic method for segmenting skin lesions in conventional macroscopic images is presented. The images are acquired with conventional cameras, without the use of a dermoscope. Automatic segmentation of skin lesions from macroscopic images is a very challenging problem due to factors such as illumination variations, irregular structural and color variations, the presence of hair, as well as the occurrence of multiple unhealthy skin regions. To address these factors, a novel iterative stochastic region-merging approach is employed to segment the regions corresponding to skin lesions from the macroscopic images, where stochastic region merging is initialized first on a pixel level, and subsequently on a region level until convergence. A region merging likelihood function based on the regional statistics is introduced to determine the merger of regions in a stochastic manner. Experimental results show that the proposed system achieves overall segmentation error of under 10% for skin lesions in macroscopic images, which is lower than that achieved by existing methods. Alexander Wong, Jacob Scharcanski, Paul W. Fieguth |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2010 | Remote sensing image synthesisabstractFor remote sensing data, the testing analysis tools is difficult since the ground-truth data are not available in many cases. To address this issue, a novel method for image synthesis is presented for use as a evaluation test-bed. Given the scale-dependent, non-stationary nature of remotely sensed data, a new modeling approach that combines a resolution-oriented hierarchical method with a regional label-oriented binary tree structure is introduced to synthesize such complex data. In this paper, we are proposing on first synthesizing a label field, which contains the complex structural characteristics, then synthesizing the texture based on the generated label field for a more accurate modeling. Experimental results using operational RADARSAT SAR sea-ice image data show that the proposed method is capable of modeling complex, nonstationary scale structures, thus making it well-suitable to produce reliable, realistic remote sensing imagery. Ying Liu 0010, Alexander Wong, Paul W. Fieguth |
IGARSS | 2 |
| 2010 | A new Bayesian source separation approach to blind decorrelation of SAR dataabstractIn this paper, a novel approach for performing blind decorrelation of SAR data is proposed. A patch-wise computation of the point-spread function (PSF) is performed directly from the SAR data to account for spatial nonstationarities present in SAR. The problem of estimating the PSF is formulated as an additive source separation problem in the frequency domain, and is subsequently solved using a Bayesian least squares estimation approach based on a Fisher-Tippett log-scatter model. Experimental results using both simulated SAR data and real RADARSAT-2 SAR sea-ice data showed that the proposed decorrelation approach can successfully learn the correct PSF and significantly reduce the correlation in SAR data. Alexander Wong, Paul W. Fieguth |
IGARSS | 1 |
| 2010 | IceSynth II: Synthesis of SAR Sea-Ice Imagery Using Region-Based Posterior SamplingabstractA novel method for synthesizing synthetic aperture radar (SAR) sea-ice imagery named IceSynth II is presented. A Markov random field model is assumed, and a conditional sampling approach is used to learn local conditional posterior probability distributions on a regional basis. Synthetic SAR sea-ice images and the associated ground-truth segmentations are generated using a region-based posterior sampling approach. Experimental results using single-polarization RADARSAT-1 and dual-polarization RADARSAT-2 SAR sea-ice imagery provided by the Canadian Ice Service show that IceSynth II is capable of producing SAR sea-ice imagery that is more realistic than existing approaches. The synthesized images are well suited for performing systematic and reliable objective evaluation of SAR sea-ice image segmentation methods. Alexander Wong, Peter Yu, Wen Zhang 0012, David A. Clausi |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2010 | CPOL: Complex phase order likelihood as a similarity measure for MR-CT registration
Alexander Wong, David A. Clausi, Paul W. Fieguth |
Medical Image Anal. | 1 |
| 2010 | Enabling scalable spectral clustering for image segmentation
Frederick Tung, Alexander Wong, David A. Clausi |
Pattern Recognit. | 2 |
| 2010 | Quasi-random nonlinear scale space
Akshaya Kumar Mishra, Alexander Wong, David A. Clausi, Paul W. Fieguth |
Pattern Recognit. Lett. | 2 |
| 2010 | AISIR: Automated inter-sensor/inter-band satellite image registration using robust complex wavelet feature representations
Alexander Wong, David A. Clausi |
Pattern Recognit. Lett. | 1 |
| 2010 | KPAC: A Kernel-Based Parametric Active Contour Method for Fast Image SegmentationabstractObject boundary detection has been a topic of keen interest to the signal processing and pattern recognition community. A popular approach for object boundary detection is parametric active contours. Existing parametric active contour approaches often suffer from slower convergence rates, difficulty dealing with complex high curvature boundaries, and are prone to being trapped in local optima in the presence of noise and background clutter. To address these problems, this paper proposes a novel kernel-based active contour (KPAC) approach, which replaces the conventional internal energy term used in existing approaches by incorporating an adaptive kernel derived for the underlying image characteristics. Experimental results demonstrate that the KPAC approach achieves state-of-the-art performance when compared to two other state-of-the-art parametric active contour approaches. Akshaya Kumar Mishra, Alexander Wong |
IEEE Signal Process. Lett. | 2 |
| 2010 | Generalized Probabilistic Scale Space for Image RestorationabstractA novel generalized sampling-based probabilistic scale space theory is proposed for image restoration. We explore extending the definition of scale space to better account for both noise and observation models, which is important for producing accurately restored images. A new class of scale-space realizations based on sampling and probability theory is introduced to realize this extended definition in the context of image restoration. Experimental results using 2-D images show that generalized sampling-based probabilistic scale-space theory can be used to produce more accurate restored images when compared with state-of-the-art scale-space formulations, particularly under situations characterized by low signal-to-noise ratios and image degradation. Alexander Wong, Akshaya Kumar Mishra |
IEEE Trans. Image Process. | 1 |
| 2010 | An adaptive Monte Carlo approach to phase-based multimodal image registrationabstractIn this paper, a novel multiresolution algorithm for registering multimodal images, using an adaptive Monte Carlo scheme is presented. At each iteration, random solution candidates are generated from a multidimensional solution space of possible geometric transformations, using an adaptive sampling approach. The generated solution candidates are evaluated based on the Pearson type-VII error between the phase moments of the images to determine the solution candidate with the lowest error residual. The multidimensional sampling distribution is refined with each iteration to produce increasingly more plausible solution candidates for the optimal alignment between the images. The proposed algorithm is efficient, robust to local optima, and does not require manual initialization or prior information about the images. Experimental results based on various real-world medical images show that the proposed method is capable of achieving higher registration accuracy than existing multimodal registration algorithms for situations, where little to no overlapping regions exist. Alexander Wong |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2010 | Extended Knowledge-Based Reasoning Approach to Spectrum Sensing for Cognitive RadioabstractIn this paper, a novel scheme for cognitive radio (CR) spectrum sensing in medium access control (MAC) layer, called as extended knowledge-based reasoning (EKBR), is proposed. The target of EKBR is to improve the fine sensing efficiency by jointly considering a number of network states and environmental statistics, including fast sensing results, short-term statistical information, channel quality, data transmission rate, and channel contention characteristics. This is for a better estimation on the optimal range of spectrum for fine sensing so as to adaptively reduce the overall channel sensing time. Performance analysis is conducted on the proposed EKBR scheme using a multidimensional absorbing Markov chain to evaluate various performance metrics of interest, such as average sensing delay (or referred to as sensing overhead in the study), average data transmission rate, and percentage of missed spectrum opportunities. Numerical results show that the proposed EKBR scheme achieves better performance than that by the state-or-the-art techniques while yielding less computation complexity and sensing overhead. Xiao-Yu Wang 0010, Alexander Wong, Pin-Han Ho |
IEEE Trans. Mob. Comput. | 2 |
| 2010 | Dynamically optimized spatiotemporal prioritization for spectrum sensing in cooperative cognitive radio
Xiao-Yu Wang 0010, Alexander Wong, Pin-Han Ho |
Wirel. Networks | 2 |
| 2009 | Stochastic Channel Prioritization for Spectrum Sensing in Cooperative Cognitive RadioabstractIn this paper, a novel cooperative stochastic channel prioritization algorithm is presented for the purpose of improving spectrum sensing efficiency in cooperative cognitive radio systems. The proposed algorithm achieves the goal by prioritizing the channels for fine sensing based on both local statistics obtained by the cognitive radio as well as long-term spatiotemporal statistics obtained from other cognitive radios. Channel priority is determined in a stochastic manner by performing statistical fusion on the local statistics and statistics from neighboring cognitive radios to obtain a biasing density from which stochastic sampling can be used to identify the likelihood of channel availability. Therefore, the individual cognitive radios collaborate to improve the likelihood of each cognitive radio in obtaining available channels. Simulation results show that the proposed cooperative stochastic channel prioritization algorithm can be used to reduce both sensing overhead and percentage of missed opportunities when implemented in a complimentary manner with existing cooperative cognitive radio systems. Xiao-Yu Wang 0010, Alexander Wong, Pin-Han Ho |
CCNC | 2 |
| 2009 | Prioritized spectrum sensing in cognitive radio based on spatiotemporal statistical fusionabstractIn this paper, a novel statistics-driven spectrum sensing algorithm is developed for improving spectrum sensing efficiency in the media access control (MAC) layer of cognitive radio (CR) systems. The proposed algorithm aims to achieve higher spectrum sensing efficiency and spectrum access opportunity by prioritizing channels for fine sensing based on the statistical likelihood of channel availability. The sensing priority is obtained by jointly exploiting the long-term spatiotemporal statistics recorded from the historical result of fine sensing, the short-term statistical information of channel condition obtained from a small-scale observation window, and the instantaneous statistical information obtained from fast sensing. Simulation results show that the proposed prioritization algorithm can achieve improved data transmission rates and reduced missed spectrum access opportunities when compared to the conventional non- prioritization spectrum sensing approach for situations where cooperative spectrum sensing is not suitable. Xiao-Yu Wang 0010, Alexander Wong, Pin-Han Ho |
WCNC | 2 |
| 2009 | Fast phase-based registration of multimodal image data
Alexander Wong, Paul W. Fieguth |
Signal Process. | 1 |
| 2009 | Phase-Adaptive Superresolution of Mammographic Images Using Complex WaveletsabstractThis correspondence describes a new superresolution approach for enhancing the resolution of mammographic images using complex wavelet frequency information. This method allows regions of interest of a mammographic image to be viewed in enhanced resolution while reducing the patient exposure to radiation. The proposed method exploits the structural characteristics of breast tissues being imaged and produces higher resolution mammographic images with sufficient visual fidelity that fine image details can be discriminated more easily. In our approach, the superresolution problem is formulated as a constrained optimization problem using a third-order Markov prior model and adapts the priors based on the phase variations of the low-resolution mammographic images. Experimental results indicate the proposed method is more effective at preserving the visual information when compared with existing resolution enhancement methods. Alexander Wong, Jacob Scharcanski |
IEEE Trans. Image Process. | 1 |
| 2008 | Efficient nonlocal-means denoising using the SVDabstractNonlocal-means (NL-means) is an image denoising method that replaces each pixel by a weighted average of all the pixels in the image. Unfortunately, the method requires the computation of the weighting terms for all possible pairs of pixels, making it computationally expensive. Some short-cuts assign a weight of zero to any pixel pairs whose neighbourhood averages are too dissimilar. In this paper, we propose an alternative strategy that uses the SVD to more efficiently eliminate pixel pairs that are dissimilar. Experiments comparing this method against other NL-means speed-up strategies show that its refined discrimination between similar and dissimilar pixel neighbourhoods significantly improves the denoising effect. Jeff Orchard, Mehran Ebrahimi, Alexander Wong |
ICIP | 3 |
| 2008 | Illumination invariant active contour-based segmentation using complex-valued waveletsabstractThis paper introduces a novel approach to the problem of active contour-based segmentation through the use of complex- valued wavelets. In traditional active contour-based segmentation techniques based on level set methods, the energy functionals are defined based on intensity gradients. This makes them highly sensitive to situations where the underlying image content is characterized by image non-homogeneities due to illumination and contrast conditions. In the proposed approach, the energy functionals used to evolve a level set function are based on the moments of phase coherence of complex-valued wavelet components. This formulation is highly invariant to non-homogeneities caused by illumination and contrast variations. Experimental results demonstrate that the proposed approach can be used to improve existing active contour-based segmentation methods under situations characterized by image non-homogeneities. Alexander Wong |
ICIP | 1 |
| 2008 | A perceptually adaptive approach to image denoising using anisotropic non-local meansabstractThis paper introduces a novel perceptually adaptive approach to image denoising using anisotropic non-local means. In the classical non-local means image denoising approach, the value of a pixel is determined based on the weighted average of other pixels, where the weights are determined based on a fixed isotropically weighted similarity function between the local neighborhoods. In the proposed algorithm, we demonstrate that noticeably improved perceptual quality can be achieved through the use of adaptive anisotropically weighted similarity functions between local neighborhoods. This is accomplished by adapting the similarity weighing function in an anisotropic manner based on the perceptual characteristics of the underlying image content derived efficiently based on the Mexican Hat wavelet. Experimental results show that the proposed method can be used to provide improved perceptual quality in the denoised image both quantitatively and qualitatively when compared to existing methods. Alexander Wong, Paul W. Fieguth, David A. Clausi |
ICIP | 1 |
| 2008 | A nonlocal-means approach to exemplar-based inpaintingabstractThis paper introduces a novel approach to the problem of image inpainting through the use of nonlocal-means. In traditional inpainting techniques, only local information around the target regions are used to fill in the missing information, which is insufficient in many cases. More recent inpainting techniques based on the concept of exemplar-based synthesis utilize nonlocal information but in a very limited way. In the proposed algorithm, we use nonlocal image information from multiple samples within the image. The contribution of each sample to the reconstruction of a target pixel is determined using an weighted similarity function and aggregated to form the missing information. Experimental results show that the proposed method yields quantitative and qualitative improvements compared to the current exemplar-based approach. The proposed approach can also be integrated into existing exemplar-based inpainting techniques to provide improved visual quality. Alexander Wong, Jeff Orchard |
ICIP | 1 |
| 2008 | Simultaneous multi-modal registration of multiple images based on multi-dimensional joint phase moment distributionsabstractIn this paper, a novel method for simultaneously registering multiple images acquired from different imaging modalities is presented. The optimal alignment is computed as the set of transformations that minimize the dispersion of the multi-dimensional joint phase moment distribution. Dispersion is measured as the cumulative quadratic orthogonal distance between samples from the joint phase moment distribution and the corresponding multi-dimensional fitted hyperline. The proposed method is designed to be computationally efficient, robust to signal non-homogeneities and noise, and maintains internal consistency amongst all images. Experimental results using real-world medical and remote sensing images show that the proposed method achieves a high level of registration accuracy when simultaneously registering multiple multimodal images. Alexander Wong |
ICPR | 1 |
| 2008 | Phase-adaptive image signal fusion using complex-valued waveletsabstractThis paper presents a novel method for medical image signal fusion using complex-valued wavelets to enhance the information content in the fused signal from a perceptual manner. The proposed method introduces an adaptively weighted aggregation of signal characteristics based on the phase characteristics of medical image signals. The proposed method exploits the phase characteristics of the image signals to adaptively accentuate important anatomical and functional characteristics captured by each image signal during the signal fusion process. Experimental results show that the proposed method can improve the visualization of important anatomical and functional characteristics from different medical image signals in the fused image signal when compared with non-adaptive image signal fusion methods. Alexander Wong, David A. Clausi, Paul W. Fieguth |
ICPR | 1 |
| 2008 | An adaptive Monte Carlo approach to nonlinear image denoisingabstractThis paper introduces a novel stochastic approach to image denoising using an adaptive Monte Carlo scheme. Random samples are generated from the image field using a spatially-adaptive importance sampling approach. Samples are then represented using Gaussian probability distributions and a sample rejection scheme is performed based on a chi2statistical hypothesis test. The remaining samples are then aggregated based on Pearson Type VII statistics to create a non-linear estimate of the denoised image. The proposed method exploits global information redundancy to suppress noise in an image. Experimental results show that the proposed method provides superior noise suppression performance both quantitatively and qualitatively when compared to the state-of-the-art image denoising methods. Alexander Wong, Akshaya Kumar Mishra, Paul W. Fieguth, David A. Clausi |
ICPR | 1 |
| 2008 | Deblocking of Block-Transform Compressed Images Using Phase-Adaptive Shifted ThresholdingabstractMany popular image compression schemes are based on block-transform coding, a technique where images are broken into small blocks of pixels prior to transformation and compression. Block-transform coding often introduces blocking artifacts which are particularly prevalent at low bit-rates due to quantization errors. A novel algorithm for deblocking block-transform compressed images is proposed in this paper. This algorithm is based on a phase-adaptive, shifted thresholding technique that estimates the original uncompressed image as the weighted sum of shifted versions of the decompressed image subjected to a threshold. An efficient integer transform is used to construct the shifted versions of the decompressed image. The aggregation weights are obtained adaptively using the local phase moment characteristics of the underlying image content. The proposed algorithm utilizes important human perceptual characteristics to provide effective image deblocking while preserving image detail. Experimental results show that the proposed algorithm is more efficient than comparable methods and yields both subjective results and peak signal-to-noise ratio (PSNR) results comparable to existing methods. Alexander Wong, William D. Bishop |
ISM | 1 |
| 2008 | Robust Edge Detection Based on Non-local Contribution of Local Frequency CharacteristicsabstractThis paper introduces a robust approach to the problem of edge detection in situations characterized by high levels of image degradation and illumination non-uniformity. Popular edge detection methods are dependent solely on local image characteristics to localize edge information. In the proposed algorithm, we demonstrate that superior edge detection performance can be achieved by utilizing nonlocal image characteristics. This is accomplished by estimating the edge strength of a given pixel based on the local frequency characteristics of every pixel within an image. The proposed method is highly robust to scenarios characterized by illumination non-uniformity and low signal-to-noise ratios. Experimental results show that using the proposed method, edge detection can be qualitatively improved over existing methods on a small set of test images. Alexander Wong, William D. Bishop |
ISM | 1 |
| 2008 | Efficient least squares fusion of MRI and CT images using a phase congruency model
Alexander Wong, William D. Bishop |
Pattern Recognit. Lett. | 1 |
| 2008 | Adaptive bilateral filtering of image signals using local phase characteristics
Alexander Wong |
Signal Process. | 1 |
| 2008 | Efficient FFT-Accelerated Approach to Invariant Optical-LIDAR RegistrationabstractThis paper presents a fast Fourier transform (FFT)-accelerated approach designed to handle many of the difficulties associated with the registration of optical and light detection and ranging (LIDAR) images. The proposed algorithm utilizes an exhaustive region correspondence search technique to determine the correspondence between regions of interest from the optical image with the LIDAR image over all translations for various rotations. The computational cost associated with exhaustive search is greatly reduced by exploiting the FFT. The substantial differences in intensity mappings between optical and LIDAR images are addressed through local feature mapping transformation optimization. Geometric distortions in the underlying images are dealt with through a geometric transformation estimation process that handles various transformations such as translation, rotation, scaling, shear, and perspective transformations. To account for mismatches caused by factors such as severe contrast differences, the proposed algorithm attempts to prune such outliers using the random sample consensus technique to improve registration accuracy. The proposed algorithm has been tested using various optical and LIDAR images and evaluated based on its registration accuracy. The results indicate that the proposed algorithm is suitable for the multimodal invariant registration of optical and LIDAR images. Alexander Wong, Jeff Orchard |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2007 | ARRSI: Automatic Registration of Remote-Sensing ImagesabstractThis paper presents the Automatic Registration of Remote-Sensing Images (ARRSI); an automatic registration system built to register satellite and aerial remotely sensed images. The system is designed specifically to address the problems associated with the registration of remotely sensed images obtained at different times and/or from different sensors. The ARRSI system is capable of handling remotely sensed images geometrically distorted by various transformations such as translation, rotation, and shear. Global and local contrast issues associated with remotely sensed images are addressed in ARRSI using control-point detection and matching processes based on a phase-congruency model. Intensity-difference issues associated with multimodal registration of remotely sensed images are addressed in ARRSI through the use of features that are invariant to intensity mappings during the control-point matching process. An adaptive control-point matching scheme is employed in ARRSI to reduce the performance issues associated with the registration of large remotely sensed images. Finally, a variation on the Random Sample and Consensus algorithm called Maximum Distance Sample Consensus is introduced in ARRSI to improve the accuracy of the transformation model between two remotely sensed images while minimizing computational overhead. The ARRSI system has been tested using various satellite and aerial remotely sensed images and evaluated based on its accuracy and computational performance. The results indicate that the registration accuracy of ARRSI is comparable to that produced by a human expert and improvement over the baseline and multimodal sum of squared differences registration techniques tested. Alexander Wong, David A. Clausi |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2006 | A Flexible Content-Based Approach to Adaptive Image CompressionabstractRecent research in image compression has focused on lossy compression algorithms. However, the baseline implementations of such algorithms generally use a universal quantization process that results in poor image quality for certain types of images, particularly mixed-content images. This paper addresses this image quality issue by presenting a new algorithm that provides flexible and customizable image quality preservation by introducing an adaptive thresholding and quantization process based on content information such as edge and texture characteristics from the actual image. The algorithm is designed to improve visual quality based on the human vision system. Experimental results from the compression of various test images show noticeable improvements both quantitatively and qualitatively relative to baseline implementations as well as other adaptive techniques Alexander Wong, William D. Bishop |
ICME | 1 |
| 2006 | Adaptive Perceptual Degradation Based on Video UsageabstractThe emergence of online digital media sales services has given rise to the issue of online digital rights management (DRM). This paper presents a novel approach to online digital rights management and video content protection through the use of adaptive perceptual degradation based on usage. The proposed system facilitates the sharing of purchased digital videos between users while offering an incentive for users to acquire original video content through legal means. This is accomplished by adaptively degrading the quality of the digital video content based on a number of usage factors such as the number of times viewed and the number of copies made. Experimental results demonstrate the effectiveness of this system in striking a balance between the freedom of file sharing and content protection Alexander Wong, William D. Bishop |
ISM | 1 |
| 2006 | Practical Content-Adaptive Subsampling for Image and Video CompressionabstractSubsampling is a commonly used technique for modern image and video compression. Existing image and video standards such as JPEG, MPEG, and H.264 provide support for uniform chroma subsampling. This paper presents an algorithm that uses adaptive luma subsampling based on a combination of three perceptually significant image characteristics (texture, edges, and brightness) to complement uniform chroma subsampling. The algorithm is computationally efficient, simple to implement, and easy to integrate into existing standards. Experimental results show that the introduction of adaptive luma Subsampling improves image quality both quantitatively and qualitatively when compared with the sole use of uniform chroma subsampling Alexander Wong, William D. Bishop |
ISM | 1 |