Philip Ogunbona

dblp:09/325 · also Philip O. Ogunbona · DBLP profile ↗
← Back
69ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0003-4119-2873ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 50 · 1 since 2021Artificial intelligence and machine learning · 29 · 3 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Video understanding and tracking · 36% Transfer learning and domain adaptation · 33% 3D vision · 8%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Medical and health informatics · 74% Bioinformatics and computational biology · 26%
Computer graphics and multimedia
6 papers
Image and video processing · 52% Multimedia analysis and retrieval · 34% Image and video coding · 14%

Topics — the 30 heaviest of 40, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action recognition
1.242018
Depth Pooling Based Large-Scale 3-D Action Recognition With Convolutional Neural Networks · IEEE Trans. Multim. 2018
Cooperative Training of Deep Aggregation Networks for RGB-D Action Recognition · AAAI 2018
Scene Flow to Action Map: A New Representation for RGB-D Based Action Recognition with Convolutional Neural Networks · CVPR 2017
Computer vision › Video understanding and tracking › action recognition › multimodal action recognition
RGB-D action recognition
0.622018
Cooperative Training of Deep Aggregation Networks for RGB-D Action Recognition · AAAI 2018
Scene Flow to Action Map: A New Representation for RGB-D Based Action Recognition with Convolutional Neural Networks · CVPR 2017
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.622018
Importance Weighted Adversarial Nets for Partial Domain Adaptation · CVPR 2018
Joint Geometrical and Statistical Alignment for Visual Domain Adaptation · CVPR 2017
Natural language and speech › Information extraction and text analysis › sentiment analysis
aspect-based sentiment analysis
0.412020
Modelling Context and Syntactical Features for Aspect-based Sentiment Analysis · ACL 2020
Medical and health informatics › clinical diagnosis › neurodegenerative disease diagnosis
alzheimer's disease diagnosis
0.422014
Discriminative Sparse Inverse Covariance Matrix: Application in Brain Functional Network Classification · CVPR 2014
Discriminative Brain Effective Connectivity Analysis for Alzheimer's Disease: A Kernel Learning Approach upon Sparse Gaussian Bayesian Network · CVPR 2013
Medical and health informatics › neuroimaging
neuroimaging analysis
0.422014
Discriminative Sparse Inverse Covariance Matrix: Application in Brain Functional Network Classification · CVPR 2014
Discriminative Brain Effective Connectivity Analysis for Alzheimer's Disease: A Kernel Learning Approach upon Sparse Gaussian Bayesian Network · CVPR 2013
Machine learning › Transfer learning and domain adaptation › domain adaptation › distribution adaptation
adversarial domain adaptation
0.312018
Importance Weighted Adversarial Nets for Partial Domain Adaptation · CVPR 2018
Machine learning › Deep learning architectures and training
convolutional neural network
0.312018
Cooperative Training of Deep Aggregation Networks for RGB-D Action Recognition · AAAI 2018
Computer vision › Video understanding and tracking
gesture recognition
0.312018
Depth Pooling Based Large-Scale 3-D Action Recognition With Convolutional Neural Networks · IEEE Trans. Multim. 2018
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
partial domain adaptation
0.312018
Importance Weighted Adversarial Nets for Partial Domain Adaptation · CVPR 2018
Image and video processing › image restoration
image dehazing
0.312018
Detection and Separation of Smoke From Single Image Frames · IEEE Trans. Image Process. 2018
Machine learning › Transfer learning and domain adaptation
distribution matching
0.312017
Joint Geometrical and Statistical Alignment for Visual Domain Adaptation · CVPR 2017
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.312017
Joint Geometrical and Statistical Alignment for Visual Domain Adaptation · CVPR 2017
Computer vision › 3D vision
scene flow estimation
0.312017
Scene Flow to Action Map: A New Representation for RGB-D Based Action Recognition with Convolutional Neural Networks · CVPR 2017
Machine learning › Transfer learning and domain adaptation › domain adaptation
visual domain adaptation
0.312017
Joint Geometrical and Statistical Alignment for Visual Domain Adaptation · CVPR 2017
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian network
0.212016
Learning Discriminative Bayesian Networks from High-Dimensional Continuous Neuroimaging Data · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Bioinformatics and computational biology › neuroscience › neuroinformatics
brain network analysis
0.212016
Learning Discriminative Bayesian Networks from High-Dimensional Continuous Neuroimaging Data · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Medical and health informatics
neuroimaging
0.212016
Learning Discriminative Bayesian Networks from High-Dimensional Continuous Neuroimaging Data · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Video understanding and tracking › action recognition › 3d action recognition
depth-based action recognition
0.212015
ConvNets-Based Action Recognition from Depth Maps through Virtual Cameras and Pseudocoloring · ACM Multimedia 2015
Computer vision › 3D vision › depth image analysis
depth map representation
0.212015
ConvNets-Based Action Recognition from Depth Maps through Virtual Cameras and Pseudocoloring · ACM Multimedia 2015
Medical and health informatics › neuroimaging › neuroimaging analysis
functional connectivity analysis
0.212014
Discriminative Sparse Inverse Covariance Matrix: Application in Brain Functional Network Classification · CVPR 2014
Multimedia analysis and retrieval › video surveillance
smoke detection
0.212014
Smoke Detection in Video: An Image Separation Approach · Int. J. Comput. Vis. 2014
Multimedia analysis and retrieval
video analysis
0.212014
Smoke Detection in Video: An Image Separation Approach · Int. J. Comput. Vis. 2014
Bioinformatics and computational biology › computational neuroscience › brain connectivity analysis
brain effective connectivity
0.212013
Discriminative Brain Effective Connectivity Analysis for Alzheimer's Disease: A Kernel Learning Approach upon Sparse Gaussian Bayesian Network · CVPR 2013
Natural language and speech › Language models and text generation › text representation
contextualized word embeddings
0.112020
Modelling Context and Syntactical Features for Aspect-based Sentiment Analysis · ACL 2020
Machine learning › Learning paradigms
class imbalance
0.112018
Importance Weighted Adversarial Nets for Partial Domain Adaptation · CVPR 2018
Computer vision › Vision and language
multimodal representation
0.112018
Cooperative Training of Deep Aggregation Networks for RGB-D Action Recognition · AAAI 2018
Image and video processing
image matting
0.112018
Detection and Separation of Smoke From Single Image Frames · IEEE Trans. Image Process. 2018
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
subspace learning
0.112017
Joint Geometrical and Statistical Alignment for Visual Domain Adaptation · CVPR 2017
Image and video coding › transform coding
DCT coefficient recovery
0.112006
Recovering DC Coefficients in Block-Based DCT · IEEE Trans. Image Process. 2006

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 1.2convex optimization · 0.5syntactic relative distance · 0.4self-attention · 0.4part-of-speech embeddings · 0.4dependency embeddings · 0.4fisher kernel · 0.4triplet loss · 0.3sparse representation · 0.3softmax loss · 0.3ranking loss · 0.3hierarchical bidirectional rank pooling · 0.3fine-tuning · 0.3dual dictionary · 0.3cooperative training · 0.3atmospheric scattering model · 0.3adversarial nets · 0.3max-margin learning · 0.2
YearPublicationVenuePosition
2026 Unrolling operator splitting in learning PDEs for object detection
abstract
Object detection presents significant challenges due to the variability in object scale, location, and orientation within images. Most state-of-the-art detectors are based on convolutional or Transformer architectures, which, while effective, often result in deep, opaque models that generalise poorly and lack interpretability. In contrast, iterative algorithms offer greater transparency and generalisation, albeit at the cost of efficiency and accuracy. In this work, we reformulate object detection as a partial differential equation (PDE)-constrained optimal control problem. This formulation exploits linear combinations of fundamental differential invariants—such as translation and rotation invariance—to embed structural priors into the learning process. We solve this problem using operator splitting via the Alternating Direction Method of Multipliers (ADMM), and unroll each optimisation step into a network layer, yielding a novel architecture: ADMM-ODNet. This approach provides a principled and interpretable alternative to conventional deep networks. Experimental results on the Corel, Pascal VOC and COCO datasets demonstrate that ADMM-ODNet outperforms leading models such as Cascade Mask R-CNN, Swin Transformer, Deformable DETR, DINO, DN-and RT-DETR, and achieves performance comparable to Plain DETR and YOLO, while requiring significantly fewer parameters.
Banu Wirawan Yohanes, Philip Ogunbona, Wanqing Li 0001
Neurocomputing2
2026 Efficient spatial pyramid for object detection
abstract
One-stage object detectors emerge as a trade-off between detection accuracy and speed. However, they do not exploit long-range context relationship and their performance can easily drop in complex scenes. In this work, we propose efficient spatial pyramid based on graph model of Laplacian with a novel graph correlation filter. This filter is designed to measure symmetric uncertainty among features to learn long-range context relationship, while preserving important features. Furthermore, we employ Winograd algorithm to reduce floating point operations significantly without decreasing detection performance, by trading-off costly multiplication operations with more additions. They enable an efficient object detection. Extensive experiments were conducted on two challenging object detection datasets, COCO and KITTI. The proposed network was compared to state-of-the-art efficient object detectors, MobileNet-SSD Lite, YOLO, MobileVIT, Tiny DSOD, and EfficientDet. Detailed convergence proof and training epoch analysis provide strong support and evidence for the achieved results improving overall detection accuracy in complex scenes with less computational resources. • Hierarchical Laplacian-driven design capturing both fine-grained and global context. • Symmetrical uncertainty filter for capturing long-range feature interactions. • Accelerated convolutional layers via Winograd minimal filtering. • Thorough empirical validation and ablation study on COCO and KITTI datasets. • ESP attains the best-in-class accuracy-to-efficiency trade-off.
Banu Wirawan Yohanes, Philip Ogunbona, Wanqing Li 0001
Pattern Recognit.2
2025 Asynchronous Joint-Based Temporal Pooling for Skeleton-Based Action Recognition
abstract
Deep neural networks for skeleton-based human action recognition (HAR) often utilize traditional averaging or maximum temporal pooling to aggregate features by treating all joints and frames equally. However, this approach can excessively aggregate less discriminative or even indiscriminative features into the final feature vectors for recognition. To address this issue, a novel method called asynchronous joint adaptive temporal pooling (AJTP) is introduced in this paper. The method aims to enhance action recognition by identifying a set of informative joints across the temporal dimension and applying a joint-based and asynchronous motion-preservative pooling rather than conventional frame-based pooling. The effectiveness of the proposed AJTP has been empirically validated by integrating it with popular Graph Convolutional Network (GCN) models on three benchmark datasets: NTU RGB+D 120, PKUMMD, and Kinetic400. The results have shown that a GCN model with AJTP substantially improves performance compared to its counterpart GCN model with conventional temporal pooling techniques. The source code is available athttps://github.com/ShanakaRG/AJTP.
Shanaka Ramesh Gunasekara, Wanqing Li 0001, Jie Yang 0009, Philip Ogunbona
IEEE Trans. Circuits Syst. Video Technol.4
2024 An attention-based CNN for automatic whole-body postural assessment
abstract
Fully automatic postural assessment is highly useful, but has been challenging. Conventional methods either require manual assessment by ergonomists or depend on special devices that are intrusive, thus being hardly feasible in daily activities and workplaces. In this work, an attention-based convolutional neural network (CNN) is developed for automatic whole-body postural assessment. The proposed network learns to identify highly relevant regions (or body parts) and extract features automatically. Risk of the posture is estimated from the extracted features accordingly. To evaluate the proposed method, a postural dataset, referred to as pH36M, is created by re-targeting Human3.6M, one of the largest publicly available datasets for pose estimation using the Rapid Entire Body Assessment (REBA) criteria. Experimental results on pH36M demonstrate that proposed method achieves promising performance in comparison to baselines and the average assessment scores are substantially aligned with human assessment with a Kappa value of 0.73.
Zewei Ding, Wanqing Li 0001, Jie Yang 0009, Philip Ogunbona
Expert Syst. Appl.4
2020 Modelling Context and Syntactical Features for Aspect-based Sentiment Analysis
abstract
The aspect-based sentiment analysis (ABSA) consists of two conceptual tasks, namely an aspect extraction and an aspect sentiment classification.Rather than considering the tasks separately, we build an end-to-end ABSA solution.Previous works in ABSA tasks did not fully leverage the importance of syntactical information.Hence, the aspect extraction model often failed to detect the boundaries of multi-word aspect terms.On the other hand, the aspect sentiment classifier was unable to account for the syntactical correlation between aspect terms and the context words.This paper explores the grammatical aspect of the sentence and employs the self-attention mechanism for syntactical learning.We combine part-of-speech embeddings, dependencybased embeddings and contextualized embeddings (e.g.BERT, RoBERTa) to enhance the performance of the aspect extractor.We also propose the syntactic relative distance to de-emphasize the adverse effects of unrelated words, having weak syntactic connection with the aspect terms.This increases the accuracy of the aspect sentiment classifier.Our solutions outperform the state-of-the-art models on SemEval-2014 dataset in both two subtasks.
Minh-Hieu Phan, Philip Ogunbona
ACL2
2020 Jointly Learning Visual Poses and Pose Lexicon for Semantic Action Recognition
abstract
A novel method for semantic action recognition through learning a pose lexicon is presented in this paper. A pose lexicon comprises a set of semantic poses, a set of visual poses, and a probabilistic mapping between the visual and semantic poses. This paper assumes that both the visual poses and mapping are hidden and proposes a method to simultaneously learn a visual pose model that estimates the likelihood of an observed video frame being generated from hidden visual poses, and a pose lexicon model establishes the probabilistic mapping between the hidden visual poses and the semantic poses parsed from textual instructions. Specifically, the proposed method consists of two-level hidden Markov models. One level represents the alignment between the visual poses and semantic poses. The other level represents a visual pose sequence, and each visual pose is modeled as a Gaussian mixture. An expectation-maximization algorithm is developed to train a pose lexicon. With the learned lexicon, action classification is formulated as a problem of finding the maximum posterior probability of a given sequence of video frames that follows a given sequence of semantic poses, constrained by the most likely visual pose and the alignment sequences. The proposed method was evaluated on MSRC-12, WorkoutSU-10, WorkoutUOW-18, Combined-15, Combined-17, and Combined-50 action datasets using cross-subject, cross-dataset, zero-shot, and seen/unseen protocols.
Lijuan Zhou 0002, Wanqing Li 0001, Philip Ogunbona, Zhengyou Zhang
IEEE Trans. Circuits Syst. Video Technol.3
2019 Unsupervised domain adaptation: A multi-task learning-based method
Jing Zhang 0017, Wanqing Li 0001, Philip Ogunbona
Knowl. Based Syst.3
2019 A real-time webcam-based method for assessing upper-body postures
Zewei Ding, Wanqing Li 0001, Philip Ogunbona
Mach. Vis. Appl.3
2018 Cooperative Training of Deep Aggregation Networks for RGB-D Action Recognition
abstract
A novel deep neural network training paradigm that exploits the conjoint information in multiple heterogeneous sources is proposed. Specifically, in a RGB-D based action recognition task, it cooperatively trains a single convolutional neural network (named c-ConvNet) on both RGB visual features and depth features, and deeply aggregates the two kinds of features for action recognition. Differently from the conventional ConvNet that learns the deep separable features for homogeneous modality-based classification with only one softmax loss function, the c-ConvNet enhances the discriminative power of the deeply learned features and weakens the undesired modality discrepancy by jointly optimizing a ranking loss and a softmax loss for both homogeneous and heterogeneous modalities. The ranking loss consists of intra-modality and cross-modality triplet losses, and it reduces both the intra-modality and cross-modality feature variations. Furthermore, the correlations between RGB and depth data are embedded in the c-ConvNet, and can be retrieved by either of the modalities and contribute to the recognition in the case even only one of the modalities is available. The proposed method was extensively evaluated on two large RGB-D action recognition datasets, ChaLearn LAP IsoGD and NTU RGB+D datasets, and one small dataset, SYSU 3D HOI, and achieved state-of-the-art results.
Pichao Wang, Wanqing Li 0001, Jun Wan 0001, Philip Ogunbona, Xinwang Liu 0002
AAAI4
2018 Importance Weighted Adversarial Nets for Partial Domain Adaptation
abstract
This paper proposes an importance weighted adversarial nets-based method for unsupervised domain adaptation, specific for partial domain adaptation where the target domain has less number of classes compared to the source domain. Previous domain adaptation methods generally assume the identical label spaces, such that reducing the distribution divergence leads to feasible knowledge transfer. However, such an assumption is no longer valid in a more realistic scenario that requires adaptation from a larger and more diverse source domain to a smaller target domain with less number of classes. This paper extends the adversarial nets-based domain adaptation and proposes a novel adversarial nets-based partial domain adaptation method to identify the source samples that are potentially from the outlier classes and, at the same time, reduce the shift of shared classes between domains.
Jing Zhang 0017, Zewei Ding, Wanqing Li 0001, Philip Ogunbona
CVPR4
2018 Cross-Validated Bandwidth Selection for Precision Matrix Estimation
abstract
Inverse covariance matrix, a.k.a. precision matrix, has wide applications in signal processing and is often estimated from training samples. The quality of estimation can be poor when the sample support is low. Banding/tapering are effective regularization approaches for covariance and precision matrix estimation but the bandwidth must be properly chosen. This paper investigates the bandwidth selection problem for banding/tapering-based precision matrix estimation. Exploiting a regression analysis interpretation of the precision matrix, we design a data-driven cross-validation (CV) method for automatically tuning the bandwidth. The effectiveness of the proposed method is demonstrated by numerical examples under a quadratic loss.
Jun Tong, Jiangtao Xi, Yanguang Yu, Philip Ogunbona
ICASSP4
2018 RGB-D-based human motion recognition with deep learning: A survey
Pichao Wang, Wanqing Li 0001, Philip Ogunbona, Jun Wan 0001, Sergio Escalera
Comput. Vis. Image Underst.3
2018 Detection and Separation of Smoke From Single Image Frames
abstract
This paper proposes novel methods for detecting and separating smoke from a single image frame. Specifically, an image formation model is derived based on the atmospheric scattering models. The separation of a frame into quasi-smoke and quasi-background components is formulated as convex optimization that solves a sparse representation problem using dual dictionaries for the smoke and background components, respectively. A novel feature is constructed as a concatenation of the respective sparse coefficients for detection. In addition, a method based on the concept of image matting is developed to separate the true smoke and background components from the smoke detection results. Extensive experiments on detection were conducted and the results showed that the proposed feature significantly outperforms existing features for smoke detection. In particular, the proposed method is able to differentiate smoke from other challenging objects (e.g. fog/haze, cloud, and so on) with similar visual appearance in a gray-scale frame. Experiments on smoke separation also demonstrated that the proposed separation method can effectively estimate/separate the true smoke and background components.
Hongda Tian, Wanqing Li 0001, Philip Ogunbona, Lei Wang 0001
IEEE Trans. Image Process.3
2018 Depth Pooling Based Large-Scale 3-D Action Recognition With Convolutional Neural Networks
abstract
This paper proposes three simple, compact yet effective representations of depth sequences, referred to respectively as dynamic depth images (DDI), dynamic depth normal images (DDNI), and dynamic depth motion normal images (DDMNI), for both isolated and continuous action recognition. These dynamic images are constructed from a segmented sequence of depth maps using hierarchical bidirectional rank pooling to effectively capture the spatial-temporal information. Specifically, DDI exploits the dynamics of postures over time, and DDNI and DDMNI exploit the 3-D structural information captured by depth maps. Upon the proposed representations, a convolutional neural network (ConvNet)-based method is developed for action recognition. The image-based representations enable us to fine-tune the existing ConvNet models trained on image data without training a large number of parameters from scratch. The proposed method achieved the state-of-art results on three large datasets, namely, the large-scale continuous gesture recognition dataset (means the Jaccard index 0.4109), the large-scale isolated gesture recognition dataset (59.21%), and the NTU RGB+D dataset (87.08% cross-subject and 84.22% cross-view) even though only the depth modality was used.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Philip Ogunbona
IEEE Trans. Multim.5
2017 Scene Flow to Action Map: A New Representation for RGB-D Based Action Recognition with Convolutional Neural Networks
abstract
Scene flow describes the motion of 3D objects in real world and potentially could be the basis of a good feature for 3D action recognition. However, its use for action recognition, especially in the context of convolutional neural networks (ConvNets), has not been previously studied. In this paper, we propose the extraction and use of scene flow for action recognition from RGB-D data. Previous works have considered the depth and RGB modalities as separate channels and extract features for later fusion. We take a different approach and consider the modalities as one entity, thus allowing feature extraction for action recognition at the beginning. Two key questions about the use of scene flow for action recognition are addressed: how to organize the scene flow vectors and how to represent the long term dynamics of videos based on scene flow. In order to calculate the scene flow correctly on the available datasets, we propose an effective self-calibration method to align the RGB and depth data spatially without knowledge of the camera parameters. Based on the scene flow vectors, we propose a new representation, namely, Scene Flow to Action Map (SFAM), that describes several long term spatio-temporal dynamics for action recognition. We adopt a channel transform kernel to transform the scene flow vectors to an optimal color space analogous to RGB. This transformation takes better advantage of the trained ConvNets models over ImageNet. Experimental results indicate that this new representation can surpass the performance of state-of-the-art methods on two large public datasets.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Philip Ogunbona
CVPR6
2017 Joint Geometrical and Statistical Alignment for Visual Domain Adaptation
abstract
This paper presents a novel unsupervised domain adaptation method for cross-domain visual recognition. We propose a unified framework that reduces the shift between domains both statistically and geometrically, referred to as Joint Geometrical and Statistical Alignment (JGSA). Specifically, we learn two coupled projections that project the source domain and target domain data into low-dimensional subspaces where the geometrical shift and distribution shift are reduced simultaneously. The objective function can be solved efficiently in a closed form. Extensive experiments have verified that the proposed method significantly outperforms several state-of-the-art domain adaptation methods on a synthetic dataset and three different real world cross-domain visual recognition tasks.
Jing Zhang 0017, Wanqing Li 0001, Philip Ogunbona
CVPR3
2017 Weakly structured information aggregation for upper-body posture assessment using ConvNets
abstract
Posture assessment aims to determine the risk associated with poor posture and thus avoid injury in subjects. Upper-body posture assessment from images offers an attractive alternative to manual methods by directly extracting relevant features for classification. A deep convolutional neural network is proposed to extract structured features from different body parts and learn shared features that are used to determine the appropriate assessment. The structured features are learned with triplet-based rank constraints based on head and torso separately. The shared feature and assessment function are learned with soft-max constraints based on posture risk measurements. Experimental evaluation on a self-collected upper-body posture dataset has verified the efficacy of the proposed method and network architecture.
Zewei Ding, Wanqing Li 0001, Pichao Wang, Philip Ogunbona
ICME4
2017 Semantic action recognition by learning a pose lexicon
Lijuan Zhou 0002, Wanqing Li 0001, Philip Ogunbona, Zhengyou Zhang
Pattern Recognit.3
2016 Learning structured dictionary based on inter-class similarity and representative margins
abstract
We consider the problem of learning a structured and discriminative dictionary based on sparse representation for classification task. The structure comprises class-shared and class-specific partitions which allows the separation of common and class-specific information in the data for classification. The resulting optimization problem was a max margin formulation that exploits the hinge loss function property. Comparative evaluation of the proposed classifier against four recent alternatives in a gender classification task indicates a 3-percenatge point improvement.
Philip Ogunbona, Wanqing Li 0001, Gordon G. Wallace
ICASSP2
2016 Learning a pose lexicon for semantic action recognition
abstract
This paper presents a novel method for learning a pose lexicon comprising semantic poses defined by textual instructions and their associated visual poses defined by visual features. The proposed method simultaneously takes two input streams, semantic poses and visual pose candidates, and statistically learns a mapping between them to construct the lexicon. With the learned lexicon, action recognition can be cast as the problem of finding the maximum translation probability of a sequence of semantic poses given a stream of visual pose candidates. Experiments evaluating pre-trained and zero-shot action recognition conducted on MSRC-12 gesture and WorkoutSu-10 exercise datasets were used to verify the efficacy of the proposed method.
Lijuan Zhou 0002, Wanqing Li 0001, Philip Ogunbona
ICME3
2016 Large-scale Isolated Gesture Recognition using Convolutional Neural Networks
abstract
This paper proposes three simple, compact yet effective representations of depth sequences, referred to respectively as Dynamic Depth Images (DDI), Dynamic Depth Normal Images (DDNI) and Dynamic Depth Motion Normal Images (DDMNI). These dynamic images are constructed from a sequence of depth maps using bidirectional rank pooling to effectively capture the spatial-temporal information. Such image-based representations enable us to fine-tune the existing ConvNets models trained on image data for classification of depth sequences, without introducing large parameters to learn. Upon the proposed representations, a convolutional Neural networks (ConvNets) based method is developed for gesture recognition and evaluated on the Large-scale Isolated Gesture Recognition at the ChaLearn Looking at People (LAP) challenge 2016. The method achieved 55.57% classification accuracy and ranked 2ndplace in this challenge but was very close to the best performance even though we only used depth data.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Philip Ogunbona
ICPR6
2016 Large-scale Continuous Gesture Recognition Using Convolutional Neural Networks
abstract
This paper addresses the problem of continuous gesture recognition from sequences of depth maps using Convolutional Neural networks (ConvNets). The proposed method first segments individual gestures from a depth sequence based on quantity of movement (QOM). For each segmented gesture, an Improved Depth Motion Map (IDMM), which converts the depth sequence into one image, is constructed and fed to a ConvNet for recognition. The IDMM effectively encodes both spatial and temporal information and allows the fine-tuning with existing ConvNet models for classification without introducing millions of parameters to learn. The proposed method is evaluated on the Large-scale Continuous Gesture Recognition of the ChaLearn Looking at People (LAP) challenge 2016. It achieved the performance of 0.2655 (Mean Jaccard Index) and ranked 3rdplace in this challenge.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Philip Ogunbona
ICPR6
2016 Learning Discriminative Bayesian Networks from High-Dimensional Continuous Neuroimaging Data
abstract
Due to its causal semantics, Bayesian networks (BN) have been widely employed to discover the underlying data relationship in exploratory studies, such as brain research. Despite its success in modeling the probability distribution of variables, BN is naturally a generative model, which is not necessarily discriminative. This may cause the ignorance of subtle but critical network changes that are of investigation values across populations. In this paper, we propose to improve the discriminative power of BN models for continuous variables from two different perspectives. This brings two general discriminative learning frameworks for Gaussian Bayesian networks (GBN). In the first framework, we employ Fisher kernel to bridge the generative models of GBN and the discriminative classifiers of SVMs, and convert the GBN parameter learning to Fisher kernel learning via minimizing a generalization error bound of SVMs. In the second framework, we employ the max-margin criterion and build it directly upon GBN models to explicitly optimize the classification performance of the GBNs. The advantages and disadvantages of the two frameworks are discussed and experimentally compared. Both of them demonstrate strong power in learning discriminative parameters of GBNs for neuroimaging based brain network analysis, as well as maintaining reasonable representation capacity. The contributions of this paper also include a new Directed Acyclic Graph (DAG) constraint with theoretical guarantee to ensure the graph validity of GBN.
Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Human detection from images and videos: A survey
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona
Pattern Recognit.3
2016 RGB-D-based action recognition datasets: A survey
Jing Zhang 0017, Wanqing Li 0001, Philip Ogunbona, Pichao Wang, Chang Tang
Pattern Recognit.3
2016 Action Recognition From Depth Maps Using Deep Convolutional Neural Networks
abstract
This paper proposes a new method, i.e., weighted hierarchical depth motion maps (WHDMM) + three-channel deep convolutional neural networks (3ConvNets), for human action recognition from depth maps on small training datasets. Three strategies are developed to leverage the capability of ConvNets in mining discriminative features for recognition. First, different viewpoints are mimicked by rotating the 3-D points of the captured depth maps. This not only synthesizes more data, but also makes the trained ConvNets view-tolerant. Second, WHDMMs at several temporal scales are constructed to encode the spatiotemporal motion patterns of actions into 2-D spatial structures. The 2-D spatial structures are further enhanced for recognition by converting the WHDMMs into pseudocolor images. Finally, the three ConvNets are initialized with the models obtained from ImageNet and fine-tuned independently on the color-coded WHDMMs constructed in three orthogonal planes. The proposed algorithm was evaluated on the MSRAction3D, MSRAction3DExt, UTKinect-Action, and MSRDailyActivity3D datasets using cross-subject protocols. In addition, the method was evaluated on the large dataset constructed from the above datasets. The proposed method achieved 2-9% better results on most of the individual datasets. Furthermore, the proposed method maintained its performance on the large dataset, whereas the performance of existing methods decreased with the increased number of actions.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Jing Zhang 0017, Chang Tang, Philip Ogunbona
IEEE Trans. Hum. Mach. Syst.6
2015 ConvNets-Based Action Recognition from Depth Maps through Virtual Cameras and Pseudocoloring
abstract
In this paper, we propose to adopt ConvNets to recognize human actions from depth maps on relatively small datasets based on Depth Motion Maps (DMMs). In particular, three strategies are developed to effectively leverage the capability of ConvNets in mining discriminative features for recognition. Firstly, different viewpoints are mimicked by rotating virtual cameras around subject represented by the 3D points of the captured depth maps. This not only synthesizes more data from the captured ones, but also makes the trained ConvNets view-tolerant. Secondly, DMMs are constructed and further enhanced for recognition by encoding them into Pseudo-RGB images, turning the spatial-temporal motion patterns into textures and edges. Lastly, through transferring learning the models originally trained over ImageNet for image classification, the three ConvNets are trained independently on the color-coded DMMs constructed in three orthogonal planes. The proposed algorithm was extensively evaluated on MSRAction3D, MSRAction3DExt and UTKinect-Action datasets and achieved the state-of-the-art results on these datasets.
Pichao Wang, Wanqing Li 0001, Zhimin Gao, Chang Tang, Jing Zhang 0017, Philip Ogunbona
ACM Multimedia6
2014 Single Image Smoke Detection
Hongda Tian, Wanqing Li 0001, Philip Ogunbona, Lei Wang 0001
ACCV (2)3
2014 Discriminative Sparse Inverse Covariance Matrix: Application in Brain Functional Network Classification
abstract
Recent studies show that mental disorders change the functional organization of the brain, which could be investigated via various imaging techniques. Analyzing such changes is becoming critical as it could provide new biomarkers for diagnosing and monitoring the progression of the diseases. Functional connectivity analysis studies the covary activity of neuronal populations in different brain regions. The sparse inverse covariance estimation (SICE), also known as graphical LASSO, is one of the most important tools for functional connectivity analysis, which estimates the interregional partial correlations of the brain. Although being increasingly used for predicting mental disorders, SICE is basically a generative method that may not necessarily perform well on classifying neuroimaging data. In this paper, we propose a learning framework to effectively improve the discriminative power of SICEs by taking advantage of the samples in the opposite class. We formulate our objective as convex optimization problems for both one-class and two-class classifications. By analyzing these optimization problems, we not only solve them efficiently in their dual form, but also gain insights into this new learning framework. The proposed framework is applied to analyzing the brain metabolic covariant networks built upon FDG-PET images for the prediction of the Alzheimer's disease, and shows significant improvement of classification performance for both one-class and two-class scenarios. Moreover, as SICE is a general method for learning undirected Gaussian graphical models, this paper has broader meanings beyond the scope of brain research.
Luping Zhou, Lei Wang 0001, Philip Ogunbona
CVPR3
2014 Max-Margin Based Learning for Discriminative Bayesian Network from Neuroimaging Data
Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen
MICCAI (3)4
2014 Smoke Detection in Video: An Image Separation Approach
Hongda Tian, Wanqing Li 0001, Lei Wang 0001, Philip Ogunbona
Int. J. Comput. Vis.4
2014 Food image classification using local appearance and global structural information
Duc Thanh Nguyen, Zhimin Zong, Philip Ogunbona, Yasmine C. Probst, Wanqing Li 0001
Neurocomputing3
2013 Discriminative Brain Effective Connectivity Analysis for Alzheimer's Disease: A Kernel Learning Approach upon Sparse Gaussian Bayesian Network
abstract
Analyzing brain networks from neuroimages is becoming a promising approach in identifying novel connectivity-based biomarkers for the Alzheimer's disease (AD). In this regard, brain ``effective connectivity" analysis, which studies the causal relationship among brain regions, is highly challenging and of many research opportunities. Most of the existing works in this field use generative methods. Despite their success in data representation and other important merits, generative methods are not necessarily discriminative, which may cause the ignorance of subtle but critical disease-induced changes. In this paper, we propose a learning-based approach that integrates the benefits of generative and discriminative methods to recover effective connectivity. In particular, we employ Fisher kernel to bridge the generative models of sparse Bayesian networks (SBN) and the discriminative classifiers of SVMs, and convert the SBN parameter learning to Fisher kernel learning via minimizing a generalization error bound of SVMs. Our method is able to simultaneously boost the discriminative power of both the generative SBN models and the SBN-induced SVM classifiers via Fisher kernel. The proposed method is tested on analyzing brain effective connectivity for AD from ADNI data, and demonstrates significant improvements over the state-of-the-art work.
Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen
CVPR4
2013 Inter-occlusion reasoning for human detection based on variational mean field
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona
Neurocomputing3
2013 A novel shape-based non-redundant local binary pattern descriptor for object detection
Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001
Pattern Recognit.2
2013 Measuring the degree of face familiarity based on extended NMF
abstract
Getting familiar with a face is an important cognitive process in human perception of faces, but little study has been reported on how to objectively measure the degree of familiarity. In this article, a method is proposed to quantitatively measure the familiarity of a face with respect to a set of reference faces that have been seen previously. The proposed method models the context-free and context-dependent forms of familiarity suggested by psychological studies and accounts for the key factors, namely exposure frequency, exposure intensity and similar exposure, that affect human perception of face familiarity. Specifically, the method divides the reference set into nonexclusive groups and measures the familiarity of a given face by aggregating the similarities of the face to the individual groups. In addition, the nonnegative matrix factorization (NMF) is extended in this paper to learn a compact and localized subspace representation for measuring the similarities of the face with respect to the individual groups. The proposed method has been evaluated through experiments that follow the protocols commonly used in psychological studies and has been compared with subjective evaluation. Results have shown that the proposed measurement is highly consistent with the subjective judgment of face familiarity. Moreover, a face recognition method is devised using the concept of face familiarity and the results on the standard FERET evaluation protocols have further verified the efficacy of the proposed familiarity measurement.
Ce Zhan, Wanqing Li 0001, Philip Ogunbona
ACM Trans. Appl. Percept.3
2012 Private Fingerprint Matching
Siamak F. Shahandashti, Reihaneh Safavi-Naini, Philip Ogunbona
ACISP3
2012 A Novel Video-Based Smoke Detection Method Using Image Separation
abstract
In the state-of-the-art video-based smoke detection methods, the representation of smoke mainly depends on the visual information in the current image frame. In the case of light smoke, the original background can be still seen and may deteriorate the characterization of smoke. The core idea of this paper is to demonstrate the superiority of using smoke component for smoke detection. In order to obtain smoke component, a blended image model is constructed, which basically is a linear combination of background and smoke components. Smoke opacity which represents a weighting of the smoke component is also defined. Based on this model, an optimization problem is posed. An algorithm is devised to solve for smoke opacity and smoke component, given an input image and the background. The resulting smoke opacity and smoke component are then used to perform the smoke detection task. The experimental results on both synthesized and real image data verify the effectiveness of the proposed method.
Hongda Tian, Wanqing Li 0001, Lei Wang 0001, Philip Ogunbona
ICME4
2012 Measuring face familiarity and its application to face recognition
abstract
The familiarity of faces is one of the key factors that come into play during human face analysis. However, there is very little research that studies face familiarity. In this paper, two methods are proposed to quantitatively measure the degree of familiarity of a face with respect to a known set. The methods are in accordance with the psychological study. In particular, non-negative matrix factorization (NMF) is extended to learn a localized non-overlapping subspace representation of commonly experienced facial patterns from known faces. The familiarity of a given face is then measured based on its reconstruction error after being projected into the learned extended NMF subspaces. A subjective study involving 50 subjects indicates the proposed familiarity measurement is in line with human judgments. Furthermore, the familiarity vector generated during the measuring process is employed for face recognition. Experiments based on the standard FERET evaluation protocol demonstrates the efficacy of the familiarity based representation for face recognition.
Ce Zhan, Wanqing Li 0001, Philip Ogunbona
WACV3
2011 Detecting humans under occlusion using variational mean field method
abstract
This paper proposes a human detection method using variational mean field approximation for occlusion reasoning. In the method, parts of human objects are detected individually using template matching. Initial detection hypotheses with spatial layout information are represented in a graphical model and refined through a Bayesian estimation. In this paper, mean field method is employed for such an estimation. The proposed method was evaluated on the popular CAVIAR-INRIA dataset. Experimental results show that the proposed algorithm is able to detect humans in severe occlusion within reasonable processing time.
Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001
ICIP2
2011 Human detection with contour-based local motion binary patterns
abstract
This paper presents a human detection method using contour- based local motion features. The local motion is encoded using a variant of the popular Local Binary Pattern (LBP) called Non-Redundant Local Binary Pattern (NRLBP) descriptor computed on the difference image of two consecutive frames. In addition, the local motion features are extracted along the human's boundary contour. Localising features on the contours has the advantage of utilizing a precise human shape description. A motivation of the proposed method is that most of informative movements are performed on boundary contours of the body parts, e.g. legs of pedestrians. Evaluation of the proposed method was conducted on the INRIA and ETH datasets. Apart from showing the importance of motion information, experimental results also showed that localising features along the object boundary contours improves the detection performance.
Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001
ICIP2
2011 Smoke detection in videos using Non-Redundant Local Binary Pattern-based features
abstract
This paper presents a novel and low complexity method for real-time video-based smoke detection. As a local texture operator, Non-Redundant Local Binary Pattern (NRLBP) is more discriminative and robust to illumination changes in comparison with original Local Binary Pattern (LBP), thus is employed to encode the appearance information of smoke. Non-Redundant Local Motion Binary Pattern (NRLMBP), which is computed on the difference image of consecutive frames, is introduced to capture the motion information of smoke. Experimental results show that NRLBP outperforms the original LBP in the smoke detection task. Furthermore, the combination of NRLBP and NRLMBP, which can be considered as a spatial-temporal descriptor of smoke, can lead to remarkable improvement on detection performance.
Hongda Tian, Wanqing Li 0001, Philip Ogunbona, Duc Thanh Nguyen, Ce Zhan
MMSP3
2011 Age estimation based on extended non-negative matrix factorization
abstract
Previous studies suggested that local appearance-based methods are more efficient than geometric-based and holistic methods for age estimation. This is mainly due to the fact that age information are usually encoded by the local features such as wrinkles and skin texture on the forehead or at the eye corners. However, the variations of theses features caused by other factors such as identity, expression, pose and lighting may be larger than that caused by aging. Thus, one of the key challenges of age estimation lies in constructing a feature space that could successfully recovers age information while ignoring other sources of variations. In this paper, non-negative matrix factorization (NMF) is extended to learn a localized non-overlapping subspace representation for age estimation. To emphasize the appearance variation in aging, one individual extended NMF subspace is learned for each age or age group. The age or age group of a given face image is then estimated based on its reconstruction error after being projected into the learned age subspaces. Furthermore, a coarse to fine scheme is employed for exact age estimation, so that the age is estimated within the pre-classified age groups. Cross-database tests are conducted using FG-NET and MORPH databases to evaluate the proposed method. Experimental results have demonstrated the efficacy of the method.
Ce Zhan, Wanqing Li 0001, Philip Ogunbona
MMSP3
2010 Human detection using local shape and Non-Redundant binary patterns
abstract
Motivated by the advantages of using shape matching technique in detecting objects in various postures and viewpoints and the discriminative power of local patterns in object recognition, this paper proposes a human detection method combining both shape and appearance cues. In particular, local shapes of the body parts are detected using template matching. Based on body parts' shapes, local appearance features are extracted. We introduce a novel local binary pattern (LBP) descriptor, called Non-Redundant LBP (NRLBP), to encode local appearance of human. The proposed method was evaluated and compared with other state-of-the-art human detection methods on two commonly used datasets: MIT and INRIA pedestrian test sets. We also performed extensive experiments on selecting appropriate parameters as well as verifying the improvement of the proposed method through all stages of the framework.
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona
ICARCV3
2010 Finding distinctive facial areas for face recognition
abstract
One of the key issues for local appearance based face recognition methods is that how to find the most discriminative facial areas. Most of the existing methods take the assumption that anatomical facial components, such as the eyes, nose, and mouth, are the most useful areas for recognition. Other more elaborate methods locate the most salient parts within the face according to a pre-specified criterion. In this paper, a novel method is proposed to identify the discriminative facial areas for face recognition. Unlike the existing methods that only analyze the given face, the proposed method identifies the distinctive areas of each individual's face by its comparison to the general population. In particular, non-negative matrix factorization (NMF) is extended to learn a localized non-overlapping subspace representation of the facial patterns from a generic face image database. In the learned subspace, the degree of distinctiveness for any facial area is measured depends on the probability of this area is belong to a general face. For evaluation, the proposed method is tested on exaggerated face images and applied in exiting face recognition systems. Experimental results demonstrate the efficiency of the proposed method.
Ce Zhan, Wanqing Li 0001, Philip Ogunbona
ICARCV3
2010 Object detection using Non-Redundant Local Binary Patterns
abstract
Local Binary Pattern (LBP) as a descriptor, has been successfully used in various object recognition tasks because of its discriminative property and computational simplicity. In this paper a variant of the LBP referred to as Non-Redundant Local Binary Pattern (NRLBP) is introduced and its application for object detection is demonstrated. Compared with the original LBP descriptor, the NRLBP has advantage of providing a more compact description of object's appearance. Furthermore, the NRLBP is more discriminative since it reflects the relative contrast between the background and foreground. The proposed descriptor is employed to encode human's appearance in a human detection task. Experimental results show that the NRLBP is robust and adaptive with changes of the background and foreground and also outperforms the original LBP in detection task.
Duc Thanh Nguyen, Zhimin Zong, Philip Ogunbona, Wanqing Li 0001
ICIP3
2010 On the Combination of Local Texture and Global Structure for Food Classification
abstract
This paper proposes a food image classification method using local textural patterns and their global structure to describe the food image. In this paper, a visual codebook of local textural patterns is created by employing Scale Invariant Feature Transformation (SIFT) interest point detector with the Local Binary Pattern (LBP) feature. In addition to describing the food image using local texture, the global structure of the food object is represented as the spatial distribution of the local textural structures and encoded using shape context. We evaluated the proposed method on the Pittsburgh Fast-Food Image (PFI) dataset. Experimental results showed that the proposed method could obtain better performance than the baseline experiment on the PFI dataset.
Zhimin Zong, Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001
ISM3
2009 An Improved Template Matching Method for Object Detection
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona
ACCV (3)3
2009 Face detection using generalised integral image features
abstract
This paper proposes generalised integral image features (GIIFs) for face detection. GIIFs provide a richer and more flexible set of features than Haar-like features. Due to the large set of possible GIIFs, a genetic algorithm is developed to select the feature space for the optimal weak classifiers. Experimental results have shown that this method is able to improve face detection accuracy.
Alister Cordiner, Philip Ogunbona, Wanqing Li 0001
ICIP2
2009 A novel template matching method for human detection
abstract
This paper proposes a novel weighted template matching method. It employs a generalized distance transform (GDT) and an orientation map (OM). The GDT allows us to weight the distance transform more on the strong edge points and the OM provides supplementary local orientation information for matching. Based on the matching method, a two-stage human detection method consisting of template matching and Bayesian verification is developed. Experimental results have shown that the proposed method can effectively reduce the false positive and false negative detection rates and perform superiorly in comparison to the conventional Chamfer matching method.
Duc Thanh Nguyen, Wanqing Li 0001, Philip Ogunbona
ICIP3
2009 Human detection based on weighted template matching
abstract
This paper proposes a new two-stage human detection method involving matching and verification. A Bayesian framework is developed to verify the matching score obtained from a weighted distance measure. Performance evaluation indicates that the proposed method is able to utilize the flexible matching scheme and produce superior true positive, true negative and low misclassification rates.
Duc Thanh Nguyen, Philip Ogunbona, Wanqing Li 0001
ICME2
2008 An efficient iterative algorithm for image thresholding
Liju Dong, Ge Yu 0001, Philip Ogunbona, Wanqing Li 0001
Pattern Recognit. Lett.3
2007 A Maximum Likelihood Watermark Decoding Scheme
abstract
Based on the observation that an attack applied on a watermarked image, from a decoding point of view, modifies the distribution of the detection values away from the ideal distribution (without attack) for corresponding watermarking scheme, we propose a generic maximum likelihood decoding scheme by approximating the distribution with a finite Gaussian mixture model. The parameters of the model are estimated using expectation-maximization algorithm. The scheme allows the decoding to be automatically adapted to attacks that the watermarked images have undergone and, in consequence, to improve the decoding accuracy. Experiments on a QIM based watermarking system have clearly verified the significant improvement of the decoding accuracy achieved by the proposed maximum likelihood decoding in comparison to conventional threshold decoding.
Wenming Lu, Wanqing Li 0001, Reihaneh Safavi-Naini, Philip Ogunbona
ICME4
2006 A Pixel-Based Robust Imagewatermarking System
abstract
Robust image watermarking systems are required to be resistant to geometric attacks in addition to common image processing tasks, such as JPEG compression. However, robustness against geometric attacks, such as rotation, scaling and translation, still remains one of the most challenging research topics in image watermarking. We propose a new pixel-based watermarking system in which a binary logo is embedded, a bit per pixel, in the pixel domain of an image. The encoder of the proposed system is based on a sliding window embedding scheme that applies the local average quantization index modulation (QEM), to achieve geometric attack robustness. The decoder employs a maximum a posteriori (MAP) estimation supported by Markov random field (MRF) model to achieve robust decoding. Additionally, we demonstrate that the proposed scheme is also robust against possible watermark removal due to JPEG compression
Wenming Lu, Wanqing Li 0001, Reihaneh Safavi-Naini, Philip Ogunbona
ICME4
2006 Image Content Annotation Based on Visual Features
abstract
Automatic image content annotation techniques attempt to explore structural visual features of images that describe image content and associate them with image semantics. In this paper, two types of concept spaces, atomic concept and collective concept spaces, are defined and the annotation problems in those spaces are formulated as feature classification and Bayesian inference, respectively. A scheme of image content annotation in this framework is presented and evaluated as an application of photo categorization using MPEG-7 VCE2 dataset and its ground truth. The experimental results show a promising performance
Lei Ye 0002, Philip Ogunbona
ISM2
2006 Recovering DC Coefficients in Block-Based DCT
abstract
It is a common approach for JPEG and MPEG encryption systems to provide higher protection for dc coefficients and less protection for ac coefficients. Some authors have employed a cryptographic encryption algorithm for the dc coefficients and left the ac coefficients to techniques based on random permutation lists which are known to be weak against known-plaintext and chosen-ciphertext attacks. In this paper we show that in block-based DCT, it is possible to recover dc coefficients from ac coefficients with reasonable image quality and show the insecurity of image encryption methods which rely on the encryption of dc values using a cryptoalgorithm. The method proposed in this paper combines dc recovery from ac coefficients and the fact that ac coefficients can be recovered using a chosen ciphertext attack. We demonstrate that a method proposed by Tang to encrypt and decrypt MPEG video can be completely broken.
Takeyuki Uehara, Reihaneh Safavi-Naini, Philip Ogunbona
IEEE Trans. Image Process.3
2005 A New Divide and Conquer Algorithm for Graph-based Image and Video Segmentation
abstract
The concept of the shortest (or minimum) spanning tree (SST)and recursive SST (RSST) of an undirected weighted graph has been successfully applied in image segmentation and edge detection. This paper presents a divide-and-conquer approach for (R)SST based image segmentation in order to overcome the problem of high computational complexity associated with conventional graph algorithms. In the simplest form, the proposed approach, block-based RSST (BRSST), first divides the image into rectangular blocks, finds the (R)SST of each block individually using conventional graph algorithms and, then, merges the (R)SSTs of all image blocks to form an (R)SST of the entire image. Efficient merging algorithms are presented in this paper. We proved a theorem showing that the (R)SST obtained by the merging algorithms is one of the (R)SST that would be found by applying to the entire image the same algorithm used for finding the (R)SST of each image block. Theoretical analysis and experimental results have shown that BRSST has significantly reduced the computational cost. In addition, an incremental BRSST is proposed for video segmentation and results are presented
Wanqing Li 0001, Mingren Shi, Philip Ogunbona
MMSP3
2004 An MPEG tolerant authentication system for video data
abstract
We propose a secure video authentication algorithm that is tolerant to visual degradation due to MPEG lossy compression to a designed level. The authentication process generates a tag that is sent with video data and the level of protection can be adjusted so that longer tags are used for higher security, and that the protection is distributed such that higher security is provided for regions of interest in the image. The computation required for authentication and verification can be largely performed as part of MPEG compression and so generation and verification of the tag can be integrated into the compression system. Calculation of the tag can be parallelized and so made fast
Takeyuki Uehara, Reihaneh Safavi-Naini, Philip Ogunbona
ICME3
2004 A secure and flexible authentication system for digital images
Takeyuki Uehara, Reihaneh Safavi-Naini, Philip Ogunbona
Multim. Syst.3
2003 Stereoscopic panoramic video generation using centro-circular projection technique
abstract
The paper presents a method of stereoscopic panoramic video generation including techniques for panorama projection, stitching and calibration for various depth planes. The methods described can be used on video sequences captured by an arrangement of multiple pairs of cameras or multiple stereoscopic cameras mounted on a regular polygonal shaped camera rig. Algorithms can also be used in combination or separately, for generating both stereoscopic and monoscopic video and still panoramas.
Chaminda Weerasinghe, Wanqing Li 0001, Philip Ogunbona
ICASSP (3)3
2002 Modelling of color cross-talk in CMOS image sensors
abstract
This paper presents a way to model the cross-talk effect in CMOS image sensors. Two algorithms are derived from the model; both of them work on the Bayer raw data and have low computational complexity. Experiments on Macbeth color chart and real images have shown the effectiveness of the modeling to eliminate the cross-talk effect and produce better quality images with traditional color interpolation and correction algorithms designed for CCD image sensors.
Wanqing Li 0001, Philip Ogunbona, Igor Kharitonenko
ICASSP2
2002 Method of color interpolation in a single sensor color camera using green channel separation
abstract
This paper presents a color interpolation algorithm for a single sensor color camera. The proposed algorithm is especially designed to solve the problem of pixel crosstalk among the pixels of different color channels. Interchannel cross-talk gives rise to blocking effects on the interpolated green plane, and also spreading of false colors into detailed structures. The proposed algorithm separates the green channel into two planes, one highly correlated with the red channel and the other with the blue channel. These separate planes are used for red and blue channel interpolation. Experiments conducted on McBeth color chart and natural images have shown that the proposed algorithm can eliminate or suppress blocking and color artifacts to produce better quality images.
Chaminda Weerasinghe, Igor Kharitonenko, Philip Ogunbona
ICASSP3
2001 2D to pseudo-3D conversion of "head and shoulder" images using feature based parametric disparity maps
abstract
This paper presents a method of converting a 2D still photo containing the head & shoulders of a human (e.g. a passport photo) to pseudo-3D, so that the depth can be perceived via stereopsis. This technology has the potential to be included in self-serve photo booths and, also as an added accessory (i.e. software package) for digital still cameras and scanners. The basis of the algorithm is to exploit the ability of the human visual system in combining monoscopic and stereoscopic cues for depth perception. Common facial features are extracted from the 2D photograph, in order to create a parametric depth map that conforms to the available monoscopic depth cues. The original 2D photograph and the created depth map are used to generate left and right views for stereoscopic viewing. The algorithm is implemented in software, and promising results are obtained.
Chaminda Weerasinghe, Philip Ogunbona, Wanqing Li 0001
ICIP (3)2
2001 Signal analysis using a multiresolution form of the singular value decomposition
abstract
This paper proposes a multiresolution form of the singular value decomposition (SVD) and shows how it may be used for signal analysis and approximation. It is well-known that the SVD has optimal decorrelation and subrank approximation properties. The multiresolution form of SVD proposed here retains those properties, and moreover, has linear computational complexity. By using the multiresolution SVD, the following important characteristics of a signal may he measured, at each of several levels of resolution: isotropy, sphericity of principal components, self-similarity under scaling, and resolution of mean-squared error into meaningful components. Theoretical calculations are provided for simple statistical models to show what might be expected. Results are provided with real images to show the usefulness of the SVD decomposition.
Ramakrishna Kakarala, Philip Ogunbona
IEEE Trans. Image Process.2
1999 Index compressed tree-structured vector quantisation
Jamshid Shanbehzadeh, Philip Ogunbona
Signal Process. Image Commun.2
1997 On the computational complexity of the LBG and PNN algorithms
abstract
This correspondence compares the computational complexity of the pair-wise nearest neighbor (PNN) and Linde-Buzo-Gray (LBG) algorithms by deriving analytical expressions for their computational times. It is shown that for a practical codebook size and training vector sequence, the LBG algorithm is indeed more computationally efficient than the PNN algorithm.
Jamshid Shanbehzadeh, Philip Ogunbona
IEEE Trans. Image Process.2
1996 Index compressed image adaptive vector quantisation
Jamshid Shanbehzadeh, Philip Ogunbona
Signal Process. Image Commun.2
1994 Comparison of "wavelet" filters and subband analysis structures for still image compression
abstract
The paper compares several subband analysis structures and filters for image compression using a generic subband quantisation and encoding method. The authors show that filter banks designed on the basis of an image correlation model outperform other filter banks. The DCT, lapped orthogonal transform (LOT), a cosine modulated filter bank, and an octave band orthogonal filter bank (discrete wavelet transform-DWT) perform in a similar manner. Wavelet filters used to implement the DWT are designed to maximise a coding gain metric. These filters significantly outperform filters designed using other criteria. Daubechies wavelet filters exhibit a coding gain close to these optimum filters and perform similarly for image compression. Some interesting properties relating filter time-width and step response ringing to the location of zeros are observed. Finally a modified DWT structure is proposed and shown to give an improvement over the DWT for some images.>
James P. Andrew, Philip Ogunbona, Frank John Paoloni
ICASSP (5)2
1994 Coding Gain and Spatial Localisation Properties of Discrete Wavelet Transform Filter Banks for Image Coding
abstract
We consider coding gain and spatial localisation properties of DWT filters for still image compression. A high coding gain, relative to a highly correlated image source model, ensures that slowly varying image areas are coded efficiently. A low spatial width ensures that anomalies such as edges can also be coded efficiently. By considering an image model we design filters that simultaneously exhibit a high coding gain and low spatial width. Several existing filters are also examined. Several filter sets are then compared for image compression using a DWT and a JPEG type quantisation method. The filters with the best spatial width and coding gain properties outperform the other filters at low bit rates, especially with regard to ringing distortion.>
James P. Andrew, Philip Ogunbona, Frank John Paoloni
ICIP (3)2