Eng-Jon Ong

dblp:16/1335 · DBLP profile ↗
← Back
29ranked-venue papers
16as first author
2since 2021 · last 2024
0000-0002-7907-6352ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 14 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 10 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Artificial intelligence
6 papers
Video understanding and tracking · 66% Face, body and person analysis · 19% Probabilistic and Bayesian machine learning · 11%
Computer graphics and multimedia
3 papers
Computer animation and physical simulation · 54% Visualization and visual analytics · 18% Image and video processing · 14%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 18 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
image retrieval
0.512021
ACTNET: End-to-End Learning of Feature Activations and Multi-stream Aggregation for Effective Instance Image Retrieval · Int. J. Comput. Vis. 2021
Information retrieval › image retrieval
instance retrieval
0.512021
ACTNET: End-to-End Learning of Feature Activations and Multi-stream Aggregation for Effective Instance Image Retrieval · Int. J. Comput. Vis. 2021
Computer vision › Video understanding and tracking
sign language recognition
0.532014
Sign Spotting Using Hierarchical Sequential Patterns with Temporal Intervals · CVPR 2014
Sign language recognition using sub-units · J. Mach. Learn. Res. 2012
Sign Language Recognition using Sequential Pattern Trees · CVPR 2012
Information retrieval › similarity search › metric space similarity search
hamming space search
0.212016
Improved Hamming Distance Search Using Variable Length Hashing · CVPR 2016
Information retrieval
similarity search
0.212016
Improved Hamming Distance Search Using Variable Length Hashing · CVPR 2016
Algorithms and data structures › data structure design › search structures
hashing
0.212016
Improved Hamming Distance Search Using Variable Length Hashing · CVPR 2016
Computer vision › Video understanding and tracking › sign language recognition
sign spotting
0.212014
Sign Spotting Using Hierarchical Sequential Patterns with Temporal Intervals · CVPR 2014
Computer vision › Video understanding and tracking
gesture recognition
0.112012
Sign language recognition using sub-units · J. Mach. Learn. Res. 2012
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
sequential classification
0.112012
Sign Language Recognition using Sequential Pattern Trees · CVPR 2012
Computer vision › Face, body and person analysis › face tracking
facial feature tracking
0.112011
Robust Facial Feature Tracking Using Shape-Constrained Multiresolution-Selected Linear Predictors · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Visualization and visual analytics › flow visualization
feature tracking
0.112011
Robust Facial Feature Tracking Using Shape-Constrained Multiresolution-Selected Linear Predictors · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer animation and physical simulation › motion control
interactive motion control
0.112011
MIMiC: Multimodal Interactive Motion Controller · IEEE Trans. Multim. 2011
Computer animation and physical simulation
motion synthesis
0.112011
MIMiC: Multimodal Interactive Motion Controller · IEEE Trans. Multim. 2011
Computer animation and physical simulation › motion synthesis › motion composition
motion transition generation
0.112011
MIMiC: Multimodal Interactive Motion Controller · IEEE Trans. Multim. 2011
Multimedia analysis and retrieval
object tracking
0.112009
Robust facial feature tracking using selected multi-resolution linear predictors · ICCV 2009
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
0.112006
Real-Time Upper Body Detection and 3D Pose Estimation in Monoscopic Images · ECCV (3) 2006
Computer vision › Face, body and person analysis
human pose estimation
0.112006
Real-Time Upper Body Detection and 3D Pose Estimation in Monoscopic Images · ECCV (3) 2006
Haptics and multimodal interaction
multimodal interaction
0.012011
MIMiC: Multimodal Interactive Motion Controller · IEEE Trans. Multim. 2011

Methods — techniques the papers use, named apart from their topics

multi-stream aggregation · 1.0variable length hashing · 0.5branch-and-bound · 0.5biased linear predictor · 0.3kernel density estimation · 0.2k-medoids clustering · 0.2dimensionality reduction · 0.2probabilistic selection · 0.2sequential interval pattern · 0.2hierarchical tree classifier · 0.2ensemble of trees · 0.2sub-unit modeling · 0.1sequential pattern tree · 0.1markov model · 0.1hidden markov model · 0.1discriminative multi-class classifier · 0.1shape-constrained multiresolution linear predictors · 0.1multi-resolution model · 0.1
YearPublicationVenuePosition
2024 Understanding the Distributions of Aggregation Layers in Deep Neural Networks
abstract
The process of aggregation is ubiquitous in almost all the deep nets' models. It functions as an important mechanism for consolidating deep features into a more compact representation while increasing the robustness to overfitting and providing spatial invariance in deep nets. In particular, the proximity of global aggregation layers to the output layers of DNNs means that aggregated features directly influence the performance of a deep net. A better understanding of this relationship can be obtained using information theoretic methods. However, this requires knowledge of the distributions of the activations of aggregation layers. To achieve this, we propose a novel mathematical formulation for analytically modeling the probability distributions of output values of layers involved with deep feature aggregation. An important outcome is our ability to analytically predict the Kullback-Leibler (KL)-divergence of output nodes in a DNN. We also experimentally verify our theoretical predictions against empirical observations across a broad range of different classification tasks and datasets.
Eng-Jon Ong, Syed Sameed Husain, Miroslaw Bober
IEEE Trans. Neural Networks Learn. Syst.1
2021 ACTNET: End-to-End Learning of Feature Activations and Multi-stream Aggregation for Effective Instance Image Retrieval
Syed Sameed Husain, Eng-Jon Ong, Miroslaw Bober
Int. J. Comput. Vis.2
2019 Deep Architectures and Ensembles for Semantic Video Classification
abstract
This paper addresses the problem of accurate semantic labeling of short videos. To this end, a multitude of three different deep nets, ranging from traditional recurrent neural 4 networks (LSTM, GRU), temporal agnostic networks (FV, VLAD, BoW), fully connected neural networks mid-stage AV fusion, and others were considered. Additionally, we also propose a residual architecture-based deep neural network (DNN) for video classification, with state-of-the-art classification performance at significantly reduced complexity. Furthermore, we propose four new approaches to diversity-driven multi-net ensembling, one based on fast correlation measure and three incorporating a DNN-based combiner. We show that significant performance gains can be achieved by ensembling diverse nets and we investigate factors contributing to high diversity. Based on the extensive YouTube8M dataset, we provide an in-depth evaluation and analysis of their behavior. We show that the performance of the ensemble is state-of-the-art achieving the highest accuracy on the YouTube8M Kaggle test data. The performance of the ensemble of classifiers was also evaluated on the HMDB51 and UCF101 datasets, and show that the resulting method achieves comparable accuracy with the state-of-the-art methods using similar input features.
Eng-Jon Ong, Syed Sameed Husain, Mikel Bober-Irizar, Miroslaw Bober
IEEE Trans. Circuits Syst. Video Technol.1
2016 Improved Hamming Distance Search Using Variable Length Hashing
abstract
This paper addresses the problem of ultra-large-scale search in Hamming spaces. There has been considerable research on generating compact binary codes in vision, for example for visual search tasks. However the issue of efficient searching through huge sets of binary codes remains largely unsolved. To this end, we propose a novel, unsupervised approach to thresholded search in Hamming space, supporting long codes (e.g. 512-bits) with a wide-range of Hamming distance radii. Our method is capable of working efficiently with billions of codes delivering between one to three orders of magnitude acceleration, as compared to prior art. This is achieved by relaxing the equal-size constraint in the Multi-Index Hashing approach, leading to multiple hash-tables with variable length hash-keys. Based on the theoretical analysis of the retrieval probabilities of multiple hash-tables we propose a novel search algorithm for obtaining a suitable set of hash-key lengths. The resulting retrieval mechanism is shown empirically to improve the efficiency over the state-of-the-art, across a range of datasets, bit-depths and retrieval thresholds.
Eng-Jon Ong, Miroslaw Bober
CVPR1
2014 Sign Spotting Using Hierarchical Sequential Patterns with Temporal Intervals
abstract
This paper tackles the problem of spotting a set of signs occuring in videos with sequences of signs. To achieve this, we propose to model the spatio-temporal signatures of a sign using an extension of sequential patterns that contain temporal intervals called Sequential Interval Patterns (SIP). We then propose a novel multi-class classifier that organises different sequential interval patterns in a hierarchical tree structure called a Hierarchical SIP Tree (HSP-Tree). This allows one to exploit any subsequence sharing that exists between different SIPs of different classes. Multiple trees are then combined together into a forest of HSP-Trees resulting in a strong classifier that can be used to spot signs. We then show how the HSP-Forest can be used to spot sequences of signs that occur in an input video. We have evaluated the method on both concatenated sequences of isolated signs and continuous sign sequences. We also show that the proposed method is superior in robustness and accuracy to a state of the art sign recogniser when applied to spotting a sequence of signs.
Eng-Jon Ong, Nicolas Pugeault, Richard Bowden
CVPR1
2012 Sign Language Recognition using Sequential Pattern Trees
abstract
This paper presents a novel, discriminative, multi-class classifier based on Sequential Pattern Trees. It is efficient to learn, compared to other Sequential Pattern methods, and scalable for use with large classifier banks. For these reasons it is well suited to Sign Language Recognition. Using deterministic robust features based on hand trajectories, sign level classifiers are built from sub-units. Results are presented both on a large lexicon single signer data set and a multi-signer Kinect™ data set. In both cases it is shown to out perform the non-discriminative Markov model approach and be equivalent to previous, more costly, Sequential Pattern (SP) techniques.
Eng-Jon Ong, Helen Cooper, Nicolas Pugeault, Richard Bowden
CVPR1
2012 Sign language recognition using sub-units
Helen Cooper, Eng-Jon Ong, Nicolas Pugeault, Richard Bowden
J. Mach. Learn. Res.2
2011 Learning Sequential Patterns for Lipreading
abstract
This paper proposes a novel machine learning algorithm (SP-Boosting) to tackle the problem of lipreading by building visual sequence classifiers based on sequential patterns. We show that an exhaustive search of optimal sequential patterns is not possible due to the immense search space, and tackle this with a novel, efficient tree-search method with a set of pruning criteria. Crucially, the pruning strategies preserve our ability to locate the optimal sequential pattern. Additionally, the tree-based search method accounts for the training set’s boosting weight distribution. This temporal search method is then integrated into the boosting framework resulting in the SP-Boosting algorithm. We also propose a novel constrained set of strong classifiers that further improves recognition accuracy. The resulting learnt classifiers are applied to lipreading by performing multi-class recognition on the OuluVS database. Experimental results show that our method achieves state of the art recognition performane, using only a small set of sequential patterns.
Eng-Jon Ong, Richard Bowden
BMVC1
2011 Visualisation and prediction of conversation interest through mined social signals
abstract
This paper introduces a novel approach to social behaviour recognition governed by the exchange of non-verbal cues between people. We conduct experiments to try and deduce distinct rules that dictate the social dynamics of people in a conversation, and utilise semi-supervised computer vision techniques to extract their social signals such as laughing and nodding. Data mining is used to deduce frequently occurring patterns of social trends between a speaker and listener in both interested and not interested social scenarios. The confidence values from rules are utilised to build a Social Dynamic Model (SDM), that can then be used for classification and visualisation. By visualising the rules generated in the SDM, we can analyse distinct social trends between an interested and not interested listener in a conversation. Results show that these distinctions can be applied generally and used to accurately predict conversational interest.
Dumebi Okwechime, Eng-Jon Ong, Andrew Gilbert, Richard Bowden
FG2
2011 Robust Facial Feature Tracking Using Shape-Constrained Multiresolution-Selected Linear Predictors
abstract
This paper proposes a learned data-driven approach for accurate, real-time tracking of facial features using only intensity information. The task of automatic facial feature tracking is nontrivial since the face is a highly deformable object with large textural variations and motion in certain regions. Existing works attempt to address these problems by either limiting themselves to tracking feature points with strong and unique visual cues (e.g., mouth and eye corners) or by incorporating a priori information that needs to be manually designed (e.g., selecting points for a shape model). The framework proposed here largely avoids the need for such restrictions by automatically identifying the optimal visual support required for tracking a single facial feature point. This automatic identification of the visual context required for tracking allows the proposed method to potentially track any point on the face. Tracking is achieved via linear predictors which provide a fast and effective method for mapping pixel intensities into tracked feature position displacements. Building upon the simplicity and strengths of linear predictors, a more robust biased linear predictor is introduced. Multiple linear predictors are then grouped into a rigid flock to further increase robustness. To improve tracking accuracy, a novel probabilistic selection method is used to identify relevant visual areas for tracking a feature point. These selected flocks are then combined into a hierarchical multiresolution LP model. Finally, we also exploit a simple shape constraint for correcting the occasional tracking failure of a minority of feature points. Experimental results show that this method performs more robustly and accurately than AAMs, with minimal training examples on example sequences that range from SD quality to Youtube quality. Additionally, an analysis of the visual support consistency across different subjects is also provided.
Eng-Jon Ong, Richard Bowden
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 MIMiC: Multimodal Interactive Motion Controller
abstract
We introduce a new algorithm for real-time interactive motion control and demonstrate its application to motion captured data, prerecorded videos, and HCI. Firstly, a data set of frames are projected into a lower dimensional space. An appearance model is learnt using a multivariate probability distribution. A novel approach to determining transition points is presented based on k-medoids, whereby appropriate points of intersection in the motion trajectory are derived as cluster centers. These points are used to segment the data into smaller subsequences. A transition matrix combined with a kernel density estimation is used to determine suitable transitions between the subsequences to develop novel motion. To facilitate real-time interactive control, conditional probabilities are used to derive motion given user commands. The user commands can come from any modality including auditory, touch, and gesture. The system is also extended to HCI using audio signals of speech in a conversation to trigger nonverbal responses from a synthetic listener in real-time. We demonstrate the flexibility of the model by presenting results ranging from data sets composed of vectorized images, 2-D, and 3-D point representations. Results show real-time interaction and plausible motion generation between different types of movement.
Dumebi Okwechime, Eng-Jon Ong, Richard Bowden
IEEE Trans. Multim.2
2010 Social Interactive Human Video Synthesis
Dumebi Okwechime, Eng-Jon Ong, Andrew Gilbert, Richard Bowden
ACCV (1)2
2009 Robust facial feature tracking using selected multi-resolution linear predictors
abstract
This paper proposes a learnt data-driven approach for accurate, real-time tracking of facial features using only intensity information. Constraints such as a-priori shape models or temporal models for dynamics are not required or used. Tracking facial features simply becomes the independent tracking of a set of points on the face. This allows us to cope with facial configurations not present in the training data. Tracking is achieved via linear predictors which provide a fast and effective method for mapping pixel-level information to tracked feature position displacements. To improve on this, a novel and robust biased linear predictor is proposed in this paper. Multiple linear predictors are grouped into a rigid flock to increase robustness. To further improve tracking accuracy, a novel probabilistic selection method is used to identify relevant visual areas for tracking a feature point. These selected flocks are then combined into a hierarchical multi-resolution LP model. Experimental results also show that this method performs more robustly and accurately than AAMs, without any a priori shape information and with minimal training examples.
Eng-Jon Ong, Yuxuan Lan, Barry-John Theobald, Richard W. Harvey, Richard Bowden
ICCV1
2009 Problem solving through imitation
Eng-Jon Ong, Liam F. Ellis, Richard Bowden
Image Vis. Comput.1
2006 Learning Distances for Arbitrary Visual Features
abstract
This paper presents a method for learning distance functions of arbitrary feature representations that is based on the concept of wormholes.We introduce wormholes and describe how it provides a method for warping the topology of visual representation spaces such that a meaningful distance between examples is available.Additionally, we show how a more general distance function can be learnt through the combination of many wormholes via an inter-wormhole network.We then demonstrate the application of the distance learning method on a variety of problems including nonlinear synthetic data, face illumination detection and the retrieval of images containing natural landscapes and man-made objects (e.g.cities).
Eng-Jon Ong, Richard Bowden
BMVC1
2006 Real-Time Upper Body Detection and 3D Pose Estimation in Monoscopic Images
Antonio S. Micilotta, Eng-Jon Ong, Richard Bowden
ECCV (3)2
2006 Learnt inverse kinematics for animation synthesis
Eng-Jon Ong, Adrian Hilton 0001
Graph. Model.1
2006 Viewpoint invariant exemplar-based 3D human tracking
Eng-Jon Ong, Antonio S. Micilotta, Richard Bowden, Adrian Hilton 0001
Comput. Vis. Image Underst.1
2005 Detection and Tracking of Humans by Probabilistic Body Part Assembly
abstract
This paper presents a probabilistic framework of assembling detected human body parts into a full 2D human configuration. The face, torso, legs and hands are detected in cluttered scenes using boosted body part detectors trained by AdaBoost. Body configurations are assembled from the detected parts using RANSAC, and a coarse heuristic is applied to eliminate obvious outliers. An a priori mixture model of upper-body configurations is used to provide a pose likelihood for each configuration. A joint-likelihood model is then determined by combining the pose, part detector and corresponding skin model likelihoods. The assembly with the highest likelihood is selected by RANSAC, and the elbow positions are inferred. This paper also illustrates the combination of skin colour likelihood and detection likelihood to further reduce false hand and face detections. 1
Antonio S. Micilotta, Eng-Jon Ong, Richard Bowden
BMVC2
2005 Learning multi-kernel distance functions using relative comparisons
Eng-Jon Ong, Richard Bowden
Pattern Recognit.1
2004 Minimal Training, Large Lexicon, Unconstrained Sign Language Recognition
abstract
This paper presents a flexible monocular system capable of recognising sign lexicons far greater in number than previous approaches. The power of the system is due to four key elements: (i) Head and hand detection based upon boosting which removes the need for temperamental colour segmentation; (ii) A body centred description of activity which overcomes issues with camera placement, calibration and user; (iii) A two stage classification in which stage I generates a high level linguistic description of activity which naturally generalises and hence reduces training; (iv) A stage II classifier bank which does not require HMMs, further reducing training requirements. The outcome of which is a system capable of running in real-time, and generating extremely high recognition rates for large lexicons with as little as a single training instance per sign. We demonstrate classification rates as high as 92 % for a lexicon of 164 words with extremely low training requirements outperforming previous approaches where thousands of training examples are required. 1
Timor Kadir, Richard Bowden, Eng-Jon Ong, Andrew Zisserman
BMVC3
2002 The dynamics of linear combinations: tracking 3D skeletons of human subjects
Eng-Jon Ong, Shaogang Gong
Image Vis. Comput.1
2001 Face distributions in similarity space under varying head pose
Jamie Sherrah, Shaogang Gong, Eng-Jon Ong
Image Vis. Comput.3
2000 Tracking Multiple People Under Occlusion Using Multiple Cameras
abstract
We describe a system for tracking multiple people with multiple cameras based on fusion of multiple cues. Face trackers are used to self-calibrate our system. Epipolar geometry and landmarks are employed to disambiguate the tracking problem. The correlation of visual information between different cameras is learnt using Support Vector Regression and Hierarchical Principal Component Analysis to estimate the subject appearance across cameras. The joint features of subjects extracted from multiple cameras are tracked and used as a model to re-track people once the subjects are lost tracking in the system. Results demonstrate that our system can deal with the occlusion. 1
Ting-Hsun Chang, Shaogang Gong, Eng-Jon Ong
BMVC3
2000 Quantifying Ambiguities in Inferring Vector-Based 3D Models
abstract
This paper presents a framework for directly addressing issues arising from self-occlusions and ambiguities due to the lack of depth information in vector-based representations. Visual data directly observed from an image are used to indirectly recover the parameters of an underlying dynamic model of an articulated object. The proposed framework allows us to learn the ambiguities of a representation from training examples. The resulting model is then used to measure the ambiguities of each estimated underlying model parameter given the available visual information. This provides an indication of how much we can “trust ” the visual data for estimating certain parts of the model. We then provide a working example of multi-view data fusion for tracking 3D skeletons of articulated objects in a multi-camera environment. 1
Eng-Jon Ong, Shaogang Gong
BMVC1
1999 A Dynamic 3D Human Model using Hybrid 2D-3D Representations in Hierarchical PCA Space
abstract
We propose a novel framework for a hybrid 2D-3D dynamic human model with which robust matching and tracking of a 3D skeleton model of a human body among multiple views can be performed. We describe a method that measures the image ambiguity at each view. The 3D skeleton model and the correspondence between the model and its 2D images are learnt using hierarchical principal component analysis. Tracking in individual views is performed based on CONDENSATION.
Eng-Jon Ong, Shaogang Gong
BMVC1
1999 Understanding Pose Discrimination in Similarity Space
abstract
Identity-independent estimation of head pose from prototype im-ages is a perplexing task, requiring pose-invariant face detection. The problem is exacerbated by changes in illumination, identity and facial position. Facial images must be transformed in such a way as to em-phasise dierences in pose, while suppressing dierences in identity. We investigate appropriate transformations for use with a similarity-to-prototypes philosophy. The results show that orientation-selective Gabor lters enhance dierences in pose, and that dierent lter ori-entations are optimal at dierent poses. In contrast, PCA was found to provide an identity-invariant representation in which similarities can be calculated more robustly. We also investigate the angular resolution at which pose changes can be resolved using our methods. An angular resolution of 10 was found to be suciently discriminable at some poses but not at others, while 20 is quite acceptable at most poses. 1
Jamie Sherrah, Shaogang Gong, Eng-Jon Ong
BMVC3
1998 Appearance-Based Face Recognition under Large Head Rotations in Depth
Shaogang Gong, Eng-Jon Ong, Peter J. Loft
ACCV (2)2
1998 Learning to Associate Faces across Views in Vector Space of Similarities to Prototypes
abstract
We present a method for learning appearance models that can be used to recognise and track both 3D head pose and identities of novel subjects with continuous head movement across the view-sphere. We describe an automatic face data acquisition system based on a magnetic sensor and a calibrated camera. The system enabled us to obtain systematically a database of face images with labelled 3D poses across a view-sphere of \\Sigma90 ffi yaw and \\Sigma30 ffi tilt at intervals of 10 ffi . The database was used to learn appearance models of unseen faces based on similarity measures to prototype faces. The method is computationally efficient and enables real-time performance with ease. 1 Introduction To be able to recognise faces of moving people not only requires the ability to label novel face images with known identities, but also needs detecting and tracking of faces over time [1]. We refer to this as the task of associating faces. We adopt the view such a task can be better achieved...
Shaogang Gong, Eng-Jon Ong, Stephen J. McKenna
BMVC2