Du Q. Huynh

dblp:35/2599 · also Du Huynh · DBLP profile ↗
← Back
52ranked-venue papers
12as first author
4since 2021 · last 2025
0000-0003-3080-9655ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 9 first-authorArtificial intelligence and machine learning · 30 · 12 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Video understanding and tracking · 59% 3D vision · 16% Face, body and person analysis · 13%
Databases, data mining, and information retrieval
3 papers
Information retrieval · 51% Recommender systems · 47% Indexing and storage engines · 2%

Topics — the 30 heaviest of 41, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
evaluation
0.912025
#REval: A Semantic Evaluation Framework for Hashtag Recommendation · IEEE Trans. Knowl. Data Eng. 2025
Recommender systems › tag recommendation
hashtag recommendation
0.912025
#REval: A Semantic Evaluation Framework for Hashtag Recommendation · IEEE Trans. Knowl. Data Eng. 2025
Computer vision › Video understanding and tracking
action recognition
0.832019
Hallucinating IDT Descriptors and I3D Optical Flow Features for Action Recognition With CNNs · ICCV 2019
Histogram of Oriented Principal Components for Cross-View Action Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2016
HOPC: Histogram of Oriented Principal Components of 3D Pointclouds for Action Recognition · ECCV (2) 2014
Computer vision › Video understanding and tracking › action recognition
human action recognition
0.412020
A Comparative Review of Recent Kinect-Based Action Recognition Algorithms · IEEE Trans. Image Process. 2020
Computer vision › Video understanding and tracking › action recognition
skeleton-based action recognition
0.412020
A Comparative Review of Recent Kinect-Based Action Recognition Algorithms · IEEE Trans. Image Process. 2020
Computer vision › Video understanding and tracking
multi-object tracking
0.312018
Multiple Pedestrian Tracking From Monocular Videos in an Interacting Multiple Model Framework · IEEE Trans. Image Process. 2018
Computer vision › Video understanding and tracking › object tracking
person tracking
0.312018
Multiple Pedestrian Tracking From Monocular Videos in an Interacting Multiple Model Framework · IEEE Trans. Image Process. 2018
Computer vision › Face, body and person analysis
face recognition
0.212016
Constrained Metric Learning by Permutation Inducing Isometries · IEEE Trans. Image Process. 2016
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.212016
Constrained Metric Learning by Permutation Inducing Isometries · IEEE Trans. Image Process. 2016
Computer vision › Image recognition and object detection › image classification
object classification
0.212016
Constrained Metric Learning by Permutation Inducing Isometries · IEEE Trans. Image Process. 2016
Computer vision › Face, body and person analysis
person re-identification
0.212016
Constrained Metric Learning by Permutation Inducing Isometries · IEEE Trans. Image Process. 2016
Computer vision › 3D vision
point cloud processing
0.212016
Histogram of Oriented Principal Components for Cross-View Action Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Video understanding and tracking
spatio-temporal features
0.212016
Histogram of Oriented Principal Components for Cross-View Action Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Video understanding and tracking › action recognition › robust action recognition
view-invariant action recognition
0.212016
Histogram of Oriented Principal Components for Cross-View Action Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Video understanding and tracking › action recognition › 3d action recognition
point cloud action recognition
0.212014
HOPC: Histogram of Oriented Principal Components of 3D Pointclouds for Action Recognition · ECCV (2) 2014
Computer vision › 3D vision › 3d human pose estimation
3d human pose tracking
0.212013
A Gaussian Process Guided Particle Filter for Tracking 3D Human Pose in Video · IEEE Trans. Image Process. 2013
Computer vision › Face, body and person analysis
human pose estimation
0.212013
A Gaussian Process Guided Particle Filter for Tracking 3D Human Pose in Video · IEEE Trans. Image Process. 2013
Computer vision › 3D vision
structure from motion
0.142003
Outlier Correction in Image Sequences for the Affine Camera · ICCV 2003
Outlier Detection in Video Sequences under Affine Projection · CVPR (1) 2001
Affine Reconstruction from Monocular Vision in the Presence of a Symmetry Plane · ICCV 1999
Multimedia analysis and retrieval › multimedia feature representation
video representation
0.112019
Hallucinating IDT Descriptors and I3D Optical Flow Features for Action Recognition With CNNs · ICCV 2019
Information retrieval › image retrieval
image indexing
0.122003
CMVF: A Novel Dimension Reduction Scheme for Efficient Indexing in A Large Image Database · SIGMOD Conference 2003
Combining multi-visual features for efficient indexing in a large image database · VLDB J. 2001
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › sequential monte carlo
particle filtering
0.012013
A Gaussian Process Guided Particle Filter for Tracking 3D Human Pose in Video · IEEE Trans. Image Process. 2013
Computer vision › 3D vision › structure from motion
factorization
0.012003
Outlier Correction in Image Sequences for the Affine Camera · ICCV 2003
Machine learning › Optimization for machine learning › low-rank optimization
robust matrix factorization
0.012003
Outlier Correction in Image Sequences for the Affine Camera · ICCV 2003
Indexing and storage engines › multidimensional indexing
dimensionality reduction for indexing
0.012003
CMVF: A Novel Dimension Reduction Scheme for Efficient Indexing in A Large Image Database · SIGMOD Conference 2003
Computer vision › 3D vision
camera calibration
0.021999
Affine Reconstruction from Monocular Vision in the Presence of a Symmetry Plane · ICCV 1999
Direct Methods for Self-Calibration of a Moving Stereo Head · ECCV (2) 1996
Computer vision › 3D vision › camera calibration
self-calibration
0.021998
Self-calibrating a Stereo Head: An Error Analysis in the Neighbourhood of Degenerate Configurations · ICCV 1998
Direct Methods for Self-Calibration of a Moving Stereo Head · ECCV (2) 1996
Computer vision › 3D vision › structure from motion › factorization
robust factorization
0.012001
Outlier Detection in Video Sequences under Affine Projection · CVPR (1) 2001
Computer vision › 3D vision › 3d reconstruction › projective reconstruction
affine reconstruction
0.011999
Affine Reconstruction from Monocular Vision in the Presence of a Symmetry Plane · ICCV 1999
Computer vision › 3D vision › range sensing
structured light
0.011999
Calibrating a Structured Light Stripe System: A Novel Approach · Int. J. Comput. Vis. 1999
Computer vision › 3D vision › 3d reconstruction › geometric reconstruction
symmetry-based reconstruction
0.011999
Affine Reconstruction from Monocular Vision in the Presence of a Symmetry Plane · ICCV 1999

Methods — techniques the papers use, named apart from their topics

word2vec · 0.9word embeddings · 0.9glove · 0.9fasttext · 0.9BERTag · 0.9BERT · 0.9improved dense trajectory · 0.8fisher vector · 0.8bag-of-words · 0.8CNN · 0.8histogram of oriented principal components · 0.4hand-crafted features · 0.4deep learning features · 0.4histogram of oriented gradients · 0.3data association · 0.3color histogram · 0.3dimension reduction · 0.0feature extraction · 0.0
YearPublicationVenuePosition
2025 #REval: A Semantic Evaluation Framework for Hashtag Recommendation
abstract
Automatic evaluation of hashtag recommendation models is a fundamental task in Twitter. In the traditional evaluation methods, the recommended hashtags from an algorithm are first compared with the ground truth hashtags for exact correspondences. The number of exact matches is then used to calculate the hit rate, hit ratio, precision, recall, or F1-score. This way of evaluating hashtag similarities is inadequate as it ignores the semantic correlation between the recommended and ground truth hashtags. To tackle this problem, we propose a novel semantic evaluation framework for hashtag recommendation, called #REval. This framework includes an internal module referred to asBERTag, which automatically learns the hashtag embeddings. We investigate on how the #REval framework performs under different word embedding methods and different numbers of synonyms and hashtags in the recommendation using our proposed #REval-hit-ratio measure. Our experiments of the proposed framework on three large datasets show that #REval gave more meaningful hashtag synonyms for hashtag recommendation evaluation. Our analysis also highlights the sensitivity of the framework to the word embedding technique, with #REval based on BERTag more superior over #REval based on Word2Vec, FastText, and GloVe.
Areej Alsini, Du Q. Huynh, Amitava Datta
IEEE Trans. Knowl. Data Eng.2
2024 Are Graph Embeddings the Panacea? - An Empirical Survey from the Data Fitness Perspective
Qiang Sun 0006, Du Q. Huynh, Mark Reynolds 0001, Wei Liu 0006
PAKDD (2)2
2021 A Vision-Based Pipeline for Vehicle Counting, Speed Estimation, and Classification
abstract
Cameras have been widely used in traffic operations. While many technologically smart camera solutions in the market can be integrated into Intelligent Transport Systems (ITS) for automated detection, monitoring and data generation, many Network Operations (a.k.a Traffic Control) Centres still use legacy camera systems as manual surveillance devices. In this paper, we demonstrate effective use of these older assets by applying computer vision techniques to extract traffic data from videos captured by legacy cameras. In our proposed vision-based pipeline, we adopt recent state-of-the-art object detectors and transfer-learning to detect vehicles, pedestrians, and cyclists from monocular videos. By weakly calibrating the camera, we demonstrate a novel application of the image-to-world homography which gives our monocular vision system the efficacy of counting vehicles by lane and estimating vehicle length and speed in real-world units. Our pipeline also includes a module which combines a convolutional neural network (CNN) classifier with projective geometry information to classify vehicles. We have tested it on videos captured at several sites with different traffic flow conditions and compared the results with the data collected by piezoelectric sensors. Our experimental results show that the proposed pipeline can process 60 frames per second for pre-recorded videos and yield high-quality metadata for further traffic analysis.
Chenghuan Liu, Du Q. Huynh, Yuchao Sun, Mark Reynolds 0001, Steve Atkinson
IEEE Trans. Intell. Transp. Syst.2
2021 PoPPL: Pedestrian Trajectory Prediction by LSTM With Automatic Route Class Clustering
abstract
Pedestrian path prediction is a very challenging problem because scenes are often crowded or contain obstacles. Existing state-of-the-art long short-term memory (LSTM)-based prediction methods have been mainly focused on analyzing the influence of other people in the neighborhood of each pedestrian while neglecting the role of potential destinations in determining a walking path. In this article, we propose classifying pedestrian trajectories into a number of route classes (RCs) and using them to describe the pedestrian movement patterns. Based on the RCs obtained from trajectory clustering, our algorithm, which we name the prediction of pedestrian paths by LSTM (PoPPL), predicts the destination regions through a bidirectional LSTM classification network in the first stage and then generates trajectories corresponding to the predicted destination regions through one of the three proposed LSTM-based architectures in the second stage. Our algorithm also outputs probabilities of multiple predicted trajectories that head toward the destination regions. We have evaluated PoPPL against other state-of-the-art methods on two public data sets. The results show that our algorithm outperforms other methods and incorporating potential destination prediction improves the trajectory prediction accuracy.
Hao Xue 0001, Du Q. Huynh, Mark Reynolds 0001
IEEE Trans. Neural Networks Learn. Syst.2
2020 Maximum Entropy Reinforced Single Object Visual Tracking
abstract
Single object visual tracking is a fundamental problem in computer vision and has many applications. Given only the location of the target of interest in the first video frame, a visual tracking algorithm must track the target until the end of the video while having to face challenging factors such as illumination change and scale variation. In this paper, we formulate this tracking problem in a framework of maximum entropy reinforcement learning where the agent is our visual tracker and the goal is to learn a tracking policy that maximises both the expected reward and its entropy so as to achieve a balance between exploitation and exploration. The aim of our tracking framework is to improve the tracking accuracy while giving the tracking agent the ability to avoid getting stuck on a nontarget object. Extensive experiments have been performed on a range of benchmarks where our method achieves state-of-the-art performance. Furthermore, we demonstrate that, in contrast to other visual trackers based on deep reinforcement learning, our method can run in real-time while maintaining high tracking accuracy.
Chenghuan Liu, Du Q. Huynh, Mark Reynolds 0001
ECAI2
2020 Take a NAP: Non-Autoregressive Prediction for Pedestrian Trajectories
Hao Xue 0001, Du Q. Huynh, Mark Reynolds 0001
ICONIP (1)2
2020 On Utilizing Communities Detected From Social Networks in Hashtag Recommendation
abstract
Personalized recommendation automatically predicts the top-y hashtags to a given tweet. Most research in the literature of hashtag recommendation focused on the content of the posts such as words and topics. Although these methods have measured the performance of hashtag recommendation on large data sets, there is a lack of analysis on how these methods perform on small communities. Motivated by the well-studied research area of community detection algorithms that aggregate strongly connected users with similar interests and behaviors, in this article, we propose a community-based hashtag recommendation framework, which studies hashtag recommendation through tweet similarity task and applies it on communities detected using the Clique percolation method, Louvain algorithm, and label propagation method. The detected communities are extracted from four social network constructions based on following, mention, hashtag, and topic. Compared to the three state-of-the-art hashtag recommendation methods, our extensive experiments show that our community-based method outperforms these methods, thus giving a higher hit rate. Our in-depth analysis demonstrates that the performance of hashtag recommendation is the best when the communities are generated using the Clique percolation method (CPM) from the network of users who share similar usage of hashtags.
Areej Alsini, Amitava Datta, Du Q. Huynh
IEEE Trans. Comput. Soc. Syst.3
2020 Toward Occlusion Handling in Visual Tracking via Probabilistic Finite State Machines
abstract
Visual tracking has been an active research area in computer vision for decades. However, the performance of existing techniques is still challenged by various factors, such as occlusion and change in appearance of the target. In this paper, we propose a novel framework based on correlation filtering and probabilistic finite state machines (FSMs) to handle occlusion. In our tracking framework, the target is partitioned into several parts whose occlusion states are automatically detected. A set of states for the target is defined in terms of the combination of the parts' occlusion states. The probabilistic FSMs are then used to model the target's state transitions so as to reduce the effect of noise in the output response maps of correlation filters. Our target model's update strategy is adaptable online depending on the estimated state of the target. Extensive experiments have been performed on several public benchmarks and the proposed algorithm achieves competitive results against state-of-the-art techniques.
Chenghuan Liu, Du Q. Huynh, Mark Reynolds 0001
IEEE Trans. Cybern.2
2020 A Comparative Review of Recent Kinect-Based Action Recognition Algorithms
abstract
Video-based human action recognition is currently one of the most active research areas in computer vision. Various research studies indicate that the performance of action recognition is highly dependent on the type of features being extracted and how the actions are represented. Since the release of the Kinect camera, a large number of Kinect-based human action recognition techniques have been proposed in the literature. However, there still does not exist a thorough comparison of these Kinect-based techniques under the grouping of feature types, such as handcrafted versus deep learning features and depth-based versus skeleton-based features. In this paper, we analyze and compare 10 recent Kinect-based algorithms for both cross-subject action recognition and cross-view action recognition using six benchmark datasets. In addition, we have implemented and improved some of these techniques and included their variants in the comparison. Our experiments show that the majority of methods perform better on cross-subject action recognition than cross-view action recognition, that the skeleton-based features are more robust for cross-view recognition than the depth-based features, and that the deep learning features are suitable for large datasets.
Lei Wang 0108, Du Q. Huynh, Piotr Koniusz
IEEE Trans. Image Process.2
2019 Hallucinating IDT Descriptors and I3D Optical Flow Features for Action Recognition With CNNs
abstract
In this paper, we revive the use of old-fashioned handcrafted video representations for action recognition and put new life into these techniques via a CNN-based hallucination step. Despite of the use of RGB and optical flow frames, the I3D model (amongst others) thrives on combining its output with the Improved Dense Trajectory (IDT) and extracted with its low-level video descriptors encoded via Bag-of-Words (BoW) and Fisher Vectors (FV). Such a fusion of CNNs and handcrafted representations is time-consuming due to pre-processing, descriptor extraction, encoding and tuning parameters. Thus, we propose an end-to-end trainable network with streams which learn the IDT-based BoW/FV representations at the training stage and are simple to integrate with the I3D model. Specifically, each stream takes I3D feature maps ahead of the last 1D conv. layer and learns to `translate' these maps to BoW/FV representations. Thus, our model can hallucinate and use such synthesized BoW/FV representations at the testing stage. We show that even features of the entire I3D optical flow stream can be hallucinated thus simplifying the pipeline. Our model saves 20-55h of computations and yields state-of-the-art results on four publicly available datasets.
Lei Wang 0108, Piotr Koniusz, Du Q. Huynh
ICCV3
2019 Loss Switching Fusion with Similarity Search for Video Classification
abstract
From video streaming to security and surveillance applications, video data play an important role in our daily living today. However, managing a large amount of video data and retrieving the most useful information for the user remain a challenging task. In this paper, we propose a novel video classification system that would benefit the scene understanding task. We define our classification problem as classifying background and foreground motions using the same feature representation for outdoor scenes. This means that the feature representation needs to be robust enough and adaptable to different classification tasks. We propose a lightweight Loss Switching Fusion Network (LSFNet) for the fusion of spatiotemporal descriptors and a similarity search scheme with soft voting to boost the classification performance. The proposed system has a variety of potential applications such as content-based video clustering, video filtering, etc. Evaluation results on two private industry datasets show that our system is robust in both classifying different background motions and detecting human motions from these background motions.
Lei Wang 0108, Du Q. Huynh, Moussa Reda Mansour
ICIP2
2019 Urban Area Vehicle Re-Identification With Self-Attention Stair Feature Fusion and Temporal Bayesian Re-Ranking
abstract
Vehicle re-identification (Re-ID) plays a key role in many smart traffic management systems. Re-identifying a vehicle can be very challenging because the differences in visual appearances between pairs of vehicles are sometimes extremely subtle if they have the same colour and the same model. Given an image of a vehicle, most existing techniques adopt a global feature representation where details may be ignored. In this paper, we propose an Self-Attention Stair Feature Fusion model to learn the discriminative features for vehicle Re-ID. The model is designed to extract multi-level features in order to capture as much small details as possible. We also propose a Temporal Bayesian Re-Ranking method to exploit the spatial-temporal information in the vehicles' travel patterns. Our algorithm has been tested against state-of-the-art techniques on popular benchmarks. The results show that our algorithm outperforms other state-of-the-art techniques by a large margin.
Chenghuan Liu, Du Q. Huynh, Mark Reynolds 0001
IJCNN2
2019 Pedestrian Trajectory Prediction Using a Social Pyramid
Hao Xue 0001, Du Q. Huynh, Mark Reynolds 0001
PRICAI (2)2
2019 Pedestrian Tracking and Stereo Matching of Tracklets for Autonomous Vehicles
abstract
The prediction of the surrounding pedestrians' walking paths is a vital part for autonomous driving systems in the aspect of traffic safety. In this paper, we propose a pipeline which tracks pedestrians captured by a stereo camera system onboard a mobile vehicle, composes the pedestrian tracklets, clusters the tracklets to form trajectories, and matches the trajectories. The output 3D pedestrian trajectories can be used for further applications such as pedestrian trajectory prediction for driverless vehicles. Our algorithm has been compared with various state-of-art pedestrian tracking methods. Our experimental results show that the visual temporal features computed by our algorithm are effective for trajectory representation and that, by incorporating tracklet clustering into the pipeline, the pedestrian tracking performance is improved.
Hao Xue 0001, Du Q. Huynh, Mark Reynolds 0001
VTC Spring2
2019 Location-Velocity Attention for Pedestrian Trajectory Prediction
abstract
Pedestrian path forecasting is crucial in applications such as smart video surveillance. It is a challenging task because of the complex crowd movement patterns in the scenes. Most of existing state-of-the-art LSTM based prediction methods require rich context like labelled static obstacles, labelled entrance/exit regions and even the background scene. Furthermore, incorporating contextual information into trajectory prediction increases the computational overhead and decreases the generalization of the prediction models across different scenes. In this paper, we propose a joint Location-Velocity Attention LSTM based method to predict trajectories. Specifically, a module is designed to tweak the LSTM network and an attention mechanism is trained to learn to optimally combine the location and the velocity information of pedestrians in the prediction process. We have evaluated our approach against other baselines and state-of-the-art methods on several publicly available datasets. The results show that it not only outperforms other prediction methods but it also has a good generalization ability.
Hao Xue 0001, Du Q. Huynh, Mark Reynolds 0001
WACV2
2018 SS-LSTM: A Hierarchical LSTM Model for Pedestrian Trajectory Prediction
abstract
Pedestrian trajectory prediction is an extremely challenging problem because of the crowdedness and clutter of the scenes. Previous deep learning LSTM-based approaches focus on the neighbourhood influence of pedestrians but ignore the scene layouts in pedestrian trajectory prediction. In this paper, a novel hierarchical LSTM-based network is proposed to consider both the influence of social neighbourhood and scene layouts. Our SS-LSTM, which stands for Social-Scene-LSTM, uses three different LSTMs to capture person, social and scene scale information. We also use a circular shape neighbourhood setting instead of the traditional rectangular shape neighbourhood in the social scale. We evaluate our proposed method against two baseline methods and a state-of-art technique on three public datasets. The results show that our method outperforms other methods and that using circular shape neighbourhood improves the prediction accuracy.
Hao Xue 0001, Du Q. Huynh, Mark Reynolds 0001
WACV2
2018 Multiple Pedestrian Tracking From Monocular Videos in an Interacting Multiple Model Framework
abstract
We present a multiple pedestrian tracking method for monocular videos captured by a fixed camera in an interacting multiple model (IMM) framework. Our tracking method involves multiple IMM trackers running in parallel, which are tied together by a robust data association component. We investigate two data association strategies which take into account both the target appearance and motion errors. We use a 4D color histogram as the appearance model for each pedestrian returned by a people detector that is based on the histogram of oriented gradients features. Short-term occlusion problems and false negative errors from the detector are dealt with using a sliding window of video frames, where tracking persists in the absence of observations. Our method has been evaluated, and compared both qualitatively and quantitatively with four state-of-the-art visual tracking methods using benchmark video databases. The experiments demonstrate that, on average, our tracking method outperforms these four methods.
Zhengqiang Jiang, Du Q. Huynh
IEEE Trans. Image Process.2
2017 Empirical Analysis of Factors Influencing Twitter Hashtag Recommendation on Detected Communities
Areej Alsini, Amitava Datta, Jianxin Li 0001, Du Q. Huynh
ADMA4
2016 Histogram of Oriented Principal Components for Cross-View Action Recognition
abstract
Existing techniques for 3D action recognition are sensitive to viewpoint variations because they extract features from depth images which are viewpoint dependent. In contrast, we directly process pointclouds for cross-view action recognition from unknown and unseen views. We propose the histogram of oriented principal components (HOPC) descriptor that is robust to noise, viewpoint, scale and action speed variations. At a 3D point, HOPC is computed by projecting the three scaled eigenvectors of the pointcloud within its local spatio-temporal support volume onto the vertices of a regular dodecahedron. HOPC is also used for the detection of spatio-temporal keypoints (STK) in 3D pointcloud sequences so that view-invariant STK descriptors (or Local HOPC descriptors) at these key locations only are used for action recognition. We also propose a global descriptor computed from the normalized spatio-temporal distribution of STKs in 4-D, which we refer to as STK-D. We have evaluated the performance of our proposed descriptors against nine existing techniques on two cross-view and three single-view human action recognition datasets. The experimental results show that our techniques provide significant improvement over state-of-the-art methods.
Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Discriminative human action classification using locality-constrained linear coding
Hossein Rahmani 0001, Du Q. Huynh, Arif Mahmood, Ajmal Mian
Pattern Recognit. Lett.2
2016 Constrained Metric Learning by Permutation Inducing Isometries
abstract
The choice of metric critically affects the performance of classification and clustering algorithms. Metric learning algorithms attempt to improve performance, by learning a more appropriate metric. Unfortunately, most of the current algorithms learn a distance function which is not invariant to rigid transformations of images. Therefore, the distances between two images and their rigidly transformed pair may differ, leading to inconsistent classification or clustering results. We propose to constrain the learned metric to be invariant to the geometry preserving transformations of images that induce permutations in the feature space. The constraint that these transformations are isometries of the metric ensures consistent results and improves accuracy. Our second contribution is a dimension reduction technique that is consistent with the isometry constraints. Our third contribution is the formulation of the isometry constrained logistic discriminant metric learning (IC-LDML) algorithm, by incorporating the isometry constraints within the objective function of the LDML algorithm. The proposed algorithm is compared with the existing techniques on the publicly available labeled faces in the wild, viewpoint-invariant pedestrian recognition, and Toy Cars data sets. The IC-LDML algorithm has outperformed existing techniques for the tasks of face recognition, person identification, and object classification by a significant margin.
Joel Bosveld, Arif Mahmood, Du Q. Huynh, Lyle Noakes
IEEE Trans. Image Process.3
2014 HOPC: Histogram of Oriented Principal Components of 3D Pointclouds for Action Recognition
Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian
ECCV (2)3
2014 Action Classification with Locality-Constrained Linear Coding
abstract
We propose an action classification algorithm which uses Locality-constrained Linear Coding (LLC) to capture discriminative information of human body variations in each spatio-temporal subsequence of a video sequence. Our proposed method divides the input video into equally spaced overlapping spatio-temporal sub sequences, each of which is decomposed into blocks and then cells. We use the Histogram of Oriented Gradient (HOG3D) feature to encode the information in each cell. We justify the use of LLC for encoding the block descriptor by demonstrating its superiority over Sparse Coding (SC). Our sequence descriptor is obtained via a logistic regression classifier with L2 regularization. We evaluate and compare our algorithm with ten state-of-the-art algorithms on five benchmark datasets. Experimental results show that, on average, our algorithm gives better accuracy than these ten algorithms.
Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian
ICPR3
2014 Real time action recognition using histograms of depth gradients and random decision forests
abstract
We propose an algorithm which combines the discriminative information from depth images as well as from 3D joint positions to achieve high action recognition accuracy. To avoid the suppression of subtle discriminative information and also to handle local occlusions, we compute a vector of many independent local features. Each feature encodes spatiotemporal variations of depth and depth gradients at a specific space-time location in the action volume. Moreover, we encode the dominant skeleton movements by computing a local 3D joint position difference histogram. For each joint, we compute a 3D space-time motion volume which we use as an importance indicator and incorporate in the feature vector for improved action discrimination. To retain only the discriminant features, we train a random decision forest (RDF). The proposed algorithm is evaluated on three standard datasets and compared with nine state-of-the-art algorithms. Experimental results show that, on the average, the proposed algorithm outperform all other algorithms in accuracy and have a processing speed of over 112 frames/second.
Hossein Rahmani 0001, Arif Mahmood, Du Q. Huynh, Ajmal Mian
WACV3
2013 Combining background subtraction and temporal persistency in pedestrian detection from static videos
abstract
This paper presents a method that incorporates background subtraction and temporal persistency with the HOG pedestrian detector to detect pedestrians from videos captured by a fixed camera. We use a codebook based method and interpolation to extract a series of foreground sub-images for the HOG detector. This allows the detector to focus on pedestrian detection on smaller image regions and thereby reduce its computational cost and lower its false positive error rate. We employ a temporal persistency constraint to overcome problems that may arise from background subtraction. Compared to other state-of-art techniques, the performance of our method is pleasantly promising.
Zhengqiang Jiang, Du Q. Huynh, William Moran 0001, Subhash Challa
ICIP2
2013 Discriminative fusion of shape and appearance features for human pose estimation
Suman Sedai, Mohammed Bennamoun, Du Q. Huynh
Pattern Recognit.3
2013 A Gaussian Process Guided Particle Filter for Tracking 3D Human Pose in Video
abstract
In this paper, we propose a hybrid method that combines Gaussian process learning, a particle filter, and annealing to track the 3D pose of a human subject in video sequences. Our approach, which we refer to as annealed Gaussian process guided particle filter, comprises two steps. In the training step, we use a supervised learning method to train a Gaussian process regressor that takes the silhouette descriptor as an input and produces multiple output poses modeled by a mixture of Gaussian distributions. In the tracking step, the output pose distributions from the Gaussian process regression are combined with the annealed particle filter to track the 3D pose in each frame of the video sequence. Our experiments show that the proposed method does not require initialization and does not lose tracking of the pose. We compare our approach with a standard annealed particle filter using the HumanEva-I dataset and with other state of the art approaches using the HumanEva-II dataset. The evaluation results show that our approach can successfully track the 3D human pose over long video sequences and give more accurate pose tracking results than the annealed particle filter.
Suman Sedai, Mohammed Bennamoun, Du Q. Huynh
IEEE Trans. Image Process.3
2011 Tracking pedestrians using smoothed colour histograms in an interacting multiple model framework
abstract
In this paper, we present a method for tracking pedestrians in video sequences captured by a fixed camera. Pedestrians are detected in every video frame using the human detector proposed by Dalal and Triggs. An interacting multiple model method is used to predict and update pedestrian trajectories from current frame to the next one. We employ a stationary model and a constant velocity model in our method to handle cases such as when a pedestrian suddenly stops or changes walking direction. We smooth the colour histogram that describes the appearance of each detected pedestrian using kernel density estimation. Our experimental results show that our tracking method outperforms one that uses the Kalman filter and colour histograms.
Zhengqiang Jiang, Du Q. Huynh, William Moran 0001, Subhash Challa
ICIP2
2011 Supervised particle filter for tracking 2D human pose in monocular video
abstract
In this paper, we propose a hybrid method that combines supervised learning and particle filtering to track the 2D pose of a human subject in monocular video sequences. Our approach, which we call a supervised particle filter method, consists of two steps: the training step and the tracking step. In the training step, we use a supervised learning method to train the regressors that take the silhouette descriptors as input and produce the 2D poses as output. In the tracking step, the output pose estimated from the regressors is combined with the particle filter to track the 2D pose in each video frame. Unlike the particle filter, our method does not require any manual initialization. We have tested our approach using the HumanEva video datasets and compared it with the standard particle filter and 2D pose estimation on individual frames. Our experimental results show that our approach can successfully track the pose over long video sequences and that it gives more accurate 2D human pose tracking than the particle filter and 2D pose estimation.
Suman Sedai, Du Q. Huynh, Mohammed Bennamoun
WACV2
2010 Localized fusion of Shape and Appearance features for 3D Human Pose Estimation
Suman Sedai, Mohammed Bennamoun, Du Q. Huynh
BMVC3
2010 Probabilistic human pose recovery from 2D images
abstract
Image based human pose recovery has many applications in different industries such as games, entertainment, physiological rehabilitation and biometrics. This paper presents a new pose estimation algorithm from monocular images based on a nonlinear mapping of human silhouettes, coded using a collection of local image moments, to the pose space using a mixture of Neural Networks (NN) regressors. All parameters are estimated automatically. Experiments and comparative results show a superior performance of the proposed method.
Farid Flitti, Mohammed Bennamoun, Du Q. Huynh, Robyn A. Owens
ICIP3
2009 Probabilistic Satellite Image Fusion
Farid Flitti, Mohammed Bennamoun, Du Q. Huynh, Amine Bermak, Christophe Collet 0001
CAIP3
2008 On-belt analysis of minerals using naturally occurring gamma radiation
abstract
We describe a method to analyze materials on a conveyor belt using natural gamma spectra collected with a BGO (Bismuth Germanate) gamma ray detector, which collects emissions from Potassium (K), Uranium (U), and Thorium (Th) in the materials. A statistical model is proposed based on a Poisson process and an approximate maximum likelihood (ML) technique via the expectation-maximization (EM) algorithm is then used to estimate the amount of each of the three elements in the material. A refinement of the statistical model is used to estimate linear drift in the detector.
William Moran 0001, Du Q. Huynh, Michael Edwards, Andrew Harris, Xuezhi Wang 0001, Barbara F. La Scala
ICASSP2
2008 Recursive structure and motion estimation from noisy uncalibrated video sequences
abstract
This paper builds on a novel framework of hybrid matching constraints for estimation of structure and recovery of camera focal length and motion, combining the advantages of both discrete and continuous methods. Our recursive method can deal with both image noise and outliers. The system is an extension of the epipolar hybrid matching constraints in conjunction with a simple structure estimation scheme using standard triangulation. The extension enables the system to deal with varying focal length of the camera. The structure obtained from some previous image frames is used to improve estimates of the camera focal length and motion for the current image frame. These are, in turn, used to refine the structure. Finally, a RANSAC outlier rejection scheme is employed to reject outlier tracks, inevitably obtained from any tracker. The performance of the proposed system is demonstrated on simulated experiments.
Du Q. Huynh, Anders Heyden
ICPR1
2006 Motion Guided Video Sequence Synchronization
Daniel Wedge, Du Q. Huynh, Peter Kovesi
ACCV (2)2
2005 Scene point constraints in camera auto-calibration: an implementational perspective
Du Q. Huynh, Anders Heyden
Image Vis. Comput.1
2004 Improving Query Effectiveness for Large Image Databases with Multiple Visual Feature Combination
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu, Du Q. Huynh
DASFAA4
2003 Outlier Correction in Image Sequences for the Affine Camera
abstract
It is widely known that, for the affine camera model, both shape and motion can be factorized directly from the so-called image measurement matrix constructed from image point coordinates. The ability to extract both shape and motion from this matrix by a single SVD operation makes this shape-from-motion approach attractive; however, it can not deal with missing feature points and, in the presence of outliers, a direct SVD to the matrix would yield highly unreliable shape and motion components. Here, we present an outlier correction scheme that iteratively updates the elements of the image measurement matrix. The magnitude and sign of the update to each element is dependent upon the residual robustly estimated in each iteration. The result is that outliers are corrected and retained, giving improved reconstruction and smaller reprojection errors. Our iterative outlier correction scheme has been applied to both synthesized and real video sequences. The results obtained are remarkably good.
Du Q. Huynh, R. Hartley, Anders Heyden
ICCV1
2003 CMVF: A Novel Dimension Reduction Scheme for Efficient Indexing in A Large Image Database
abstract
No abstract available.
Jialie Shen 0001, Anne H. H. Ngu, John Shepherd 0001, Du Q. Huynh, Quan Z. Sheng
SIGMOD Conference4
2002 Robust factorization for the affine camera: Analysis and comparison
abstract
Based on our previous work on the use of subspace dis-tances for the outlier deiection problem in video sequences under afFne projection. this paper reports ourfurther anal-ysis of the problem and presents fwo algorithms for com-puting the reprojection errors of imagefeatures in the out-lier detection process. Extensive experiments on real video sequences have been conducted to verrfi the performance ofthe algorithms. The key contributions ofthe paper are presentation offhe relationship befween subspace distances and reprojection errors and demonstration thar repmjec-tion errors can be estimated without erplicitly computing the projective structure. 1.
Du Q. Huynh, Anders Heyden
ICARCV1
2001 Outlier Detection in Video Sequences under Affine Projection
abstract
A novel robust method for outlier detection in structure and motion recovery for affine cameras is presented. It is an extension of the well-known Tomasi-Kanade factorization technique (C. Tomasi T. Kanade, 1992) designed to handle outliers. It can also be seen as an importation of the LMedS technique or RANSAC into the factorization framework. Based on the computation of distances between subspaces, it relates closely with the subspace-based factorization methods for the perspective case presented by G. Sparr (1996) and others and the subspace-based-factorization for affine cameras with missing data by D. Jacobs (1997). Key features of the method presented are its ability to compare different subspaces and the complete automation of the detection and elimination of outliers. Its performance and effectiveness are demonstrated by experiments involving simulated and real video sequences.
Du Q. Huynh, Anders Heyden
CVPR (1)1
2001 Combining multi-visual features for efficient indexing in a large image database
Anne H. H. Ngu, Quan Z. Sheng, Du Q. Huynh, Ron Lei
VLDB J.3
2000 The Cross Ratio: A Revisit to its Probability Density Function
abstract
The cross ratio has wide applications in computer vision because of its invariance under projective transformation. In active vision where the projections of quadruples of collinear landmark points in the scene are tracked in the image sequence for robot localisation or online camera calibration, one often needs to compute cross ratios from noisy image data for some subsequent operations. Being able to assess the reliability of each computed cross ratio value against a known level of image noise is therefore of importance. This aim motivates our research to derive the probability density function (p.d.f.) of the cross ratio based on the normality assumption of the associated random variables and to investigate into empirical cases where this assumption fails to hold. Although an analytical formula for the general p.d.f. of the cross ratio has not been achieved, our research results show that (i) the distance between the closest pair of collinear points is a significant factor that determines the shape of the p.d.f. of the cross ratio and (ii) a good estimate of the cross ratio can be obtained if the points of the quadruple are sufficiently far apart.
Du Q. Huynh
BMVC1
2000 Semi-Automatic Metric Reconstruction of Buildings from Self-Calibration: Preliminary Results on the Evaluation of a Linear Camera Self-Calibration Method
abstract
We investigate the linear self-calibration method proposed by Newsam et al. (1996) for our project on 3D reconstruction of architectural buildings. This self-calibration method assumes that the principal point is known, the camera has square pixels and has no skew. It allows 3D shape to be reconstructed from two images while giving the camera the freedom to vary its focal length. In this paper, we evaluate the focal lengths obtained from their method with both synthetic data and real data. In real data where known 3D data are available, Tsai's (1987) calibration method is used for comparison. Our experimental results show that the focal lengths from the two methods differed by less than 5% and the reconstructed 3D shape was very good in that angles were well preserved.
Du Q. Huynh, Y. S. Chou, Hung-Tat Tsui
ICPR1
1999 Affine Reconstruction from Monocular Vision in the Presence of a Symmetry Plane
abstract
This paper reports a closed-form solution for reconstructing a scene up to an affine transformation from a single image in the presence of a symmetry plane. Unlike scene reconstruction in stereo vision, the affine reconstruction process discussed in this paper does not require any knowledge about camera parameters or camera orientation relative to the scene, so camera self-calibration is totally eliminated. By setting in the scene a plane mirror which creates lateral symmetric world points for an uncalibrated, perspective camera to capture, the linear equations involved in the reconstruction process can be derived from two sets of similar triangles. The affine reconstruction is relative to an arbitrary affine coordinated frame implicitly defined on the mirror plane. Also involved in the process are the estimation of the epipole and recovery of the image-to-mirror plane homography. Implementation on estimating the epipole is detailed. A real experiment is presented to demonstrate the reconstruction.
Du Q. Huynh
ICCV1
1999 Calibrating a Structured Light Stripe System: A Novel Approach
Du Q. Huynh, Robyn A. Owens, Peter E. Hartmann
Int. J. Comput. Vis.1
1998 Self-calibrating a Stereo Head: An Error Analysis in the Neighbourhood of Degenerate Configurations
abstract
We show that the self-calibration of a stereo head corresponding points in an image pair is in certain circumstances prone to considerable error. A novel error analysis reveals that the automated determination of relative orientation and focal length is adversely affected when the cameras verge inwards a similar amount, and when the principal point locations have a horizontal error. This analysis is facilitated by the adoption of closed-form solutions for self-calibration from previous work of the authors. It is also shown that estimation of the fundamental matrix associated with a stereo head image pair is improved when a domain-specific parameterisation and associated computational techniques are adopted. Experiments conducted with such image pairs suggest that, given cognisance of sensitive configurations and adoption of the revised method of fundamental matrix estimation, robust reconstructions are attainable. This is demonstrated on the problem of metrically reconstructing a scene from two pairs of images obtained by an uncalibrated stereo head undergoing unknown ground-plane motion.
Lourdes Agapito, Du Q. Huynh, Michael J. Brooks
ICCV2
1998 Euclidean reconstruction from an image triplet: a sensitivity analysis
abstract
This paper studies the sensitivity in Euclidean reconstruction from an image triplet taken by an uncalibrated camera mounted on a robot arm. The idea of such a reconstruction is closely related to that proposed by Zisserman et al. (1995). In this paper, we focus on an intermediate step of the reconstruction procedure which requires estimating the screw axis that corresponds to the defective eigenvector of a 4/spl times/4 matrix. Hundreds of the conducted synthetic tests show that the algorithm is very sensitive to image noise and perturbations on camera motions and that if the matrix is perturbed by Gaussian noise then the reliability of the computed screw axis can be estimated.
Du Q. Huynh
ICPR1
1998 Towards robust metric reconstruction via a dynamic uncalibrated stereo head
Michael J. Brooks, Lourdes Agapito, Du Q. Huynh, Luis Baumela
Image Vis. Comput.3
1997 Calibration of a Structured Light System: A Projective Approach
abstract
We present in this paper a novel calibration method that uses cross ratio to compute world points falling onto any given light stripe plane of a structured light system. We show that, by using 4 known non-coplanar sets of 3 collinear world points, the direct 4/spl times/3 image-to-world transformation matrix for each light stripe plane can also be recovered from plane-to-plane homography. Preliminary experiments conducted with a calibration target and a mannequin suggest that this novel calibration method is robust and is applicable to many shape measurement task.
Du Q. Huynh
CVPR1
1996 Direct Methods for Self-Calibration of a Moving Stereo Head
Michael J. Brooks, Lourdes Agapito, Du Q. Huynh, Luis Baumela
ECCV (2)3
1994 Line labelling and region segmentation in stereo image pairs
Du Q. Huynh, Robyn A. Owens
Image Vis. Comput.1