Neil Robertson 0002

dblp:26/7169 · also Neil M. Robertson, Neil Martin Robertson · DBLP profile ↗
← Back
56ranked-venue papers
4as first author
9since 2021 · last 2023
0000-0003-2461-8799ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 30 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 SMKD: Selective Mutual Knowledge Distillation
abstract
Mutual knowledge distillation (MKD) is a technique used to transfer knowledge between multiple models in a collaborative manner. However, it is important to note that not all knowledge is accurate or reliable, particularly under challenging conditions such as label noise, which can lead to models that memorize undesired information. This problem can be addressed by improving the reliability of the knowledge source, as well as selectively selecting reliable knowledge for distillation. While making a model more reliable is a widely studied topic, selective MKD has received less attention. To address this, we propose a new framework called selective mutual knowledge distillation (SMKD). The key component of SMKD is a generic knowledge selection formulation, which allows for either static or progressive selection thresholds. Additionally, SMKD covers two special cases: using no knowledge and using all knowledge, resulting in a unified MKD framework. We present extensive experimental results to demonstrate the effectiveness of SMKD and justify its design.
Xinshao Wang, Neil Robertson 0002, David A. Clifton, Christoph Meinel, Haojin Yang 0001
IJCNN3
2023 Discrepancy-Guided Domain-Adaptive Data Augmentation
abstract
Data augmentation has been observed playing a crucial role in achieving better generalization in many machine learning tasks, especially in unsupervised domain adaptation (DA). It is particularly effective on visual object recognition tasks as images are high-dimensional with an enormous range of variations that can be simulated. Existing data augmentation techniques, however, are not explicitly designed to address the differences between different domains. Expert knowledge about the data is required, as well as manual efforts in finding the optimal parameters. In this article, we propose a novel domain-adaptive augmentation method by making use of a state-of-the-art style transfer method and domain discrepancy measurement. Specifically, we measure the discrepancy between source and target domains, and use it as a guide to augment the original source samples using style transferred source-to-target samples. The proposed domain-adaptive augmentation method is data and model agnostic that can be easily incorporated with state-of-the-art DA algorithms. We show empirically that, by using this domain-adaptive augmentation, we are able to gradually reduce the discrepancy between the source and target samples, and further boost the adaptation performance using different DA algorithms on three popular domain adaption datasets.
Jian Gao 0018, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
IEEE Trans. Neural Networks Learn. Syst.5
2022 Efficient One-Stage Video Object Detection by Exploiting Temporal Consistency
Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
ECCV (35)4
2022 TDViT: Temporal Dilated Video Transformer for Dense Video Tasks
Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
ECCV (35)4
2022 Ranked List Loss for Deep Metric Learning
abstract
The objective of deep metric learning (DML) is to learn embeddings that can capture semantic similarity and dissimilarity information among data points. Existing pairwise or tripletwise loss functions used in DML are known to suffer from slow convergence due to a large proportion of trivial pairs or triplets as the model improves. To improve this, ranking-motivated structured losses are proposed recently to incorporate multiple examples and exploit the structured information among them. They converge faster and achieve state-of-the-art performance. In this work, we unveil two limitations of existing ranking-motivated structured losses and propose a novel ranked list loss to solve both of them. First, given a query, only a fraction of data points is incorporated to build the similarity structure. Consequently, some useful examples are ignored and the structure is less informative. To address this, we propose to build a set-based similarity structure by exploiting all instances in the gallery. The learning setting can be interpreted as few-shot retrieval: given a mini-batch, every example is iteratively used as a query, and the rest ones compose the gallery to search, i.e., the support set in few-shot setting. The rest examples are split into a positive set and a negative set. For every mini-batch, the learning objective of ranked list loss is to make the query closer to the positive set than to the negative set by a margin. Second, previous methods aim to pull positive pairs as close as possible in the embedding space. As a result, the intraclass data distribution tends to be extremely compressed. In contrast, we propose to learn a hypersphere for each class in order to preserve useful similarity structure inside it, which functions as regularisation. Extensive experiments demonstrate the superiority of our proposal by comparing with the state-of-the-art methods on the fine-grained image retrieval task. Our source code is available online: https://github.com/XinshaoAmosWang/Ranked-List-Loss-for-DML.
Xinshao Wang, Yang Hua 0001, Elyor Kodirov, Neil Robertson 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 MAMBA: Multi-level Aggregation via Memory Bank for Video Object Detection
abstract
State-of-the-art video object detection methods maintain a memory structure, either a sliding window or a memory queue, to enhance the current frame using attention mechanisms. However, we argue that these memory structures are not efficient or sufficient because of two implied operations: (1) concatenating all features in memory for enhancement, leading to a heavy computational cost; (2) frame-wise memory updating, preventing the memory from capturing more temporal information. In this paper, we propose a multi-level aggregation architecture via memory bank called MAMBA. Specifically, our memory bank employs two novel operations to eliminate disadvantages of existing methods: (1) light-weight key-set construction which can significantly reduce the computational cost; (2) fine-grained feature-wise updating strategy which enables our method to utilize knowledge from the whole video. To better enhance features from complementary levels, i.e., feature maps and proposals, we further propose a generalized enhancement operation (GEO) to aggregate multi-level features in a unified manner. We conduct extensive evaluations on the challenging ImageNetVID dataset. Compared with existing state-of-the-art methods, our method achieves superior performance in terms of both speed and accuracy. More remarkably, MAMBA achieves mAP of 83.7%/84.6% at 12.6/9.1 FPS with ResNet-101.
Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
AAAI4
2021 Temporal Meta-Adaptor for Video Object Detection
Yang Hua 0001, Jian Gao 0018, Neil Robertson 0002
BMVC5
2021 ProSelfLC: Progressive Self Label Correction for Training Robust Deep Neural Networks
abstract
To train robust deep neural networks (DNNs), we systematically study several target modification approaches, which include output regularisation, self and non-self label correction (LC). Two key issues are discovered: (1) Self LC is the most appealing as it exploits its own knowledge and requires no extra models. However, how to automatically decide the trust degree of a learner as training goes is not well answered in the literature? (2) Some methods penalise while the others reward low-entropy predictions, prompting us to ask which one is better?To resolve the first issue, taking two well-accepted propositions–deep neural networks learn meaningful patterns before fitting noise [3] and minimum entropy regularisation principle [10]–we propose a novel end-to-end method named ProSelfLC, which is designed according to learning time and entropy. Specifically, given a data point, we progressively increase trust in its predicted label distribution versus its annotated one if a model has been trained for enough time and the prediction is of low entropy (high confidence). For the second issue, according to ProSelfLC, we empirically prove that it is better to redefine a meaningful low-entropy status and optimise the learner toward it. This serves as a defence of entropy minimisation.We demonstrate the effectiveness of ProSelfLC through extensive experiments in both clean and noisy settings. The source code is available at https://github.com/XinshaoAmosWang/ProSelfLC-CVPR2021.
Xinshao Wang, Yang Hua 0001, Elyor Kodirov, David A. Clifton, Neil Robertson 0002
CVPR5
2021 Fine-Grained Pose Temporal Memory Module for Video Pose Estimation and Tracking
abstract
The task of video pose estimation and tracking has been largely improved with the development of image pose estimation recently. However, there are still many challenging cases, such as body part occlusion, fast body motion, camera zooming, and complex background. Most existing methods generally use the temporal information to get more precise human bounding boxes or just use it in the tracking stage, but they fail to improve the accuracy of pose estimation tasks. To better solve these problems and utilize the temporal information efficiently and effectively, we present a novel structure, called pose temporal memory module, which is flexible to be transferred into top-down pose estimation frameworks. The temporal information stored in the pose temporal memory is aggregated into the current frame feature in our proposed module. We also transfer compositional de-attention (CoDA) to solve the unique keypoint occlusion problem in this task and propose a novel keypoint feature replacement to recover the extreme error detection under fine-grained keypoint-level guidance. To verify the generality and effectiveness of our proposed method, we integrate our module into two widely used pose estimation frameworks and obtain notable improvement on the PoseTrack dataset with only a few extra computing resources.
Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ICASSP6
2020 Reinforcing Neural Network Stability with Attractor Dynamics
abstract
Recent approaches interpret deep neural works (DNNs) as dynamical systems, drawing the connection between stability in forward propagation and generalization of DNNs. In this paper, we take a step further to be the first to reinforce this stability of DNNs without changing their original structure and verify the impact of the reinforced stability on the network representation from various aspects. More specifically, we reinforce stability by modeling attractor dynamics of a DNN and propose relu-max attractor network (RMAN), a light-weight module readily to be deployed on state-of-the-art ResNet-like networks. RMAN is only needed during training so as to modify a ResNet's attractor dynamics by minimizing an energy function together with the loss of the original learning task. Through intensive experiments, we show that RMAN-modified attractor dynamics bring a more structured representation space to ResNet and its variants, and more importantly improve the generalization ability of ResNet-like networks in supervised tasks due to reinforced stability.
Hanming Deng, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
AAAI6
2020 Imbalance Robust Softmax for Deep Embeeding Learning
Hao Zhu 0010, Guosheng Hu, Neil Robertson 0002
ACCV (5)5
2020 Reducing Distributional Uncertainty by Mutual Information Maximisation and Transferable Feature Learning
Jian Gao 0018, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
ECCV (23)5
2020 Differentiable Automatic Data Augmentation
Yonggang Li 0001, Guosheng Hu, Yongtao Wang, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang
ECCV (22)5
2020 Improving Detection And Recognition Of Degraded Faces By Discriminative Feature Restoration Using GAN
abstract
Face detection and recognition in the wild is currently one of the most interesting and challenging problems. Many algorithms with high performance have already been proposed and applied in real-world applications. However, the problem of detecting and recognising degraded faces from low-quality images and videos mostly remains unsolved. In this paper, we present an algorithm capable of recovering facial features from low-quality videos and images. The resulting output image boosts the performance of existing face detection and recognition algorithms. It contains an effective method involving metric learning and different loss function components operating on different parts of the generator. This enhances the degraded faces by restoring their lost features rather than its perceptual quality. Our approach has been experimentally proven to enhance face detection and recognition, e.g., the face detection rate is improved by 3.08% for S3FD [1] and the area under the ROC curve for recognition is improved by 2.55% for ArcFace [2] on the SCFace dataset.
Soumya Shubhra Ghosh, Yang Hua 0001, Sankha S. Mukherjee, Neil Robertson 0002
ICIP4
2020 Deep Metric Learning for Proteomics
abstract
Deep learning has become an innovative tool for predicting the properties of a protein. However, obtaining an accurate predictive model using deep learning methods typically requires a large amount of labelled data, which is expensive and time-consuming to accumulate. Even when optimised, these algorithms are often black boxes, which make it challenging to interpret the decision-making processes that lead to the final prediction. Therefore, there is a demand for innovative modelling techniques that overcome these drawbacks within the space of bioinformatic deep learning. To address these issues, we have designed a modelling scheme that utilises techniques from computer vision. Specifically, we explore how triplet-networks can form a robust model architecture that is capable of learning and ranking proteins from just a few labelled examples. We evaluate our model on a variety of downstream tasks, including peak absorption wavelength, enantioselectivity, plasma membrane localisation, and thermostability. The embedded representations produced by this method show considerable improvement when compared to previous baseline models. Finally, to emphasise that this is an example of white-box deep learning, we visualised the features produced by the algorithm to gain a better understanding as to how the network reaches its prediction for each protein property.
Mark Lennox, Neil Robertson 0002, Barry Devereux
ICMLA2
2020 VTT: Long-term Visual Tracking with Transformers
abstract
Long-term visual tracking is a challenging problem. State-of-the-art long-term trackers, e.g., GlobalTrack, utilize region proposal networks (RPNs) to generate target proposals. However, the performance of the trackers is affected by occlusions and large scale or ratio variations. To address these issues, in this paper, we are the first to propose a novel architecture with transformers for long-term visual tracking. Specifically, the proposed Visual Tracking Transformer (VTT) utilizes a transformer encoder-decoder architecture for aggregating global information to deal with occlusion and large scale or ratio variation. Furthermore, it also shows better discriminative power against instance-level distractors without the need for extra labeling and hard-sample mining. We conduct extensive experiments on three large-scale long-term tracking datasets and have achieved state-of-the-art performance.
Tianling Bian, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ICPR6
2020 Attention-Based Two-Stream Convolutional Networks for Face Spoofing Detection
abstract
Since the human face preserves the richest information for recognizing individuals, face recognition has been widely investigated and achieved great success in various applications in the past decades. However, face spoofing attacks (e.g., face video replay attack) remain a threat to modern face recognition systems. Though many effective methods have been proposed for anti-spoofing, we find that the performance of many existing methods is degraded by illuminations. It motivates us to develop illumination-invariant methods for anti-spoofing. In this paper, we propose a two-stream convolutional neural network (TSCNN), which works on two complementary spaces: RGB space (original imaging space) and multi-scale retinex (MSR) space (illumination-invariant space). Specifically, the RGB space contains the detailed facial textures, yet it is sensitive to illumination; MSR is invariant to illumination, yet it contains less detailed facial information. In addition, the MSR images can effectively capture the high-frequency information, which is discriminative for face spoofing detection. Images from two spaces are fed to the TSCNN to learn the discriminative features for anti-spoofing. To effectively fuse the features from two sources (RGB and MSR), we propose an attention-based fusion method, which can effectively capture the complementarity of two features. We evaluate the proposed framework on various databases, i.e., CASIA-FASD, REPLAY-ATTACK, and OULU, and achieve very competitive performance. To further verify the generalization capacity of the proposed strategies, we conduct cross-database experiments, and the results show the great effectiveness of our method.
Haonan Chen 0003, Guosheng Hu, Zhen Lei 0001, Yaowu Chen, Neil Robertson 0002, Stan Z. Li
IEEE Trans. Inf. Forensics Secur.5
2019 Deep Metric Learning by Online Soft Mining and Class-Aware Attention
abstract
Deep metric learning aims to learn a deep embedding that can capture the semantic similarity of data points. Given the availability of massive training samples, deep metric learning is known to suffer from slow convergence due to a large fraction of trivial samples. Therefore, most existing methods generally resort to sample mining strategies for selecting nontrivial samples to accelerate convergence and improve performance. In this work, we identify two critical limitations of the sample mining methods, and provide solutions for both of them. First, previous mining methods assign one binary score to each sample, i.e., dropping or keeping it, so they only selects a subset of relevant samples in a mini-batch. Therefore, we propose a novel sample mining method, called Online Soft Mining (OSM), which assigns one continuous score to each sample to make use of all samples in the mini-batch. OSM learns extended manifolds that preserve useful intraclass variances by focusing on more similar positives. Second, the existing methods are easily influenced by outliers as they are generally included in the mined subset. To address this, we introduce Class-Aware Attention (CAA) that assigns little attention to abnormal data samples. Furthermore, by combining OSM and CAA, we propose a novel weighted contrastive loss to learn discriminative embeddings. Extensive experiments on two fine-grained visual categorisation datasets and two video-based person re-identification benchmarks show that our method significantly outperforms the state-of-the-art.
Xinshao Wang, Yang Hua 0001, Elyor Kodirov, Guosheng Hu, Neil Robertson 0002
AAAI5
2019 SC-RANK: Improving Convolutional Image Captioning with Self-Critical Learning and Ranking Metric-based Reward
Shiyang Yan, Yang Hua 0001, Neil Robertson 0002
BMVC3
2019 Semantic Alignment: Finding Semantically Consistent Ground-Truth for Facial Landmark Detection
abstract
Recently, deep learning based facial landmark detection has achieved great success. Despite this, we notice that the semantic ambiguity greatly degrades the detection performance. Specifically, the semantic ambiguity means that some landmarks (e.g. those evenly distributed along the face contour) do not have clear and accurate definition, causing inconsistent annotations by annotators. Accordingly, these inconsistent annotations, which are usually provided by public databases, commonly work as the ground-truth to supervise network training, leading to the degraded accuracy. To our knowledge, little research has investigated this problem. In this paper, we propose a novel probabilistic model which introduces a latent variable, i.e. the `real' ground-truth which is semantically consistent, to optimize. This framework couples two parts (1) training landmark detection CNN and (2) searching the `real' ground-truth. These two parts are alternatively optimized: the searched `real' ground-truth supervises the CNN training; and the trained CNN assists the searching of `real' ground-truth. In addition, to recover the unconfidently predicted landmarks due to occlusion and low quality, we propose a global heatmap correction unit (GHCU) to correct outliers by considering the global face shape as a constraint. Extensive experiments on both image-based (300W and AFLW) and video-based (300-VW) databases demonstrate that our method effectively improves the landmark detection accuracy and achieves the state of the art performance.
Zhiwei Liu 0004, Xiangyu Zhu 0001, Guosheng Hu, Haiyun Guo, Ming Tang 0001, Zhen Lei 0001, Neil Robertson 0002, Jinqiao Wang
CVPR7
2019 Ranked List Loss for Deep Metric Learning
abstract
The objective of deep metric learning (DML) is to learn embeddings that can capture semantic similarity information among data points. Existing pairwise or tripletwise loss functions used in DML are known to suffer from slow convergence due to a large proportion of trivial pairs or triplets as the model improves. To improve this, rankingmotivated structured losses are proposed recently to incorporate multiple examples and exploit the structured information among them. They converge faster and achieve state-of-the-art performance. In this work, we present two limitations of existing ranking-motivated structured losses and propose a novel ranked list loss to solve both of them. First, given a query, only a fraction of data points is incorporated to build the similarity structure. Consequently, some useful examples are ignored and the structure is less informative. To address this, we propose to build a setbased similarity structure by exploiting all instances in the gallery. The samples are split into a positive set and a negative set. Our objective is to make the query closer to the positive set than to the negative set by a margin. Second, previous methods aim to pull positive pairs as close as possible in the embedding space. As a result, the intraclass data distribution might be dropped. In contrast, we propose to learn a hypersphere for each class in order to preserve the similarity structure inside it. Our extensive experiments show that the proposed method achieves state-of-the-art performance on three widely used benchmarks.
Xinshao Wang, Yang Hua 0001, Elyor Kodirov, Guosheng Hu, Romain Garnier, Neil Robertson 0002
CVPR6
2019 Object Guided External Memory Network for Video Object Detection
abstract
Video object detection is more challenging than image object detection because of the deteriorated frame quality. To enhance the feature representation, state-of-the-art methods propagate temporal information into the deteriorated frame by aligning and aggregating entire feature maps from multiple nearby frames. However, restricted by feature map's low storage-efficiency and vulnerable content-address allocation, long-term temporal information is not fully stressed by these methods. In this work, we propose the first object guided external memory network for online video object detection. Storage-efficiency is handled by object guided hard-attention to selectively store valuable features, and long-term information is protected when stored in an addressable external data matrix. A set of read/write operations are designed to accurately propagate/allocate and delete multi-level memory feature under object guidance. We evaluate our method on the ImageNet VID dataset and achieve state-of-the-art performance as well as good speed-accuracy tradeoff. Furthermore, by visualizing the external memory, we show the detailed object-level reasoning process across frames.
Hanming Deng, Yang Hua 0001, Tao Song 0003, Zongpu Zhang, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ICCV7
2019 Unsupervised Video Summarization with Attentive Conditional Generative Adversarial Networks
abstract
With the rapid growth of video data, video summarization technique plays a key role in reducing people's efforts to explore the content of videos by generating concise but informative summaries. Though supervised video summarization approaches have been well studied and achieved state-of-the-art performance, unsupervised methods are still highly demanded due to the intrinsic difficulty of obtaining high-quality annotations. In this paper, we propose a novel yet simple unsupervised video summarization method with attentive conditional Generative Adversarial Networks (GANs). Firstly, we build our framework upon Generative Adversarial Networks in an unsupervised manner. Specifically, the generator produces high-level weighted frame features and predicts frame-level importance scores, while the discriminator tries to distinguish between weighted frame features and raw frame features. Furthermore, we utilize a conditional feature selector to guide GAN model to focus on more important temporal regions of the whole video frames. Secondly, we are the first to introduce the frame-level multi-head self-attention for video summarization, which learns long-range temporal dependencies along the whole video sequence and overcomes the local constraints of recurrent units, e.g., LSTMs. Extensive evaluations on two datasets, SumMe and TVSum, show that our proposed framework surpasses state-of-the-art unsupervised methods by a large margin, and even outperforms most of the supervised methods. Additionally, we also conduct the ablation study to unveil the influence of each component and parameter settings in our framework.
Xufeng He, Yang Hua 0001, Tao Song 0003, Zongpu Zhang, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ACM Multimedia7
2019 GAN-Based Pose-Aware Regulation for Video-Based Person Re-Identification
abstract
Video-based person re-identification deals with the inherent difficulty of matching sequences with different length, unregulated, and incomplete target pose/viewpoint structure. Common approaches operate either by reducing the problem to the still images case, facing a significant information loss, or by exploiting inter-sequence temporal dependencies as in Siamese Recurrent Neural Networks or in gait analysis. However, in all cases, the inter-sequences pose/viewpoint misalignment is considered, and the existing spatial approaches are mostly limited to the still images context. To this end, we propose a novel approach that can exploit more effectively the rich video information, by accounting for the role that the changing pose/viewpoint factor plays in the sequences matching process. In particular, our approach consists of two components. The first one attempts to complement the original pose-incomplete information carried by the sequences with synthetic GAN-generated images, and fuse their features vectors into a more discriminative viewpoint-insensitive embedding, namely Weighted Fusion (WF). Another one performs an explicit pose-based alignment of sequence pairs to promote coherent feature matching, namely Weighted-Pose Regulation (WPR). Extensive experiments on two large video-based benchmark datasets show that our approach outperforms considerably existing methods.
Alessandro Borgia, Yang Hua 0001, Elyor Kodirov, Neil Robertson 0002
WACV4
2019 IEGAN: Multi-Purpose Perceptual Quality Image Enhancement Using Generative Adversarial Network
abstract
Despite the breakthroughs in quality of image enhancement, an end-to-end solution for simultaneous recovery of the finer texture details and sharpness for degraded images with low resolution is still unsolved. Some existing approaches focus on minimizing the pixel-wise reconstruction error which results in a high peak signal-to-noise ratio. The enhanced images fail to provide high-frequency details and are perceptually unsatisfying, i.e., they fail to match the quality expected in a photo-realistic image. In this paper, we present Image Enhancement Generative Adversarial Network (IEGAN), a versatile framework capable of inferring photo-realistic natural images for both artifact removal and super-resolution simultaneously. Moreover, we propose a new loss function consisting of a combination of reconstruction loss, feature loss and an edge loss counterpart. The feature loss helps to push the output image to the natural image manifold and the edge loss preserves the sharpness of the output image. The reconstruction loss provides low-level semantic information to the generator regarding the quality of the generated images compared to the original. Our approach has been experimentally proven to recover photo-realistic textures from heavily compressed low-resolution images on public benchmarks and our proposed high-resolution World100 dataset.
Soumya Shubhra Ghosh, Yang Hua 0001, Sankha S. Mukherjee, Neil Robertson 0002
WACV4
2019 Computational Load Balancing on the Edge in Absence of Cloud and Fog
abstract
Mobile Cloud Computing or Fog computing refers to offloading computationally intensive algorithms from a mobile device to the cloud or an intermediate cloud in order to save resources e.g., time and energy in the mobile device. This paper proposes new solutions for situations when the cloud or fog is not available. First, the sensor network is modelled using a network of queues, then a linear programming technique is used to make scheduling decisions. Various centralized and distributed algorithms are then proposed, which improves overall system performance. Extensive simulations show slightly higher energy usage in comparison to the baseline non-offloading case, however, the job completion rate is significantly improved, the efficiency score metric shows the extra energy usage is justified. The algorithms have been simulated in various environments including high and low bandwidth, partial connectivity, and different rate of information exchanges to study the pros and cons of the proposed algorithms.
Saurav Sthapit, John S. Thompson, Neil Robertson 0002, James R. Hopgood
IEEE Trans. Mob. Comput.3
2018 Deep Multi-task Learning to Recognise Subtle Facial Expressions of Mental States
Guosheng Hu, Li Liu 0004, Yang Hua 0001, Zhihong Zhang 0001, Fumin Shen, Ling Shao 0001, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang
ECCV (12)10
2018 Deep Stock Representation Learning: From Candlestick Charts to Investment Decisions
abstract
We propose a novel investment decision strategy (IDS) based on deep learning. The performance of many IDSs is affected by stock similarity. Most existing stock similarity measurements have the problems: (a) The linear nature of many measurements cannot capture nonlinear stock dynamics; (b) The estimation of many similarity metrics (e.g. covariance) needs very long period historic data (e.g. 3K days) which cannot represent current market effectively; (c) They cannot capture translation-invariance. To solve these problems, we apply Convolutional AutoEncoder to learn a stock representation, based on which we propose a novel portfolio construction strategy by: (i) using the deeply learned representation and modularity optimisation to cluster stocks and identify diverse sectors, (ii) picking stocks within each cluster according to their Sharpe ratio (Sharpe 1994). Overall this strategy provides low-risk high-return portfolios. We use the Financial Times Stock Exchange 100 Index (FTSE 100) data for evaluation. Results show our portfolio outperforms FTSE 100 index and many well known funds in terms of total return in 2000 trading days.
Guosheng Hu, Kai Yang 0031, Flood Sung, Zhihong Zhang 0001, Neil Robertson 0002, Timothy M. Hospedales, Qiangwei Miemie
ICASSP9
2018 Tracking-assisted Weakly Supervised Online Visual Object Segmentation in Unconstrained Videos
abstract
This paper tackles the task of online video object segmentation with weak supervision, i.e., labeling the target object and background with pixel-level accuracy in unconstrained videos, given only one bounding box information in the first frame. We present a novel tracking-assisted visual object segmentation framework to achieve this. On the one hand, initialized with a given bounding box in the first frame, the auxiliary object tracking module guides the segmentation module frame by frame by providing motion and region information, which is usually missing in semi-supervised methods. Moreover, compared with the unsupervised approach, our approach with such minimum supervision can focus on the target object without bringing unrelated objects into the final results. On the other hand, the video object segmentation module also improves the robustness of the visual object tracking module by pixel-level localization and objectness information. Thus, segmentation and tracking in our framework can mutually help each other in an online manner. To verify the generality and effectiveness of the proposed framework, we evaluate our weakly supervised method on two cross-domain datasets, i.e., the DAVIS and VOT2016 datasets, with the same configuration and parameter setting. Experimental results show the top performance of our method, which is even better than the leading semi-supervised methods. Furthermore, we conduct the extensive ablation study on our approach to investigate the influence of each component and main parameters.
Zongpu Zhang, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ACM Multimedia6
2018 Cross-View Discriminative Feature Learning for Person Re-Identification
abstract
The viewpoint variability across a network of non-overlapping cameras is a challenging problem affecting person re-identification performance. In this paper, we investigate how to mitigate the cross-view ambiguity by learning highly discriminative deep features under the supervision of a novel loss function. The proposed objective is made up of two terms, the steering meta center term and the enhancing centers dispersion term, that steer the training process to mining effective intra-class and inter-class relationships in the feature domain of the identities. The effect of our loss supervision is to generate a more expanded feature space of compact classes where the overall level of the inter-identities' interference is reduced. Compared with the existing metric learning techniques, this approach has the advantage of achieving a better optimization because it jointly learns the embedding and the metric contextually. Our technique, by dismissing side-sources of performance gain, proves to enhance the CNN invariance to viewpoint without incurring increased training complexity (like in Siamese or triplet networks) and outperforms many related state-of-the-art techniques on Market-1501 and CUHK03.
Alessandro Borgia, Yang Hua 0001, Elyor Kodirov, Neil Robertson 0002
IEEE Trans. Image Process.4
2017 Balanced sensor management across multiple time instances via l-1/l-infinity norm minimization
abstract
In this paper, we propose a solution to the sensor management problem over multiple time instances that balances the accuracy of the sensor network estimation with its utilization. We show how this problem reduces to a binary optimization problem for which we give a convex relaxation based solution that involves the minimization of a regularized ℓ∞reweighted ℓ1norm. We show experimentally the behavior of the proposed algorithm and compare it with previous methods from the literature.
Cristian Rusu, John S. Thompson, Neil Robertson 0002
ICASSP3
2017 Attribute-Enhanced Face Recognition with Neural Tensor Fusion Networks
abstract
Deep learning has achieved great success in face recognition, however deep-learned features still have limited invariance to strong intra-personal variations such as large pose changes. It is observed that some facial attributes (e.g. eyebrow thickness, gender) are robust to such variations. We present the first work to systematically explore how the fusion of face recognition features (FRF) and facial attribute features (FAF) can enhance face recognition performance in various challenging scenarios. Despite the promise of FAF, we find that in practice existing fusion methods fail to leverage FAF to boost face recognition performance in some challenging scenarios. Thus, we develop a powerful tensor-based framework which formulates feature fusion as a tensor optimisation problem. It is nontrivial to directly optimise this tensor due to the large number of parameters to optimise. To solve this problem, we establish a theoretical equivalence between low-rank tensor optimisation and a two-stream gated neural network. This equivalence allows tractable learning using standard neural network optimisation tools, leading to accurate and stable optimisation. Experimental results show the fused feature works better than individual features, thus proving for the first time that facial attributes aid face recognition. We achieve state-of-the-art performance on three popular databases: MultiPIE (cross pose, lighting and expression), CASIA NIR-VIS2.0 (cross-modality environment) and LFW (uncontrolled environment).
Guosheng Hu, Yang Hua 0001, Zhihong Zhang 0001, Sankha S. Mukherjee, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang
ICCV8
2016 Face Recognition Using a Unified 3D Morphable Model
Guosheng Hu, Fei Yan 0001, Chi-Ho Chan, Weihong Deng, William J. Christmas, Josef Kittler, Neil Robertson 0002
ECCV (8)7
2016 Robust indoor speaker recognition in a network of audio and video sensors
abstract
Situational awareness is achieved naturally by the human senses of sight and hearing in combination. Automatic scene understanding aims at replicating this human ability using microphones and cameras in cooperation. In this paper, audio and video signals are fused and integrated at different levels of semantic abstractions. We detect and track a speaker who is relatively unconstrained, i.e., free to move indoors within an area larger than the comparable reported work, which is usually limited to round table meetings. The system is relatively simple: consisting of just 4 microphone pairs and a single camera. Results show that the overall multimodal tracker is more reliable than single modality systems, tolerating large occlusions and cross-talk. System evaluation is performed on both single and multi-modality tracking. The performance improvement given by the audio–video integration and fusion is quantified in terms of tracking precision and accuracy as well as speaker diarisation error rate and precision–recall (recognition). Improvements vs. the closest works are evaluated: 56% sound source localisation computational cost over an audio only system, 8% speaker diarisation error rate over an audio only speaker recognition unit and 36% on the precision–recall metric over an audio–video dominant speaker recognition method.
Eleonora D'Arca, Neil Robertson 0002, James R. Hopgood
Signal Process.2
2016 Video Anomaly Detection in Real Time on a Power-Aware Heterogeneous Platform
abstract
Field-programmable gate arrays (FPGAs) and graphics processing units (GPUs) are often used when real-time performance in video processing is required. An accelerated processor is chosen based on task-specific priorities (power consumption, processing time, and detection accuracy), and this decision is normally made once at design time. All three of the characteristics are important, particularly in battery-powered systems. Here, we propose a method for moving selection of processing platform from a single design-time choice to a continuous run-time one. We implement Histogram of Oriented Gradients (HOG) for cars and people and Mixture-of-Gaussians motion detectors running across FPGA, GPU, and central processing unit in a heterogeneous system. We use this to detect illegally parked vehicles in urban scenes. Power, time, and accuracy information for each detector is characterized. An anomaly measure is assigned to each detected object based on its trajectory and location compared with learned contextual movement patterns. This drives processor and implementation selection so that scenes with high behavioral anomalies are processed with faster but more power-hungry implementations, but routine or static time periods are processed with power-optimized and less accurate slower versions. Real-time performance is evaluated on video data sets including i-LIDS. Compared with power-optimized static selection, automatic dynamic implementation mapping is 10% more accurate, but draws 12-W of extra power in our testbed desktop system.
Calum G. Blair, Neil Robertson 0002
IEEE Trans. Circuits Syst. Video Technol.2
2015 Identifying anomalous objects in SAS imagery using uncertainty
Calum G. Blair, John S. Thompson, Neil Robertson 0002
FUSION3
2015 Joint classification of actions with matrix completion
abstract
Action classification is one of the crucial research areas with multitude applications. It has witnessed significant developments over last decade. In this paper, we propose to jointly classify actions from more than a single class using Matrix completion. Matrix-completion methods can handle the deficiencies in data very effectively resulting in improved classification accuracy. Features and labels from data are concatenated to form a big matrix with unknown or missing entries in the place of test data labels. Matrix-completion methods fill up these entries using tools from convex optimization resulting in classification. We show that the proposed method achieves improved performance over the recent works on two human action datasets including most popular Weizmann dataset and recently released and more realistic UCF-101 dataset.
Sushma Bomma, Neil Robertson 0002
ICIP2
2015 Instantaneous real-time head pose at a distance
abstract
In this paper we focus on robust, real-time human head pose estimation in low resolution RGB data without any smoothing motion priors e.g. direction of motion. Our main contributions lie in three major areas. First, we show that a generative Deep Belief Network model can be learned on human head data from multiple types of data sources. These sources have similar underlying data that are not necessarily labelled or have the same kind of ground truth. Second, we perform discriminative training using multiple disparate supervisory labels to fine tune the model for head pose estimation. Third, we present state-of-the-art results on two publicly available datasets using this new approach. Our implementation computes head pose for a head image in 0.8 milliseconds, making it real-time and highly scalable.
Sankha S. Mukherjee, Rolf Baxter, Neil Robertson 0002
ICIP3
2015 Recovering background regions in videos of cluttered urban scenes
abstract
In this paper we present a novel segmentation technique and adapt it for enhanced background recovery in crowded urban scenes. Building on an initial superpixels representation, smaller regions are merged depending on their perceived similarity to develop larger regions. Based on the observation that human activity tends to be quite structured, during this process we exploit emerging foreground context (tracks of people) to influence the segmentation process via Bayesian priors. These priors incorporate both temporal and spatial smoothing. We validate the approach on benchmarked urban datasets and show that our method improves on established image segmentation methods.
Iain Rodger, Barry Connor, Neil Robertson 0002
ICIP3
2015 On covariate factor detection and removal for robust gait recognition
abstract
We propose a novel bolt-on module capable of boosting the robustness of various single compact 2D gait representations. Gait recognition is negatively influenced by covariate factors including clothing and time which alter the natural gait appearance and motion. Contrary to traditional gait recognition, our bolt-on module remedies this by a dedicated covariate factor detection and removal procedure which we quantitatively and qualitatively evaluate. The fundamental concept of the bolt-on module is founded on exploiting the pixel-wise composition of covariate factors. Results demonstrate how our bolt-on module is a powerful component leading to significant improvements across gait representations and datasets yielding state-of-the-art results.
Tenika P. Whytock, Alexander G. Belyaev, Neil Robertson 0002
Mach. Vis. Appl.3
2015 Human behaviour recognition in data-scarce domains
abstract
This paper presents the novel theory for performing multi-agent activity recognition without requiring large training corpora. The reduced need for data means that robust probabilistic recognition can be performed within domains where annotated datasets are traditionally unavailable. Complex human activities are composed from sequences of underlying primitive activities. We do not assume that the exact temporal ordering of primitives is necessary, so can represent complex activity using an unordered bag. Our three-tier architecture comprises low-level video tracking, event analysis and high-level inference. High-level inference is performed using a new, cascading extension of the Rao–Blackwellised Particle Filter. Simulated annealing is used to identify pairs of agents involved in multi-agent activity. We validate our framework using the benchmarked PETS 2006 video surveillance dataset and our own sequences, and achieve a mean recognition F-Score of 0.82. Our approach achieves a mean improvement of 17% over a Hidden Markov Model baseline.
Rolf Baxter, Neil Robertson 0002, David M. Lane
Pattern Recognit.2
2015 An Adaptive Motion Model for Person Tracking with Instantaneous Head-Pose Features
abstract
This letter presents novel behaviour-based tracking of people in low-resolution using instantaneous priors mediated by head-pose. We extend the Kalman Filter to adaptively combine motion information with an instantaneous prior belief about where the person will go based on where they are currently looking. We apply this new method to pedestrian surveillance, using automatically-derived head pose estimates, although the theory is not limited to head-pose priors. We perform a statistical analysis of pedestrian gazing behaviour and demonstrate tracking performance on a set of simulated and real pedestrian observations. We show that by using instantaneous `intentional' priors our algorithm significantly outperforms a standard Kalman Filter on comprehensive test data.
Rolf Baxter, Michael J. V. Leach, Sankha S. Mukherjee, Neil Robertson 0002
IEEE Signal Process. Lett.4
2015 Deep Head Pose: Gaze-Direction Estimation in Multimodal Video
abstract
In this paper we present a convolutional neural network (CNN)-based model for human head pose estimation in low-resolution multi-modal RGB-D data. We pose the problem as one of classification of human gazing direction. We further fine-tune a regressor based on the learned deep classifier. Next we combine the two models (classification and regression) to estimate approximate regression confidence. We present state-of-the-art results in datasets that span the range of high-resolution human robot interaction (close up faces plus depth information) data to challenging low resolution outdoor surveillance data. We build upon our robust head-pose estimation and further introduce a new visual attention model to recover interaction with the environment . Using this probabilistic model, we show that many higher level scene understanding like human-human/scene interaction detection can be achieved. Our solution runs in real-time on commercial hardware.
Sankha S. Mukherjee, Neil Robertson 0002
IEEE Trans. Multim.2
2014 Look who's talking: Detecting the dominant speaker in a cluttered scenario
abstract
In this work we propose a novel method to automatically detect and localise the dominant speaker in an enclosed scenario by means of audio and video cues. The underpinning idea is that gesturing means speaking, so observing motions means observing an audio signal. To the best of our knowledge state-of-the-art algorithms are focussed on stationary motion scenarios and close-up scenes where only one audio source exists, whereas we enlarge the extent of the method to larger field of views and cluttered scenarios including multiple non-stationary moving speakers. In such contexts, moving objects which are not correlated to the dominant audio may exist and their motion may incorrectly drive the audio-video (AV) correlation estimation. This suggests extra localisation data may be fused at decision level to avoid detecting false positives. In this work, we learn Mel-frequency cepstral coefficients (MFCC) coefficients and correlate them to the optical flow. We also exploit the audio and video signals to estimate the position of the actual speaker, narrowing down the visual space of search, hence reducing the probability of incurring in a wrong voice-to-pixel region association. We compare our work with a state-of-the-art existing algorithm and show on real datasets a 36% precision improvement in localising a moving dominant speaker through occlusions and speech interferences.
Eleonora D'Arca, Neil Robertson 0002, James R. Hopgood
ICASSP2
2014 Depth Sensor Placement for Human Robot Cooperation
abstract
Continuous sensing of the environment from a mobile robot perspective can prevent harmful collisions between human and mobile service robots. However, the overall collision avoidance performance depends strongly on the optimal placement of multiple depth sensors on the mobile robot and maintains flexibility of the working area. In this paper, we present a novel approach to optimal sensor placement based on the visibility of the human in the robot environment combined with a quantified risk of collision. Human visibility is determined by ray tracing from all possible camera positions on the robot surface, quantifying safety based on the speed and direction of the robot throughout a pre-determined task. A cost function based on discrete cells is formulated and solved numerically for two scenarios of increasing complexity, using a CUDA implementation to reduce computation time.
Max Stähr, Andrew M. Wallace, Neil Robertson 0002
ICINCO (2)3
2014 Contextual anomaly detection in crowded surveillance scenes
abstract
This work addresses the problem of detecting human behavioural anomalies in crowded surveillance environments. We focus in particular on the problem of detecting subtle anomalies in a behaviourally heterogeneous surveillance scene. To reach this goal we implement a novel unsupervised context-aware process. We propose and evaluate a method of utilising social context and scene context to improve behaviour analysis. We find that in a crowded scene the application of Mutual Information based social context permits the ability to prevent self-justifying groups and propagate anomalies in a social network, granting a greater anomaly detection capability. Scene context uniformly improves the detection of anomalies in both datasets. The strength of our contextual features is demonstrated by the detection of subtly abnormal behaviours, which otherwise remain indistinguishable from normal behaviour.
Michael J. V. Leach, Ed P. Sparks, Neil Robertson 0002
Pattern Recognit. Lett.3
2013 Outlines of Objects Detection by Analogy
Asma Bellili, Slimane Larabi, Neil Robertson 0002
CAIP (1)3
2013 Using The Voice Spectrum For Improved Tracking Of People In A Joint Audio-Video Scheme
abstract
In this paper we present a new solution to the problem of speaker tracking among people where occlusions occur (disappearance and non-speaking). In a normal conversation between two or more people, we learn speaker mel-cepstral coefficients (MFCC) and incorporate this information into a sequential Bayesian audio-video position tracker. The joint video-to-audio data association step is thus improved and we achieve robust person recognition which in turn aids tracking performance. We provide comprehensive evaluation via simulations and real data quoting tracking accuracy, precision and diarisation error rate (DER) compared to ground truth. For simulate and real experiments in an open space the trajectory tracking performance increases by 20% measured against ground truth using our approach. As a further enhancement versus the state-of-the-art, speaker identity recognition at a distance is improved by 20% by exploiting audio-video localisation cues.
Eleonora D'Arca, Neil Robertson 0002, James R. Hopgood
ICASSP2
2013 Sparse representation based action and gesture recognition
abstract
In this paper we present a solution to the problem of action and gesture recognition using sparse representations. The dictionary is modelled as a simple concatenation of features computed for each action or gesture class from the training data, and test data is classified by finding sparse representation of the test video features over this dictionary. Our method does not impose any explicit training procedure on the dictionary. We experiment our model with two kinds of features, by projecting (i) Gait Energy Images (GEIs) and (ii) Motion-descriptors, to a lower dimension using Random projection. Experiments have shown 100% recognition rate on standard datasets and are compared to the results obtained with widely used SVM classifier.
Sushma Bomma, Paolo Favaro, Neil Robertson 0002
ICIP3
2013 Video tracking through occlusions by fast audio source localisation
abstract
In this paper we present a novel audio-visual speaker detection and localisation algorithm. Audio source position estimates are computed by a novel stochastic region contraction (SRC) audio search algorithm for accurate speaker localisation. This audio search algorithm is aided by available video information (stochastic region contraction with height estimation (SRC-HE)) which estimates head heights over the whole scene and gives a speed improvement of 56% over SRC. We finally combine audio and video data in a Kalman filter (KF) which fuses person-position likelihoods and tracks the speaker. Our system is composed of a single video camera and 16 microphones. We validate the approach on the problem of video occlusion i.e. two people having a conversation have to be detected and localised at a distance (as in surveillance scenarios vs. enclosed meeting rooms). We show video occlusion can be resolved and speakers can be correctly detected/localised in real data. Moreover, SRC-HE based joint audio-video (AV) speaker tracking outperforms the one based on the original SRC by 16% and 4% in terms of multi object tracking precision (MOTP) and multi object tracking accuracy (MOTA). Speaker change detection improves by 11% over SRC.
Eleonora D'Arca, Ashley Hughes, Neil Robertson 0002, James R. Hopgood
ICIP3
2012 A proposed gesture set for the control of industrial collaborative robots
abstract
Human-Robot Interaction is one of the key challenges in collaborative autonomous robotics. However, no standardised framework allowing either efficient portability to an actual industrial use nor comparison benchmarking exists. This work proposes, implements and evaluates such a set of common ground rules. We present the design constraints between different groups of requirements and a technical solution for automatic recognition using imaging hardware. The Human to Robot and Robot to Human Communications concepts are illustrated on the real industrial scenario: we focus on the definition of a set of gestures for Human to Robot communication in automotive manufacturing. The case study outlines the need for a defined set of gestures for establishing a basic communication with the collaborative robot. First, the gestures are designed to respect the social acceptance principle. Second, a gesture recognition algorithm based on Dynamic Time Warping is used to demonstrate the feasibility of discriminating those gestures by automatic processing. Evaluation of our technique shows low confusion and high accuracy with this method.
Paolo Barattini, Claire Morand, Neil Robertson 0002
RO-MAN3
2010 Probabilistic Behaviour Signatures: Feature-based behaviour recognition in data-scarce domains
Rolf Baxter, Neil Robertson 0002, David M. Lane
FUSION2
2009 Aerial image segmentation for flood risk analysis
abstract
This paper presents a technique for image segmentation. We demonstrate its efficacy for classsifying high-resolution aerial images. The application is peak water flow estimation in a river catchment in the city of Zurich and the data covers a large rural and urban setting. The output of the segmentation process is used as input to a hydrological model. We introduce a combined, probabilistic, segmentation approach based on colour (the LAB colour space is used), texture (using entropy) and image features (gradients). Classification rates for natural land surfaces and man-made structures are up to 90% and 85% respectively. When the automatic segmentation result is compared to the official land use data and reclassified for use in GIS we achieve an overall classification accuracy of 70%. This new classification is tested on the WetSpa hydrological model and the resulting flow estimate compares favourably with that computed from hand-classified land use data.
Neil Robertson 0002, Tak Chan
ICIP1
2006 Estimating Gaze Direction from Low-Resolution Faces in Video
Neil Robertson 0002, Ian D. Reid 0001
ECCV (2)1
2006 A general method for human activity recognition in video
Neil Robertson 0002, Ian D. Reid 0001
Comput. Vis. Image Underst.1
2005 Behaviour Understanding in Video: A Combined Method
abstract
In this paper we develop a system for human behaviour recognition in video sequences. Human behaviour is modelled as a stochastic sequence of actions. Actions are described by a feature vector comprising both trajectory information (position and velocity), and a set of local motion descriptors. Action recognition is achieved via probabilistic search of image feature databases representing previously seen actions. A HMM which encodes the rules of the scene is used to smooth sequences of actions. High-level behaviour recognition is achieved by computing the likelihood that a set of predefined hidden Markov models explains the current action sequence. Thus, human actions and behaviour are represented using a hierarchy of abstraction: from simple actions, to actions with spatio-temporal context, to action sequences and finally general behaviours. While the upper levels all use (parametric) Bayes networks and belief propagation, the lowest level uses nonparametric sampling from a previously learned database of actions. The combined method represents a general framework for human behaviour modelling. In this paper we demonstrate the results chiefly on broadcast tennis sequences for automated video annotation.
Neil Robertson 0002, Ian D. Reid 0001
ICCV1