Anthony R. Dick

dblp:77/4964 · DBLP profile ↗
← Back
55ranked-venue papers
5as first author
1since 2021 · last 2024
0000-0001-9049-7345ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 4 first-authorArtificial intelligence and machine learning · 40 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
26 papers
Vision and language · 35% Video understanding and tracking · 16% 3D vision · 10%
Computer graphics and multimedia
7 papers
Geometric modeling and processing · 43% Image and video processing · 35% Multimedia analysis and retrieval · 8%
Databases, data mining, and information retrieval
2 papers
Data mining · 70% Information retrieval · 30%
Theoretical computer science
3 papers
Mathematical optimization · 100%

Topics — the 30 heaviest of 74, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
visual question answering
1.862018
Image Captioning and Visual Question Answering Based on Attributes and External Knowledge · IEEE Trans. Pattern Anal. Mach. Intell. 2018
FVQA: Fact-Based Visual Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Visual Question Answering With Memory-Augmented Networks · CVPR 2018
Computer vision › Video understanding and tracking
object tracking
0.952016
Online Metric-Weighted Linear Representations for Robust Visual Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Visual Tracking With Spatio-Temporal Dempster-Shafer Information Fusion · IEEE Trans. Image Process. 2013
Incremental Learning of 3D-DCT Compact Representations for Robust Visual Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Computer vision › Vision and language
vision-language model
0.812024
BLURD: Benchmarking and Learning using a Unified Rendering and Diffusion Model · NeurIPS 2024
Computer vision › Vision and language › visual question answering
knowledge-based visual question answering
0.622018
Image Captioning and Visual Question Answering Based on Attributes and External Knowledge · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Explicit Knowledge-based Reasoning for Visual Question Answering · IJCAI 2017
Computer vision › 3D vision › geometric deep learning › set learning
set prediction
0.622018
Joint Learning of Set Cardinality and State Distribution · AAAI 2018
DeepSetNet: Predicting Sets with Deep Neural Networks · ICCV 2017
Computer vision › Vision and language
image captioning
0.622018
Image Captioning and Visual Question Answering Based on Attributes and External Knowledge · IEEE Trans. Pattern Anal. Mach. Intell. 2018
What Value Do Explicit High Level Concepts Have in Vision to Language Problems? · CVPR 2016
Computer vision › Video understanding and tracking
multi-object tracking
0.522017
Online Multi-Target Tracking Using Recurrent Neural Networks · AAAI 2017
Joint Probabilistic Data Association Revisited · ICCV 2015
Geometric modeling and processing
3d reconstruction
0.432015
Part-based modelling of compound scenes from images · CVPR 2015
Interactive modelling for AR applications · ISMAR 2010
VideoTrace: rapid interactive scene modelling from video · ACM Trans. Graph. 2007
Computer vision › Video understanding and tracking › object tracking
appearance modeling
0.322013
Incremental Learning of 3D-DCT Compact Representations for Robust Visual Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Learning Compact Binary Codes for Visual Tracking · CVPR 2013
Computer vision › Vision and language › image captioning › fine-grained image captioning
attribute-based image captioning
0.312018
Image Captioning and Visual Question Answering Based on Attributes and External Knowledge · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.312018
FVQA: Fact-Based Visual Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Computer vision › Vision and language › visual question answering
fact-based visual question answering
0.312018
FVQA: Fact-Based Visual Question Answering · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Machine learning › Deep learning architectures and training
memory-augmented neural networks
0.312018
Visual Question Answering With Memory-Augmented Networks · CVPR 2018
Machine learning › Learning paradigms
multi-label classification
0.312018
Joint Learning of Set Cardinality and State Distribution · AAAI 2018
Computer vision › Video understanding and tracking › multi-object tracking
data association
0.312017
Online Multi-Target Tracking Using Recurrent Neural Networks · AAAI 2017
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
dirichlet process mixture model
0.312017
Infinite Variational Autoencoder for Semi-Supervised Learning · CVPR 2017
Machine learning › Learning paradigms
semi-supervised learning
0.312017
Infinite Variational Autoencoder for Semi-Supervised Learning · CVPR 2017
Machine learning › Generative modeling
variational autoencoder
0.312017
Infinite Variational Autoencoder for Semi-Supervised Learning · CVPR 2017
Mathematical optimization › combinatorial optimization › assignment problem
quadratic assignment problem
0.312017
Data-Driven Approximations to NP-Hard Problems · AAAI 2017
Mathematical optimization › combinatorial optimization › vehicle routing
traveling salesman problem
0.312017
Data-Driven Approximations to NP-Hard Problems · AAAI 2017
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge base
0.212016
Ask Me Anything: Free-Form Visual Question Answering Based on Knowledge from External Sources · CVPR 2016
Robotics › Robot manipulation › object perception
object identification
0.212016
Online Metric-Weighted Linear Representations for Robust Visual Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Face, body and person analysis
person re-identification
0.212016
Joint Probabilistic Matching Using m-Best Solutions · CVPR 2016
Computer vision › 3D vision › feature matching
point correspondence
0.212016
Joint Probabilistic Matching Using m-Best Solutions · CVPR 2016
Machine learning › Generative modeling
diffusion model
0.212024
BLURD: Benchmarking and Learning using a Unified Rendering and Diffusion Model · NeurIPS 2024
Machine learning › Generative modeling › diffusion model › latent diffusion model
stable diffusion
0.212024
BLURD: Benchmarking and Learning using a Unified Rendering and Diffusion Model · NeurIPS 2024
Computer vision › Segmentation and scene understanding › scene understanding
semantic scene understanding
0.212015
A fast, modular scene understanding system using context-aware object detection · ICRA 2015
Geometric modeling and processing › shape modeling › 3d object modeling
part-based modeling
0.212015
Part-based modelling of compound scenes from images · CVPR 2015
Data mining
clustering
0.212014
Context-Aware Hypergraph Construction for Robust Spectral Clustering · IEEE Trans. Knowl. Data Eng. 2014
Data mining › clustering › spectral clustering
hypergraph spectral clustering
0.212014
Context-Aware Hypergraph Construction for Robust Spectral Clustering · IEEE Trans. Knowl. Data Eng. 2014

Methods — techniques the papers use, named apart from their topics

recurrent neural network · 1.7stable diffusion · 1.53d rendering · 1.5m-best solutions · 0.7integer linear programming · 0.7convolutional neural network · 0.6permutation-invariant architecture · 0.3large-margin learning · 0.3external memory · 0.3deep learning · 0.3column generation · 0.3attention · 0.3quadratic cost function · 0.2marginal distribution approximation · 0.2sparse estimation · 0.2combinatorial optimization · 0.2spectral clustering · 0.2hypergraph construction · 0.2
YearPublicationVenuePosition
2024 BLURD: Benchmarking and Learning using a Unified Rendering and Diffusion Model
abstract
Recent advancements in pre-trained vision models have made them pivotal in computer vision, emphasizing the need for their thorough evaluation and benchmarking. This evaluation needs to consider various factors of variation, their potential biases, shortcuts, and inaccuracies that ultimately lead to disparate performance in models. Such evaluations are conventionally done using either synthetic data from 2D or 3D rendering software or real-world images in controlled settings. Synthetic methods offer full control and flexibility, while real-world methods are limited by high costs and less adaptability. Moreover, 3D rendering can't yet fully replicate real photography, creating a realism gap.In this paper, we introduce BLURD--Benchmarking and Learning using a Unified Rendering and Diffusion Model--a novel method combining 3D rendering and Stable Diffusion to bridge this gap in representation learning. With BLURD we create a new family of datasets that allow for the creation of both 3D rendered and photo-realistic images with identical factors. BLURD, therefore, provides deeper insights into the representations learned by various CLIP backbones. The source code for creating the BLURD datasets is available at https://github.com/squaringTheCircle/BLURD
Boris Repasky, Ehsan Abbasnejad, Anthony R. Dick
NeurIPS3
2018 Joint Learning of Set Cardinality and State Distribution
abstract
We present a novel approach for learning to predict sets using deep learning. In recent years, deep neural networks have shown remarkable results in computer vision, natural language processing and other related problems. Despite their success,traditional architectures suffer from a serious limitation in that they are built to deal with structured input and output data,i.e. vectors or matrices. Many real-world problems, however, are naturally described as sets, rather than vectors. Existing techniques that allow for sequential data, such as recurrent neural networks, typically heavily depend on the input and output order and do not guarantee a valid solution. Here, we derive in a principled way, a mathematical formulation for set prediction where the output is permutation invariant. In particular, our approach jointly learns both the cardinality and the state distribution of the target set. We demonstrate the validity of our method on the task of multi-label image classification and achieve a new state of the art on the PASCAL VOC and MS COCO datasets.
Seyed Hamid Rezatofighi, Anton Milan, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001
AAAI4
2018 Active Learning from Noisy Tagged Images
Ehsan Abbasnejad, Anthony R. Dick, Qinfeng Shi, Anton van den Hengel
BMVC2
2018 Visual Question Answering With Memory-Augmented Networks
abstract
In this paper, we exploit memory-augmented neural networks to predict accurate answers to visual questions, even when those answers rarely occur in the training set. The memory network incorporates both internal and external memory blocks and selectively pays attention to each training exemplar. We show that memory-augmented neural networks are able to maintain a relatively long-term memory of scarce training exemplars, which is important for visual question answering due to the heavy-tailed distribution of answers in a general VQA setting. Experimental results in two large-scale benchmark datasets show the favorable performance of the proposed algorithm with the comparison to state of the art.
Chao Ma 0004, Chunhua Shen, Anthony R. Dick, Qi Wu 0001, Peng Wang 0023, Anton van den Hengel, Ian D. Reid 0001
CVPR3
2018 FVQA: Fact-Based Visual Question Answering
abstract
Visual Question Answering (VQA) has attracted much attention in both computer vision and natural language processing communities, not least because it offers insight into the relationships between two important sources of information. Current datasets, and the models built upon them, have focused on questions which are answerable by direct analysis of the question and image alone. The set of such questions that require no external information to answer is interesting, but very limited. It excludes questions which require common sense, or basic factual knowledge to answer, for example. Here we introduce FVQA (Fact-based VQA), a VQA dataset which requires, and supports, much deeper reasoning. FVQA primarily contains questions that require external information to answer. We thus extend a conventional visual question answering dataset, which contains image-question-answer triplets, through additional image-question-answer-supporting fact tuples. Each supporting-fact is represented as a structural triplet, such as .
Peng Wang 0015, Qi Wu 0001, Chunhua Shen, Anthony R. Dick, Anton van den Hengel
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Image Captioning and Visual Question Answering Based on Attributes and External Knowledge
abstract
Much of the recent progress in Vision-to-Language problems has been achieved through a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This approach does not explicitly represent high-level semantic concepts, but rather seeks to progress directly from image features to text. In this paper we first propose a method of incorporating high-level concepts into the successful CNN-RNN approach, and show that it achieves a significant improvement on the state-of-the-art in both image captioning and visual question answering. We further show that the same mechanism can be used to incorporate external knowledge, which is critically important for answering high level visual questions. Specifically, we design a visual question answering model that combines an internal representation of the content of an image with information extracted from a general knowledge base to answer a broad range of image-based questions. It particularly allows questions to be asked where the image alone does not contain the information required to select the appropriate answer. Our final model achieves the best reported results for both image captioning and visual question answering on several of the major benchmark datasets.
Qi Wu 0001, Chunhua Shen, Peng Wang 0015, Anthony R. Dick, Anton van den Hengel
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 Online Multi-Target Tracking Using Recurrent Neural Networks
abstract
We present a novel approach to online multi-target tracking based on recurrent neural networks (RNNs). Tracking multiple objects in real-world scenes involves many challenges, including a) an a-priori unknown and time-varying number of targets, b) a continuous state estimation of all present targets, and c) a discrete combinatorial problem of data association. Most previous methods involve complex models that require tedious tuning of parameters. Here, we propose for the first time, an end-to-end learning approach for online multi-target tracking. Existing deep learning methods are not designed for the above challenges and cannot be trivially applied to the task. Our solution addresses all of the above points in a principled way. Experiments on both synthetic and real data show promising results obtained at ~300 Hz on a standard CPU, and pave the way towards future research in this direction.
Anton Milan, Seyed Hamid Rezatofighi, Anthony R. Dick, Ian D. Reid 0001, Konrad Schindler
AAAI3
2017 Data-Driven Approximations to NP-Hard Problems
abstract
There exist a number of problem classes for which obtaining the exact solution becomes exponentially expensive with increasing problem size. The quadratic assignment problem (QAP) or the travelling salesman problem (TSP) are just two examples of such NP-hard problems. In practice, approximate algorithms are employed to obtain a suboptimal solution, where one must face a trade-off between computational complexity and solution quality. In this paper, we propose to learn to solve these problem from approximate examples, using recurrent neural networks (RNNs). Surprisingly, such architectures are capable of producing highly accurate solutions at minimal computational cost. Moreover, we introduce a simple, yet effective technique for improving the initial (weak) training set by incorporating the objective cost into the training procedure. We demonstrate the functionality of our approach on three exemplar applications: marginal distributions of a joint matching space, feature point matching and the travelling salesman problem. We show encouraging results on synthetic and real data in all three cases.
Anton Milan, Seyed Hamid Rezatofighi, Ravi Garg, Anthony R. Dick, Ian D. Reid 0001
AAAI4
2017 Large-Scale Camera Network Topology Estimation by Lighting Variation
Michael Zhu, Anthony R. Dick, Anton van den Hengel
ACIVS2
2017 Infinite Variational Autoencoder for Semi-Supervised Learning
abstract
This paper presents an infinite variational autoencoder (VAE) whose capacity adapts to suit the input data. This is achieved using a mixture model where the mixing coefficients are modeled by a Dirichlet process, allowing us to integrate over the coefficients when performing inference. Critically, this then allows us to automatically vary the number of autoencoders in the mixture based on the data. Experiments show the flexibility of our method, particularly for semi-supervised learning, where only a small number of training samples are available.
Ehsan Abbasnejad, Anthony R. Dick, Anton van den Hengel
CVPR2
2017 DeepSetNet: Predicting Sets with Deep Neural Networks
abstract
This paper addresses the task of set prediction using deep learning. This is important because the output of many computer vision tasks, including image tagging and object detection, are naturally expressed as sets of entities rather than vectors. As opposed to a vector, the size of a set is not fixed in advance, and it is invariant to the ordering of entities within it. We define a likelihood for a set distribution and learn its parameters using a deep neural network. We also derive a loss for predicting a discrete distribution corresponding to set cardinality. Set prediction is demonstrated on the problem of multi-class image classification. Moreover, we show that the proposed cardinality loss can also trivially be applied to the tasks of object counting and pedestrian detection. Our approach outperforms existing methods in all three cases on standard datasets.
Seyed Hamid Rezatofighi, Anton Milan, Ehsan Abbasnejad, Anthony R. Dick, Ian D. Reid 0001
ICCV5
2017 Explicit Knowledge-based Reasoning for Visual Question Answering
abstract
We describe a method for visual question answering which is capable of reasoning about an image on the basis of information extracted from a large-scale knowledge base. The method not only answers natural language questions using concepts not contained in the image, but can explain the reasoning by which it developed its answer. It is capable of answering far more complex questions than the predominant long short-term memory-based approach, and outperforms it significantly in testing. We also provide a dataset and a protocol by which to evaluate general visual question answering methods.
Peng Wang 0015, Qi Wu 0001, Chunhua Shen, Anthony R. Dick, Anton van den Hengel
IJCAI4
2017 Visual question answering: A survey of methods and datasets
Qi Wu 0001, Damien Teney, Peng Wang 0015, Chunhua Shen, Anthony R. Dick, Anton van den Hengel
Comput. Vis. Image Underst.5
2016 Joint Probabilistic Matching Using m-Best Solutions
abstract
Matching between two sets of objects is typically approached by finding the object pairs that collectively maximize the joint matching score. In this paper, we argue that this single solution does not necessarily lead to the optimal matching accuracy and that general one-to-one assignment problems can be improved by considering multiple hypotheses before computing the final similarity measure. To that end, we propose to utilize the marginal distributionsfor each entity. Previously, this idea has been neglected mainly because exact marginalization is intractable due to a combinatorial number of all possible matching permutations. Here, we propose a generic approach to efficiently approximate the marginal distributions by exploiting the m-best solutions of the original problem. This approach not only improves the matching solution, but also provides more accurate ranking of the results, because of the extra information included in the marginal distribution. We validate our claim on two distinct objectives: (i) person re-identification and temporal matching modeled as an integer linear program, and (ii) feature point matching using a quadratic cost function. Our experiments confirm that marginalization indeed leads to superior performance compared to the single (nearly) optimal solution, yielding state-of-the-art results in both applications on standard benchmarks.
Seyed Hamid Rezatofighi, Anton Milan, Zhen Zhang 0008, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001
CVPR5
2016 What Value Do Explicit High Level Concepts Have in Vision to Language Problems?
abstract
Much recent progress in Vision-to-Language (V2L) problems has been achieved through a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This approach does not explicitly represent high-level semantic concepts, but rather seeks to progress directly from image features to text. In this paper we investigate whether this direct approach succeeds due to, or despite, the fact that it avoids the explicit representation of high-level information. We propose a method of incorporating high-level concepts into the successful CNN-RNN approach, and show that it achieves a significant improvement on the state-of-the-art in both image captioning and visual question answering. We also show that the same mechanism can be used to introduce external semantic information and that doing so further improves performance. We achieve the best reported results on both image captioning and VQA on several benchmark datasets, and provide an analysis of the value of explicit high-level concepts in V2L problems.
Qi Wu 0001, Chunhua Shen, Lingqiao Liu, Anthony R. Dick, Anton van den Hengel
CVPR4
2016 Ask Me Anything: Free-Form Visual Question Answering Based on Knowledge from External Sources
abstract
We propose a method for visual question answering which combines an internal representation of the content of an image with information extracted from a general knowledge base to answer a broad range of image-based questions. This allows more complex questions to be answered using the predominant neural network-based approach than has previously been possible. It particularly allows questions to be asked about the contents of an image, even when the image itself does not contain the whole answer. The method constructs a textual representation of the semantic content of an image, and merges it with textual information sourced from a knowledge base, to develop a deeper understanding of the scene viewed. Priming a recurrent neural network with this combined information, and the submitted question, leads to a very flexible visual question answering approach. We are specifically able to answer questions posed in natural language, that refer to information not contained in the image. We demonstrate the effectiveness of our model on two publicly available datasets, Toronto COCO-QA [23] and VQA [1] and show that it produces the best reported results in both cases.
Qi Wu 0001, Peng Wang 0015, Chunhua Shen, Anthony R. Dick, Anton van den Hengel
CVPR4
2016 Online Metric-Weighted Linear Representations for Robust Visual Tracking
abstract
In this paper, we propose a visual tracker based on a metric-weighted linear representation of appearance. In order to capture the interdependence of different feature dimensions, we develop two online distance metric learning methods using proximity comparison information and structured output learning. The learned metric is then incorporated into a linear representation of appearance. We show that online distance metric learning significantly improves the robustness of the tracker, especially on those sequences exhibiting drastic appearance changes. In order to bound growth in the number of training samples, we design a time-weighted reservoir sampling method. Moreover, we enable our tracker to automatically perform object identification during the process of object tracking, by introducing a collection of static template samples belonging to several object classes of interest. Object identification results for an entire video sequence are achieved by systematically combining the tracking information and visual recognition at each frame. Experimental results on challenging video sequences demonstrate the effectiveness of the method for both inter-frame tracking and object identification.
Xi Li 0001, Chunhua Shen, Anthony R. Dick, Zhongfei Zhang, Yueting Zhuang
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Part-based modelling of compound scenes from images
abstract
We propose a method to recover the structure of a compound scene from multiple silhouettes. Structure is expressed as a collection of 3D primitives chosen from a predefined library, each with an associated pose. This has several advantages over a volume or mesh representation both for estimation and the utility of the recovered model. The main challenge in recovering such a model is the combinatorial number of possible arrangements of parts. We address this issue by exploiting the intrinsic structure and sparsity of the problem, and show that our method scales to scenes constructed from large libraries of parts.
Anton van den Hengel, Chris Russell 0001, Anthony R. Dick, John W. Bastian, Daniel Pooley, Lachlan Fleming, Lourdes Agapito
CVPR3
2015 Joint Probabilistic Data Association Revisited
abstract
In this paper, we revisit the joint probabilistic data association (JPDA) technique and propose a novel solution based on recent developments in finding the m-best solutions to an integer linear program. The key advantage of this approach is that it makes JPDA computationally tractable in applications with high target and/or clutter density, such as spot tracking in fluorescence microscopy sequences and pedestrian tracking in surveillance footage. We also show that our JPDA algorithm embedded in a simple tracking framework is surprisingly competitive with state-of-the-art global tracking methods in these two applications, while needing considerably less processing time.
Seyed Hamid Rezatofighi, Anton Milan, Zhen Zhang 0008, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001
ICCV5
2015 A fast, modular scene understanding system using context-aware object detection
abstract
We propose a semantic scene understanding system that is suitable for real robotic operations. The system solves different tasks (semantic segmentation and object detections) in an opportunistic and distributed fashion but still allows communication between modules to improve their respective performances. We propose the use of the semantic space to improve specific out-of-the-box object detectors and an update model to take the evidence from different detection into account in the semantic segmentation process. Our proposal is evaluated with the KITTI dataset, on the object detection benchmark and on five different sequences manually annotated for the semantic segmentation task, demonstrating the efficacy of our approach.
Cesar Dario Cadena Lerma, Anthony R. Dick, Ian D. Reid 0001
ICRA2
2014 Context Based Re-ranking for Object Retrieval
Yanzhi Chen, Anthony R. Dick, Xi Li 0001, Rhys Hill
ACCV (1)2
2014 Ranking consistency for image matching and object retrieval
Yanzhi Chen, Xi Li 0001, Anthony R. Dick, Rhys Hill
Pattern Recognit.3
2014 Compass Rose: A Rotational Robust Signature for Optical Flow Computation
abstract
This paper proposes a new image signature, called Compass Rose, which is particularly suited to optical flow computation. It is differentiable, fast to compute, and robust to additive illumination changes, translation, and fast rotation. We design a sparse flow computation system based on the invariance of the Compass Rose signatures. This is then extended to dense motion estimation by the addition of an optional diffusion step. Quantitative testing on several benchmark sequences shows the Compass Rose attains higher accuracy than the traditional flow signatures under a range of conditions. Finally, we demonstrate its application to human motion estimation, which is challenging for optical flow methods due to fast limb rotation.
Yan Niu, Anthony R. Dick, Michael J. Brooks
IEEE Trans. Circuits Syst. Video Technol.2
2014 Context-Aware Hypergraph Construction for Robust Spectral Clustering
abstract
Spectral clustering is a powerful tool for unsupervised data analysis. In this paper, we propose a context-aware hypergraph similarity measure (CAHSM), which leads to robust spectral clustering in the case of noisy data. We construct three types of hypergraphs-the pairwise hypergraph, the k-nearest-neighbor (kNN) hypergraph, and the high-order over-clustering hypergraph. The pairwise hypergraph captures the pairwise similarity of data points; the kNNhypergraph captures the neighborhood of each point; and the clustering hypergraph encodes high-order contexts within the dataset. By combining the affinity information from these three hypergraphs, the CAHSM algorithm is able to explore the intrinsic topological information of the dataset. Therefore, data clustering using CAHSM tends to be more robust. Considering the intra-cluster compactness and the inter-cluster separability of vertices, we further design a discriminative hypergraph partitioning criterion (DHPC). Using both CAHSM and DHPC, a robust spectral clustering algorithm is developed. Theoretical analysis and experimental evaluation demonstrate the effectiveness and robustness of the proposed algorithm.
Xi Li 0001, Weiming Hu 0004, Chunhua Shen, Anthony R. Dick, Zhongfei Zhang
IEEE Trans. Knowl. Data Eng.4
2013 Learning Compact Binary Codes for Visual Tracking
abstract
A key problem in visual tracking is to represent the appearance of an object in a way that is robust to visual changes. To attain this robustness, increasingly complex models are used to capture appearance variations. However, such models can be difficult to maintain accurately and efficiently. In this paper, we propose a visual tracker in which objects are represented by compact and discriminative binary codes. This representation can be processed very efficiently, and is capable of effectively fusing information from multiple cues. An incremental discriminative learner is then used to construct an appearance model that optimally separates the object from its surrounds. Furthermore, we design a hyper graph propagation method to capture the contextual information on samples, which further improves the tracking accuracy. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracker.
Xi Li 0001, Chunhua Shen, Anthony R. Dick, Anton van den Hengel
CVPR3
2013 Contextual Hypergraph Modeling for Salient Object Detection
abstract
Salient object detection aims to locate objects that capture human attention within images. Previous approaches often pose this as a problem of image contrast analysis. In this work, we model an image as a hyper graph that utilizes a set of hyper edges to capture the contextual properties of image pixels or regions. As a result, the problem of salient object detection becomes one of finding salient vertices and hyper edges in the hyper graph. The main advantage of hyper graph modeling is that it takes into account each pixel's (or region's) affinity with its neighborhood as well as its separation from image background. Furthermore, we propose an alternative approach based on center-versus-surround contextual contrast analysis, which performs salient object detection by optimizing a cost-sensitive support vector machine (SVM) objective function. Experimental results on four challenging datasets demonstrate the effectiveness of the proposed approaches against the state-of-the-art approaches to salient object detection.
Xi Li 0001, Yao Li 0003, Chunhua Shen, Anthony R. Dick, Anton van den Hengel
ICCV4
2013 Learning Hash Functions Using Column Generation
abstract
Fast nearest neighbor searching is becoming an increasingly important tool in solving many large-scale problems. Recently a number of approaches to learning data-dependent hash functions have been developed. In this work, we propose a column generation based method for learning data-dependent hash functions on the basis of proximity comparison information. Given a set of triplets that encode the pairwise proximity comparison information, our method learns hash functions that preserve the relative comparison relationships in the data as well as possible within the large-margin learning framework. The learning procedure is implemented using column generation and hence is named CGHash. At each iteration of the column generation procedure, the best hash function is selected. Unlike most other hashing methods, our method generalizes to new data points naturally; and has a training objective which is convex, thus ensuring that the global optimum can be identified. Experiments demonstrate that the proposed method learns compact binary codes and that its retrieval performance compares favorably with state-of-the-art methods when tested on a few benchmark datasets.
Xi Li 0001, Guosheng Lin, Chunhua Shen, Anton van den Hengel, Anthony R. Dick
ICML (1)5
2013 Adaptive earth movers distance-based Bayesian multi-target tracking
abstract
This study describes a complete system for multiple‐target tracking in image sequences. The target appearance is represented as a set of weighted clusters in colour space. This is in contrast to the more typical use of colour histograms to model target appearance. The use of clusters allows a more flexible and accurate representation of the target, which demonstrates the benefits for tracking. However, it also introduces a number of computational difficulties, as calculating and matching cluster signatures are both computationally intensive tasks. To overcome this, the authors introduce a new formulation of incremental medoid‐shift clustering that operates faster than mean shift in multi‐target tracking scenarios. This matching scheme is integrated into a Bayesian tracking framework. Particle filters, a special case of Bayesian filters where the state variables are non‐linear and non‐Gaussian, are used in this study. An adaptive model update procedure is proposed for the cluster signature representation to handle target changes with time. The model update procedure is demonstrated to work successfully on a synthetic dataset and then on real datasets. Successful tracking results are shown on public datasets. Both qualitative and quantitative evaluations have been carried out to demonstrate the improved performance of the proposed multi‐target tracking system. A higher tracking accuracy in long image sequences has been achieved compared to other standard tracking methods.
Pankaj Kumar 0001, Anthony R. Dick
IET Comput. Vis.2
2013 Spatially aware feature selection and weighting for object retrieval
Yanzhi Chen, Anthony R. Dick, Xi Li 0001, Anton van den Hengel
Image Vis. Comput.2
2013 Incremental Learning of 3D-DCT Compact Representations for Robust Visual Tracking
abstract
Visual tracking usually requires an object appearance model that is robust to changing illumination, pose, and other factors encountered in video. Many recent trackers utilize appearance samples in previous frames to form the bases upon which the object appearance model is built. This approach has the following limitations: 1) The bases are data driven, so they can be easily corrupted, and 2) it is difficult to robustly update the bases in challenging situations. In this paper, we construct an appearance model using the 3D discrete cosine transform (3D-DCT). The 3D-DCT is based on a set of cosine basis functions which are determined by the dimensions of the 3D signal and thus independent of the input video data. In addition, the 3D-DCT can generate a compact energy spectrum whose high-frequency coefficients are sparse if the appearance samples are similar. By discarding these high-frequency coefficients, we simultaneously obtain a compact 3D-DCT-based object representation and a signal reconstruction-based similarity measure (reflecting the information loss from signal reconstruction). To efficiently update the object representation, we propose an incremental 3D-DCT algorithm which decomposes the 3D-DCT into successive operations of the 2D discrete cosine transform (2D-DCT) and 1D discrete cosine transform (1D-DCT) on the input video data. As a result, the incremental 3D-DCT algorithm only needs to compute the 2D-DCT for newly added frames as well as the 1D-DCT along the third dimension, which significantly reduces the computational complexity. Based on this incremental 3D-DCT algorithm, we design a discriminative criterion to evaluate the likelihood of a test sample belonging to the foreground object. We then embed the discriminative criterion into a particle filtering framework for object state inference over time. Experimental results demonstrate the effectiveness and robustness of the proposed tracker.
Xi Li 0001, Anthony R. Dick, Chunhua Shen, Anton van den Hengel, Hanzi Wang
IEEE Trans. Pattern Anal. Mach. Intell.2
2013 Visual Tracking With Spatio-Temporal Dempster-Shafer Information Fusion
abstract
A key problem in visual tracking is how to effectively combine spatio-temporal visual information from throughout a video to accurately estimate the state of an object. We address this problem by incorporating Dempster-Shafer (DS) information fusion into the tracking approach. To implement this fusion task, the entire image sequence is partitioned into spatially and temporally adjacent subsequences. A support vector machine (SVM) classifier is trained for object/nonobject classification on each of these subsequences, the outputs of which act as separate data sources. To combine the discriminative information from these classifiers, we further present a spatio-temporal weighted DS (STWDS) scheme. In addition, temporally adjacent sources are likely to share discriminative information on object/nonobject classification. To use such information, an adaptive SVM learning scheme is designed to transfer discriminative information across sources. Finally, the corresponding DS belief function of the STWDS scheme is embedded into a Bayesian tracking model. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracking approach.
Xi Li 0001, Anthony R. Dick, Chunhua Shen, Zhongfei Zhang, Anton van den Hengel, Hanzi Wang
IEEE Trans. Image Process.2
2013 A survey of appearance models in visual object tracking
abstract
Visual object tracking is a significant computer vision task which can be applied to many domains, such as visual surveillance, human computer interaction, and video compression. Despite extensive research on this topic, it still suffers from difficulties in handling complex object appearance changes caused by factors such as illumination variation, partial occlusion, shape deformation, and camera motion. Therefore, effective modeling of the 2D appearance of tracked objects is a key issue for the success of a visual tracker. In the literature, researchers have proposed a variety of 2D appearance models. To help readers swiftly learn the recent advances in 2D appearance models for visual object tracking, we contribute this survey, which provides a detailed review of the existing 2D appearance models. In particular, this survey takes a module-based architecture that enables readers to easily grasp the key points of visual object tracking. In this survey, we first decompose the problem of appearance modeling into two different processing stages: visual representation and statistical modeling. Then, different 2D appearance models are categorized and discussed with respect to their composition modules. Finally, we address several issues of interest as well as the remaining challenges for future research on this topic. The contributions of this survey are fourfold. First, we review the literature of visual representations according to their feature-construction mechanisms (i.e., local and global). Second, the existing statistical modeling schemes for tracking-by-detection are reviewed according to their model-construction mechanisms: generative, discriminative, and hybrid generative-discriminative. Third, each type of visual representations or statistical modeling techniques is analyzed and discussed from a theoretical or practical viewpoint. Fourth, the existing benchmark resources (e.g., source codes and video datasets) are examined in this survey.
Xi Li 0001, Weiming Hu 0004, Chunhua Shen, Zhongfei Zhang, Anthony R. Dick, Anton van den Hengel
ACM Trans. Intell. Syst. Technol.5
2012 Non-sparse linear representations for visual tracking with online reservoir metric learning
abstract
Most sparse linear representation-based trackers need to solve a computationally expensive li-regularized optimization problem. To address this problem, we propose a visual tracker based on non-sparse linear representations, which admit an efficient closed-form solution without sacrificing accuracy. Moreover, in order to capture the correlation information between different feature dimensions, we learn a Mahalanobis distance metric in an online fashion and incorporate the learned metric into the optimization problem for obtaining the linear representation. We show that online metric learning using proximity comparison significantly improves the robustness of the tracking, especially on those sequences exhibiting drastic appearance changes. Furthermore, in order to prevent the unbounded growth in the number of training samples for the metric learning, we design a time-weighted reservoir sampling method to maintain and update limited-sized foreground and background sample buffers for balancing sample diversity and adaptability. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracker.
Xi Li 0001, Chunhua Shen, Qinfeng Shi, Anthony R. Dick, Anton van den Hengel
CVPR4
2012 Boosting Object Retrieval With Group Queries
abstract
Given a query image of an object, object retrieval aims to return all images from a corpus that depict the same object. Inevitably, the accuracy of the result depends strongly on the quality of the query image. Several measures have been taken to improve retrieval result quality, including the addition of a bounding box to the query, the mining of highly ranked results for more views of the object, and spatial consistency re-ranking. In this letter, we propose a discriminative criterion for improving result quality. This criterion lends itself to the addition of extra query data, and we show that multiple query images can be combined to produce enhanced results. Experiments compare the performance of the method to state-of-the-art in object retrieval, and show how performance is lifted by the inclusion of further query images.
Yanzhi Chen, Xi Li 0001, Anthony R. Dick, Anton van den Hengel
IEEE Signal Process. Lett.3
2012 Locally Oriented Optical Flow Computation
abstract
This paper proposes the use of an adaptive locally oriented coordinate frame when calculating an optical flow field. The coordinate frame is aligned with the least curvature direction in a local window about each pixel. This has advantages to both fitting the flow field to the image data and in imposing smoothness constraints between neighboring pixels. In terms of fitting, robustness is obtained to a wider variety of image motions due to the extra invariance provided by the coordinate frame. Smoothness constraints are naturally propagated along image boundaries which often correspond to motion boundaries. In addition, moving objects can be efficiently segmented in the least curvature direction. We show experimentally the benefits of the method and demonstrate robustness to fast rotational motion, such as what often occurs in human motion.
Yan Niu, Anthony R. Dick, Michael J. Brooks
IEEE Trans. Image Process.2
2011 Graph mode-based contextual kernels for robust SVM tracking
abstract
Visual tracking has been typically solved as a binary classification problem. Most existing trackers only consider the pairwise interactions between samples, and thereby ignore the higher-order contextual interactions, which may lead to the sensitivity to complicated factors such as noises, outliers, background clutters and so on. In this paper, we propose a visual tracker based on support vector machines (SVMs), for which a novel graph mode-based contextual kernel is designed to effectively capture the higher-order contextual information from samples. To do so, we first create a visual graph whose similarity matrix is determined by a baseline visual kernel. Second, a set of high-order contexts are discovered in the visual graph. The problem of discovering these high-order contexts is solved by seeking modes of the visual graph. Each graph mode corresponds to a vertex community termed as a high-order context. Third, we construct a contextual kernel that effectively captures the interaction information between the high-order contexts. Finally, this contextual kernel is embedded into SVMs for robust tracking. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracker.
Xi Li 0001, Anthony R. Dick, Hanzi Wang, Chunhua Shen, Anton van den Hengel
ICCV2
2010 Interactive modelling for AR applications
abstract
We present a method for estimating the 3D shape of an object from a sequence of images captured by a hand-held device. The method is well suited to augmented reality applications in that minimal user interaction is required, and the models generated are of an appropriate form. The method proceeds by segmenting the object in every image as it is captured and using the calculated silhouette to update the current shape estimate. In contrast to previous silhouette-based modelling approaches, however, the segmentation process is informed by a 3D prior based on the previous shape estimate. A voting scheme is also introduced in order to compensate for the inevitable noise in the camera position estimates. The combination of the voting scheme with the closed-loop segmentation process provides a robust and flexible shape estimation method. We demonstrate the approach on a number of scenes where segmentation without a 3D prior would be challenging.
John W. Bastian, Ben Ward, Rhys Hill, Anton van den Hengel, Anthony R. Dick
ISMAR5
2009 Contradiction and Correlation for Camera Overlap Estimation
abstract
An accurate estimate of camera overlap is a key enabler for efficient network-wide surveillance processing (e.g. inter-camera tracking), especially in large-scale surveillance networks. Techniques based on contradictions in pair-wise occupancy data, such as the exclusion approach, have advantages in robustness and efficiency that make them particularly well suited for large surveillance networks. Correlation techniques share some of these advantages,but have a better understood statistical basis. This paper evaluates a set of contradiction and correlation techniques, using a novel metric, search space precision-recall. This metric reflects the activity-based overlap estimation required for camera handover, such as would be used in inter-camera tracking.Results are reported for a range of networks, including a 24-camera network setup in an office space, where the exclusion estimator showed the best performance.
Alex Cichowski, Christopher S. Madden, Anton van den Hengel, Rhys Hill, Henry Detmold, Anthony R. Dick
AVSS6
2009 In situ image-based modeling
abstract
We present an interactive image-based modelling method for generating 3D models within an augmented reality system. Applying real time camera tracking, and high-level automated image analysis, enables more powerful modelling interactions than have previously been possible. The result is an immersive modelling process which generates accurate three dimensional models of real objects efficiently and effectively. In demonstrating the modelling process on a range of indoor and outdoor scenes, we show the flexibility it offers in enabling augmented reality applications in previously unseen environments.
Anton van den Hengel, Rhys Hill, Ben Ward, Anthony R. Dick
ISMAR4
2009 Automatic camera placement for large scale surveillance networks
abstract
Automatic placement of surveillance cameras in arbitrary buildings is a challenging task, and also one that is essential for efficient deployment of large scale surveillance networks. Existing approaches for automatic camera placement are either limited to a small number of cameras, or constrained in terms of the building layouts to which they can be applied. This paper describes a new method for determining the best placement for large numbers of cameras within arbitrary building layouts. The method takes as input a 3D model of the building, and uses a genetic algorithm to find a placement that optimises coverage and (if desired) overlap between cameras. Results are reported for an implementation of the method, including its application to a wide variety of complex buildings, both real and synthetic.
Anton van den Hengel, Rhys Hill, Ben Ward, Alex Cichowski, Henry Detmold, Christopher S. Madden, Anthony R. Dick, John W. Bastian
WACV7
2009 SecondSkin: An interactive method for appearance transfer
abstract
Abstract SecondSkin estimates an appearance model for an object visible in a video sequence, without the need for complex interaction or any calibration apparatus. This model can then be transferred to other objects, allowing a non‐expert user to insert a synthetic object into a real video sequence so that its appearance matches that of an existing object, and changes appropriately throughout the sequence. As the method does not require any prior knowledge about the scene, the lighting conditions, or the camera, it is applicable to video which was not captured with this purpose in mind. However, this lack of prior knowledge precludes the recovery of separate lighting and surface reflectance information. The SecondSkin appearance model therefore combines these factors. The appearance model does require a dominant light‐source direction, which we estimate via a novel process involving a small amount of user interaction. The resulting model estimate provides exactly the information required to transfer the appearance of the original object to new geometry composited into the same video sequence.
Anton van den Hengel, Dylan Sale, Anthony R. Dick
Comput. Graph. Forum3
2008 Multiple target tracking with an efficient compact colour correlogram
abstract
A robust approach to detection and tracking of multiple moving targets from a moving camera is presented. The main novelty of this approach is that objects are represented using efficient compact form of the colour correlogram. Like previous correlograms, this encodes both spatial pattern and appearance information about the target. However it is less complex to compute, making it applicable to real time target tracking. The correlogram representation is incorporated into a particle filter tracking framework. Robustness to camera motion is obtained by identifying homographies linking adjacent frames, and using them to align corresponding image areas to a common local reference frame. We demonstrate successful detection and tracking in real life image sequences.
Pankaj Kumar 0001, Anthony R. Dick, Michael J. Brooks
ICARCV2
2007 Finding Camera Overlap in Large Surveillance Networks
Anton van den Hengel, Anthony R. Dick, Henry Detmold, Alex Cichowski, Rhys Hill
ACCV (1)2
2007 Adaptive Multiple Object Tracking Using Colour and Segmentation Cues
Pankaj Kumar 0001, Michael J. Brooks, Anthony R. Dick
ACCV (1)3
2007 VideoTrace: rapid interactive scene modelling from video
abstract
VideoTrace is a system for interactively generating realistic 3D models of objects from video---models that might be inserted into a video game, a simulation environment, or another video sequence. The user interacts with VideoTrace by tracing the shape of the object to be modelled over one or more frames of the video. By interpreting the sketch drawn by the user in light of 3D information obtained from computer vision techniques, a small number of simple 2D interactions can be used to generate a realistic 3D model. Each of the sketching operations in VideoTrace provides an intuitive and powerful means of modelling shape from video, and executes quickly enough to be used interactively. Immediate feedback allows the user to model rapidly those parts of the scene which are of interest and to the level of detail required. The combination of automated and manual reconstruction allows VideoTrace to model parts of the scene not visible, and to succeed in cases where purely automated approaches would fail.
Anton van den Hengel, Anthony R. Dick, Thorsten Thormählen, Ben Ward, Philip Torr 0001
ACM Trans. Graph.2
2006 Scalable Surveillance Software Architecture
abstract
Video surveillance is a key technology for enhanced protection of facilities such as airports and power stations from various types of threat. Networks of thousands of IP-based cameras are now possible, but current surveillance methodologies become increasingly ineffective as the number of cameras grows. Constructing software that efficiently and reliably deals with networks of this size is a distributed information processing problem as much as it is a video interpretation challenge. This paper demonstrates a software architecture approach to the construction of large scale surveillance network software and explores the implications for instantiating surveillance algorithms at such a scale. A novel architecture for video surveillance is presented, and its efficacy demonstrated through application to an important class of surveillance algorithms.
Henry Detmold, Anthony R. Dick, Katrina Falkner, David S. Munro, Anton van den Hengel, Ronald Morrison
AVSS2
2006 Activity Topology Estimation for Large Networks of Cameras
abstract
Estimating the paths that moving objects can take through the fields of view of possibly non-overlapping cameras, also known as their activity topology, is an important step in the effective interpretation of surveillance video. Existing approaches to this problem involve tracking moving objects within cameras, and then attempting to link tracks across views. In contrast we propose an approach which begins by assuming all camera views are potentially linked, and successively eliminates camera topologies that are contradicted by observed motion. Over time, the true patterns of motion emerge as those which are not contradicted by the evidence. These patterns may then be used to initialise a finer level search using other approaches if required. This method thus represents an efficient and effective way to learn activity topology for a large network of cameras, particularly with a limited amount of data.
Anton van den Hengel, Anthony R. Dick, Rhys Hill
AVSS2
2006 Building Models of Regular Scenes from Structure and Motion
abstract
This paper describes a method for generating a model-based reconstruction of a scene from image data. The method uses the camera models and point cloud typically generated by a structure-and-motion process as a starting point for developing a higher level model of the scene. The method relies on the user to provide a minimal amount of structural seeding information from which more complex geometry is extrapolated. The regularity typically present in man-made environments is used to minimise the interaction required, but also to improve the accuracy of fit. We demonstrate model based reconstructions obtained using this method. 1
Anton van den Hengel, Anthony R. Dick, Thorsten Thormählen, Ben Ward, Philip Torr 0001
BMVC2
2004 Modelling and Interpretation of Architecture from Several Images
Anthony R. Dick, Philip Torr 0001, Roberto Cipolla
Int. J. Comput. Vis.1
2002 A Bayesian Estimation of Building Shape Using MCMC
Anthony R. Dick, Philip Torr 0001, Roberto Cipolla
ECCV (2)1
2001 Combining Single View Recognition and Multiple View Stereo for Architectural Scenes
Anthony R. Dick, Philip Torr 0001, Simon J. Ruffle, Roberto Cipolla
ICCV1
2000 Automatic 3D Modelling of Architecture
abstract
This paper describes a system which automatically derives 3D models of ar-chitectural scenes from multiple images. This system differs from previous structure from motion algorithms in that it explicitly makes use of strong ge-ometric constraints such as perpendicularity and verticality which are likely to be found in architecture. Structure is compactly represented as a piece-wise planar model which is initialised automatically by segmenting a feature-based reconstruction. An efficient technique for evaluation of model likeli-hood is also presented, which allows a rapid search through a large number of 3D models. 1
Anthony R. Dick, Philip Torr 0001, Roberto Cipolla
BMVC1
2000 Layer Extraction with a Bayesian Model of Shapes
Philip Torr 0001, Anthony R. Dick, Roberto Cipolla
ECCV (2)2
1999 Model Refinement from Planar Parallax
abstract
This paper presents a system for refining the accuracy and realism of coarse piecewise planar models from an uncalibrated sequence of images. First, dense depth maps are estimated by aligning a planar region of a scene in each image, approximating camera calibration, and generating dense planar paral-lax. These depth maps are then robustly fused to obtain incrementally refined surface estimates. It is envisaged that this system will extend the modelling capability of existing systems [3] which generate simple, piecewise planar architectural models. 1
Anthony R. Dick, Roberto Cipolla
BMVC1
1998 Multiresolution stereo image matching using complex wavelets
abstract
This paper describes a multiresolution image-matching strategy, based on the complex discrete wavelet transform (CDWT), to derive a dense disparity field with hierarchical (coarse-to-fine) refinement. The CDWT feature space efficiently provides fractionally accurate matching results which are robust to typical image formation perturbations such as offsets, global scaling, and additive noise. At each level of the hierarchy, the disparity field is regularised to provide a global compromise between feature similarity and disparity field continuity, resulting in feature-sensitive smoothing. The algorithm is well suited to analysing facial images, for which we demonstrate striking reconstruction results.
Julian Magarey, Anthony R. Dick
ICPR2