Ognjen Arandjelovic

dblp:98/5718 · DBLP profile ↗
← Back
79ranked-venue papers
41as first author
15since 2021 · last 2026
0000-0002-9314-194XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 62 · 34 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 24 first-author · 5 since 2021Databases, data management, data science and information retrieval · 8Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSecurity and privacy · 2 · 1 first-author · 1 since 2021Theory of computation · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
YearPublicationVenuePosition
2026 Attacking hard-label large vision-language models with model-sensitive adversarial patch designs
Nian Ai, Guangke Chen, Xiaowen Cai 0001, Zhongliang Guo 0001, Daizong Liu, Pan Zhou 0001, Ognjen Arandjelovic
Pattern Recognit.7
2026 Artwork protection against unauthorized neural style transfer and aesthetic color distance metric
Zhongliang Guo 0001, Yifei Qian, Shuai Zhao 0007, Junhao Dong 0001, Ognjen Arandjelovic, Lei Fang 0001, Chun Pong Lau 0001
Pattern Recognit.6
2026 Multimodal latent emotion recognition from micro-expression and physiological signal
Liangfei Zhang, Yifei Qian, Ognjen Arandjelovic, Hongjiang Xiao
Pattern Recognit.3
2025 LLM Agents Making Agent Tools
abstract
Georg Wölflein, Dyke Ferber, Daniel Truhn, Ognjen Arandjelovic, Jakob Nikolas Kather. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Georg Wölflein, Dyke Ferber, Daniel Truhn, Ognjen Arandjelovic, Jakob Nikolas Kather
ACL (1)4
2025 Perspective-assisted prototype-based learning for semi-supervised crowd counting
Yifei Qian, Liangfei Zhang, Zhongliang Guo 0001, Xiaopeng Hong, Ognjen Arandjelovic, Carl Donovan
Pattern Recognit.5
2025 A Gray-Box Attack Against Latent Diffusion Model-Based Image Editing by Posterior Collapse
abstract
Recent advancements in Latent Diffusion Models (LDMs) have revolutionized image synthesis and manipulation, raising significant concerns about data misappropriation and intellectual property infringement. While adversarial attacks have been extensively explored as a protective measure against such misuse of generative AI, current approaches are severely limited by their heavy reliance on model-specific knowledge and substantial computational costs. Drawing inspiration from the posterior collapse phenomenon observed in VAE training, we propose the Posterior Collapse Attack (PCA), a novel framework for protecting images from unauthorized manipulation. Through comprehensive theoretical analysis and empirical validation, we identify two distinct collapse phenomena during VAE inference: diffusion collapse and concentration collapse. Based on this discovery, we design a unified loss function that can flexibly achieve both types of collapse through parameter adjustment, each corresponding to different protection objectives in preventing image manipulation. Our method significantly reduces dependence on model-specific knowledge by requiring access to only the VAE encoder, which constitutes less than 4% of LDM parameters. Notably, PCA achieves prompt-invariant protection by operating on the VAE encoder before text conditioning occurs, eliminating the need for empty prompt optimization required by existing methods. This minimal requirement enables PCA to maintain adequate transferability across various VAE-based LDM architectures while effectively preventing unauthorized image editing. Extensive experiments show PCA outperforms existing techniques in protection effectiveness, computational efficiency (runtime and VRAM), and generalization across VAE-based LDM variants. Our code is available at https://github.com/ZhongliangGuo/PosteriorCollapseAttack.
Zhongliang Guo 0001, Chun Tong Lei, Lei Fang 0001, Shuai Zhao 0007, Yifei Qian, Zeyu Wang 0010, Cunjian Chen, Ognjen Arandjelovic, Chun Pong Lau 0001
IEEE Trans. Inf. Forensics Secur.9
2024 A White-Box False Positive Adversarial Attack Method on Contrastive Loss Based Offline Handwritten Signature Verification Models
abstract
In this paper, we tackle the challenge of white-box false positive adversarial attacks on contrastive loss based offline handwritten signature verification models. We propose a novel attack method that treats the attack as a style transfer between closely related but distinct writing styles. To guide the generation of deceptive images, we introduce two new loss functions that enhance the attack success rate by perturbing the Euclidean distance between the embedding vectors of the original and synthesized samples, while ensuring minimal perturbations by reducing the difference between the generated image and the original image. Our method demonstrates state-of-the-art performance in white-box attacks on contrastive loss based offline handwritten signature verification models, as evidenced by our experiments. The key contributions of this paper include a novel false positive attack method, two new loss functions, effective style transfer in handwriting styles, and superior performance in white-box false positive attacks compared to other white-box attack methods.
Zhongliang Guo 0001, Yifei Qian, Ognjen Arandjelovic, Lei Fang 0001
AISTATS4
2024 Artwork Protection Against Neural Style Transfer Using Locally Adaptive Adversarial Color Attack
abstract
Neural style transfer (NST) generates new images by combining the style of one image with the content of another. However, unauthorized NST can exploit artwork, raising concerns about artists’ rights and motivating the development of proactive protection methods. We propose Locally Adaptive Adversarial Color Attack (LAACA), empowering artists to protect their artwork from unauthorized style transfer by processing before public release. By delving into the intricacies of human visual perception and the role of different frequency components, our method strategically introduces frequency-adaptive perturbations in the image. These perturbations significantly degrade the generation quality of NST while maintaining an acceptable level of visual change in the original image, ensuring that potential infringers are discouraged from using the protected artworks, because of its bad NST generation quality. Additionally, existing metrics often overlook the importance of color fidelity in evaluating color-mattered tasks, such as the quality of NST-generated images, which is crucial in the context of artistic works. To comprehensively assess the color-mattered tasks, we propose the Aesthetic Color Distance Metric (ACDM), designed to quantify the color difference of images pre- and post-manipulations. Experimental results confirm that attacking NST using LAACA results in visually inferior style transfer, and the ACDM can efficiently measure color-mattered tasks. By providing artists with a tool to safeguard their intellectual property, our work relieves the socio-technical challenges posed by the misuse of NST in the art community.
Zhongliang Guo 0001, Junhao Dong 0001, Yifei Qian, Ziheng Guo, Ognjen Arandjelovic, Lei Fang 0001
ECAI9
2024 Would I Lie To You? Inference Time Alignment of Language Models using Direct Preference Heads
abstract
Pre-trained Language Models (LMs) exhibit strong zero-shot and in-context learning capabilities; however, their behaviors are often difficult to control. By utilizing Reinforcement Learning from Human Feedback (RLHF), it is possible to fine-tune unsupervised LMs to follow instructions and produce outputs that reflect human preferences. Despite its benefits, RLHF has been shown to potentially harm a language model's reasoning capabilities and introduce artifacts such as hallucinations where the model may fabricate facts. To address this issue we introduce Direct Preference Heads (DPH), a fine-tuning framework that enables LMs to learn human preference signals through an auxiliary reward head without directly affecting the output distribution of the language modeling head. We perform a theoretical analysis of our objective function and find strong ties to Conservative Direct Preference Optimization (cDPO). Finally we evaluate our models on GLUE, RACE, and the GPT4All evaluation suite and demonstrate that our method produces models which achieve higher scores than those fine-tuned with Supervised Fine-Tuning (SFT) or Direct Preference Optimization (DPO) alone.
Avelina Asada Hadji-Kyriacou, Ognjen Arandjelovic
NeurIPS2
2024 Semi-Supervised Crowd Counting With Contextual Modeling: Facilitating Holistic Understanding of Crowd Scenes
abstract
To alleviate the heavy annotation burden for training a reliable crowd counting model and thus make the model more practicable and accurate by being able to benefit from more data, this paper presents a new semi-supervised method based on the mean teacher framework. When there is a scarcity of labeled data available, the model is prone to overfit local patches. Within such contexts, the conventional approach of solely improving the accuracy of local patch predictions through unlabeled data proves inadequate. Consequently, we propose a more nuanced approach: fostering the model’s intrinsic ‘subitizing’ capability. This ability allows the model to accurately estimate the count in regions by leveraging its understanding of the crowd scenes, mirroring the human cognitive process. To achieve this goal, we apply masking on unlabeled data, guiding the model to make predictions for these masked patches based on the holistic cues. Furthermore, to help with feature learning, herein we incorporate a fine-grained density classification task. Our method is general and applicable to most existing crowd counting methods as it doesn’t have strict structural or loss constraints. In addition, we observe that the model trained with our framework shows strong contextual modeling capabilities, which allows it to make robust predictions even when some local details of patches are lost. Our method achieves the state-of-the-art performance, surpassing previous approaches by a large margin on challenging benchmarks such as ShanghaiTech A and UCF-QNRF. The code is available at: https://github.com/cha15yq/MRC-Crowd.
Yifei Qian, Xiaopeng Hong, Zhongliang Guo 0001, Ognjen Arandjelovic, Carl Donovan
IEEE Trans. Circuits Syst. Video Technol.4
2023 HoechstGAN: Virtual Lymphocyte Staining Using Generative Adversarial Networks
abstract
The presence and density of specific types of immune cells are important to understand a patient’s immune response to cancer. However, immunofluorescence staining required to identify T cell subtypes is expensive, time-consuming, and rarely performed in clinical settings. We present a framework to virtually stain Hoechst images (which are cheap and widespread) with both CD3 and CD8 to identify T cell subtypes in clear cell renal cell carcinoma using generative adversarial networks. Our proposed method jointly learns both staining tasks, incentivising the network to incorporate mutually beneficial information from each task. We devise a novel metric to quantify the virtual staining quality, and use it to evaluate our method.
Georg Wölflein, In Hwa Um, David J. Harrison, Ognjen Arandjelovic
WACV4
2022 Segmentation Assisted U-shaped Multi-scale Transformer for Crowd Counting
Yifei Qian, Liangfei Zhang, Xiaopeng Hong, Carl Donovan, Ognjen Arandjelovic
BMVC5
2022 Data Efficient Support Vector Machine Training Using the Minimum Description Length Principle
abstract
Support vector machines (SVMs) are established as highly successful classifiers in a broad range of applications, including numerous medical ones. Nevertheless, their current employment is restricted by a limitation in the manner in which they are trained, most often the training-validation-test or kfold cross-validation approaches, which are wasteful both in terms of the use of the available data as well as computational resources. This is a particularly important consideration in many medical problems, in which data availability is low (be it because of the inherent difficulty in obtaining sufficient data, or because of practical reasons, e.g. pertaining to privacy and data sharing). In this paper we propose a novel approach to training SVMs which does not suffer from the aforementioned limitation, which is at the same time much more rigorous in nature, being built upon solid information theoretic grounds. Specifically, we show how the training process, that is the process of hyperparameter inference, can be formulated as a search for the optimal model under the minimum description length (MDL) criterion, allowing for theory rather than empiricism driven selection and removing the need for validation data. The effectiveness and superiority of our approach are demonstrated on the Wisconsin Diagnostic Breast Cancer Data Set.
Harsh Singh, Ognjen Arandjelovic
ICASSP2
2022 Believe the HiPe: Hierarchical perturbation for fast, robust, and model-agnostic saliency mapping
abstract
Understanding the predictions made by Artificial Intelligence (AI) systems is becoming more and more important as deep learning models are used for increasingly complex and high-stakes tasks. Saliency mapping – a popular visual attribution method – is one important tool for this, but existing formulations are limited by either computational cost or architectural constraints. We therefore propose Hierarchical Perturbation, a very fast and completely model-agnostic method for interpreting model predictions with robust saliency maps. Using standard benchmarks and datasets, we show that our saliency maps are of competitive or superior quality to those generated by existing model-agnostic methods – and are over 20× faster to compute.
Jessica Rumbelow, Ognjen Arandjelovic, David J. Harrison
Pattern Recognit.2
2022 Short and Long Range Relation Based Spatio-Temporal Transformer for Micro-Expression Recognition
abstract
Being spontaneous, micro-expressions are useful in the inference of a person's true emotions even if an attempt is made to conceal them. Due to their short duration and low intensity, the recognition of micro-expressions is a difficult task in affective computing. The early work based on handcrafted spatio-temporal features which showed some promise, has recently been superseded by different deep learning approaches which now compete for the state of the art performance. Nevertheless, the problem of capturing both local and global spatio-temporal patterns remains challenging. To this end, herein we propose a novel spatio-temporal transformer architecture – to the best of our knowledge, the first purely transformer based approach (i.e., void of any convolutional network use) for micro-expression recognition. The architecture comprises a spatial encoder which learns spatial patterns, a temporal aggregator for temporal dimension analysis, and a classification head. A comprehensive evaluation on three widely used spontaneous micro-expression data sets, namely SMIC-HS, CASME II and SAMM, shows that the proposed approach consistently outperforms the state of the art, and is the first framework in the published literature on micro-expression recognition to achieve the unweighted F1-score greater than 0.9 on any of the aforementioned data sets. The source code is available athttps://github.com/Vision-Intelligence-and-Robots-Group/SLSTT.
Liangfei Zhang, Xiaopeng Hong, Ognjen Arandjelovic, Guoying Zhao 0001
IEEE Trans. Affect. Comput.3
2018 Automatic Semantic Labelling of Images by Their Content Using Non-Parametric Bayesian Machine Learning and Image Search Using Synthetically Generated Image Collages
abstract
In this paper, we describe a novel algorithm for fully unsupervised discovery of meaningful object categories in images, and their semantic labelling. Our approach works bottom-up, starting with superpixel segmentation and the representation of images as bags of quantized superpixel based visual words. Inference over these is achieved automatically using hierarchical non-parametric Bayesian learning. This visual learning process is built upon by the association of class labels using two complementary methods. The first of these is primarily suited for conventional object categories, whereas the second one, which employs a reverse image search using synthetically generated superpixel collages, is used for amorphous image content (such as 'grass', 'sky', etc.). The proposed algorithm does not rely on any dataset specific parameter tuning and (by exploiting auxiliary Google retrievable "big data"' in the wild) can be applied on any kind of corpora (no labelling, weak labelling, or fully annotated images).
Michael Niemeyer, Ognjen Arandjelovic
DSAA2
2018 Employing Domain Specific Discriminative Information to Address Inherent Limitations of the LBP Descriptor in Face Recognition
abstract
The local binary patern (LBP) descriptor and its derivatives have a demonstrated track record of good performance in face recognition. Nevertheless the original descriptor, the framework within which it is employed, and the aforementioned improvements of these in the existing literature, all suffer from a number of inherent limitations. In this work we highlight these and propose novel ways of addressing them in a principled fashion. Specifically, we introduce (i) gradient based weighting of local descriptor contributions to region based histograms as a means of avoiding data smoothing by non-discriminative image loci, and (ii) Gaussian fuzzy region membership as a means of achieving robustness to registration errors. Importantly, the nature of these contributions allows the proposed techniques to be combined with the existing extensions to the LBP descriptor thus making them universally recommendable. Effectiveness is demonstrated on the notoriously challenging Extended Yale B face corpus.
Junjie Fan, Ognjen Arandjelovic
IJCNN2
2018 Discovering topic structures of a temporally evolving document corpus
abstract
In this paper we describe a novel framework for the discovery of the topical content of a data corpus, and the tracking of its complex structural changes across the temporal dimension. In contrast to previous work our model does not impose a prior on the rate at which documents are added to the corpus nor does it adopt the Markovian assumption which overly restricts the type of changes that the model can capture. Our key technical contribution is a framework based on (i) discretization of time into epochs, (ii) epoch-wise topic discovery using a hierarchical Dirichlet process-based model, and (iii) a temporal similarity graph which allows for the modelling of complex topic changes: emergence and disappearance, evolution, splitting, and merging. The power of the proposed framework is demonstrated on two medical literature corpora concerned with the autism spectrum disorder (ASD) and the metabolic syndrome (MetS)—both increasingly important research subjects with significant social and healthcare consequences. In addition to the collected ASD and metabolic syndrome literature corpora which we made freely available, our contribution also includes an extensive empirical analysis of the proposed framework. We describe a detailed and careful examination of the effects that our algorithms’s free parameters have on its output and discuss the significance of the findings both in the context of the practical application of our algorithm as well as in the context of the existing body of work on temporal topic analysis. Our quantitative analysis is followed by several qualitative case studies highly relevant to the current research on ASD and MetS, on which our algorithm is shown to capture well the actual developments in these fields.
Adham Beykikhoshk, Ognjen Arandjelovic, Dinh Q. Phung, Svetha Venkatesh
Knowl. Inf. Syst.2
2018 Reimagining the central challenge of face recognition: Turning a problem into an advantage
Ognjen Arandjelovic
Pattern Recognit.1
2017 Ancient Roman Coin Retrieval: A Systematic Examination of the Effects of Coin Grade
Callum Fare, Ognjen Arandjelovic
ECIR2
2017 Towards computer vision based ancient coin recognition in the wild - Automatic reliable image preprocessing and normalization
abstract
As an attractive area of application in the sphere of cultural heritage, in recent years automatic analysis of ancient coins has been attracting an increasing amount of research attention from the computer vision community. Recent work has demonstrated that the existing state of the art performs extremely poorly when applied on images acquired in realistic conditions. One of the reasons behind this lies in the (often implicit) assumptions made by many of the proposed algorithms - a lack of background clutter, and a uniform scale, orientation, and translation of coins across different images. These assumptions are not satisfied by default and before any further progress in the realm of more complex analysis is made, a robust method capable of preprocessing and normalizing images of coins acquired `in the wild' is needed. In this paper we introduce an algorithm capable of localizing and accurately segmenting out a coin from a cluttered image acquired by an amateur collector. Specifically, we propose a two stage approach which first uses a simple shape hypothesis to localize the coin roughly and then arrives at the final, accurate result by refining this initial estimate using a statistical model learnt from large amounts of data. Our results on data collected `in the wild' demonstrate excellent accuracy even when the proposed algorithm is applied on highly challenging images.
Brandon Conn, Ognjen Arandjelovic
IJCNN2
2017 Information and knowing when to forget it
abstract
In this paper we propose several novel approaches for incorporating forgetting mechanisms into sequential prediction based machine learning algorithms. The broad premise of our work, supported and motivated in part by recent findings stemming from neurology research on the development of human brains, is that knowledge acquisition and forgetting are complementary processes, and that learning can (perhaps unintuitively) benefit from the latter too. We demonstrate that if forgetting is implemented in a purposeful and date driven manner, there are a number of benefits which can be gained from discarding information. The framework we introduce is a general one and can be used with any baseline predictor of choice. Hence in this sense it is best described as a meta-algorithm. The method we described was developed through a series of steps which increase the adaptability of the model, while being data driven. We first discussed a weakly adaptive forgetting process which we termed passive forgetting. A fully adaptive framework, which we termed active forgetting was developed by enveloping a passive forgetting process with a monitoring, self-aware module which detects contextual changes and makes a statistically informed choice when the model parameters should be abruptly rather than gradually updated. The effectiveness of the proposed meta-framework was demonstrated on two real world data sets concerned with challenges of major practical importance: those of predicting currency exchange rates and daily temperatures. On both tasks our approach was shown to be highly effective, reducing prediction errors by nearly 40%.
Ognjen Arandjelovic
IJCNN2
2017 Light Curve Analysis From Kepler Spacecraft Collected Data
abstract
Although scarce, previous work on the application of machine learning and data mining techniques on large corpora of astronomical data has produced promising results. For example, on the task of detecting so-called Kepler objects of interest (KOIs), a range of different `off the shelf' classifiers has demonstrated outstanding performance. These rather preliminary research efforts motivate further exploration of this data domain. In the present work we focus on the analysis of threshold crossing events (TCEs) extracted from photometric data acquired by the Kepler spacecraft. We show that the task of classifying TCEs as being effected by actual planetary transits as opposed to confounding astrophysical phenomena is significantly more challenging than that of KOI detection, with different classifiers exhibiting vastly different performances. Nevertheless, the best performing classifier type, the random forest, achieved excellent accuracy, correctly predicting in approximately 96% of the cases. Our results and analysis should illuminate further efforts into the development of more sophisticated, automatic techniques, and encourage additional work in the area.
Eduardo Nigri, Ognjen Arandjelovic
ICMR2
2016 Learnt Quasi-Transitive Similarity for Retrieval from Large Collections of Faces
abstract
We are interested in identity-based retrieval of face sets from large unlabelled collections acquired in uncontrolled environments. Given a baseline algorithm for measuring the similarity of two face sets, the meta-algorithm introduced in this paper seeks to leverage the structure of the data corpus to make the best use of the available baseline. In particular, we show how partial transitivity of inter-personal similarity can be exploited to improve the retrieval of particularly challenging sets which poorly match the query under the baseline measure. We: (i) describe the use of proxy sets as a means of computing the similarity between two sets, (ii) introduce transitivity meta-features based on the similarity of salient modes of appearance variation between sets, (iii) show how quasi-transitivity can be learnt from such features without any labelling or manual intervention, and (iv) demonstrate the effectiveness of the proposed methodology through experiments on the notoriously challenging YouTube database.
Ognjen Arandjelovic
CVPR1
2016 Analysing the History of Autism Spectrum Disorder Using Topic Models
abstract
We describe a novel framework for the discovery of underlying topics of a longitudinal collection of scholarly data, and the tracking of their lifetime and popularity over time. Unlike the social media or news data where the underlying topics evolve over time, the topic nuances in science result in new scientific directions to emerge. Therefore, we model the longitudinal literature data with a new approach that uses topics which remain identifiable over the course of time. Current studies either disregard the time dimension or treat it as an exchangeable covariate when they fix the topics over time or do not share the topics over epochs when they model the time naturally. We address these issues by adopting a non-parametric Bayesian approach. We assume the data is partially exchangeable and divide it into consecutive epochs. Then, by fixing the topics in a recurrent Chinese restaurant franchise, we impose a static topical structure on the corpus such that the topics are shared across epochs and the documents within epochs. We demonstrate the effectiveness of the proposed framework on a collection of medical literature related to autism spectrum disorder. We collect a large corpus of publications and carefully examine two important research issues of the domain as case studies. Moreover, we make the results of our experiment and the source code of the model, freely available to the public. This aids other researchers to analyse our results or apply the model to their data collections.
Adham Beykikhoshk, Dinh Q. Phung, Ognjen Arandjelovic, Svetha Venkatesh
DSAA3
2016 Highly Accurate Gaze Estimation Using a Consumer RGB-D Sensor
Reza Shoja Ghiass, Ognjen Arandjelovic
IJCAI2
2016 Achieving stable subspace clustering by post-processing generic clustering results
abstract
We propose an effective subspace selection scheme as a post-processing step to improve results obtained by sparse subspace clustering (SSC). Our method starts by the computation of stable subspaces using a novel random sampling scheme. Thus constructed preliminary subspaces are used to identify the initially incorrectly clustered data points and then to reassign them to more suitable clusters based on their goodness-of-fit to the preliminary model. To improve the robustness of the algorithm, we use a dominant nearest subspace classification scheme that controls the level of sensitivity against reassignment. We demonstrate that our algorithm is convergent and superior to the direct application of a generic alternative such as principal component analysis. On several popular datasets for motion segmentation and face clustering pervasively used in the sparse subspace clustering literature the proposed method is shown to reduce greatly the incidence of clustering errors while introducing negligible disturbance to the data points already correctly clustered.
Duc-Son Pham 0001, Ognjen Arandjelovic, Svetha Venkatesh
IJCNN2
2016 Descriptor transition tables for object retrieval using unconstrained cluttered video acquired using a consumer level handheld mobile device
abstract
Visual recognition and vision based retrieval of objects from large databases are tasks with a wide spectrum of potential applications. In this paper we propose a novel recognition method from video sequences suitable for retrieval from databases acquired in highly unconstrained conditions e.g. using a mobile consumer-level device such as a phone. On the lowest level, we represent each sequence as a 3D mesh of densely packed local appearance descriptors. While image plane geometry is captured implicitly by a large overlap of neighbouring regions from which the descriptors are extracted, 3D information is extracted by means of a descriptor transition table, learnt from a single sequence for each known gallery object. These allow us to connect local descriptors along the 3rddimension (which corresponds to viewpoint changes), thus resulting in a set of variable length Markov chains for each video. The matching of two sets of such chains is formulated as a statistical hypothesis test, whereby a subset of each is chosen to maximize the likelihood that the corresponding video sequences show the same object. The effectiveness of the proposed algorithm is empirically evaluated on the Amsterdam Library of Object Images and a new highly challenging video data set acquired using a mobile phone. On both data sets our method is shown to be successful in recognition in the presence of background clutter and large viewpoint changes.
Warren Rieutort-Louis, Ognjen Arandjelovic
IJCNN2
2016 Weighted Linear Fusion of Multimodal Data: A Reasonable Baseline?
abstract
The ever-increasing demand for reliable inference capable of handling unpredictable challenges of practical application in the real world, has made research on information fusion of major importance. There are few fields of application and research where this is more evident than in the sphere of multimedia which by its very nature inherently involves the use of multiple modalities, be it for learning, prediction, or human-computer interaction, say. In the development of the most common type, score-level fusion algorithms, it is virtually without an exception desirable to have as a reference starting point a simple and universally sound baseline benchmark which newly developed approaches can be compared to. One of the most pervasively used methods is that of weighted linear fusion. It has cemented itself as the default off-the-shelf baseline owing to its simplicity of implementation, interpretability, and surprisingly competitive performance across a wide range of application domains and information source types. In this paper I argue that despite this track record, weighted linear fusion is not a good baseline on the grounds that there is an equally simple and interpretable alternative - namely quadratic mean-based fusion - which is theoretically more principled and which is more successful in practice. I argue the former from first principles and demonstrate the latter using a series of experiments on a diverse set of fusion problems: computer vision-based object recognition, arrhythmia detection, and fatality prediction in motor vehicle accidents.
Ognjen Arandjelovic
ACM Multimedia1
2016 On the discovery of hospital admission patterns - a clarification
abstract
Abstract Contact: [email protected]
Ognjen Arandjelovic
Bioinform.1
2016 CCTV Scene Perspective Distortion Estimation From Low-Level Motion Features
abstract
Our aim is to estimate the perspective-effected geometric distortion of a scene from a video feed. In contrast to most related previous work, in this task we are constrained to use low-level spatiotemporally local motion features only. This particular challenge arises in many semiautomatic surveillance systems that alert a human operator to potential abnormalities in the scene. Low-level spatiotemporally local motion features are sparse (and thus require comparatively little storage space) and sufficiently powerful in the context of video abnormality detection to reduce the need for human intervention by more than 100-fold. This paper introduces three significant contributions. First, we describe a dense algorithm for perspective estimation, which uses motion features to estimate the perspective distortion at each image locus and then polls all such local estimates to arrive at the globally best estimate. Second, we also present an alternative coarse algorithm that subdivides the image frame into blocks and uses motion features to derive block-specific motion characteristics and constrain the relationships between these characteristics, with the perspective estimate emerging as a result of a global optimization scheme. Third, we report the results of an evaluation using nine large sets acquired using existing closed-circuit television cameras, not installed specifically for the purposes of this paper. Our findings demonstrate that both proposed methods are successful, their accuracy matching that of human labeling using complete visual data (by the constraints of the setup unavailable to our algorithms).
Ognjen Arandjelovic, Duc-Son Pham 0001, Svetha Venkatesh
IEEE Trans. Circuits Syst. Video Technol.1
2015 Sample-Targeted Clinical Trial Adaptation
abstract
Clinical trial adaptation refers to any adjustment of the trial protocol after the onset of the trial. The main goal is to make the process of introducing new medical interventions to patients more efficient by reducing the cost and the time associated with evaluating their safety and efficacy. The principal question is how should adaptation be performed so as to minimize the chance of distorting the outcome of the trial. We propose a novel method for achieving this. Unlike previous work our approach focuses on trial adaptation by sample size adjustment. We adopt a recently proposed stratification framework based on collected auxiliary data and show that this information together with the primary measured variables can be used to make a probabilistically informed choice of the particular sub-group a sample should be removed from. Experiments on simulated data are used to illustrate the effectiveness of our method and its application in practice.
Ognjen Arandjelovic
AAAI1
2015 Overcoming Data Scarcity of Twitter: Using Tweets as Bootstrap with Application to Autism-Related Topic Content Analysis
abstract
Notwithstanding recent work which has demonstrated the potential of using Twitter messages for content-specific data mining and analysis, the depth of such analysis is inherently limited by the scarcity of data imposed by the 140 character tweet limit. In this paper we describe a novel approach for targeted knowledge exploration which uses tweet content analysis as a preliminary step. This step is used to bootstrap more sophisticated data collection from directly related but much richer content sources. In particular we demonstrate that valuable information can be collected by following URLs included in tweets. We automatically extract content from the corresponding web pages and treating each web page as a document linked to the original tweet show how a temporal topic model based on a hierarchical Dirichlet process can be used to track the evolution of a complex topic structure of a Twitter community. Using autism-related tweets we demonstrate that our method is capable of capturing a much more meaningful picture of information exchange than user-chosen hashtags.
Adham Beykikhoshk, Ognjen Arandjelovic, Dinh Q. Phung, Svetha Venkatesh
ASONAM2
2015 Automatic vehicle tracking and recognition from aerial image sequences
abstract
This paper addresses the problem of automated vehicle tracking and recognition from aerial image sequences. Motivated by its successes in the existing literature we focus on the use of linear appearance subspaces to describe multi-view object appearance and highlight the challenges involved in their application as a part of a practical system. A working solution which includes steps for data extraction and normalization is described. In experiments on real-world data the proposed methodology achieved promising results with a high correct recognition rate and few, meaningful errors (type II errors whereby genuinely similar targets are sometimes being confused with one another). Directions for future research and possible improvements of the proposed method are discussed.
Ognjen Arandjelovic
AVSS1
2015 Groupwise Registration of Aerial Images
Ognjen Arandjelovic, Duc-Son Pham 0001, Svetha Venkatesh
IJCAI1
2015 The adaptable buffer algorithm for high quantile estimation in non-stationary data streams
abstract
The need to estimate a particular quantile of a distribution is an important problem which frequently arises in many computer vision and signal processing applications. For example, our work was motivated by the requirements of many semi-automatic surveillance analytics systems which detect abnormalities in close-circuit television (CCTV) footage using statistical models of low-level motion features. In this paper we specifically address the problem of estimating the running quantile of a data stream with non-stationary stochasticity when the memory for storing observations is limited. We make several major contributions: (i) we derive an important theoretical result which shows that the change in the quantile of a stream is constrained regardless of the stochastic properties of data, (ii) we describe a set of high-level design goals for an effective estimation algorithm that emerge as a consequence of our theoretical findings, (iii) we introduce a novel algorithm which implements the aforementioned design goals by retaining a sample of data values in a manner adaptive to changes in the distribution of data and progressively narrowing down its focus in the periods of quasi-stationary stochasticity, and (iv) we present a comprehensive evaluation of the proposed algorithm and compare it with the existing methods in the literature on both synthetic data sets and three large ‘real-world’ streams acquired in the course of operation of an existing commercial surveillance system. Our findings convincingly demonstrate that the proposed method is highly successful and vastly outperforms the existing alternatives, especially when the target quantile is high valued and the available buffer capacity severely limited.
Ognjen Arandjelovic, Duc-Son Pham 0001, Svetha Venkatesh
IJCNN1
2015 Hierarchical Dirichlet Process for Tracking Complex Topical Structure Evolution and Its Application to Autism Research Literature
Adham Beykikhoshk, Ognjen Arandjelovic, Svetha Venkatesh, Dinh Q. Phung
PAKDD (1)2
2015 Discovering hospital admission patterns using models learnt from electronic hospital records
abstract
MOTIVATION: Electronic medical records, nowadays routinely collected in many developed countries, open a new avenue for medical knowledge acquisition. In this article, this vast amount of information is used to develop a novel model for hospital admission type prediction. RESULTS: I introduce a novel model for hospital admission-type prediction based on the representation of a patient's medical history in the form of a binary history vector. This representation is motivated using empirical evidence from previous work and validated using a large data corpus of medical records from a local hospital. The proposed model allows exploration, visualization and patient-specific prognosis making in an intuitive and readily understood manner. Its power is demonstrated using a large, real-world data corpus collected by a local hospital on which it is shown to outperform previous state-of-the-art in the literature, achieving over 82% accuracy in the prediction of the first future diagnosis. The model was vastly superior for long-term prognosis as well, outperforming previous work in 82% of the cases, while producing comparable performance in the remaining 18% of the cases. AVAILABILITY AND IMPLEMENTATION: Full Matlab source code is freely available for download at: http://ognjen-arandjelovic.t15.org/data/dprog.zip.
Ognjen Arandjelovic
Bioinform.1
2015 Efficient and accurate set-based registration of time-separated aerial images
Ognjen Arandjelovic, Duc-Son Pham 0001, Svetha Venkatesh
Pattern Recognit.1
2015 Two Maximum Entropy-Based Algorithms for Running Quantile Estimation in Nonstationary Data Streams
abstract
The need to estimate a particular quantile of a distribution is an important problem that frequently arises in many computer vision and signal processing applications. For example, our work was motivated by the requirements of many semiautomatic surveillance analytics systems that detect abnormalities in close-circuit television footage using statistical models of low-level motion features. In this paper, we specifically address the problem of estimating the running quantile of a data stream when the memory for storing observations is limited. We make the following several major contributions: 1) we highlight the limitations of approaches previously described in the literature that make them unsuitable for nonstationary streams; 2) we describe a novel principle for the utilization of the available storage space; 3) we introduce two novel algorithms that exploit the proposed principle in different ways; and 4) we present a comprehensive evaluation and analysis of the proposed algorithms and the existing methods in the literature on both synthetic data sets and three large real-world streams acquired in the course of operation of an existing commercial surveillance system. Our findings convincingly demonstrate that both of the proposed methods are highly successful and vastly outperform the existing alternatives. We show that the better of the two algorithms (data-aligned histogram) exhibits far superior performance in comparison with the previously described methods, achieving more than 10 times lower estimate errors on real-world data, even when its available working memory is an order of magnitude smaller.
Ognjen Arandjelovic, Duc-Son Pham 0001, Svetha Venkatesh
IEEE Trans. Circuits Syst. Video Technol.1
2015 Detection of Dynamic Background Due to Swaying Movements From Motion Features
abstract
Dynamically changing background (dynamic background) still presents a great challenge to many motion-based video surveillance systems. In the context of event detection, it is a major source of false alarms. There is a strong need from the security industry either to detect and suppress these false alarms, or dampen the effects of background changes, so as to increase the sensitivity to meaningful events of interest. In this paper, we restrict our focus to one of the most common causes of dynamic background changes: 1) that of swaying tree branches and 2) their shadows under windy conditions. Considering the ultimate goal in a video analytics pipeline, we formulate a new dynamic background detection problem as a signal processing alternative to the previously described but unreliable computer vision-based approaches. Within this new framework, we directly reduce the number of false alarms by testing if the detected events are due to characteristic background motions. In addition, we introduce a new data set suitable for the evaluation of dynamic background detection. It consists of real-world events detected by a commercial surveillance system from two static surveillance cameras. The research question we address is whether dynamic background can be detected reliably and efficiently using simple motion features and in the presence of similar but meaningful events, such as loitering. Inspired by the tree aerodynamics theory, we propose a novel method named local variation persistence (LVP), that captures the key characteristics of swaying motions. The method is posed as a convex optimization problem, whose variable is the local variation. We derive a computationally efficient algorithm for solving the optimization problem, the solution of which is then used to form a powerful detection statistic. On our newly collected data set, we demonstrate that the proposed LVP achieves excellent detection results and outperforms the best alternative adapted from existing art in the dynamic background literature.
Duc-Son Pham 0001, Ognjen Arandjelovic, Svetha Venkatesh
IEEE Trans. Image Process.2
2014 Data-mining twitter and the autism spectrum disorder: A Pilot study
abstract
The autism spectrum disorder (ASD) is increasingly being recognized as a major public health issue which affects approximately 0.5-0.6% of the population. Promoting the general awareness of the disorder, increasing the engagement with the affected individuals and their carers, and understanding the success of penetration of the current clinical recommendations in the target communities, is crucial in driving research as well as policy. The aim of the present work is to investigate if Twitter, as a highly popular platform for information exchange, can be used as a data-mining source which could aid in the aforementioned challenges. Specifically, using a large data set of harvested tweets, we present a series of experiments which examine a range of linguistic and semantic aspects of messages posted by individuals interested in ASD. Our findings, the first of their nature in the published scientific literature, strongly motivate additional research on this topic and present a methodological basis for further work.
Adham Beykikhoshk, Ognjen Arandjelovic, Dinh Q. Phung, Svetha Venkatesh, Terry Caelli
ASONAM2
2014 A framework for improving the performance of verification algorithms with a low false positive rate requirement and limited training data
abstract
In this paper we address the problem of matching patterns in the so-called verification setting in which a novel, query pattern is verified against a single training pattern: the decision sought is whether the two match (i.e. belong to the same class) or not. Unlike previous work which has universally focused on the development of more discriminative distance functions between patterns, here we consider the equally important and pervasive task of selecting a distance threshold which fits a particular operational requirement - specifically, the target false positive rate (FPR). First, we argue on theoretical grounds that a data-driven approach is inherently ill-conditioned when the desired FPR is low, because by the very nature of the challenge only a small portion of training data affects or is affected by the desired threshold. This leads us to propose a general, statistical model-based method instead. Our approach is based on the interpretation of an inter-pattern distance as implicitly defining a pattern embedding which approximately distributes patterns according to an isotropic multi-variate normal distribution in some space. This interpretation is then used to show that the distribution of training interpattern distances is the non-central χ2distribution, differently parameterized for each class. Thus, to make the class-specific threshold choice we propose a novel analysis-by-synthesis iterative algorithm which estimates the three free parameters of the model (for each class) using task-specific constraints. The validity of the premises of our work and the effectiveness of the proposed method are demonstrated by applying the method to the task of set-based face verification on a large database of pseudo-random head motion videos.
Ognjen Arandjelovic
IJCB1
2014 Stream Quantiles via Maximal Entropy Histograms
Ognjen Arandjelovic, Duc-Son Pham 0001, Svetha Venkatesh
ICONIP (2)1
2014 A Unified Framework for Thermal Face Recognition
Reza Shoja Ghiass, Ognjen Arandjelovic, Abdelhakim Bendada, Xavier Maldague
ICONIP (2)2
2014 Discriminative extended canonical correlation analysis for pattern set matching
Ognjen Arandjelovic
Mach. Learn.1
2014 Hallucinating optimal high-dimensional subspaces
Ognjen Arandjelovic
Pattern Recognit.1
2014 Infrared face recognition: A comprehensive review of methodologies and databases
Reza Shoja Ghiass, Ognjen Arandjelovic, Abdelhakim Bendada, Xavier Maldague
Pattern Recognit.2
2013 Vesselness Features and the Inverse Compositional AAM for Robust Face Recognition Using Thermal IR
abstract
Over the course of the last decade, infrared (IR) and particularly thermal IR imaging based face recognition has emerged as a promising complement to conventional, visible spectrum based approaches which continue to struggle when applied in the real world. While inherently insensitive to visible spectrum illumination changes, IR images introduce specific challenges of their own, most notably sensitivity to factors which affect facial heat emission patterns, e.g. emotional state, ambient temperature, and alcohol intake. In addition, facial expression and pose changes are more difficult to correct in IR images because they are less rich in high frequency detail which is an important cue for fitting any deformable model. In this paper we describe a novel method which addresses these major challenges. Specifically, to normalize for pose and facial expression changes we generate a synthetic frontal image of a face in a canonical, neutral facial expression from an image of the face in an arbitrary pose and facial expression. This is achieved by piecewise affine warping which follows active appearance model (AAM) fitting. This is the first publication which explores the use of an AAM on thermal IR images; we propose a pre-processing step which enhances detail in thermal images, making AAM convergence faster and more accurate. To overcome the problem of thermal IR image sensitivity to the exact pattern of facial temperature emissions we describe a representation based on reliable anatomical features. In contrast to previous approaches, our representation is not binary; rather, our method accounts for the reliability of the extracted features. This makes the proposed representation much more robust both to pose and scale changes. The effectiveness of the proposed approach is demonstrated on the largest public database of thermal IR images of faces on which it achieved 100% identification rate, significantly outperforming previously described methods.
Reza Shoja Ghiass, Ognjen Arandjelovic, Abdelhakim Bendada, Xavier Maldague
AAAI2
2013 Discriminative k-means clustering
abstract
The k-means algorithm is a partitional clustering method. Over 60 years old, it has been successfully used for a variety of problems. The popularity of k-means is in large part a consequence of its simplicity and efficiency. In this paper we are inspired by these appealing properties of k-means in the development of a clustering algorithm which accepts the notion of “positively” and “negatively” labelled data. The goal is to discover the cluster structure of both positive and negative data in a manner which allows for the discrimination between the two sets. The usefulness of this idea is demonstrated practically on the problem of face recognition, where the task of learning the scope of a person's appearance should be done in a manner which allows this face to be differentiated from others.
Ognjen Arandjelovic
IJCNN1
2013 Illumination-invariant face recognition from a single image across extreme pose using a dual dimension AAM ensemble in the thermal infrared spectrum
abstract
Over the course of the last decade, infrared (IR) and particularly thermal IR imaging based face recognition has emerged as a promising complement to conventional, visible spectrum based approaches which continue to struggle when applied in practice. While inherently insensitive to visible spectrum illumination changes, IR data introduces specific challenges of its own, most notably sensitivity to factors which affect facial heat emission patterns, e.g. emotional state, ambient temperature, and alcohol intake. In addition, facial expression and pose changes are more difficult to correct in IR images because they are less rich in high frequency detail which is an important cue for fitting any deformable model. In this paper we describe a novel method which addresses these major challenges. Specifically, when comparing two thermal IR images of faces, we mutually normalize their poses and facial expressions by using an active appearance model (AAM) to generate synthetic images of the two faces with a neutral facial expression and in the same view (the average of the two input views). This is achieved by piecewise affine warping which follows AAM fitting. A major contribution of our work is the use of an AAM ensemble in which each AAM is specialized to a particular range of poses and a particular region of the thermal IR face space. Combined with the contributions from our previous work which addressed the problem of reliable AAM fitting in the thermal IR spectrum, and the development of a person-specific representation robust to transient changes in the pattern of facial temperature emissions, the proposed ensemble framework accurately matches faces across the full range of yaw from frontal to profile, even in the presence of scale variation (e.g. due to the varying distance of a subject from the camera). The effectiveness of the proposed approach is demonstrated on the largest public database of thermal IR images of faces and a newly acquired data set of thermal IR motion videos. Our approach achieved perfect recognition performance on both data sets, significantly outperforming the current state of the art methods even when they are trained with multiple images spanning a range of head views.
Reza Shoja Ghiass, Ognjen Arandjelovic, Abdelhakim Bendada, Xavier Maldague
IJCNN2
2013 Infrared face recognition: A literature review
abstract
Automatic face recognition (AFR) is an area with immense practical potential which includes a wide range of commercial and law enforcement applications, and it continues to be one of the most active research areas of computer vision. Even after over three decades of intense research, the state-of-the-art in AFR continues to improve, benefiting from advances in a range of different fields including image processing, pattern recognition, computer graphics and physiology. However, systems based on visible spectrum images continue to face challenges in the presence of illumination, pose and expression changes, as well as facial disguises, all of which can significantly decrease their accuracy. Amongst various approaches which have been proposed in an attempt to overcome these limitations, the use of infrared (IR) imaging has emerged as a particularly promising research direction. This paper presents a comprehensive and timely review of the literature on this subject.
Reza Shoja Ghiass, Ognjen Arandjelovic, Abdelhakim Bendada, Xavier Maldague
IJCNN2
2013 Achieving robust face recognition from video by combining a weak photometric model and a learnt generic face invariant
Ognjen Arandjelovic, Roberto Cipolla
Pattern Recognit.1
2012 Gradient Edge Map Features for Frontal Face Recognition under Extreme Illumination Changes
abstract
Gradient edge map features for frontal face recognition under extreme Gradient edge map features for frontal face recognition under extreme illumination changes illumination changes
Ognjen Arandjelovic
BMVC1
2012 Object Matching Using Boundary Descriptors
abstract
The problem of object recognition is of immense practical importance and potential, and the last decade has witnessed a number of breakthroughs in the state of the art. Most of the past object recognition work focuses on textured objects and local appearance descriptors extracted around salient points in an image. These methods fail in the matching of smooth, untextured objects for which salient point detection does not produce robust results. The recently proposed bag of boundaries (BoB) method is the first to directly address this problem. Since the texture of smooth objects is largely uninformative, BoB focuses on describing and matching objects based on their post-segmentation boundaries. Herein we address three major weaknesses of this work. The first of these is the uniform treatment of all boundary segments. Instead, we describe a method for detecting the locations and scales of salient boundary segments. Secondly, while the BoB method uses an image based elementary descriptor (HoGs + occupancy matrix), we propose a more compact descriptor based on the local profile of boundary normals’ directions. Lastly, we conduct a far more systematic evaluation, both of the bag of boundaries method and the method proposed here. Using a large public database, we demonstrate that our method exhibits greater robustness while at the same time achieving a major computational saving – object representation is extracted from an image in only 6% of the time needed to extract a bag of boundaries, and the storage requirement is similarly reduced to less than 8%.
Ognjen Arandjelovic
BMVC1
2012 Reading Ancient Coins: Automatically Identifying Denarii Using Obverse Legend Seeded Retrieval
Ognjen Arandjelovic
ECCV (4)1
2012 Assessing Blinding in Clinical Trials
abstract
The interaction between the patient's expected outcome of an intervention and the inherent effects of that intervention can have extraordinary effects. Thus in clinical trials an effort is made to conceal the nature of the administered intervention from the participants in the trial i.e. to blind it. Yet, in practice perfect blinding is impossible to ensure or even verify. The current standard is follow up the trial with an auxiliary questionnaire, which allows trial participants to express their belief concerning the assigned intervention and which is used to compute a measure of the extent of blinding in the trial. If the estimated extent of blinding exceeds a threshold the trial is deemed sufficiently blinded; otherwise, the trial is deemed to have failed. In this paper we make several important contributions. Firstly, we identify a series of fundamental problems of the aforesaid practice and discuss them in context of the most commonly used blinding measures. Secondly, motivated by the highlighted problems, we formulate a novel method for handling imperfectly blinded trials. We too adopt a post-trial feedback questionnaire but interpret the collected data using an original approach, fundamentally different from those previously proposed. Unlike previous approaches, ours is void of any ad hoc free parameters, is robust to small changes in auxiliary data and is not predicated on any strong assumptions used to interpret participants' feedback.
Ognjen Arandjelovic
NIPS1
2012 Computationally efficient application of the generic shape-illumination invariant to face recognition from video
Ognjen Arandjelovic
Pattern Recognit.1
2012 Colour invariants under a non-linear photometric camera model and their application to face recognition from video
Ognjen Arandjelovic
Pattern Recognit.1
2010 Accurate and Efficient Face Recognition from Video
abstract
As a problem of high practical appeal but outstanding challenges, computer-based face recognition remains a topic of extensive research attention.In this paper we are specifically interested in the task of identifying a person from multiple training and query images.Thus, a novel method is proposed which advances the state-of-the-art in setbased face recognition.Our method is based on a previously described invariant in the form of generic shape-illumination effects.The contributions include: (i) an analysis of computational demands of the original method and a demonstration of its practical limitations, (ii) a novel representation of personal appearance in the form of linked mixture models in image and pose-signature spaces, and (iii) an efficient (in terms of storage needs and matching time) manifold re-illumination algorithm based on the aforementioned representation.An evaluation and comparison of the proposed method with the original generic shape-illumination algorithm shows that comparably high recognition rates are achieved on a large data set (1.5% error on 700 face sets containing 100 individuals and extreme illumination variation) with a dramatic improvement in matching speed (over 700 times for sets containing 1600 faces) and storage requirements (independent of the number of training images).
Ognjen Arandjelovic
BMVC1
2010 Recognition from Appearance Subspaces across Image Sets of Variable Scale
abstract
Linear subspace representations of appearance variation are pervasive in computer vision. In this paper we address the problem of robustly matching them (computing the similarity between them) when they correspond to sets of images of different (possibly greatly so) scales. We show that the naive solution of projecting the low-scale subspace into the high-scale image space is inadequate, especially at large scale discrepancies. A successful approach is proposed instead. It consists of (i) an interpolated projection of the low-scale subspace into the high-scale space, which is followed by (ii) a rotation of this initial estimate within the bounds of the imposed “downsampling constraint”. The optimal rotation is found in the closed-form which best aligns the high-scale reconstruction of the low-scale subspace with the reference it is compared to. The proposed method is evaluated on the problem of matching sets of face appearances under varying illumination. In comparison to the naive matching, our algorithm is shown to greatly increase the separation of between-class and within-class similarities, as well as produce far more meaningful modes of common appearance on which the match score is based.
Ognjen Arandjelovic
BMVC1
2010 Automatic attribution of ancient Roman imperial coins
abstract
Classification of coins is an important but laborious aspect of numismatics - the field that studies coins and currency. It is particularly challenging in the case of ancient coins. Due to the way they were manufactured, as well as wear from use and exposure to chemicals in the soil, the same ancient coin type can exhibit great variability in appearance. We demonstrate that geometry-free models of appearance do not perform better than chance on this task and that only a small improvement is gained by previously proposed models of combined appearance and geometry. Thus, our first major contribution is a new type of feature which is efficient in terms of computational time and storage requirements, and which effectively captures geometric configurations between descriptors corresponding to local features. Our second contribution is a description of a fully automatic system based on the proposed features, which robustly localizes, segments out and classifies coins from cluttered images. We also describe a large database of ancient coins that we collected and which will be made publicly available. Finally, we report the results of empirical comparison of different coin matching techniques. The features proposed in this paper are found to greatly outperform existing methods.
Ognjen Arandjelovic
CVPR1
2010 Thermal and reflectance based personal identification methodology under variable illumination
Ognjen Arandjelovic, Riad I. Hammoud, Roberto Cipolla
Pattern Recognit.1
2009 Unfolding a Face: From Singular to Manifold
Ognjen Arandjelovic
ACCV (3)1
2009 A pose-wise linear illumination manifold model for face recognition using video
Ognjen Arandjelovic, Roberto Cipolla
Comput. Vis. Image Underst.1
2009 A methodology for rapid illumination-invariant face recognition using image processing filters
Ognjen Arandjelovic, Roberto Cipolla
Comput. Vis. Image Underst.1
2008 Crowd Detection from Still Images
abstract
The analysis of human crowds has widespread uses from law enforcement to urban engineering and traffic management. All of these require a crowd to first be detected, which is the problem addressed in this paper. Given an image, the algorithm we propose segments it into crowd and non-crowd regions. The main idea is to capture two key properties of crowds: (i) on a narrow scale, its basic element should look like a human (only weakly so, due to low resolution, occlusion, clothing variation etc.), while (ii) on a larger scale, a crowd inherently contains repetitive appearance elements. Our method exploits this by building a pyramid of sliding windows and quan-tifying how “crowd-like ” each level of the pyramid is using an underlying statistical model based on quantized SIFT features. The two aforementioned crowd properties are captured by the resulting feature vector of window re-sponses, describing the degree of crowd-like appearance around an image location as the surrounding spatial extent is increased. 1
Ognjen Arandjelovic
BMVC1
2008 Colour invariants for machine face recognition
abstract
Illumination invariance remains the most researched, yet the most challenging aspect of automatic face recognition. In this paper we investigate the discriminative power of colour-based invariants in the presence of large illumination changes between training and test data, when appearance changes due to cast shadows and non-Lambertian effects are significant. Specifically, there are three main contributions: (i) we employ a more sophisticated photometric model of the camera and show how its parameters can be estimated, (ii) we derive several novel colour-based face invariants, and (iii) on a large database of video sequences we examine and evaluate the largest number of colour-based representations in the literature. Our results suggest that colour invariants do have a substantial discriminative power which may increase the robustness and accuracy of recognition from low resolution images.
Ognjen Arandjelovic, Roberto Cipolla
FG1
2007 Boosted manifold principal angles for image set-based recognition
Tae-Kyun Kim 0001, Ognjen Arandjelovic, Roberto Cipolla
Pattern Recognit.2
2006 On Person Authentication by Fusing Visual and Thermal Face Biometrics
abstract
Recognition algorithms that use data obtained by imaging faces in the thermal spectrum are promising in achieving invariance to extreme illumination changes that are often present in practice. In this paper we analyze the performance of a recently proposed face recognition algorithm that combines visual and thermal modalities by decision level fusion. We examine (i) the effects of the proposed data preprocessing in each domain, (ii) the contribution to improved recognition of different types of features, (iii) the importance of prescription glasses detection, in the context of both 1-to-N and 1-to-1 matching (recognition vs. verification performance). Finally, we discuss the significance of our results and, in particular, identify a number of limitations of the current state-of-the-art and propose promising directions for future research.
Ognjen Arandjelovic, Riad I. Hammoud, Roberto Cipolla
AVSS1
2006 Automatic Cast Listing in Feature-Length Films with Anisotropic Manifold Space
abstract
Our goal is to automatically determine the cast of a feature-length film. This is challenging because the cast size is not known, with appearance changes of faces caused by extrinsic imaging factors (illumination, pose, expression) often greater than due to differing identities. The main contribution of this paper is an algorithm for clustering over face appearance manifolds. Specifically: (i) we develop a novel algorithm for exploiting coherence of dissimilarities between manifolds, (ii) we show how to estimate the optimal dataset-specific discriminant manifold starting from a generic one, and (iii) we describe a fully automatic, practical system based on the proposed algorithm. The performance of the system is evaluated on well-known featurelength films and situation comedies on which it is shown to produce good results.
Ognjen Arandjelovic, Roberto Cipolla
CVPR (2)1
2006 Face Recognition from Video Using the Generic Shape-Illumination Manifold
Ognjen Arandjelovic, Roberto Cipolla
ECCV (4)1
2006 Semantic Photo Synthesis
abstract
Abstract Composite images are synthesized from existing photographs by artists who make concept art, e.g., storyboards for movies or architectural planning. Current techniques allow an artist to fabricate such an image by digitally splicing parts of stock photographs. While these images serve mainly to “quickly”convey how a scene should look, their production is laborious. We propose a technique that allows a person to design a new photograph with substantially less effort. This paper presents a method that generates a composite image when a user types in nouns, such as “boat”and “sand.”The artist can optionally design an intended image by specifying other constraints. Our algorithm formulates the constraints as queries to search an automatically annotated image database. The desired photograph, not a collage, is then synthesized using graph‐cut optimization, optionally allowing for further user interaction to edit or choose among alternative generated photos. An implementation of our approach, shown in the associated video, demonstrates our contributions of (1) a method for creating specific images with minimal human effort, and (2) a combined algorithm for automatically building an image library with semantic annotations from any photo collection.
Matthew Johnson 0003, Gabriel J. Brostow, Jamie Shotton, Ognjen Arandjelovic, Vivek Kwatra, Roberto Cipolla
Comput. Graph. Forum4
2006 An information-theoretic approach to face recognition from face motion manifolds
Ognjen Arandjelovic, Roberto Cipolla
Image Vis. Comput.1
2005 Incremental Learning of Temporally-Coherent Gaussian Mixture Models
abstract
In this paper we address the problem of learning Gaussian Mixture Models (GMMs) incrementally.Unlike previous approaches which universally assume that new data comes in blocks representable by GMMs which are then merged with the current model estimate, our method works for the case when novel data points arrive oneby-one, while requiring little additional memory.We keep only two GMMs in the memory and no historical data.The current fit is updated with the assumption that the number of components is fixed, which is increased (or reduced) when enough evidence for a new component is seen.This is deduced from the change from the oldest fit of the same complexity, termed the Historical GMM, the concept of which is central to our method.The performance of the proposed method is demonstrated qualitatively and quantitatively on several synthetic data sets and video sequences of faces acquired in realistic imaging conditions.
Ognjen Arandjelovic, Roberto Cipolla
BMVC1
2005 Learning over Sets using Boosted Manifold Principal Angles (BoMPA)
abstract
In this paper we address the problem of classifying vector sets. We motivate and introduce a novel method based on comparisons between corresponding vector subspaces. In particular, there are two main areas of novelty: (i) we extend the concept of principal angles between linear subspaces to manifolds with arbitrary nonlinearities; (ii) it is demonstrated how boosting can be used for application-optimal principal angle fusion. The strengths of the proposed method are empirically demonstrated on the task of automatic face recognition (AFR), in which it is shown to outperform state-of-the-art methods in the literature.
Tae-Kyun Kim 0001, Ognjen Arandjelovic, Roberto Cipolla
BMVC2
2005 Face Recognition with Image Sets Using Manifold Density Divergence
abstract
In many automatic face recognition applications, a set of a person's face images is available rather than a single image. In this paper, we describe a novel method for face recognition using image sets. We propose a flexible, semi-parametric model for learning probability densities confined to highly non-linear but intrinsically low-dimensional manifolds. The model leads to a statistical formulation of the recognition problem in terms of minimizing the divergence between densities estimated on these manifolds. The proposed method is evaluated on a large data set, acquired in realistic imaging conditions with severe illumination variation. Our algorithm is shown to match the best and outperform other state-of-the-art algorithms in the literature, achieving 94% recognition rate on average.
Ognjen Arandjelovic, Gregory Shakhnarovich, John Fisher, Roberto Cipolla, Trevor Darrell
CVPR (1)1
2005 Automatic Face Recognition for Film Character Retrieval in Feature-Length Films
abstract
The objective of this work is to recognize all the frontal faces of a character in the closed world of a movie or situation comedy, given a small number of query faces. This is challenging because faces in a feature-length film are relatively uncontrolled with a wide variability of scale, pose, illumination, and expressions, and also may be partially occluded. We develop a recognition method based on a cascade of processing steps that normalize for the effects of the changing imaging environment. In particular there are three areas of novelty: (i) we suppress the background surrounding the face, enabling the maximum area of the face to be retained for recognition rather than a subset; (ii) we include a pose refinement step to optimize the registration between the test image and face exemplar; and (iii) we use robust distance to a sub-space to allow for partial occlusion and expression change. The method is applied and evaluated on several feature length films. It is demonstrated that high recall rates (over 92%) can be achieved whilst maintaining good precision (over 93%).
Ognjen Arandjelovic, Andrew Zisserman
CVPR (1)1
2004 An Illumination Invariant Face Recognition System for Access Control using Video
abstract
Illumination and pose invariance are the most challenging aspects of face recognition. In this paper we describe a fully automatic face recognition system that uses video information to achieve illumination and pose robustness. In the proposed method, highly nonlinear manifolds of face motion are approximated using three Gaussian pose clusters. Pose robustness is achieved by comparing the corresponding pose clusters and probabilistically combining the results to derive a measure of similarity between two manifolds. Illumination is normalized on a per-pose basis. Region-based gamma intensity correction is used to correct for coarse illumination changes, while further refinement is achieved by combining a learnt linear manifold of illumination variation with constraints on face pattern distribution, derived from video. Comparative experimental evaluation is presented and the proposed method is shown to greatly outperform state-of-the-art algorithms. Consistent recognition rates of 94-100% are achieved across dramatic changes in illumination. 1
Ognjen Arandjelovic, Roberto Cipolla
BMVC1