VLDB 2026 Research / reviewers in the wild / expert
Timothy M. Hospedales
dblp:32/3545
· DBLP profile ↗
194ranked-venue papers
12as first author
75since 2021 · last 2026
0000-0003-4867-7486ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 169 · 10 first-author · 67 since 2021Graphics, computer vision, multimedia, augmented reality and games · 118 · 4 first-author · 35 since 2021Systems, architecture and hardware · 4Databases, data management, data science and information retrieval · 3 · 3 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedP²EFT: Federated Learning to Personalize PEFT for Multilingual LLMsabstractFederated learning (FL) has enabled training of multilingual large language models (LLMs) on diverse and decentralized multilingual data, especially on low-resource languages. To improve client-specific performance, personalization via the use of parameter-efficient fine-tuning (PEFT) modules such as LoRA is common. This involves a personalization strategy (PS), such as the design of the PEFT adapter structures (e.g., in which layers to add LoRAs and what ranks) and choice of hyperparameters (e.g., learning rates) for fine-tuning. Instead of manual PS configuration, we propose FedP²EFT, a federated learning-to-personalize method for multilingual LLMs in cross-device FL settings. Unlike most existing PEFT structure selection methods, which are prone to overfitting low-data regimes, FedP²EFT collaboratively learns the optimal personalized PEFT structure for each client via Bayesian sparse rank selection. Evaluations on both simulated and real-world multilingual FL benchmarks demonstrate that FedP²EFT largely outperforms existing personalized fine-tuning methods, while complementing other existing FL methods. Royson Lee, Minyoung Kim 0001, Fady Rezk, Rui Li 0052, Stylianos I. Venieris, Timothy M. Hospedales |
AAAI | 6 |
| 2026 | MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in VideoLMs for Multimodal Sarcasm Detection
Anisha Saha, Varsha Suresh, Timothy M. Hospedales, Vera Demberg |
LREC | 3 |
| 2026 | Boosting Multimodal Chain of Thought Reasoning by Selective Mixture of Experts
Qilei Li, Shitong Sun, Da Li 0001, Timothy M. Hospedales, Shaogang Gong |
Pattern Recognit. | 4 |
| 2025 | A Stochastic Approach to Bi-Level Optimization for Hyperparameter Optimization and Meta LearningabstractWe tackle the general differentiable meta learning problem that is ubiquitous in modern deep learning, including hyperparameter optimization, loss function learning, few-shot learning and more. These problems are often formalized as Bi-Level Optimizations (BLO). We introduce a novel perspective by turning a given BLO problem into a stochastic optimization, where the inner loss function becomes a smooth probability distribution, and the outer loss becomes an expected loss over the inner distribution. To solve this stochastic optimization, we adopt Stochastic Gradient Langevin Dynamics (SGLD) MCMC to sample inner distribution, and propose a recurrent algorithm to compute the MC-estimated hypergradient. Our derivation is similar to forward-mode differentiation, but we introduce a new first-order approximation that makes it feasible for large models without needing to store huge Jacobian matrices. The main benefits are two fold: i) Our stochastic formulation takes into account uncertainty, which makes the method robust to suboptimal inner optimization or non-unique multiple inner minima due to overparametrization; ii) Compared to existing methods that often exhibit unstable behavior and hyperparameter sensitivity in practice, our method leads to considerably more reliable solutions. We demonstrate that the new approach achieves promising results on diverse meta learning problems and easily scales to learning 87M hyper-parameters in the case of Vision Transformers. Minyoung Kim 0001, Timothy M. Hospedales |
AAAI | 2 |
| 2025 | FW-Merging: Scaling Model Merging with Frank-Wolfe Optimization
Hao Mark Chen, Shell Xu Hu, Wayne Luk, Timothy M. Hospedales, Hongxiang Fan |
ICCV | 4 |
| 2025 | ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron PruningabstractWhile large-scale text-to-image diffusion models have demonstrated impressive image-generation capabilities, there are significant concerns about their potential misuse for generating unsafe content, violating copyright, and perpetuating societal biases. Recently, the text-to-image generation community has begun addressing these concerns by editing or unlearning undesired concepts from pre-trained models. However, these methods often involve data-intensive and inefficient fine-tuning or utilize various forms of token remapping, rendering them susceptible to adversarial jailbreaks. In this paper, we present a simple and effective training-free approach, ConceptPrune, wherein we first identify critical regions within pre-trained models responsible for generating undesirable concepts, thereby facilitating straightforward concept unlearning via weight pruning. Experiments across a range of concepts including artistic styles, nudity, and object erasure demonstrate that target concepts can be efficiently erased by pruning a tiny fraction, approximately 0.12% of total weights, enabling multi-concept erasure and robustness against various white-box and black-box adversarial attacks. Ruchika Chavhan, Da Li 0001, Timothy M. Hospedales |
ICLR | 3 |
| 2025 | LiFT: Learning to Fine-Tune via Bayesian Parameter Efficient Meta Fine-TuningabstractWe tackle the problem of parameter-efficient fine-tuning (PEFT) of a pre-trained large deep model on many different but related tasks. Instead of the simple but strong baseline strategy of task-wise independent fine-tuning, we aim to meta-learn the core shared information that can be used for unseen test tasks to improve the prediction performance further. That is, we propose a method for {\em learning-to-fine-tune} (LiFT). LiFT introduces a novel hierarchical Bayesian model that can be superior to both existing general meta learning algorithms like MAML and recent LoRA zoo mixing approaches such as LoRA-Retriever and model-based clustering. In our Bayesian model, the parameters of the task-specific LoRA modules are regarded as random variables where these task-wise LoRA modules are governed/regularized by higher-level latent random variables, which represents the prior of the LoRA modules that capture the shared information across all training tasks. To make the posterior inference feasible, we propose a novel SGLD-Gibbs sampling algorithm that is computationally efficient. To represent the posterior samples from the SGLD-Gibbs, we propose an online EM algorithm that maintains a Gaussian mixture representation for the posterior in an online manner in the course of iterative posterior sampling. We demonstrate the effectiveness of LiFT on NLP and vision multi-task meta learning benchmarks. Minyoung Kim 0001, Timothy M. Hospedales |
ICLR | 2 |
| 2025 | VL-ICL Bench: The Devil in the Details of Multimodal In-Context LearningabstractLarge language models (LLMs) famously exhibit emergent in-context learning (ICL) - the ability to rapidly adapt to new tasks using few-shot examples provided as a prompt, without updating the model's weights. Built on top of LLMs, vision large language models (VLLMs) have advanced significantly in areas such as recognition, reasoning, and grounding. However, investigations into multimodal ICL have predominantly focused on few-shot visual question answering (VQA), and image captioning, which we will show neither exploit the strengths of ICL, nor test its limitations. The broader capabilities and limitations of multimodal ICL remain under-explored. In this study, we introduce a comprehensive benchmark VL-ICL Bench for multimodal in-context learning, encompassing a broad spectrum of tasks that involve both images and text as inputs and outputs, and different types of challenges, from {perception to reasoning and long context length}. We evaluate the abilities of state-of-the-art VLLMs against this benchmark suite, revealing their diverse strengths and weaknesses, and showing that even the most advanced models, such as GPT-4, find the tasks challenging. By highlighting a range of new ICL tasks, and the associated strengths and limitations of existing models, we hope that our dataset will inspire future work on enhancing the in-context learning capabilities of VLLMs, as well as inspire new applications that leverage VLLM ICL. Project page is at https://ys-zong.github.io/VL-ICL/ Yongshuo Zong, Ondrej Bohdal, Timothy M. Hospedales |
ICLR | 3 |
| 2025 | HyperIV: Real-time Implied Volatility SmoothingabstractWe propose HyperIV, a novel approach for real-time implied volatility smoothing that eliminates the need for traditional calibration procedures. Our method employs a hypernetwork to generate parameters for a compact neural network that constructs complete volatility surfaces within 2 milliseconds, using only 9 market observations. Moreover, the generated surfaces are guaranteed to be free of static arbitrage. Extensive experiments across 8 index options demonstrate that HyperIV achieves superior accuracy compared to existing methods while maintaining computational efficiency. The model also exhibits strong cross-asset generalization capabilities, indicating broader applicability across different market instruments. These key features -- rapid adaptation to market conditions, guaranteed absence of arbitrage, and minimal data requirements -- make HyperIV particularly valuable for real-time trading applications. We make code available at https://github.com/qmfin/hyperiv. Yongxin Yang, Chao Shu, Timothy M. Hospedales |
ICML | 4 |
| 2025 | MemControl: Mitigating Memorization in Diffusion Models via Automated Parameter SelectionabstractDiffusion models excel in generating images that closely resemble their training data but are also susceptible to data memorization, raising privacy, ethical, and legal concerns, particularly in sensitive domains such as medical imaging. We hypothesize that this memorization stems from the overparameterization of deep models and propose that regularizing model capacity during fine-tuning can mitigate this issue. Firstly, we empirically show that regulating the model capacity via Parameter-efficient fine-tuning (PEFT) mitigates memorization to some extent, however, it further requires the identification of the exact parameter subsets to be fine-tuned for high-quality generation. To identify these subsets, we introduce a bilevel optimization framework, MemControl, that automates parameter selection using memorization and generation quality metrics as rewards during fine-tuning. The parameter subsets discovered through MemControl achieve a superior tradeoff between generation quality and memorization. For the task of medical image generation, our approach outperforms existing state-of-the-art memorization mitigation strategies by fine-tuning as few as 0.019% of model parameters. Moreover, we demonstrate that the discovered parameter subsets are transferable to non-medical domains. Our framework is scalable to large datasets, agnostic to reward functions, and can be integrated with existing approaches for further memorization mitigation. To the best of our knowledge, this is the first study to empirically evaluate memorization in medical images and propose a targeted yet universal mitigation strategy. The code is available at this https URL. Raman Dutt, Ondrej Bohdal, Pedro Sanchez, Sotirios A. Tsaftaris, Timothy M. Hospedales |
WACV | 5 |
| 2025 | FedHB: Hierarchical Bayesian Federated LearningabstractWe propose a novel hierarchical Bayesian approach to Federated Learning (FL), where our model reasonably describes the generative process of clients' local data via hierarchical Bayesian modeling: constituting random variables of local models for clients that are governed by a higher-level global variate. Interestingly, the variational inference in our Bayesian model leads to an optimisation problem whose block-coordinate descent solution becomes a distributed algorithm that is separable over clients and allows them not to reveal their own private data at all, thus fully compatible with FL. We also highlight that our block-coordinate algorithm has particular forms that subsume the well-known FL algorithms including Fed-Avg and Fed-Prox as special cases. Beyond introducing novel modeling and derivations, we also offer convergence analysis showing that our block-coordinate FL algorithm converges to an (local) optimum of the objective at the rate of $O(1/\sqrt{t})$, the same rate as regular (centralised) SGD, as well as the generalisation error analysis where we prove that the test error of our model on unseen data is guaranteed to vanish as we increase the training data size, thus asymptotically optimal. Minyoung Kim 0001, Timothy M. Hospedales |
J. Mach. Learn. Res. | 2 |
| 2025 | Self-Supervised Multimodal Learning: A SurveyabstractMultimodal learning, which aims to understand and analyze information from multiple modalities, has achieved substantial progress in the supervised regime in recent years. However, the heavy dependence on data paired with expensive human annotations impedes scaling up models. Meanwhile, given the availability of large-scale unannotated data in the wild, self-supervised learning has become an attractive strategy to alleviate the annotation bottleneck. Building on these two directions, self-supervised multimodal learning (SSML) provides ways to learn from raw multimodal data. In this survey, we provide a comprehensive review of the state-of-the-art in SSML, in which we elucidate three major challenges intrinsic to self-supervised learning with multimodal data: 1) learning representations from multimodal data without labels, 2) fusion of different modalities, and 3) learning with unaligned data. We then detail existing solutions to these challenges. Specifically, we consider 1) objectives for learning from multimodal unlabeled data via self-supervision, 2) model architectures from the perspective of different multimodal fusion strategies, and 3) pair-free learning strategies for coarse-grained and fine-grained alignment. We also review real-world applications of SSML algorithms in diverse fields, such as healthcare, remote sensing, and machine translation. Finally, we discuss challenges and future directions for SSML. Yongshuo Zong, Oisin Mac Aodha, Timothy M. Hospedales |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | SketchINR: A First Look into Sketches as Implicit Neural RepresentationsabstractWe propose SketchINR, to advance the representation of vector sketches with implicit neural models. A variable length vector sketch is compressed into a latent space of fixed dimension that implicitly encodes the underlying shape as a function of time and strokes. The learned function predicts the xy point coordinates in a sketch at each time and stroke. Despite its simplicity, SketchINR outperforms existing representations at multiple tasks: (i) Encoding an entire sketch dataset into a fixed size latent vector, SketchINR gives 60× and 10× data compression over raster and vector sketches, respectively. (ii) SketchINR's auto-decoder provides a much higher-fidelity representation than other learned vector sketch representations, and is uniquely able to scale to complex vector sketches such as FS-COCO. (iii) SketchINR supports parallelisation that can decode/render ∼100× faster than other learned vector representations such as SketchRNN. (iv) SketchINR, for the first time, emulates the human ability to reproduce a sketch with varying abstraction in terms of number and complexity of strokes. As a first look at implicit sketches, SketchINR's compact high-fidelity representation will support future work in modelling long and complex sketches. Hmrishav Bandyopadhyay, Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Aneeshan Sain, Tao Xiang 0002, Timothy M. Hospedales, Yi-Zhe Song |
CVPR | 6 |
| 2024 | DemoFusion: Democratising High-Resolution Image Generation With No $$$abstractHigh-resolution image generation with Generative Artificial Intelligence (GenAl) has immense potential but, due to the enormous capital investment required for training, it is increasimgly centralised to a few large corporations, and hidden behind paywalls. This paper aims to democratise high-resolution GenAl by advancing the frontier of high-resolution generation while remaining accessible to a broad audience. We demonstrate that existing Latent Diffusion Models (LDMs) possess untapped potential for higher-resolution image generation. Our novel DemoFusion framework seamlessly extends open-source GenAl models, employing Progressive Upscaling, Skip Residual, and Di-lated Sampling mechanisms to achieve higher-resolution image generation. The progressive nature of DemoFusion requires more passes, but the intermediate results can serve as “previews”, facilitating rapid prompt iteration. Ruoyi Du, Dongliang Chang, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma |
CVPR | 3 |
| 2024 | FairTune: Optimizing Parameter Efficient Fine Tuning for Fairness in Medical Image AnalysisabstractTraining models with robust group fairness properties is crucial in ethically sensitive application areas such as medical diagnosis. Despite the growing body of work aiming to minimise demographic bias in AI, this problem remains challenging. A key reason for this challenge is the fairness generalisation gap: High-capacity deep learning models can fit all training data nearly perfectly, and thus also exhibit perfect fairness during training. In this case, bias emerges only during testing when generalisation performance differs across sub-groups. This motivates us to take a bi-level optimisation perspective on fair learning: Optimising the learning strategy based on validation fairness. Specifically, we consider the highly effective workflow of adapting pre-trained models to downstream medical imaging tasks using parameter-efficient fine-tuning (PEFT) techniques. There is a trade-off between updating more parameters, enabling a better fit to the task of interest vs. fewer parameters, potentially reducing the generalisation gap. To manage this tradeoff, we propose FairTune, a framework to optimise the choice of PEFT parameters with respect to fairness. We demonstrate empirically that FairTune leads to improved fairness on a range of medical imaging datasets. The code is available at https://github.com/Raman1121/FairTune. Raman Dutt, Ondrej Bohdal, Sotirios A. Tsaftaris, Timothy M. Hospedales |
ICLR | 4 |
| 2024 | Neural Fine-Tuning Search for Few-Shot LearningabstractIn few-shot recognition, a classifier that has been trained on one set of classes is required to rapidly adapt and generalize to a disjoint, novel set of classes. To that end, recent studies have shown the efficacy of fine-tuning with carefully-crafted adaptation architectures. However this raises the question of: How can one design the optimal adaptation strategy? In this paper, we study this question through the lens of neural architecture search (NAS). Given a pre-trained neural network, our algorithm discovers the optimal arrangement of adapters, which layers to keep frozen, and which to fine-tune. We demonstrate the generality of our NAS method by applying it to both residual networks and vision transformers and report state-of-the-art performance on Meta-Dataset and Meta-Album. Panagiotis Eustratiadis, Lukasz Dudziak, Da Li 0001, Timothy M. Hospedales |
ICLR | 4 |
| 2024 | A Hierarchical Bayesian Model for Few-Shot Meta LearningabstractWe propose a novel hierarchical Bayesian model for the few-shot meta learning problem. We consider episode-wise random variables to model episode-specific generative processes, where these local random variables are governed by a higher-level global random variable. The global variable captures information shared across episodes, while controlling how much the model needs to be adapted to new episodes in a principled Bayesian manner. Within our framework, prediction on a novel episode/task can be seen as a Bayesian inference problem. For tractable training, we need to be able to relate each local episode-specific solution to the global higher-level parameters. We propose a Normal-Inverse-Wishart model, for which establishing this local-global relationship becomes feasible due to the approximate closed-form solutions for the local posterior distributions. The resulting algorithm is more attractive than the MAML in that it does not maintain a costly computational graph for the sequence of gradient descent steps in an episode. Our approach is also different from existing Bayesian meta learning methods in that rather than modeling a single random variable for all episodes, it leverages a hierarchical structure that exploits the local-global relationships desirable for principled Bayesian learning with many related tasks. Minyoung Kim 0001, Timothy M. Hospedales |
ICLR | 2 |
| 2024 | Recurrent Early Exits for Federated Learning with Heterogeneous ClientsabstractFederated learning (FL) has enabled distributed learning of a model across multiple clients in a privacy-preserving manner. One of the main challenges of FL is to accommodate clients with varying hardware capacities; clients have differing compute and memory requirements. To tackle this challenge, recent state-of-the-art approaches leverage the use of early exits. Nonetheless, these approaches fall short of mitigating the challenges of joint learning multiple exit classifiers, often relying on hand-picked heuristic solutions for knowledge distillation among classifiers and/or utilizing additional layers for weaker classifiers. In this work, instead of utilizing multiple classifiers, we propose a recurrent early exit approach named ReeFL that fuses features from different sub-models into a single shared classifier. Specifically, we use a transformer-based early-exit module shared among sub-models to i) better exploit multi-layer feature representations for task-specific prediction and ii) modulate the feature representation of the backbone model for subsequent predictions. We additionally present a per-client self-distillation approach where the best sub-model is automatically selected as the teacher of the other sub-models at each client. Our experiments on standard image and speech classification benchmarks across various emerging federated fine-tuning baselines demonstrate ReeFL effectiveness over previous works. Royson Lee, Javier Fernández-Marqués, Shell Xu Hu, Da Li 0001, Stefanos Laskaridis, Lukasz Dudziak, Timothy M. Hospedales, Ferenc Huszar, Nicholas D. Lane |
ICML | 7 |
| 2024 | Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language ModelsabstractCurrent vision large language models (VLLMs) exhibit remarkable capabilities yet are prone to generate harmful content and are vulnerable to even the simplest jailbreaking attacks. Our initial analysis finds that this is due to the presence of harmful data during vision-language instruction fine-tuning, and that VLLM fine-tuning can cause forgetting of safety alignment previously learned by the underpinning LLM. To address this issue, we first curate a vision-language safe instruction-following dataset VLGuard covering various harmful categories. Our experiments demonstrate that integrating this dataset into standard vision-language fine-tuning or utilizing it for post-hoc fine-tuning effectively safety aligns VLLMs. This alignment is achieved with minimal impact on, or even enhancement of, the models' helpfulness. The versatility of our safety fine-tuning dataset makes it a valuable resource for safety-testing existing VLLMs, training new models or safeguarding pre-trained VLLMs. Empirical results demonstrate that fine-tuned VLLMs effectively reject unsafe instructions and substantially reduce the success rates of several black-box adversarial attacks, which approach zero in many cases. The code and dataset will be open-sourced. Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang, Timothy M. Hospedales |
ICML | 5 |
| 2024 | Fool Your (Vision and) Language Model with Embarrassingly Simple PermutationsabstractLarge language and vision-language models are rapidly being deployed in practice thanks to their impressive capabilities in instruction following, in-context learning, and so on. This raises an urgent need to carefully analyse their robustness so that stakeholders can understand if and when such models are trustworthy enough to be relied upon in any given application. In this paper, we highlight a specific vulnerability in popular models, namely permutation sensitivity in multiple-choice question answering (MCQA). Specifically, we show empirically that popular models are vulnerable to adversarial permutation in answer sets for multiple-choice prompting, which is surprising as models should ideally be as invariant to prompt permutation as humans are. These vulnerabilities persist across various model sizes, and exist in very recent language and vision-language models. Code to reproduce all experiments is provided in supplementary materials. Yongshuo Zong, Tingyang Yu, Ruchika Chavhan, Bingchen Zhao, Timothy M. Hospedales |
ICML | 5 |
| 2024 | From Pixels to Waveforms: Evaluating Pre-trained Image Models for Few-Shot Audio ClassificationabstractWithin machine learning literature, few-shot learning has emerged as a crucial tool in many use-cases, enabling adaptation to new tasks quickly with limited supervision. While the majority of this research has concentrated on the image domain, recent efforts have extended its scope to others. A prevalent strategy in addressing few-shot learning involves harnessing pre-trained models—whether supervised or otherwise—as feature extractors. These models, paired with a lightweight trainable head, demonstrate remarkable efficacy, often achieving near or state-of-the-art results. However, this method encounters challenges in domains like audio, where the number of available and regularly published pre-trained models is comparatively low. Given that audio signals can be represented as spectrograms, that are akin to traditional imagery, a fundamental question arises: Can off-the-shelf pre-trained image models prove beneficial for few-shot audio classification? Additionally, can insights from the image domain guide model and approach selection in the audio domain? Our investigation yields diverse insights, showcasing the effectiveness of both supervised and self-supervised image-pre-trained models for few-shot audio classification. Alongside this, we identify strong relationships between the common Cross-Domain few-shot imagery learning settings and few-shot audio performance. Calum Heggan, Timothy M. Hospedales, Sam Budgett, Mehrdad Yaghoobi |
IJCNN | 2 |
| 2024 | A Bayesian Approach to Data Point SelectionabstractData point selection (DPS) is becoming a critical topic in deep learning due to the ease of acquiring uncurated training data compared to the difficulty of obtaining curated or processed data.
Existing approaches to DPS are predominantly based on a bi-level optimisation (BLO) formulation, which is demanding in terms of memory and computation, and exhibits some theoretical defects regarding minibatches.
Thus, we propose a novel Bayesian approach to DPS. We view the DPS problem as posterior inference in a novel Bayesian model where the posterior distributions of the instance-wise weights and the main neural network parameters are inferred under a reasonable prior and likelihood model.
We employ stochastic gradient Langevin MCMC sampling to learn the main network and instance-wise weights jointly, ensuring convergence even with minibatches. Our update equation is comparable to the widely used SGD and much more efficient than existing BLO-based methods. Through controlled experiments in both the vision and language domains, we present the proof-of-concept. Additionally, we demonstrate that our method scales effectively to large language models and facilitates automated per-task optimization for instruction fine-tuning datasets. Xinnuo Xu, Minyoung Kim 0001, Royson Lee, Brais Martínez, Timothy M. Hospedales |
NeurIPS | 5 |
| 2024 | Feed-Forward Latent Domain AdaptationabstractWe study a new highly-practical problem setting that enables resource-constrained edge devices to adapt a pre-trained model to their local data distributions. Recognizing that device’s data are likely to come from multiple latent domains that include a mixture of unlabelled domain-relevant and domain-irrelevant examples, we focus on the comparatively under-studied problem of latent domain adaptation. Considering limitations of edge devices, we aim to only use a pre-trained model and adapt it in a feed-forward way, without using back-propagation and without access to the source data. Modelling these realistic constraints bring us to the novel and practically important problem setting of feed-forward latent domain adaptation. Our solution is to meta-learn a network capable of embedding the mixed-relevance target dataset and dynamically adapting inference for target examples using cross-attention. The resulting framework leads to consistent improvements over strong ERM baselines. We also show that our framework sometimes even improves on the upper bound of domain-supervised adaptation, where only domain-relevant instances are provided for adaptation. This suggests that human annotated domain labels may not always be optimal, and raises the possibility of doing better through automated instance selection. Ondrej Bohdal, Da Li 0001, Shell Xu Hu, Timothy M. Hospedales |
WACV | 4 |
| 2024 | Meta-Learned Kernel For Blind Super-Resolution Kernel EstimationabstractRecent image degradation estimation methods have enabled single-image super-resolution (SR) approaches to better upsample real-world images. Among these methods, explicit kernel estimation approaches have demonstrated unprecedented performance at handling unknown degradations. Nonetheless, a number of limitations constrain their efficacy when used by downstream SR models. Specifically, this family of methods yields i) excessive inference time due to long per-image adaptation times and ii) inferior image fidelity due to kernel mismatch. In this work, we introduce a learning-to-learn approach that meta-learns from the information contained in a distribution of images, thereby enabling significantly faster adaptation to new images with substantially improved performance in both kernel estimation and image fidelity. Specifically, we meta-train a kernelgenerating GAN, named MetaKernelGAN, on a range of tasks, such that when a new image is presented, the generator starts from an informed kernel estimate and the discriminator starts with a strong capability to distinguish between patch distributions. Compared with state-of-the-art methods, our experiments show that MetaKernelGAN better estimates the magnitude and covariance of the kernel, leading to state-of-the-art blind SR results within a similar computational regime when combined with a non-blind SR model. Through supervised learning of an unsupervised learner, our method maintains the generalizability of the unsupervised learner, improves the optimization stability of kernel estimation, and hence image adaptation, and leads to a faster inference with a speedup between 14.24 to 102.1× over existing methods.0 Royson Lee, Rui Li 0052, Stylianos I. Venieris, Timothy M. Hospedales, Ferenc Huszar, Nicholas D. Lane |
WACV | 4 |
| 2024 | Editorial: Learning With Fewer Labels in Computer VisionabstractUndoubtedly, Deep Neural Networks (DNNs), from AlexNet to ResNet to Transformer, have sparked revolutionary advancements in diverse computer vision tasks. The scale of DNNs has grown exponentially due to the rapid development of computational resources. Despite the tremendous success, DNNs typically depend on massive amounts of training data (especially the recent various foundation models) to achieve high performance and are brittle in that their performance can degrade severely with small changes in their operating environment. Generally, collecting massive-scale training datasets is costly or even infeasible, as for certain fields, only very limited or no examples at all can be gathered. Nevertheless, collecting, labeling, and vetting massive amounts of practical training data is certainly difficult and expensive, as it requires the painstaking efforts of experienced human annotators or experts, and in many cases, prohibitively costly or impossible due to some reason, such as privacy, safety or ethic issues. Li Liu 0002, Timothy M. Hospedales, Yann LeCun, Mingsheng Long, Jiebo Luo 0001, Wanli Ouyang, Matti Pietikäinen, Tinne Tuytelaars |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Sketch-based Video Object Segmentation: Benchmark and Analysis
Ruolin Yang 0001, Da Li 0001, Conghui Hu, Timothy M. Hospedales, Honggang Zhang 0002, Yi-Zhe Song |
BMVC | 4 |
| 2023 | Meta Omnium: A Benchmark for General-Purpose Learning-to-LearnabstractMeta-learning and other approaches to few-shot learning are widely studied for image recognition, and are increasingly applied to other vision tasks such as pose estimation and dense prediction. This naturally raises the question of whether there is any fewshot metalearning algorithm capable of generalizing across these diverse task types? To support the community in answering this question, we introduce Meta Omnium, a dataset-of-datasets spanning multiple vision tasks including recognition, keypoint localization, semantic segmentation and regression. We experiment with popular fewshot metalearning baselines and analyze their ability to generalize across tasks and to transfer knowledge between them. Meta Omnium enables metalearning researchers to evaluate model generalization to a much wider array of tasks than previously possible, and provides a single framework for evaluating meta-learners across a wide suite of vision applications in a consistent manner. Code and dataset are available at https://github.com/edi-meta-learning/meta-omnium. Ondrej Bohdal, Yinbing Tian, Yongshuo Zong, Ruchika Chavhan, Da Li 0001, Henry Gouk, Timothy M. Hospedales |
CVPR | 8 |
| 2023 | An Erudite Fine-Grained Visual Classification ModelabstractCurrent fine-grained visual classification (FGVC) models are isolated. In practice, we first need to identify the coarse-grained label of an object, then select the corresponding FGVC model for recognition. This hinders the application of FGVC algorithms in real-life scenarios. In this paper, we propose an erudite FGVC model jointly trained by several different datasets11In this paper, different datasets mean different fine-grained visual classification datasets., which can efficiently and accurately predict an object's fine-grained label across the combined label space. We found through a pilot study that positive and negative transfers co-occur when different datasets are mixed for training, i.e., the knowledge from other datasets is not always useful. Therefore, we first propose a feature disentanglement module and a feature re-fusion module to reduce negative transfer and boost positive transfer between different datasets. In detail, we reduce negative transfer by decoupling the deep features through many dataset-specific feature extractors. Subsequently, these are channel-wise re-fused to facilitate positive transfer. Finally, we propose a meta-learning based dataset-agnostic spatial attention layer to take full advantage of the multi-dataset training data, given that localisation is dataset-agnostic between different datasets. Experimental results across 11 different mixed-datasets built on four different FGVC datasets demonstrate the effectiveness of the proposed method. Furthermore, the proposed method can be easily combined with existing FGVC methods to obtain state-of-the-art results. Our code is available at https://github.com/PRIS-CV/An-Erudite-FGVC-Model. Dongliang Chang, Yujun Tong, Ruoyi Du, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma |
CVPR | 4 |
| 2023 | On-the-Fly Category DiscoveryabstractAlthough machines have surpassed humans on visual recognition problems, they are still limited to providing closed-set answers. Unlike machines, humans can cognize novel categories at the first observation. Novel category discovery (NCD) techniques, transferring knowledge from seen categories to distinguish unseen categories, aim to bridge the gap. However, current NCD methods assume a transductive learning and offline inference paradigm, which restricts them to a predefined query set and renders them unable to deliver instant feedback. In this paper, we study on-the-fly category discovery (OCD) aimed at making the model instantaneously aware of novel category samples (i.e., enabling inductive learning and streaming inference). We first design a hash coding-based expandable recognition model as a practical baseline. Afterwards, noticing the sensitivity of hash codes to intra-category variance, we further propose a novel Sign-Magnitude dIsentangLEment (SMILE) architecture to alleviate the disturbance it brings. Our experimental results demonstrate the superiority of SMILE against our baseline model and prior art. Our code is available at https://github.com/PRIS-CV/On-the-fly-Category-Discovery. Ruoyi Du, Dongliang Chang, Kongming Liang, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma |
CVPR | 4 |
| 2023 | Zero-Shot Everything Sketch-Based Image Retrieval, and in Explainable StyleabstractThis paper studies the problem of zero-short sketch-based image retrieval (ZS-SBIR), however with two significant differentiators to prior art (i) we tackle all variants (inter-category, intracategory, and cross datasets) of ZS-SBIR with just one network (“everything”), and (ii) we would really like to understand how this sketch-photo matching operates (“explainable”). Our key innovation lies with the realization that such a cross-modal matching problem could be reduced to comparisons of groups of key local patches - akin to the seasoned “bag-of-words” paradigm. Just with this change, we are able to achieve both of the aforementioned goals, with the added benefit of no longer requiring external semantic knowledge. Technically, ours is a transformer-based cross-modal network, with three novel components (i) a self-attention module with a learnable tokenizer to produce visual tokens that correspond to the most informative local regions, (ii) a cross-attention module to compute local correspondences between the visual tokens across two modalities, and finally (iii) a kernel-based relation network to assemble local putative matches and produce an overall similarity metric for a sketch-photo pair. Experiments show ours indeed delivers superior performances across all ZS-SBIR settings. The all important explainable goal is elegantly achieved by visualizing cross-modal token correspondences, and for the first time, via sketch to photo synthesis by universal replacement of all matched photo patches. Code and model are available at https://github.com/buptLinfy/ZSE-SBIR. Fengyin Lin, Mingkang Li 0003, Da Li 0001, Timothy M. Hospedales, Yi-Zhe Song, Yonggang Qi |
CVPR | 4 |
| 2023 | An Evaluation of Self-supervised Learning for Portfolio Diversification
Yongxin Yang, Timothy M. Hospedales |
ICANN (3) | 2 |
| 2023 | Quality Diversity for Visual Pre-TrainingabstractModels pre-trained on large datasets such as ImageNet provide the de-facto standard for transfer learning, with both supervised and self-supervised approaches proving effective. However, emerging evidence suggests that any single pre-trained feature will not perform well on diverse downstream tasks. Each pre-training strategy encodes a certain inductive bias, which may suit some downstream tasks but not others. Notably, the augmentations used in both supervised and self-supervised training lead to features with high invariance to spatial and appearance transformations. This renders them sub-optimal for tasks that demand sensitivity to these factors. In this paper we develop a feature that better supports diverse downstream tasks by providing a diverse set of sensitivities and invariances. In particular, we are inspired by Quality-Diversity in evolution, to define a pre-training objective that requires high quality yet diverse features — where diversity is defined in terms of transformation (in)variances. Our framework plugs in to both supervised and self-supervised pre-training, and produces a small ensemble of features. We further show how downstream tasks can easily and efficiently select their preferred (in)variances. Both empirical and theoretical analysis show the efficacy of our representation and transfer learning approach for diverse downstream tasks. Code available at https://github.com/ruchikachavhan/quality-diversity-pretraining.git Ruchika Chavhan, Henry Gouk, Da Li 0001, Timothy M. Hospedales |
ICCV | 4 |
| 2023 | Task-aware Adaptive Learning for Cross-domain Few-shot LearningabstractAlthough existing few-shot learning works yield promising results for in-domain queries, they still suffer from weak cross-domain generalization. Limited support data requires effective knowledge transfer, but domain-shift makes this harder. Towards this emerging challenge, researchers improved adaptation by introducing task-specific parameters, which are directly optimized and estimated for each task. However, adding a fixed number of additional parameters fails to consider the diverse domain shifts between target tasks and the source domain, limiting efficacy. In this paper, we first observe the dependence of task-specific parameter configuration on the target task. Abundant task-specific parameters may over-fit, and insufficient task-specific parameters may result in under-adaptation – but the optimal task-specific configuration varies for different test tasks. Based on these findings, we propose the Task-aware Adaptive Network (TA2-Net), which is trained by reinforcement learning to adaptively estimate the optimal task-specific parameter configuration for each test task. It learns, for example, that tasks with significant domain-shift usually have a larger need for task-specific parameters for adaptation. We evaluate our model on Meta-dataset. Empirical results show that our model outperforms existing state-of-the-art methods. Our code is available at https://github.com/PRIS-CV/TA2-Net. Yurong Guo 0001, Ruoyi Du, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma |
ICCV | 4 |
| 2023 | ChiroDiff: Modelling chirographic data with Diffusion Models
Ayan Das 0003, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song |
ICLR | 3 |
| 2023 | Amortised Invariance Learning for Contrastive Self-Supervision
Ruchika Chavhan, Jan Stühmer, Calum Heggan, Mehrdad Yaghoobi, Timothy M. Hospedales |
ICLR | 5 |
| 2023 | Learning where and when to reason in neuro-symbolic inference
Cristina Cornelio, Jan Stühmer, Shell Xu Hu, Timothy M. Hospedales |
ICLR | 4 |
| 2023 | Domain Generalisation via Domain Adaptation: An Adversarial Fourier Amplitude Approach
Minyoung Kim 0001, Da Li 0001, Timothy M. Hospedales |
ICLR | 3 |
| 2023 | MEDFAIR: Benchmarking Fairness for Medical Imaging
Yongshuo Zong, Yongxin Yang, Timothy M. Hospedales |
ICLR | 3 |
| 2023 | MT-SLVR: Multi-Task Self-Supervised Learning for Transformation In(Variant) RepresentationsabstractContrastive self-supervised learning has gained attention for its ability to create high-quality representations from large unlabelled data sets. A key reason that these powerful features enable data-efficient learning of downstream tasks is that they provide augmentation invariance, which is often a useful inductive bias. However, the amount and type of invariances preferred is not known apriori, and varies across different downstream tasks. We therefore propose a multi-task self-supervised framework (MT-SLVR) that learns both variant and invariant features in a parameter-efficient manner. Our multi-task representation provides a strong and flexible feature that benefits diverse downstream tasks. We evaluate our approach on few-shot classification tasks drawn from a variety of audio domains and demonstrate improved classification performance on all of them Calum Heggan, Timothy M. Hospedales, Sam Budgett, Mehrdad Yaghoobi |
INTERSPEECH | 2 |
| 2023 | BayesTune: Bayesian Sparse Deep Model Fine-tuningabstractDeep learning practice is increasingly driven by powerful foundation models (FM), pre-trained at scale and then fine-tuned for specific tasks of interest. A key property of this workflow is the efficacy of performing sparse or parameter-efficient fine-tuning, meaning that by updating only a tiny fraction of the whole FM parameters on a downstream task can lead to surprisingly good performance, often even superior to a full model update. However, it is not clear what is the optimal and principled way to select which parameters to update. Although a growing number of sparse fine-tuning ideas have been proposed, they are mostly not satisfactory, relying on hand-crafted heuristics or heavy approximation. In this paper we propose a novel Bayesian sparse fine-tuning algorithm: we place a (sparse) Laplace prior for each parameter of the FM, with the mean equal to the initial value and the scale parameter having a hyper-prior that encourages small scale. Roughly speaking, the posterior means of the scale parameters indicate how important it is to update the corresponding parameter away from its initial value when solving the downstream task. Given the sparse prior, most scale parameters are small a posteriori, and the few large-valued scale parameters identify those FM parameters that crucially need to be updated away from their initial values. Based on this, we can threshold the scale parameters to decide which parameters to update or freeze, leading to a principled sparse fine-tuning strategy. To efficiently infer the posterior distribution of the scale parameters, we adopt the Langevin MCMC sampler, requiring only two times the complexity of the vanilla SGD. Tested on popular NLP benchmarks as well as the VTAB vision tasks, our approach shows significant improvement over the state-of-the-arts (e.g., 1% point higher than the best SOTA when fine-tuning RoBERTa for GLUE and SuperGLUE benchmarks). Minyoung Kim 0001, Timothy M. Hospedales |
NeurIPS | 2 |
| 2023 | FedL2P: Federated Learning to PersonalizeabstractFederated learning (FL) research has made progress in developing algorithms for distributed learning of global models, as well as algorithms for local personalization of those common models to the specifics of each client’s local data distribution. However, different FL problems may require different personalization strategies, and it may not even be possible to define an effective one-size-fits-all personalization strategy for all clients: Depending on how similar each client’s optimal predictor is to that of the global model, different personalization strategies may be preferred. In this paper, we consider the federated meta-learning problem of learning personalization strategies. Specifically, we consider meta-nets that induce the batch-norm and learning rate parameters for each client given local data statistics. By learning these meta-nets through FL, we allow the whole FL network to collaborate in learning a customized personalization strategy for each client. Empirical results show that this framework improves on a range of standard hand-crafted personalization baselines in both label and feature shift situations. Royson Lee, Minyoung Kim 0001, Da Li 0001, Xinchi Qiu, Timothy M. Hospedales, Ferenc Huszar, Nicholas D. Lane |
NeurIPS | 5 |
| 2023 | Mixture of Normalizing Flows for European Option PricingabstractWe present a mixture of normalizing flows (MoNF) approach to European option pricing with guarantees that its estimations are free from static arbitrage. In contrast to many existing methods that meet economic rationality constraints (e.g., non-arbitrage) by introducing auxiliary losses, our solution meets those constraints exactly by design. To achieve this, we propose to build a model for risk neutral density using normalizing flows, which results in a pricing model, instead of modelling the option pricing function directly. First, we convert the constraints for direct pricing models to the constraints for models backed by risk neutral density estimation, then we design a specific NF architecture that meets these constraints. Furthermore, we find that employing a mixture of such normalizing flows improves the performance significantly, compared to using a deeper single NF. Finally, we present a mechanism to regularise the proposed model, and this regularisation can serve as a bridge between our method and any sample-based mathematical finance method. The evaluations on five option datasets show superiority of our method compared to mathematical finance solutions and some other neural networks based methods. The code is available at \url{https://github.com/qmfin/MoNF}. Yongxin Yang, Timothy M. Hospedales |
UAI | 2 |
| 2023 | Accelerating Self-Supervised Learning via Efficient Training StrategiesabstractRecently the focus of the computer vision community has shifted from expensive supervised learning towards self-supervised learning of visual representations. While the performance gap between supervised and self-supervised has been narrowing, the time for training self-supervised deep networks remains an order of magnitude larger than its supervised counterparts, which hinders progress, imposes carbon cost, and limits societal benefits to institutions with substantial resources. Motivated by these issues, this paper investigates reducing the training time of recent self-supervised methods by various model-agnostic strategies that have not been used for this problem. In particular, we study three strategies: an extendable cyclic learning rate schedule, a matching progressive augmentation magnitude and image resolutions schedule, and a hard positive mining strategy based on augmentation difficulty. We show that all three methods combined lead up to 2.7 times speed-up in the training time of several self-supervised methods while retaining comparable performance to the standard self-supervised learning setting. Mustafa Taha Koçyigit, Timothy M. Hospedales, Hakan Bilen |
WACV | 2 |
| 2023 | Deep Learning for Free-Hand Sketch: A SurveyabstractFree-hand sketches are highly illustrative, and have been widely used by humans to depict objects or stories from ancient times to the present. The recent prevalence of touchscreen devices has made sketch creation a much easier task than ever and consequently made sketch-oriented applications increasingly popular. The progress of deep learning has immensely benefited free-hand sketch research and applications. This paper presents a comprehensive survey of the deep learning techniques oriented at free-hand sketch data, and the applications that they enable. The main contents of this survey include: (i) A discussion of the intrinsic traits and unique challenges of free-hand sketch, to highlight the essential differences between sketch data and other data modalities, e.g., natural photos. (ii) A review of the developments of free-hand sketch research in the deep learning era, by surveying existing datasets, research topics, and the state-of-the-art methods through a detailed taxonomy and experimental evaluation. (iii) Promotion of future work via a discussion of bottlenecks, open problems, and potential research directions for the community. Peng Xu 0005, Timothy M. Hospedales, Qiyue Yin, Yi-Zhe Song, Tao Xiang 0002, Liang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Uncertainty-Aware Source-Free Domain Adaptive Semantic SegmentationabstractSource-Free Domain Adaptation (SFDA) is becoming topical to address the challenge of distribution shift between training and deployment data, while also relaxing the requirement of source data availability during target domain adaptation. In this paper, we focus on SFDA for semantic segmentation, in which pseudo labeling based target domain self-training is a common solution. However, pseudo labels generated by the source models are particularly unreliable on the target domain data due to the domain shift issue. Therefore, we propose to use Bayesian Neural Network (BNN) to improve the target self-training by better estimating and exploiting pseudo-label uncertainty. With the uncertainty estimation of BNNs, we introduce two novel self-training based components: Uncertainty-aware Online Teacher-Student Learning (UOTSL) and Uncertainty-aware FeatureMix (UFM). Extensive experiments on two popular benchmarks, GTA 5 → Cityscapes and SYNTHIA → Cityscapes, show the superiority of our proposed method with mIoU gains of 3.6% and 5.7% over the state-of-the-art respectively. Zhihe Lu, Da Li 0001, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
IEEE Trans. Image Process. | 5 |
| 2022 | Why Do Self-Supervised Models Transfer? On the Impact of Invariance on Downstream Tasks
Linus Ericsson, Henry Gouk, Timothy M. Hospedales |
BMVC | 3 |
| 2022 | Towards Unsupervised Sketch-based Image Retrieval
Conghui Hu, Yongxin Yang, Yunpeng Li 0004, Timothy M. Hospedales, Yi-Zhe Song |
BMVC | 4 |
| 2022 | Pushing the Limits of Simple Pipelines for Few-Shot Learning: External Data and Fine-Tuning Make a DifferenceabstractFew-shot learning (FSL) is an important and topical problem in computer vision that has motivated extensive research into numerous methods spanning from sophisticated metalearning methods to simple transfer learning baselines. We seek to push the limits of a simple-but-effective pipeline for real-worldfew-shot image classification in practice. To this end, we explore few-shot learning from the perspective of neural architecture, as well as a three stage pipeline of pre-training on external data, meta-training with labelled few-shot tasks, and task-specific fine-tuning on unseen tasks. We investigate questions such as: ① How pre-training on external data benefits FSL? ② How state of the art transformer architectures can be exploited? and ③ How to best exploit finetuning? Ultimately, we show that a simple transformer-based pipeline yields surprisingly good performance on standard benchmarks such as Mini-ImageNet, CIFAR-FS, CDFSL and Meta-Dataset. Our code is available at https://hushell.github.io/pmf. Shell Xu Hu, Da Li 0001, Jan Stühmer, Minyoung Kim 0001, Timothy M. Hospedales |
CVPR | 5 |
| 2022 | MetaAudio: A Few-Shot Audio Classification Benchmark
Calum Heggan, Sam Budgett, Timothy M. Hospedales, Mehrdad Yaghoobi |
ICANN (1) | 3 |
| 2022 | SketchODE: Learning neural sketch representation in continuous time
Ayan Das 0003, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song |
ICLR | 3 |
| 2022 | Visual Representation Learning over Latent Domains
Lucas Deecke, Timothy M. Hospedales, Hakan Bilen |
ICLR | 2 |
| 2022 | Online Hyperparameter Meta-Learning with Hypergradient Distillation
Haebeom Lee, Hayeon Lee, Jaewoong Shin, Eunho Yang, Timothy M. Hospedales, Sung Ju Hwang |
ICLR | 5 |
| 2022 | Loss Function Learning for Domain Generalization by Implicit GradientabstractGeneralising robustly to distribution shift is a major challenge that is pervasive across most real-world applications of machine learning. A recent study highlighted that many advanced algorithms proposed to tackle such domain generalisation (DG) fail to outperform a properly tuned empirical risk minimisation (ERM) baseline. We take a different approach, and explore the impact of the ERM loss function on out-of-domain generalisation. In particular, we introduce a novel meta-learning approach to loss function search based on implicit gradient. This enables us to discover a general purpose parametric loss function that provides a drop-in replacement for cross-entropy. Our loss can be used in standard training pipelines to efficiently train robust models using any neural architecture on new datasets. The results show that it clearly surpasses cross-entropy, enables simple ERM to outperform some more complicated prior DG methods, and provides state-of-the-art performance across a variety of DG benchmarks. Furthermore, unlike most existing DG approaches, our setup applies to the most practical setting of single-source domain generalisation, on which we show significant improvement. Boyan Gao, Henry Gouk, Yongxin Yang, Timothy M. Hospedales |
ICML | 4 |
| 2022 | Fisher SAM: Information Geometry and Sharpness Aware MinimisationabstractRecent sharpness-aware minimisation (SAM) is known to find flat minima which is beneficial for better generalisation with improved robustness. SAM essentially modifies the loss function by the maximum loss value within the small neighborhood around the current iterate. However, it uses the Euclidean ball to define the neighborhood, which can be less accurate since loss functions for neural networks are typically defined over probability distributions (e.g., class predictive probabilities), rendering the parameter space no more Euclidean. In this paper we consider the information geometry of the model parameter space when defining the neighborhood, namely replacing SAM’s Euclidean balls with ellipsoids induced by the Fisher information. Our approach, dubbed Fisher SAM, defines more accurate neighborhood structures that conform to the intrinsic metric of the underlying statistical manifold. For instance, SAM may probe the worst-case loss value at either a too nearby or inappropriately distant point due to the ignorance of the parameter space geometry, which is avoided by our Fisher SAM. Another recent Adaptive SAM approach that stretches/shrinks the Euclidean ball in accordance with the scales of the parameter magnitudes, might be dangerous, potentially destroying the neighborhood structure even severely. We demonstrate the improved performance of the proposed Fisher SAM on several benchmark datasets/tasks. Minyoung Kim 0001, Da Li 0001, Shell Xu Hu, Timothy M. Hospedales |
ICML | 4 |
| 2022 | Meta-Learning in Neural Networks: A SurveyabstractThe field of meta-learning, or learning-to-learn, has seen a dramatic rise in interest in recent years. Contrary to conventional approaches to AI where tasks are solved from scratch using a fixed learning algorithm, meta-learning aims to improve the learning algorithm itself, given the experience of multiple learning episodes. This paradigm provides an opportunity to tackle many conventional challenges of deep learning, including data and computation bottlenecks, as well as generalization. This survey describes the contemporary meta-learning landscape. We first discuss definitions of meta-learning and position it with respect to related fields, such as transfer learning and hyperparameter optimization. We then propose a new taxonomy that provides a more comprehensive breakdown of the space of meta-learning methods today. We survey promising applications and successes of meta-learning such as few-shot learning and reinforcement learning. Finally, we discuss outstanding challenges and promising areas for future research. Timothy M. Hospedales, Antreas Antoniou, Paul Micaelli, Amos J. Storkey |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Neural-Symbolic Integration: A Compositional PerspectiveabstractDespite significant progress in the development of neural-symbolic frameworks, the question of how to integrate a neural and a symbolic system in a compositional manner remains open. Our work seeks to fill this gap by treating these two systems as black boxes to be integrated as modules into a single architecture, without making assumptions on their internal structure and semantics. Instead, we expect only that each module exposes certain methods for accessing the functions that the module implements: the symbolic module exposes a deduction method for computing the function's output on a given input, and an abduction method for computing the function's inputs for a given output; the neural module exposes a deduction method for computing the function's output on a given input, and an induction method for updating the function given input-output training instances. We are, then, able to show that a symbolic module --- with any choice for syntax and semantics, as long as the deduction and abduction methods are exposed --- can be cleanly integrated with a neural module, and facilitate the latter's efficient training, achieving empirical performance that exceeds that of previous work. Efthymia Tsamoura, Timothy M. Hospedales, Loizos Michael |
AAAI | 2 |
| 2021 | Simple and Effective Stochastic Neural NetworksabstractStochastic neural networks (SNNs) are currently topical, with several paradigms being actively investigated including dropout, Bayesian neural networks, variational information bottleneck (VIB) and noise regularized learning. These neural network variants impact several major considerations, including generalization, network compression, robustness against adversarial attack and label noise, and model calibration. However, many existing networks are complicated and expensive to train, and/or only address one or two of these practical considerations. In this paper we propose a simple and effective stochastic neural network (SE-SNN) architecture for discriminative learning by directly modeling activation uncertainty and encouraging high activation variability. Compared to existing SNNs, our SE-SNN is simpler to implement and faster to train, and produces state of the art results on network compression by pruning, adversarial defense, learning with label noise, and model calibration. Yongxin Yang, Da Li 0001, Timothy M. Hospedales, Tao Xiang 0002 |
AAAI | 4 |
| 2021 | Robust Domain Randomised Reinforcement Learning through Peer-to-Peer DistillationabstractIn reinforcement learning, domain randomisation is a popular technique for learning general policies that are robust to new environments and domain-shifts at deployment. However, naively aggregating information from randomised domains may lead to high variances in gradient estimation and sub-optimal policies. To address this issue, we present a peer-to-peer online distillation strategy for reinforcement learning termed P2PDRL, where multiple learning agents are each assigned to a different environment, and then exchange knowledge through mutual regularisation based on Kullback–Leibler divergence. Our experiments on continuous control tasks show that P2PDRL enables robust learning across a wider randomisation distribution than baselines, and more robust generalisation performance to new environments at testing. Chenyang Zhao 0007, Timothy M. Hospedales |
ACML | 2 |
| 2021 | Defensive Tensorization
Adrian Bulat, Jean Kossaifi, Sourav Bhattacharya, Yannis Panagakis, Timothy M. Hospedales, Georgios Tzimiropoulos, Nicholas D. Lane, Maja Pantic |
BMVC | 5 |
| 2021 | Tensor Composition Net for Visual Relationship Prediction
Yuting Qiang, Yongxin Yang, Yanwen Guo 0001, Timothy M. Hospedales |
BMVC | 5 |
| 2021 | Cloud2Curve: Generation and Vectorization of Parametric SketchesabstractAnalysis of human sketches in deep learning has advanced immensely through the use of waypoint-sequences rather than raster-graphic representations. We further aim to model sketches as a sequence of low-dimensional parametric curves. To this end, we propose an inverse graphics framework capable of approximating a raster or waypoint based stroke encoded as a point-cloud with a variable-degree Bézier curve. Building on this module, we present Cloud2Curve, a generative model for scalable high-resolution vector sketches that can be trained end-to-end using point-cloud data alone. As a consequence, our model is also capable of deterministic vectorization which can map novel raster or waypoint based sketches to their corresponding high-resolution scalable Bézier equivalent. We evaluate the generation and vectorization capabilities of our model on Quick, Draw! and K-MNIST datasets. Ayan Das 0003, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song |
CVPR | 3 |
| 2021 | Vectorization and Rasterization: Self-Supervised Learning for Sketch and HandwritingabstractSelf-supervised learning has gained prominence due to its efficacy at learning powerful representations from unlabelled data that achieve excellent performance on many challenging downstream tasks. However, supervision-free pre-text tasks are challenging to design and usually modality specific. Although there is a rich literature of self-supervised methods for either spatial (such as images) or temporal data (sound or text) modalities, a common pretext task that benefits both modalities is largely missing. In this paper, we are interested in defining a self-supervised pre-text task for sketches and handwriting data. This data is uniquely characterised by its existence in dual modalities of rasterized images and vector coordinate sequences. We address and exploit this dual representation by proposing two novel cross-modal translation pre-text tasks for self-supervised feature learning: Vectorization and Rasterization. Vectorization learns to map image space to vector coordinates and rasterization maps vector coordinates to image space. We show that our learned encoder modules benefit both raster-based and vector-based downstream approaches to analysing hand-drawn data. Empirical evidence shows that our novel pre-text tasks surpass existing single and multi-modal self-supervision methods. Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song |
CVPR | 4 |
| 2021 | How Well Do Self-Supervised Models Transfer?abstractSelf-supervised visual representation learning has seen huge progress recently, but no large scale evaluation has compared the many models now available. We evaluate the transfer performance of 13 top self-supervised models on 40 downstream tasks, including many-shot and few-shot recognition, object detection, and dense prediction. We compare their performance to a supervised baseline and show that on most tasks the best self-supervised models outperform supervision, confirming the recently observed trend in the literature. We find ImageNet Top-1 accuracy to be highly correlated with transfer to many-shot recognition, but increasingly less so for few-shot, object detection and dense prediction. No single self-supervised method dominates overall, suggesting that universal pre-training is still unsolved. Our analysis of features suggests that top self-supervised learners fail to preserve colour information as well as supervised alternatives, but tend to induce better classifier calibration, and less attentive overfitting than supervised learners. Linus Ericsson, Henry Gouk, Timothy M. Hospedales |
CVPR | 3 |
| 2021 | NewtonianVAE: Proportional Control and Goal Identification From Pixels via Physical Latent SpacesabstractLearning low-dimensional latent state space dynamics models has proven powerful for enabling vision-based planning and learning for control. We introduce a latent dynamics learning framework that is uniquely designed to induce proportional controlability in the latent space, thus enabling the use of simple and well-known PID controllers. We show that our learned dynamics model enables proportional control from pixels, dramatically simplifies and accelerates behavioural cloning of vision-based controllers, and provides interpretable goal discovery when applied to imitation learning of switching controllers from demonstration. Notably, such proportional controlability also allows for robust path following from visual demonstrations using Dynamic Movement Primitives in the learned latent space. Miguel Jaques, Michael Burke, Timothy M. Hospedales |
CVPR | 3 |
| 2021 | Searching for Robustness: Loss Learning for Noisy Classification TasksabstractWe present a "learning to learn" approach for discovering white-box classification loss functions that are robust to label noise in the training data. We parameterise a flexible family of loss functions using Taylor polynomials, and apply evolutionary strategies to search for noise-robust losses in this space. To learn re-usable loss functions that can apply to new tasks, our fitness function scores their performance in aggregate across a range of training datasets and architectures. The resulting white-box loss provides a simple and fast "plug-and-play" module that enables effective label-noise-robust learning in diverse downstream tasks, without requiring a special training procedure or network architecture. The efficacy of our loss is demonstrated on a variety of datasets with both synthetic and real label noise, where we compare favourably to prior work. Boyan Gao, Henry Gouk, Timothy M. Hospedales |
ICCV | 3 |
| 2021 | A Simple Feature Augmentation for Domain GeneralizationabstractThe topical domain generalization (DG) problem asks trained models to perform well on an unseen target domain with different data statistics from the source training domains. In computer vision, data augmentation has proven one of the most effective ways of better exploiting the source data to improve domain generalization. However, existing approaches primarily rely on image-space data augmentation, which requires careful augmentation design, and provides limited diversity of augmented data. We argue that feature augmentation is a more promising direction for DG. We find that an extremely simple technique of perturbing the feature embedding with Gaussian noise during training leads to a classifier with domain-generalization performance comparable to existing state of the art. To model more meaningful statistics reflective of cross-domain variability, we further estimate the full class-conditional feature covariance matrix iteratively during training. Subsequent joint stochastic feature augmentation provides an effective domain randomization method, perturbing features in the directions of intra-class/cross-domain variability. We verify our proposed method on three standard domain generalization benchmarks, Digit-DG, VLCS and PACS, and show it is outperforming or comparable to the state of the art in all setups, together with experimental analysis to illustrate how our method works towards training a robust generalisable model. Da Li 0001, Wei Li 0132, Shaogang Gong, Yanwei Fu 0001, Timothy M. Hospedales |
ICCV | 6 |
| 2021 | Shallow Bayesian Meta Learning for Real-World Few-Shot RecognitionabstractMany state-of-the-art few-shot learners focus on developing effective training procedures for feature representations, before using simple (e.g., nearest centroid) classifiers. We take an approach that is agnostic to the features used, and focus exclusively on meta-learning the final classifier layer. Specifically, we introduce MetaQDA, a Bayesian meta-learning generalisation of the classic quadratic discriminant analysis. This approach has several benefits of interest to practitioners: meta-learning is fast and memory efficient, without the need to fine-tune features. It is agnostic to the off-the-shelf features chosen, and thus will continue to benefit from future advances in feature representations. Empirically, it leads to excellent performance in cross-domain few-shot learning, class-incremental few-shot learning, and crucially for real-world applications, the Bayesian formulation leads to state-of-the-art uncertainty calibration in predictions. Debin Meng, Henry Gouk, Timothy M. Hospedales |
ICCV | 4 |
| 2021 | Interpreting Knowledge Graph Relation Representation from Word Embeddings
Carl Allen, Ivana Balazevic, Timothy M. Hospedales |
ICLR | 3 |
| 2021 | Distance-Based Regularisation of Deep Networks for Fine-Tuning
Henry Gouk, Timothy M. Hospedales, Massimiliano Pontil |
ICLR | 2 |
| 2021 | Weight-covariance alignment for adversarially robust neural networksabstractStochastic Neural Networks (SNNs) that inject noise into their hidden layers have recently been shown to achieve strong robustness against adversarial attacks. However, existing SNNs are usually heuristically motivated, and often rely on adversarial training, which is computationally costly. We propose a new SNN that achieves state-of-the-art performance without relying on adversarial training, and enjoys solid theoretical justification. Specifically, while existing SNNs inject learned or hand-tuned isotropic noise, our SNN learns an anisotropic noise distribution to optimize a learning-theoretic bound on adversarial robustness. We evaluate our method on a number of popular benchmarks, show that it can be applied to different architectures, and that it provides robustness to a variety of white-box and black-box attacks, while being simple and fast to train compared to existing alternatives. Panagiotis Eustratiadis, Henry Gouk, Da Li 0001, Timothy M. Hospedales |
ICML | 4 |
| 2021 | EvoGrad: Efficient Gradient-Based Meta-Learning and Hyperparameter OptimizationabstractGradient-based meta-learning and hyperparameter optimization have seen significant progress recently, enabling practical end-to-end training of neural networks together with many hyperparameters. Nevertheless, existing approaches are relatively expensive as they need to compute second-order derivatives and store a longer computational graph. This cost prevents scaling them to larger network architectures. We present EvoGrad, a new approach to meta-learning that draws upon evolutionary techniques to more efficiently compute hypergradients. EvoGrad estimates hypergradient with respect to hyperparameters without calculating second-order gradients, or storing a longer computational graph, leading to significant improvements in efficiency. We evaluate EvoGrad on three substantial recent meta-learning applications, namely cross-domain few-shot learning with feature-wise transformations, noisy label learning with Meta-Weight-Net and low-resource cross-lingual learning with meta representation transformation. The results show that EvoGrad significantly improves efficiency and enables scaling meta-learning to bigger architectures such as from ResNet10 to ResNet34. Ondrej Bohdal, Yongxin Yang, Timothy M. Hospedales |
NeurIPS | 3 |
| 2021 | Fine-Grained Instance-Level Sketch-Based Image Retrieval
Qian Yu 0002, Jifei Song, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
Int. J. Comput. Vis. | 5 |
| 2021 | On Learning Semantic Representations for Large-Scale Abstract SketchesabstractIn this paper, we focus on learning semantic representations for large-scale highly abstract sketches that were produced by the practical sketch-based application rather than the excessively well dawn sketches obtained by crowd-sourcing. We propose a dual-branch CNN-RNN network architecture to represent sketches, which simultaneously encodes both the static and temporal patterns of sketch strokes. Based on this architecture, we further explore learning the sketch-oriented semantic representations in two practical settings, i.e., hashing retrieval and zero-shot recognition on million-scale highly abstract sketches produced by practical online interactions. Specifically, we use our dual-branch architecture as a universal representation framework to design two sketch-specific deep models: (i) We propose a deep hashing model for sketch retrieval, where a novel hashing loss is specifically designed to further accommodate both the abstract and messy traits of sketches. (ii) We propose a deep embedding model for sketch zero-shot recognition, via collecting a large-scale edge-map dataset and proposing to extract a set of semantic vectors from edge-maps as the semantic knowledge for sketch zero-shot domain alignment. Both deep models are evaluated by comprehensive experiments on million-scale abstract sketches produced by a global online game QuickDraw and outperform state-of-the-art competitors. Peng Xu 0005, Yongye Huang, Tongtong Yuan, Tao Xiang 0002, Timothy M. Hospedales, Yi-Zhe Song, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Fine-Grained Instance-Level Sketch-Based Video RetrievalabstractExisting sketch-analysis work studies sketches depicting static objects or scenes. In this work, we propose a novel cross-modal retrieval problem of fine-grained instance-level sketch-based video retrieval (FG-SBVR), where a sketch sequence is used as a query to retrieve a specific target video instance. Compared with sketch-based still image retrieval, and coarse-grained category-level video retrieval, this is more challenging as both visual appearance and motion need to be simultaneously matched at a fine-grained level. We contribute the first FG-SBVR dataset with rich annotations. We then introduce a novel multi-stream multi-modality deep network to perform FG-SBVR under both strong and weakly supervised settings. The key component of the network is a relation module, designed to prevent model overfitting given scarce training data. We show that this model significantly outperforms a number of existing state-of-the-art models designed for video analysis. Peng Xu 0005, Kun Liu 0016, Tao Xiang 0002, Timothy M. Hospedales, Zhanyu Ma, Jun Guo 0002, Yi-Zhe Song |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Toward Fine-Grained Sketch-Based 3D Shape RetrievalabstractIn this paper we study, for the first time, the problem of fine-grained sketch-based 3D shape retrieval. We advocate the use of sketches as a fine-grained input modality to retrieve 3D shapes at instance-level - e.g., given a sketch of a chair, we set out to retrieve a specific chair from a gallery of all chairs. Fine-grained sketch-based 3D shape retrieval (FG-SBSR) has not been possible till now due to a lack of datasets that exhibit one-to-one sketch-3D correspondences. The first key contribution of this paper is two new datasets, consisting a total of 4,680 sketch-3D pairings from two object categories. Even with the datasets, FG-SBSR is still highly challenging because (i) the inherent domain gap between 2D sketch and 3D shape is large, and (ii) retrieval needs to be conducted at the instance level instead of the coarse category level matching as in traditional SBSR. Thus, the second contribution of the paper is the first cross-modal deep embedding model for FG-SBSR, which specifically tackles the unique challenges presented by this new problem. Core to the deep embedding model is a novel cross-modal view attention module which automatically computes the optimal combination of 2D projections of a 3D shape given a query sketch. Anran Qi, Yulia Gryaditskaya, Jifei Song, Yongxin Yang, Yonggang Qi, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song |
IEEE Trans. Image Process. | 6 |
| 2020 | Index Tracking with Cardinality Constraints: A Stochastic Neural Networks Approach
Yu Zheng 0032, Bowei Chen 0001, Timothy M. Hospedales, Yongxin Yang |
AAAI | 3 |
| 2020 | Deep Domain-Adversarial Image Generation for Domain GeneralisationabstractMachine learning models typically suffer from the domain shift problem when trained on a source dataset and evaluated on a target dataset of different distribution. To overcome this problem, domain generalisation (DG) methods aim to leverage data from multiple source domains so that a trained model can generalise to unseen domains. In this paper, we propose a novel DG approach based on Deep Domain-Adversarial Image Generation (DDAIG). Specifically, DDAIG consists of three components, namely a label classifier, a domain classifier and a domain transformation network (DoTNet). The goal for DoTNet is to map the source training data to unseen domains. This is achieved by having a learning objective formulated to ensure that the generated data can be correctly classified by the label classifier while fooling the domain classifier. By augmenting the source training data with the generated unseen domain data, we can make the label classifier more robust to unknown domain changes. Extensive experiments on four DG datasets demonstrate the effectiveness of our approach. Kaiyang Zhou, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002 |
AAAI | 3 |
| 2020 | ALBA: Reinforcement Learning for Video Object Segmentation
Shreyank N. Gowda, Panagiotis Eustratiadis, Timothy M. Hospedales, Laura Sevilla-Lara |
BMVC | 3 |
| 2020 | Sketch Less for More: On-the-Fly Fine-Grained Sketch-Based Image RetrievalabstractFine-grained sketch-based image retrieval (FG-SBIR) addresses the problem of retrieving a particular photo instance given a user's query sketch. Its widespread applicability is however hindered by the fact that drawing a sketch takes time, and most people struggle to draw a complete and faithful sketch. In this paper, we reformulate the conventional FG-SBIR framework to tackle these challenges, with the ultimate goal of retrieving the target photo with the least number of strokes possible. We further propose an on-the-fly design that starts retrieving as soon as the user starts drawing. To accomplish this, we devise a reinforcement learning based cross-modal retrieval framework that directly optimizes rank of the ground-truth photo over a complete sketch drawing episode. Additionally, we introduce a novel reward scheme that circumvents the problems related to irrelevant sketch strokes, and thus provides us with a more consistent rank list during the retrieval. We achieve superior early-retrieval efficiency over state-of-the-art methods and alternative baselines on two publicly available fine-grained sketch retrieval datasets. Ayan Kumar Bhunia, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song |
CVPR | 3 |
| 2020 | Factorized Higher-Order CNNs With an Application to Spatio-Temporal Emotion EstimationabstractTraining deep neural networks with spatio-temporal (i.e., 3D) or multidimensional convolutions of higher-order is computationally challenging due to millions of unknown parameters across dozens of layers. To alleviate this, one approach is to apply low-rank tensor decompositions to convolution kernels in order to compress the network and reduce its number of parameters. Alternatively, new convolutional blocks, such as MobileNet, can be directly designed for efficiency. In this paper, we unify these two approaches by proposing a tensor factorization framework for efficient multidimensional (separable) convolutions of higher-order. Interestingly, the proposed framework enables a novel higher-order transduction, allowing to train a network on a given domain (e.g., 2D images or N-dimensional data in general) and using transduction to generalize to higher-order data such as videos (or (N+K)--dimensional data in general), capturing for instance temporal dynamics while preserving the learnt spatial information. We apply the proposed methodology, coined CP-Higher-Order Convolution (HO-CPConv), to spatio-temporal facial emotion analysis. Most existing facial affect models focus on static imagery and discard all temporal information. This is due to the above-mentioned burden of training 3D convolutional nets and the lack of large bodies of video data annotated by experts. We address both issues with our proposed framework. Initial training is first done on static imagery before using transduction to generalize to the temporal domain. We demonstrate superior performance on three challenging large scale affect estimation datasets, AffectNet, SEWA, and AFEW-VA. Jean Kossaifi, Antoine Toisoul, Adrian Bulat, Yannis Panagakis, Timothy M. Hospedales, Maja Pantic |
CVPR | 5 |
| 2020 | Solving Mixed-Modal Jigsaw Puzzle for Fine-Grained Sketch-Based Image RetrievalabstractImageNet pre-training has long been considered crucial by the fine-grained sketch-based image retrieval (FG-SBIR) community due to the lack of large sketch-photo paired datasets for FG-SBIR training. In this paper, we propose a self-supervised alternative for representation pre-training. Specifically, we consider the jigsaw puzzle game of recomposing images from shuffled parts. We identify two key facets of jigsaw task design that are required for effective FG-SBIR pre-training. The first is formulating the puzzle in a mixed-modality fashion. Second we show that framing the optimisation as permutation matrix inference via Sinkhorn iterations is more effective than the common classifier formulation of Jigsaw self-supervision. Experiments show that this self-supervised pre-training strategy significantly outperforms the standard ImageNet-based pipeline across all four product-level FG-SBIR benchmarks. Interestingly it also leads to improved cross-category generalisation across both pre-train/fine-tune and fine-tune/testing stages. Kaiyue Pang, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song |
CVPR | 3 |
| 2020 | Incremental Few-Shot Object DetectionabstractExisting object detection methods typically rely on the availability of abundant labelled training samples per class and offline model training in a batch mode. These requirements substantially limit their scalability to open-ended accommodation of novel classes with limited labelled training data, both in terms of model accuracy and training efficiency during deployment. We present the first study aiming to go beyond these limitations by considering the Incremental Few-Shot Detection (iFSD) problem setting, where new classes must be registered incrementally (without revisiting base classes) and with few examples. To this end we propose OpeN-ended Centre nEt (ONCE), a detector designed for incrementally learning to detect novel class objects with few examples. This is achieved by an elegant adaptation of the efficient CentreNet detector to the few-shot learning scenario, and meta-learning a class-wise code generator model for registering novel classes. ONCE fully respects the incremental learning paradigm, with novel class registration requiring only a single forward pass of few-shot training samples, and no access to base classes - thus making it suitable for deployment on embedded devices, etc. Extensive experiments conducted on both the standard object detection (COCO, PASCAL VOC) and fashion landmark detection (DeepFashion2) tasks show the feasibility of iFSD for the first time, opening an interesting and very important line of research. Juan-Manuel Pérez-Rúa, Xiatian Zhu, Timothy M. Hospedales, Tao Xiang 0002 |
CVPR | 3 |
| 2020 | BézierSketch: A Generative Model for Scalable Vector Sketches
Ayan Das 0003, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song |
ECCV (26) | 3 |
| 2020 | Online Meta-learning for Multi-source and Semi-supervised Domain Adaptation
Da Li 0001, Timothy M. Hospedales |
ECCV (16) | 2 |
| 2020 | Differentiable Automatic Data Augmentation
Yonggang Li 0001, Guosheng Hu, Yongtao Wang, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang |
ECCV (22) | 4 |
| 2020 | Learning to Generate Novel Domains for Domain Generalization
Kaiyang Zhou, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002 |
ECCV (16) | 3 |
| 2020 | Deep Clustering for Domain AdaptationabstractWe address the heterogeneous domain adaptation task: adapting a classifier trained on data from one domain to operate on another domain that also has a different label space. We consider two settings that both exhibit label scarcity of some form—one where only unlabelled data is available, and another where a small volume of labelled data is available in addition to the unlabelled data. Our method is based on two specialisations of a recently proposed approach for deep clustering. It is shown that our approach noticeably outperforms other methods based on deep clustering in both the fully unsupervised and the semi-supervised settings. Boyan Gao, Yongxin Yang, Henry Gouk, Timothy M. Hospedales |
ICASSP | 4 |
| 2020 | Deep Clusteringwith Concrete K-MeansabstractWe address the problem of simultaneously learning a k-means clustering and deep feature representation from unlabelled data, which is of interest due to the potential for deep k-means to outperform traditional two-step feature extraction and shallow clustering strategies. We achieve this by developing a gradient estimator for the non-differentiable k-means objective via the Gumbel-Softmax reparameterisation trick. In contrast to previous attempts at deep clustering, our concrete k-means model can be optimised with respect to the canonical k-means objective and is easily trained end-to-end without resorting to time consuming alternating optimisation techniques. We demonstrate the efficacy of our method on standard clustering benchmarks. Boyan Gao, Yongxin Yang, Henry Gouk, Timothy M. Hospedales |
ICASSP | 4 |
| 2020 | Diversity and Sparsity: A New Perspective on Index TrackingabstractWe address the problem of partial index tracking, replicating a benchmark index using a small number of assets. Accurate tracking with a sparse portfolio is extensively studied as a classic finance problem. However in practice, a tracking portfolio must also be diverse in order to minimise risk-a requirement which has only been dealt with by ad-hoc methods before. We introduce the first index tracking method that explicitly optimises both diversity and sparsity in a single joint framework. Diversity is realised by a regulariser based on pairwise similarity of assets, and we demonstrate that learning similarity from data can outperform some existing heuristics. Finally, we show that the way we model diversity leads to an easy solution for sparsity, allowing both constraints to be optimised easily and efficiently. we run out-of-sample back-testing for a long interval of 15 years (2003-2018), and the results demonstrate the superiority of the proposed algorithm. Yu Zheng 0032, Timothy M. Hospedales, Yongxin Yang |
ICASSP | 2 |
| 2020 | Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from Video
Miguel Jaques, Michael Burke, Timothy M. Hospedales |
ICLR | 3 |
| 2020 | Full-Scale Continuous Synthetic Sonar Data Generation with Markov Conditional Generative Adversarial Networks*abstractDeployment and operation of autonomous underwater vehicles is expensive and time-consuming. High-quality realistic sonar data simulation could be of benefit to multiple applications, including training of human operators for post-mission analysis, as well as tuning and validation of autonomous target recognition (ATR) systems for underwater vehicles. Producing realistic synthetic sonar imagery is a challenging problem as the model has to account for specific artefacts of real acoustic sensors, vehicle attitude, and a variety of environmental factors. We propose a novel method for generating realistic-looking sonar side-scans of full-length missions, called Markov Conditional pix2pix (MC-pix2pix). Quantitative assessment results confirm that the quality of the produced data is almost indistinguishable from real. Furthermore, we show that bootstrapping ATR systems with MC-pix2pix data can improve the performance. Synthetic data is generated 18 times faster than real acquisition speed, with full user control over the topography of the generated data. Marija Jegorova, Antti Ilari Karjalainen, Jose Vazquez, Timothy M. Hospedales |
ICRA | 4 |
| 2020 | RelationNet2: Deep Comparison Network for Few-Shot LearningabstractFew-shot deep learning is a topical challenge area for scaling visual recognition to open ended growth of unseen new classes with limited labeled examples. A promising approach is based on metric learning, which trains a deep embedding to support image similarity matching. Our insight is that effective general purpose matching requires non-linear comparison of features at multiple abstraction levels. We thus propose a new deep comparison network comprised of embedding and relation modules that learn multiple non-linear distance metrics based on different levels of features simultaneously. Furthermore, to reduce over-fitting and enable the use of deeper embeddings, we represent images as distributions rather than vectors via learning parameterized Gaussian noise regularization. The resulting network achieves excellent performance on both miniImageNet and tieredImageNet. Yuting Qiang, Flood Sung, Yongxin Yang, Timothy M. Hospedales |
IJCNN | 5 |
| 2020 | Adversarial Generation of Informative Trajectories for Dynamics System IdentificationabstractDnamic System Identification approaches usually heavily rely on evolutionary and gradient-based optimisation techniques to produce optimal excitation trajectories for determining the physical parameters of robot platforms. Current optimisation techniques tend to generate single trajectories. This is expensive, and intractable for longer trajectories, thus limiting their efficacy for system identification. We propose to tackle this issue by using multiple shorter cyclic trajectories, which can be generated in parallel, and subsequently combined together to achieve the same effect as a longer trajectory. Crucially, we show how to scale this approach even further by increasing the generation speed and quality of the dataset through the use of generative adversarial network (GAN) based architectures to produce large databases of valid and diverse excitation trajectories. To the best of our knowledge, this is the first robotics work to explore system identification with multiple cyclic trajectories and to develop GAN-based techniques for scaleably producing excitation trajectories that are diverse in both control parameter and inertial parameter spaces. We show that our approach dramatically accelerates trajectory optimisation, while simultaneously providing more accurate system identification than the conventional approach. Marija Jegorova, Joshua Smith 0002, Michael N. Mistry, Timothy M. Hospedales |
IROS | 4 |
| 2020 | Online Meta-Critic Learning for Off-Policy Actor-Critic MethodsabstractOff-Policy Actor-Critic (OffP-AC) methods have proven successful in a variety of continuous control tasks. Normally, the critic's action-value function is updated using temporal-difference, and the critic in turn provides a loss for the actor that trains it to take actions with higher expected return. In this paper, we introduce a flexible and augmented meta-critic that observes the learning process and meta-learns an additional loss for the actor that accelerates and improves actor-critic learning. Compared to existing meta-learning algorithms, meta-critic is rapidly learned online for a single task, rather than slowly over a family of tasks. Crucially, our meta-critic is designed for off-policy based learners, which currently provide state-of-the-art reinforcement learning sample efficiency. We demonstrate that online meta-critic learning benefits to a variety of continuous control tasks when combined with contemporary OffP-AC methods DDPG, TD3 and SAC. Wei Zhou 0107, Yiying Li, Yongxin Yang, Huaimin Wang 0001, Timothy M. Hospedales |
NeurIPS | 5 |
| 2020 | Editorial: Special Issue on Machine Vision with Deep Learning
Ling Shao 0001, Hubert P. H. Shum, Timothy M. Hospedales |
Int. J. Comput. Vis. | 3 |
| 2020 | Inverse Visual Question Answering: A New Benchmark and VQA Diagnosis ToolabstractIn recent years, visual question answering (VQA) has become topical. The premise of VQA's significance as a benchmark in AI, is that both the image and textual question need to be well understood and mutually grounded in order to infer the correct answer. However, current VQA models perhaps 'understand' less than initially hoped, and instead master the easier task of exploiting cues given away in the question and biases in the answer distribution [1]. In this paper we propose the inverse problem of VQA (iVQA). The iVQA task is to generate a question that corresponds to a given image and answer pair. We propose a variational iVQA model that can generate diverse, grammatically correct and content correlated questions that match the given answer. Based on this model, we show that iVQA is an interesting benchmark for visuo-linguistic understanding, and a more challenging alternative to VQA because an iVQA model needs to understand the image better to be successful. As a second contribution, we show how to use iVQA in a novel reinforcement learning framework to diagnose any existing VQA model by way of exposing its belief set: the set of question-answer pairs that the VQA model would predict true for a given image. This provides a completely new window into what VQA models 'believe' about images. We show that existing VQA models have more erroneous beliefs than previously thought, revealing their intrinsic weaknesses. Suggestions are then made on how to address these weaknesses going forward. Feng Liu 0036, Tao Xiang 0002, Timothy M. Hospedales, Wankou Yang, Changyin Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Sketch-a-Segmenter: Sketch-Based Photo Segmenter GenerationabstractGiven pixel-level annotated data, traditional photo segmentation techniques have achieved promising results. However, these photo segmentation models can only identify objects in categories for which data annotation and training have been carried out. This limitation has inspired recent work on few-shot and zero-shot learning for image segmentation. In this paper, we show the value of sketch for photo segmentation, in particular as a transferable representation to describe a concept to be segmented. We show, for the first time, that it is possible to generate a photo-segmentation model of a novel category using just a single sketch and furthermore exploit the unique fine-grained characteristics of sketch to produce more detailed segmentation. More specifically, we propose a sketch-based photo segmentation method that takes sketch as input and synthesizes the weights required for a neural network to segment the corresponding region of a given photo. Our framework can be applied at both the category-level and the instance-level, and fine-grained input sketches provide more accurate segmentation in the latter. This framework generalizes across categories via sketch and thus provides an alternative to zero-shot learning when segmenting a photo from a category without annotated training data. To investigate the instance-level relationship across sketch and photo, we create the SketchySeg dataset which contains segmentation annotations for photos corresponding to paired sketches in the Sketchy Dataset. Conghui Hu, Da Li 0001, Yongxin Yang, Timothy M. Hospedales, Yi-Zhe Song |
IEEE Trans. Image Process. | 4 |
| 2020 | Deep Ranking for Image Zero-Shot Multi-Label ClassificationabstractDuring the past decade, both multi-label learning and zero-shot learning have attracted huge research attention, and significant progress has been made. Multi-label learning algorithms aim to predict multiple labels given one instance, while most existing zero-shot learning approaches target at predicting a single testing label for each unseen class via transferring knowledge from auxiliary seen classes to target unseen classes. However, relatively less effort has been made on predicting multiple labels in the zero-shot setting, which is nevertheless a quite challenging task. In this work, we investigate and formalize a flexible framework consisting of two components, i.e., visual-semantic embedding and zero-shot multi-label prediction. First, we present a deep regression model to project the visual features into the semantic space, which explicitly exploits the correlations in the intermediate semantic layer of word vectors and makes label prediction possible. Then, we formulate the label prediction problem as a pairwise one and employ Ranking SVM to seek the unique multi-label correlations in the embedding space. Furthermore, we provide a transductive multi-label zeroshot prediction approach that exploits the testing data manifold structure. We demonstrate the effectiveness of the proposed approach on three popular multi-label datasets with state-of-theart performance obtained on both conventional and generalized ZSL settings. Zhong Ji, Biying Cui, Yu-Gang Jiang 0001, Tao Xiang 0002, Timothy M. Hospedales, Yanwei Fu 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Pixelor: a competitive sketching AI agent. so you think you can sketch?abstractWe present the first competitive drawing agent Pixelor that exhibits humanlevel performance at a Pictionary-like sketching game, where the participant whose sketch is recognized first is a winner. Our AI agent can autonomously sketch a given visual concept, and achieve a recognizable rendition as quickly or faster than a human competitor. The key to victory for the agent's goal is to learn the optimal stroke sequencing strategies that generate the most recognizable and distinguishable strokes first. Training Pixelor is done in two steps. First, we infer the stroke order that maximizes early recognizability of human training sketches. Second, this order is used to supervise the training of a sequence-to-sequence stroke generator. Our key technical contributions are a tractable search of the exponential space of orderings using neural sorting; and an improved Seq2Seq Wasserstein (S2S-WAE) generator that uses an optimal-transport loss to accommodate the multi-modal nature of the optimal stroke distribution. Our analysis shows that Pixelor is better than the human players of the Quick, Draw! game, under both AI and human judging of early recognition. To analyze the impact of human competitors' strategies, we conducted a further human study with participants being given unlimited thinking time and training in early recognizability by feedback from an AI judge. The study shows that humans do gradually improve their strategies with training, but overall Pixelor still matches human performance. The code and the dataset are available at http://sketchx.ai/pixelor. Ayan Kumar Bhunia, Ayan Das 0003, Umar Riaz Muhammad, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002, Yulia Gryaditskaya, Yi-Zhe Song |
ACM Trans. Graph. | 5 |
| 2019 | Disjoint Label Space Transfer Learning with Common Factorised SpaceabstractIn this paper, a unified approach is presented to transfer learning that addresses several source and target domain labelspace and annotation assumptions with a single model. It is particularly effective in handling a challenging case, where source and target label-spaces are disjoint, and outperforms alternatives in both unsupervised and semi-supervised settings. The key ingredient is a common representation termed Common Factorised Space. It is shared between source and target domains, and trained with an unsupervised factorisation loss and a graph-based loss. With a wide range of experiments, we demonstrate the flexibility, relevance and efficacy of our method, both in the challenging cases with disjoint label spaces, and in the more conventional cases such as unsupervised domain adaptation, where the source and target domains share the same label-sets. Xiaobin Chang, Yongxin Yang, Tao Xiang 0002, Timothy M. Hospedales |
AAAI | 4 |
| 2019 | Frustratingly Easy Person Re-Identification: Generalizing Person Re-ID in Practice
Jieru Jia, Qiuqi Ruan, Timothy M. Hospedales |
BMVC | 3 |
| 2019 | Generalising Fine-Grained Sketch-Based Image RetrievalabstractFine-grained sketch-based image retrieval (FG-SBIR) addresses matching specific photo instance using free-hand sketch as a query modality. Existing models aim to learn an embedding space in which sketch and photo can be directly compared. While successful, they require instance-level pairing within each coarse-grained category as annotated training data. Since the learned embedding space is domain-specific, these models do not generalise well across categories. This limits the practical applicability of FG-SBIR. In this paper, we identify cross-category generalisation for FG-SBIR as a domain generalisation problem, and propose the first solution. Our key contribution is a novel unsupervised learning approach to model a universal manifold of prototypical visual sketch traits. This manifold can then be used to paramaterise the learning of a sketch/photo representation. Model adaptation to novel categories then becomes automatic via embedding the novel sketch in the manifold and updating the representation and retrieval function accordingly. Experiments on the two largest FG-SBIR datasets, Sketchy and QMUL-Shoe-V2, demonstrate the efficacy of our approach in enabling cross-category generalisation of FG-SBIR. Kaiyue Pang, Ke Li 0004, Yongxin Yang, Honggang Zhang 0002, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song |
CVPR | 5 |
| 2019 | Generalizable Person Re-Identification by Domain-Invariant Mapping NetworkabstractWe aim to learn a domain generalizable person re-identification (ReID) model. When such a model is trained on a set of source domains (ReID datasets collected from different camera networks), it can be directly applied to any new unseen dataset for effective ReID without any model updating. Despite its practical value in real-world deployments, generalizable ReID has seldom been studied. In this work, a novel deep ReID model termed Domain-Invariant Mapping Network (DIMN) is proposed. DIMN is designed to learn a mapping between a person image and its identity classifier, i.e., it produces a classifier using a single shot. To make the model domain-invariant, we follow a meta-learning pipeline and sample a subset of source domain training tasks during each training episode. However, the model is significantly different from conventional meta-learning methods in that: (1) no model updating is required for the target domain, (2) different training tasks share a memory bank for maintaining both scalability and discrimination ability, and (3) it can be used to match an arbitrary number of identities in a target domain. Extensive experiments on a newly proposed large-scale ReID domain generalization benchmark show that our DIMN significantly outperforms alternative domain generalization or meta-learning methods. Jifei Song, Yongxin Yang, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
CVPR | 5 |
| 2019 | TuckER: Tensor Factorization for Knowledge Graph CompletionabstractIvana Balazevic, Carl Allen, Timothy Hospedales. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ivana Balazevic, Carl Allen, Timothy M. Hospedales |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Episodic Training for Domain GeneralizationabstractDomain generalization (DG) is the challenging and topical problem of learning models that generalize to novel testing domains with different statistics than a set of known training domains. The simple approach of aggregating data from all source domains and training a single deep neural network end-to-end on all the data provides a surprisingly strong baseline that surpasses many prior published methods. In this paper we build on this strong baseline by designing an episodic training procedure that trains a single deep network in a way that exposes it to the domain shift that characterises a novel domain at runtime. Specifically, we decompose a deep network into feature extractor and classifier components, and then train each component by simulating it interacting with a partner who is badly tuned for the current domain. This makes both components more robust, ultimately leading to our networks producing state-of-the-art performance on three DG benchmarks. Furthermore, we consider the pervasive workflow of using an ImageNet trained CNN as a fixed feature extractor for downstream recognition tasks. Using the Visual Decathlon benchmark, we demonstrate that our episodic-DG training improves the performance of such a general purpose feature extractor by explicitly training a feature for robustness to novel problems. This shows that DG training can benefit standard practice in computer vision. Da Li 0001, Jianshu Zhang 0001, Yongxin Yang, Cong Liu 0006, Yi-Zhe Song, Timothy M. Hospedales |
ICCV | 6 |
| 2019 | Goal-Driven Sequential Data AbstractionabstractAutomatic data abstraction is an important capability for both benchmarking machine intelligence and supporting summarization applications. In the former one asks whether a machine can `understand' enough about the meaning of input data to produce a meaningful but more compact abstraction. In the latter this capability is exploited for saving space or human time by summarizing the essence of input data. In this paper we study a general reinforcement learning based framework for learning to abstract sequential data in a goal-driven way. The ability to define different abstraction goals uniquely allows different aspects of the input data to be preserved according to the ultimate purpose of the abstraction. Our reinforcement learning objective does not require human-defined examples of ideal abstraction. Importantly our model processes the input sequence holistically without being constrained by the original input order. Our framework is also domain agnostic -- we demonstrate applications to sketch, video and text data and achieve promising results in all domains. Umar Riaz Muhammad, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002, Yi-Zhe Song |
ICCV | 3 |
| 2019 | Robust Person Re-Identification by Modelling Feature UncertaintyabstractWe aim to learn deep person re-identification (ReID) models that are robust against noisy training data. Two types of noise are prevalent in practice: (1) label noise caused by human annotator errors and (2) data outliers caused by person detector errors or occlusion. Both types of noise pose serious problems for training ReID models, yet have been largely ignored so far. In this paper, we propose a novel deep network termed DistributionNet for robust ReID. Instead of representing each person image as a feature vector, DistributionNet models it as a Gaussian distribution with its variance representing the uncertainty of the extracted features. A carefully designed loss is formulated in DistributionNet to unevenly allocate uncertainty across training samples. Consequently, noisy samples are assigned large variance/uncertainty, which effectively alleviates their negative impacts on model fitting. Extensive experiments demonstrate that our model is more effective than alternative noise-robust deep models. The source code is available at: https://github.com/TianyuanYu/DistributionNet. Da Li 0001, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002 |
ICCV | 4 |
| 2019 | Analogies Explained: Towards Understanding Word EmbeddingsabstractWord embeddings generated by neural network methods such as word2vec (W2V) are well known to exhibit seemingly linear behaviour, e.g. the embeddings of analogy “woman is to queen as man is to king” approximately describe a parallelogram. This property is particularly intriguing since the embeddings are not trained to achieve it. Several explanations have been proposed, but each introduces assumptions that do not hold in practice. We derive a probabilistically grounded definition of paraphrasing that we re-interpret as word transformation, a mathematical description of “$w_x$ is to $w_y$”. From these concepts we prove existence of linear relationship between W2V-type embeddings that underlie the analogical phenomenon, identifying explicit error terms. Carl Allen, Timothy M. Hospedales |
ICML | 2 |
| 2019 | Feature-Critic Networks for Heterogeneous Domain GeneralizationabstractThe well known domain shift issue causes model performance to degrade when deployed to a new target domain with different statistics to training. Domain adaptation techniques alleviate this, but need some instances from the target domain to drive adaptation. Domain generalisation is the recently topical problem of learning a model that generalises to unseen domains out of the box, and various approaches aim to train a domain-invariant feature extractor, typically by adding some manually designed losses. In this work, we propose a learning to learn approach, where the auxiliary loss that helps generalisation is itself learned. Beyond conventional domain generalisation, we consider a more challenging setting of heterogeneous domain generalisation, where the unseen domains do not share label space with the seen ones, and the goal is to train a feature representation that is useful off-the-shelf for novel data and novel categories. Experimental evaluation demonstrates that our method outperforms state-of-the-art solutions in both settings. Yiying Li, Yongxin Yang, Wei Zhou 0107, Timothy M. Hospedales |
ICML | 4 |
| 2019 | Learning-driven Coarse-to-Fine Articulated Robot TrackingabstractIn this work we present an articulated tracking approach for robotic manipulators, which relies only on visual cues from colour and depth images to estimate the robot's state when interacting with or being occluded by its environment. We hypothesise that articulated model fitting approaches can only achieve accurate tracking if subpixel-level accurate correspondences between observed and estimated state can be established. Previous work in this area has exclusively relied on either discriminative depth information or colour edge correspondences as tracking objective and required initialisation from joint encoders. In this paper we propose a coarse-to-fine articulated state estimator, which relies only on visual cues from colour edges and learned depth keypoints, and which is initialised from a robot state distribution predicted from a depth image. We evaluate our approach on four RGB-D sequences showing a KUICA LWR arm with a Schunk SDH2 hand interacting with its environment and demonstrate that this combined keypoint and edge tracking objective can estimate the palm position with an average error of 2. 5cm without using any joint encoder sensing. Christian Rauch 0002, Vladimir Ivan, Timothy M. Hospedales, Jamie Shotton, Maurice Fallon |
ICRA | 3 |
| 2019 | What the Vec? Towards Probabilistically Grounded EmbeddingsabstractWord2Vec (W2V) and Glove are popular word embedding algorithms that perform well on a variety of natural language processing tasks. The algorithms are fast, efficient and their embeddings widely used. Moreover, the W2V algorithm has recently been adopted in the field of graph embedding, where it underpins several leading algorithms. However, despite their ubiquity and the relative simplicity of their common architecture, what the embedding parameters of W2V and Glove learn, and why that it useful in downstream tasks largely remains a mystery. We show that different interactions of PMI vectors encode semantic properties that can be captured in low dimensional word embeddings by suitable projection, theoretically explaining why the embeddings of W2V and Glove work, and, in turn, revealing an interesting mathematical interconnection between the semantic relationships of relatedness, similarity, paraphrase and analogy. Carl Allen, Ivana Balazevic, Timothy M. Hospedales |
NeurIPS | 3 |
| 2019 | Multi-relational Poincaré Graph EmbeddingsabstractHyperbolic embeddings have recently gained attention in machine learning due to their ability to represent hierarchical data more accurately and succinctly than their Euclidean analogues. However, multi-relational knowledge graphs often exhibit multiple simultaneous hierarchies, which current hyperbolic models do not capture. To address this, we propose a model that embeds multi-relational graph data in the Poincaré ball model of hyperbolic space. Our Multi-Relational Poincaré model (MuRP) learns relation-specific parameters to transform entity embeddings by Möbius matrix-vector multiplication and Möbius addition. Experiments on the hierarchical WN18RR knowledge graph show that our Poincaré embeddings outperform their Euclidean counterpart and existing embedding methods on the link prediction task, particularly at lower dimensionality. Ivana Balazevic, Carl Allen, Timothy M. Hospedales |
NeurIPS | 3 |
| 2019 | Toward Deep Universal Sketch Perceptual GrouperabstractHuman free-hand sketches provide the useful data for studying human perceptual grouping, where the grouping principles such as the Gestalt laws of grouping are naturally in play during both the perception and sketching stages. In this paper, we make the first attempt to develop a universal sketch perceptual grouper. That is, a grouper that can be applied to sketches of any category created with any drawing style and ability, to group constituent strokes/segments into semantically meaningful object parts. The first obstacle to achieving this goal is the lack of large-scale datasets with grouping annotation. To overcome this, we contribute the largest sketch perceptual grouping dataset to date, consisting of 20 000 unique sketches evenly distributed over 25 object categories. Furthermore, we propose a novel deep perceptual grouping model learned with both generative and discriminative losses. The generative loss improves the generalization ability of the model, while the discriminative loss guarantees both local and global grouping consistency. Extensive experiments demonstrate that the proposed grouper significantly outperforms the state-of-the-art competitors. In addition, we show that our grouper is useful for a number of sketch analysis tasks, including sketch semantic segmentation, synthesis, and fine-grained sketch-based image retrieval. Ke Li 0004, Kaiyue Pang, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Honggang Zhang 0002 |
IEEE Trans. Image Process. | 5 |
| 2018 | Learning to Generalize: Meta-Learning for Domain GeneralizationabstractDomain shift refers to the well known problem that a model trained in one source domain performs poorly when appliedto a target domain with different statistics. Domain Generalization (DG) techniques attempt to alleviate this issue by producing models which by design generalize well to novel testing domains. We propose a novel meta-learning method for domain generalization. Rather than designing a specific model that is robust to domain shift as in most previous DG work, we propose a model agnostic training procedure for DG. Our algorithm simulates train/test domain shift during training by synthesizing virtual testing domains within each mini-batch. The meta-optimization objective requires that steps to improve training domain performance should also improve testing domain performance. This meta-learning procedure trains models with good generalization ability to novel domains. We evaluate our method and achieve state of the art results on a recent cross-domain image classification benchmark, as well demonstrating its potential on two classic reinforcement learning tasks. Da Li 0001, Yongxin Yang, Yi-Zhe Song, Timothy M. Hospedales |
AAAI | 4 |
| 2018 | SketchMate: Deep Hashing for Million-Scale Human Sketch RetrievalabstractWe propose a deep hashing framework for sketch retrieval that, for the first time, works on a multi-million scale human sketch dataset. Leveraging on this large dataset, we explore a few sketch-specific traits that were otherwise under-studied in prior literature. Instead of following the conventional sketch recognition task, we introduce the novel problem of sketch hashing retrieval which is not only more challenging, but also offers a better testbed for large-scale sketch analysis, since: (i) more fine-grained sketch feature learning is required to accommodate the large variations in style and Abstraction, and (ii) a compact binary code needs to be learned at the same time to enable efficient retrieval. Key to our network design is the embedding of unique characteristics of human sketch, where (i) a two-branch CNN-RNN architecture is adapted to explore the temporal ordering of strokes, and (ii) a novel hashing loss is specifically designed to accommodate both the temporal and Abstract traits of sketches. By working with a 3.8M sketch dataset, we show that state-of-the-art hashing models specifically engineered for static images fail to perform well on temporal sketch data. Our network on the other hand not only offers the best retrieval performance on various code sizes, but also yields the best generalization performance under a zero-shot setting and when re-purposed for sketch recognition. Such superior performances effectively demonstrate the benefit of our sketch-specific design. Peng Xu 0005, Yongye Huang, Tongtong Yuan, Kaiyue Pang, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Zhanyu Ma, Jun Guo 0002 |
CVPR | 7 |
| 2018 | Multi-Level Factorisation Net for Person Re-IdentificationabstractKey to effective person re-identification (Re-ID) is modelling discriminative and view-invariant factors of person appearance at both high and low semantic levels. Recently developed deep Re-ID models either learn a holistic single semantic level feature representation and/or require laborious human annotation of these factors as attributes. We propose Multi-Level Factorisation Net (MLFN), a novel network architecture that factorises the visual appearance of a person into latent discriminative factors at multiple semantic levels without manual annotation. MLFN is composed of multiple stacked blocks. Each block contains multiple factor modules to model latent factors at a specific level, and factor selection modules that dynamically select the factor modules to interpret the content of each input image. The outputs of the factor selection modules also provide a compact latent factor descriptor that is complementary to the conventional deeply learned features. MLFN achieves state-of-the-art results on three Re-ID datasets, as well as compelling results on the general object categorisation CIFAR-100 dataset. Xiaobin Chang, Timothy M. Hospedales, Tao Xiang 0002 |
CVPR | 2 |
| 2018 | Scalable and Effective Deep CCA via Soft DecorrelationabstractRecently the widely used multi-view learning model, Canonical Correlation Analysis (CCA) has been generalised to the non-linear setting via deep neural networks. Existing deep CCA models typically first decorrelate the feature dimensions of each view before the different views are maximally correlated in a common latent space. This feature decorrelation is achieved by enforcing an exact decorrelation constraint; these models are thus computationally expensive due to the matrix inversion or SVD operations required for exact decorrelation at each training iteration. Furthermore, the decorrelation step is often separated from the gradient descent based optimisation, resulting in sub-optimal solutions. We propose a novel deep CCA model Soft CCA to overcome these problems. Specifically, exact decorrelation is replaced by soft decorrelation via a mini-batch based Stochastic Decorrelation Loss (SDL) to be optimised jointly with the other training objectives. Extensive experiments show that the proposed soft CCA is more effective and efficient than existing deep CCA models. In addition, our SDL loss can be applied to other deep models beyond multi-view learning, and obtains superior performance compared to existing decorrelation losses. Xiaobin Chang, Tao Xiang 0002, Timothy M. Hospedales |
CVPR | 3 |
| 2018 | Sketch-a-Classifier: Sketch-Based Photo Classifier GenerationabstractContemporary deep learning techniques have made image recognition a reasonably reliable technology. However training effective photo classifiers typically takes numerous examples which limits image recognition's scalability and applicability to scenarios where images may not be available. This has motivated investigation into zero-shot learning, which addresses the issue via knowledge transfer from other modalities such as text. In this paper we investigate an alternative approach of synthesizing image classifiers: almost directly from a user's imagination, via freehand sketch. This approach doesn't require the category to be nameable or describable via attributes as per zero-shot learning. We achieve this via training a model regression network to map from free-hand sketch space to the space of photo classifiers. It turns out that this mapping can be learned in a category-agnostic way, allowing photo classifiers for new categories to be synthesized by user with no need for annotated training photos. We also demonstrate that this modality of classifier generation can also be used to enhance the granularity of an existing photo classifier, or as a complement to name-based zero-shot learning. Conghui Hu, Da Li 0001, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
CVPR | 5 |
| 2018 | IVQA: Inverse Visual Question AnsweringabstractWe propose the inverse problem of Visual question answering (iVQA), and explore its suitability as a benchmark for visuo-linguistic understanding. The iVQA task is to generate a question that corresponds to a given image and answer pair. Since the answers are less informative than the questions, and the questions have less learnable bias, an iVQA model needs to better understand the image to be successful than a VQA model. We pose question generation as a multi-modal dynamic inference process and propose an iVQA model that can gradually adjust its focus of attention guided by both a partially generated question and the answer. For evaluation, apart from existing linguistic metrics, we propose a new ranking metric. This metric compares the ground truth question's rank among a list of distractors, which allows the drawbacks of different algorithms and sources of error to be studied. Experimental results show that our model can generate diverse, grammatically correct and content correlated questions that match the given answer. Feng Liu 0036, Tao Xiang 0002, Timothy M. Hospedales, Wankou Yang, Changyin Sun 0001 |
CVPR | 3 |
| 2018 | Learning Deep Sketch Abstraction
Umar Riaz Muhammad, Yongxin Yang, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
CVPR | 5 |
| 2018 | Learning to Sketch With Shortcut Cycle ConsistencyabstractTo see is to sketch - free-hand sketching naturally builds ties between human and machine vision. In this paper, we present a novel approach for translating an object photo to a sketch, mimicking the human sketching process. This is an extremely challenging task because the photo and sketch domains differ significantly. Furthermore, human sketches exhibit various levels of sophistication and abstraction even when depicting the same object instance in a reference photo. This means that even if photo-sketch pairs are available, they only provide weak supervision signal to learn a translation model. Compared with existing supervised approaches that solve the problem of D(E(photo)) → sketch), where E(·) and D(·) denote encoder and decoder respectively, we take advantage of the inverse problem (e.g., D(E(sketch) → photo), and combine with the unsupervised learning tasks of within-domain reconstruction, all within a multi-task learning framework. Compared with existing unsupervised approaches based on cycle consistency (i.e., D(E(D(E(photo)))) → photo), we introduce a shortcut consistency enforced at the encoder bottleneck (e.g., D(E(photo)) → photo) to exploit the additional self-supervision. Both qualitative and quantitative results show that the proposed model is superior to a number of state-of-the-art alternatives. We also show that the synthetic sketches can be used to train a better fine-grained sketch-based image retrieval (FG-SBIR) model, effectively alleviating the problem of sketch data scarcity. Jifei Song, Kaiyue Pang, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
CVPR | 5 |
| 2018 | Learning to Compare: Relation Network for Few-Shot LearningabstractWe present a conceptually simple, flexible, and general framework for few-shot learning, where a classifier must learn to recognise new classes given only few examples from each. Our method, called the Relation Network (RN), is trained end-to-end from scratch. During meta-learning, it learns to learn a deep distance metric to compare a small number of images within episodes, each of which is designed to simulate the few-shot setting. Once trained, a RN is able to classify images of new classes by computing relation scores between query images and the few examples of each new class without further updating the network. Besides providing improved performance on few-shot learning, our framework is easily extended to zero-shot learning. Extensive experiments on five benchmarks demonstrate that our simple approach provides a unified and effective approach for both of these two tasks. Flood Sung, Yongxin Yang, Li Zhang 0040, Tao Xiang 0002, Philip Torr 0001, Timothy M. Hospedales |
CVPR | 6 |
| 2018 | Deep Mutual LearningabstractModel distillation is an effective and widely used technique to transfer knowledge from a teacher to a student network. The typical application is to transfer from a powerful large network or ensemble to a small network, in order to meet the low-memory or fast execution requirements. In this paper, we present a deep mutual learning (DML) strategy. Different from the one-way transfer between a static pre-defined teacher and a student in model distillation, with DML, an ensemble of students learn collaboratively and teach each other throughout the training process. Our experiments show that a variety of network architectures benefit from mutual learning and achieve compelling results on both category and instance recognition tasks. Surprisingly, it is revealed that no prior powerful teacher network is necessary - mutual learning of a collection of simple student networks works, and moreover outperforms distillation from a more powerful yet static teacher. Ying Zhang 0021, Tao Xiang 0002, Timothy M. Hospedales, Huchuan Lu |
CVPR | 3 |
| 2018 | Deep Multi-task Learning to Recognise Subtle Facial Expressions of Mental States
Guosheng Hu, Li Liu 0004, Yang Hua 0001, Zhihong Zhang 0001, Fumin Shen, Ling Shao 0001, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang |
ECCV (12) | 9 |
| 2018 | Universal Sketch Perceptual Grouping
Ke Li 0004, Kaiyue Pang, Jifei Song, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Honggang Zhang 0002 |
ECCV (8) | 6 |
| 2018 | Deep Factorised Inverse-Sketching
Kaiyue Pang, Da Li 0001, Jifei Song, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
ECCV (15) | 6 |
| 2018 | Learning Unsupervised Word Translations Without AdversariesabstractWord translation, or bilingual dictionary induction, is an important capability that impacts many multilingual language processing tasks.Recent research has shown that word translation can be achieved in an unsupervised manner, without parallel seed dictionaries or aligned corpora.However, state of the art methods for unsupervised bilingual dictionary induction are based on generative adversarial models, and as such suffer from their well known problems of instability and hyperparameter sensitivity.We present a statistical dependency-based approach to bilingual dictionary induction that is unsupervised -no seed dictionary or parallel corpora required; and introduces no adversary -therefore being much easier to train.Our method performs comparably to adversarial alternatives and outperforms prior non-adversarial methods. Tanmoy Mukherjee, Makoto Yamada, Timothy M. Hospedales |
EMNLP | 3 |
| 2018 | Deep Stock Representation Learning: From Candlestick Charts to Investment DecisionsabstractWe propose a novel investment decision strategy (IDS) based on deep learning. The performance of many IDSs is affected by stock similarity. Most existing stock similarity measurements have the problems: (a) The linear nature of many measurements cannot capture nonlinear stock dynamics; (b) The estimation of many similarity metrics (e.g. covariance) needs very long period historic data (e.g. 3K days) which cannot represent current market effectively; (c) They cannot capture translation-invariance. To solve these problems, we apply Convolutional AutoEncoder to learn a stock representation, based on which we propose a novel portfolio construction strategy by: (i) using the deeply learned representation and modularity optimisation to cluster stocks and identify diverse sectors, (ii) picking stocks within each cluster according to their Sharpe ratio (Sharpe 1994). Overall this strategy provides low-risk high-return portfolios. We use the Financial Times Stock Exchange 100 Index (FTSE 100) data for evaluation. Results show our portfolio outperforms FTSE 100 index and many well known funds in terms of total return in 2000 trading days. Guosheng Hu, Kai Yang 0031, Flood Sung, Zhihong Zhang 0001, Neil Robertson 0002, Timothy M. Hospedales, Qiangwei Miemie |
ICASSP | 10 |
| 2018 | Dynamic Ensemble Active Learning: A Non-Stationary Bandit with Expert AdviceabstractActive learning aims to reduce annotation cost by predicting which samples are useful for a human teacher to label. However it has become clear there is no best active learning algorithm. Inspired by various philosophies about what constitutes a good criteria, different algorithms perform well on different datasets. This has motivated research into ensembles of active learners that learn what constitutes a good criteria in a given scenario, typically via multi-armed bandit algorithms. Though algorithm ensembles can lead to better results, they overlook the fact that not only does algorithm efficacy vary across datasets, but also during a single active learning session. That is, the best criteria is non-stationary. This breaks existing algorithms' guarantees and hampers their performance in practice. In this paper, we propose dynamic ensemble active learning as a more general and promising research direction. We develop a dynamic ensemble active learner based on a non-stationary multi-armed bandit with expert advice algorithm. Our dynamic ensemble selects the right criteria at each step of active learning. It has theoretical guarantees, and shows encouraging results on 13 popular datasets. Kunkun Pang, Mingzhi Dong, Yang Wu 0001, Timothy M. Hospedales |
ICPR | 4 |
| 2018 | Visual Articulated Tracking in the Presence of OcclusionsabstractThis paper focuses on visual tracking of a robotic manipulator during manipulation. In this situation, tracking is prone to failure when visual distractions are created by the object being manipulated and the clutter in the environment. Current state-of-the-art approaches, which typically rely on model-fitting using Iterative Closest Point (ICP), fail in the presence of distracting data points and are unable to recover. Meanwhile, discriminative methods which are trained only to distinguish parts of the tracked object can also fail in these scenarios as data points from the occlusions are incorrectly classified as being from the manipulator. We instead propose to use the per-pixel data-to-model associations provided from a random forest to avoid local minima during model fitting. By training the random forest with artificial occlusions we can achieve increased robustness to occlusion and clutter present in the scene. We do this without specific knowledge about the type or location of the manipulated object. Our approach is demonstrated by using dense depth data from an RGB-D camera to track a robotic manipulator during manipulation and in presence of occlusions. Christian Rauch 0002, Timothy M. Hospedales, Jamie Shotton, Maurice Fallon |
ICRA | 2 |
| 2018 | Frankenstein: Learning Deep Face Representations Using Small DataabstractDeep convolutional neural networks have recently proven extremely effective for difficult face recognition problems in uncontrolled settings. To train such networks, very large training sets are needed with millions of labeled images. For some applications, such as near-infrared (NIR) face recognition, such large training data sets are not publicly available and difficult to collect. In this paper, we propose a method to generate very large training data sets of synthetic images by compositing real face images in a given data set. We show that this method enables to learn models from as few as 10 000 training images, which perform on par with models trained from 500 000 images. Using our approach, we also obtain state-of-the-art results on the CASIA NIR-VIS2.0 heterogeneous face recognition data set. Guosheng Hu, Xiaojiang Peng, Yongxin Yang, Timothy M. Hospedales, Jakob Verbeek |
IEEE Trans. Image Process. | 4 |
| 2017 | Gated Neural Networks for Option Pricing: Rationality by DesignabstractWe propose a neural network approach to price EU call options that significantly outperforms some existing pricing models and comes with guarantees that its predictions are economically reasonable. To achieve this, we introduce a class of gated neural networks that automatically learn to divide-and-conquer the problem space for robust and accurate pricing. We then derive instantiations of these networks that are 'rational by design' in terms of naturally encoding a valid call option surface that enforces no arbitrage principles. This integration of human insight within data-driven learning provides significantly better generalisation in pricing performance due to the encoded inductive bias in the learning, guarantees sanity in the model's predictions, and provides econometrically useful byproduct such as risk neutral density. Yongxin Yang, Yu Zheng 0032, Timothy M. Hospedales |
AAAI | 3 |
| 2017 | Now You See Me: Deep Face Hallucination for Unviewed Sketches
Conghui Hu, Da Li 0001, Yi-Zhe Song, Timothy M. Hospedales |
BMVC | 4 |
| 2017 | Cross-domain Generative Learning for Fine-Grained Sketch-Based Image Retrieval
Kaiyue Pang, Yi-Zhe Song, Tony Xiang, Timothy M. Hospedales |
BMVC | 4 |
| 2017 | Fine-Grained Image Retrieval: the Text/Sketch Input Dilemma
Jifei Song, Yi-Zhe Song, Tony Xiang, Timothy M. Hospedales |
BMVC | 4 |
| 2017 | Semantic Regularisation for Recurrent Image Annotation
Feng Liu 0036, Tao Xiang 0002, Timothy M. Hospedales, Wankou Yang, Changyin Sun 0001 |
CVPR | 3 |
| 2017 | Attribute-Enhanced Face Recognition with Neural Tensor Fusion NetworksabstractDeep learning has achieved great success in face recognition, however deep-learned features still have limited invariance to strong intra-personal variations such as large pose changes. It is observed that some facial attributes (e.g. eyebrow thickness, gender) are robust to such variations. We present the first work to systematically explore how the fusion of face recognition features (FRF) and facial attribute features (FAF) can enhance face recognition performance in various challenging scenarios. Despite the promise of FAF, we find that in practice existing fusion methods fail to leverage FAF to boost face recognition performance in some challenging scenarios. Thus, we develop a powerful tensor-based framework which formulates feature fusion as a tensor optimisation problem. It is nontrivial to directly optimise this tensor due to the large number of parameters to optimise. To solve this problem, we establish a theoretical equivalence between low-rank tensor optimisation and a two-stream gated neural network. This equivalence allows tractable learning using standard neural network optimisation tools, leading to accurate and stable optimisation. Experimental results show the fused feature works better than individual features, thus proving for the first time that facial attributes aid face recognition. We achieve state-of-the-art performance on three popular databases: MultiPIE (cross pose, lighting and expression), CASIA NIR-VIS2.0 (cross-modality environment) and LFW (uncontrolled environment). Guosheng Hu, Yang Hua 0001, Zhihong Zhang 0001, Sankha S. Mukherjee, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang |
ICCV | 7 |
| 2017 | Deeper, Broader and Artier Domain GeneralizationabstractThe problem of domain generalization is to learn from multiple training domains, and extract a domain-agnostic model that can then be applied to an unseen domain. Domain generalization (DG) has a clear motivation in contexts where there are target domains with distinct characteristics, yet sparse data for training. For example recognition in sketch images, which are distinctly more abstract and rarer than photos. Nevertheless, DG methods have primarily been evaluated on photo-only benchmarks focusing on alleviating the dataset bias where both problems of domain distinctiveness and data sparsity can be minimal. We argue that these benchmarks are overly straightforward, and show that simple deep learning baselines perform surprisingly well on them. In this paper, we make two main contributions: Firstly, we build upon the favorable domain shift-robust properties of deep learning methods, and develop a low-rank parameterized CNN model for end-to-end DG learning. Secondly, we develop a DG benchmark dataset covering photo, sketch, cartoon and painting domains. This is both more practically relevant, and harder (bigger domain shift) than existing benchmarks. The results show that our method outperforms existing DG alternatives, and our dataset provides a more significant DG challenge to drive future research. Da Li 0001, Yongxin Yang, Yi-Zhe Song, Timothy M. Hospedales |
ICCV | 4 |
| 2017 | Deep Spatial-Semantic Attention for Fine-Grained Sketch-Based Image RetrievalabstractHuman sketches are unique in being able to capture both the spatial topology of a visual object, as well as its subtle appearance details. Fine-grained sketch-based image retrieval (FG-SBIR) importantly leverages on such fine-grained characteristics of sketches to conduct instance-level retrieval of photos. Nevertheless, human sketches are often highly abstract and iconic, resulting in severe misalignments with candidate photos which in turn make subtle visual detail matching difficult. Existing FG-SBIR approaches focus only on coarse holistic matching via deep cross-domain representation learning, yet ignore explicitly accounting for fine-grained details and their spatial context. In this paper, a novel deep FG-SBIR model is proposed which differs significantly from the existing models in that: (1) It is spatially aware, achieved by introducing an attention module that is sensitive to the spatial position of visual details: (2) It combines coarse and fine semantic information via a shortcut connection fusion block: and (3) It models feature correlation and is robust to misalignments between the extracted features across the two domains by introducing a novel higher-order learnable energy function (HOLEF) based loss. Extensive experiments show that the proposed deep spatial-semantic attention model significantly outperforms the state-of-the-art. Jifei Song, Qian Yu 0002, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
ICCV | 5 |
| 2017 | Transferring CNNS to multi-instance multi-label classification on small datasetsabstractImage tagging is a well known challenge in image processing. It is typically addressed through multi-instance multi-label (MIML) classification methodologies. Convolutional Neural Networks (CNNs) possess great potential to perform well on MIML tasks, since multi-level convolution and max pooling coincide with the multi-instance setting and the sharing of hidden representation may benefit multi-label modeling. However, CNNs usually require a large amount of carefully labeled data for training, which is hard to obtain in many real applications. In this paper, we propose a new approach for transferring pre-trained deep networks such as VGG16 on Imagenet to small MIML tasks. We extract features from each group of the network layers and apply multiple binary classifiers to them for multi-label prediction. Moreover, we adopt an L1-norm regularized Logistic Regression (L1LR) to find the most effective features for learning the multi-label classifiers. The experiment results on two most-widely used and relatively small benchmark MIML image datasets demonstrate that the proposed approach can substantially outperform the state-of-the-art algorithms, in terms of all popular performance metrics. Mingzhi Dong, Kunkun Pang, Yang Wu 0001, Jing-Hao Xue, Timothy M. Hospedales, Tsukasa Ogasawara |
ICIP | 5 |
| 2017 | Deep Multi-task Representation Learning: A Tensor Factorisation Approach
Yongxin Yang, Timothy M. Hospedales |
ICLR (Poster) | 2 |
| 2017 | Tensor Based Knowledge Transfer Across Skill Categories for Robot ControlabstractAdvances in hardware and learning for control are enabling robots to perform increasingly dextrous and dynamic control tasks. These skills typically require a prohibitive amount of exploration for reinforcement learning, and so are commonly achieved by imitation learning from manual demonstration. The costly non-scalable nature of manual demonstration has motivated work into skill generalisation, e.g., through contextual policies and options. Despite good results, existing work along these lines is limited to generalising across variants of one skill such as throwing an object to different locations. In this paper we go significantly further and investigate generalisation across qualitatively different classes of control skills. In particular, we introduce a class of neural network controllers that can realise four distinct skill classes: reaching, object throwing, casting, and ball-in-cup. By factorising the weights of the neural network, we are able to extract transferrable latent skills, that enable dramatic acceleration of learning in cross-task transfer. With a suitable curriculum, this allows us to learn challenging dextrous control tasks like ball-in-cup from scratch with pure reinforcement learning. Chenyang Zhao 0007, Timothy M. Hospedales, Freek Stulp, Olivier Sigaud |
IJCAI | 2 |
| 2017 | Free-Hand Sketch Synthesis with Deformable Stroke ModelsabstractWe present a generative model which can automatically summarize the stroke composition of free-hand sketches of a given category. When our model is fit to a collection of sketches with similar poses, it discovers and learns the structure and appearance of a set of coherent parts, with each part represented by a group of strokes. It represents both consistent (topology) as well as diverse aspects (structure and appearance variations) of each sketch category. Key to the success of our model are important insights learned from a comprehensive study performed on human stroke data. By fitting this model to images, we are able to synthesize visually similar and pleasant free-hand sketches. Yi Li 0004, Yi-Zhe Song, Timothy M. Hospedales, Shaogang Gong |
Int. J. Comput. Vis. | 3 |
| 2017 | Transductive Zero-Shot Action Recognition by Word-Vector Embedding
Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
Int. J. Comput. Vis. | 2 |
| 2017 | Sketch-a-Net: A Deep Neural Network that Beats Humans
Qian Yu 0002, Yongxin Yang, Feng Liu 0036, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
Int. J. Comput. Vis. | 6 |
| 2017 | Weakly-Supervised Image Annotation and Segmentation with Objects and AttributesabstractWe propose to model complex visual scenes using a non-parametric Bayesian model learned from weakly labelled images abundant on media sharing sites such as Flickr. Given weak image-level annotations of objects and attributes without locations or associations between them, our model aims to learn the appearance of object and attribute classes as well as their association on each object instance. Once learned, given an image, our model can be deployed to tackle a number of vision problems in a joint and coherent manner, including recognising objects in the scene (automatic object annotation), describing objects using their attributes (attribute prediction and association), and localising and delineating the objects (object detection and semantic segmentation). This is achieved by developing a novel Weakly Supervised Markov Random Field Stacked Indian Buffet Process (WS-MRF-SIBP) that models objects and attributes as latent factors and explicitly captures their correlations within and across superpixels. Extensive experiments on benchmark datasets demonstrate that our weakly supervised model significantly outperforms weakly supervised alternatives and is often comparable with existing strongly supervised models on a variety of tasks including semantic segmentation, automatic image annotation and retrieval based on object-attribute associations. Zhiyuan Shi 0001, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Discovery of Shared Semantic Spaces for Multiscene Video Query and SummarizationabstractThe growing rate of public space closed-circuit television (CCTV) installations has generated a need for automated methods for exploiting video surveillance data, including scene understanding, query, behavior annotation, and summarization. For this reason, extensive research has been performed on surveillance scene understanding and analysis. However, most studies have considered single scenes or groups of adjacent scenes. The semantic similarity between different but related scenes (e.g., many different traffic scenes of a similar layout) is not generally exploited to improve any automated surveillance tasks and reduce manual effort. Exploiting commonality and sharing any supervised annotations between different scenes is, however, challenging due to the following reason: some scenes are totally unrelated and thus any information sharing between them would be detrimental, whereas others may share only a subset of common activities and thus information sharing is only useful if it is selective. Moreover, semantically similar activities that should be modeled together and shared across scenes may have quite different pixel-level appearances in each scene. To address these issues, we develop a new framework for distributed multiple-scene global understanding that clusters surveillance scenes by their ability to explain each other's behaviors and further discovers which subset of activities are shared versus scene specific within each cluster. We show how to use this structured representation of multiple scenes to improve common surveillance tasks, including scene activity understanding, cross-scene query-by-example, behavior classification with reduced supervised labeling requirements, and video summarization. In each case, we demonstrate how our multiscene model improves on a collection of standard single-scene models and a flat model of all scenes. Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Synergistic Instance-Level Subspace Alignment for Fine-Grained Sketch-Based Image RetrievalabstractWe study the problem of fine-grained sketch-based image retrieval. By performing instance-level (rather than category-level) retrieval, it embodies a timely and practical application, particularly with the ubiquitous availability of touchscreens. Three factors contribute to the challenging nature of the problem: 1) free-hand sketches are inherently abstract and iconic, making visual comparisons with photos difficult; 2) sketches and photos are in two different visual domains, i.e., black and white lines versus color pixels; and 3) fine-grained distinctions are especially challenging when executed across domain and abstraction-level. To address these challenges, we propose to bridge the image-sketch gap both at the high level via parts and attributes, as well as at the low level via introducing a new domain alignment method. More specifically, first, we contribute a data set with 304 photos and 912 sketches, where each sketch and image is annotated with its semantic parts and associated part-level attributes. With the help of this data set, second, we investigate how strongly supervised deformable part-based models can be learned that subsequently enable automatic detection of part-level attributes, and provide pose-aligned sketch-image comparisons. To reduce the sketch-image gap when comparing low-level features, third, we also propose a novel method for instance-level domain-alignment that exploits both subspace and instance-level cues to better align the domains. Finally, fourth, these are combined in a matching framework integrating aligned low-level features, mid-level geometric structure, and high-level semantic attributes. Extensive experiments conducted on our new data set demonstrate effectiveness of the proposed method. Ke Li 0004, Kaiyue Pang, Yi-Zhe Song, Timothy M. Hospedales, Tao Xiang 0002, Honggang Zhang 0002 |
IEEE Trans. Image Process. | 4 |
| 2016 | L1 Graph Based Sparse Model for Label De-noising
Xiaobin Chang, Tao Xiang 0002, Timothy M. Hospedales |
BMVC | 3 |
| 2016 | Deep Multi-task Attribute-driven Ranking for Fine-grained Sketch-based Image Retrieval
Jifei Song, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Xiang Ruan |
BMVC | 4 |
| 2016 | ForgetMeNot: Memory-Aware Forensic Facial Sketch MatchingabstractWe investigate whether it is possible to improve the performance of automated facial forensic sketch matching by learning from examples of facial forgetting over time. Forensic facial sketch recognition is a key capability for law enforcement, but remains an unsolved problem. It is extremely challenging because there are three distinct contributors to the domain gap between forensic sketches and photos: The well-studied sketch-photo modality gap, and the less studied gaps due to (i) the forgetting process of the eye-witness and (ii) their inability to elucidate their memory. In this paper, we address the memory problem head on by introducing a database of 400 forensic sketches created at different time-delays. Based on this database we build a model to reverse the forgetting process. Surprisingly, we show that it is possible to systematically "un-forget" facial details. Moreover, it is possible to apply this model to dramatically improve forensic sketch recognition in practice: we achieve the state of the art results when matching 195 benchmark forensic sketches against corresponding photos and a 10,030 mugshot database. Shuxin Ouyang, Timothy M. Hospedales, Yi-Zhe Song, Xueming Li 0002 |
CVPR | 2 |
| 2016 | Multivariate Regression on the Grassmannian for Predicting Novel DomainsabstractWe study the problem of predicting how to recognise visual objects in novel domains with neither labelled nor unlabelled training data. Domain adaptation is now an established research area due to its value in ameliorating the issue of domain shift between train and test data. However, it is conventionally assumed that domains are discrete entities, and that at least unlabelled data is provided in testing domains. In this paper, we consider the case where domains are parametrised by a vector of continuous values (e.g., time, lighting or view angle). We aim to use such domain metadata to predict novel domains for recognition. This allows a recognition model to be pre-calibrated for a new domain in advance (e.g., future time or view angle) without waiting for data collection and re-training. We achieve this by posing the problem as one of multivariate regression on the Grassmannian, where we regress a domain's subspace (point on the Grassmannian) against an independent vector of domain parameters. We derive two novel methodologies to achieve this challenging task: a direct kernel regression from RM ! G, and an indirect method with better extrapolation properties. We evaluate our methods on two crossdomain visual recognition benchmarks, where they perform close to the upper bound of full data domain adaptation. This demonstrates that data is not necessary for domain adaptation if a domain can be parametrically described. Yongxin Yang, Timothy M. Hospedales |
CVPR | 2 |
| 2016 | Sketch Me That ShoeabstractWe investigate the problem of fine-grained sketch-based image retrieval (SBIR), where free-hand human sketches are used as queries to perform instance-level retrieval of images. This is an extremely challenging task because (i) visual comparisons not only need to be fine-grained but also executed cross-domain, (ii) free-hand (finger) sketches are highly abstract, making fine-grained matching harder, and most importantly (iii) annotated cross-domain sketch-photo datasets required for training are scarce, challenging many state-of-the-art machine learning techniques. In this paper, for the first time, we address all these challenges, providing a step towards the capabilities that would underpin a commercial sketch-based image retrieval application. We introduce a new database of 1,432 sketchphoto pairs from two categories with 32,000 fine-grained triplet ranking annotations. We then develop a deep tripletranking model for instance-level SBIR with a novel data augmentation and staged pre-training strategy to alleviate the issue of insufficient fine-grained training data. Extensive experiments are carried out to contribute a variety of insights into the challenges of data sufficiency and over-fitting avoidance when training deep networks for finegrained cross-domain ranking tasks. Qian Yu 0002, Feng Liu 0036, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Chen Change Loy |
CVPR | 5 |
| 2016 | Multi-Task Zero-Shot Action Recognition with Prioritised Data Augmentation
Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
ECCV (2) | 2 |
| 2016 | Gaussian Visual-Linguistic Embedding for Zero-Shot RecognitionabstractAn exciting outcome of research at the intersection of language and vision is that of zeroshot learning (ZSL).ZSL promises to scale visual recognition by borrowing distributed semantic models learned from linguistic corpora and turning them into visual recognition models.However the popular word-vector DSM embeddings are relatively impoverished in their expressivity as they model each word as a single vector point.In this paper we explore word-distribution embeddings for ZSL.We present a visual-linguistic mapping for ZSL in the case where words and visual categories are both represented by distributions.Experiments show improved results on ZSL benchmarks due to this better exploiting of intra-concept variability in each modality Tanmoy Mukherjee, Timothy M. Hospedales |
EMNLP | 2 |
| 2016 | Emerging Topics in Learning from Noisy and Missing DataabstractWhile vital for handling most multimedia and computer vision problems, collecting large scale fully annotated datasets is a resource-consuming, often unaffordable task. Indeed, on the one hand datasets need to be large and variate enough so that learning strategies can successfully exploit the variability inherently present in real data, but on the other hand they should be small enough so that they can be fully annotated at a reasonable cost. With the overwhelming success of (deep) learning methods, the traditional problem of balancing between dataset dimensions and resources needed for annotations became a full-fledged dilemma. In this context, methodological approaches able to deal with partially described data sets represent a one-of-a-kind opportunity to find the right balance between data variability and resource-consumption in annotation. These include methods able to deal with noisy, weak or partial annotations. In this tutorial we will present several recent methodologies addressing different visual tasks under the assumption of noisy, weakly annotated data sets. Xavier Alameda-Pineda, Timothy M. Hospedales, Elisa Ricci 0001, Nicu Sebe, Xiaogang Wang 0001 |
ACM Multimedia | 2 |
| 2016 | Fine-grained sketch-based image retrieval: The role of part-aware attributesabstractWe study the problem of fine-grained sketch-based image retrieval. By performing instance-level (rather than category-level) retrieval, it embodies a timely and practical application, particularly with the ubiquitous availability of touchscreens. Three factors contribute to the challenging nature of the problem: (i) free-hand sketches are inherently abstract and iconic, making visual comparisons with photos more difficult, (ii) sketches and photos are in two different visual domains, i.e. black and white lines vs. color pixels, and (iii) fine-grained distinctions are especially challenging when executed across domain and abstraction-level. To address this, we propose to detect visual attributes at part-level, in order to build a new representation that not only captures fine-grained characteristics but also traverses across visual domains. More specifically, (i) we propose a dataset with 304 photos and 912 sketches, where each sketch and photo is annotated with its semantic parts and associated part-level attributes, and with the help of this dataset, we investigate (ii) how strongly-supervised deformable part-based models can be learned that subsequently enable automatic detection of part-level attributes, and (iii) a novel matching framework that synergistically integrates low-level features, mid-level geometric structure and high-level semantic attributes to boost retrieval performance. Extensive experiments conducted on our new dataset demonstrate value of the proposed method. Ke Li 0004, Kaiyue Pang, Yi-Zhe Song, Timothy M. Hospedales, Honggang Zhang 0002, Yichuan Hu |
WACV | 4 |
| 2016 | When and where to transfer for Bayesian network parameter learning
Yun Zhou 0001, Timothy M. Hospedales, Norman E. Fenton |
Expert Syst. Appl. | 2 |
| 2016 | A survey on heterogeneous face recognition: Sketch, infra-red, 3D and low-resolution
Shuxin Ouyang, Timothy M. Hospedales, Yi-Zhe Song, Xueming Li 0002, Chen Change Loy, Xiaogang Wang 0001 |
Image Vis. Comput. | 2 |
| 2016 | Robust Subjective Visual Property Prediction from Crowdsourced Pairwise LabelsabstractThe problem of estimating subjective visual properties from image and video has attracted increasing interest. A subjective visual property is useful either on its own (e.g. image and video interestingness) or as an intermediate representation for visual recognition (e.g. a relative attribute). Due to its ambiguous nature, annotating the value of a subjective visual property for learning a prediction model is challenging. To make the annotation more reliable, recent studies employ crowdsourcing tools to collect pairwise comparison labels. However, using crowdsourced data also introduces outliers. Existing methods rely on majority voting to prune the annotation outliers/errors. They thus require a large amount of pairwise labels to be collected. More importantly as a local outlier detection method, majority voting is ineffective in identifying outliers that can cause global ranking inconsistencies. In this paper, we propose a more principled way to identify annotation outliers by formulating the subjective visual property prediction task as a unified robust learning to rank problem, tackling both the outlier detection and learning to rank jointly. This differs from existing methods in that (1) the proposed method integrates local pairwise comparison labels together to minimise a cost that corresponds to global inconsistency of ranking order, and (2) the outlier detection and learning to rank problems are solved jointly. This not only leads to better detection of annotation outliers but also enables learning with extremely sparse annotations. Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Jiechao Xiong, Shaogang Gong, Yizhou Wang 0001, Yuan Yao 0011 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Sketch-a-Net that Beats HumansabstractWe propose a multi-scale multi-channel deep neural network framework that, for the first time, yields sketch recognition performance surpassing that of humans. Our superior performance is a result of explicitly embedding the unique characteristics of sketches in our model: (i) a network architecture designed for sketch rather than natural photo statistics, (ii) a multi-channel generalisation that encodes sequential ordering in the sketching process, and (iii) a multi-scale network ensemble with joint Bayesian fusion that accounts for the different levels of abstraction exhibited in free-hand sketches. We show that state-of-the-art deep networks specifically engineered for photos of natural objects fail to perform well on sketch recognition, regardless whether they are trained using photo or sketch. Our network on the other hand not only delivers the best performance on the largest human sketch dataset to date, but also is small in size making efficient training possible using just CPUs. Qian Yu 0002, Yongxin Yang, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales |
BMVC | 5 |
| 2015 | Making better use of edges via perceptual groupingabstractWe propose a perceptual grouping framework that organizes image edges into meaningful structures and demonstrate its usefulness on various computer vision tasks. Our grouper formulates edge grouping as a graph partition problem, where a learning to rank method is developed to encode probabilities of candidate edge pairs. In particular, RankSVM is employed for the first time to combine multiple Gestalt principles as cue for edge grouping. Afterwards, an edge grouping based object proposal measure is introduced that yields proposals comparable to state-of-the-art alternatives. We further show how human-like sketches can be generated from edge groupings and consequently used to deliver state-of-the-art sketch-based image retrieval performance. Last but not least, we tackle the problem of freehand human sketch segmentation by utilizing the proposed grouper to cluster strokes into semantic object parts. Yonggang Qi, Yi-Zhe Song, Tao Xiang 0002, Honggang Zhang 0002, Timothy M. Hospedales, Yi Li 0004, Jun Guo 0002 |
CVPR | 5 |
| 2015 | Transferring a semantic representation for person re-identification and searchabstractLearning semantic attributes for person re-identification and description-based person search has gained increasing interest due to attributes' great potential as a pose and view-invariant representation. However, existing attribute-centric approaches have thus far underperformed state-of-the-art conventional approaches. This is due to their nonscalable need for extensive domain (camera) specific annotation. In this paper we present a new semantic attribute learning approach for person re-identification and search. Our model is trained on existing fashion photography datasets - either weakly or strongly labelled. It can then be transferred and adapted to provide a powerful semantic description of surveillance person detections, without requiring any surveillance domain supervision. The resulting representation is useful for both unsupervised and supervised person re-identification, achieving state-of-the-art and near state-of-the-art performance respectively. Furthermore, as a semantic representation it allows description-based person search to be integrated within the same framework. Zhiyuan Shi 0001, Timothy M. Hospedales, Tao Xiang 0002 |
CVPR | 2 |
| 2015 | Semantic embedding space for zero-shot action recognitionabstractThe number of categories for action recognition is growing rapidly. It is thus becoming increasingly hard to collect sufficient training data to learn conventional models for each category. This issue may be ameliorated by the increasingly popular “zero-shot learning” (ZSL) paradigm. In this framework a mapping is constructed between visual features and a human interpretable semantic description of each category, allowing categories to be recognised in the absence of any training data. Existing ZSL studies focus primarily on image data, and attribute-based semantic representations. In this paper, we address zero-shot recognition in contemporary video action recognition tasks, using semantic word vector space as the common space to embed videos and category labels. This is more challenging because the mapping between the semantic space and space-time features of videos containing complex actions is more complex and harder to learn. We demonstrate that a simple self-training and data augmentation strategy can significantly improve the efficacy of this mapping. Experiments on human action datasets including HMDB51 and UCF101 demonstrate that our approach achieves the state-of-the-art zero-shot action recognition performance. Xun Xu 0002, Timothy M. Hospedales, Shaogang Gong |
ICIP | 2 |
| 2015 | Probabilistic Graphical Models Parameter Learning with Transferred Prior and Constraints
Yun Zhou 0001, Norman E. Fenton, Timothy M. Hospedales, Martin Neil |
UAI | 3 |
| 2015 | Free-hand sketch recognition by multi-kernel feature learning
Yi Li 0004, Timothy M. Hospedales, Yi-Zhe Song, Shaogang Gong |
Comput. Vis. Image Underst. | 2 |
| 2015 | Transductive Multi-View Zero-Shot LearningabstractMost existing zero-shot learning approaches exploit transfer learning via an intermediate semantic representation shared between an annotated auxiliary dataset and a target dataset with different classes and no annotation. A projection from a low-level feature space to the semantic representation space is learned from the auxiliary dataset and applied without adaptation to the target dataset. In this paper we identify two inherent limitations with these approaches. First, due to having disjoint and potentially unrelated classes, the projection functions learned from the auxiliary dataset/domain are biased when applied directly to the target dataset/domain. We call this problem the projection domain shift problem and propose a novel framework, transductive multi-view embedding, to solve it. The second limitation is the prototype sparsity problem which refers to the fact that for each target class, only a single prototype is available for zero-shot learning given a semantic representation. To overcome this problem, a novel heterogeneous multi-view hypergraph label propagation method is formulated for zero-shot learning in the transductive embedding space. It effectively exploits the complementary information offered by different semantic representations and takes advantage of the manifold structures of multiple representation spaces in a coherent manner. We demonstrate through extensive experiments that the proposed approach (1) rectifies the projection shift between the auxiliary and target domains, (2) exploits the complementarity of multiple semantic representations, (3) significantly outperforms existing methods for both zero-shot and N-shot recognition on three image and video benchmark datasets, and (4) enables novel cross-view annotation tasks. Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Bayesian Joint Modelling for Object Localisation in Weakly Labelled ImagesabstractWe address the problem of localisation of objects as bounding boxes in images and videos with weak labels. This weakly supervised object localisation problem has been tackled in the past using discriminative models where each object class is localised independently from other classes. In this paper, a novel framework based on Bayesian joint topic modelling is proposed, which differs significantly from the existing ones in that: (1) All foreground object classes are modelled jointly in a single generative model that encodes multiple object co-existence so that "explaining away" inference can resolve ambiguity and lead to better learning and localisation. (2) Image backgrounds are shared across classes to better learn varying surroundings and "push out" objects of interest. (3) Our model can be learned with a mixture of weakly labelled and unlabelled data, allowing the large volume of unlabelled images on the Internet to be exploited for learning. Moreover, the Bayesian formulation enables the exploitation of various types of prior knowledge to compensate for the limited supervision offered by weakly labelled data, as well as Bayesian domain adaptation for transfer learning. Extensive experiments on the PASCAL VOC, ImageNet and YouTube-Object videos datasets demonstrate the effectiveness of our Bayesian joint model for weakly supervised object localisation. Zhiyuan Shi 0001, Timothy M. Hospedales, Tao Xiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Cross-Modal Face Matching: Beyond Viewed Sketches
Shuxin Ouyang, Timothy M. Hospedales, Yi-Zhe Song, Xueming Li 0002 |
ACCV (2) | 2 |
| 2014 | Open-world Person Re-Identification by Multi-Label Assignment Inference
Brais Cancela, Timothy M. Hospedales, Shaogang Gong |
BMVC | 2 |
| 2014 | Transductive Multi-label Zero-shot Learning
Yanwei Fu 0001, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong |
BMVC | 3 |
| 2014 | Re-id: Hunting Attributes in the Wild
Ryan Layne, Timothy M. Hospedales, Shaogang Gong |
BMVC | 2 |
| 2014 | Intra-category sketch-based image retrieval by matching deformable part models
Yi Li 0004, Timothy M. Hospedales, Yi-Zhe Song, Shaogang Gong |
BMVC | 2 |
| 2014 | Transductive Multi-view Embedding for Zero-Shot Recognition and Annotation
Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Zhenyong Fu, Shaogang Gong |
ECCV (2) | 2 |
| 2014 | Interestingness Prediction by Robust Learning to Rank
Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong, Yuan Yao 0011 |
ECCV (2) | 2 |
| 2014 | Weakly Supervised Learning of Objects, Attributes and Their Associations
Zhiyuan Shi 0001, Yongxin Yang, Timothy M. Hospedales, Tao Xiang 0002 |
ECCV (2) | 3 |
| 2014 | Learning Multimodal Latent AttributesabstractThe rapid development of social media sharing has created a huge demand for automatic media classification and annotation techniques. Attribute learning has emerged as a promising paradigm for bridging the semantic gap and addressing data sparsity via transferring attribute knowledge in object recognition and relatively simple action classification. In this paper, we address the task of attribute learning for understanding multimedia data with sparse and incomplete labels. In particular, we focus on videos of social group activities, which are particularly challenging and topical examples of this task because of their multimodal content and complex and unstructured nature relative to the density of annotations. To solve this problem, we 1) introduce a concept of semilatent attribute space, expressing user-defined and latent attributes in a unified framework, and 2) propose a novel scalable probabilistic topic model for learning multimodal semilatent attributes, which dramatically reduces requirements for an exhaustive accurate attribute ontology and expensive annotation effort. We show that our framework is able to exploit latent attributes to outperform contemporary approaches for addressing a variety of realistic multimedia sparse data learning tasks including: multitask learning, learning with label noise, N-shot transfer learning, and importantly zero-shot learning. Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Bayesian Joint Topic Modelling for Weakly Supervised Object LocalisationabstractWe address the problem of localisation of objects as bounding boxes in images with weak labels. This weakly supervised object localisation problem has been tackled in the past using discriminative models where each object class is localised independently from other classes. We propose a novel framework based on Bayesian joint topic modelling. Our framework has three distinctive advantages over previous works: (1) All object classes and image backgrounds are modelled jointly together in a single generative model so that "explaining away" inference can resolve ambiguity and lead to better learning and localisation. (2) The Bayesian formulation of the model enables easy integration of prior knowledge about object appearance to compensate for limited supervision. (3) Our model can be learned with a mixture of weakly labelled and unlabelled data, allowing the large volume of unlabelled images on the Internet to be exploited for learning. Extensive experiments on the challenging VOC dataset demonstrate that our approach outperforms the state-of-the-art competitors. Zhiyuan Shi 0001, Timothy M. Hospedales, Tao Xiang 0002 |
ICCV | 2 |
| 2013 | Finding Rare Classes: Active Learning with Generative and Discriminative ModelsabstractDiscovering rare categories and classifying new instances of them are important data mining issues in many fields, but fully supervised learning of a rare class classifier is prohibitively costly in labeling effort. There has therefore been increasing interest both in active discovery: to identify new classes quickly, and active learning: to train classifiers with minimal supervision. These goals occur together in practice and are intrinsically related because examples of each class are required to train a classifier. Nevertheless, very few studies have tried to optimise them together, meaning that data mining for rare classes in new domains makes inefficient use of human supervision. Developing active learning algorithms to optimise both rare class discovery and classification simultaneously is challenging because discovery and classification have conflicting requirements in query criteria. In this paper, we address these issues with two contributions: a unified active learning model to jointly discover new categories and learn to classify them by adapting query criteria online; and a classifier combination algorithm that switches generative and discriminative classifiers as learning progresses. Extensive evaluation on a batch of standard UCI and vision data sets demonstrates the superiority of this approach over existing methods. Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2012 | Person Re-identification by AttributesabstractVisually identifying a target individual reliably in a crowded environment observed by a distributed camera network is critical to a variety of tasks in managing business information, border control, and crime prevention. Automatic re-identification of a human candidate from public space CCTV video is challenging due to spatiotemporal visual feature variations and strong visual similarity between different people, compounded by low-resolution and poor quality video data. In this work, we propose a novel method for re-identification that learns a selection and weighting of mid-level semantic attributes to describe people. Specifically, the model learns an attribute-centric, parts-based feature representation. This differs from and complements existing low-level features for re-identification that rely purely on bottom-up statistics for feature selection, which are limited in discriminating and identifying reliably visual appearances of target people appearing in different camera views under certain degrees of occlusion due to crowdedness. Our experiments demonstrate the effectiveness of our approach compared to existing feature representations when applied to benchmarking datasets. 1 Ryan Layne, Timothy M. Hospedales, Shaogang Gong |
BMVC | 2 |
| 2012 | Stream-based joint exploration-exploitation active learningabstractLearning from streams of evolving and unbounded data is an important problem, for example in visual surveillance or internet scale data. For such large and evolving real-world data, exhaustive supervision is impractical, particularly so when the full space of classes is not known in advance therefore joint class discovery (exploration) and boundary learning (exploitation) becomes critical. Active learning has shown promise in jointly optimising exploration-exploitation with minimal human supervision. However, existing active learning methods either rely on heuristic multi-criteria weighting or are limited to batch processing. In this paper, we present a new unified framework for joint exploration-exploitation active learning in streams without any heuristic weighting. Extensive evaluation on classification of various image and surveillance video datasets demonstrates the superiority of our framework over existing methods. Chen Change Loy, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong |
CVPR | 2 |
| 2012 | Attribute Learning for Understanding Unstructured Social Activity
Yanwei Fu 0001, Timothy M. Hospedales, Tao Xiang 0002, Shaogang Gong |
ECCV (4) | 2 |
| 2012 | A Unifying Theory of Active Discovery and Learning
Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
ECCV (5) | 1 |
| 2012 | Video Behaviour Mining Using a Dynamic Topic Model
Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
Int. J. Comput. Vis. | 1 |
| 2011 | Learning Tags from Unsegmented Videos of Multiple Human ActionsabstractProviding methods to support semantic interaction with growing volumes of video data is an increasingly important challenge for data mining. To this end, there has been some success in recognition of simple objects and actions in video, however most of this work requires strongly supervised training data. The supervision cost of these approaches therefore renders them economically non-scalable for real world applications. In this paper we address the problem of learning to annotate and retrieve semantic tags of human actions in realistic video data with sparsely provided tags of semantically salient activities. This is challenging because of (1) the multi-label nature of the learning problem and (2) realistic videos are often dominated by (semantically uninteresting) background activity un-supported by any tags of interest, leading to a strong irrelevant data problem. To address these challenges, we introduce a new topic model based approach to video tag annotation. Our model simultaneously learns a low dimensional representation of the video data, which dimensions are semantically relevant (supported by tags), and how to annotate videos with tags. Experimental evaluation on three different video action/activity datasets demonstrate the challenge of this problem, and value of our contribution. Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
ICDM | 1 |
| 2011 | Finding Rare Classes: Adapting Generative and Discriminative Models in Active Learning
Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
PAKDD (2) | 1 |
| 2011 | Identifying Rare and Subtle Behaviors: A Weakly Supervised Joint Topic ModelabstractOne of the most interesting and desired capabilities for automated video behavior analysis is the identification of rarely occurring and subtle behaviors. This is of practical value because dangerous or illegal activities often have few or possibly only one prior example to learn from and are often subtle. Rare and subtle behavior learning is challenging for two reasons: (1) Contemporary modeling approaches require more data and supervision than may be available and (2) the most interesting and potentially critical rare behaviors are often visually subtle-occurring among more obvious typical behaviors or being defined by only small spatio-temporal deviations from typical behaviors. In this paper, we introduce a novel weakly supervised joint topic model which addresses these issues. Specifically, we introduce a multiclass topic model with partially shared latent structure and associated learning and inference algorithms. These contributions will permit modeling of behaviors from as few as one example, even without localization by the user and when occurring in clutter, and subsequent classification and localization of such behaviors online and in real time. We extensively validate our approach on two standard public-space data sets, where it clearly outperforms a batch of contemporary alternatives. Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Learning Rare Behaviours
Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
ACCV (2) | 2 |
| 2009 | A Unified Bayesian Framework for Adaptive Visual TrackingabstractTracking is regarded as one of the most fundamental tasks in computer vision. It is used in many computer vision applications in fields such as surveillance, robotic navigation and 3D reconstruction to name but a few. Despite decades of research, the goal of fully automatic tracking of arbitrary types of objects in real world conditions is still an open problem. In this paper, we take a step toward the goal of general real-world tracking, and demonstrate a unified generative model for Bayesian multifeature, adaptive target tracking, or AMFT for short (Adaptive Multiple Feature Tracker). We derive a unified generative model for multi-sensory adaptive tracking which cleanly integrates tracking and the modeling of appearance change across multiple features in the same framework. The unified multi-feature observation model ensures that if one feature is not confident, e.g., color after an object crosses into a region of shadow, it is automatically down-weighted in its contribution to the appearance model update. In this way, without pre-training of specific object models, we achieve an extensible tracker for general object types, robust to real-world problems of clutter, appearance/lighting change and target model drift. The standard modeling assumptions made by a non-adaptive generative model are illustrated by the probabilistic graphical model in Figure 1(a). The unknown target state (e.g., location, size, velocity) xt is assumed to change with time t according to some process parameterized by A. At every time t, we make some noisy observations zt of the target xt (e.g., raw image or color histograms). The target is then tracked online by computing the posterior, p(xt |z1:t) over the true target location recursively. In the case of the Kalman filter (KF), all the distributions involved are Gaussian. In the case of the particle filter (PF), all the distributions involved are represented non-parametrically by a set of samples [1]. The true target model, e.g., the appearance or color histogram to search for, is assumed to be part of the parameters H, i.e., it is known and fixed by an operator or initialized by some external process. In many cases however, the true appearance of the target H may change significantly in time, e.g., the appearance changes when a subject moves between shade and sunlight. This is the case for outdoor surveillance applications and is the motivation for this research. Adaptive trackers [2, 3, 4, 6] have been proposed to update the target appearance online in various heuristic ways. We can formalise this more general modeling assumption generatively, by the generalized dynamic Bayesian network illustrated in Figure 1(b). In contrast to Figure 1(a), the true target model which was previously included in the fixed parameters H, is now included as the the initial condition y0 of a dynamic latent variable yt , formalizing the modeling assumption that the target appearance can change over time. In addition to the target state xt , the target appearance yt will therefore be incrementally and recursively updated as part of the process of inferring the latent variables in this model p(xt ,yt |z1:t). The latent space is of course now greatly expanded, and poses a more challenging inference problem than that of Figure 1(a). In Section 2 of the paper, we detail the specific parametric form of the model and an efficient inference algorithm. We evaluate our method (AMFT) against three contemporary trackers: A standard single feature particle filter (PF), mean-shift (MS) [5] and incremental visual tracking (IVT) [4]. The PF and MS trackers are non-adaptive color-based trackers, while IVT aims for pose and illumination change robustness by performing online adaptation in a subspace appearance model. Note that the AMFT, PF and IVT trackers track object scale, but MS does not. We evaluated these methods on a series of challenging video clips exhibiting a wide variety of data and object types for tracking, including far-field indoor and outdoor pedestrians with and without carried objects, vehicle tracking, and near-field indoor face trackH 1 x2 x3 Emanuel Zelniker, Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
BMVC | 2 |
| 2009 | A Markov Clustering Topic Model for mining behaviour in videoabstractThis paper addresses the problem of fully automated mining of public space video data. A novel Markov Clustering Topic Model (MCTM) is introduced which builds on existing Dynamic Bayesian Network models (e.g. HMMs) and Bayesian topic models (e.g. Latent Dirichlet Allocation), and overcomes their drawbacks on accuracy, robustness and computational efficiency. Specifically, our model profiles complex dynamic scenes by robustly clustering visual events into activities and these activities into global behaviours, and correlates behaviours over time. A collapsed Gibbs sampler is derived for offline learning with unlabeled training data, and significantly, a new approximation to online Bayesian inference is formulated to enable dynamic scene understanding and behaviour mining in new video data online in real-time. The strength of this model is demonstrated by unsupervised learning of dynamic scene models, mining behaviours and detecting salient events in three complex and crowded public scenes. Timothy M. Hospedales, Shaogang Gong, Tao Xiang 0002 |
ICCV | 1 |
| 2008 | An Adaptive Machine DirectorabstractWe model the class of problem faced by a video broadcast director, who must act as an active perception agent to select a view of interest to a human from a range of possibilities. Real-time learning of a broadcast direction policy is achieved by efficient online Bayesian learning of the model’s parameters based on intermittent user feedback. In contrast to existing machine direction systems, which are dedicated to a particular scenario, our novel approach allows flexible learning of direction policies for novel domains or for viewerspecific preferences. We illustrate the flexibility of our approach by applying our model to a selection of scenarios with audio-visual input including teleconferencing, meetings and dance entertainment. 1 Timothy M. Hospedales, Oliver Williams |
BMVC | 1 |
| 2008 | Implications of Noise and Neural Heterogeneity for Vestibulo-Ocular Reflex FidelityabstractThe vestibulo-ocular reflex (VOR) is characterized by a short-latency, high-fidelity eye movement response to head rotations at frequencies up to 20 Hz. Electrophysiological studies of medial vestibular nucleus (MVN) neurons, however, show that their response to sinusoidal currents above 10 to 12 Hz is highly nonlinear and distorted by aliasing for all but very small current amplitudes. How can this system function in vivo when single cell response cannot explain its operation? Here we show that the necessary wide VOR frequency response may be achieved not by firing rate encoding of head velocity in single neurons, but in the integrated population response of asynchronously firing, intrinsically active neurons. Diffusive synaptic noise and the pacemaker-driven, intrinsic firing of MVN cells synergistically maintain asynchronous, spontaneous spiking in a population of model MVN neurons over a wide range of input signal amplitudes and frequencies. Response fidelity is further improved by a reciprocal inhibitory link between two MVN populations, mimicking the vestibular commissural system in vivo, but only if asynchrony is maintained by noise and pacemaker inputs. These results provide a previously missing explanation for the full range of VOR function and a novel account of the role of the intrinsic pacemaker conductances in MVN cells. The values of diffusive noise and pacemaker currents that give optimal response fidelity yield firing statistics similar to those in vivo, suggesting that the in vivo network is tuned to optimal performance. While theoretical studies have argued that noise and population heterogeneity can improve coding, to our knowledge this is the first evidence indicating that these parameters are indeed tuned to optimize coding fidelity in a neural control system in vivo. Timothy M. Hospedales, Mark C. W. van Rossum, Bruce P. Graham, Mayank B. Dutia |
Neural Comput. | 1 |
| 2008 | Structure Inference for Bayesian Multisensory Scene UnderstandingabstractWe investigate a solution to the problem of multi-sensor scene understanding by formulating it in the framework of Bayesian model selection and structure inference. Humans robustly associate multimodal data as appropriate, but previous modelling work has focused largely on optimal fusion, leaving segregation unaccounted for and unexploited by machine perception systems. We illustrate a unifying, Bayesian solution to multi-sensor perception and tracking which accounts for both integration and segregation by explicit probabilistic reasoning about data association in a temporal context. Such explicit inference of multimodal data association is also of intrinsic interest for higher level understanding of multisensory data. We illustrate this using a probabilistic implementation of data association in a multi-party audio-visual scenario, where unsupervised learning and structure inference is used to automatically segment, associate and track individual subjects in audiovisual sequences. Indeed, the structure inference based framework introduced in this work provides the theoretical foundation needed to satisfactorily explain many confounding results in human psychophysics experiments involving multimodal cue integration and association. Timothy M. Hospedales, Sethu Vijayakumar |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Structure Inference for Bayesian Multisensory Perception and Tracking
Timothy M. Hospedales, Joel J. Cartwright, Sethu Vijayakumar |
IJCAI | 1 |