VLDB 2026 Research / reviewers in the wild / expert
Graham W. Taylor
dblp:17/1633 · also Graham Taylor 0001
· DBLP profile ↗
70ranked-venue papers
10as first author
24since 2021 · last 2025
0000-0001-5867-3652ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 61 · 8 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bridging the AI Gap: Evaluating the Impact of an AI Education Program for Caregivers on Parental LeaveabstractArtificial Intelligence (AI) literacy is increasingly important across many fields, yet caregivers remain underrepresented in AI-related fields due to a combination of systemic and individual barriers. To address this, the Caregivers and Machine Learning (C&ML) program developed and delivered an accessible AI education program to caregivers on parental leave. Two cohorts participated in this 6-week interprofessional program, featuring fundamental machine learning concepts, hands-on programming assignments, and a capstone project. This study examines the program's impact on participants, focusing on their motivations and barriers before, during, and after the program as outcomes after completion. Post-program surveys and semi-structured interviews highlight that caregivers often face barriers such as the rapid pace of AI, discrimination, and balancing caregiving responsibilities with learning new skills. The C&ML program's flexible structure and personalized support network were critical in enabling participants to fully engage in the program, leading to significant improvements in their knowledge of ML and increased confidence in applying these skills. After completing the program, 20\% of participants transitioned into AI-related roles or pursued further education. This research highlights the value of targeted, inclusive educational programs for underrepresented groups and provides practical recommendations for refining future AI training programs for caregivers. Kristina L. Kupferschmidt, Flora Wan, Juan Carrasquilla Alvarez, Dora Gaviria Castaño, Graham W. Taylor, Sedef Akinli Koçak |
AAAI | 5 |
| 2025 | CLIBD: Bridging Vision and Genomics for Biodiversity Monitoring at ScaleabstractMeasuring biodiversity is crucial for understanding ecosystem health. While prior works have developed machine learning models for taxonomic classification of photographic images and DNA separately, in this work, we introduce a multi-modal approach combining both, using CLIP-style contrastive learning to align images, barcode DNA, and text-based representations of taxonomic labels in a unified embedding space. This allows for accurate classification of both known and unknown insect species without task-specific fine-tuning, leveraging contrastive learning for the first time to fuse DNA and image data. Our method surpasses previous single-modality approaches in accuracy by over 8% on zero-shot learning tasks, showcasing its effectiveness in biodiversity studies. ZeMing Gong, Austin T. Wang, Xiaoliang Huo, Joakim Bruslund Haurum, Scott C. Lowe, Graham W. Taylor, Angel X. Chang |
ICLR | 6 |
| 2025 | LAST SToP for Modeling Asynchronous Time SeriesabstractWe present a novel prompt design for Large Language Models (LLMs) tailored to Asynchronous Time Series. Unlike regular time series, which assume values at evenly spaced time points, asynchronous time series consist of timestamped events occurring at irregular intervals, each described in natural language. Our approach effectively utilizes the rich natural language of event descriptions, allowing LLMs to benefit from their broad world knowledge for reasoning across different domains and tasks. This allows us to extend the scope of asynchronous time series analysis beyond forecasting to include tasks like anomaly detection and data imputation. We further introduce Stochastic Soft Prompting, a novel prompt-tuning mechanism that significantly improves model performance, outperforming existing finetuning methods such as QLORA. Through extensive experiments on real-world datasets, we demonstrate that our approach achieves state-of-the-art performance across different tasks and datasets. Thibaut Durand, Graham W. Taylor, Lilian W. Bialokozowicz |
ICML | 3 |
| 2025 | Adapting Prediction Sets to Distribution Shifts Without LabelsabstractRecently there has been a surge of interest to deploy confidence set predictions rather than point predictions in machine learning. Unfortunately, the effectiveness of such prediction sets is frequently impaired by distribution shifts in practice, and the challenge is often compounded by the lack of ground truth labels at test time. Focusing on a standard set-valued prediction framework called conformal prediction (CP), this paper studies how to improve its practical performance using only unlabeled data from the shifted test domain. This is achieved by two new methods called $\texttt{ECP}$ and $\texttt{E{\small A}CP}$, whose main idea is to adjust the score function in CP according to its base model’s own uncertainty evaluation. Through extensive experiments on a number of large-scale datasets and neural network architectures, we show that our methods provide consistent improvement over existing baselines and nearly match the performance of fully supervised methods. Kevin Kasa, Graham W. Taylor |
UAI | 4 |
| 2025 | LoSA: Long-Short-Range Adapter for Scaling End-to-End Temporal Action LocalizationabstractTemporal Action Localization (TAL) involves localizing and classifying action snippets in an untrimmed video. The emergence of large video foundation models has led RGB-only video backbones to outperform previous methods needing both RGB and optical flow modalities. Leveraging these large models is often limited to training only the TAL head due to the prohibitively large GPU memory required to adapt the video backbone for TAL. To overcome this limitation, we introduce LoSA, the first memory-and-parameter-efficient backbone adapter designed specifically for TAL to handle untrimmed videos. LoSA specializes for TAL by introducing Long-Short-range Adapters that adapt the intermediate layers of the video backbone over different temporal ranges. These adapters run parallel to the video backbone to significantly reduce memory footprint. LoSA also includes Long-Short-range Gated Fusion that strategically combines the output of these adapters from the video backbone layers to enhance the video features provided to the TAL head. Experiments show that LoSA significantly outperforms all existing methods on standard TAL benchmarks, THUMOS-14 and ActivityNet-v1.3, by scaling end-to-end backbone adaptation to billion-parameter-plus models like VideoMAEv2 (ViT-g) and leveraging them beyond head-only transfer learning. Gaurav Mittal, Ahmed Magooda, Ye Yu 0003, Graham W. Taylor |
WACV | 5 |
| 2024 | Open-Vocabulary Temporal Action Localization using Multimodal Guidance
Aditya Arora, Sanath Narayan, Salman Khan 0001, Fahad Shahbaz Khan, Graham W. Taylor |
BMVC | 6 |
| 2024 | Agglomerative Token Clustering
Joakim Bruslund Haurum, Sergio Escalera, Graham W. Taylor, Thomas B. Moeslund |
ECCV (57) | 3 |
| 2024 | BIOSCAN-5M: A Multimodal Dataset for Insect BiodiversityabstractAs part of an ongoing worldwide effort to comprehend and monitor insect biodiversity, this paper presents the BIOSCAN-5M Insect dataset to the machine learning community and establish several benchmark tasks. BIOSCAN-5M is a comprehensive dataset containing multi-modal information for over 5 million insect specimens, and it significantly expands existing image-based biological datasets by including taxonomic labels, raw nucleotide barcode sequences, assigned barcode index numbers, geographical, and size information. We propose three benchmark experiments to demonstrate the impact of the multi-modal data types on the classification and clustering accuracy. First, we pretrain a masked language model on the DNA barcode sequences of the BIOSCAN-5M dataset, and demonstrate the impact of using this large reference library on species- and genus-level classification performance. Second, we propose a zero-shot transfer learning task applied to images and DNA barcodes to cluster feature embeddings obtained from self-supervised learning, to investigate whether meaningful clusters can be derived from these representation embeddings. Third, we benchmark multi-modality by performing contrastive learning on DNA barcodes, image data, and taxonomic information. This yields a general shared embedding space enabling taxonomic classification using multiple types of information and modalities. The code repository of the BIOSCAN-5M Insect dataset is available at https://github.com/bioscan-ml/BIOSCAN-5M. Zahra Gharaee, Scott C. Lowe, ZeMing Gong, Pablo Millan Arias, Nicholas Pellegrino, Austin T. Wang, Joakim Bruslund Haurum, Iuliia Zarubiieva, Lila Kari, Dirk Steinke, Graham W. Taylor, Paul W. Fieguth, Angel X. Chang |
NeurIPS | 11 |
| 2024 | GCNet: Probing self-similarity learning for Generalized Counting Network
Mingjie Wang 0002, Yande Li, Graham W. Taylor, Minglun Gong |
Pattern Recognit. | 4 |
| 2023 | Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision TransformersabstractVision Transformers (ViT) have shown competitive advantages in terms of performance compared to convolutional neural networks (CNNs), though they often come with high computational costs. To this end, previous methods explore different attention patterns by limiting a fixed number of spatially nearby tokens to accelerate the ViT's multi-head self-attention (MHSA) operations. However, such structured attention patterns limit the token-to-token connections to their spatial relevance, which disregards learned semantic connections from a full attention mask. In this work, we propose an approach to learn instance-dependent attention patterns, by devising a lightweight connectivity predictor module that estimates the connectivity score of each pair of tokens. Intuitively, two tokens have high connectivity scores if the features are considered relevant either spatially or semantically. As each token only attends to a small number of other tokens, the binarized connectivity masks are often very sparse by nature and therefore provide the opportunity to reduce network FLOPs via sparse computations. Equipped with the learned unstructured attention pattern, sparse attention ViT (Sparsifiner) produces a superior Pareto frontier between FLOPs and top-1 accuracy on ImageNet compared to token sparsity. Our method reduces 48% ~ 69% FLOPs of MHSA while the accuracy drop is within 0.4%. We also show that combining attention and token sparsity reduces ViT FLOPs by over 60%. Cong Wei 0001, Brendan Duke, Ruowei Jiang, Parham Aarabi, Graham W. Taylor, Florian Shkurti |
CVPR | 5 |
| 2023 | The Catalog Problem: Clustering and Ordering Variable-Sized SetsabstractPrediction of a $\textbf{varying number}$ of $\textbf{ordered clusters}$ from sets of $\textbf{any cardinality}$ is a challenging task for neural networks, combining elements of set representation, clustering and learning to order. This task arises in many diverse areas, ranging from medical triage and early discharge, through machine part management and multi-channel signal analysis for petroleum exploration to product catalog structure prediction. This paper focuses on that last area, which exemplifies a number of challenges inherent to adaptive ordered clustering, referred to further as the eponymous $\textit{Catalog Problem}$. These include learning variable cluster constraints, exhibiting relational reasoning and managing combinatorial complexity. Despite progress in both neural clustering and set-to-sequence methods, no joint, fully differentiable model exists to-date. We develop such a modular architecture, referred to further as Neural Ordered Clusters (NOC), enhance it with a specific mechanism for learning cluster-level cardinality constraints, and provide a robust comparison of its performance in relation to alternative models. We test our method on three datasets, including synthetic catalog structures and PROCAT, a dataset of real-world catalogs consisting of over 1.5M products, achieving state-of-the-art results on a new, more challenging formulation of the underlying problem, which has not been addressed before. Additionally, we examine the network’s ability to learn higher-order interactions. Mateusz Jurewicz, Graham W. Taylor, Leon Derczynski |
ICML | 2 |
| 2023 | A Step Towards Worldwide Biodiversity Assessment: The BIOSCAN-1M Insect DatasetabstractIn an effort to catalog insect biodiversity, we propose a new large dataset of hand-labelled insect images, the BIOSCAN-1M Insect Dataset. Each record is taxonomically classified by an expert, and also has associated genetic information including raw nucleotide barcode sequences and assigned barcode index numbers, which are genetic-based proxies for species classification. This paper presents a curated million-image dataset, primarily to train computer-vision models capable of providing image-based taxonomic assessment, however, the dataset also presents compelling characteristics, the study of which would be of interest to the broader machine learning community. Driven by the biological nature inherent to the dataset, a characteristic long-tailed class-imbalance distribution is exhibited. Furthermore, taxonomic labelling is a hierarchical classification scheme, presenting a highly fine-grained classification problem at lower levels. Beyond spurring interest in biodiversity research within the machine learning community, progress on creating an image-based taxonomic classifier will also further the ultimate goal of all BIOSCAN research: to lay the foundation for a comprehensive survey of global biodiversity. This paper introduces the dataset and explores the classification task through the implementation and analysis of a baseline classifier. The code repository of the BIOSCAN-1M-Insect dataset is available at https://github.com/zahrag/BIOSCAN-1M Zahra Gharaee, ZeMing Gong, Nicholas Pellegrino, Iuliia Zarubiieva, Joakim Bruslund Haurum, Scott C. Lowe, Jaclyn T. A. McKeown, Chris C. Y. Ho, Joschka McLeod, Yi-Yun C. Wei, Jireh Agda, Sujeevan Ratnasingham, Dirk Steinke, Angel X. Chang, Graham W. Taylor, Paul W. Fieguth |
NeurIPS | 15 |
| 2022 | DelphAI: A human-centered approach to time-series forecastingabstractWhen applying machine learning (ML) based techniques to time-series forecasting applications, there are many domain-specific considerations that can be integrated into model development to improve the likelihood of successful real-world translation. A human-centered approach, that involves end-users, has the potential to address commonly cited concerns such as algorithmic trust, explainability, and fairness. We present the DelphAI framework as an example of a practical human-centered approach to ML-based time-series forecasting for applications where end-users have little familiarity with ML techniques. The proposed socio-technical methodology incorporates essential domain knowledge through stakeholder participation into the development of predictive models. We advocate that the application of user-centered design principles can improve downstream translation and address other ethical concerns associated with ML-based forecasting. Kristina L. Kupferschmidt, Joshua August Gus Skorburg, Graham W. Taylor |
IEEE Big Data | 3 |
| 2022 | On Evaluation Metrics for Graph Generative Models
Rylee Thompson, Boris Knyazev 0001, Elaheh Ghalebi, Jungtaek Kim 0001, Graham W. Taylor |
ICLR | 5 |
| 2021 | SSTVOS: Sparse Spatiotemporal Transformers for Video Object SegmentationabstractIn this paper we introduce a Transformer-based approach to video object segmentation (VOS). To address compounding error and scalability issues of prior work, we propose a scalable, end-to-end method for VOS called Sparse Spatiotemporal Transformers (SST). SST extracts per-pixel representations for each object in a video using sparse attention over spatiotemporal features. Our attention-based formulation for VOS allows a model to learn to attend over a history of multiple frames and provides suitable inductive bias for performing correspondence-like computations necessary for solving motion segmentation. We demonstrate the effectiveness of attention-based over recurrent networks in the spatiotemporal domain. Our method achieves competitive results on YouTube-VOS and DAVIS 2017 with improved scalability and robustness to occlusions compared with the state of the art. Code is available at https://github.com/dukebw/SSTVOS. Brendan Duke, Abdalla Ahmed, Christian Wolf 0001, Parham Aarabi, Graham W. Taylor |
CVPR | 5 |
| 2021 | LOHO: Latent Optimization of Hairstyles via OrthogonalizationabstractHairstyle transfer is challenging due to hair structure differences in the source and target hair. Therefore, we propose Latent Optimization of Hairstyles via Orthogonalization (LOHO), an optimization-based approach using GAN inversion to infill missing hair structure details in latent space during hairstyle transfer. Our approach decomposes hair into three attributes: perceptual structure, appearance, and style, and includes tailored losses to model each of these attributes independently. Furthermore, we propose two-stage optimization and gradient orthogonalization to enable disentangled latent space optimization of our hair attributes. Using LOHO for latent space manipulation, users can synthesize novel photorealistic images by manipulating hair attributes either individually or jointly, transferring the desired attributes from reference hairstyles. LOHO achieves a superior FID compared with the current state-of-the-art (SOTA) for hairstyle transfer. Additionally, LOHO preserves the subject’s identity comparably well according to PSNR and SSIM when compared to SOTA image embedding pipelines. Code is available at https://github.com/dukebw/LOHO. Rohit Saha, Brendan Duke, Florian Shkurti, Graham W. Taylor, Parham Aarabi |
CVPR | 4 |
| 2021 | Unconstrained Scene Generation with Locally Conditioned Radiance FieldsabstractWe tackle the challenge of learning a distribution over complex, realistic, indoor scenes. In this paper, we introduce Generative Scene Networks (GSN), which learns to decompose scenes into a collection of many local radiance fields that can be rendered from a free moving camera. Our model can be used as a prior to generate new scenes, or to complete a scene given only sparse 2D observations. Recent work has shown that generative models of radiance fields can capture properties such as multi-view consistency and view-dependent lighting. However, these models are specialized for constrained viewing of single objects, such as cars or faces. Due to the size and complexity of realistic indoor environments, existing models lack the representational capacity to adequately capture them. Our decomposition scheme scales to larger and more complex scenes while preserving details and diversity, and the learned prior enables high-quality rendering from viewpoints that are significantly different from observed viewpoints. When compared to existing models, GSN produces quantitatively higher-quality scene renderings across several different scene datasets. Terrance Devries, Miguel Ángel Bautista 0001, Nitish Srivastava, Graham W. Taylor, Joshua M. Susskind |
ICCV | 4 |
| 2021 | Generative Compositional Augmentations for Scene Graph PredictionabstractInferring objects and their relationships from an image in the form of a scene graph is useful in many applications at the intersection of vision and language. We consider a challenging problem of compositional generalization that emerges in this task due to a long tail data distribution. Current scene graph generation models are trained on a tiny fraction of the distribution corresponding to the most frequent compositions, e.g.. However, test images might contain zero- and few-shot compositions of objects and relationships, e.g.. Despite each of the object categories and the predicate (e.g. ‘on’) being frequent in the training data, the models often fail to properly understand such unseen or rare compositions. To improve generalization, it is natural to attempt increasing the diversity of the training distribution. However, in the graph domain this is non-trivial. To that end, we propose a method to synthesize rare yet plausible scene graphs by perturbing real ones. We then propose and empirically study a model based on conditional generative adversarial networks (GANs) that allows us to generate visual features of perturbed scene graphs and learn from them in a joint fashion. When evaluated on the Visual Genome dataset, our approach yields marginal, but consistent improvements in zero- and few-shot metrics. We analyze the limitations of our approach indicating promising directions for future research. Boris Knyazev 0001, Harm de Vries, Catalina Cangea, Graham W. Taylor, Aaron C. Courville, Eugene Belilovsky |
ICCV | 4 |
| 2021 | Context-aware Scene Graph Generation with Seq2Seq TransformersabstractScene graph generation is an important task in computer vision aimed at improving the semantic understanding of the visual world. In this task, the model needs to detect objects and predict visual relationships between them. Most of the existing models predict relationships in parallel assuming their independence. While there are different ways to capture these dependencies, we explore a conditional approach motivated by the sequence-to-sequence (Seq2Seq) formalism. Different from the previous research, our proposed model predicts visual relationships one at a time in an autoregressive manner by explicitly conditioning on the already predicted relationships. Drawing from translation models in NLP, we propose an encoder-decoder model built using Transformers where the encoder captures global context and long range interactions. The decoder then makes sequential predictions by conditioning on the scene graph constructed so far. In addition, we introduce a novel reinforcement learning-based training strategy tailored to Seq2Seq scene graph generation. By using a self-critical policy gradient training approach with Monte Carlo search we directly optimize for the (mean) recall metrics and bridge the gap between training and evaluation. Experimental results on two public benchmark datasets demonstrate that our Seq2Seq learning approach achieves strong empirical performance, outperforming previous state-of-the-art, while remaining efficient in terms of training and inference time. Full code for this work is available here: https://github.com/layer6ai-labs/SGG-Seq2Seq. Yichao Lu, Himanshu Rai, Boris Knyazev 0001, Guangwei Yu, Shashank Shekhar 0005, Graham W. Taylor, Maksims Volkovs |
ICCV | 7 |
| 2021 | CARE-AI special session on AI ethicsabstractSummary form only given. A complete record of the panel discussion was not made available for publication as part of the conference proceedings. This special session organized by the Centre for Advancing Responsible and Ethical Artificial Intelligence (CARE-AI) consists of two 90-minute parts, focusing on two groups at the frontline of AI Ethics: students and start-up founders. Part 1 is a student-led AI Ethics paper presentation and critique: two students from the Philosophy program will present original work, “Analyzing Distrust in Human Interactions with AI,” and “Enactivism and Modelling Human Behaviour in AI,” (20 min); each presentation will be followed by a prepared critique from a student in the Collaborative Specialization in AI (10 min) and a 15-minute general discussion with the audience. Part 2 is an AI Ethics start-up showcase: 5 Canadian start-up companies (whose products or services either present an AI Ethics dilemma or propose a solution) will present 5-minute pitches, which will each be followed by 5 minutes of expert commentary and 5 minutes of open discussion. Clair Baleshta, Dylan White, Glen Reavie, Alysha Cooper, Graham W. Taylor, Joshua August Gus Skorburg, David Van Bruwaene, Sarah Gignac, Chris Schmidt, Laura McDonald, Patricia Thaine, Chloë Ryan, Rency Luan |
ISTAS | 5 |
| 2021 | One Health Informatics and the stewardship of complex systemsabstractThis session explores how Complex Adaptive Systems provide a framework for analyzing important social, biological, and environmental systems in One Health. Anthropogenic disturbances, many of which are technological, pose a threat to key ecological and sociological processes. They lead us to consider questions such as: Is artificial intelligence a saviour or a demon? What are the political, ethical, and scientific implications for One Health? How might the Global Burden of Disease (human), the Global Burden of Animal Diseases (GBADs) and other Global Burdens constitute a broader “One Health Burdens of Disease” and provide an evidence-base for One Health decisions? It will be necessary to address different data challenges in the developed and developing worlds, many of which are ethical and political, not just technical. Panelists will discuss the GBADs approach to data sharing, including how FAIR-principled metadata can be used to create trustworthy data systems and how the Data Governance Handbook provides important guidance for communicating data sharing principles to data contributors and users. Each panelist will provide a 5-10 minute “primer” talk which will introduce and link the key themes. This will be followed by a moderated panel discussion with opportunities for the audience to pose questions. Graham W. Taylor, Theresa Bernardo, Deborah A. Stacey, Kassy Raymond, Rozita Dara 0001, Samira Yousefinaghani, Ethan Pike |
ISTAS | 1 |
| 2021 | Brick-by-Brick: Combinatorial Construction with Deep Reinforcement LearningabstractDiscovering a solution in a combinatorial space is prevalent in many real-world problems but it is also challenging due to diverse complex constraints and the vast number of possible combinations. To address such a problem, we introduce a novel formulation, combinatorial construction, which requires a building agent to assemble unit primitives (i.e., LEGO bricks) sequentially -- every connection between two bricks must follow a fixed rule, while no bricks mutually overlap. To construct a target object, we provide incomplete knowledge about the desired target (i.e., 2D images) instead of exact and explicit volumetric information to the agent. This problem requires a comprehensive understanding of partial information and long-term planning to append a brick sequentially, which leads us to employ reinforcement learning. The approach has to consider a variable-sized action space where a large number of invalid actions, which would cause overlap between bricks, exist. To resolve these issues, our model, dubbed Brick-by-Brick, adopts an action validity prediction network that efficiently filters invalid actions for an actor-critic network. We demonstrate that the proposed method successfully learns to construct an unseen object conditioned on a single image or multiple views of a target object. Hyunsoo Chung, Jungtaek Kim 0001, Boris Knyazev 0001, Jinhwi Lee, Graham W. Taylor, Jaesik Park, Minsu Cho |
NeurIPS | 5 |
| 2021 | Parameter Prediction for Unseen Deep ArchitecturesabstractDeep learning has been successful in automating the design of features in machine learning pipelines. However, the algorithms optimizing neural network parameters remain largely hand-designed and computationally inefficient. We study if we can use deep learning to directly predict these parameters by exploiting the past knowledge of training other networks. We introduce a large-scale dataset of diverse computational graphs of neural architectures - DeepNets-1M - and use it to explore parameter prediction on CIFAR-10 and ImageNet. By leveraging advances in graph neural networks, we propose a hypernetwork that can predict performant parameters in a single forward pass taking a fraction of a second, even on a CPU. The proposed model achieves surprisingly good performance on unseen and diverse networks. For example, it is able to predict all 24 million parameters of a ResNet-50 achieving a 60% accuracy on CIFAR-10. On ImageNet, top-5 accuracy of some of our networks approaches 50%. Our task along with the model and results can potentially lead to a new, more computationally efficient paradigm of training networks. Our model also learns a strong representation of neural architectures enabling their analysis. Boris Knyazev 0001, Michal Drozdzal, Graham W. Taylor, Adriana Romero-Soriano |
NeurIPS | 3 |
| 2021 | Learn, Generate, Rank, Explain: A Case Study of Visual Explanation by Generative Machine LearningabstractWhile the computer vision problem of searching for activities in videos is usually addressed by using discriminative models, their decisions tend to be opaque and difficult for people to understand. We propose a case study of a novel machine learning approach for generative searching and ranking of motion capture activities with visual explanation. Instead of directly ranking videos in the database given a text query, our approach uses a variant of Generative Adversarial Networks (GANs) to generate exemplars based on the query and uses them to search for the activity of interest in a large database. Our model is able to achieve comparable results to its discriminative counterpart, while being able to dynamically generate visual explanations. In addition to our searching and ranking method, we present an explanation interface that enables the user to successfully explore the model’s explanations and its confidence by revealing query-based, model-generated motion capture clips that contributed to the model’s decision. Finally, we conducted a user study with 44 participants to show that by using our model and interface, participants benefit from a deeper understanding of the model’s conceptualization of the search query. We discovered that the XAI system yielded a comparable level of efficiency, accuracy, and user-machine synchronization as its black-box counterpart, if the user exhibited a high level of trust for AI explanation. Chris Kim, Graham W. Taylor, Mohamed R. Amer |
ACM Trans. Interact. Intell. Syst. | 4 |
| 2020 | Graph Density-Aware Losses for Novel Compositions in Scene Graph Generation
Boris Knyazev 0001, Harm de Vries, Catalina Cangea, Graham W. Taylor, Aaron C. Courville, Eugene Belilovsky |
BMVC | 4 |
| 2020 | Learning Permutation Invariant Representations Using Memory Networks
Shivam Kalra, Mohammed Adnan, Graham W. Taylor, Hamid R. Tizhoosh |
ECCV (29) | 3 |
| 2020 | ProxyNCA++: Revisiting and Revitalizing Proxy Neighborhood Component Analysis
Eu Wern Teh, Terrance Devries, Graham W. Taylor |
ECCV (24) | 3 |
| 2020 | Modular Length Control for Sentence Generation
Katya Kudashkina, Peter Wittek, Jamie Kiros, Graham W. Taylor |
ESANN | 4 |
| 2020 | Enabling Continual Learning with Differentiable Hebbian PlasticityabstractContinual learning is the problem of sequentially learning new tasks or knowledge while protecting previously acquired knowledge. However, catastrophic forgetting poses a grand challenge for neural networks performing such learning process. Thus, neural networks that are deployed in the real world often struggle in scenarios where the data distribution is non-stationary (concept drift), imbalanced, or not always fully available, i.e., rare edge cases. We propose a Differentiable Hebbian Consolidation model which is composed of a Differentiable Hebbian Plasticity (DHP) Softmax layer that adds a rapid learning plastic component (compressed episodic memory) to the fixed (slow changing) parameters of the softmax output layer; enabling learned representations to be retained for a longer timescale. We demonstrate the flexibility of our method by integrating well-known task-specific synaptic consolidation methods to penalize changes in the slow weights that are important for each target task. We evaluate our approach on the Permuted MNIST, Split MNIST and Vision Datasets Mixture benchmarks, and introduce an imbalanced variant of Permuted MNIST - a dataset that combines the challenges of class imbalance and concept drift. Our proposed model requires no additional hyperparameters and outperforms comparable baselines by reducing forgetting. Vithursan Thangarasa, Thomas Miconi, Graham W. Taylor |
IJCNN | 3 |
| 2020 | Instance Selection for GANsabstractRecent advances in Generative Adversarial Networks (GANs) have led to their widespread adoption for the purposes of generating high quality synthetic imagery. While capable of generating photo-realistic images, these models often produce unrealistic samples which fall outside of the data manifold. Several recently proposed techniques attempt to avoid spurious samples, either by rejecting them after generation, or by truncating the model's latent space. While effective, these methods are inefficient, as a large fraction of training time and model capacity are dedicated towards samples that will ultimately go unused. In this work we propose a novel approach to improve sample quality: altering the training dataset via instance selection before model training has taken place. By refining the empirical data distribution before training, we redirect model capacity towards high-density regions, which ultimately improves sample fidelity, lowers model capacity requirements, and significantly reduces training time. Code is available at https://github.com/uoguelph-mlrg/instanceselectionfor_gans. Terrance Devries, Michal Drozdzal, Graham W. Taylor |
NeurIPS | 3 |
| 2020 | Multisource Domain Adaptation for Remote Sensing Using Deep Neural NetworksabstractIn applying machine learning to remote sensing problems, it is often the case that multiple training data sources, known as domains, are available for the same task. It is sample-inefficient to train separate models per domain, which motivates learning a single model from multiple sources. For example, the local climate zone (LCZ) classification problem that aims to produce per-pixel classifications of surface structure from remotely sensed images of urban and rural environments. These classification maps need to be generated for different cities at different times. To do this efficiently, available training data from different sources (i.e., cities) must be adapted for the task at hand. However, multisource domain adaptation (MDA) is a challenging problem and is particularly apparent when there are significant changes in the data distribution among these sources. In this article, we propose a scalable yet simple adaptive MDA (AMDA) framework to address this problem. AMDA is also capable of dealing with imbalanced data distributions among the sources more effectively than existing baselines. We also extend two techniques originally proposed for domain expansion (DE) to the task of DA. AMDA and the extended DE techniques are implemented and evaluated on the LCZ classification problem. Despite its simplicity, AMDA is able to achieve more than 12% improvement over the baseline. Ahmed Elshamli, Graham W. Taylor, Shawki Areibi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Image Classification with Hierarchical Multigraph Networks
Boris Knyazev 0001, Mohamed R. Amer, Graham W. Taylor |
BMVC | 4 |
| 2019 | Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic InstructionabstractConditional text-to-image generation is an active area of research, with many possible applications. Existing research has primarily focused on generating a single image from available conditioning information in one step. One practical extension beyond one-step generation is a system that generates an image iteratively, conditioned on ongoing linguistic input or feedback. This is significantly more challenging than one-step generation tasks, as such a system must understand the contents of its generated images with respect to the feedback history, the current feedback, as well as the interactions among concepts present in the feedback history. In this work, we present a recurrent image generation model which takes into account both the generated output up to the current step as well as all past instructions for generation. We show that our model is able to generate the background, add new objects, and apply simple transformations to existing objects. We believe our approach is an important step toward interactive generation. Code and data is available at: https://www.microsoft.com/en-us/research/project/generative-neural-visual-artist-geneva/. Alaaeldin El-Nouby, Shikhar Sharma 0001, Hannes Schulz, R. Devon Hjelm, Layla El Asri, Samira Ebrahimi Kahou, Yoshua Bengio, Graham W. Taylor |
ICCV | 8 |
| 2019 | Understanding Attention and Generalization in Graph Neural NetworksabstractWe aim to better understand attention over nodes in graph neural networks (GNNs) and identify factors influencing its effectiveness. We particularly focus on the ability of attention GNNs to generalize to larger, more complex or noisy graphs. Motivated by insights from the work on Graph Isomorphism Networks, we design simple graph reasoning tasks that allow us to study attention in a controlled environment. We find that under typical conditions the effect of attention is negligible or even harmful, but under certain conditions it provides an exceptional gain in performance of more than 60% in some of our classification tasks. Satisfying these conditions in practice is challenging and often requires optimal initialization or supervised training of attention. We propose an alternative recipe and train attention in a weakly-supervised fashion that approaches the performance of supervised models, and, compared to unsupervised models, improves results on several synthetic as well as real datasets. Source code and datasets are available at https://github.com/bknyaz/graphattentionpool. Boris Knyazev 0001, Graham W. Taylor, Mohamed R. Amer |
NeurIPS | 2 |
| 2019 | Classification and Re-Identification of Fruit Fly Individuals Across Days With Convolutional Neural NetworksabstractFruit flies have commonly been used as model organisms to understand various biological processes from genetics and development to social behavior. When studying multiple flies in an experiment, maintaining the identities of individual flies is often essential. The field of "computational ethology" has provided several methods that help biologists in this regard. Unmarked fly identities can be maintained by tracking each fly throughout the experiment, either manually correcting errors or computing and storing the statistics at each frame, but this becomes infeasible for experiments spanning several days. A popular alternative is to mark insects (with tags, paint or wing-clipping), but this has shown to have adverse effects on many of the same behaviors under study. To solve the above problems, we developed a method that can augment existing tracking software with a deep learning-based approach to recognize and re-identify individual fruit flies across days. This method should prove invaluable to biologists as it promises to revolutionize our ability to perform long-term experiments. To the best of our knowledge, this is the first work to investigate and successfully re-identify unmarked fruit flies across multiple days. We acquired our own dataset, which has around 3M images of 60 different fruit fly individuals, acquired in independent sets of 20. A vast majority (95%) of the flies have greater than 90% accuracy on a per-frame basis, and a majority of those (74%) have over 98% accuracy. With our method, less than 1 fly in 20 (on average) should be difficult to identify without resorting to parsimonious methods (i.e. examining which ID has not yet been assigned). This high classification accuracy demonstrates that we are able to successfully re-identify unmarked flies across days. Nihal Murali, Jon Schneider, Joel Levine, Graham W. Taylor |
WACV | 4 |
| 2018 | Self-Paced Learning with Adaptive Deep Visual Embeddings
Vithursan Thangarasa, Graham W. Taylor |
BMVC | 2 |
| 2018 | Glimpse Clouds: Human Activity Recognition From Unstructured Feature PointsabstractWe propose a method for human activity recognition from RGB data that does not rely on any pose information during test time, and does not explicitly calculate pose information internally. Instead, a visual attention module learns to predict glimpse sequences in each frame. These glimpses correspond to interest points in the scene that are relevant to the classified activities. No spatial coherence is forced on the glimpse locations, which gives the attention module liberty to explore different points at each frame and better optimize the process of scrutinizing visual information. Tracking and sequentially integrating this kind of unstructured data is a challenge, which we address by separating the set of glimpses from a set of recurrent tracking/recognition workers. These workers receive glimpses, jointly performing subsequent motion tracking and activity prediction. The glimpses are soft-assigned to the workers, optimizing coherence of the assignments in space, time and feature space using an external memory module. No hard decisions are taken, i.e. each glimpse point is assigned to all existing workers, albeit with different importance. Our methods outperform the state-of-the-art on the largest human activity recognition dataset available to-date, NTU RGB+D, and on the Northwestern-UCLA Multiview Action 3D Dataset. Fabien Baradel, Christian Wolf 0001, Julien Mille, Graham W. Taylor |
CVPR | 4 |
| 2018 | Attacking Binarized Neural Networks
Angus Galloway, Graham W. Taylor, Medhat A. Moussa |
ICLR (Poster) | 2 |
| 2018 | Quantitatively Evaluating GANs With Divergences Proposed for Training
Daniel Jiwoong Im, Graham W. Taylor, Kristin Branson |
ICLR (Poster) | 3 |
| 2018 | Distributed Sensor Network for Indirect Occupancy Measurement in Smart BuildingsabstractMany areas like smart buildings, crowd flow, action recognition, and assisted living rely on occupancy information. Although the use of smart cameras can alleviate the problem and provide accurate occupancy information, at the same time it can be cost prohibitive, invasive and not easy to scale or generalize to different environments. An alternative solution should bring similar accuracy while minimizing the previous problems. This work presents a candidate wireless sensor network for indirect occupancy measurements in a smart building. A prototype was built, that consists of CO2, temperature and humidity sensors. The prototype was placed in a complex indoor environment to collect data for a week. Then, the data was analyzed to examine potential correlation between sensor data and occupancy information. According to experimental results, CO2can be used for indirect occupancy measurements. Colin Brennan, Graham W. Taylor, Petros Spachos |
IWCMC | 2 |
| 2018 | Stochastic Layer-Wise Precision in Deep Neural Networks
Griffin Lacey, Graham W. Taylor, Shawki Areibi |
UAI | 2 |
| 2018 | Bayesian optimization on graph-structured search spaces: Optimizing deep multimodal fusion architectures
Dhanesh Ramachandram, Michal Lisicki, Timothy J. Shields, Mohamed R. Amer, Graham W. Taylor |
Neurocomputing | 5 |
| 2017 | Structure optimization for deep multimodal fusion networks using graph-induced kernels
Dhanesh Ramachandram, Michal Lisicki, Timothy J. Shields, Mohamed R. Amer, Graham W. Taylor |
ESANN | 5 |
| 2017 | Modout: Learning Multi-Modal Architectures by Stochastic RegularizationabstractModel selection methods based on stochastic regularization have been widely used in deep learning due to their simplicity and effectiveness. The well-known Dropout method treats all units, visible or hidden, in the same way, thus ignoring any a priori information related to grouping or structure. Such structure is present in multi-modal learning applications such as affect analysis and gesture recognition, where subsets of units may correspond to individual modalities. Here we describe Modout, a model selection method based on stochastic regularization, which is particularly useful in the multi-modal setting. Different from other forms of stochastic regularization, it is capable of learning whether or when to fuse two modalities in a layer, which is usually considered to be an architectural hyper-parameter by deep learning researchers and practitioners. Modout is evaluated on two real multi-modal datasets. The results indicate improved performance compared to other forms of stochastic regularization. The result on the Montalbano dataset shows that learning a fusion structure by Modout is on par with a state-of-the-art carefully designed architecture. Fan Li 0005, Natalia Neverova, Christian Wolf 0001, Graham W. Taylor |
FG | 4 |
| 2017 | Hand pose estimation through semi-supervised and weakly-supervised learning
Natalia Neverova, Christian Wolf 0001, Florian Nebout, Graham W. Taylor |
Comput. Vis. Image Underst. | 4 |
| 2016 | Caffeinated FPGAs: FPGA framework For Convolutional Neural NetworksabstractConvolutional Neural Networks (CNNs) have gained significant traction in the field of machine learning, particularly due to their high accuracy in visual recognition. Recent works have pushed the performance of GPU implementations of CNNs showing significant improvements in their classification and training times. With these improvements, many frameworks have become available for implementing CNNs on both CPUs and GPUs, with no support for FPGA implementations. In this work we present a modified version of the popular CNN framework Caffe, with FPGA support. This allows for classification using CNN models and specialized FPGA implementations with the flexibility of reprogramming the device when necessary, seamless memory transactions between host and device, simple-to-use test benches, and the ability to create pipelined layer implementations. To validate the framework, we use the Xilinx SDAccel environment to implement an FPGA-based Winograd convolution engine and show that it can be used alongside other layers running on a host processor to run several popular CNNs (AlexNet, GoogleNet, VGG A, Overfeat). The results show that our framework achieves 50 GFLOPS across 3×3 convolutions in the benchmarks. This is achieved within a practical framework, which will aid in future development of FPGA-based CNNs. Roberto DiCecco, Griffin Lacey, Jasmina Vasiljevic, Paul Chow, Graham W. Taylor, Shawki Areibi |
FPT | 5 |
| 2016 | Learning a metric for class-conditional KNNabstractNaïve Bayes Nearest Neighbour (NBNN) is a simple and effective framework which addresses many of the pitfalls of K-Nearest Neighbour (KNN) classification. It has yielded competitive results on several computer vision benchmarks. Its central tenet is that during NN search, a query is not compared to every example in a database, ignoring class information. Instead, NN searches are performed within each class, generating a score per class. A key problem with NN techniques, including NBNN, is that they fail when the data representation does not capture perceptual (e.g. class-based) similarity. NBNN circumvents this by using independent engineered descriptors (e.g. SIFT). To extend its applicability outside of image-based domains, we propose to learn a metric which captures perceptual similarity. Similar to how Neighbourhood Components Analysis optimizes a differentiable form of KNN classification, we propose “Class Conditional” metric learning (CCML), which optimizes a soft form of the NBNN selection rule. Typical metric learning algorithms learn either a global or local metric. However, our proposed method can be adjusted to a particular level of locality by tuning a single parameter. An empirical evaluation on classification and retrieval tasks demonstrates that our proposed method clearly outperforms existing learned distance metrics across a variety of image and non-image datasets. Daniel Jiwoong Im, Graham W. Taylor |
IJCNN | 2 |
| 2016 | ModDrop: Adaptive Multi-Modal Gesture RecognitionabstractWe present a method for gesture detection and localisation based on multi-scale and multi-modal deep learning. Each visual modality captures spatial information at a particular spatial scale (such as motion of the upper body or a hand), and the whole system operates at three temporal scales. Key to our technique is a training strategy which exploits: i) careful initialization of individual modalities; and ii) gradual fusion involving random dropping of separate channels (dubbed ModDrop) for learning cross-modality correlations while preserving uniqueness of each modality-specific representation. We present experiments on the ChaLearn 2014 Looking at People Challenge gesture recognition track, in which we placed first out of 17 teams. Fusing multiple modalities at several spatial and temporal scales leads to a significant increase in recognition rates, allowing the model to compensate for errors of the individual classifiers as well as noise in the separate channels. Furthermore, the proposed ModDrop training technique ensures robustness of the classifier to missing signals in one or several channels to produce meaningful predictions from any number of available modalities. In addition, we demonstrate the applicability of the proposed fusion scheme to modalities of arbitrary nature by experiments on the same dataset augmented with audio. Natalia Neverova, Christian Wolf 0001, Graham W. Taylor, Florian Nebout |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | An Empirical Investigation of Minimum Probability Flow Learning Under Different Connectivity Patterns
Daniel Jiwoong Im, Ethan Buchman, Graham W. Taylor |
ECML/PKDD (1) | 3 |
| 2015 | Scoring and Classifying with Gated Auto-Encoders
Daniel Jiwoong Im, Graham W. Taylor |
ECML/PKDD (1) | 2 |
| 2015 | Semisupervised Hyperspectral Image Classification via Neighborhood Graph LearningabstractIn problems where labeled data are scarce, semisupervised learning (SSL) techniques are an attractive framework that can exploit both labeled and unlabeled data. These approaches typically rely on a smoothness assumption such that examples that are similar in input space should also be similar in label space. In many domains, such as remotely sensed hyperspectral image (HSI) classification, the data violate this assumption. In response, we propose a general method by which a neighborhood graph used in SSL is learned using binary classifiers that are trained to predict whether a pair of pixels shares the same label. Working within the framework of semisupervised neural networks (SSNNs), we show that our approach improves on the performance of the SSNN on two HSI data sets. Daniel Jiwoong Im, Graham W. Taylor |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | Hand Segmentation with Structured Convolutional Learning
Natalia Neverova, Christian Wolf 0001, Graham W. Taylor, Florian Nebout |
ACCV (3) | 3 |
| 2014 | Human body part estimation from depth images via spatially-constrained deep learning
Mingyuan Jiu, Christian Wolf 0001, Graham W. Taylor, Atilla Baskurt |
Pattern Recognit. Lett. | 3 |
| 2012 | Analysis of the IJCNN 2011 UTL challenge
Isabelle Guyon, Gideon Dror, Vincent Lemaire 0001, Daniel L. Silver, Graham W. Taylor, David W. Aha |
Neural Networks | 5 |
| 2011 | Learning invariance through imitationabstractSupervised methods for learning an embedding aim to map high-dimensional images to a space in which perceptually similar observations have high measurable similarity. Most approaches rely on binary similarity, typically defined by class membership where labels are expensive to obtain and/or difficult to define. In this paper we propose crowd-sourcing similar images by soliciting human imitations. We exploit temporal coherence in video to generate additional pairwise graded similarities between the user-contributed imitations. We introduce two methods for learning nonlinear, invariant mappings that exploit graded similarities. We learn a model that is highly effective at matching people in similar pose. It exhibits remarkable invariance to identity, clothing, background, lighting, shift and scale. Graham W. Taylor, Ian Spiro, Christoph Bregler, Rob Fergus |
CVPR | 1 |
| 2011 | Adaptive deconvolutional networks for mid and high level feature learningabstractWe present a hierarchical model that learns image decompositions via alternating layers of convolutional sparse coding and max pooling. When trained on natural images, the layers of our model capture image information in a variety of forms: low-level edges, mid-level edge junctions, high-level object parts and complete objects. To build our model we rely on a novel inference scheme that ensures each layer reconstructs the input, rather than just the output of the layer directly beneath, as is common with existing hierarchical approaches. This makes it possible to learn multiple layers of representation and we show models with 4 layers, trained on images from the Caltech-101 and 256 datasets. When combined with a standard classifier, features extracted from these models outperform SIFT, as well as representations from other feature learning methods. Matthew D. Zeiler, Graham W. Taylor, Rob Fergus |
ICCV | 2 |
| 2011 | Unsupervised and transfer learning challengeabstractWe organized a data mining challenge in “unsupervised and transfer learning” (the UTL challenge), in collaboration with the DARPA Deep Learning program. The goal of this year's challenge was to learn good data representations that can be re-used across tasks by building models that capture regularities of the input space. The representations provided by the participants were evaluated by the organizers on supervised learning “target tasks”, which were unknown to the participants. In a first phase of the challenge, the competitors were given only unlabeled data to learn their data representation. In a second phase of the challenge, the competitors were also provided with a limited amount of labeled data from “source tasks”, distinct from the “target tasks”. We made available large datasets from various application domains: handwriting recognition, image recognition, video processing, text processing, and ecology. The results indicate that learned data representation yield results significantly better than what can be achieved with raw data or data preprocessed with standard normalizations and functional transforms. The UTL challenge is part of the IJCNN 2011 competition program1. The website of the challenge remains open for submission of new methods beyond the termination of the challenge as a resource for students and researchers2. Isabelle Guyon, Gideon Dror, Vincent Lemaire 0001, Graham W. Taylor, David W. Aha |
IJCNN | 4 |
| 2011 | Facial Expression Transfer with Input-Output Temporal Restricted Boltzmann MachinesabstractWe present a type of Temporal Restricted Boltzmann Machine that defines a probability distribution over an output sequence conditional on an input sequence. It shares the desirable properties of RBMs: efficient exact inference, an exponentially more expressive latent state than HMMs, and the ability to model nonlinear structure and dynamics. We apply our model to a challenging real-world graphics problem: facial expression transfer. Our results demonstrate improved performance over several baselines modeling high-dimensional 2D and 3D data. Matthew D. Zeiler, Graham W. Taylor, Leonid Sigal, Iain A. Matthews, Rob Fergus |
NIPS | 2 |
| 2011 | Two Distributed-State Models For Generating High-Dimensional Time Series
Graham W. Taylor, Geoffrey E. Hinton, Sam T. Roweis |
J. Mach. Learn. Res. | 1 |
| 2010 | Dynamical binary latent variable models for 3D human pose trackingabstractWe introduce a new class of probabilistic latent variable model called the Implicit Mixture of Conditional Restricted Boltzmann Machines (imCRBM) for use in human pose tracking. Key properties of the imCRBM are as follows: (1) learning is linear in the number of training exemplars so it can be learned from large datasets; (2) it learns coherent models of multiple activities; (3) it automatically discovers atomic “movemes” and (4) it can infer transitions between activities, even when such transitions are not present in the training set. We describe the model and how it is learned and we demonstrate its use in the context of Bayesian filtering for multi-view and monocular pose tracking. The model handles difficult scenarios including multiple activities and transitions among activities. We report state-of-the-art results on the HumanEva dataset. Graham W. Taylor, Leonid Sigal, David J. Fleet, Geoffrey E. Hinton |
CVPR | 1 |
| 2010 | Deconvolutional networksabstractBuilding robust low and mid-level image representations, beyond edge primitives, is a long-standing goal in vision. Many existing feature detectors spatially pool edge information which destroys cues such as edge intersections, parallelism and symmetry. We present a learning framework where features that capture these mid-level cues spontaneously emerge from image data. Our approach is based on the convolutional decomposition of images under a spar-sity constraint and is totally unsupervised. By building a hierarchy of such decompositions we can learn rich feature sets that are a robust image representation for both the analysis and synthesis of images. Matthew D. Zeiler, Dilip Krishnan, Graham W. Taylor, Rob Fergus |
CVPR | 3 |
| 2010 | Convolutional Learning of Spatio-temporal Features
Graham W. Taylor, Rob Fergus, Yann LeCun, Christoph Bregler |
ECCV (6) | 1 |
| 2010 | Body Motion Analysis for Multi-modal Identity VerificationabstractThis paper shows how “Body Motion Signature Analysis” - a new “soft-biometrics” technique - can be used for identity verification. It is able to extract motion features from the upper body of people and estimates so called “super-features” for input to a classifier. We demonstrate how this new technique can be used to identify people just based on their motion, or it can be used to significantly improve “hard-biometrics” techniques. For example, face verification achieves on this domain 6.45% Equal Error Rate (EER), and the combined verification performance of motion features and face reduces the error to 4.96% using an adaptive score-level integration method. The more ambiguous motion-only performance is 17.1% EER. Graham W. Taylor, Kirill Smolskiy, Christoph Bregler |
ICPR | 2 |
| 2010 | Pose-Sensitive Embedding by Nonlinear NCA RegressionabstractThis paper tackles the complex problem of visually matching people in similar pose but with different clothes, background, and other appearance changes. We achieve this with a novel method for learning a nonlinear embedding based on several extensions to the Neighborhood Component Analysis (NCA) framework. Our method is convolutional, enabling it to scale to realistically-sized images. By cheaply labeling the head and hands in large video databases through Amazon Mechanical Turk (a crowd-sourcing service), we can use the task of localizing the head and hands as a proxy for determining body pose. We apply our method to challenging real-world data and show that it can generalize beyond hand localization to infer a more general notion of body pose. We evaluate our method quantitatively against other embedding methods. We also demonstrate that real-world performance can be improved through the use of synthetic data. Graham W. Taylor, Rob Fergus, Ian Spiro, Christoph Bregler |
NIPS | 1 |
| 2009 | Modeling pigeon behavior using a Conditional Restricted Boltzmann Machine
Matthew D. Zeiler, Graham W. Taylor, Nikolaus F. Troje, Geoffrey E. Hinton |
ESANN | 2 |
| 2009 | Factored conditional restricted Boltzmann Machines for modeling motion styleabstractThe Conditional Restricted Boltzmann Machine (CRBM) is a recently proposed model for time series that has a rich, distributed hidden state and permits simple, exact inference. We present a new model, based on the CRBM that preserves its most important computational properties and includes multiplicative three-way interactions that allow the effective interaction weight between two units to be modulated by the dynamic state of a third unit. We factor the three-way weight tensor implied by the multiplicative model, reducing the number of parameters from O(N3) to O(N2). The result is an efficient, compact model whose effectiveness we demonstrate by modeling human motion. Like the CRBM, our model can capture diverse styles of motion with a single set of parameters, and the three-way interactions greatly improve the model's ability to blend motion styles or to transition smoothly among them. Graham W. Taylor, Geoffrey E. Hinton |
ICML | 1 |
| 2009 | Products of Hidden Markov Models: It Takes N>1 to Tango
Graham W. Taylor, Geoffrey E. Hinton |
UAI | 1 |
| 2008 | Maximum a posteriori ICA: Applying prior knowledge to the separation of acoustic sourcesabstractIndependent component analysis (ICA) for convolutive mixtures is often applied in the frequency domain due to the desirable decoupling into independent instantaneous mixtures per frequency bin. This approach suffers from a well-known scaling and permutation ambiguity. Existing methods perform a computation-heavy and sometimes unreliable phase of post-processing which typically makes use of knowledge regarding the geometry of the sensors post-ICA. In this paper, we propose a natural way to incorporate a priori knowledge of the unmixing matrix in the form of a prior distribution. This softly constrains ICA in a manner that avoids the permutation problem, and also allows us to integrate information about the environment, such as likely user configurations, into ICA using a unified statistical framework. Maximum a priori ICA easily follows from the maximum likelihood derivation of ICA. Its effectiveness is demonstrated through a series of experiments on convolutive mixtures of speech signals. Graham W. Taylor, Michael L. Seltzer, Alex Acero |
ICASSP | 1 |
| 2008 | The Recurrent Temporal Restricted Boltzmann MachineabstractThe Temporal Restricted Boltzmann Machine (TRBM) is a probabilistic model for sequences that is able to successfully model (i.e., generate nice-looking samples of) several very high dimensional sequences, such as motion capture data and the pixels of low resolution videos of balls bouncing in a box. The major disadvantage of the TRBM is that exact inference is extremely hard, since even computing a Gibbs update for a single variable of the posterior is exponentially expensive. This difficulty has necessitated the use of a heuristic inference procedure, that nonetheless was accurate enough for successful learning. In this paper we introduce the Recurrent TRBM, which is a very slight modification of the TRBM for which exact inference is very easy and exact gradient learning is almost tractable. We demonstrate that the RTRBM is better than an analogous TRBM at generating motion capture and videos of bouncing balls. Ilya Sutskever, Geoffrey E. Hinton, Graham W. Taylor |
NIPS | 3 |
| 2006 | Modeling Human Motion Using Binary Latent VariablesabstractWe propose a non-linear generative model for human motion data that uses an undirected model with binary latent variables and real-valued "visible" variables that represent joint angles. The latent and visible variables at each time step receive directed connections from the visible variables at the last few time-steps. Such an architecture makes on-line inference efficient and allows us to use a simple approximate learning procedure. After training, the model finds a single set of parameters that simultaneously capture several different kinds of motion. We demonstrate the power of our approach by synthesizing various motion sequences and by performing on-line filling in of data lost during motion capture. Website: http://www.cs.toronto.edu/gwtaylor/publications/nips2006mhmublv/ Graham W. Taylor, Geoffrey E. Hinton, Sam T. Roweis |
NIPS | 1 |