EDBT 2026 Demo / reviewers in the wild / expert
Christopher Kanan
dblp:14/8653 · also Chris Kanan
· DBLP profile ↗
41ranked-venue papers
3as first author
14since 2021 · last 2025
0000-0002-6412-995XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Controlling Neural Collapse Enhances Out-of-Distribution Detection and Transfer LearningabstractOut-of-distribution (OOD) detection and OOD generalization are widely studied in Deep Neural Networks (DNNs), yet their relationship remains poorly understood. We empirically show that the degree of Neural Collapse (NC) in a network layer is inversely related with these objectives: stronger NC improves OOD detection but degrades generalization, while weaker NC enhances generalization at the cost of detection. This trade-off suggests that a single feature space cannot simultaneously achieve both tasks. To address this, we develop a theoretical framework linking NC to OOD detection and generalization. We show that entropy regularization mitigates NC to improve generalization, while a fixed Simplex ETF projector enforces NC for better detection. Based on these insights, we propose a method to control NC at different DNN layers. In experiments, our method excels at both tasks across OOD datasets and DNN architectures. Md Yousuf Harun, Jhair Gallardo, Christopher Kanan |
ICML | 3 |
| 2025 | Dynamic Sparse Training of Diagonally Sparse NetworksabstractRecent advances in Dynamic Sparse Training (DST) have pushed the frontier of sparse neural network training in structured and unstructured contexts, matching dense-model performance while drastically reducing parameter counts to facilitate model scaling. However, unstructured sparsity often fails to translate into practical speedups on modern hardware. To address this shortcoming, we propose DynaDiag, a novel structured sparse-to-sparse DST method that performs at par with unstructured sparsity. DynaDiag enforces a diagonal sparsity pattern throughout training and preserves sparse computation in forward and backward passes. We further leverage the diagonal structure to accelerate computation via a custom CUDA kernel, rendering the method hardware-friendly. Empirical evaluations on diverse neural architectures demonstrate that our method maintains accuracy on par with unstructured counterparts while benefiting from tangible computational gains. Notably, with 90\% sparse linear layers in ViTs, we observe up to a 3.13x speedup in online inference without sacrificing model performance and a 1.59x speedup in training on a GPU compared to equivalent unstructured layers. Abhishek Tyagi, Arjun Iyer, William H. Renninger, Christopher Kanan, Yuhao Zhu 0001 |
ICML | 4 |
| 2024 | What Variables Affect Out-of-Distribution Generalization in Pretrained Models?abstractEmbeddings produced by pre-trained deep neural networks (DNNs) are widely used; however, their efficacy for downstream tasks can vary widely. We study the factors influencing transferability and out-of-distribution (OOD) generalization of pre-trained DNN embeddings through the lens of the tunnel effect hypothesis, which is closely related to intermediate neural collapse. This hypothesis suggests that deeper DNN layers compress representations and hinder OOD generalization. Contrary to earlier work, our experiments show this is not a universal phenomenon. We comprehensively investigate the impact of DNN architecture, training data, image resolution, and augmentations on transferability. We identify that training with high-resolution datasets containing many classes greatly reduces representation compression and improves transferability. Our results emphasize the danger of generalizing findings from toy datasets to broader contexts. Md Yousuf Harun, Kyungbok Lee, Gianmarco J. Gallardo, Giri Krishnan, Christopher Kanan |
NeurIPS | 5 |
| 2024 | Learning to Evaluate the Artness of AI-Generated ImagesabstractAssessing the artness of AI-generated images continues to be a challenge within the realm of image generation. Most existing metrics cannot be used to perform instance-level and reference-free artness evaluation. This paper presents ArtScore, a metric designed to evaluate the degree to which an image resembles authentic artworks by artists (or conversely photographs), thereby offering a novel approach to artness assessment. We first blend pre-trained models for photo and artwork generation, resulting in a series of mixed models. Subsequently, we utilize these mixed models to generate images exhibiting varying degrees of artness with pseudo-annotations. Each photorealistic image has a corresponding artistic counterpart and a series of interpolated images that range from realistic to artistic. This dataset is then employed to train a neural network that learns to estimate quantized artness levels of arbitrary images. Extensive experiments reveal that the artness levels predicted by ArtScorealign more closely with human artistic evaluation than existing evaluation metrics, such as Gram loss and ArtFID. Junyu Chen 0005, Jie An 0002, Hanjia Lyu, Christopher Kanan, Jiebo Luo 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Spatial and Temporal Attention-Based Emotion Estimation on HRI-AVC DatasetabstractMany attempts have been made at estimating discrete emotions (calmness, anxiety, boredom, surprise, anger) and continuous emotional measures commonly used in psychology, namely ‘valence’ (The pleasantness of the emotion being displayed) and ‘arousal’ (The intensity of the emotion being displayed). Existing methods to estimate arousal and valence rely on learning from data sets, where an expert annotator labels every image frame. Access to an expert annotator is not always possible, and the annotation can also be tedious. Hence it is more practical to obtain self-reported arousal and valence values directly from the human in a real-time Human-Robot collaborative setting. Hence this paper provides an emotion data set (HRI-AVC) obtained while conducting a human-robot interaction (HRI) task. The self-reported pair of labels in this data set is associated with a set of image frames. This paper also proposes a spatial and temporal attention-based network to estimate arousal and valence from this set of image frames. The results show that an attention-based network can estimate valence and arousal on the HRI-AVC data set even when Arousal and Valence values are unavailable per frame. Karthik Subramanian, Saurav Singh, Justin Namba, Jamison Heard, Christopher Kanan, Ferat Sahin |
SMC | 5 |
| 2023 | Semantic Segmentation with Active Semi-Supervised LearningabstractUsing deep learning, we now have the ability to create exceptionally good semantic segmentation systems; however, collecting the prerequisite pixel-wise annotations for training images remains expensive and time-consuming. Therefore, it would be ideal to minimize the number of human annotations needed when creating a new dataset. Here, we address this problem by proposing a novel algorithm that combines active learning and semi-supervised learning. Active learning is an approach for identifying the best unlabeled samples to annotate. While there has been work on active learning for segmentation, most methods require annotating all pixel objects in each image, rather than only the most informative regions. We argue that this is inefficient. Instead, our active learning approach aims to minimize the number of annotations per-image. Our method is enriched with semi-supervised learning, where we use pseudo labels generated with a teacher-student framework to identify image regions that help disambiguate confused classes. We also integrate mechanisms that enable better performance on imbalanced label distributions, which have not been studied previously for active learning in semantic segmentation. In experiments on the CamVid and CityScapes datasets, our method obtains over 95% of the network’s performance on the full-training set using less than 17% of the training data, whereas the previous state of the art required 40% of the training data. Aneesh Rangnekar, Christopher Kanan, Matthew J. Hoffman 0001 |
WACV | 2 |
| 2022 | Can I see an Example? Active Learning the Long Tail of Attributes and Relations
Tyler L. Hayes, Maximilian Nickel, Christopher Kanan, Ludovic Denoyer, Arthur Szlam |
BMVC | 3 |
| 2022 | Semantic Segmentation with Active Semi-Supervised Representation Learning
Aneesh Rangnekar, Christopher Kanan, Matthew J. Hoffman 0001 |
BMVC | 2 |
| 2022 | OccamNets: Mitigating Dataset Bias by Favoring Simpler Hypotheses
Robik Shrestha, Kushal Kafle, Christopher Kanan |
ECCV (20) | 3 |
| 2022 | Detecting Out-Of-Context Objects Using Graph Contextual Reasoning NetworkabstractThis paper presents an approach for detecting out-of-context (OOC) objects in images. Given an image with a set of objects, our goal is to determine if an object is inconsistent with the contextual relations and detect the OOC object with a bounding box. In this work, we consider common contextual relations such as co-occurrence relations, the relative size of an object with respect to other objects, and the position of the object in the scene. We posit that contextual cues are useful to determine object labels for in-context objects and inconsistent context cues are detrimental to determining object labels for out-of-context objects. To realize this hypothesis, we propose a graph contextual reasoning network (GCRN) to detect OOC objects. GCRN consists of two separate graphs to predict object labels based on the contextual cues in the image: 1) a representation graph to learn object features based on the neighboring objects and 2) a context graph to explicitly capture contextual cues from the neighboring objects. GCRN explicitly captures the contextual cues to improve the detection of in-context objects and identify objects that violate contextual relations. In order to evaluate our approach, we create a large-scale dataset by adding OOC object instances to the COCO images. We also evaluate on recent OCD benchmark. Our results show that GCRN outperforms competitive baselines in detecting OOC objects and correctly detecting in-context objects. Code and data: https://nusci.csl.sri.com/project/trinity-ooc Manoj Acharya, Kaushik Koneripalli, Susmit Jha, Christopher Kanan, Ajay Divakaran |
IJCAI | 5 |
| 2022 | An Investigation of Critical Issues in Bias Mitigation TechniquesabstractA critical problem in deep learning is that systems learn inappropriate biases, resulting in their inability to perform well on minority groups. This has led to the creation of multiple algorithms that endeavor to mitigate bias. However, it is not clear how effective these methods are. This is because study protocols differ among papers, systems are tested on datasets that fail to test many forms of bias, and systems have access to hidden knowledge or are tuned specifically to the test set. To address this, we introduce an improved evaluation protocol, sensible metrics, and a new dataset, which enables us to ask and answer critical questions about bias mitigation algorithms. We evaluate seven state-of-the-art algorithms using the same network architecture and hyperparameter selection policy across three benchmark datasets. We introduce a new dataset called Biased MNIST that enables assessment of robustness to multiple bias sources. We use Biased MNIST and a visual question answering (VQA) benchmark to assess robustness to hidden biases. Rather than only tuning to the test set distribution, we study robustness across different tuning distributions, which is critical because for many applications the test distribution may not be known during development. We find that algorithms exploit hidden biases, are unable to scale to multiple forms of bias, and are highly sensitive to the choice of tuning set. Based on our findings, we implore the community to adopt more rigorous assessment of future bias mitigation methods. All data, code, and results are publicly available1. Robik Shrestha, Kushal Kafle, Christopher Kanan |
WACV | 3 |
| 2022 | EllSeg-Gen, towards Domain Generalization for Head-Mounted EyetrackingabstractThe study of human gaze behavior in natural contexts requires algorithms for gaze estimation that are robust to a wide range of imaging conditions. However, algorithms often fail to identify features such as the iris and pupil centroid in the presence of reflective artifacts and occlusions. Previous work has shown that convolutional networks excel at extracting gaze features despite the presence of such artifacts. However, these networks often perform poorly on data unseen during training. This work follows the intuition that jointly training a convolutional network with multiple datasets learns a generalized representation of eye parts. We compare the performance of a single model trained with multiple datasets against a pool of models trained on individual datasets. Results indicate that models tested on datasets in which eye images exhibit higher appearance variability benefit from multiset training. In contrast, dataset-specific models generalize better onto eye images with lower appearance variability. Rakshit Sunil Kothari, Reynold J. Bailey, Christopher Kanan, Jeff B. Pelz, Gabriel J. Diaz |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2021 | Self-Supervised Training Enhances Online Continual Learning
Gianmarco J. Gallardo, Tyler L. Hayes, Christopher Kanan |
BMVC | 3 |
| 2021 | Replay in Deep Learning: Current Approaches and Missing Biological ElementsabstractReplay is the reactivation of one or more neural patterns that are similar to the activation patterns experienced during past waking experiences. Replay was first observed in biological neural networks during sleep, and it is now thought to play a critical role in memory formation, retrieval, and consolidation. Replay-like mechanisms have been incorporated in deep artificial neural networks that learn over time to avoid catastrophic forgetting of previous knowledge. Replay algorithms have been successfully used in a wide range of deep learning methods within supervised, unsupervised, and reinforcement learning paradigms. In this letter, we provide the first comprehensive comparison between replay in the mammalian brain and replay in artificial neural networks. We identify multiple aspects of biological replay that are missing in deep learning systems and hypothesize how they could be used to improve artificial neural networks. Tyler L. Hayes, Giri P. Krishnan, Maxim Bazhenov, Hava T. Siegelmann, Terrence J. Sejnowski, Christopher Kanan |
Neural Comput. | 6 |
| 2020 | A negative case analysis of visual grounding methods for VQAabstractExisting Visual Question Answering (VQA) methods tend to exploit dataset biases and spurious statistical correlations, instead of producing right answers for the right reasons. To address this issue, recent bias mitigation methods for VQA propose to incorporate visual cues (e.g., human attention maps) to better ground the VQA models, showcasing impressive gains. However, we show that the performance improvements are not a result of improved visual grounding, but a regularization effect which prevents over-fitting to linguistic priors. For instance, we find that it is not actually necessary to provide proper, human-based cues; random, insensible cues also result in similar improvements. Based on this observation, we propose a simpler regularization scheme that does not require any external annotations and yet achieves near state-of-the-art performance on VQA-CPv2. Robik Shrestha, Kushal Kafle, Christopher Kanan |
ACL | 3 |
| 2020 | RODEO: Replay for Online Object Detection
Manoj Acharya, Tyler L. Hayes, Christopher Kanan |
BMVC | 3 |
| 2020 | REMIND Your Neural Network to Prevent Catastrophic Forgetting
Tyler L. Hayes, Kushal Kafle, Robik Shrestha, Manoj Acharya, Christopher Kanan |
ECCV (8) | 5 |
| 2020 | On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawabstractOut-of-distribution (OOD) testing is increasingly popular for evaluating a machine learning system's ability to generalize beyond the biases of a training set. OOD benchmarks are designed to present a different joint distribution of data and labels between training and test time. VQA-CP has become the standard OOD benchmark for visual question answering, but we discovered three troubling practices in its current use. First, most published methods rely on explicit knowledge of the construction of the OOD splits. They often rely on inverting'' the distribution of labels, e.g. answering mostlyyes'' when the common training answer was ``no''. Second, the OOD test set is used for model selection. Third, a model's in-domain performance is assessed after retraining it on in-domain splits (VQA v2) that exhibit a more balanced distribution of labels. These three practices defeat the objective of evaluating generalization, and put into question the value of methods specifically designed for this dataset. We show that embarrassingly-simple methods, including one that generates answers at random, surpass the state of the art on some question types. We provide short- and long-term solutions to avoid these pitfalls and realize the benefits of OOD evaluation. Damien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha, Christopher Kanan, Anton van den Hengel |
NeurIPS | 5 |
| 2020 | Answering Questions about Data Visualizations using Efficient Bimodal FusionabstractChart question answering (CQA) is a newly proposed visual question answering (VQA) task where an algorithm must answer questions about data visualizations, e.g. bar charts, pie charts, and line graphs. CQA requires capabilities that natural-image VQA algorithms lack: fine-grained measurements, optical character recognition, and handling out-of-vocabulary words in both questions and answers. Without modifications, state-of-the-art VQA algorithms perform poorly on this task. Here, we propose a novel CQA algorithm called parallel recurrent fusion of image and language (PReFIL). PReFIL first learns bimodal embeddings by fusing question and image features and then intelligently aggregates these learned embeddings to answer the given question. Despite its simplicity, PReFIL greatly surpasses state-of-the art systems and human baselines on both the FigureQA and DVQA datasets. Additionally, we demonstrate that PReFIL can be used to reconstruct tables by asking a series of questions about a chart. Kushal Kafle, Robik Shrestha, Brian L. Price, Scott Cohen, Christopher Kanan |
WACV | 5 |
| 2020 | AeroRIT: A New Scene for Hyperspectral Image AnalysisabstractWe investigate applying convolutional neural network (CNN) architecture to facilitate aerial hyperspectral scene understanding and present a new hyperspectral data set, AeroRIT, which is large enough for CNN training. To date, the majority of hyperspectral airborne has been confined to various subcategories of vegetation and roads, and this scene introduces two new categories: buildings and cars. To the best of our knowledge, this is the first comprehensive large-scale hyperspectral scene with nearly seven-million pixel annotations for identifying cars, roads, and buildings. We compare the performance of the three popular architectures-SegNet, U-Net, and Res-U-Net, for scene understanding and object identification via the task of dense semantic segmentation to establish a benchmark for the scene. To further strengthen the network, we add squeeze and excitation blocks for better channel interactions and use self-supervised learning for better encoder initialization. Aerial hyperspectral image analysis has been restricted to small data sets with limited train/test splits capabilities, and we believe that AeroRIT will help advance the research in the field with a more complex object distribution to perform well on. The full data set, with flight lines in radiance and reflectance domains, is available for download at https://github.com/aneesh3108/AeroRIT. This data set is the first step toward developing robust algorithms for hyperspectral airborne sensing that can robustly perform advanced tasks such as vehicle tracking and occlusion handling. Aneesh Rangnekar, Nilay Mokashi, Emmett J. Ientilucci, Christopher Kanan, Matthew J. Hoffman 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | TallyQA: Answering Complex Counting QuestionsabstractMost counting questions in visual question answering (VQA) datasets are simple and require no more than object detection. Here, we study algorithms for complex counting questions that involve relationships between objects, attribute identification, reasoning, and more. To do this, we created TallyQA, the world’s largest dataset for open-ended counting. We propose a new algorithm for counting that uses relation networks with region proposals. Our method lets relation networks be efficiently used with high-resolution imagery. It yields stateof-the-art results compared to baseline and recent systems on both TallyQA and the HowMany-QA benchmark. Manoj Acharya, Kushal Kafle, Christopher Kanan |
AAAI | 3 |
| 2019 | Answer Them All! Toward Universal Visual Question Answering ModelsabstractVisual Question Answering (VQA) research is split into two camps: the first focuses on VQA datasets that require natural image understanding and the second focuses on synthetic datasets that test reasoning. A good VQA algorithm should be capable of both, but only a few VQA algorithms are tested in this manner. We compare five state-of-the-art VQA algorithms across eight VQA datasets covering both domains. To make the comparison fair, all of the models are standardized as much as possible, e.g., they use the same visual features, answer vocabularies, etc. We find that methods do not generalize across the two domains. To address this problem, we propose a new VQA algorithm that rivals or exceeds the state-of-the-art for both domains. Robik Shrestha, Kushal Kafle, Christopher Kanan |
CVPR | 3 |
| 2019 | Memory Efficient Experience Replay for Streaming LearningabstractIn supervised machine learning, an agent is typically trained once and then deployed. While this works well for static settings, robots often operate in changing environments and must quickly learn new things from data streams. In this paradigm, known as streaming learning, a learner is trained online, in a single pass, from a data stream that cannot be assumed to be independent and identically distributed (iid). Streaming learning will cause conventional deep neural networks (DNNs) to fail for two reasons: 1) they need multiple passes through the entire dataset; and 2) non-iid data will cause catastrophic forgetting. An old fix to both of these issues is rehearsal. To learn a new example, rehearsal mixes it with previous examples, and then this mixture is used to update the DNN. Full rehearsal is slow and memory intensive because it stores all previously observed examples, and its effectiveness for preventing catastrophic forgetting has not been studied in modern DNNs. Here, we describe the ExStream algorithm for memory efficient rehearsal and compare it to alternatives. We find that full rehearsal can eliminate catastrophic forgetting in a variety of streaming learning settings, with ExStream performing well using far less memory and computation. Tyler L. Hayes, Nathan D. Cahill, Christopher Kanan |
ICRA | 3 |
| 2019 | Continual lifelong learning with neural networks: A reviewabstractHumans and animals have the ability to continually acquire, fine-tune, and transfer knowledge and skills throughout their lifespan. This ability, referred to as lifelong learning, is mediated by a rich set of neurocognitive mechanisms that together contribute to the development and specialization of our sensorimotor skills as well as to long-term memory consolidation and retrieval. Consequently, lifelong learning capabilities are crucial for computational learning systems and autonomous agents interacting in the real world and processing continuous streams of information. However, lifelong learning remains a long-standing challenge for machine learning and neural network models since the continual acquisition of incrementally available information from non-stationary data distributions generally leads to catastrophic forgetting or interference. This limitation represents a major drawback for state-of-the-art deep neural network models that typically learn representations from stationary batches of training data, thus without accounting for situations in which information becomes incrementally available over time. In this review, we critically summarize the main challenges linked to lifelong learning for artificial learning systems and compare existing neural network approaches that alleviate, to different extents, catastrophic forgetting. Although significant advances have been made in domain-specific learning with neural networks, extensive research efforts are required for the development of robust lifelong learning on autonomous agents and robots. We discuss well-established and emerging research motivated by lifelong learning factors in biological systems such as structural plasticity, memory replay, curriculum and transfer learning, intrinsic motivation, and multisensory integration. German Ignacio Parisi, Ronald Kemker, Jose L. Part, Christopher Kanan, Stefan Wermter |
Neural Networks | 4 |
| 2018 | Measuring Catastrophic Forgetting in Neural NetworksabstractDeep neural networks are used in many state-of-the-art systems for machine perception. Once a network is trained to do a specific task, e.g., bird classification, it cannot easily be trained to do new tasks, e.g., incrementally learning to recognize additional bird species or learning an entirely different task such as flower recognition. When new tasks are added, typical deep neural networks are prone to catastrophically forgetting previous tasks. Networks that are capable of assimilating new information incrementally, much like how humans form new memories over time, will be more efficient than re-training the model from scratch each time a new task needs to be learned. There have been multiple attempts to develop schemes that mitigate catastrophic forgetting, but these methods have not been directly compared, the tests used to evaluate them vary considerably, and these methods have only been evaluated on small-scale problems (e.g., MNIST). In this paper, we introduce new metrics and benchmarks for directly comparing five different mechanisms designed to mitigate catastrophic forgetting in neural networks: regularization, ensembling, rehearsal, dual-memory, and sparse-coding. Our experiments on real-world images and sounds show that the mechanism(s) that are critical for optimal performance vary based on the incremental training paradigm and type of data being used, but they all demonstrate that the catastrophic forgetting problem is not yet solved. Ronald Kemker, Marc McClure, Angelina Abitino, Tyler L. Hayes, Christopher Kanan |
AAAI | 5 |
| 2018 | Characterizing the Temporal Dynamics of Information in Visually Guided Predictive Control Using LSTM Recurrent Neural Networks
Kamran Binaee, Anna Starynska, Jeff B. Pelz, Christopher Kanan, Gabriel J. Diaz |
CogSci | 4 |
| 2018 | DVQA: Understanding Data Visualizations via Question AnsweringabstractBar charts are an effective way to convey numeric information, but today's algorithms cannot parse them. Existing methods fail when faced with even minor variations in appearance. Here, we present DVQA, a dataset that tests many aspects of bar chart understanding in a question answering framework. Unlike visual question answering (VQA), DVQA requires processing words and answers that are unique to a particular bar chart. State-of-the-art VQA algorithms perform poorly on DVQA, and we propose two strong baselines that perform considerably better. Our work will enable algorithms to automatically extract numeric and semantic information from vast quantities of bar charts found in scientific publications, Internet articles, business reports, and many other areas. Kushal Kafle, Brian L. Price, Scott Cohen, Christopher Kanan |
CVPR | 4 |
| 2018 | FearNet: Brain-Inspired Model for Incremental Learning
Ronald Kemker, Christopher Kanan |
ICLR (Poster) | 2 |
| 2018 | Low-Shot Learning for the Semantic Segmentation of Remote Sensing ImageryabstractRecent advances in computer vision using deep learning with RGB imagery (e.g., object recognition and detection) have been made possible thanks to the development of large annotated RGB image data sets. In contrast, multispectral image (MSI) and hyperspectral image (HSI) data sets contain far fewer labeled images, in part due to the wide variety of sensors used. These annotations are especially limited for semantic segmentation, or pixelwise classification, of remote sensing imagery because it is labor intensive to generate image annotations. Low-shot learning algorithms can make effective inferences despite smaller amounts of annotated data. In this paper, we study low-shot learning using self-taught feature learning for semantic segmentation. We introduce: 1) an improved self-taught feature learning framework for HSI and MSI data and 2) a semisupervised classification algorithm. When these are combined, they achieve the state-of-the-art performance on remote sensing data sets that have little annotated training data available. These low-shot learning frameworks will reduce the manual image annotation burden and improve semantic segmentation performance for remote sensing imagery. Ronald Kemker, Ryan Luu, Christopher Kanan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | An Analysis of Visual Question Answering AlgorithmsabstractIn visual question answering (VQA), an algorithm must answer text-based questions about images. While multiple datasets for VQA have been created since late 2014, they all have flaws in both their content and the way algorithms are evaluated on them. As a result, evaluation scores are inflated and predominantly determined by answering easier questions, making it difficult to compare different methods. In this paper, we analyze existing VQA algorithms using a new dataset called the Task Driven Image Understanding Challenge (TDIUC), which has over 1.6 million questions organized into 12 different categories. We also introduce questions that are meaningless for a given image to force a VQA system to reason about image content. We propose new evaluation schemes that compensate for over-represented question-types and make it easier to study the strengths and weaknesses of algorithms. We analyze the performance of both baseline and state-of-the-art VQA models, including multi-modal compact bilinear pooling (MCB), neural module networks, and recurrent answering units. Our experiments establish how attention helps certain categories more than others, determine which models work better than others, and explain how simple models (e.g. MLP) can surpass more complex models (MCB) by simply learning to answer large, easy question categories. Kushal Kafle, Christopher Kanan |
ICCV | 2 |
| 2017 | Data Augmentation for Visual Question AnsweringabstractData augmentation is widely used to train deep neural networks for image classification tasks.Simply flipping images can help learning by increasing the number of training images by a factor of two.However, data augmentation in natural language processing is much less studied.Here, we describe two methods for data augmentation for Visual Question Answering (VQA).The first uses existing semantic annotations to generate new questions.The second method is a generative approach using recurrent neural networks.Experiments show the proposed schemes improve performance of baseline and state-of-the-art VQA algorithms. Kushal Kafle, Mohammed A. Yousefhussien, Christopher Kanan |
INLG | 3 |
| 2017 | Robotic grasp detection using deep convolutional neural networksabstractDeep learning has significantly advanced computer vision and natural language processing. While there have been some successes in robotics using deep learning, it has not been widely adopted. In this paper, we present a novel robotic grasp detection system that predicts the best grasping pose of a parallel-plate robotic gripper for novel objects using the RGB-D image of the scene. The proposed model uses a deep convolutional neural network to extract features from the scene and then uses a shallow convolutional neural network to predict the grasp configuration for the object of interest. Our multi-modal model achieved an accuracy of 89.21% on the standard Cornell Grasp Dataset and runs at real-time speeds. This redefines the state-of-the-art for robotic grasp detection. Sulabh Kumra, Christopher Kanan |
IROS | 2 |
| 2017 | Visual question answering: Datasets, algorithms, and future challenges
Kushal Kafle, Christopher Kanan |
Comput. Vis. Image Underst. | 2 |
| 2017 | Self-Taught Feature Learning for Hyperspectral Image ClassificationabstractIn this paper, we study self-taught learning for hyperspectral image (HSI) classification. Supervised deep learning methods are currently state of the art for many machine learning problems, but these methods require large quantities of labeled data to be effective. Unfortunately, existing labeled HSI benchmarks are too small to directly train a deep supervised network. Alternatively, we used self-taught learning, which is an unsupervised method to learn feature extracting frameworks from unlabeled hyperspectral imagery. These models learn how to extract generalizable features by training on sufficiently large quantities of unlabeled data that are distinct from the target data set. Once trained, these models can extract features from smaller labeled target data sets. We studied two self-taught learning frameworks for HSI classification. The first is a shallow approach that uses independent component analysis and the second is a three-layer stacked convolutional autoencoder. Our models are applied to the Indian Pines, Salinas Valley, and Pavia University data sets, which were captured by two separate sensors at different altitudes. Despite large variation in scene type, our algorithms achieve state-of-the-art results across all the three data sets. Ronald Kemker, Christopher Kanan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Automatic scanpath generation with deep recurrent neural networksabstractMany computer vision algorithms are biologically inspired and designed based on the human visual system. Convolutional neural networks (CNNs) are similarly inspired by the primary visual cortex in the human brain. However, the key difference between current visual models and the human visual system is how the visual information is gathered and processed. We make eye movements to collect information from the environment for navigation and task performance. We also make specific eye movements to important regions in the stimulus to perform the task-at-hand quickly and efficiently. Researchers have used expert scanpaths to train novices for improving the accuracy of visual search tasks. One of the limitations of such a system is that we need an expert to examine each visual stimuli beforehand to generate the scanpaths. In order to extend the idea of gaze guidance to a new unseen stimulus, there is a need for a computational model that can automatically generate expert-like scanpaths. We propose a model for automatic scanpath generation using a convolutional neural network (CNN) and long short-term memory (LSTM) modules. Our model uses LSTMs due to the temporal nature of eye movement data (scanpaths) where the system makes fixation predictions based on previous locations examined. Daniel Simon, Srinivas Sridharan 0001, Shagan Sah, Raymond W. Ptucha, Christopher Kanan, Reynold J. Bailey |
SAP | 5 |
| 2016 | Answer-Type Prediction for Visual Question AnsweringabstractRecently, algorithms for object recognition and related tasks have become sufficiently proficient that new vision tasks can now be pursued. In this paper, we build a system capable of answering open-ended text-based questions about images, which is known as Visual Question Answering (VQA). Our approach's key insight is that we can predict the form of the answer from the question. We formulate our solution in a Bayesian framework. When our approach is combined with a discriminative model, the combined model achieves state-of-the-art results on four benchmark datasets for open-ended VQA: DAQUAR, COCO-QA, The VQA Dataset, and Visual7W. Kushal Kafle, Christopher Kanan |
CVPR | 2 |
| 2016 | Online tracking using saliencyabstractWhen tracking small moving objects, primates use smooth pursuit eye movements to keep a target in the center of the field of view. In this paper, we propose the Smooth Pursuit tracking algorithm, which uses three kinds of saliency maps to perform online target tracking: appearance, location, and motion. In addition to tracking single targets, our method can track multiple targets with little additional overhead. The appearance saliency map uses deep convolutional neural network features along with gnostic fields, a brain-inspired model for object recognition. The location saliency map predicts where the object will move next. Finally, the motion saliency map indicates which objects are moving in the scene. We combine all three saliency maps into a smooth pursuit map, which is used to generate bounding boxes for tracked objects. We evaluate our algorithm and others from the literature on a vehicle tracking task. Our approach achieves the best overall performance, including being the only method we tested capable of handling long-term occlusions. Mohammed A. Yousefhussien, N. Andrew Browning, Christopher Kanan |
WACV | 3 |
| 2015 | Modeling the Object Recognition Pathway: A Deep Hierarchical Model Using Gnostic Fields
Panqu Wang, Garrison W. Cottrell, Christopher Kanan |
CogSci | 3 |
| 2014 | Predicting an observer's task using multi-fixation pattern analysisabstractSince Yarbus's seminal work in 1965, vision scientists have argued that people's eye movement patterns differ depending upon their task. This suggests that we may be able to infer a person's task (or mental state) from their eye movements alone. Recently, this was attempted by Greene et al. [2012] in a Yarbus-like replication study; however, they were unable to successfully predict the task given to their observer. We reanalyze their data, and show that by using more powerful algorithms it is possible to predict the observer's task. We also used our algorithms to infer the image being viewed by an observer and their identity. More generally, we show how off-the-shelf algorithms from machine learning can be used to make inferences from an observer's eye movements, using an approach we call Multi-Fixation Pattern Analysis (MFPA). Christopher Kanan, Nicholas A. Ray, Dina N. F. Bseiso, Janet Hui-wen Hsiao, Garrison W. Cottrell |
ETRA | 1 |
| 2014 | Fine-grained object recognition with Gnostic FieldsabstractMuch object recognition research is concerned with basic-level classification, in which objects differ greatly in visual shape and appearance, e.g., desk vs duck. In contrast, fine-grained classification involves recognizing objects at a subordinate level, e.g., Wood duck vs Mallard duck. At the basic-level objects tend to differ greatly in shape and appearance, but these differences are usually much more subtle at the subordinate level, making fine-grained classification especially challenging. In this work, we show that Gnostic Fields, a brain-inspired model of object categorization, excel at fine-grained recognition. Gnostic Fields exceeded state-of-the-art methods on benchmark bird classification and dog breed recognition datasets, achieving a relative improvement on the Caltech-UCSD Bird-200 (CUB-200) dataset of 30.5% over the state-of-the-art and a 25.5% relative improvement on the Stanford Dogs dataset. We also demonstrate that Gnostic Fields can be sped up, enabling real-time classification in less than 70 ms per image. Christopher Kanan |
WACV | 1 |
| 2010 | Robust classification of objects, faces, and flowers using natural image statisticsabstractClassification of images in many category datasbets has rapidly improved in recent years. However, systems that perform well on particular datasets typically have one or more limitations such as a failure to generalize across visual tasks (e.g., requiring a face detector or extensive retuning of parameters), insufficient translation invariance, inability to cope with partial views and occlusion, or significant performance degradation as the number of classes is increased. Here we attempt to overcome these challenges using a model that combines sequential visual attention using fixations with sparse coding. The model's biologically-inspired filters are acquired using unsupervised learning applied to natural image patches. Using only a single feature type, our approach achieves 78.5% accuracy on Caltech-101 and 75.2% on the 102 Flowers dataset when trained on 30 instances per class and it achieves 92.7% accuracy on the AR Face database with 1 training instance per person. The same features and parameters are used across these datasets to illustrate its robust performance. Christopher Kanan, Garrison W. Cottrell |
CVPR | 1 |