Amlan Kar

dblp:190/7073 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-4540-6007ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
17 papers
Segmentation and scene understanding · 19% Generative modeling · 13% 3D vision · 12%
Computer graphics and multimedia
7 papers
Visual content generation and editing · 47% Geometric modeling and processing · 22% Image and video processing · 12%

Topics — the 30 heaviest of 55, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model reasoning
inference-time reasoning
0.912025
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions · EMNLP 2025
Natural language and speech › Language models and text generation › test-time scaling
inference-time search
0.912025
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions · EMNLP 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.912025
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions · EMNLP 2025
Computer vision › Vision and language
visual reasoning
0.912025
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions · EMNLP 2025
Machine learning › Generative modeling
3d generative model
0.812024
Outdoor Scene Extrapolation with Hierarchical Generative Cellular Automata · CVPR 2024
Computer vision › 3D vision › 3d generation
3d scene generation
0.812024
Outdoor Scene Extrapolation with Hierarchical Generative Cellular Automata · CVPR 2024
Computer vision › Segmentation and scene understanding
instance segmentation
0.722019
Fast Interactive Object Annotation With Curve-GCN · CVPR 2019
Efficient Interactive Annotation of Segmentation Datasets With Polygon-RNN++ · CVPR 2018
Machine learning › Generative modeling
diffusion model
0.712023
DreamTeacher: Pretraining Image Backbones with Deep Generative Models · ICCV 2023
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.712023
DreamTeacher: Pretraining Image Backbones with Deep Generative Models · ICCV 2023
Computer vision › 3D vision
hand-object interaction
0.612022
EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations · NeurIPS 2022
Computer vision › Video understanding and tracking
video object segmentation
0.612022
EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations · NeurIPS 2022
Natural language and speech › Information extraction and text analysis
data annotation
0.512021
Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets · CVPR 2021
Computer vision › Image recognition and object detection
image classification
0.512021
Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets · CVPR 2021
Machine learning › Efficient and distributed learning › data-efficient learning
label-efficient learning
0.512021
Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets · CVPR 2021
Machine learning › Learning paradigms
semi-supervised learning
0.512021
Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets · CVPR 2021
Visual content generation and editing
3d scene generation
0.512021
ATISS: Autoregressive Transformers for Indoor Scene Synthesis · NeurIPS 2021
Visual content generation and editing › 3d scene generation
indoor scene synthesis
0.512021
ATISS: Autoregressive Transformers for Indoor Scene Synthesis · NeurIPS 2021
Computer vision › Segmentation and scene understanding
semantic segmentation
0.522019
Creative Flow+ Dataset · CVPR 2019
Devil Is in the Edges: Learning Semantic Boundaries From Noisy Annotations · CVPR 2019
Computer vision › 3D vision
3d object detection
0.412020
Learning to Evaluate Perception Models Using Planner-Centric Metrics · CVPR 2020
Robotics › Autonomous driving › perception › perception systems
perception evaluation
0.412020
Learning to Evaluate Perception Models Using Planner-Centric Metrics · CVPR 2020
Machine learning › Generative modeling
synthetic data generation
0.412020
Meta-Sim2: Unsupervised Learning of Scene Structure for Synthetic Data Generation · ECCV (17) 2020
Machine learning › Learning paradigms
unsupervised learning
0.412020
Meta-Sim2: Unsupervised Learning of Scene Structure for Synthetic Data Generation · ECCV (17) 2020
Image and video processing › color image processing
color representation
0.412020
Nonlinear color triads for approximation, learning and direct manipulation of color distributions · ACM Trans. Graph. 2020
Machine learning › Efficient and distributed learning
active learning
0.412019
Learning to Caption Images Through a Lifetime by Asking Questions · ICCV 2019
Computer vision › Segmentation and scene understanding
boundary detection
0.412019
Devil Is in the Edges: Learning Semantic Boundaries From Noisy Annotations · CVPR 2019
Computer vision › Vision and language
image captioning
0.412019
Learning to Caption Images Through a Lifetime by Asking Questions · ICCV 2019
Computer vision › Segmentation and scene understanding
interactive segmentation
0.412019
Object Instance Annotation With Deep Extreme Level Set Evolution · CVPR 2019
Computer vision › Segmentation and scene understanding › image segmentation
level set segmentation
0.412019
Object Instance Annotation With Deep Extreme Level Set Evolution · CVPR 2019
Computer vision › Segmentation and scene understanding
object segmentation
0.412019
Object Instance Annotation With Deep Extreme Level Set Evolution · CVPR 2019
Computer vision › 3D vision › motion estimation
optical flow
0.412019
Creative Flow+ Dataset · CVPR 2019

Methods — techniques the papers use, named apart from their topics

neural network · 1.5sub-question decomposition · 0.9monte carlo tree search · 0.9generative cellular automata · 0.8coarse-to-fine generation · 0.8knowledge distillation · 0.7generative model · 0.7diffusion model · 0.7convolutional neural network · 0.6annotation pipeline · 0.6self-supervised learning · 0.5online joint probabilistic modeling · 0.5human annotation simulation · 0.5autoregressive transformer · 0.5deep learning · 0.4uncertainty reasoning · 0.4spline-based annotation · 0.4sequential generative model · 0.4
YearPublicationVenuePosition
2025 Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
abstract
Recent research in vision-language models (VLMs) has centered around the possibility of equipping them with implicit long-form chainof-thought reasoning-akin to the success observed in language models-via distillation and reinforcement learning.But what about the non-reasoning models already trained and deployed across the internet?Should we simply abandon them, or is there hope for a search mechanism that can elicit hidden knowledge and induce long reasoning traces-without any additional training or supervision?In this paper, we explore this possibility using a Monte Carlo Tree Search (MCTS)-inspired algorithm, which injects subquestion-subanswer pairs into the model's output stream.We show that framing reasoning as a search process-where subquestions act as latent decisions within a broader inference trajectory-helps the model "connect the dots" between fragmented knowledge and produce extended reasoning traces in nonreasoning models.We evaluate our method across three benchmarks and observe consistent improvements.Notably, our approach yields a 2% overall improvement on MMMU-PRO, including a significant 9% gain in Liberal Arts.Problem (P) What is the position of the blue car in the image?A: parked on the sidewalk B: driving on the road C: parked on the grass D: parked on the road CoT: To determine the position of the blue car in the image, let's analyze the scene: * The image shows a busy urban street with several vehicles.* The blue car is located on the right side of the image.* It is positioned on what appears to be a paved surface, which is consistent with a road.* The car is not on the grass or the sidewalk, as those areas are clearly distinguishable in the image.* The car is stationary, suggesting it is parked.Given these observations, the blue car is parked on the road.
David Acuna, Ximing Lu, Jaehun Jung, Hyunwoo Kim 0002, Amlan Kar, Sanja Fidler, Yejin Choi 0001
EMNLP5
2024 Outdoor Scene Extrapolation with Hierarchical Generative Cellular Automata
abstract
We aim to generate fine-grained 3D geometry from large-scale sparse LiDAR scans, abundantly captured by autonomous vehicles (AV). Contrary to prior work on AV scene completion, we aim to extrapolate fine geometry from unlabeled and beyond spatial limits of LiDAR scans, taking a step towards generating realistic, high-resolution simulation-ready 3D street environments. We propose hierarchical Generative Cellular Automata (hGCA), a spatially scalable conditional 3D generative model, which grows geometry recursively with local kernels following [46, 47], in a coarse-to-fine manner, equipped with a light-weight planner to induce global consistency. Experiments on synthetic scenes show that hGCA generates plausible scene geometry with higher fidelity and completeness compared to state-of-the-art baselines. Our model generalizes strongly from sim-to-real, qualitatively outperforming baselines on the Waymo-open dataset. We also show anecdotal evidence of the ability to create novel objects from real-world geometric cues even when trained on limited synthetic content. More results and details can be found on our project page.
Dongsu Zhang, Francis Williams, Zan Gojcic, Karsten Kreis, Sanja Fidler, Young Min Kim 0001, Amlan Kar
CVPR7
2023 DreamTeacher: Pretraining Image Backbones with Deep Generative Models
abstract
In this work, we introduce a self-supervised feature representation learning framework DreamTeacher that utilizes generative networks for pre-training downstream image backbones. We propose to distill knowledge from a trained generative model into standard image backbones that have been well engineered for specific perception tasks. We investigate two types of knowledge distillation: 1) distilling learned generative features onto target image backbones as an alternative to pretraining these backbones on large labeled datasets such as ImageNet, and 2) distilling labels obtained from generative networks with task heads onto logits of target backbones. We perform extensive analyses on multiple generative models, dense prediction benchmarks, and several pre-training regimes. We empirically find that our DreamTeacher significantly outperforms existing self-supervised representation learning approaches across the board. Unsupervised ImageNet pre-training with DreamTeacher leads to significant improvements over ImageNet classification pre-training on downstream datasets, showcasing generative models, and diffusion generative models specifically, as a promising approach to representation learning on large, diverse datasets without requiring manual annotation.
Daiqing Li, Huan Ling, Amlan Kar, David Acuna, Seung Wook Kim 0001, Karsten Kreis, Antonio Torralba 0001, Sanja Fidler
ICCV3
2022 EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations
abstract
We introduce VISOR, a new dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video. VISOR annotates videos from EPIC-KITCHENS, which comes with a new set of challenges not encountered in current video segmentation datasets. Specifically, we need to ensure both short- and long-term consistency of pixel-level annotations as objects undergo transformative interactions, e.g. an onion is peeled, diced and cooked - where we aim to obtain accurate pixel-level annotations of the peel, onion pieces, chopping board, knife, pan, as well as the acting hands. VISOR introduces an annotation pipeline, AI-powered in parts, for scalability and quality. In total, we publicly release 272K manual semantic masks of 257 object classes, 9.9M interpolated dense masks, 67K hand-object relations, covering 36 hours of 179 untrimmed videos. Along with the annotations, we introduce three challenges in video object segmentation, interaction understanding and long-term reasoning.For data, code and leaderboards: http://epic-kitchens.github.io/VISOR
Ahmad Darkhalil, Dandan Shan, Bin Zhu 0006, Amlan Kar, Richard E. L. Higgins, Sanja Fidler, David F. Fouhey, Dima Damen
NeurIPS5
2021 Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets
abstract
Data is the engine of modern computer vision, which necessitates collecting large-scale datasets. This is expensive, and guaranteeing the quality of the labels is a major challenge. In this paper, we investigate efficient annotation strategies for collecting multi-class classification labels for a large collection of images. While methods that exploit learnt models for labeling exist, a surprisingly prevalent approach is to query humans for a fixed number of labels per datum and aggregate them, which is expensive. Building on prior work on online joint probabilistic modeling of human annotations and machine-generated beliefs, we propose modifications and best practices aimed at minimizing human labeling effort. Specifically, we make use of advances in self-supervised learning, view annotation as a semi-supervised learning problem, identify and mitigate pitfalls and ablate several key design choices to propose effective guidelines for labeling. Our analysis is done in a more realistic simulation that involves querying human la-belers, which uncovers issues with evaluation using existing worker simulation methods. Simulated experiments on a 125k image subset of the ImageNet100 show that it can be annotated to 80% top-1 accuracy with 0.35 annotations per image on average, a 2.7x and 6.7x improvement over prior work and manual annotation, respectively.1
Yuan-Hong Liao, Amlan Kar, Sanja Fidler
CVPR2
2021 ATISS: Autoregressive Transformers for Indoor Scene Synthesis
abstract
The ability to synthesize realistic and diverse indoor furniture layouts automatically or based on partial input, unlocks many applications, from better interactive 3D tools to data synthesis for training and simulation. In this paper, we present ATISS, a novel autoregressive transformer architecture for creating diverse and plausible synthetic indoor environments, given only the room type and its floor plan. In contrast to prior work, which poses scene synthesis as sequence generation, our model generates rooms as unordered sets of objects. We argue that this formulation is more natural, as it makes ATISS generally useful beyond fully automatic room layout synthesis. For example, the same trained model can be used in interactive applications for general scene completion, partial room re-arrangement with any objects specified by the user, as well as object suggestions for any partial room. To enable this, our model leverages the permutation equivariance of the transformer when conditioning on the partial scene, and is trained to be permutation-invariant across object orderings. Our model is trained end-to-end as an autoregressive generative model using only labeled 3D bounding boxes as supervision. Evaluations on four room types in the 3D-FRONT dataset demonstrate that our model consistently generates plausible room layouts that are more realistic than existing methods.In addition, it has fewer parameters, is simpler to implement and train and runs up to 8 times faster than existing methods.
Despoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis, Andreas Geiger 0001, Sanja Fidler
NeurIPS2
2020 Learning to Evaluate Perception Models Using Planner-Centric Metrics
abstract
Variants of accuracy and precision are the gold-standard by which the computer vision community measures progress of perception algorithms. One reason for the ubiquity of these metrics is that they are largely task-agnostic; we in general seek to detect zero false negatives or positives. The downside of these metrics is that, at worst, they penalize all incorrect detections equally without conditioning on the task or scene, and at best, heuristics need to be chosen to ensure that different mistakes count differently. In this paper, we propose a principled metric for 3D object detection specifically for the task of self-driving. The core idea behind our metric is to isolate the task of object detection and measure the impact the produced detections would induce on the downstream task of driving. Without hand-designing it to, we find that our metric penalizes many of the mistakes that other metrics penalize by design. In addition, our metric downweighs detections based on additional factors such as distance from a detection to the ego car and the speed of the detection in intuitive ways that other detection metrics do not. For human evaluation, we generate scenes in which standard metrics and our metric disagree and find that humans side with our metric 79% of the time. Our project page including an evaluation server can be found at https://nv-tlabs.github.io/detection-relevance.
Jonah Philion, Amlan Kar, Sanja Fidler
CVPR2
2020 Meta-Sim2: Unsupervised Learning of Scene Structure for Synthetic Data Generation
Jeevan Devaranjan, Amlan Kar, Sanja Fidler
ECCV (17)2
2020 Interactive Annotation of 3D Object Geometry Using 2D Scribbles
Tianchang Shen, Jun Gao 0004, Amlan Kar, Sanja Fidler
ECCV (17)3
2020 Federated Simulation for Medical Imaging
Daiqing Li, Amlan Kar, Nishant Ravikumar, Alejandro F. Frangi, Sanja Fidler
MICCAI (1)2
2020 Nonlinear color triads for approximation, learning and direct manipulation of color distributions
abstract
We present nonlinear color triads, an extension of color gradients able to approximate a variety of natural color distributions that have no standard interactive representation. We derive a method to fit this compact parametric representation to existing images and show its power for tasks such as image editing and compression. Our color triad formulation can also be included in standard deep learning architectures, facilitating further research.
Maria Shugrina, Amlan Kar, Sanja Fidler, Karan Singh 0004
ACM Trans. Graph.2
2019 Devil Is in the Edges: Learning Semantic Boundaries From Noisy Annotations
abstract
We tackle the problem of semantic boundary prediction, which aims to identify pixels that belong to object(class) boundaries. We notice that relevant datasets consist of a significant level of label noise, reflecting the fact that precise annotations are laborious to get and thus annotators trade-off quality with efficiency. We aim to learn sharp and precise semantic boundaries by explicitly reasoning about annotation noise during training. We propose a simple new layer and loss that can be used with existing learning-based boundary detectors. Our layer/loss enforces the detector to predict a maximum response along the normal direction at an edge, while also regularizing its direction. We further reason about true object boundaries during training using a level set formulation, which allows the network to learn from misaligned labels in an end-to-end fashion. Experiments show that we improve over the CASENet backbone network by more than 4% in terms of MF(ODS) and 18.61% in terms of AP, outperforming all current state-of-the-art methods including those that deal with alignment. Furthermore, we show that our learned network can be used to significantly improve coarse segmentation labels, lending itself as an efficient way to label new data.
David Acuna, Amlan Kar, Sanja Fidler
CVPR2
2019 Fast Interactive Object Annotation With Curve-GCN
abstract
Manually labeling objects by tracing their boundaries is a laborious process. In Polygon-RNN++, the authors proposed Polygon-RNN that produces polygonal annotations in a recurrent manner using a CNN-RNN architecture, allowing interactive correction via humans-in-the-loop. We propose a new framework that alleviates the sequential nature of Polygon-RNN, by predicting all vertices simultaneously using a Graph Convolutional Network (GCN). Our model is trained end-to-end, and runs in real time. It supports object annotation by either polygons or splines, facilitating labeling efficiency for both line-based and curved objects. We show that Curve-GCN outperforms all existing approaches in automatic mode, including the powerful DeepLab, and is significantly more efficient in interactive mode than Polygon-RNN++. Our model runs at 29.3ms in automatic, and 2.6ms in interactive mode, making it 10x and 100x faster than Polygon-RNN++.
Huan Ling, Jun Gao 0004, Amlan Kar, Wenzheng Chen, Sanja Fidler
CVPR3
2019 Creative Flow+ Dataset
abstract
We present the Creative Flow+ Dataset, the first diverse multi-style artistic video dataset richly labeled with per-pixel optical flow, occlusions, correspondences, segmentation labels, normals, and depth. Our dataset includes 3000 animated sequences rendered using styles randomly selected from 40 textured line styles and 38 shading styles, spanning the range between flat cartoon fill and wildly sketchy shading. Our dataset includes 124K+ train set frames and 10K test set frames rendered at 1500x1500 resolution, far surpassing the largest available optical flow datasets in size. While modern techniques for tasks such as optical flow estimation achieve impressive performance on realistic images and video, today there is no way to gauge their performance on non-photorealistic images. Creative Flow+ poses a new challenge to generalize real-world Computer Vision to messy stylized content. We show that learning-based optical flow methods fail to generalize to this data and struggle to compete with classical approaches, and invite new research in this area. Our dataset and a new optical flow benchmark will be publicly available at: www.cs.toronto.edu/creativeflow/. We further release the complete dataset creation pipeline, allowing the community to generate and stylize their own data on demand.
Maria Shugrina, Ziheng Liang, Amlan Kar, Jiaman Li, Angad Singh, Karan Singh 0004, Sanja Fidler
CVPR3
2019 Object Instance Annotation With Deep Extreme Level Set Evolution
abstract
In this paper, we tackle the task of interactive object segmentation. We revive the old ideas on level set segmentation which framed object annotation as curve evolution. Carefully designed energy functions ensured that the curve was well aligned with image boundaries, and generally "well behaved". The Level Set Method can handle objects with complex shapes and topological changes such as merging and splitting, thus able to deal with occluded objects and objects with holes. We propose Deep Extreme Level Set Evolution that combines powerful CNN models with level set optimization in an end-to-end fashion. Our method learns to predict evolution parameters conditioned on the image and evolves the predicted initial contour to produce the final result. We make our model interactive by incorporating user clicks on the extreme boundary points, following DEXTR. We show that our approach significantly outperforms DEXTR on the static Cityscapes dataset and the video segmentation benchmark DAVIS, and performs on par on PASCAL and SBD.
David Acuna, Huan Ling, Amlan Kar, Sanja Fidler
CVPR4
2019 Neural Turtle Graphics for Modeling City Road Layouts
abstract
We propose Neural Turtle Graphics (NTG), a novel generative model for spatial graphs, and demonstrate its applications in modeling city road layouts. Specifically, we represent the road layout using a graph where nodes in the graph represent control points and edges in the graph represents road segments. NTG is a sequential generative model parameterized by a neural network. It iteratively generates a new node and an edge connecting to an existing node conditioned on the current graph. We train NTG on Open Street Map data and show it outperforms existing approaches using a set of diverse performance metrics. Moreover, our method allows users to control styles of generated road layouts mimicking existing cities as well as to sketch a part of the city road layout to be synthesized. In addition to synthesis, the proposed NTG finds uses in an analytical task of aerial road parsing. Experimental results show that it achieves state-of-the-art performance on the SpaceNet dataset.
Hang Chu, Daiqing Li, David Acuna, Amlan Kar, Maria Shugrina, Xinkai Wei, Ming-Yu Liu 0001, Antonio Torralba 0001, Sanja Fidler
ICCV4
2019 Meta-Sim: Learning to Generate Synthetic Datasets
abstract
Training models to high-end performance requires availability of large labeled datasets, which are expensive to get. The goal of our work is to automatically synthesize labeled datasets that are relevant for a downstream task. We propose Meta-Sim, which learns a generative model of synthetic scenes, and obtain images as well as its corresponding ground-truth via a graphics engine. We parametrize our dataset generator with a neural network, which learns to modify attributes of scene graphs obtained from probabilistic scene grammars, so as to minimize the distribution gap between its rendered outputs and target data. If the real dataset comes with a small labeled validation set, we additionally aim to optimize a meta-objective, i.e. downstream task performance. Experiments show that the proposed method can greatly improve content generation quality over a human-engineered probabilistic scene grammar, both qualitatively and quantitatively as measured by performance on a downstream task.
Amlan Kar, Aayush Prakash, Ming-Yu Liu 0001, Eric Cameracci, Justin Yuan, Matt Rusiniak, David Acuna, Antonio Torralba 0001, Sanja Fidler
ICCV1
2019 Learning to Caption Images Through a Lifetime by Asking Questions
abstract
In order to bring artificial agents into our lives, we will need to go beyond supervised learning on closed datasets to having the ability to continuously expand knowledge. Inspired by a student learning in a classroom, we present an agent that can continuously learn by posing natural language questions to humans. Our agent is composed of three interacting modules, one that performs captioning, another that generates questions and a decision maker that learns when to ask questions by implicitly reasoning about the uncertainty of the agent and expertise of the teacher. As compared to current active learning methods which query images for full captions, our agent is able to ask pointed questions to improve the generated captions. The agent trains on the improved captions, expanding its knowledge. We show that our approach achieves better performance using less human supervision than the baselines on the challenging MSCOCO dataset.
Tingke Shen, Amlan Kar, Sanja Fidler
ICCV2
2018 Efficient Interactive Annotation of Segmentation Datasets With Polygon-RNN++
abstract
Manually labeling datasets with object masks is extremely time consuming. In this work, we follow the idea of Polygon-RNN [4] to produce polygonal annotations of objects interactively using humans-in-the-loop. We introduce several important improvements to the model: 1) we design a new CNN encoder architecture, 2) show how to effectively train the model with Reinforcement Learning, and 3) significantly increase the output resolution using a Graph Neural Network, allowing the model to accurately annotate high-resolution objects in images. Extensive evaluation on the Cityscapes dataset [8] shows that our model, which we refer to as Polygon-RNN++, significantly outperforms the original model in both automatic (10% absolute and 16% relative improvement in mean IoU) and interactive modes (requiring 50% fewer clicks by annotators). We further analyze the cross-domain scenario in which our model is trained on one dataset, and used out of the box on datasets from varying domains. The results show that Polygon-RNN++ exhibits powerful generalization capabilities, achieving significant improvements over existing pixel-wise methods. Using simple online fine-tuning we further achieve a high reduction in annotation time for new datasets, moving a step closer towards an interactive annotation tool to be used in practice.
David Acuna, Huan Ling, Amlan Kar, Sanja Fidler
CVPR3
2017 AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos
abstract
We propose a novel method for temporally pooling frames in a video for the task of human action recognition. The method is motivated by the observation that there are only a small number of frames which, together, contain sufficient information to discriminate an action class present in a video, from the rest. The proposed method learns to pool such discriminative and informative frames, while discarding a majority of the non-informative frames in a single temporal scan of the video. Our algorithm does so by continuously predicting the discriminative importance of each video frame and subsequently pooling them in a deep learning framework. We show the effectiveness of our proposed pooling method on standard benchmarks where it consistently improves on baseline pooling methods, with both RGB and optical flow based Convolutional networks. Further, in combination with complementary video representations, we show results that are competitive with respect to the state-of-the-art results on two challenging and publicly available benchmark datasets.
Amlan Kar, Nishant Rai, Karan Sikka, Gaurav Sharma 0004
CVPR1