VLDB 2026 Research / reviewers in the wild / expert
Amlan Kar
dblp:190/7073
· DBLP profile ↗
20ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-4540-6007ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
17 papers |
Segmentation and scene understanding · 19% Generative modeling · 13% 3D vision · 12% | |
| Computer graphics and multimedia
7 papers |
Visual content generation and editing · 47% Geometric modeling and processing · 22% Image and video processing · 12% |
Topics — the 30 heaviest of 55, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model reasoning
inference-time reasoning |
0.9 | 1 | 2025 | Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions · EMNLP 2025 |
Natural language and speech › Language models and text generation › test-time scaling
inference-time search |
0.9 | 1 | 2025 | Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions · EMNLP 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search |
0.9 | 1 | 2025 | Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions · EMNLP 2025 |
Computer vision › Vision and language
visual reasoning |
0.9 | 1 | 2025 | Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions · EMNLP 2025 |
Machine learning › Generative modeling
3d generative model |
0.8 | 1 | 2024 | Outdoor Scene Extrapolation with Hierarchical Generative Cellular Automata · CVPR 2024 |
Computer vision › 3D vision › 3d generation
3d scene generation |
0.8 | 1 | 2024 | Outdoor Scene Extrapolation with Hierarchical Generative Cellular Automata · CVPR 2024 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.7 | 2 | 2019 | Fast Interactive Object Annotation With Curve-GCN · CVPR 2019 Efficient Interactive Annotation of Segmentation Datasets With Polygon-RNN++ · CVPR 2018 |
Machine learning › Generative modeling
diffusion model |
0.7 | 1 | 2023 | DreamTeacher: Pretraining Image Backbones with Deep Generative Models · ICCV 2023 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.7 | 1 | 2023 | DreamTeacher: Pretraining Image Backbones with Deep Generative Models · ICCV 2023 |
Computer vision › 3D vision
hand-object interaction |
0.6 | 1 | 2022 | EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations · NeurIPS 2022 |
Computer vision › Video understanding and tracking
video object segmentation |
0.6 | 1 | 2022 | EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations · NeurIPS 2022 |
Natural language and speech › Information extraction and text analysis
data annotation |
0.5 | 1 | 2021 | Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets · CVPR 2021 |
Computer vision › Image recognition and object detection
image classification |
0.5 | 1 | 2021 | Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets · CVPR 2021 |
Machine learning › Efficient and distributed learning › data-efficient learning
label-efficient learning |
0.5 | 1 | 2021 | Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets · CVPR 2021 |
Machine learning › Learning paradigms
semi-supervised learning |
0.5 | 1 | 2021 | Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets · CVPR 2021 |
Visual content generation and editing
3d scene generation |
0.5 | 1 | 2021 | ATISS: Autoregressive Transformers for Indoor Scene Synthesis · NeurIPS 2021 |
Visual content generation and editing › 3d scene generation
indoor scene synthesis |
0.5 | 1 | 2021 | ATISS: Autoregressive Transformers for Indoor Scene Synthesis · NeurIPS 2021 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.5 | 2 | 2019 | Creative Flow+ Dataset · CVPR 2019 Devil Is in the Edges: Learning Semantic Boundaries From Noisy Annotations · CVPR 2019 |
Computer vision › 3D vision
3d object detection |
0.4 | 1 | 2020 | Learning to Evaluate Perception Models Using Planner-Centric Metrics · CVPR 2020 |
Robotics › Autonomous driving › perception › perception systems
perception evaluation |
0.4 | 1 | 2020 | Learning to Evaluate Perception Models Using Planner-Centric Metrics · CVPR 2020 |
Machine learning › Generative modeling
synthetic data generation |
0.4 | 1 | 2020 | Meta-Sim2: Unsupervised Learning of Scene Structure for Synthetic Data Generation · ECCV (17) 2020 |
Machine learning › Learning paradigms
unsupervised learning |
0.4 | 1 | 2020 | Meta-Sim2: Unsupervised Learning of Scene Structure for Synthetic Data Generation · ECCV (17) 2020 |
Image and video processing › color image processing
color representation |
0.4 | 1 | 2020 | Nonlinear color triads for approximation, learning and direct manipulation of color distributions · ACM Trans. Graph. 2020 |
Machine learning › Efficient and distributed learning
active learning |
0.4 | 1 | 2019 | Learning to Caption Images Through a Lifetime by Asking Questions · ICCV 2019 |
Computer vision › Segmentation and scene understanding
boundary detection |
0.4 | 1 | 2019 | Devil Is in the Edges: Learning Semantic Boundaries From Noisy Annotations · CVPR 2019 |
Computer vision › Vision and language
image captioning |
0.4 | 1 | 2019 | Learning to Caption Images Through a Lifetime by Asking Questions · ICCV 2019 |
Computer vision › Segmentation and scene understanding
interactive segmentation |
0.4 | 1 | 2019 | Object Instance Annotation With Deep Extreme Level Set Evolution · CVPR 2019 |
Computer vision › Segmentation and scene understanding › image segmentation
level set segmentation |
0.4 | 1 | 2019 | Object Instance Annotation With Deep Extreme Level Set Evolution · CVPR 2019 |
Computer vision › Segmentation and scene understanding
object segmentation |
0.4 | 1 | 2019 | Object Instance Annotation With Deep Extreme Level Set Evolution · CVPR 2019 |
Computer vision › 3D vision › motion estimation
optical flow |
0.4 | 1 | 2019 | Creative Flow+ Dataset · CVPR 2019 |
Methods — techniques the papers use, named apart from their topics
neural network · 1.5sub-question decomposition · 0.9monte carlo tree search · 0.9generative cellular automata · 0.8coarse-to-fine generation · 0.8knowledge distillation · 0.7generative model · 0.7diffusion model · 0.7convolutional neural network · 0.6annotation pipeline · 0.6self-supervised learning · 0.5online joint probabilistic modeling · 0.5human annotation simulation · 0.5autoregressive transformer · 0.5deep learning · 0.4uncertainty reasoning · 0.4spline-based annotation · 0.4sequential generative model · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right QuestionsabstractRecent research in vision-language models (VLMs) has centered around the possibility of equipping them with implicit long-form chainof-thought reasoning-akin to the success observed in language models-via distillation and reinforcement learning.But what about the non-reasoning models already trained and deployed across the internet?Should we simply abandon them, or is there hope for a search mechanism that can elicit hidden knowledge and induce long reasoning traces-without any additional training or supervision?In this paper, we explore this possibility using a Monte Carlo Tree Search (MCTS)-inspired algorithm, which injects subquestion-subanswer pairs into the model's output stream.We show that framing reasoning as a search process-where subquestions act as latent decisions within a broader inference trajectory-helps the model "connect the dots" between fragmented knowledge and produce extended reasoning traces in nonreasoning models.We evaluate our method across three benchmarks and observe consistent improvements.Notably, our approach yields a 2% overall improvement on MMMU-PRO, including a significant 9% gain in Liberal Arts.Problem (P) What is the position of the blue car in the image?A: parked on the sidewalk B: driving on the road C: parked on the grass D: parked on the road CoT: To determine the position of the blue car in the image, let's analyze the scene: * The image shows a busy urban street with several vehicles.* The blue car is located on the right side of the image.* It is positioned on what appears to be a paved surface, which is consistent with a road.* The car is not on the grass or the sidewalk, as those areas are clearly distinguishable in the image.* The car is stationary, suggesting it is parked.Given these observations, the blue car is parked on the road. David Acuna, Ximing Lu, Jaehun Jung, Hyunwoo Kim 0002, Amlan Kar, Sanja Fidler, Yejin Choi 0001 |
EMNLP | 5 |
| 2024 | Outdoor Scene Extrapolation with Hierarchical Generative Cellular AutomataabstractWe aim to generate fine-grained 3D geometry from large-scale sparse LiDAR scans, abundantly captured by autonomous vehicles (AV). Contrary to prior work on AV scene completion, we aim to extrapolate fine geometry from unlabeled and beyond spatial limits of LiDAR scans, taking a step towards generating realistic, high-resolution simulation-ready 3D street environments. We propose hierarchical Generative Cellular Automata (hGCA), a spatially scalable conditional 3D generative model, which grows geometry recursively with local kernels following [46, 47], in a coarse-to-fine manner, equipped with a light-weight planner to induce global consistency. Experiments on synthetic scenes show that hGCA generates plausible scene geometry with higher fidelity and completeness compared to state-of-the-art baselines. Our model generalizes strongly from sim-to-real, qualitatively outperforming baselines on the Waymo-open dataset. We also show anecdotal evidence of the ability to create novel objects from real-world geometric cues even when trained on limited synthetic content. More results and details can be found on our project page. Dongsu Zhang, Francis Williams, Zan Gojcic, Karsten Kreis, Sanja Fidler, Young Min Kim 0001, Amlan Kar |
CVPR | 7 |
| 2023 | DreamTeacher: Pretraining Image Backbones with Deep Generative ModelsabstractIn this work, we introduce a self-supervised feature representation learning framework DreamTeacher that utilizes generative networks for pre-training downstream image backbones. We propose to distill knowledge from a trained generative model into standard image backbones that have been well engineered for specific perception tasks. We investigate two types of knowledge distillation: 1) distilling learned generative features onto target image backbones as an alternative to pretraining these backbones on large labeled datasets such as ImageNet, and 2) distilling labels obtained from generative networks with task heads onto logits of target backbones. We perform extensive analyses on multiple generative models, dense prediction benchmarks, and several pre-training regimes. We empirically find that our DreamTeacher significantly outperforms existing self-supervised representation learning approaches across the board. Unsupervised ImageNet pre-training with DreamTeacher leads to significant improvements over ImageNet classification pre-training on downstream datasets, showcasing generative models, and diffusion generative models specifically, as a promising approach to representation learning on large, diverse datasets without requiring manual annotation. Daiqing Li, Huan Ling, Amlan Kar, David Acuna, Seung Wook Kim 0001, Karsten Kreis, Antonio Torralba 0001, Sanja Fidler |
ICCV | 3 |
| 2022 | EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object RelationsabstractWe introduce VISOR, a new dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video. VISOR annotates videos from EPIC-KITCHENS, which comes with a new set of challenges not encountered in current video segmentation datasets. Specifically, we need to ensure both short- and long-term consistency of pixel-level annotations as objects undergo transformative interactions, e.g. an onion is peeled, diced and cooked - where we aim to obtain accurate pixel-level annotations of the peel, onion pieces, chopping board, knife, pan, as well as the acting hands. VISOR introduces an annotation pipeline, AI-powered in parts, for scalability and quality. In total, we publicly release 272K manual semantic masks of 257 object classes, 9.9M interpolated dense masks, 67K hand-object relations, covering 36 hours of 179 untrimmed videos. Along with the annotations, we introduce three challenges in video object segmentation, interaction understanding and long-term reasoning.For data, code and leaderboards: http://epic-kitchens.github.io/VISOR Ahmad Darkhalil, Dandan Shan, Bin Zhu 0006, Amlan Kar, Richard E. L. Higgins, Sanja Fidler, David F. Fouhey, Dima Damen |
NeurIPS | 5 |
| 2021 | Towards Good Practices for Efficiently Annotating Large-Scale Image Classification DatasetsabstractData is the engine of modern computer vision, which necessitates collecting large-scale datasets. This is expensive, and guaranteeing the quality of the labels is a major challenge. In this paper, we investigate efficient annotation strategies for collecting multi-class classification labels for a large collection of images. While methods that exploit learnt models for labeling exist, a surprisingly prevalent approach is to query humans for a fixed number of labels per datum and aggregate them, which is expensive. Building on prior work on online joint probabilistic modeling of human annotations and machine-generated beliefs, we propose modifications and best practices aimed at minimizing human labeling effort. Specifically, we make use of advances in self-supervised learning, view annotation as a semi-supervised learning problem, identify and mitigate pitfalls and ablate several key design choices to propose effective guidelines for labeling. Our analysis is done in a more realistic simulation that involves querying human la-belers, which uncovers issues with evaluation using existing worker simulation methods. Simulated experiments on a 125k image subset of the ImageNet100 show that it can be annotated to 80% top-1 accuracy with 0.35 annotations per image on average, a 2.7x and 6.7x improvement over prior work and manual annotation, respectively.1 Yuan-Hong Liao, Amlan Kar, Sanja Fidler |
CVPR | 2 |
| 2021 | ATISS: Autoregressive Transformers for Indoor Scene SynthesisabstractThe ability to synthesize realistic and diverse indoor furniture layouts automatically or based on partial input, unlocks many applications, from better interactive 3D tools to data synthesis for training and simulation. In this paper, we present ATISS, a novel autoregressive transformer architecture for creating diverse and plausible synthetic indoor environments, given only the room type and its floor plan. In contrast to prior work, which poses scene synthesis as sequence generation, our model generates rooms as unordered sets of objects. We argue that this formulation is more natural, as it makes ATISS generally useful beyond fully automatic room layout synthesis. For example, the same trained model can be used in interactive applications for general scene completion, partial room re-arrangement with any objects specified by the user, as well as object suggestions for any partial room. To enable this, our model leverages the permutation equivariance of the transformer when conditioning on the partial scene, and is trained to be permutation-invariant across object orderings. Our model is trained end-to-end as an autoregressive generative model using only labeled 3D bounding boxes as supervision. Evaluations on four room types in the 3D-FRONT dataset demonstrate that our model consistently generates plausible room layouts that are more realistic than existing methods.In addition, it has fewer parameters, is simpler to implement and train and runs up to 8 times faster than existing methods. Despoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis, Andreas Geiger 0001, Sanja Fidler |
NeurIPS | 2 |
| 2020 | Learning to Evaluate Perception Models Using Planner-Centric MetricsabstractVariants of accuracy and precision are the gold-standard by which the computer vision community measures progress of perception algorithms. One reason for the ubiquity of these metrics is that they are largely task-agnostic; we in general seek to detect zero false negatives or positives. The downside of these metrics is that, at worst, they penalize all incorrect detections equally without conditioning on the task or scene, and at best, heuristics need to be chosen to ensure that different mistakes count differently. In this paper, we propose a principled metric for 3D object detection specifically for the task of self-driving. The core idea behind our metric is to isolate the task of object detection and measure the impact the produced detections would induce on the downstream task of driving. Without hand-designing it to, we find that our metric penalizes many of the mistakes that other metrics penalize by design. In addition, our metric downweighs detections based on additional factors such as distance from a detection to the ego car and the speed of the detection in intuitive ways that other detection metrics do not. For human evaluation, we generate scenes in which standard metrics and our metric disagree and find that humans side with our metric 79% of the time. Our project page including an evaluation server can be found at https://nv-tlabs.github.io/detection-relevance. Jonah Philion, Amlan Kar, Sanja Fidler |
CVPR | 2 |
| 2020 | Meta-Sim2: Unsupervised Learning of Scene Structure for Synthetic Data Generation
Jeevan Devaranjan, Amlan Kar, Sanja Fidler |
ECCV (17) | 2 |
| 2020 | Interactive Annotation of 3D Object Geometry Using 2D Scribbles
Tianchang Shen, Jun Gao 0004, Amlan Kar, Sanja Fidler |
ECCV (17) | 3 |
| 2020 | Federated Simulation for Medical Imaging
Daiqing Li, Amlan Kar, Nishant Ravikumar, Alejandro F. Frangi, Sanja Fidler |
MICCAI (1) | 2 |
| 2020 | Nonlinear color triads for approximation, learning and direct manipulation of color distributionsabstractWe present nonlinear color triads, an extension of color gradients able to approximate a variety of natural color distributions that have no standard interactive representation. We derive a method to fit this compact parametric representation to existing images and show its power for tasks such as image editing and compression. Our color triad formulation can also be included in standard deep learning architectures, facilitating further research. Maria Shugrina, Amlan Kar, Sanja Fidler, Karan Singh 0004 |
ACM Trans. Graph. | 2 |
| 2019 | Devil Is in the Edges: Learning Semantic Boundaries From Noisy AnnotationsabstractWe tackle the problem of semantic boundary prediction, which aims to identify pixels that belong to object(class) boundaries. We notice that relevant datasets consist of a significant level of label noise, reflecting the fact that precise annotations are laborious to get and thus annotators trade-off quality with efficiency. We aim to learn sharp and precise semantic boundaries by explicitly reasoning about annotation noise during training. We propose a simple new layer and loss that can be used with existing learning-based boundary detectors. Our layer/loss enforces the detector to predict a maximum response along the normal direction at an edge, while also regularizing its direction. We further reason about true object boundaries during training using a level set formulation, which allows the network to learn from misaligned labels in an end-to-end fashion. Experiments show that we improve over the CASENet backbone network by more than 4% in terms of MF(ODS) and 18.61% in terms of AP, outperforming all current state-of-the-art methods including those that deal with alignment. Furthermore, we show that our learned network can be used to significantly improve coarse segmentation labels, lending itself as an efficient way to label new data. David Acuna, Amlan Kar, Sanja Fidler |
CVPR | 2 |
| 2019 | Fast Interactive Object Annotation With Curve-GCNabstractManually labeling objects by tracing their boundaries is a laborious process. In Polygon-RNN++, the authors proposed Polygon-RNN that produces polygonal annotations in a recurrent manner using a CNN-RNN architecture, allowing interactive correction via humans-in-the-loop. We propose a new framework that alleviates the sequential nature of Polygon-RNN, by predicting all vertices simultaneously using a Graph Convolutional Network (GCN). Our model is trained end-to-end, and runs in real time. It supports object annotation by either polygons or splines, facilitating labeling efficiency for both line-based and curved objects. We show that Curve-GCN outperforms all existing approaches in automatic mode, including the powerful DeepLab, and is significantly more efficient in interactive mode than Polygon-RNN++. Our model runs at 29.3ms in automatic, and 2.6ms in interactive mode, making it 10x and 100x faster than Polygon-RNN++. Huan Ling, Jun Gao 0004, Amlan Kar, Wenzheng Chen, Sanja Fidler |
CVPR | 3 |
| 2019 | Creative Flow+ DatasetabstractWe present the Creative Flow+ Dataset, the first diverse multi-style artistic video dataset richly labeled with per-pixel optical flow, occlusions, correspondences, segmentation labels, normals, and depth. Our dataset includes 3000 animated sequences rendered using styles randomly selected from 40 textured line styles and 38 shading styles, spanning the range between flat cartoon fill and wildly sketchy shading. Our dataset includes 124K+ train set frames and 10K test set frames rendered at 1500x1500 resolution, far surpassing the largest available optical flow datasets in size. While modern techniques for tasks such as optical flow estimation achieve impressive performance on realistic images and video, today there is no way to gauge their performance on non-photorealistic images. Creative Flow+ poses a new challenge to generalize real-world Computer Vision to messy stylized content. We show that learning-based optical flow methods fail to generalize to this data and struggle to compete with classical approaches, and invite new research in this area. Our dataset and a new optical flow benchmark will be publicly available at: www.cs.toronto.edu/creativeflow/. We further release the complete dataset creation pipeline, allowing the community to generate and stylize their own data on demand. Maria Shugrina, Ziheng Liang, Amlan Kar, Jiaman Li, Angad Singh, Karan Singh 0004, Sanja Fidler |
CVPR | 3 |
| 2019 | Object Instance Annotation With Deep Extreme Level Set EvolutionabstractIn this paper, we tackle the task of interactive object segmentation. We revive the old ideas on level set segmentation which framed object annotation as curve evolution. Carefully designed energy functions ensured that the curve was well aligned with image boundaries, and generally "well behaved". The Level Set Method can handle objects with complex shapes and topological changes such as merging and splitting, thus able to deal with occluded objects and objects with holes. We propose Deep Extreme Level Set Evolution that combines powerful CNN models with level set optimization in an end-to-end fashion. Our method learns to predict evolution parameters conditioned on the image and evolves the predicted initial contour to produce the final result. We make our model interactive by incorporating user clicks on the extreme boundary points, following DEXTR. We show that our approach significantly outperforms DEXTR on the static Cityscapes dataset and the video segmentation benchmark DAVIS, and performs on par on PASCAL and SBD. David Acuna, Huan Ling, Amlan Kar, Sanja Fidler |
CVPR | 4 |
| 2019 | Neural Turtle Graphics for Modeling City Road LayoutsabstractWe propose Neural Turtle Graphics (NTG), a novel generative model for spatial graphs, and demonstrate its applications in modeling city road layouts. Specifically, we represent the road layout using a graph where nodes in the graph represent control points and edges in the graph represents road segments. NTG is a sequential generative model parameterized by a neural network. It iteratively generates a new node and an edge connecting to an existing node conditioned on the current graph. We train NTG on Open Street Map data and show it outperforms existing approaches using a set of diverse performance metrics. Moreover, our method allows users to control styles of generated road layouts mimicking existing cities as well as to sketch a part of the city road layout to be synthesized. In addition to synthesis, the proposed NTG finds uses in an analytical task of aerial road parsing. Experimental results show that it achieves state-of-the-art performance on the SpaceNet dataset. Hang Chu, Daiqing Li, David Acuna, Amlan Kar, Maria Shugrina, Xinkai Wei, Ming-Yu Liu 0001, Antonio Torralba 0001, Sanja Fidler |
ICCV | 4 |
| 2019 | Meta-Sim: Learning to Generate Synthetic DatasetsabstractTraining models to high-end performance requires availability of large labeled datasets, which are expensive to get. The goal of our work is to automatically synthesize labeled datasets that are relevant for a downstream task. We propose Meta-Sim, which learns a generative model of synthetic scenes, and obtain images as well as its corresponding ground-truth via a graphics engine. We parametrize our dataset generator with a neural network, which learns to modify attributes of scene graphs obtained from probabilistic scene grammars, so as to minimize the distribution gap between its rendered outputs and target data. If the real dataset comes with a small labeled validation set, we additionally aim to optimize a meta-objective, i.e. downstream task performance. Experiments show that the proposed method can greatly improve content generation quality over a human-engineered probabilistic scene grammar, both qualitatively and quantitatively as measured by performance on a downstream task. Amlan Kar, Aayush Prakash, Ming-Yu Liu 0001, Eric Cameracci, Justin Yuan, Matt Rusiniak, David Acuna, Antonio Torralba 0001, Sanja Fidler |
ICCV | 1 |
| 2019 | Learning to Caption Images Through a Lifetime by Asking QuestionsabstractIn order to bring artificial agents into our lives, we will need to go beyond supervised learning on closed datasets to having the ability to continuously expand knowledge. Inspired by a student learning in a classroom, we present an agent that can continuously learn by posing natural language questions to humans. Our agent is composed of three interacting modules, one that performs captioning, another that generates questions and a decision maker that learns when to ask questions by implicitly reasoning about the uncertainty of the agent and expertise of the teacher. As compared to current active learning methods which query images for full captions, our agent is able to ask pointed questions to improve the generated captions. The agent trains on the improved captions, expanding its knowledge. We show that our approach achieves better performance using less human supervision than the baselines on the challenging MSCOCO dataset. Tingke Shen, Amlan Kar, Sanja Fidler |
ICCV | 2 |
| 2018 | Efficient Interactive Annotation of Segmentation Datasets With Polygon-RNN++abstractManually labeling datasets with object masks is extremely time consuming. In this work, we follow the idea of Polygon-RNN [4] to produce polygonal annotations of objects interactively using humans-in-the-loop. We introduce several important improvements to the model: 1) we design a new CNN encoder architecture, 2) show how to effectively train the model with Reinforcement Learning, and 3) significantly increase the output resolution using a Graph Neural Network, allowing the model to accurately annotate high-resolution objects in images. Extensive evaluation on the Cityscapes dataset [8] shows that our model, which we refer to as Polygon-RNN++, significantly outperforms the original model in both automatic (10% absolute and 16% relative improvement in mean IoU) and interactive modes (requiring 50% fewer clicks by annotators). We further analyze the cross-domain scenario in which our model is trained on one dataset, and used out of the box on datasets from varying domains. The results show that Polygon-RNN++ exhibits powerful generalization capabilities, achieving significant improvements over existing pixel-wise methods. Using simple online fine-tuning we further achieve a high reduction in annotation time for new datasets, moving a step closer towards an interactive annotation tool to be used in practice. David Acuna, Huan Ling, Amlan Kar, Sanja Fidler |
CVPR | 3 |
| 2017 | AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in VideosabstractWe propose a novel method for temporally pooling frames in a video for the task of human action recognition. The method is motivated by the observation that there are only a small number of frames which, together, contain sufficient information to discriminate an action class present in a video, from the rest. The proposed method learns to pool such discriminative and informative frames, while discarding a majority of the non-informative frames in a single temporal scan of the video. Our algorithm does so by continuously predicting the discriminative importance of each video frame and subsequently pooling them in a deep learning framework. We show the effectiveness of our proposed pooling method on standard benchmarks where it consistently improves on baseline pooling methods, with both RGB and optical flow based Convolutional networks. Further, in combination with complementary video representations, we show results that are competitive with respect to the state-of-the-art results on two challenging and publicly available benchmark datasets. Amlan Kar, Nishant Rai, Karan Sikka, Gaurav Sharma 0004 |
CVPR | 1 |