EDBT 2026 Demo / reviewers in the wild / expert
Kevin D. Tang
dblp:90/8652
· DBLP profile ↗
11ranked-venue papers
7as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Image recognition and object detection · 34% Video understanding and tracking · 24% Transfer learning and domain adaptation · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 28 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection › object localization
co-localization |
0.4 | 2 | 2014 | Efficient Image and Video Co-localization with Frank-Wolfe Algorithm · ECCV (6) 2014 Co-localization in Real-World Images · CVPR 2014 |
Computer vision › Video understanding and tracking › event recognition
complex event detection |
0.3 | 2 | 2013 | Combining the Right Features for Complex Event Recognition · ICCV 2013 Learning latent temporal structure for complex event detection · CVPR 2012 |
Computer vision › Vision and language › image captioning
dense captioning |
0.3 | 1 | 2017 | Dense Captioning with Joint Inference and Visual Context · CVPR 2017 |
Computer vision › Image recognition and object detection › visual concept recognition
visual concept detection |
0.3 | 1 | 2017 | Dense Captioning with Joint Inference and Visual Context · CVPR 2017 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2015 | Improving Image Classification with Location Context · ICCV 2015 |
Computer vision › Video understanding and tracking › temporal modeling
temporal context learning |
0.2 | 1 | 2015 | Learning Temporal Embeddings for Complex Video Analysis · ICCV 2015 |
Machine learning › Representation and self-supervised learning › representation learning › embedding learning
temporal embedding |
0.2 | 1 | 2015 | Learning Temporal Embeddings for Complex Video Analysis · ICCV 2015 |
Computer vision › Video understanding and tracking
video representation learning |
0.2 | 1 | 2015 | Learning Temporal Embeddings for Complex Video Analysis · ICCV 2015 |
Computer vision › Image recognition and object detection
object localization |
0.2 | 1 | 2014 | Co-localization in Real-World Images · CVPR 2014 |
Computer vision › Image recognition and object detection › object localization
weakly supervised object localization |
0.2 | 1 | 2014 | Co-localization in Real-World Images · CVPR 2014 |
Machine learning › Deep learning architectures and training
feature fusion |
0.2 | 1 | 2013 | Combining the Right Features for Complex Event Recognition · ICCV 2013 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.2 | 1 | 2013 | Discriminative Segment Annotation in Weakly Labeled Video · CVPR 2013 |
Machine learning › Trustworthy machine learning › robustness
robust learning |
0.2 | 1 | 2013 | Discriminative Segment Annotation in Weakly Labeled Video · CVPR 2013 |
Computer vision › Segmentation and scene understanding
video segmentation |
0.2 | 1 | 2013 | Discriminative Segment Annotation in Weakly Labeled Video · CVPR 2013 |
Computer vision › Video understanding and tracking
action segmentation |
0.1 | 1 | 2012 | Learning latent temporal structure for complex event detection · CVPR 2012 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › visual domain adaptation
image-to-video adaptation |
0.1 | 1 | 2012 | Shifting Weights: Adapting Object Detectors from Image to Video · NIPS 2012 |
Computer vision › Image recognition and object detection
object detection |
0.1 | 1 | 2012 | Shifting Weights: Adapting Object Detectors from Image to Video · NIPS 2012 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.1 | 1 | 2012 | Shifting Weights: Adapting Object Detectors from Image to Video · NIPS 2012 |
Computer vision › Video understanding and tracking
video object detection |
0.1 | 1 | 2012 | Shifting Weights: Adapting Object Detectors from Image to Video · NIPS 2012 |
Computer vision › Image recognition and object detection › image classification
object classification |
0.1 | 1 | 2010 | Optimizing one-shot recognition with micro-set learning · CVPR 2010 |
Machine learning › Transfer learning and domain adaptation › few-shot learning › few-shot image classification
one-shot recognition |
0.1 | 1 | 2010 | Optimizing one-shot recognition with micro-set learning · CVPR 2010 |
Natural language and speech › Information extraction and text analysis
geolocation |
0.1 | 1 | 2015 | Improving Image Classification with Location Context · ICCV 2015 |
Computer vision › Video understanding and tracking
video classification |
0.1 | 1 | 2015 | Learning Temporal Embeddings for Complex Video Analysis · ICCV 2015 |
Information retrieval › multimedia analysis and retrieval
video retrieval |
0.1 | 1 | 2015 | Learning Temporal Embeddings for Complex Video Analysis · ICCV 2015 |
Machine learning › Optimization for machine learning › constrained optimization
frank-wolfe algorithm |
0.1 | 1 | 2014 | Efficient Image and Video Co-localization with Frank-Wolfe Algorithm · ECCV (6) 2014 |
Computer vision › Segmentation and scene understanding › image segmentation
pixel-level segmentation |
0.0 | 1 | 2013 | Discriminative Segment Annotation in Weakly Labeled Video · CVPR 2013 |
Multimedia analysis and retrieval › event detection
video event detection |
0.0 | 1 | 2013 | Combining the Right Features for Complex Event Recognition · ICCV 2013 |
Robotics › Motion planning and robot control › robot control
trajectory tracking |
0.0 | 1 | 2012 | Shifting Weights: Adapting Object Detectors from Image to Video · NIPS 2012 |
Methods — techniques the papers use, named apart from their topics
hard negative mining · 0.4data augmentation · 0.4joint inference · 0.3context fusion · 0.3feature pooling · 0.2convolutional neural network · 0.2joint image-box formulation · 0.2frank-wolfe algorithm · 0.2convex quadratic programming · 0.2score-based structure learning · 0.2parallelizable optimization · 0.2hierarchical feature combination · 0.2and-or graph · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Dense Captioning with Joint Inference and Visual ContextabstractDense captioning is a newly emerging computer vision topic for understanding images with dense language descriptions. The goal is to densely detect visual concepts (e.g., objects, object parts, and interactions between them) from images, labeling each with a short descriptive phrase. We identify two key challenges of dense captioning that need to be properly addressed when tackling the problem. First, dense visual concept annotations in each image are associated with highly overlapping target regions, making accurate localization of each visual concept challenging. Second, the large amount of visual concepts makes it hard to recognize each of them by appearance alone. We propose a new model pipeline based on two novel ideas, joint inference and context fusion, to alleviate these two challenges. We design our model architecture in a methodical manner and thoroughly evaluate the variations in architecture. Our final model, compact and efficient, achieves state-of-the-art accuracy on Visual Genome [23] for dense captioning with a relative gain of 73% compared to the previous best algorithm. Qualitative experiments also reveal the semantic capabilities of our model in dense captioning. Kevin D. Tang, Jianchao Yang, Li-Jia Li 0001 |
CVPR | 2 |
| 2015 | Learning Temporal Embeddings for Complex Video AnalysisabstractIn this paper, we propose to learn temporal embeddings of video frames for complex video analysis. Large quantities of unlabeled video data can be easily obtained from the Internet. These videos possess the implicit weak label that they are sequences of temporally and semantically coherent images. We leverage this information to learn temporal embeddings for video frames by associating frames with the temporal context that they appear in. To do this, we propose a scheme for incorporating temporal context based on past and future frames in videos, and compare this to other contextual representations. In addition, we show how data augmentation using multi-resolution samples and hard negatives helps to significantly improve the quality of the learned embeddings. We evaluate various design decisions for learning temporal embeddings, and show that our embeddings can improve performance for multiple video tasks such as retrieval, classification, and temporal order recovery in unconstrained Internet video. Vignesh Ramanathan, Kevin D. Tang, Greg Mori, Li Fei-Fei 0001 |
ICCV | 2 |
| 2015 | Improving Image Classification with Location ContextabstractWith the widespread availability of cellphones and cameras that have GPS capabilities, it is common for images being uploaded to the Internet today to have GPS coordinates associated with them. In addition to research that tries to predict GPS coordinates from visual features, this also opens up the door to problems that are conditioned on the availability of GPS coordinates. In this work, we tackle the problem of performing image classification with location context, in which we are given the GPS coordinates for images in both the train and test phases. We explore different ways of encoding and extracting features from the GPS coordinates, and show how to naturally incorporate these features into a Convolutional Neural Network (CNN), the current state-of-the-art for most image classification and recognition problems. We also show how it is possible to simultaneously learn the optimal pooling radii for a subset of our features within the CNN framework. To evaluate our model and to help promote research in this area, we identify a set of location-sensitive concepts and annotate a subset of the Yahoo Flickr Creative Commons 100M dataset that has GPS coordinates with these concepts, which we make publicly available. By leveraging location context, we are able to achieve almost a 7% gain in mean average precision. Kevin D. Tang, Manohar Paluri, Li Fei-Fei 0001, Rob Fergus, Lubomir D. Bourdev |
ICCV | 1 |
| 2014 | Co-localization in Real-World ImagesabstractIn this paper, we tackle the problem of co-localization in real-world images. Co-localization is the problem of simultaneously localizing (with bounding boxes) objects of the same class across a set of distinct images. Although similar problems such as co-segmentation and weakly supervised localization have been previously studied, we focus on being able to perform co-localization in real-world settings, which are typically characterized by large amounts of intra-class variation, inter-class diversity, and annotation noise. To address these issues, we present a joint image-box formulation for solving the co-localization problem, and show how it can be relaxed to a convex quadratic program which can be efficiently solved. We perform an extensive evaluation of our method compared to previous state-of-the-art approaches on the challenging PASCAL VOC 2007 and Object Discovery datasets. In addition, we also present a large-scale study of co-localization on ImageNet, involving ground-truth annotations for 3, 624 classes and approximately 1 million images. Kevin D. Tang, Armand Joulin, Li-Jia Li 0001, Li Fei-Fei 0001 |
CVPR | 1 |
| 2014 | Efficient Image and Video Co-localization with Frank-Wolfe Algorithm
Armand Joulin, Kevin D. Tang, Li Fei-Fei 0001 |
ECCV (6) | 2 |
| 2013 | Discriminative Segment Annotation in Weakly Labeled VideoabstractThe ubiquitous availability of Internet video offers the vision community the exciting opportunity to directly learn localized visual concepts from real-world imagery. Unfortunately, most such attempts are doomed because traditional approaches are ill-suited, both in terms of their computational characteristics and their inability to robustly contend with the label noise that plagues uncurated Internet content. We present CRANE, a weakly supervised algorithm that is specifically designed to learn under such conditions. First, we exploit the asymmetric availability of real-world training data, where small numbers of positive videos tagged with the concept are supplemented with large quantities of unreliable negative data. Second, we ensure that CRANE is robust to label noise, both in terms of tagged videos that fail to contain the concept as well as occasional negative videos that do. Finally, CRANE is highly parallelizable, making it practical to deploy at large scale without sacrificing the quality of the learned solution. Although CRANE is general, this paper focuses on segment annotation, where we show state-of-the-art pixel-level segmentation results on two datasets, one of which includes a training set of spatiotemporal segments from more than 20,000 videos. Kevin D. Tang, Rahul Sukthankar, Jay Yagnik, Li Fei-Fei 0001 |
CVPR | 1 |
| 2013 | Combining the Right Features for Complex Event RecognitionabstractIn this paper, we tackle the problem of combining features extracted from video for complex event recognition. Feature combination is an especially relevant task in video data, as there are many features we can extract, ranging from image features computed from individual frames to video features that take temporal information into account. To combine features effectively, we propose a method that is able to be selective of different subsets of features, as some features or feature combinations may be uninformative for certain classes. We introduce a hierarchical method for combining features based on the AND/OR graph structure, where nodes in the graph represent combinations of different sets of features. Our method automatically learns the structure of the AND/OR graph using score-based structure learning, and we introduce an inference procedure that is able to efficiently compute structure scores. We present promising results and analysis on the difficult and large-scale 2011 TRECVID Multimedia Event Detection dataset. Kevin D. Tang, Bangpeng Yao, Li Fei-Fei 0001, Daphne Koller |
ICCV | 1 |
| 2012 | Learning latent temporal structure for complex event detectionabstractIn this paper, we tackle the problem of understanding the temporal structure of complex events in highly varying videos obtained from the Internet. Towards this goal, we utilize a conditional model trained in a max-margin framework that is able to automatically discover discriminative and interesting segments of video, while simultaneously achieving competitive accuracies on difficult detection and recognition tasks. We introduce latent variables over the frames of a video, and allow our algorithm to discover and assign sequences of states that are most discriminative for the event. Our model is based on the variable-duration hidden Markov model, and models durations of states in addition to the transitions between states. The simplicity of our model allows us to perform fast, exact inference using dynamic programming, which is extremely important when we set our sights on being able to process a very large number of videos quickly and efficiently. We show promising results on the Olympic Sports dataset [16] and the 2011 TRECVID Multimedia Event Detection task [18]. We also illustrate and visualize the semantic understanding capabilities of our model. Kevin D. Tang, Li Fei-Fei 0001, Daphne Koller |
CVPR | 1 |
| 2012 | Shifting Weights: Adapting Object Detectors from Image to VideoabstractTypical object detectors trained on images perform poorly on video, as there is a clear distinction in domain between the two types of data. In this paper, we tackle the problem of adapting object detectors learned from images to work well on videos. We treat the problem as one of unsupervised domain adaptation, in which we are given labeled data from the source domain (image), but only unlabeled data from the target domain (video). Our approach, self-paced domain adaptation, seeks to iteratively adapt the detector by re-training the detector with automatically discovered target domain examples, starting with the easiest first. At each iteration, the algorithm adapts by considering an increased number of target domain examples, and a decreased number of source domain examples. To discover target domain examples from the vast amount of video data, we introduce a simple, robust approach that scores trajectory tracks instead of bounding boxes. We also show how rich and expressive features specific to the target domain can be incorporated under the same framework. We show promising results on the 2011 TRECVID Multimedia Event Detection and LabelMe Video datasets that illustrate the benefit of our approach to adapt object detectors to video. Kevin D. Tang, Vignesh Ramanathan, Li Fei-Fei 0001, Daphne Koller |
NIPS | 1 |
| 2010 | Optimizing one-shot recognition with micro-set learningabstractFor object category recognition to scale beyond a small number of classes, it is important that algorithms be able to learn from a small amount of labeled data per additional class. One-shot recognition aims to apply the knowledge gained from a set of categories with plentiful data to categories for which only a single exemplar is available for each. As with earlier efforts motivated by transfer learning, we seek an internal representation for the domain that generalizes across classes. However, in contrast to existing work, we formulate the problem in a fundamentally new manner by optimizing the internal representation for the one-shot task using the notion of micro-sets. A micro-set is a sample of data that contains only a single instance of each category, sampled from the pool of available data, which serves as a mechanism to force the learned representation to explicitly address the variability and noise inherent in the one-shot recognition task. We optimize our learned domain features so that they minimize an expected loss over micro-sets drawn from the training set and show that these features generalize effectively to previously unseen categories. We detail a discriminative approach for optimizing one-shot recognition using micro-sets and present experiments on the Animals with Attributes and Caltech-101 datasets that demonstrate the benefits of our formulation. Kevin D. Tang, Marshall F. Tappen, Rahul Sukthankar, Christoph H. Lampert |
CVPR | 1 |
| 2010 | Towards computational models of kinship verificationabstractWe tackle the challenge of kinship verification using novel feature extraction and selection methods, automatically classifying pairs of face images as “related” or “unrelated” (in terms of kinship). First, we conducted a controlled online search to collect frontal face images of 150 pairs of public figures and celebrities, along with images of their parents or children. Next, we propose and evaluate a set of low-level image features for this classification problem. After selecting the most discriminative inherited facial features, we demonstrate a classification accuracy of 70.67% on a test set of image pairs using K-Nearest-Neighbors. Finally, we present an evaluation of human performance on this problem. Ruogu Fang, Kevin D. Tang, Noah Snavely, Tsuhan Chen |
ICIP | 2 |