David J. Weiss

dblp:06/7944 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
1since 2021 · last 2022
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Autonomous driving · 29% Image recognition and object detection · 15% Segmentation and scene understanding · 14%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving › trajectory prediction
multi-agent trajectory prediction
0.612022
Scene Transformer: A unified architecture for predicting future trajectories of multiple agents · ICLR 2022
Robotics › Autonomous driving
trajectory prediction
0.612022
Scene Transformer: A unified architecture for predicting future trajectories of multiple agents · ICLR 2022
Computer vision › Face, body and person analysis
human pose estimation
0.442013
Dynamic Structured Model Selection · ICCV 2013
Parsing human motion with stretchable models · CVPR 2011
Learning Adaptive Value of Information for Structured Prediction · NIPS 2013
Computer vision › Image recognition and object detection
attribute recognition
0.212014
Understanding Objects in Detail with Fine-Grained Attributes · CVPR 2014
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.212014
Understanding Objects in Detail with Fine-Grained Attributes · CVPR 2014
Computer vision › Image recognition and object detection › object detection
part-based object detection
0.212014
Understanding Objects in Detail with Fine-Grained Attributes · CVPR 2014
Machine learning › Efficient and distributed learning
adaptive computation
0.212013
Learning Adaptive Value of Information for Structured Prediction · NIPS 2013
Machine learning › Efficient and distributed learning › adaptive computation
dynamic model selection
0.212013
Dynamic Structured Model Selection · ICCV 2013
Machine learning › Efficient and distributed learning
inference efficiency
0.212013
Dynamic Structured Model Selection · ICCV 2013
Computer vision › Segmentation and scene understanding
object segmentation
0.212013
SCALPEL: Segmentation Cascades with Localized Priors and Efficient Learning · CVPR 2013
Computer vision › Segmentation and scene understanding › image segmentation › region-based segmentation
region merging
0.212013
SCALPEL: Segmentation Cascades with Localized Priors and Efficient Learning · CVPR 2013
Computer vision › Segmentation and scene understanding
shape prior
0.212013
SCALPEL: Segmentation Cascades with Localized Priors and Efficient Learning · CVPR 2013
Machine learning › Reinforcement learning
value function approximation
0.212013
Learning Adaptive Value of Information for Structured Prediction · NIPS 2013
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › decision making under uncertainty
value of information
0.212013
Learning Adaptive Value of Information for Structured Prediction · NIPS 2013
Computer vision › Face, body and person analysis › human pose estimation
video pose estimation
0.212013
Dynamic Structured Model Selection · ICCV 2013
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.112010
Sidestepping Intractable Inference with Structured Ensemble Cascades · NIPS 2010
Machine learning › Probabilistic and Bayesian machine learning
structured prediction
0.112010
Sidestepping Intractable Inference with Structured Ensemble Cascades · NIPS 2010
Recommender systems › collaborative filtering
matrix factorization
0.112010
Mixed Membership Matrix Factorization · ICML 2010
Computer vision › Segmentation and scene understanding
part segmentation
0.112014
Understanding Objects in Detail with Fine-Grained Attributes · CVPR 2014
Computer vision › Image recognition and object detection › text recognition
optical character recognition
0.012013
Learning Adaptive Value of Information for Structured Prediction · NIPS 2013

Methods — techniques the papers use, named apart from their topics

transformer architecture · 0.6part-based pooling · 0.2coarse-to-fine cascade · 0.2value function approximation · 0.2structured learning · 0.2reinforcement learning · 0.2introspection · 0.2cascade learning · 0.2ensemble of tractable models · 0.1dual decomposition · 0.1
YearPublicationVenuePosition
2022 Scene Transformer: A unified architecture for predicting future trajectories of multiple agents
Jiquan Ngiam, Vijay Vasudevan, Benjamin Caine, Hao-Tien Chiang, Jeffrey Ling, Rebecca Roelofs, Alex Bewley, Chenxi Liu 0001, Ashish Venugopal, David J. Weiss, Benjamin Sapp, Jonathon Shlens
ICLR11
2014 Understanding Objects in Detail with Fine-Grained Attributes
abstract
We study the problem of understanding objects in detail, intended as recognizing a wide array of fine-grained object attributes. To this end, we introduce a dataset of 7, 413 airplanes annotated in detail with parts and their attributes, leveraging images donated by airplane spotters and crowd-sourcing both the design and collection of the detailed annotations. We provide a number of insights that should help researchers interested in designing fine-grained datasets for other basic level categories. We show that the collected data can be used to study the relation between part detection and attribute prediction by diagnosing the performance of classifiers that pool information from different parts of an object. We note that the prediction of certain attributes can benefit substantially from accurate part detection. We also show that, differently from previous results in object detection, employing a large number of part templates can improve detection accuracy at the expenses of detection speed. We finally propose a coarse-to-fine approach to speed up detection through a hierarchical cascade algorithm.
Andrea Vedaldi, Siddharth Mahendran, Stavros Tsogkas, Subhransu Maji, Ross B. Girshick, Juho Kannala, Esa Rahtu, Iasonas Kokkinos, Matthew B. Blaschko, David J. Weiss, Ben Taskar, Karen Simonyan, Naomi Saphra, Sammy Mohamed
CVPR10
2013 SCALPEL: Segmentation Cascades with Localized Priors and Efficient Learning
abstract
We propose SCALPEL, a flexible method for object segmentation that integrates rich region-merging cues with mid- and high-level information about object layout, class, and scale into the segmentation process. Unlike competing approaches, SCALPEL uses a cascade of bottom-up segmentation models that is capable of learning to ignore boundaries early on, yet use them as a stopping criterion once the object has been mostly segmented. Furthermore, we show how such cascades can be learned efficiently. When paired with a novel method that generates better localized shape priors than our competitors, our method leads to a concise, accurate set of segmentation proposals, these proposals are more accurate on the PASCAL VOC2010 dataset than state-of-the-art methods that use re-ranking to filter much larger bags of proposals. The code for our algorithm is available online.
David J. Weiss, Ben Taskar
CVPR1
2013 Dynamic Structured Model Selection
abstract
In many cases, the predictive power of structured models for for complex vision tasks is limited by a trade-off between the expressiveness and the computational tractability of the model. However, choosing this trade-off statically a priori is sub optimal, as images and videos in different settings vary tremendously in complexity. On the other hand, choosing the trade-off dynamically requires knowledge about the accuracy of different structured models on any given example. In this work, we propose a novel two-tier architecture that provides dynamic speed/accuracy trade-offs through a simple type of introspection. Our approach, which we call dynamic structured model selection (DMS), leverages typically intractable features in structured learning problems in order to automatically determine' which of several models should be used at test-time in order to maximize accuracy under a fixed budgetary constraint. We demonstrate DMS on two sequential modeling vision tasks, and we establish a new state-of-the-art in human pose estimation in video with an implementation that is roughly 23× faster than the previous standard implementation.
David J. Weiss, Benjamin Sapp, Ben Taskar
ICCV1
2013 Learning Adaptive Value of Information for Structured Prediction
abstract
Discriminative methods for learning structured models have enabled wide-spread use of very rich feature representations. However, the computational cost of feature extraction is prohibitive for large-scale or time-sensitive applications, often dominating the cost of inference in the models. Significant efforts have been devoted to sparsity-based model selection to decrease this cost. Such feature selection methods control computation statically and miss the opportunity to fine-tune feature extraction to each input at run-time. We address the key challenge of learning to control fine-grained feature extraction adaptively, exploiting non-homogeneity of the data. We propose an architecture that uses a rich feedback loop between extraction and prediction. The run-time control policy is learned using efficient value-function approximation, which adaptively determines the value of information of features at the level of individual variables for each input. We demonstrate significant speedups over state-of-the-art methods on two challenging datasets. For articulated pose estimation in video, we achieve a more accurate state-of-the-art model that is simultaneously 4$\times$ faster while using only a small fraction of possible features, with similar results on an OCR task.
David J. Weiss, Ben Taskar
NIPS1
2011 Parsing human motion with stretchable models
abstract
We address the problem of articulated human pose estimation in videos using an ensemble of tractable models with rich appearance, shape, contour and motion cues. In previous articulated pose estimation work on unconstrained videos, using temporal coupling of limb positions has made little to no difference in performance over parsing frames individually. One crucial reason for this is that joint parsing of multiple articulated parts over time involves intractable inference and learning problems, and previous work has resorted to approximate inference and simplified models. We overcome these computational and modeling limitations using an ensemble of tractable submodels which couple locations of body joints within and across frames using expressive cues. Each submodel is responsible for tracking a single joint through time (e.g., left elbow) and also models the spatial arrangement of all joints in a single frame. Because of the tree structure of each submodel, we can perform efficient exact inference and use rich temporal features that depend on image appearance, e.g., color tracking and optical flow contours. We propose and experimentally investigate a hierarchy of submodel combination methods, and we find that a highly efficient max-marginal combination method outperforms much slower (by orders of magnitude) approximate inference using dual decomposition. We apply our pose model on a new video dataset of highly varied and articulated poses from TV shows. We show significant quantitative and qualitative improvements over state-of-the-art single-frame pose estimation approaches.
Benjamin Sapp, David J. Weiss, Ben Taskar
CVPR2
2010 Mixed Membership Matrix Factorization
Lester Mackey, David J. Weiss, Michael I. Jordan
ICML2
2010 Sidestepping Intractable Inference with Structured Ensemble Cascades
abstract
For many structured prediction problems, complex models often require adopting approximate inference techniques such as variational methods or sampling, which generally provide no satisfactory accuracy guarantees. In this work, we propose sidestepping intractable inference altogether by learning ensembles of tractable sub-models as part of a structured prediction cascade. We focus in particular on problems with high-treewidth and large state-spaces, which occur in many computer vision tasks. Unlike other variational methods, our ensembles do not enforce agreement between sub-models, but filter the space of possible outputs by simply adding and thresholding the max-marginals of each constituent model. Our framework jointly estimates parameters for all models in the ensemble for each level of the cascade by minimizing a novel, convex loss function, yet requires only a linear increase in computation over learning or inference in a single tractable sub-model. We provide a generalization bound on the filtering loss of the ensemble as a theoretical justification of our approach, and we evaluate our method on both synthetic data and the task of estimating articulated human pose from challenging videos. We find that our approach significantly outperforms loopy belief propagation on the synthetic data and a state-of-the-art model on the pose estimation/tracking problem.
David J. Weiss, Benjamin Sapp, Ben Taskar
NIPS1