Brandon Rothrock

dblp:50/890 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
0since 2021 · last 2020
0000-0003-2237-6589ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-authorSystems, architecture and hardware · 4Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Video understanding and tracking · 61% 3D vision · 15% Multi-agent systems · 10%
Human-computer interaction and pervasive computing
4 papers
User interface design and tools · 36% Interaction techniques and input · 29% Haptics and multimodal interaction · 29%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
activity recognition
0.522017
Privacy-Preserving Human Activity Recognition from Extreme Low Resolution · AAAI 2017
Joint inference of groups, events and human roles in aerial videos · CVPR 2015
Computer vision › 3D vision
hand-object interaction
0.312018
Unsupervised Learning of Hierarchical Models for Hand-Object Interactions · ICRA 2018
Image and video processing
super-resolution
0.312017
Privacy-Preserving Human Activity Recognition from Extreme Low Resolution · AAAI 2017
Computer vision › Video understanding and tracking › video analytics
aerial video analysis
0.212015
Joint inference of groups, events and human roles in aerial videos · CVPR 2015
Computer vision › Video understanding and tracking › egocentric video understanding
first-person activity recognition
0.212015
Pooled motion features for first-person videos · CVPR 2015
Computer vision › Video understanding and tracking › activity recognition
human activity recognition
0.212015
Pooled motion features for first-person videos · CVPR 2015
Computer vision › Video understanding and tracking
motion representation
0.212015
Pooled motion features for first-person videos · CVPR 2015
Knowledge, reasoning and agents › Multi-agent systems › task allocation
role assignment
0.212015
Joint inference of groups, events and human roles in aerial videos · CVPR 2015
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
background segmentation
0.212013
Integrating Grammar and Segmentation for Human Pose Estimation · CVPR 2013
Computer vision › Face, body and person analysis
human pose estimation
0.212013
Integrating Grammar and Segmentation for Human Pose Estimation · CVPR 2013
User interface design and tools
user interface generation
0.122006
Huddle: automatically generating interfaces for systems of multiple connected appliances · UIST 2006
UNIFORM: automatically generating consistent remote control user interfaces · CHI 2006
Haptics and multimodal interaction
tactile sensing
0.112018
Unsupervised Learning of Hierarchical Models for Hand-Object Interactions · ICRA 2018
Privacy and data protection › privacy-preserving computation › privacy-preserving multimedia processing
privacy-preserving video
0.112017
Privacy-Preserving Human Activity Recognition from Extreme Low Resolution · AAAI 2017
Interaction techniques and input
text entry
0.112006
Few-key text entry revisited: mnemonic gestures on four keys · CHI 2006
Interaction techniques and input › non-visual interaction
eyes-free interaction
0.012006
Few-key text entry revisited: mnemonic gestures on four keys · CHI 2006
Interaction techniques and input
remote control interface
0.012006
UNIFORM: automatically generating consistent remote control user interfaces · CHI 2006

Methods — techniques the papers use, named apart from their topics

low-resolution transformation learning · 0.9inverse super resolution · 0.9unsupervised learning · 0.7temporal and-or graph · 0.7dynamic time alignment kernel · 0.7spatiotemporal AND-OR graph · 0.2markov chain monte carlo · 0.2histogram of optical flow · 0.2dynamic programming · 0.2convolutional neural network · 0.2similarity identification · 0.1longitudinal study · 0.1content flow modeling · 0.1KSPC · 0.1
YearPublicationVenuePosition
2020 Vision-Based Gesture Recognition in Human-Robot Teams Using Synthetic Data
abstract
Building successful collaboration between humans and robots requires efficient, effective, and natural communication. Here we study a RGB-based deep learning approach for controlling robots through gestures (e.g., "follow me"). To address the challenge of collecting high-quality annotated data from human subjects, synthetic data is considered for this domain. We contribute a dataset of gestures that includes real videos with human subjects and synthetic videos from our custom simulator. A solution is presented for gesture recognition based on the state-of-the-art I3D model. Comprehensive testing was conducted to optimize the parameters for this model. Finally, to gather insight on the value of synthetic data, several experiments are described that systematically study the properties of synthetic data (e.g., gesture variations, character variety, generalization to new gestures). We discuss practical implications for the design of effective human-robot collaboration and the usefulness of synthetic data for deep learning.
Celso de Melo, Brandon Rothrock, Prudhvi Gurram, Oytun Ulutan, B. S. Manjunath
IROS2
2018 Human Causal Transfer: Challenges for Deep Reinforcement Learning
Mark Edmonds, James Kubricht, Colin Summers, Yixin Zhu 0001, Brandon Rothrock, Song-Chun Zhu, Hongjing Lu
CogSci5
2018 Unsupervised Learning of Hierarchical Models for Hand-Object Interactions
abstract
Contact forces of the hand are visually unobservable, but play a crucial role in understanding hand-object interactions. In this paper, we propose an unsupervised learning approach for manipulation event segmentation and manipulation event parsing. The proposed framework incorporates hand pose kinematics and contact forces using a low-cost easy-to-replicate tactile glove. We use a temporal grammar model to capture the hierarchical structure of events, integrating extracted force vectors from the raw sensory input of poses and forces. The temporal grammar is represented as a temporal And-Or graph (T-AOG), which can be induced in an unsupervised manner. We obtain the event labeling sequences by measuring the similarity between segments using the Dynamic Time Alignment Kernel (DTAK). Experimental results show that our method achieves high accuracy in manipulation event segmentation, recognition and parsing by utilizing both pose and force data.
Xu Xie 0001, Hangxin Liu, Mark Edmonds, Feng Gao 0013, Siyuan Qi, Yixin Zhu 0001, Brandon Rothrock, Song-Chun Zhu
ICRA7
2017 Privacy-Preserving Human Activity Recognition from Extreme Low Resolution
abstract
Privacy protection from surreptitious video recordings is an important societal challenge. We desire a computer vision system (e.g., a robot) that can recognize human activities and assist our daily life, yet ensure that it is not recording video that may invade our privacy. This paper presents a fundamental approach to address such contradicting objectives: human activity recognition while only using extreme low-resolution (e.g., 16x12) anonymized videos. We introduce the paradigm of inverse super resolution (ISR), the concept of learning the optimal set of image transformations to generate multiple low-resolution (LR) training videos from a single video. Our ISR learns different types of sub-pixel transformations optimized for the activity classification, allowing the classifier to best take advantage of existing high-resolution videos (e.g., YouTube videos) by creating multiple LR training videos tailored for the problem. We experimentally confirm that the paradigm of inverse super resolution is able to benefit activity recognition from extreme low-resolution videos.
Michael S. Ryoo, Brandon Rothrock, Charles Fleming, Hyun Jong Yang
AAAI2
2017 Feeling the force: Integrating force and pose for fluent discovery through imitation learning to open medicine bottles
abstract
Learning complex robot manipulation policies for real-world objects is challenging, often requiring significant tuning within controlled environments. In this paper, we learn a manipulation model to execute tasks with multiple stages and variable structure, which typically are not suitable for most robot manipulation approaches. The model is learned from human demonstration using a tactile glove that measures both hand pose and contact forces. The tactile glove enables observation of visually latent changes in the scene, specifically the forces imposed to unlock the child-safety mechanisms of medicine bottles. From these observations, we learn an action planner through both a top-down stochastic grammar model (And-Or graph) to represent the compositional nature of the task sequence and a bottom-up discriminative model from the observed poses and forces. These two terms are combined during planning to select the next optimal action. We present a method for transferring this human-specific knowledge onto a robot platform and demonstrate that the robot can perform successful manipulations of unseen objects with similar task structure.
Mark Edmonds, Feng Gao 0013, Xu Xie 0001, Hangxin Liu, Siyuan Qi, Yixin Zhu 0001, Brandon Rothrock, Song-Chun Zhu
IROS7
2017 A glove-based system for studying hand-object manipulation via joint pose and force sensing
abstract
We present a design of an easy-to-replicate glove-based system that can reliably perform simultaneous hand pose and force sensing in real time, for the purpose of collecting human hand data during fine manipulative actions. The design consists of a sensory glove that is capable of jointly collecting data of finger poses, hand poses, as well as forces on palm and each phalanx. Specifically, the sensory glove employs a network of 15 IMUs to measure the rotations between individual phalanxes. Hand pose is then reconstructed using forward kinematics. Contact forces on the palm and each phalanx are measured by 6 customized force sensors made from Velostat, a piezoresistive material whose force-voltage relation is investigated. We further develop an open-source software pipeline consisting of drivers and processing code and a system for visualizing hand actions that is compatible with the popular Raspberry Pi architecture. In our experiment, we conduct a series of evaluations that quantitatively characterize both individual sensors and the overall system, proving the effectiveness of the proposed design.
Hangxin Liu, Xu Xie 0001, Matt Millar, Mark Edmonds, Feng Gao 0013, Yixin Zhu 0001, Veronica J. Santos, Brandon Rothrock, Song-Chun Zhu
IROS8
2015 Pooled motion features for first-person videos
abstract
In this paper, we present a new feature representation for first-person videos. In first-person video understanding (e.g., activity recognition), it is very important to capture both entire scene dynamics (i.e., egomotion) and salient local motion observed in videos. We describe a representation framework based on time series pooling, which is designed to ab] short-term/long-term changes in feature descriptor elements. The idea is to keep track of how descriptor values are changing over time and summarize them to represent motion in the activity video. The framework is general, handling any types of per-frame feature descriptors including conventional motion descriptors like histogram of optical flows (HOF) as well as appearance descriptors from more recent convolutional neural networks (CNN). We experimentally confirm that our approach clearly outperforms previous feature representations including bag-of-visual-words and improved Fisher vector (IFV) when using identical underlying feature descriptors. We also confirm that our feature representation has superior performance to existing state-of-the-art features like local spatio-temporal features and Improved Trajectory Features (originally developed for 3rd-person videos) when handling first-person videos. Multiple first-person activity datasets were tested under various settings to confirm these findings.
Michael S. Ryoo, Brandon Rothrock, Larry H. Matthies
CVPR2
2015 Joint inference of groups, events and human roles in aerial videos
abstract
With the advent of drones, aerial video analysis becomes increasingly important; yet, it has received scant attention in the literature. This paper addresses a new problem of parsing low-resolution aerial videos of large spatial areas, in terms of 1) grouping, 2) recognizing events and 3) assigning roles to people engaged in events. We propose a novel framework aimed at conducting joint inference of the above tasks, as reasoning about each in isolation typically fails in our setting. Given noisy tracklets of people and detections of large objects and scene surfaces (e.g., building, grass), we use a spatiotemporal AND-OR graph to drive our joint inference, using Markov Chain Monte Carlo and dynamic programming. We also introduce a new formalism of spatiotemporal templates characterizing latent sub-events. For evaluation, we have collected and released a new aerial videos dataset using a hex-rotor flying over picnic areas rich with group events. Our results demonstrate that we successfully address above inference tasks under challenging conditions.
Tianmin Shu, Dan Xie 0005, Brandon Rothrock, Sinisa Todorovic, Song-Chun Zhu
CVPR3
2013 Integrating Grammar and Segmentation for Human Pose Estimation
abstract
In this paper we present a compositional and-or graph grammar model for human pose estimation. Our model has three distinguishing features: (i) large appearance differences between people are handled compositionally by allowing parts or collections of parts to be substituted with alternative variants, (ii) each variant is a sub-model that can define its own articulated geometry and context-sensitive compatibility with neighboring part variants, and (iii) background region segmentation is incorporated into the part appearance models to better estimate the contrast of a part region from its surroundings, and improve resilience to background clutter. The resulting integrated framework is trained discriminatively in a max-margin framework using an efficient and exact inference algorithm. We present experimental evaluation of our model on two popular datasets, and show performance improvements over the state-of-art on both benchmarks.
Brandon Rothrock, Song-Chun Zhu
CVPR1
2006 UNIFORM: automatically generating consistent remote control user interfaces
abstract
A problem with many of today's appliance interfaces is that they are inconsistent. For example, the procedure for setting the time on alarm clocks and VCRs differs, even among different models made by the same manufacturer. Finding particular functions can also be a challenge, because appliances often organize their features differently. This paper presents a system, called Uniform, which approaches this problem by automatically generating remote control interfaces that take into account previous interfaces that the user has seen during the generation process. Uniform is able to automatically identify similarities between different devices and users may specify additional similarities. The similarity information allows the interface generator to use the same type of controls for similar functions, place similar functions so that they can be found with the same navigation steps, and create interfaces that have a similar visual appearance.
Jeffrey Nichols 0001, Brad A. Myers, Brandon Rothrock
CHI3
2006 Few-key text entry revisited: mnemonic gestures on four keys
abstract
We present a new 4-key text entry method that, unlike most few-key methods, is gestural instead of selection-based. Importantly, its gestures mimic the writing of Roman letters for high learnability. We compare this new 4-key method to predominant 3-key and 5-key methods theoretically using KSPC and empirically using a longitudinal study of 5 subjects over 10 sessions. The study includes an evaluation of the 4-key method without any on-screen visualization-an impossible condition for the selection-based methods. Our results show that the new 4-key method is quickly learned, becoming faster than the 3-key and 5-key methods after just ~10 minutes of writing, although it produces more errors. Interestingly, removing a visualization of the gestures being made causes no detriment to the 4-key method, which is an advantage for eyes-free text entry.
Jacob O. Wobbrock, Brad A. Myers, Brandon Rothrock
CHI3
2006 Scheduling with Uncertain Resources: Collaboration with the User
abstract
We describe a scheduling system that supports collaboration between the user and automated optimizer. It enables the user to monitor the optimizer decisions, make any of the decisions manually, and leave the other decisions to the system. Furthermore, it identifies the tasks that require the user's participation, and asks for assistance with these tasks.
Eugene Fink, Ulas Bardak, Brandon Rothrock, Jaime G. Carbonell
SMC3
2006 Huddle: automatically generating interfaces for systems of multiple connected appliances
abstract
Systems of connected appliances, such as home theaters and presentation rooms, are becoming commonplace in our homes and workplaces. These systems are often difficult to use, in part because users must determine how to split the tasks they wish to perform into sub-tasks for each appliance and then find the particular functions of each appliance to complete their sub-tasks. This paper describes Huddle, a new system that automatically generates task-based interfaces for a system of multiple appliances based on models of the content flow within the multi-appliance system.
Jeffrey Nichols 0001, Brandon Rothrock, Polo Chau, Brad A. Myers
UIST2