Randi Cabezas

dblp:150/4244 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Virtual and augmented reality · 70% Computer animation and physical simulation · 19% Computational photography and imaging · 8%
Artificial intelligence
4 papers
3D vision · 32% Face, body and person analysis · 29% Probabilistic and Bayesian machine learning · 19%
Human-computer interaction and pervasive computing
1 paper
Interaction techniques and input · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Virtual and augmented reality › immersive interaction
hand tracking
1.022022
UmeTrack: Unified multi-view end-to-end hand tracking for VR · SIGGRAPH Asia 2022
MEgATrack: monochrome egocentric articulated hand-tracking for virtual reality · ACM Trans. Graph. 2020
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation
0.612022
UmeTrack: Unified multi-view end-to-end hand tracking for VR · SIGGRAPH Asia 2022
Virtual and augmented reality › immersive interaction
VR interaction
0.612022
UmeTrack: Unified multi-view end-to-end hand tracking for VR · SIGGRAPH Asia 2022
Computer animation and physical simulation
character animation
0.412020
MEgATrack: monochrome egocentric articulated hand-tracking for virtual reality · ACM Trans. Graph. 2020
Computer vision › 3D vision
3d scene reconstruction
0.212015
Semantically-Aware Aerial Reconstruction from Multi-modal Data · ICCV 2015
Computer vision › 3D vision › 3d scene reconstruction
aerial 3d reconstruction
0.212015
Semantically-Aware Aerial Reconstruction from Multi-modal Data · ICCV 2015
Machine learning › Generative modeling › generative model
probabilistic generative model
0.212015
Semantically-Aware Aerial Reconstruction from Multi-modal Data · ICCV 2015
Computer vision › 3D vision
3d reconstruction
0.212014
Aerial Reconstructions via Probabilistic Data Fusion · CVPR 2014
Machine learning › Probabilistic and Bayesian machine learning
bayesian data fusion
0.212014
Aerial Reconstructions via Probabilistic Data Fusion · CVPR 2014
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian nonparametric model
0.212014
Bayesian Nonparametric Intrinsic Image Decomposition · ECCV (4) 2014
Computer vision › Vision and language
multimodal fusion
0.212014
Aerial Reconstructions via Probabilistic Data Fusion · CVPR 2014
Computational photography and imaging
intrinsic image decomposition
0.212014
Bayesian Nonparametric Intrinsic Image Decomposition · ECCV (4) 2014
Interaction techniques and input
gesture input
0.212022
UmeTrack: Unified multi-view end-to-end hand tracking for VR · SIGGRAPH Asia 2022
Computational science and engineering
parallel computing
0.112014
Aerial Reconstructions via Probabilistic Data Fusion · CVPR 2014

Methods — techniques the papers use, named apart from their topics

neural network · 2.2end-to-end differentiable framework · 1.7probabilistic generative model · 0.4LiDAR fusion · 0.4keypoint estimation · 0.4detection-by-tracking · 0.4probabilistic graphical model · 0.4dense reconstruction · 0.4bayesian nonparametric modeling · 0.4
YearPublicationVenuePosition
2022 UmeTrack: Unified multi-view end-to-end hand tracking for VR
abstract
Real-time tracking of 3D hand pose in world space is a challenging problem and plays an important role in VR interaction. Existing work in this space are limited to either producing root-relative (versus world space) 3D pose or rely on multiple stages such as generating heatmaps and kinematic optimization to obtain 3D pose. Moreover, the typical VR scenario, which involves multi-view tracking from wide field of view (FOV) cameras is seldom addressed by these methods. In this paper, we present a unified end-to-end differentiable framework for multi-view, multi-frame hand tracking that directly predicts 3D hand pose in world space. We demonstrate the benefits of end-to-end differentiabilty by extending our framework with downstream tasks such as jitter reduction and pinch prediction. To demonstrate the efficacy of our model, we further present a new large-scale egocentric hand pose dataset that consists of both real and synthetic data. Experiments show that our system trained on this dataset handles various challenging interactive motions, and has been successfully applied to real-time VR applications.
Shangchen Han, Po-Chen Wu, Linguang Zhang, Weiguang Si, Peizhao Zhang, Yujun Cai, Tomas Hodan, Randi Cabezas, Luan Tran, Muzaffer Akbay, Tsz-Ho Yu, Cem Keskin, Robert Wang 0002
SIGGRAPH Asia11
2020 MEgATrack: monochrome egocentric articulated hand-tracking for virtual reality
abstract
We present a system for real-time hand-tracking to drive virtual and augmented reality (VR/AR) experiences. Using four fisheye monochrome cameras, our system generates accurate and low-jitter 3D hand motion across a large working volume for a diverse set of users. We achieve this by proposing neural network architectures for detecting hands and estimating hand keypoint locations. Our hand detection network robustly handles a variety of real world environments. The keypoint estimation network leverages tracking history to produce spatially and temporally consistent poses. We design scalable, semi-automated mechanisms to collect a large and diverse set of ground truth data using a combination of manual annotation and automated tracking. Additionally, we introduce a detection-by-tracking method that increases smoothness while reducing the computational cost; the optimized system runs at 60Hz on PC and 30Hz on a mobile processor. Together, these contributions yield a practical system for capturing a user's hands and is the default feature on the Oculus Quest VR headset powering input and social presence.
Shangchen Han, Randi Cabezas, Christopher D. Twigg, Peizhao Zhang, Jeff Petkau, Tsz-Ho Yu, Chun-Jung Tai, Muzaffer Akbay, Asaf Nitzan, Gang Dong, Yuting Ye, Lingling Tao, Chengde Wan, Robert Wang 0002
ACM Trans. Graph.3
2015 Semantically-Aware Aerial Reconstruction from Multi-modal Data
abstract
We consider a methodology for integrating multiple sensors along with semantic information to enhance scene representations. We propose a probabilistic generative model for inferring semantically-informed aerial reconstructions from multi-modal data within a consistent mathematical framework. The approach, called Semantically-Aware Aerial Reconstruction (SAAR), not only exploits inferred scene geometry, appearance, and semantic observations to obtain a meaningful categorization of the data, but also extends previously proposed methods by imposing structure on the prior over geometry, appearance, and semantic labels. This leads to more accurate reconstructions and the ability to fill in missing contextual labels via joint sensor and semantic information. We introduce a new multi-modal synthetic dataset in order to provide quantitative performance analysis. Additionally, we apply the model to real-world data and exploit OpenStreetMap as a source of semantic observations. We show quantitative improvements in reconstruction accuracy of large-scale urban scenes from the combination of LiDAR, aerial photography, and semantic data. Furthermore, we demonstrate the model's ability to fill in for missing sensed data, leading to more interpretable reconstructions.
Randi Cabezas, Julian Straub, John W. Fisher III
ICCV1
2014 Aerial Reconstructions via Probabilistic Data Fusion
abstract
We propose an integrated probabilistic model for multi-modal fusion of aerial imagery, LiDAR data, and (optional) GPS measurements. The model allows for analysis and dense reconstruction (in terms of both geometry and appearance) of large 3D scenes. An advantage of the approach is that it explicitly models uncertainty and allows for missing data. As compared with image-based methods, dense reconstructions of complex urban scenes are feasible with fewer observations. Moreover, the proposed model allows one to estimate absolute scale and orientation and reason about other aspects of the scene, e.g., detection of moving objects. As formulated, the model lends itself to massively-parallel computing. We exploit this in an efficient inference scheme that utilizes both general purpose and domain-specific hardware components. We demonstrate results on large-scale reconstruction of urban terrain from LiDAR and aerial photography data.
Randi Cabezas, Oren Freifeld, Guy Rosman, John W. Fisher III
CVPR1
2014 Bayesian Nonparametric Intrinsic Image Decomposition
Jason Chang 0001, Randi Cabezas, John W. Fisher III
ECCV (4)2