EDBT 2026 Demo / reviewers in the wild / expert
Matthew Turk 0001
dblp:61/3381 · also Matthew A. Turk
· DBLP profile ↗
93ranked-venue papers
8as first author
2since 2021 · last 2024
0000-0002-4198-8401ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 48 · 6 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 17 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
29 papers |
3D vision · 43% Generative modeling · 9% Robot navigation and mapping · 8% | |
| Computer graphics and multimedia
18 papers |
Virtual and augmented reality · 54% Computational photography and imaging · 22% Rendering · 13% | |
| Human-computer interaction and pervasive computing
5 papers |
Collaborative and social computing · 43% Immersive interaction · 31% Interaction techniques and input · 26% |
Topics — the 30 heaviest of 79, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Virtual and augmented reality › augmented reality
collaborative augmented reality |
0.9 | 4 | 2016 | Anchoring 2D gesture annotations in augmented reality · VR 2016 PPV: Pixel-Point-Volume Segmentation for Object Referencing in Collaborative Augmented Reality · ISMAR 2016 2D-3D Co-segmentation for AR-based Remote Collaboration · ISMAR 2015 |
Computer vision › 3D vision
structure from motion |
0.9 | 5 | 2015 | Theia: A Fast and Scalable Structure-from-Motion Library · ACM Multimedia 2015 Optimizing the Viewing Graph for Structure-from-Motion · ICCV 2015 Computing similarity transformations from only image correspondences · CVPR 2015 |
Machine learning › Trustworthy machine learning › fairness › fairness in generative models
bias in generative models |
0.8 | 1 | 2024 | TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models · ECCV (79) 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models · ECCV (79) 2024 |
Computer vision › 3D vision
camera pose estimation |
0.6 | 3 | 2015 | Efficient Computation of Absolute Pose for Gravity-Aware Augmented Reality · ISMAR 2015 Computing similarity transformations from only image correspondences · CVPR 2015 gDLS: A Scalable Solution to the Generalized Pose and Scale Problem · ECCV (4) 2014 |
Robotics › Robot navigation and mapping
SLAM |
0.5 | 3 | 2014 | Model Estimation and Selection towardsUnconstrained Real-Time Tracking and Mapping · IEEE Trans. Vis. Comput. Graph. 2014 Improved outdoor augmented reality through "Globalization" · ISMAR 2013 Live tracking and mapping from both general and rotation-only camera motion · ISMAR 2012 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.4 | 4 | 2017 | User-Perspective AR Magic Lens from Gradient-Based IBR and Semi-Dense Stereo · IEEE Trans. Vis. Comput. Graph. 2017 Multiflash Stereopsis: Depth-Edge-Preserving Stereo with Small Baseline Illumination · IEEE Trans. Pattern Anal. Mach. Intell. 2008 Discontinuity Preserving Stereo with Small Baseline Multi-Flash Illumination · ICCV 2005 |
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue |
0.4 | 1 | 2019 | What Should I Ask? Using Conversationally Informative Rewards for Goal-oriented Visual Dialog · ACL (1) 2019 |
Natural language and speech › Question answering and dialogue systems
visual dialog |
0.4 | 1 | 2019 | What Should I Ask? Using Conversationally Informative Rewards for Goal-oriented Visual Dialog · ACL (1) 2019 |
Machine learning › Probabilistic and Bayesian machine learning
decision boundary estimation |
0.3 | 1 | 2018 | CLEAR: Cumulative LEARning for One-Shot One-Class Image Recognition · CVPR 2018 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.3 | 1 | 2018 | CLEAR: Cumulative LEARning for One-Shot One-Class Image Recognition · CVPR 2018 |
Machine learning › Time series and sequential data › anomaly detection
one-class classification |
0.3 | 1 | 2018 | CLEAR: Cumulative LEARning for One-Shot One-Class Image Recognition · CVPR 2018 |
Computer vision › 3D vision › stereo vision › stereo matching
semi-dense stereo |
0.3 | 1 | 2017 | User-Perspective AR Magic Lens from Gradient-Based IBR and Semi-Dense Stereo · IEEE Trans. Vis. Comput. Graph. 2017 |
Rendering
image-based rendering |
0.3 | 1 | 2017 | User-Perspective AR Magic Lens from Gradient-Based IBR and Semi-Dense Stereo · IEEE Trans. Vis. Comput. Graph. 2017 |
Collaborative and social computing
remote collaboration |
0.3 | 2 | 2016 | World-stabilized annotations and virtual scene navigation for remote collaboration · UIST 2014 Anchoring 2D gesture annotations in augmented reality · VR 2016 |
Computer vision › Segmentation and scene understanding
3d segmentation |
0.2 | 1 | 2016 | PPV: Pixel-Point-Volume Segmentation for Object Referencing in Collaborative Augmented Reality · ISMAR 2016 |
Computer vision › Segmentation and scene understanding › 3d segmentation
supervoxel segmentation |
0.2 | 1 | 2016 | PPV: Pixel-Point-Volume Segmentation for Object Referencing in Collaborative Augmented Reality · ISMAR 2016 |
Computer vision › 3D vision › camera pose estimation
absolute pose estimation |
0.2 | 1 | 2015 | Efficient Computation of Absolute Pose for Gravity-Aware Augmented Reality · ISMAR 2015 |
Computer vision › Segmentation and scene understanding › image segmentation
co-segmentation |
0.2 | 1 | 2015 | 2D-3D Co-segmentation for AR-based Remote Collaboration · ISMAR 2015 |
Computer vision › 3D vision › structure from motion
large-scale reconstruction |
0.2 | 1 | 2015 | Theia: A Fast and Scalable Structure-from-Motion Library · ACM Multimedia 2015 |
Computer vision › 3D vision › multi-view geometry
view graph optimization |
0.2 | 1 | 2015 | Optimizing the Viewing Graph for Structure-from-Motion · ICCV 2015 |
Computer vision › 3D vision › camera pose estimation
camera tracking |
0.2 | 1 | 2014 | Model Estimation and Selection towardsUnconstrained Real-Time Tracking and Mapping · IEEE Trans. Vis. Comput. Graph. 2014 |
Robotics › Robot navigation and mapping › SLAM › visual SLAM
keyframe-based SLAM |
0.2 | 1 | 2014 | Model Estimation and Selection towardsUnconstrained Real-Time Tracking and Mapping · IEEE Trans. Vis. Comput. Graph. 2014 |
Immersive interaction › augmented reality
augmented reality annotation |
0.2 | 1 | 2014 | World-stabilized annotations and virtual scene navigation for remote collaboration · UIST 2014 |
Virtual and augmented reality
augmented reality |
0.2 | 3 | 2017 | User-Perspective AR Magic Lens from Gradient-Based IBR and Semi-Dense Stereo · IEEE Trans. Vis. Comput. Graph. 2017 Improved outdoor augmented reality through "Globalization" · ISMAR 2013 Live tracking and mapping from both general and rotation-only camera motion · ISMAR 2012 |
Computational photography and imaging › illumination analysis
computational illumination |
0.2 | 2 | 2009 | A projector-camera setup for geometry-invariant frequency demultiplexing · CVPR 2009 Characterizing the shadow space of camera-light pairs · CVPR 2008 |
Computer vision › 3D vision
depth estimation |
0.2 | 2 | 2008 | Multiflash Stereopsis: Depth-Edge-Preserving Stereo with Small Baseline Illumination · IEEE Trans. Pattern Anal. Mach. Intell. 2008 Characterizing the shadow space of camera-light pairs · CVPR 2008 |
Machine learning › Generative modeling › diffusion model
guided sampling |
0.2 | 1 | 2013 | SWIGS: A Swift Guided Sampling Method · CVPR 2013 |
Computer vision › 3D vision › robust estimation
RANSAC |
0.2 | 1 | 2013 | SWIGS: A Swift Guided Sampling Method · CVPR 2013 |
Computer vision › 3D vision › robust estimation
robust model estimation |
0.2 | 1 | 2013 | SWIGS: A Swift Guided Sampling Method · CVPR 2013 |
Methods — techniques the papers use, named apart from their topics
vertical vanishing point detection · 0.7patchmatch · 0.6user study · 0.5surface normals · 0.5graph cuts · 0.5gesture classification · 0.5energy function · 0.5reinforcement learning · 0.4rational speech act · 0.4information gain · 0.4visual tracking · 0.4geometric robust information criterion · 0.3transfer learning · 0.3convolutional neural network · 0.3graph cut · 0.2structure-from-motion pipeline · 0.2quadratic eigenvalue problem · 0.2IMU sensing · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models
Aditya Chinchure, Pushkar Shukla, Gaurav Bhatt, Kiri Salij, Kartik Hosanagar, Leonid Sigal, Matthew Turk 0001 |
ECCV (79) | 7 |
| 2021 | The 5th Recognizing Families in the Wild Data Challenge: Predicting Kinship from FacesabstractRecognizing Families In the Wild (RFIW), held as a data challenge in conjunction with the 16thIEEE International Conference on Automatic Face and Gesture Recognition (FG), is a large-scale, multi-track visual kinship recognition evaluation. For the fifth edition of RFIW, we continue to attract scholars, bring together professionals, publish new work, and discuss prospects. In this paper, we summarize submissions for the three tasks of this year's RFIW: specifically, we review the results for kinship verification, tri-subject verification, and family member search and retrieval. We look at the RFIW problem, share current efforts, and make recommendations for promising future directions. Joseph P. Robinson, Can Qin, Ming Shao, Matthew Turk 0001, Rama Chellappa, Yun Fu 0001 |
FG | 4 |
| 2020 | BLT: Balancing Long-Tailed Datasets with Adversarially-Perturbed Images
Jedrzej Kozerawski, Victor Fragoso, Nikolaos Karianakis, Gaurav Mittal, Matthew Turk 0001 |
ACCV (3) | 5 |
| 2020 | Recognizing Families In the Wild (RFIW): The 4th EditionabstractRecognizing Families In the Wild (RFIW)- an annual large-scale, multi-track automatic kinship recognition evaluation- supports various visual kin-based problems on scales much higher than ever before. Organized in conjunction with the as a Challenge, RFIW provides a platform for publishing original work and the gathering of experts for a discussion of the next steps. This paper summarizes the supported tasks (i.e., kinship verification, tri-subject verification, and search & retrieval of missing children) in the evaluation protocols, which include the practical motivation, technical background, data splits, metrics, and benchmark results. Furthermore, top submissions (i.e., leader-board stats) are listed and reviewed as a high-level analysis on the state of the problem. In the end, the purpose of this paper is to describe the 2020 RFIW challenge, end-to-end, along with forecasts in promising future directions. Joseph P. Robinson, Yu Yin 0001, Zaid Khan 0001, Ming Shao, Si-Yu Xia, Michael Stopa, Samson Timoner, Matthew Turk 0001, Rama Chellappa, Yun Fu 0001 |
FG | 8 |
| 2019 | What Should I Ask? Using Conversationally Informative Rewards for Goal-oriented Visual DialogabstractThe ability to engage in goal-oriented conversations has allowed humans to gain knowledge, reduce uncertainty, and perform tasks more efficiently.Artificial agents, however, are still far behind humans in having goaldriven conversations.In this work, we focus on the task of goal-oriented visual dialogue, aiming to automatically generate a series of questions about an image with a single objective.This task is challenging, since these questions must not only be consistent with a strategy to achieve a goal, but also consider the contextual information in the image.We propose an end-to-end goal-oriented visual dialogue system, that combines reinforcement learning with regularized information gain.Unlike previous approaches that have been proposed for the task, our work is motivated by the Rational Speech Act framework, which models the process of human inquiry to reach a goal.We test the two versions of our model on the GuessWhat?! dataset, obtaining significant results that outperform the current state-of-the-art models in the task of generating questions to find an undisclosed object in an image. Pushkar Shukla, Carlos E. L. Elmadjian, Richika Sharan, Vivek Kulkarni, Matthew Turk 0001, William Yang Wang |
ACL (1) | 5 |
| 2019 | Multimodal Classification of EEG During Physical ActivityabstractBrain Computer Interfaces (BCIs) typically utilize electroencephalography (EEG) to enable control of a computer through brain signals. However, EEG is susceptible to a large amount of noise, especially from muscle activity, making it difficult to use in ubiquitous computing environments where mobility and physicality are important features. In this work, we present a novel multimodal approach for classifying the P300 event related potential (ERP) component by coupling EEG signals with nonscalp electrodes (NSE) that measure ocular and muscle artifacts. We demonstrate the effectiveness of our approach on a new dataset where the P300 signal was evoked with participants on a stationary bike under three conditions of physical activity: rest, low-intensity, and high-intensity exercise. We show that intensity of physical activity impacts the performance of both our proposed model and existing state-of-the-art models. After incorporating signals from nonscalp electrodes our proposed model performs significantly better for the physical activity conditions. Our results suggest that the incorporation of additional modalities related to eye-movements and muscle activity may improve the efficacy of mobile EEG-based BCI systems, creating the potential for ubiquitous BCI. Yi Ding 0010, Brandon Huynh, Aiwen Xu, Tom Bullock, Hubert Cecotti, Matthew Turk 0001, Barry Giesbrecht, Tobias Höllerer |
ICMI | 6 |
| 2018 | CLEAR: Cumulative LEARning for One-Shot One-Class Image RecognitionabstractThis work addresses the novel problem of one-shot one-class classification. The goal is to estimate a classification decision boundary for a novel class based on a single image example. Our method exploits transfer learning to model the transformation from a representation of the input, extracted by a Convolutional Neural Network, to a classification decision boundary. We use a deep neural network to learn this transformation from a large labelled dataset of images and their associated class decision boundaries generated from ImageNet, and then apply the learned decision boundary to classify subsequent query images. We tested our approach on several benchmark datasets and significantly outperformed the baseline methods. Jedrzej Kozerawski, Matthew Turk 0001 |
CVPR | 2 |
| 2018 | Gaze and head pointing for hands-free text entry: applicability to ultra-small virtual keyboardsabstractWith the proliferation of small-screen computing devices, there has been a continuous trend in reducing the size of interface elements. In virtual keyboards, this allows for more characters in a layout and additional function widgets. However, vision-based interfaces (VBIs) have only been investigated with large (e.g., full-screen) keyboards. To understand how key size reduction affects the accuracy and speed performance of text entry VBIs, we evaluated gaze-controlled VBI (g-VBI) and head-controlled VBI (h-VBI) with unconventionally small (0.4°, 0.6°, 0.8° and 1°) keys. Novices (N = 26) yielded significantly more accurate and fast text production with h-VBI than with g-VBI, while the performance of experts (N = 12) for both VBIs was nearly equal when a 0.8--1° key size was used. We discuss advantages and limitations of the VBIs for typing with ultra-small keyboards and emphasize relevant factors for designing such systems. Yulia Gizatdinova, Oleg Spakov, Outi Tuisku, Matthew Turk 0001, Veikko Surakka |
ETRA | 4 |
| 2018 | Hybrid orbiting-to-photos in 3D reconstructed visual realityabstractVirtually navigating through photos from a 3D image-based reconstruction has recently become very popular in many applications. In this paper, we consider a particular virtual travel maneuver that is important for this type of virtual navigation---orbiting to photos that can see a point-of-interest (POI). The main challenge with this particular type of orbiting is how to give appropriate feedback to the user regarding the existence and information of each photo in 3D while allowing the user to manipulate three degrees-of-freedom (DoF) for orbiting around the POI. We present a hybrid approach that combines features from two baselines---proxy plane and thumbnail approaches. Experimental results indicate that users rated our hybrid approach more favorably for several qualitative questionnaire statements, and that the hybrid approach is preferred over both baselines for outdoor scenes. Benjamin Nuernberger, Tobias Höllerer, Matthew Turk 0001 |
VRST | 3 |
| 2018 | Illumination for 360 degree camerasabstractAdditional illumination improves the capture of omnidirectional 360° video and images, especially for dark or high-contrast environments. There is no "behind" for 360° cameras, so the placement of lights is a problem. We explore ways to position lights on some 360° cameras, and propose two good locations. Ismo Rakkolainen, Roope Raisamo, Matthew Turk 0001, Tobias Höllerer |
VRST | 3 |
| 2017 | ANSAC: Adaptive Non-Minimal Sample and Consensus
Victor Fragoso, Chris Sweeney, Pradeep Sen, Matthew Turk 0001 |
BMVC | 4 |
| 2017 | Evaluating snapping-to-photos virtual travel interfaces for 3D reconstructed visual realityabstractNavigating through a virtual, 3D reconstructed scene has recently become very important in many applications. A popular approach is to virtually travel to the photos used in reconstructing the scene; such an approach may be generally termed a "snapping-to-photos" virtual travel interface. While previous work has either used fully constrained interfaces (always at the photos) or minimally constrained interfaces (free-flight navigation), in this paper we introduce new snapping-to-photos interfaces that lie in between these two extremes. Our snapping-to-photos interfaces snap the view to a photo in 3D based on viewpoint similarity and optionally the user's mouse cursor or finger-tap position. Experimental results, with both indoor and outdoor scene reconstructions, found that our snapping-to-photos interfaces are preferred over the baseline fully constrained-to-photos interface, that there exist differences between indoor and outdoor scenes, and that users preferred and were able to reach target photos better with click-to-snap point-of-interest snapping compared to automatic point-of-view snapping. Benjamin Nuernberger, Matthew Turk 0001, Tobias Höllerer |
VRST | 2 |
| 2017 | Densification of Semi-Dense Reconstructions for Novel View Generation of Live ScenesabstractIn this paper, we consider the problem of rendering novel views of a live unprepared scene from video input, important to many application scenarios (such as telepresence and remote collaboration). We present an optimization approach to improving incomplete scene reconstructions captured in real time with a single moving monocular camera. We take semi-dense depth maps and convert them into a dense scene model, suitable for rendering plausible novel views of the scene using conventional image-based rendering. Our implementation densifies depth maps at the rate they are generated, and enables us to generate novel views of live scenes with no pre-capture or preprocessing. In evaluations comparing with other approaches, our method performs well even on difficult scenes, and results in higher-quality novel views. Domagoj Baricevic, Tobias Höllerer, Matthew Turk 0001 |
WACV | 3 |
| 2017 | User-Perspective AR Magic Lens from Gradient-Based IBR and Semi-Dense StereoabstractWe present a new approach to rendering a geometrically-correct user-perspective view for a magic lens interface, based on leveraging the gradients in the real world scene. Our approach couples a recent gradient-domain image-based rendering method with a novel semi-dense stereo matching algorithm. Our stereo algorithm borrows ideas from PatchMatch, and adapts them to semi-dense stereo. This approach is implemented in a prototype device build from off-the-shelf hardware, with no active depth sensing. Despite the limited depth data, we achieve high-quality rendering for the user-perspective magic lens. Domagoj Baricevic, Tobias Höllerer, Pradeep Sen, Matthew Turk 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2016 | Large Scale SfM with the Distributed Camera ModelabstractWe introduce the distributed camera model, a novel model for Structure-from-Motion (SfM). This model describes image observations in terms of light rays with ray origins and directions rather than pixels. As such, the proposed model is capable of describing a single camera or multiple cameras simultaneously as the collection of all light rays observed. We show how the distributed camera model is a generalization of the standard camera model and we describe a general formulation and solution to the absolute camera pose problem that works for standard or distributed cameras. The proposed method computes a solution that is up to 8 times more efficient and robust to rotation singularities in comparison with gDLS[21]. Finally, this method is used in an novel large-scale incremental SfM pipeline where distributed cameras are accurately and robustly merged together. This pipeline is a direct generalization of traditional incremental SfM, however, instead of incrementally adding one camera at a time to grow the reconstruction the reconstruction is grown by adding a distributed camera. Our pipeline produces highly accurate reconstructions efficiently by avoiding the need for many bundle adjustment iterations and is capable of computing a 3D model of Rome from over 15,000 images in just 22 minutes. Chris Sweeney, Victor Fragoso, Tobias Höllerer, Matthew Turk 0001 |
3DV | 4 |
| 2016 | One-class slab support vector machineabstractThis work introduces the one-class slab SVM (OCSSVM), a one-class classifier that aims at improving the performance of the one-class SVM. The proposed strategy reduces the false positive rate and increases the accuracy of detecting instances from novel classes. To this end, it uses two parallel hyperplanes to learn the normal region of the decision scores of the target class. OCSSVM extends one-class SVM since it can scale and learn non-linear decision functions via kernel methods. The experiments on two publicly available datasets show that OCSSVM can consistently outperform the one-class SVM and perform comparable to or better than other state-of-the-art one-class classifiers. Victor Fragoso, Walter J. Scheirer, João Pedro Hespanha, Matthew Turk 0001 |
ICPR | 4 |
| 2016 | PPV: Pixel-Point-Volume Segmentation for Object Referencing in Collaborative Augmented RealityabstractWe present a method for collaborative augmented reality (AR) that enables users from different viewpoints to interpret object references specified via 2D on-screen circling gestures. Based on a user's 2D drawing annotation, the method segments out the userselected object using an incomplete or imperfect scene model and the color image from the drawing viewpoint. Specifically, we propose a novel segmentation algorithm that utilizes both 2D and 3D scene cues, structured into a three-layer graph of pixels, 3D points, and volumes (supervoxels), solved via standard graph cut algorithms. This segmentation enables an appropriate rendering of the user's 2D annotation from other viewpoints in 3D augmented reality. Results demonstrate the superiority of the proposed method over existing methods. Kuo-Chin Lien, Benjamin Nuernberger, Tobias Höllerer, Matthew Turk 0001 |
ISMAR | 4 |
| 2016 | Anchoring 2D gesture annotations in augmented realityabstractAugmented reality enhanced collaboration systems often allow users to draw 2D gesture annotations onto video feeds to help collaborators to complete physical tasks. This works well for static cameras, but for movable cameras, perspective effects cause problems when trying to render 2D annotations from a new viewpoint in 3D. In this paper, we present a new approach towards solving this problem by using gesture enhanced annotations. By first classifying which type of gesture the user drew, we show that it is possible to render annotations in 3D in a way that conforms more to the original intention of the user than with traditional methods. We first determined a generic vocabulary of important 2D gestures for remote collaboration by running an Amazon Mechanical Turk study with 88 participants. Next, we designed a novel system to automatically handle the top two 2D gesture annotations - arrows and circles. Arrows are handled by identifying their anchor points and using surface normals for better perspective rendering. For circles, we designed a novel energy function to help infer the object of interest using both 2D image cues and 3D geometric cues. Results indicate that our approach outperforms previous methods in terms of better conveying the original drawing's meaning from different viewpoints. Benjamin Nuernberger, Kuo-Chin Lien, Tobias Höllerer, Matthew Turk 0001 |
VR | 4 |
| 2016 | Multi-view gesture annotations in image-based 3D reconstructed scenesabstractWe present a novel 2D gesture annotation method for use in image-based 3D reconstructed scenes with applications in collaborative virtual and augmented reality. Image-based reconstructions allow users to virtually explore a remote environment using image-based rendering techniques. To collaborate with other users, either synchronously or asynchronously, simple 2D gesture annotations can be used to convey spatial information to another user. Unfortunately, prior methods are either unable to disambiguate such 2D annotations in 3D from novel viewpoints or require relatively dense reconstructions of the environment. Benjamin Nuernberger, Kuo-Chin Lien, Lennon Grinta, Chris Sweeney, Matthew Turk 0001, Tobias Höllerer |
VRST | 5 |
| 2016 | A compact, wide-FOV optical design for head-mounted displaysabstractWe present a new optical design for head-mounted displays (HMD) which has an exceptionally wide field of view (FOV). It can cover even the full human FOV. It is based on seamless lenses and screens curved around the eyes. The proof-of-concept prototypes are promising, and one of them far exceeds the human FOV, although the effective FOV is limited by the anatomy of the human head. The presented optical design has advantages such as compactness, light weight, low cost and super-wide FOV with high resolution. Even though this is still work-in-progress and display functionality is not yet implemented, it suggests a feasible way to significantly expand the FOV of HMDs. Ismo Rakkolainen, Matthew Turk 0001, Tobias Höllerer |
VRST | 2 |
| 2015 | On Preserving Structure in Stereo Seam CarvingabstractThe major objective of image retargeting algorithms is to preserve the viewer's perception while adjusting the aspect ratio of an image. This means that an ideal retargeting algorithm has to be able to preserve high-level semantics and avoid generating low-level image distortion. Stereoscopic image retargeting poses a even more challenging problem in that the 3D perception has to be preserved as well. In this paper, we propose an algorithm based on high-order two-view co-labeling to simultaneously retarget a given stereo pair and preserve its 2D as well as 3D quality. Our experimental results qualitatively demonstrate the improved ability of preserving 2D image structures in both views. In addition, we show quantitatively that our algorithm improves upon the state-of-the-art up to 85% in terms of a measurement based on depth distortion. Kuo-Chin Lien, Matthew Turk 0001 |
3DV | 2 |
| 2015 | Computing similarity transformations from only image correspondencesabstractWe propose a novel solution for computing the relative pose between two generalized cameras that includes reconciling the internal scale of the generalized cameras. This approach can be used to compute a similarity transformation between two coordinate systems, making it useful for loop closure in visual odometry and registering multiple structure from motion reconstructions together. In contrast to alternative similarity transformation methods, our approach uses 2D-2D image correspondences thus is not subject to the depth uncertainty that often arises with 3D points. We utilize a known vertical direction (which may be easily obtained from IMU data or vertical vanishing point detection) of the generalized cameras to solve the generalized relative pose and scale problem as an efficient Quadratic Eigenvalue Problem. To our knowledge, this is the first method for computing similarity transformations that does not require any 3D information. Our experiments on synthetic and real data demonstrate that this leads to improved performance compared to methods that use 3D-3D or 2D-3D correspondences, especially as the depth of the scene increases. Chris Sweeney, Laurent Kneip, Tobias Höllerer, Matthew Turk 0001 |
CVPR | 4 |
| 2015 | Optimizing the Viewing Graph for Structure-from-MotionabstractThe viewing graph represents a set of views that are related by pairwise relative geometries. In the context of Structure-from-Motion (SfM), the viewing graph is the input to the incremental or global estimation pipeline. Much effort has been put towards developing robust algorithms to overcome potentially inaccurate relative geometries in the viewing graph during SfM. In this paper, we take a fundamentally different approach to SfM and instead focus on improving the quality of the viewing graph before applying SfM. Our main contribution is a novel optimization that improves the quality of the relative geometries in the viewing graph by enforcing loop consistency constraints with the epipolar point transfer. We show that this optimization greatly improves the accuracy of relative poses in the viewing graph and removes the need for filtering steps or robust algorithms typically used in global SfM methods. In addition, the optimized viewing graph can be used to efficiently calibrate cameras at scale. We combine our viewing graph optimization and focal length calibration into a global SfM pipeline that is more efficient than existing approaches. To our knowledge, ours is the first global SfM pipeline capable of handling uncalibrated image sets. Chris Sweeney, Torsten Sattler, Tobias Höllerer, Matthew Turk 0001, Marc Pollefeys |
ICCV | 4 |
| 2015 | High-order regularization for stereo color editingabstractThis paper pioneers a method for local color editing on stereo image pairs. We generalize the conventional edit propagation framework to stereo views by introducing recent advances in the field of image segmentation, thus allowing a user's edits in one view to be simultaneously performed in the other view. This new formulation maintains consistent editing quality in both views and avoids singularities by solving a well regularized linear system for edit propagation. Kuo-Chin Lien, Jerry D. Gibson, Matthew Turk 0001 |
ICIP | 3 |
| 2015 | 2D-3D Co-segmentation for AR-based Remote CollaborationabstractIn Augmented Reality (AR) based remote collaboration, a remote user can draw a 2D annotation that emphasizes an object of interest to guide a local user accomplishing a task. This annotation is typically performed only once and then sticks to the selected object in the local user's view, independent of his or her camera movement. In this paper, we present an algorithm to segment the selected object, including its occluded surfaces, such that the 2D selection can be appropriately interpreted in 3D and rendered as a useful AR annotation even when the local user moves and significantly changes the viewpoint. Kuo-Chin Lien, Benjamin Nuernberger, Matthew Turk 0001, Tobias Höllerer |
ISMAR | 3 |
| 2015 | Efficient Computation of Absolute Pose for Gravity-Aware Augmented RealityabstractWe propose a novel formulation for determining the absolute pose of a single or multi-camera system given a known vertical direction. The vertical direction may be easily obtained by detecting the vertical vanishing points with computer vision techniques, or with the aid of IMU sensor measurements from a smartphone. Our solver is general and able to compute absolute camera pose from two 2D-3D correspondences for single or multi-camera systems. We run several synthetic experiments that demonstrate our algorithm's improved robustness to image and IMU noise compared to the current state of the art. Additionally, we run an image localization experiment that demonstrates the accuracy of our algorithm in real-world scenarios. Finally, we show that our algorithm provides increased performance for real-time model-based tracking compared to solvers that do not utilize the vertical direction and show our algorithm in use with an augmented reality application running on a Google Tango tablet. Chris Sweeney, John Flynn, Benjamin Nuernberger, Matthew Turk 0001, Tobias Höllerer |
ISMAR | 4 |
| 2015 | Spatio-Temporal Detection of Divided Attention in Reading Applications Using EEG and Eye TrackingabstractReading is central to learning and communicating, however, divided attention in the form of distraction may be present in learning environments, resulting in a limited understanding of the reading material. This paper presents a novel system that can spatio-temporally detect divided attention in users during two different reading applications: typical document reading and speed reading. Eye tracking and electroencephalography (EEG) monitor the user during reading and provide a classifier with data to decide the user's attention state. The multimodal data informs the system where the user was distracted spatially in the user interface and when the user was distracted. Classification was evaluated with two exploratory experiments. The first experiment was designed to divide the user's attention with a multitasking scenario. The second experiment was designed to divide the users attention by simulating a real-world scenario where the reader is interrupted by unpredictable audio distractions. Results from both experiments show that divided attention may be detected spatio-temporally well above chance on a single-trial basis. Mathieu Rodrigue, Jungah Son, Barry Giesbrecht, Matthew Turk 0001, Tobias Höllerer |
IUI | 4 |
| 2015 | Theia: A Fast and Scalable Structure-from-Motion LibraryabstractIn this paper, we have presented a comprehensive multi-view geometry library, Theia, that focuses on large-scale SfM. In addition to state-of-the-art scalable SfM pipelines, the library provides numerous tools that are useful for students, researchers, and industry experts in the field of multi-view geometry. Theia contains clean code that is well documented (with code comments and the website) and easy to extend. The modular design allows for users to easily implement and experiment with new algorithms within our current pipeline without having to implement a full end-to-end SfM pipeline themselves. Theia has already gathered a large number of diverse users from universities, startups, and industry and we hope to continue to gather users and active contributors from the open-source community. Chris Sweeney, Tobias Höllerer, Matthew Turk 0001 |
ACM Multimedia | 3 |
| 2015 | Composition Context PhotographyabstractCameras are becoming increasingly aware of the picture-taking context, collecting extra information around the act of photographing. This contextual information enables the computational generation of a wide range of enhanced photographic outputs, effectively expanding the imaging experience provided by consumer cameras. Computer vision and computational photography techniques can be applied to provide image composites, such as panoramas, high dynamic range images, and stroboscopic images, as well as automatically selecting individual alternative frames. Our technology can be integrated into point-and shoot cameras, and it effectively expands the photographic possibilities for casual and amateur users, who often rely on automatic camera modes. Daniel A. Vaquero, Matthew Turk 0001 |
WACV | 2 |
| 2015 | Brief Introduction to the Special Issue on Behavior Understanding for Arts and EntertainmentabstractThis editorial introduction describes the aims and scope of the special issue of the ACM Transactions on Interactive Intelligent Systems on Behavior Understanding for Arts and Entertainment, which is being published in issues 2 and 3 of volume 5 of the journal. Here we offer a brief introduction to the use of behavior analysis for interactive systems that involve creativity in either the creator or the consumer of a work of art. We then characterize each of the five articles included in this first part of the special issue, which span a wide range of applications. Albert Ali Salah, Hayley Hung, Oya Aran, Hatice Gunes, Matthew Turk 0001 |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2015 | Behavior Understanding for Arts and EntertainmentabstractThis editorial introduction complements the shorter introduction to the first part of the two-part special issue on Behavior Understanding for Arts and Entertainment. It offers a more expansive discussion of the use of behavior analysis for interactive systems that involve creativity, either for the producer or the consumer of such a system. We first summarise the two articles that appear in this second part of the special issue. We then discuss general questions and challenges in this domain that were suggested by the entire set of seven articles of the special issue and by the comments of the reviewers of these articles. Albert Ali Salah, Hayley Hung, Oya Aran, Hatice Gunes, Matthew Turk 0001 |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2014 | Solving for Relative Pose with a Partially Known Rotation is a Quadratic Eigenvalue ProblemabstractWe propose a novel formulation of minimal case solutions for determining the relative pose of perspective and generalized cameras given a partially known rotation, namely, a known axis of rotation. An axis of rotation may be easily obtained by detecting vertical vanishing points with computer vision techniques, or with the aid of sensor measurements from a smart phone. Given a known axis of rotation, our algorithms solve for the angle of rotation around the known axis along with the unknown translation. We formulate these relative pose problems as Quadratic Eigen value Problems which are very simple to construct. We run several experiments on synthetic and real data to compare our methods to the current state-of-the-art algorithms. Our methods provide several advantages over alternatives methods, including efficiency and accuracy, particularly in the presence of image and sensor noise as is often the case for mobile devices. Chris Sweeney, John Flynn, Matthew Turk 0001 |
3DV | 3 |
| 2014 | gDLS: A Scalable Solution to the Generalized Pose and Scale Problem
Chris Sweeney, Victor Fragoso, Tobias Höllerer, Matthew Turk 0001 |
ECCV (4) | 4 |
| 2014 | Non-Visual Navigation Using Combined Audio Music and Haptic CuesabstractWhile a great deal of work has been done exploring non-visual navigation interfaces using audio and haptic cues, little is known about the combination of the two. We investigate combining different state-of-the-art interfaces for communicating direction and distance information using vibrotactile and audio music cues, limiting ourselves to interfaces that are possible with current off-the-shelf smartphones. We use experimental logs, subjective task load questionnaires, and user comments to see how users' perceived performance, objective performance, and acceptance of the system varied for different combinations. Users' perceived performance did not differ much between the unimodal and multimodal interfaces, but a few users commented that the multimodal interfaces added some cognitive load. Objective performance showed that some multimodal combinations resulted in significantly less direction or distance error over some of the unimodal ones, especially the purely haptic interface. Based on these findings we propose a few design considerations for multimodal haptic/audio navigation interfaces. Emily Fujimoto, Matthew Turk 0001 |
ICMI | 2 |
| 2014 | World-stabilized annotations and virtual scene navigation for remote collaborationabstractWe present a system that supports an augmented shared visual space for live mobile remote collaboration on physical tasks. The remote user can explore the scene independently of the local user's current camera position and can communicate via spatial annotations that are immediately visible to the local user in augmented reality. Our system operates on off-the-shelf hardware and uses real-time visual tracking and modeling, thus not requiring any preparation or instrumentation of the environment. It creates a synergy between video conferencing and remote scene exploration under a unique coherent interface. To evaluate the collaboration with our system, we conducted an extensive outdoor user study with 60 participants comparing our system with two baseline interfaces. Our results indicate an overwhelming user preference (80%) for our system, a high level of usability, as well as performance benefits compared with one of the two baselines. Steffen Gauglitz, Benjamin Nuernberger, Matthew Turk 0001, Tobias Höllerer |
UIST | 3 |
| 2014 | User-perspective augmented reality magic lens from gradientsabstractIn this paper we present a new approach to creating a geometrically-correct user-perspective magic lens and a prototype device implementing the approach. Our prototype uses just standard color cameras, with no active depth sensing. We achieve this by pairing a recent gradient domain image-based rendering method with a novel semi-dense stereo matching algorithm inspired by PatchMatch. Our stereo algorithm is simple but fast and accurate within its search area. The resulting system is a real-time magic lens that displays the correct user perspective with a high-quality rendering, despite the lack of a dense disparity map. Domagoj Baricevic, Tobias Höllerer, Pradeep Sen, Matthew Turk 0001 |
VRST | 4 |
| 2014 | In touch with the remote world: remote collaboration with augmented reality drawings and virtual navigationabstractAugmented reality annotations and virtual scene navigation add new dimensions to remote collaboration. In this paper, we present a touchscreen interface for creating freehand drawings as world-stabilized annotations and for virtually navigating a scene reconstructed live in 3D, all in the context of live remote collaboration. Two main focuses of this work are (1) automatically inferring depth for 2D drawings in 3D space, for which we evaluate four possible alternatives, and (2) gesture-based virtual navigation designed specifically to incorporate constraints arising from partially modeled remote scenes. We evaluate these elements via qualitative user studies, which in addition provide insights regarding the design of individual visual feedback elements and the need to visualize the direction of drawings. Steffen Gauglitz, Benjamin Nuernberger, Matthew Turk 0001, Tobias Höllerer |
VRST | 3 |
| 2014 | Multimodal interaction: A review
Matthew Turk 0001 |
Pattern Recognit. Lett. | 1 |
| 2014 | Model Estimation and Selection towardsUnconstrained Real-Time Tracking and MappingabstractWe present an approach and prototype implementation to initialization-free real-time tracking and mapping that supports any type of camera motion in 3D environments, that is, parallax-inducing as well as rotation-only motions. Our approach effectively behaves like a keyframe-based Simultaneous Localization and Mapping system or a panorama tracking and mapping system, depending on the camera movement. It seamlessly switches between the two modes and is thus able to track and map through arbitrary sequences of parallax-inducing and rotation-only camera movements. The system integrates both model-based and model-free tracking, automatically choosing between the two depending on the situation, and subsequently uses the "Geometric Robust Information Criterion" to decide whether the current camera motion can best be represented as a parallax-inducing motion or a rotation-only motion. It continues to collect and map data after tracking failure by creating separate tracks which are later merged if they are found to overlap. This is in contrast to most existing tracking and mapping systems, which suspend tracking and mapping and thus discard valuable data until relocalization with respect to the initial map is successful. We tested our prototype implementation on a variety of video sequences, successfully tracking through different camera motions and fully automatically building combinations of panoramas and 3D structure. Steffen Gauglitz, Chris Sweeney, Jonathan Ventura, Matthew Turk 0001, Tobias Höllerer |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2014 | Mono-spectrum marker: an AR marker robust to image blur and defocus
Masahiro Toyoura, Haruhito Aruga, Matthew Turk 0001, Xiaoyang Mao |
Vis. Comput. | 3 |
| 2013 | SWIGS: A Swift Guided Sampling MethodabstractWe present SWIGS, a Swift and efficient Guided Sampling method for robust model estimation from image feature correspondences. Our method leverages the accuracy of our new confidence measure (MR-Rayleigh), which assigns a correctness-confidence to a putative correspondence in an online fashion. MR-Rayleigh is inspired by Meta-Recognition (MR), an algorithm that aims to predict when a classifier's outcome is correct. We demonstrate that by using a Rayleigh distribution, the prediction accuracy of MR can be improved considerably. Our experiments show that MR-Rayleigh tends to predict better than the often-used Lowe's ratio, Brown's ratio, and the standard MR under a range of imaging conditions. Furthermore, our homography estimation experiment demonstrates that SWIGS performs similarly or better than other guided sampling methods while requiring fewer iterations, leading to fast and accurate model estimates. Victor Fragoso, Matthew Turk 0001 |
CVPR | 2 |
| 2013 | Detecting Markers in Blurred and Defocused ImagesabstractPlanar markers enable an augmented reality (AR) system to estimate the pose of objects from images containing them. However, conventional markers are difficult to detect in blurred or defocused images. We propose a new marker and a new detection and identification method that is designed to work under such conditions. The problem of conventional markers is that their patterns consist of high-frequency components such as sharp edges which are attenuated in blurred or defocused images. Our marker consists of a single low-frequency component. We call it a mono-spectrum marker. The mono-spectrum marker can be detected in real time with a GPU. In experiments, we confirm that the mono-spectrum marker can be accurately detected in blurred and defocused images in real time. Using these markers can increase the performance and robustness of AR systems and other vision applications that require detection or tracking of defined markers. Masahiro Toyoura, Haruhito Aruga, Matthew Turk 0001, Xiaoyang Mao |
CW | 3 |
| 2013 | EVSAC: Accelerating Hypotheses Generation by Modeling Matching Scores with Extreme Value TheoryabstractAlgorithms based on RANSAC that estimate models using feature correspondences between images can slow down tremendously when the percentage of correct correspondences (inliers) is small. In this paper, we present a probabilistic parametric model that allows us to assign confidence values for each matching correspondence and therefore accelerates the generation of hypothesis models for RANSAC under these conditions. Our framework leverages Extreme Value Theory to accurately model the statistics of matching scores produced by a nearest-neighbor feature matcher. Using a new algorithm based on this model, we are able to estimate accurate hypotheses with RANSAC at low inlier ratios significantly faster than previous state-of-the-art approaches, while still performing comparably when the number of inliers is large. We present results of homography and fundamental matrix estimation experiments for both SIFT and SURF matches that demonstrate that our method leads to accurate and fast model estimations. Victor Fragoso, Pradeep Sen, Sergio Rodríguez, Matthew Turk 0001 |
ICCV | 4 |
| 2013 | Improved outdoor augmented reality through "Globalization"abstractDespite the major interest in live tracking and mapping (e.g., SLAM), the field of augmented reality has yet to truly make use of the rich data provided from large-scale reconstructions generated by structure from motion. This dissertation focuses on extensible tracking and mapping for large-scale reconstructions that enables SfM and SLAM to operate cooperatively to mutually enhance the performance. We describe a multi-user, collaborative augmented reality system that will collectively extend and enhance reconstructions of urban environments at city-scales. Contrary to current outdoor augmented reality systems, this system is capable of continuous tracking through areas previously modeled as well as new, undiscovered areas. Further, we describe a new process called globalization that propagates new visual information back to the global model. Globalization allows for continuous updating of the 3D models with visual data from live users, providing data to fill coverage gaps that are common in 3D reconstructions and to provide the most current view of an environment as it changes over time. The proposed research is a crucial step toward enabling users to augment urban environments with location-specific information at any location in the world for a truly global augmented reality. Chris Sweeney, Tobias Höllerer, Matthew Turk 0001 |
ISMAR | 3 |
| 2013 | Machine learning in motion analysis: New advances
Matti Pietikäinen, Matthew Turk 0001, Liang Wang 0001, Guoying Zhao 0001, Li Cheng 0001 |
Image Vis. Comput. | 2 |
| 2013 | Over twenty years of eigenfacesabstractThe inaugural ACM Multimedia Conference coincided with a surge of interest in computer vision technologies for detecting and recognizing people and their activities in images and video. Face recognition was the first of these topics to broadly engage the vision and multimedia research communities. The Eigenfaces approach was, deservedly or not, the method that captured much of the initial attention, and it continues to be taught and used as a benchmark over 20 years later. This article is a brief personal view of the genesis of Eigenfaces for face recognition and its relevance to the multimedia community. Matthew Turk 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2012 | Locating binary features for keypoint recognition using noncooperative gamesabstractMany applications in computer vision rely on determining the correspondence between two images that share an overlapping region. One way to establish this correspondence is by matching binary keypoint descriptors from both images. Although, these descriptors are efficiently computed with bits produced by an arrangement of binary features (pattern), their matching performance falls short in comparison with other more elaborated descriptors such as SIFT. We present an approach based on noncooperative game theory for computing the locations of every binary feature in a pattern, improving the performance of binary-feature-based matchers. We propose a simultaneous two-player zero-sum game in which a maximizer wants to increase a payoff by selecting the possible locations for the features; a minimizer wants to decrease the payoff by selecting a pair of keypoints to confuse the maximizer; and the payoff matrix is computed from the pixel intensities across the pixel neighborhood of the keypoints. We use the best locations from the obtained maximizer's optimal policy for locating every binary feature in the pattern. Our evaluation of this approach coupled with Ferns shows an improvement in matching keypoints, in particular those with similar texture. Moreover, our approach improves the matching performance when fewer bits are required. Victor Fragoso, Matthew Turk 0001, João Pedro Hespanha |
ICIP | 2 |
| 2012 | A hand-held AR magic lens with user-perspective renderingabstractIn this paper we present a user study evaluating the benefits of geometrically correct user-perspective rendering using an Augmented Reality (AR) magic lens. In simulation we compared a user-perspective magic lens against the common device-perspective magic lens on both phone-sized and tablet-sized displays. Our results indicate that a tablet-sized display allows for significantly faster performance of a selection task and that a user-perspective lens has benefits over a device-perspective lens for a selection task. Based on these promising results, we created a proof-of-concept prototype, engineered with current off-the-shelf devices and software. To our knowledge, this is the first geometrically correct user-perspective magic lens. Domagoj Baricevic, Cha Lee, Matthew Turk 0001, Tobias Höllerer, Doug A. Bowman |
ISMAR | 3 |
| 2012 | Live tracking and mapping from both general and rotation-only camera motionabstractWe present an approach to real-time tracking and mapping that supports any type of camera motion in 3D environments, that is, general (parallax-inducing) as well as rotation-only (degenerate) motions. Our approach effectively generalizes both a panorama mapping and tracking system and a keyframe-based Simultaneous Localization and Mapping (SLAM) system, behaving like one or the other depending on the camera movement. It seamlessly switches between the two and is thus able to track and map through arbitrary sequences of general and rotation-only camera movements. Key elements of our approach are to design each system component such that it is compatible with both panoramic data and Structure-from-Motion data, and the use of the `Geometric Robust Information Criterion' to decide whether the transformation between a given pair of frames can best be modeled with an essential matrix E, or with a homography H. Further key features are that no separate initialization step is needed, that the reconstruction is unbiased, and that the system continues to collect and map data after tracking failure, thus creating separate tracks which are later merged if they overlap. The latter is in contrast to most existing tracking and mapping systems, which suspend tracking and mapping, thus discarding valuable data, while trying to relocalize the camera with respect to the initial map. We tested our system on a variety of video sequences, successfully tracking through different camera motions and fully automatically building panoramas as well as 3D structures. Steffen Gauglitz, Chris Sweeney, Jonathan Ventura, Matthew Turk 0001, Tobias Höllerer |
ISMAR | 4 |
| 2012 | Integrating the physical environment into mobile remote collaborationabstractWe describe a framework and prototype implementation for unobtrusive mobile remote collaboration on tasks that involve the physical environment. Our system uses the Augmented Reality paradigm and model-free, markerless visual tracking to facilitate decoupled, live updated views of the environment and world-stabilized annotations while supporting a moving camera and unknown, unprepared environments. In order to evaluate our concept and prototype, we conducted a user study with 48 participants in which a remote expert instructed a local user to operate a mock-up airplane cockpit. Users performed significantly better with our prototype (40.8 tasks completed on average) as well as with static annotations (37.3) than without annotations (28.9). 79% of the users preferred our prototype despite noticeably imperfect tracking. Steffen Gauglitz, Cha Lee, Matthew Turk 0001, Tobias Höllerer |
Mobile HCI | 3 |
| 2012 | Introduction to the Special Issue on Mobile Vision
Gang Hua 0001, Yun Fu 0001, Matthew Turk 0001, Marc Pollefeys, Zhengyou Zhang |
Int. J. Comput. Vis. | 3 |
| 2012 | A New Biased Discriminant Analysis Using Composite Vectors for Eye DetectionabstractWe propose a new biased discriminant analysis (BDA) using composite vectors for eye detection. A composite vector consists of several pixels inside a window on an image. The covariance of composite vectors is obtained from their inner product and can be considered as a generalization of the covariance of pixels. The proposed composite BDA (C-BDA) method is a BDA using the covariance of composite vectors. We construct a hybrid cascade detector for eye detection, using Haar-like features in the earlier stages and composite features obtained from C-BDA in the later stages. The proposed detector runs in real time; its execution time is 5.5 ms on a typical PC. The experimental results for the CMU PIE database and our own real-world data set show that the proposed detector provides robust performance to several kinds of variations such as facial pose, illumination, eyeglasses, and partial occlusion. On the whole, the detection rate per pair of eyes is 98.0% for the 3604 face images of the CMU PIE database and 95.1% for the 2331 face images of the real-world data set. In particular, it provides a 99.7% detection rate for the 2120 CMU PIE images without glasses. Face recognition performance is also investigated using the eye coordinates from the proposed detector. The recognition results for the real-world data set show that the proposed detector gives similar performance to the method using manually located eye coordinates, showing that the accuracy of the proposed eye detector is comparable with that of the ground-truth data. Chunghoon Kim, Sang-Il Choi, Matthew Turk 0001, Chong-Ho Choi |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2011 | Improving Keypoint Orientation AssignmentabstractDetection and description of local image features has proven to be a powerful paradigm for a variety of applications in computer vision. Often, this process includes an orientation assignment step to render the overall process invariant to in-plane rotation. In this paper, we review several different existing algorithms and propose two novel, efficient methods for orientation assignment. The first method exhibits a very good speedperformance trade-off; the second is capable of multiple orientations and performs comparable to SIFT’s orientation assignment while being significantly cheaper. Additionally, we improve one of the existing orientation assignment methods by generalizing it. All algorithms are evaluated empirically under a variety of conditions and in combination with six keypoint detectors. Steffen Gauglitz, Matthew Turk 0001, Tobias Höllerer |
BMVC | 2 |
| 2011 | Illumination demultiplexing from a single imageabstractA class of techniques in computer vision and graphics is based on capturing multiple images of a scene under different illumination conditions. These techniques explore variations in illumination from image to image to extract interesting information about the scene. However, their applicability to dynamic environments is limited due to the need for robust motion compensation algorithms. To overcome this issue, we propose a method to separate multiple illuminants from a single image. Given an image of a scene simultaneously illuminated by multiple light sources, our method generates individual images as if they had been illuminated by each of the light sources separately. To facilitate the illumination separation process, we encode each light source with a distinct sinusoidal pattern, strategically selected given the relative position of each light with respect to the camera, such that the observed sinusoids become independent of the scene geometry. The individual illuminants are then demultiplexed by analyzing local frequencies. We show applications of our approach in image-based relighting, photometric stereo, and multiflash imaging. Christine Chen, Daniel A. Vaquero, Matthew Turk 0001 |
ICCV | 3 |
| 2011 | Efficiently selecting spatially distributed keypoints for visual trackingabstractWe describe an algorithm dubbed Suppression via Disk Covering (SDC) to efficiently select a set of strong, spatially distributed key-points, and we show that selecting keypoint in this way significantly improves visual tracking. We also describe two efficient implementation schemes for the popular Adaptive Non-Maximal Suppression algorithm, and show empirically that SDC is significantly faster while providing the same improvements with respect to tracking robustness. In our particular application, using SDC to filter the output of an inexpensive (but, by itself, less reliable) keypoint detector (FAST) results in higher tracking robustness at significantly lower total cost than using a computationally more expensive detector. Steffen Gauglitz, Luca Foschini 0002, Matthew Turk 0001, Tobias Höllerer |
ICIP | 3 |
| 2011 | Multisensory embedded pose estimationabstractWe present a multisensory method for estimating the transformation of a mobile phone between two images taken from its camera. Pose estimation is a necessary step for applications such as 3D reconstruction and panorama construction, but detecting and matching robust features can be computationally expensive. In this paper we propose a method for combining the inertial sensors (accelerometers and gyroscopes) of a mobile phone with its camera to provide a fast and accurate pose estimation. We use the inertial based pose to warp two images into the same perspective frame. We then employ an adaptive FAST feature detector and image patches, normalized with respect to illumination, as feature descriptors. After the warping the images are approximately aligned with each other so the search for matching key-points also becomes faster and in certain cases more reliable. Our results show that by incorporating the inertial sensors we can considerably speed up the process of detecting and matching key-points between two images, which is the most time consuming step of the pose estimation. Eyrun Eyjolfsdottir, Matthew Turk 0001 |
WACV | 2 |
| 2011 | TranslatAR: A mobile augmented reality translatorabstractWe present a mobile augmented reality (AR) translation system, using a smartphone's camera and touchscreen, that requires the user to simply tap on the word of interest once in order to produce a translation, presented as an AR overlay. The translation seamlessly replaces the original text in the live camera stream, matching background and foreground colors estimated from the source images. For this purpose, we developed an efficient algorithm for accurately detecting the location and orientation of the text in a live camera stream that is robust to perspective distortion, and we combine it with OCR and a text-to-text translation engine. Our experimental results, using the ICDAR 2003 dataset and our own set of video sequences, quantify the accuracy of our detection and analyze the sources of failure among the system's components. With the OCR and translation running in a background thread, the system runs at 26 fps on a current generation smartphone (Nokia N900) and offers a particularly easy-to-use and simple method for translation, especially in situations in which typing or correct pronunciation (for systems with speech input) is cumbersome or impossible. Victor Fragoso, Steffen Gauglitz, Shane Zamora, Jim Kleban, Matthew Turk 0001 |
WACV | 5 |
| 2011 | Car-Rec: A real time car recognition systemabstractRecent advances in computer vision have significantly reduced the difficulty of object classification and recognition. Robust feature detector and descriptor algorithms are particularly useful, forming the basis for many recognition and classification applications. These algorithms have been used in divergent bag-of-words and structural matching approaches. This work demonstrates a recognition application, based upon the SURF feature descriptor algorithm, which fuses bag-of-words and structural verification techniques. The resulting system is applied to the domain of car recognition and achieves accurate (>; 90%) and real-time performance when searching databases containing thousands of images. Daniel Marcus Jang, Matthew Turk 0001 |
WACV | 2 |
| 2011 | Generalized autofocusabstractAll-in-focus imaging is a computational photography technique that produces images free of defocus blur by capturing a stack of images focused at different distances and merging them into a single sharp result. Current approaches assume that images have been captured offline, and that a reasonably powerful computer is available to process them. In contrast, we focus on the problem of how to capture such input stacks in an efficient and scene-adaptive fashion. Inspired by passive autofocus techniques, which select a single best plane of focus in the scene, we propose a method to automatically select a minimal set of images, focused at different depths, such that all objects in a given scene are in focus in at least one image. We aim to minimize both the amount of time spent metering the scene and capturing the images, and the total amount of high-resolution data that is captured. The algorithm first analyzes a set of low-resolution sharpness measurements of the scene while continuously varying the focus distance of the lens. From these measurements, we estimate the final lens positions required to capture all objects in the scene in acceptable focus. We demonstrate the use of our technique in a mobile computational photography scenario, where it is essential to minimize image capture time (as the camera is typically handheld) and processing time (as the computation and energy resources are limited). Daniel A. Vaquero, Natasha Gelfand, Marius Tico, Kari Pulli, Matthew Turk 0001 |
WACV | 5 |
| 2011 | Evaluation of Interest Point Detectors and Feature Descriptors for Visual Tracking
Steffen Gauglitz, Tobias Höllerer, Matthew Turk 0001 |
Int. J. Comput. Vis. | 3 |
| 2011 | Introduction to the Special Section on Real-World Face RecognitionabstractThe motivations for organizing this special section were to better address the challenges of face recognition in real-world scenarios, to promote systematic research and evaluation of promising methods and systems, to provide a snapshot of where we are in this domain, and to stimulate discussion about future directions. We solicited original contributions of research on all aspects of real-world face recognition, including: the design of robust face similarity features and metrics; robust face clustering and sorting algorithms; novel user interaction models and face recognition algorithms for face tagging; novel applications of web face recognition; novel computational paradigms for face recognition; challenges in large scale face recognition tasks, e.g., on the Internet; face recognition with contextual information; face recognition benchmarks and evaluation methodology for moderately controlled or uncontrolled environments; and video face recognition. We received 42 original submissions, four of which were rejected without review; the other 38 papers entered the normal review process. Each paper was reviewed by three reviewers who are experts in their respective topics. More than 100 expert reviewers have been involved in the review process. The papers were equally distributed among the guest editors. A final decision for each paper was made by at least two guest editors assigned to it. To avoid conflict of interest, no guest editor submitted any papers to this special section. Gang Hua 0001, Ming-Hsuan Yang 0001, Erik G. Learned-Miller, Yi Ma 0001, Matthew Turk 0001, David J. Kriegman, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2010 | Computational Illumination
Matthew Turk 0001 |
CIARP | 1 |
| 2010 | Human Activity Recognition Using Local Shape DescriptorsabstractWe propose a method for human activity recognition in videos, based on shape analysis. We define local shape descriptors for interest points on the detected contour of the human action and build an action descriptor using a Bag of Features method. We also use the temporal relation among matching interest points across successive video frames. Further, an SVM is trained on these action descriptors to classify the activity in the scene. The method is invariant to the length of the video sequence, and hence it is suitable in online activity recognition. We have demonstrated the results on an action database consisting of nine actions like walk, jump, bend, etc., by twenty people, in indoor and outdoor scenarios. The proposed method achieves an accuracy of 87%, and is comparable to other state-of-the-art methods. Sharath Venkatesha, Matthew Turk 0001 |
ICPR | 2 |
| 2009 | Shape classification through structured learning of matching measuresabstractMany traditional methods for shape classification involve establishing point correspondences between shapes to produce matching scores, which are in turn used as similarity measures for classification. Learning techniques have been applied only in the second stage of this process, after the matching scores have been obtained. In this paper, instead of simply taking for granted the scores obtained by matching and then learning a classifier, we learn the matching scores themselves so as to produce shape similarity scores that minimize the classification loss. The solution is based on a max-margin formulation in the structured prediction setting. Experiments in shape databases reveal that such an integrated learning algorithm substantially improves on existing methods. Longbin Chen, Julian J. McAuley, Rogério Feris, Tibério S. Caetano, Matthew Turk 0001 |
CVPR | 5 |
| 2009 | A projector-camera setup for geometry-invariant frequency demultiplexingabstractConsider a projector-camera setup where a sinusoidal pattern is projected onto the scene, and an image of the objects imprinted with the pattern is captured by the camera. In this configuration, the local frequency of the sinusoidal pattern as seen by the camera is a function of both the frequency of the projected sinusoid and the local geometry of objects in the scene. We observe that, by strategically placing the projector and the camera in canonical configuration and projecting sinusoidal patterns aligned with the epipolar lines, the frequency of the sinusoids seen in the image becomes invariant to the local object geometry. This property allows us to design systems composed of a camera and multiple projectors, which can be used to capture a single image of a scene illuminated by all projectors at the same time, and then demultiplex the frequencies generated by each individual projector separately. We show how imaging systems like those can be used to segment, from a single image, the shadows cast by each individual projector - an application that we call coded shadow photography. The method is useful to extend the applicability of techniques that rely on the analysis of shadows cast by multiple light sources placed at different positions, as the individual shadows captured at distinct instants of time now can be obtained from a single shot, enabling the processing of dynamic scenes. Daniel A. Vaquero, Ramesh Raskar, Rogério Feris, Matthew Turk 0001 |
CVPR | 4 |
| 2009 | Attribute-based people search in surveillance environmentsabstractWe propose a novel framework for searching for people in surveillance environments. Rather than relying on face recognition technology, which is known to be sensitive to typical surveillance conditions such as lighting changes, face pose variation, and low-resolution imagery, we approach the problem in a different way: we search for people based on a parsing of human parts and their attributes, including facial hair, eyewear, clothing color, etc. These attributes can be extracted using detectors learned from large amounts of training data. A complete system that implements our framework is presented. At the interface, the user can specify a set of personal characteristics, and the system then retrieves events that match the provided description. For example, a possible query is ¿show me the bald people who entered a given building last Saturday wearing a red shirt and sunglasses.¿ This capability is useful in several applications, such as finding suspects or missing people. To evaluate the performance of our approach, we present extensive experiments on a set of images collected from the Internet, on infrared imagery, and on two-and-a-half months of video from a real surveillance environment. We are not aware of any similar surveillance system capable of automatically finding people in video based on their fine-grained body parts and attributes. Daniel A. Vaquero, Rogério Feris, Duan Tran, Lisa M. Brown, Arun Hampapur, Matthew Turk 0001 |
WACV | 6 |
| 2008 | Characterizing the shadow space of camera-light pairsabstractWe present a theoretical analysis for characterizing the shadows cast by a point light source given its relative position to the camera. In particular, we analyze the epipolar geometry of camera-light pairs, including unusual camera-light configurations such as light sources aligned with the camerapsilas optical axis as well as convenient arrangements such as lights placed in the camera plane. A mathematical characterization of the shadows is derived to determine the orientations and locations of depth discontinuities when projected onto the image plane that could potentially be associated with cast shadows. The resulting theory is applied to compute a lower bound on the number of lights needed to extract all depth discontinuities from a general scene using a multiflash camera. We also provide a characterization of which discontinuities are missed and which are correctly detected by the algorithm, and a foundation for choosing an optimal light placement. Experiments with depth edges computed using two-flash setups and a four-flash setup illustrate the theory, and an additional configuration with a flash at the camerapsilas center of projection is exploited as a solution for some degenerate cases. Daniel A. Vaquero, Rogério Feris, Matthew Turk 0001, Ramesh Raskar |
CVPR | 3 |
| 2008 | Biased discriminant analysis using composite vectors for eye detectionabstractWe propose a new discriminant analysis using composite vectors for eye detection. A composite vector consists of a number of pixels inside a window on an image. The covariance of composite vectors is obtained from their inner product and can be considered as a generalized form of the covariance of pixels. The proposed C-BDA is a biased discriminant analysis using the covariance of composite vectors. In the hybrid cascade detector constructed for eye detection, Haar-like features are used in the earlier stages and composite features obtained from C-BDA are used in the later stages. The experimental results for the CMU and Yale databases show that the proposed detector provides robust performance to several kinds of variations such as facial pose, illumination, and closed eyes. In particular, it provides a 99.4% detection rate for the CMU images without glasses. Chunghoon Kim, Matthew Turk 0001, Chong-Ho Choi |
FG | 2 |
| 2008 | Using structured light for efficient depth edge detection
Cheolhwon Kim, Jaekeun Na, Juneho Yi, Matthew Turk 0001 |
Image Vis. Comput. | 5 |
| 2008 | Multiflash Stereopsis: Depth-Edge-Preserving Stereo with Small Baseline IlluminationabstractTraditional stereo matching algorithms are limited in their ability to produce accurate results near depth discontinuities, due to partial occlusions and violation of smoothness constraints. In this paper, we use small baseline multi-flash illumination to produce a rich set of feature maps that enable acquisition of discontinuity preserving point correspondences. First, from a single multi-flash camera, we formulate a qualitative depth map using a gradient domain method that encodes object relative distances. Then, in a multiview setup, we exploit shadows created by light sources to compute an occlusion map. Finally, we demonstrate the usefulness of these feature maps by incorporating them into two different dense stereo correspondence algorithms, the first based on local search and the second based on belief propagation. Experimental results show that our enhanced stereo algorithms are able to extract high quality, discontinuity preserving correspondence maps from scenes that are extremely challenging for conventional stereo methods. We also demonstrate that small baseline illumination can be useful to handle specular reflections in stereo imagery. Different from most existing active illumination techniques, our method is simple, inexpensive, compact, and requires no calibration of light sources. Rogério Feris, Ramesh Raskar, Longbin Chen, Kar-Han Tan, Matthew Turk 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2007 | The Hierarchical Isometric Self-Organizing Map for Manifold RepresentationabstractWe present an algorithm, Hierarchical ISOmetric Self-Organizing Map (H-ISOSOM), for a concise, organized manifold representation of complex, non-linear, large scale, high-dimensional input data in a low dimensional space. The main contribution of our algorithm is threefold. First, we modify the previous ISOSOM algorithm by a local linear interpolation (LLl) technique, which maps the data samples from low dimensional space back to high dimensional space and makes the complete mapping pseudo-invertible. The modified-ISOSOM (M-ISOSOM) follows the global geometric structure of the data, and also preserves local geometric relations to reduce the nonlinear mapping distortion and make the learning more accurate. Second, we propose the H-ISOSOM algorithm for the computational complexity problem of Isomap, SOM and LLI and the nonlinear complexity problem of the highly twisted manifold. H-ISOSOM learns an organized structure of a non-convex, large scale manifold and represents it by a set of hierarchical organized maps. The hierarchical structure follows a coarse-to-fine strategy. According to the coarse global structure, it "unfolds " the manifold at the coarse level and decomposes the sample data into small patches, then iteratively learns the nonlinearity of each patch in finer levels. The algorithm simultaneously reorganizes and clusters the data samples in a low dimensional space to obtain the concise representation. Third, we give quantitative comparisons of the proposed method with similar methods on standard data sets. Finally, we apply H-ISOSOM to the problem of appearance-based hand pose estimation. Encouraging experimental results validate the effectiveness and efficiency of H-ISOSOM. Haiying Guan, Matthew Turk 0001 |
CVPR | 2 |
| 2006 | Automatic Hot Spot Detection and Segmentation in Whole Body FDG-PET ImagesabstractWe present a system for automatic hot spots detection and segmentation in whole body FDG-PET images. The main contribution of our system is threefold. First, it has a novel body-section labeling module based on spatial hidden-Markov models (HMM); this allows different processing policies to be applied in different body sections. Second, the competition diffusion (CD) segmentation algorithm, which takes into account body-section information, converts the binary thresholding results to probabilistic interpretation and detects hot-spot region candidates. Third, a recursive intensity mode-seeking algorithm finds hot spot centers efficiently, and given these centers, a clinically meaningful protocol is proposed to accurately quantify hot spot volumes. Experimental results show that our system works robustly despite the large variations in clinical PET images. Haiying Guan, Toshiro Kubota, Sharon X. Huang, Xiang Sean Zhou, Matthew Turk 0001 |
ICIP | 5 |
| 2006 | Manifold based analysis of facial expression
Ya Chang, Changbo Hu, Rogério Feris, Matthew Turk 0001 |
Image Vis. Comput. | 4 |
| 2006 | Local approach for face verification in polar frequency domain
Yossi Zana, Roberto Marcondes Cesar Junior, Rogério Feris, Matthew Turk 0001 |
Image Vis. Comput. | 4 |
| 2005 | Hand Tracking with Flocks of FeaturesabstractTracking hands in live video is a challenging task: the hand appearance can change too rapidly for appearance-based trackers to work, and color-based trackers (that do not rely on geometry) have to make limiting assumptions about the background color. This article shows the results of hand tracking with "Flocks of Features", a tracking method that combines motion cues and a learned foreground color distribution to achieve fast and robust 2D tracking of highly articulated objects. Many independent image artifacts are tracked from one frame to the next, adhering only to local constraints. This concept is borrowed from nature since these tracks mimic the flight of flocking birds - exhibiting local individualism and variability while maintaining a clustered entirety. Hand tracking has important applications for interaction with wearable computers, for intuitive manipulation of virtual objects, for detection of activity signatures, and much more. Tracking with Flocks of Features is not limited to hands - any articulated or appearance-changing object can benefit from this multi-cue tracking method. Mathias Kölsch, Matthew Turk 0001 |
CVPR (2) | 2 |
| 2005 | Discontinuity Preserving Stereo with Small Baseline Multi-Flash IlluminationabstractCurrently, sharp discontinuities in depth and partial occlusions in multiview imaging systems pose serious challenges for many dense correspondence algorithms. However, it is important for 3D reconstruction methods to preserve depth edges as they correspond to important shape features like silhouettes which are critical for understanding the structure of a scene. In this paper, we show how active illumination algorithms can produce a rich set of feature maps that are useful in dense 3D reconstruction. We start by showing a method to compute a qualitative depth map from a single camera, which encodes object relative distances and can be used as a prior for stereo. In a multiview setup, we show that along with depth edges, binocular half-occluded pixels can also be explicitly and reliably labeled. To demonstrate the usefulness of these feature maps, we show how they can be used in two different algorithms for dense stereo correspondence. Our experimental results show that our enhanced stereo algorithms are able to extract high quality, discontinuity preserving correspondence maps from scenes that are extremely challenging for conventional stereo methods. Rogério Feris, Ramesh Raskar, Longbin Chen, Kar-Han Tan, Matthew Turk 0001 |
ICCV | 5 |
| 2005 | Automatic Head-size Equalization in Panorama Images for Video ConferencingabstractIn panorama images captured by omni-directional cameras during video conferencing, the image sizes of the people around the conference table are not uniform due to the varying distances to the camera. Spatially varying-uniform (SVU) scaling functions have been proposed to warp a panorama image smoothly such that the participants have similar sizes on the image. To generate the SVU function, one needs to segment the table boundaries, which was generated manually in the previous work. In this paper, we propose a robust algorithm to automatically segment the table boundaries. To ensure the robustness, we apply a symmetry voting scheme to filter out noisy points on the edge map. Trigonometry and quadratic fitting methods are developed to fit a continuous curve to the remaining edge points. We report experimental results on both synthetic and real images. Ya Chang, Ross Cutler, Zicheng Liu 0001, Zhengyou Zhang, Alex Acero, Matthew Turk 0001 |
ICME | 6 |
| 2005 | Non-negative matrix factorization framework for face recognitionabstractNon-negative Matrix Factorization (NMF) is a part-based image representation method which adds a non-negativity constraint to matrix factorization. NMF is compatible with the intuitive notion of combining parts to form a whole face. In this paper, we propose a framework of face recognition by adding NMF constraint and classifier constraints to matrix factorization to get both intuitive features and good recognition results. Based on the framework, we present two novel subspace methods: Fisher Non-negative Matrix Factorization (FNMF) and PCA Non-negative Matrix Factorization (PNMF). FNMF adds both the non-negative constraint and the Fisher constraint to matrix factorization. The Fisher constraint maximizes the between-class scatter and minimizes the within-class scatter of face samples. Subsequently, FNMF improves the capability of face recognition. PNMF adds the non-negative constraint and characteristics of PCA, such as maximizing the variance of output coordinates, orthogonal bases, etc. to matrix factorization. Therefore, we can get intuitive features and desirable PCA characteristics. Our experiments show that FNMF and PNMF achieve better face recognition performance than NMF and Local NMF. Yunde Jia, Changbo Hu, Matthew Turk 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2005 | Effective Representation Using ICA for Face Recognition Robust to Local Distortion and Partial OcclusionabstractThe performance of face recognition methods using subspace projection is directly related to the characteristics of their basis images, especially in the cases of local distortion or partial occlusion. In order for a subspace projection method to be robust to local distortion and partial occlusion, the basis images generated by the method should exhibit a part-based local representation. We propose an effective part-based local representation method named locally salient ICA (LS-ICA) method for face recognition that is robust to local distortion and partial occlusion. The LS-ICA method only employs locally salient information from important facial parts in order to maximize the benefit of applying the idea of "recognition by parts." It creates part-based local basis images by imposing additional localization constraint in the process of computing ICA architecture I basis images. We have contrasted the LS-ICA method with other part-based representations such as LNMF (Localized Nonnegative Matrix Factorization) and LFA (Local Feature Analysis). Experimental results show that the LS-ICA method performs better than PCA, ICA architecture I, ICA architecture II, LFA, and LNMF methods, especially in the cases of partial occlusions and local distortions. Jongsun Kim, Jongmoo Choi, Juneho Yi, Matthew Turk 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2004 | Probabilistic Expression Analysis on Manifolds
Ya Chang, Changbo Hu, Matthew Turk 0001 |
CVPR (2) | 3 |
| 2004 | Multimodal transformed social interactionabstractUnderstanding human-human interaction is fundamental to the long-term pursuit of powerful and natural multimodal interfaces. Nonverbal communication, including body posture, gesture, facial expression, and eye gaze, is an important aspect of human-human interaction. We introduce a paradigm for studying multimodal and nonverbal communication in collaborative virtual environments (CVEs) called Transformed Social Interaction (TSI), in which a user's visual representation is rendered in a way that strategically filters selected communication behaviors in order to change the nature of a social interaction. To achieve this, TSI must employ technology to detect, recognize, and manipulate behaviors of interest, such as facial expressions, gestures, and eye gaze. In [13] we presented a TSI experiment called non-zero-sum gaze (NZSG) to determine the effect of manipulated eye gaze on persuasion in a small group setting. Eye gaze was manipulated so that each participant in a three-person CVE received eye gaze from a presenter that was normal, less than normal, or greater than normal. We review this experiment and discuss the implications of TSI for multimodal interfaces. Matthew Turk 0001, Jeremy N. Bailenson, Andrew C. Beall, Jim Blascovich, Rosanna E. Guadagno |
ICMI | 1 |
| 2004 | Vision-Based Interfaces for MobilityabstractVision-based user interfaces are a feasible and advantageous modality for wearable computers. To substantiate this claim, we present a robust real-time hand gesture recognition system that is capable of being the sole input provider for a demonstration application. It achieves usability and interactivity even when both the head-worn camera and the object of interest are in motion. We describe a set of general gesture-based interaction styles and explore their characteristics in terms of task suitability and the computer vision algorithms required for their recognition. Preliminary evaluation of our prototype system leads to the conclusion that vision-based interfaces have achieved the maturity necessary to help overcome some limitations of more traditional mobile user interfaces. Mathias Kölsch, Matthew Turk 0001, Tobias Höllerer |
MobiQuitous | 2 |
| 2004 | Non-photorealistic camera: depth edge detection and stylized rendering using multi-flash imagingabstractWe present a non-photorealistic rendering approach to capture and convey shape features of real-world scenes. We use a camera with multiple flashes that are strategically positioned to cast shadows along depth discontinuities in the scene. The projective-geometric relationship of the camera-flash setup is then exploited to detect depth discontinuities and distinguish them from intensity edges due to material discontinuities.We introduce depiction methods that utilize the detected edge features to generate stylized static and animated images. We can highlight the detected features, suppress unnecessary details or combine features from multiple images. The resulting images more clearly convey the 3D structure of the imaged scenes.We take a very different approach to capturing geometric features of a scene than traditional approaches that require reconstructing a 3D model. This results in a method that is both surprisingly simple and computationally efficient. The entire hardware/software setup can conceivably be packaged into a self-contained device no larger than existing digital cameras. Ramesh Raskar, Kar-Han Tan, Rogério Feris, Jingyi Yu 0001, Matthew Turk 0001 |
ACM Trans. Graph. | 5 |
| 2003 | Active Wavelet Networks for Face AlignmentabstractThe active appearance model (AAM) algorithm has proved to be a successful method for face alignment and synthesis. By elegantly combining both shape and texture models, AAM allows fast and robust deformable image matching. However, the method is sensitive to partial occlusions and illumination changes. In such cases, the PCA-based texture model causes the reconstruction error to be globally spread over the image. In this paper, we propose a new method for face alignment called active wavelet networks (AWN), which replaces the AAM texture model by a wavelet network representation. Since we consider spatially localized wavelets for modeling texture, our method shows more robustness against partial occlusions and some illumination changes. 1 Changbo Hu, Rogério Feris, Matthew Turk 0001 |
BMVC | 3 |
| 2000 | Gesture Modeling and Recognition Using Finite State MachinesabstractWe propose a state-based approach to gesture learning and recognition. Using spatial clustering and temporal alignment, each gesture is defined to be an ordered sequence of states in spatial-temporal space. The 2D image positions of the centers of the head and both hands of the user are used as features; these are located by a color-based tracking method. From training data of a given gesture, we first learn the spatial information and then group the data into segments that are automatically aligned temporally. The temporal information is further integrated to build a finite state machine (FSM) recognizer. Each gesture has a FSM corresponding to it. The computational efficiency of the FSM recognizers allows us to achieve real-time on-line performance. We apply this technique to build an experimental system that plays a game of "Simon Says" with the user. Pengyu Hong, Thomas S. Huang, Matthew Turk 0001 |
FG | 3 |
| 2000 | Constructing Finite State Machines for Fast Gesture RecognitionabstractProposes an approach to 2D gesture recognition that models each gesture as a finite state machine (FSM) in the spatial-temporal space. The model construction works in a semi-automatic way. The structure of the model is first manually decided based on the observation of the spatial topology of the data. The model is refined iteratively between two stages: data segmentation and model training. We incorporate a modified Knuth-Morris-Pratt algorithm recognition procedure to speed up recognition. The computational efficiency of the FSM recognizers allows real-time online performance to be achieved. Pengyu Hong, Thomas S. Huang, Matthew Turk 0001 |
ICPR | 3 |
| 1999 | Tracking Self-Occluding Articulated Objects in Dense Disparity MapsabstractIn this paper, we present an algorithm for real-time tracking of articulated structures in dense disparity maps derived from stereo image sequences. A statistical image formation model that accounts for occlusions plays the central role in our tracking approach. This graphical model (a Bayesian network) assumes that the range image of each part of the structure is formed by drawing the depth candidates from a 3-D Gaussian distribution. The advantage over the classical mixture of Gaussians is that our model takes into account occlusions by picking the minimum depth (which could be regarded as a probabilistic version of z-buffering). The model also enforces articulation constraints among the parts of the structure. The tracking problem is formulated as an inference problem in the image formation model. This model can be extended and used for other tasks in addition to the one described in the paper and can also be used for estimating probability distribution functions instead of the ML estimates of the tracked parameters. For the purposes of real-time tracking, we used certain approximations in the inference process, which resulted in a real-time two-stage inference algorithm. We were able to successfully track upper human body motion in real time and in the presence of self-occlusions. Nebojsa Jojic, Matthew Turk 0001, Thomas S. Huang |
ICCV | 2 |
| 1998 | View-Based Interpretation of Real-Time Optical Flow for Gesture Recognition
Ross Cutler, Matthew Turk 0001 |
FG | 2 |
| 1996 | Visual Interaction With Lifelike CharactersabstractThis paper explores the use of fast, simple computer vision techniques to add compelling visual capabilities to social user interfaces. Social interfaces involve the user in natural dialog with animated, "lifelike" characters. However, current systems employ spoken language as the only input modality. Used effectively, vision can greatly enhance the user's experience interacting with these characters. In addition, vision can provide key information to help manage the dialog and to aid the speech recognition process. We describe constraints imposed by the conversational environment and present a set of "interactive-time" vision routines that begin to support the user's expectations of a seeing character. A control structure is presented which chooses among the vision routines based on the current state of the character, the conversation and the visual environment. These capabilities are beginning to be integrated into the Persona lifelike character project. Matthew Turk 0001 |
FG | 1 |
| 1991 | Face recognition using eigenfacesabstractAn approach to the detection and identification of human faces is presented, and a working, near-real-time face recognition system which tracks a subject's head and then recognizes the person by comparing characteristics of the face to those of known individuals is described. This approach treats face recognition as a two-dimensional recognition problem, taking advantage of the fact that faces are normally upright and thus may be described by a small set of 2-D characteristic views. Face images are projected onto a feature space ('face space') that best encodes the variation among known face images. The face space is defined by the 'eigenfaces', which are the eigenvectors of the set of faces; they do not necessarily correspond to isolated features such as eyes, ears, and noses. The framework provides the ability to learn to recognize new faces in an unsupervised manner.> Matthew Turk 0001, Alex Pentland |
CVPR | 1 |
| 1989 | A simple, real-time range cameraabstractA simple imaging range sensor is described, based on the measurement of focal error, as described by A. Pentland (1982 and 1987). The current implementation can produce range over a 1 m/sup 3/ workspace with a measured standard error of 2.5% (4.5 significant bits of data). The system is implemented using relatively inexpensive commercial image-processing equipment. Experience shows that this ranging technique can be both economical and practical for tasks which require quick and reliable but coarse estimates of range. Examples of such tasks are initial target acquisition or obtaining the initial coarse estimate of stereo disparity in a coarse-to-fine stereo algorithm.> Alex Pentland, Trevor Darrell, Matthew Turk 0001 |
CVPR | 3 |
| 1988 | VITS-A Vision System for Autonomous Land Vehicle NavigationabstractA description is given of VITS (for vision task sequencer), the vision system for the autonomous land vehicle (ALV) Alvin, addressing in particular the task of road-following. The ALV vision system builds symbolic descriptions of road and obstacle boundaries using both video and range sensors. The authors discuss various road segmentation methods for video-based road-following, along with approaches to boundary extraction and transformation of boundaries in the image plane into a vehicle-centered three-dimensional scene model.> Matthew Turk 0001, David G. Morgenthaler, Keith D. Gremban, Martin Marra |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1987 | Video road-following for the autonomous land vehicleabstractWe describe the vision system for Alvin, the Autonomous Land Vehicle, addressing in particular the task of road-following. The system builds symbolic descriptions of the road and obstacle boundaries using both video and range sensors. Road segmentation methods are described for video-based road-following, along with approaches to boundary extraction and the transformation of boundaries in the image plane into a vehicle-centered three dimensional scene model. The ALV has performed public road-following demonstrations, traveling distances up to 4.5 km at speeds up to 20 km/hr along a paved road, equipped with an RGB video camera with pan/tilt control and a laser range scanner. Matthew Turk 0001, David G. Morgenthaler, Keith D. Gremban, Martin Marra |
ICRA | 1 |