VLDB 2026 Research / reviewers in the wild / expert
Stephen DiVerdi
dblp:03/2531 · also Stephen J. DiVerdi
· DBLP profile ↗
54ranked-venue papers
14as first author
5since 2021 · last 2025
0000-0002-6694-3381ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 41 · 14 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 23 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VidSTR: Automatic Spatiotemporal Retargeting of Speech-Driven Video Compositions
Joshua Kong Yang, Mackenzie Leake, Jeff Huang 0002, Stephen DiVerdi |
CHI | 4 |
| 2025 | A Scaffold-Based Tool for Product Design Variations in Virtual Reality
Stephen DiVerdi, Yotam I. Gingold |
CHI | 2 |
| 2025 | MaDCoW: Marginal Distortion Correction for Wide-Angle Photography with Arbitrary ObjectsabstractWe introduce MaDCoW, a method for correcting marginal distortion of arbitrary objects in wide-angle photography. People often use wide-angle photography—it is the default in smartphone cameras—but very-wide-fields-of-view produce distorted object appearance in image margins. In our system, a user annotates straight lines and regions of interest. MaDCoW solves for a separate linear perspective projection for each region and then jointly solves for a distortion-minimizing projection for the whole photograph. We show that MaDCoW can produce good results in cases where previous methods yield visible distortions. Kevin Zhang 0003, Jia-Bin Huang 0001, Jose Echevarria, Stephen DiVerdi, Aaron Hertzmann |
CVPR | 4 |
| 2022 | ZoomShop: Depth-Aware Editing of Photographic CompositionabstractAbstract We present ZoomShop, a photographic composition editing tool for adjusting relative size, position, and foreshortening of scene elements. Given an image and corresponding depth map as input, ZoomShop combines a novel non‐linear camera model and a depth‐aware image warp to reproject and deform the image. Users can isolate objects by selecting depth ranges and adjust their scale and foreshortening, which controls the paths of the camera rays through the scene. Users can also select 2D image regions and translate them, which determines the objective function in the image warp optimization. We demonstrate that ZoomShop can be used to achieve useful compositional goals, such as making a distant object more prominent while preserving foreground scenery, or making objects both larger and closer together so they still fit in the frame. Sean J. Liu, Maneesh Agrawala, Stephen DiVerdi, Aaron Hertzmann |
Comput. Graph. Forum | 3 |
| 2021 | ScaffoldSketch: Accurate Industrial Design Drawing in VRabstractWe present an approach to in-air design drawing based on the two-stage approach common in 2D design drawing practice. The primary challenge to 3D drawing in-air is the accuracy of users’ strokes. Beautifying or auto-correcting an arbitrary drawing in 2D or 3D is challenging due to ambiguities stemming from many possible interpretations of a stroke. A similar challenge appears when drawing freehand on paper in the real world. 2D design drawing practice (as taught in industrial design school) addresses this by decomposing the process of creating realistic 2D projections of 3D shapes. Designers first create scaffold or construction lines. When drawing shape or structure curves, designers are guided by the scaffolds. Our key insight is that accurate industrial design drawing in 3D becomes tractable when decomposed into auto-correcting scaffold strokes, which have simple relationships with one another, followed by auto-correcting shape strokes with respect to the scaffold strokes. We demonstrate our approach’s effectiveness with an expert study involving industrial designers. Stephen DiVerdi, Akshay Sharma, Yotam I. Gingold |
UIST | 2 |
| 2020 | Deep Multi Depth Panoramas for View Synthesis
Kai-En Lin, Zexiang Xu, Ben Mildenhall, Pratul P. Srinivasan, Yannick Hold-Geoffroy, Stephen DiVerdi, Qi Sun 0003, Kalyan Sunkavalli, Ravi Ramamoorthi |
ECCV (13) | 6 |
| 2020 | Texture Hallucination for Large-Factor Painting Super-Resolution
Yulun Zhang 0001, Stephen DiVerdi, Jose Echevarria, Yun Fu 0001 |
ECCV (7) | 3 |
| 2020 | View-Dependent Effects for 360° Virtual Reality Videoabstract"View-dependent effects'' have parameters that change with the user's view and are rendered dynamically at runtime. They can be used to simulate physical phenomena such as exposure adaptation, as well as for dramatic purposes such as vignettes. We present a technique for adding view-dependent effects to 360 degree video, by interpolating spatial keyframes across an equirectangular video to control effect parameters during playback. An in-headset authoring tool is used to configure effect parameters and set keyframe positions. We evaluate the utility of view-dependent effects with expert 360 degree filmmakers and the perception of the effects with a general audience. Results show that experts find view-dependent effects desirable for their creative purposes and that these effects can evoke novel experiences in an audience. Jeremy Hartmann, Stephen DiVerdi, Cuong Nguyen 0003, Daniel Vogel 0001 |
UIST | 2 |
| 2020 | TransceiVR: Bridging Asymmetrical Communication Between VR Users and External CollaboratorsabstractVirtual Reality (VR) users often need to work with other users, who observe them outside of VR using an external display. Communication between them is difficult; the VR user cannot see the external user's gestures, and the external user cannot see VR scene elements outside of the VR user's view. We carried out formative interviews with experts to understand these asymmetrical interactions and identify their goals and challenges. From this, we identify high-level system design goals to facilitate asymmetrical interactions and a corresponding space of implementation approaches based on the level of programmatic access to a VR application. We present TransceiVR, a system that utilizes VR platform APIs to enable asymmetric communication interfaces for third-party applications without requiring source code access. TransceiVR allows external users to explore the VR scene spatially or temporally, to annotate elements in the VR scene at correct depths, and to discuss via a shared static virtual display. An initial co-located user evaluation with 10 pairs shows that our system makes asymmetric collaborations in VR more effective and successful in terms of task time, error rate, and task load index. An informal evaluation with a remote expert gives additional insight on utility of features for real world tasks. Balasaravanan Thoravi Kumaravel, Cuong Nguyen 0003, Stephen DiVerdi, Björn Hartmann |
UIST | 3 |
| 2020 | RealitySketch: Embedding Responsive Graphics and Visualizations in AR through Dynamic SketchingabstractWe present RealitySketch, an augmented reality interface for sketching interactive graphics and visualizations. In recent years, an increasing number of AR sketching tools enable users to draw and embed sketches in the real world. However, with the current tools, sketched contents are inherently static, floating in mid-air without responding to the real world. This paper introduces a new way to embed dynamic and responsive graphics in the real world. In RealitySketch, the user draws graphical elements on a mobile AR screen and binds them with physical objects in real-time and improvisational ways, so that the sketched elements dynamically move with the corresponding physical motion. The user can also quickly visualize and analyze real-world phenomena through responsive graph plots or interactive visualizations. This paper contributes to a set of interaction techniques that enable capturing, parameterizing, and visualizing real-world motion without pre-defined programs and configurations. Finally, we demonstrate our tool with several application scenarios, including physics education, sports training, and in-situ tangible interfaces. Ryo Suzuki 0001, Rubaiat Habib Kazi, Li-Yi Wei, Stephen DiVerdi, Wilmot Li, Daniel Leithinger |
UIST | 4 |
| 2020 | Slicing-Volume: Hybrid 3D/2D Multi-target Selection Technique for Dense Virtual Environmentsabstract3D selection in dense VR environments (e.g., point clouds) is extremely challenging due to occlusion and imprecise mid-air input modalities (e.g., 3D controllers and hand gestures). In this paper, we propose "Slicing-Volume", a hybrid selection technique that enables simultaneous 3D interaction in mid-air, and a 2D pen-and-tablet metaphor in VR. Inspired by well-known slicing plane techniques in data visualization, our technique consists of a 3D volume that encloses target objects in mid-air, which are then projected to a 2D tablet view for precise selection on a tangible physical surface. While slicing techniques and tablets-in-VR have been previously explored, in this paper, we evaluated the potential of this hybrid approach to improve accuracy in highly occluded selection tasks, comparing different multimodal interactions (e.g., Mid-air, Virtual Tablet and Real Tablet). Our results showed that our hybrid technique significantly improved overall accuracy of selection compared to Mid-air selection only, thanks to the added haptic feedback given by the physical tablet surface, rather than the added visualization given by the tablet view. Roberto A. Montaño-Murillo, Cuong Nguyen 0003, Rubaiat Habib Kazi, Sriram Subramanian, Stephen DiVerdi, Diego Martínez 0001 |
VR | 5 |
| 2020 | Foreword to the Special Section on the 8th ACM/EG Expressive symposium (Expressive 2019)
Stephen DiVerdi, Craig S. Kaplan, Angus G. Forbes, Chiara Eva Catalano |
Comput. Graph. | 1 |
| 2019 | TutoriVR: A Video-Based Tutorial System for Design Applications in Virtual RealityabstractVirtual Reality painting is a form of 3D-painting done in a Virtual Reality (VR) space. Being a relatively new kind of art form, there is a growing interest within the creative practices community to learn it. Currently, most users learn using community posted 2D-videos on the internet, which are a screencast recording of the painting process by an instructor. While such an approach may suffice for teaching 2D-software tools, these videos by themselves fail in delivering crucial details that required by the user to understand actions in a VR space. We conduct a formative study to identify challenges faced by users in learning to VR-paint using such video-based tutorials. Informed by results of this study, we develop a VR-embedded tutorial system that supplements video tutorials with 3D and contextual aids directly in the user's VR environment. An exploratory evaluation showed users were positive about the system and were able to use the proposed system to recreate painting tasks in VR. Balasaravanan Thoravi Kumaravel, Cuong Nguyen 0003, Stephen DiVerdi, Björn Hartmann |
CHI | 3 |
| 2019 | View-Dependent Video Textures for 360° VideoabstractA major concern for filmmakers creating 360° video is ensuring that the viewer does not miss important narrative elements because they are looking in the wrong direction. This paper introduces gated clips which do not play the video past a gate time until a filmmaker-defined viewer gaze condition is met, such as looking at a specific region of interest (ROI). Until the condition is met, we seamlessly loop video playback using view-dependent video textures, a new variant of standard video textures that adapt the looping behavior to the portion of the scene that is within the viewer's field of view. We use our desktop GUI to edit live action and computer animated 360° videos. In a user study with casual viewers, participants prefer our looping videos over the standard versions and are able to successfully see all of the looping videos' ROIs without fear of missing important narrative content. Sean J. Liu, Maneesh Agrawala, Stephen DiVerdi, Aaron Hertzmann |
UIST | 3 |
| 2019 | Learning A Stroke-Based Representation for FontsabstractAbstract Designing fonts and typefaces is a difficult process for both beginner and expert typographers. Existing workflows require the designer to create every glyph, while adhering to many loosely defined design suggestions to achieve an aesthetically appealing and coherent character set. This process can be significantly simplified by exploiting the similar structure character glyphs present across different fonts and the shared stylistic elements within the same font. To capture these correlations, we propose learning a stroke‐based font representation from a collection of existing typefaces. To enable this, we develop a stroke‐based geometric model for glyphs, a fitting procedure to reparametrize arbitrary fonts to our representation. We demonstrate the effectiveness of our model through a manifold learning technique that estimates a low‐dimensional font space. Our representation captures a wide range of everyday fonts with topological variations and naturally handles discrete and continuous variations, such as presence and absence of stylistic elements as well as slants and weights. We show that our learned representation can be used for iteratively improving fit quality, as well as exploratory style applications such as completing a font from a subset of observed glyphs, interpolating or adding and removing stylistic elements in existing fonts. Elena Sizikova, Amit Bermano, Vladimir G. Kim, Stephen DiVerdi, Aaron Hertzmann, Thomas A. Funkhouser |
Comput. Graph. Forum | 4 |
| 2019 | Motion parallax for 360° RGBD videoabstractWe present a method for adding parallax and real-time playback of 360° videos in Virtual Reality headsets. In current video players, the playback does not respond to translational head movement, which reduces the feeling of immersion, and causes motion sickness for some viewers. Given a 360° video and its corresponding depth (provided by current stereo 360° stitching algorithms), a naive image-based rendering approach would use the depth to generate a 3D mesh around the viewer, then translate it appropriately as the viewer moves their head. However, this approach breaks at depth discontinuities, showing visible distortions, whereas cutting the mesh at such discontinuities leads to ragged silhouettes and holes at disocclusions. We address these issues by improving the given initial depth map to yield cleaner, more natural silhouettes. We rely on a three-layer scene representation, made up of a foreground layer and two static background layers, to handle disocclusions by propagating information from multiple frames for the first background layer, and then inpainting for the second one. Our system works with input from many of today's most popular 360° stereo capture devices (e.g., Yi Halo or GoPro Odyssey), and works well even if the original video does not provide depth information. Our user studies confirm that our method provides a more compelling viewing experience than without parallax, increasing immersion while reducing discomfort and nausea. Ana Serrano, Inchul Kim 0001, Stephen DiVerdi, Diego Gutierrez, Aaron Hertzmann, Belén Masiá |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | Pigmento: Pigment-Based Image Analysis and EditingabstractThe colorful appearance of a physical painting is determined by the distribution of paint pigments across the canvas, which we model as a per-pixel mixture of a small number of pigments with multispectral absorption and scattering coefficients. We present an algorithm to efficiently recover this structure from an RGB image, yielding a plausible set of pigments and a low RGB reconstruction error. We show that under certain circumstances we are able to recover pigments that are close to ground truth, while in all cases our results are always plausible. Using our decomposition, we repose standard digital image editing operations as operations in pigment space rather than RGB, with interestingly novel results. We demonstrate tonal adjustments, selection masking, cut-copy-paste, recoloring, palette summarization, and edge enhancement. Jianchao Tan, Stephen DiVerdi, Jingwan Lu, Yotam I. Gingold |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2018 | Multimodal Prediction and Personalization of Photo Edits with Deep Generative ModelsabstractProfessional-grade software applications are powerful but complicated – expert users can achieve impressive results, but novices often struggle to complete even basic tasks. Photo editing is a prime example: after loading a photo, the user is confronted with an array of cryptic sliders like "clarity", "temp", and "highlights". An automatically generated suggestion could help, but there is no single "correct" edit for a given image – different experts may make very different aesthetic decisions when faced with the same image, and a single expert may make different choices depending on the intended use of the image (or on a whim). We therefore want a system that can propose multiple diverse, high-quality edits while also learning from and adapting to a user’s aesthetic preferences. In this work, we develop a statistical model that meets these objectives. Our model builds on recent advances in neural network generative modeling and scalable inference, and uses hierarchical structure to learn editing patterns across many diverse users. Empirically, we find that our model outperforms other approaches on this challenging multimodal prediction task. Ardavan Saeedi, Matthew Hoffman 0001, Stephen DiVerdi, Asma Ghandeharioun, Matthew J. Johnson 0002, Ryan P. Adams |
AISTATS | 3 |
| 2018 | Depth Conflict Reduction for Stereo VR Video InterfacesabstractApplications for viewing and editing 360° video often render user interface (UI) elements on top of the video. For stereoscopic video, in which the perceived depth varies over the image, the perceived depth of the video can conflict with that of the UI elements, creating discomfort and making it hard to shift focus. To address this problem, we explore two new techniques that adjust the UI rendering based on the video content. The first technique dynamically adjusts the perceived depth of the UI to avoid depth conflict, and the second blurs the video in a halo around the UI. We conduct a user study to assess the effectiveness of these techniques in two stereoscopic VR video tasks: video watching with subtitles, and video search. Cuong Nguyen 0003, Stephen DiVerdi, Aaron Hertzmann, Feng Liu 0015 |
CHI | 2 |
| 2018 | Immersive Trip ReportsabstractSince the advent of consumer photography, tourists and hikers have made photo records of their trips to share later. Aside from being kept as memories, photo presentations such as slideshows are also shown to others who have not visited the location to try to convey the experience.However, a slideshow alone is limited in conveying the broader spatial context, and thus the feeling of presence in beautiful natural scenery is lost. We address this by presenting the photographs as part of an immersive experience. We introduce an automated pipeline for aligning photographs with a digital terrain model. From this geographic registration, we produce immersive presentations which are viewed either passively as a video, or interactively in virtual reality. Our experimental evaluation verifies that this new mode of presentation successfully conveys the spatial context of the scene and is enjoyable to users. Jan Brejcha, Michal Lukác, Stephen DiVerdi, Martin Cadík |
UIST | 4 |
| 2017 | Vremiere: In-Headset Virtual Reality Video EditingabstractCreative professionals are creating Virtual Reality (VR) experiences today by capturing spherical videos, but video editing is still done primarily in traditional 2D desktop GUI applications such as Premiere. These interfaces provide limited capabilities for previewing content in a VR headset or for directly manipulating the spherical video in an intuitive way. As a result, editors must alternate between editing on the desktop and previewing in the headset, which is tedious and interrupts the creative process. We demonstrate an application that enables a user to directly edit spherical video while fully immersed in a VR headset. We first interviewed professional VR filmmakers to understand current practice and derived a suitable workflow for in-headset VR video editing. We then developed a prototype system implementing this new workflow. Our system is built upon a familiar timeline design, but is enhanced with custom widgets to enable intuitive editing of spherical video inside the headset. We conducted an expert review study and found that with our prototype, experts were able to edit videos entirely within the headset. Experts also found our interface and widgets useful, providing intuitive controls for their editing needs. Cuong Nguyen 0003, Stephen DiVerdi, Aaron Hertzmann, Feng Liu 0015 |
CHI | 2 |
| 2017 | CollaVR: Collaborative In-Headset Review for VR VideoabstractCollaborative review and feedback is an important part of conventional filmmaking and now Virtual Reality (VR) video production as well. However, conventional collaborative review practices do not easily translate to VR video because VR video is normally viewed in a headset, which makes it difficult to align gaze, share context, and take notes. This paper presents CollaVR, an application that enables multiple users to review a VR video together while wearing headsets. We interviewed VR video professionals to distill key considerations in reviewing VR video. Based on these insights, we developed a set of networked tools that enable filmmakers to collaborate and review video in real-time. We conducted a preliminary expert study to solicit feedback from VR video professionals about our system and assess their usage of the system with and without collaboration features. Cuong Nguyen 0003, Stephen DiVerdi, Aaron Hertzmann, Feng Liu 0015 |
UIST | 2 |
| 2017 | VoCo: text-based insertion and replacement in audio narrationabstractEditing audio narration using conventional software typically involves many painstaking low-level manipulations. Some state of the art systems allow the editor to work in a text transcript of the narration, and perform select, cut, copy and paste operations directly in the transcript; these operations are then automatically applied to the waveform in a straightforward manner. However, an obvious gap in the text-based interface is the ability to type new words not appearing in the transcript, for example inserting a new word for emphasis or replacing a misspoken word. While high-quality voice synthesizers exist today, the challenge is to synthesize the new word in a voice that matches the rest of the narration. This paper presents a system that can synthesize a new word or short phrase such that it blends seamlessly in the context of the existing narration. Our approach is to use a text to speech synthesizer to say the word in a generic voice, and then use voice conversion to convert it into a voice that matches the narration. Offering a range of degrees of control to the editor, our interface supports fully automatic synthesis, selection among a candidate set of alternative pronunciations, fine control over edit placements and pitch profiles, and even guidance by the editors own voice. The paper presents studies showing that the output of our method is preferred over baseline methods and often indistinguishable from the original voice. Zeyu Jin, Gautham J. Mysore, Stephen DiVerdi, Jingwan Lu, Adam Finkelstein |
ACM Trans. Graph. | 3 |
| 2017 | Playful palette: an interactive parametric color mixer for artistsabstractWe present Playful Palette, a color picker interface for digital paint programs that derives intuition from oil paint and watercolor palettes, but extends them with digital features. A Playful Palette is a set of blobs of color that blend together to create gradients and gamuts. They can be directly manipulated to explore arrangements and harmonies. All edits are non-destructive, and an infinite history allows previous palettes to be revisited and modified, recoloring the painting. The Playful Palette design is motivated by a pilot study of how artists use paint palettes, and we evaluate the final design with a set of traditional and digital media painters to demonstrate that Playful Palette is effective both at enabling artists' color tasks, and at amplifying their creativity. Maria Shugrina, Jingwan Lu, Stephen DiVerdi |
ACM Trans. Graph. | 3 |
| 2016 | Cute: A concatenative method for voice conversion using exemplar-based unit selectionabstractState-of-the art voice conversion methods re-synthesize voice from spectral representations such as MFCCs and STRAIGHT, thereby introducing muffled artifacts. We propose a method that circumvents this concern using concatenative synthesis coupled with exemplar-based unit selection. Given parallel speech from source and target speakers as well as a new query from the source, our method stitches together pieces of the target voice. It optimizes for three goals: matching the query, using long consecutive segments, and smooth transitions between the segments. To achieve these goals, we perform unit selection at the frame level and introduce triphone-based preselection that greatly reduces computation and enforces selection of long, contiguous pieces. Our experiments show that the proposed method has better quality than baseline methods, while preserving high individuality. Zeyu Jin, Adam Finkelstein, Stephen DiVerdi, Jingwan Lu, Gautham J. Mysore |
ICASSP | 3 |
| 2016 | Geometric calibration for mobile, stereo, autofocus camerasabstractMobile, stereo, autofocus cameras present unique challenges for robust and efficient depth estimation. Specifically, existing approaches for calibration of the stereo camera intrinsic and extrinsic parameters are inadequate because of per-shot changes in the configuration, long-term mechanical drift, extremely constrained manufacturing processes, and the requirement for real-world robustness. We present a hybrid strategy that combines a single-photo of a calibration grid in an offline step with online blind refinement, which satisfies all of these goals. Stephen DiVerdi, Jonathan T. Barron |
WACV | 1 |
| 2015 | IsoMatch: Creating Informative Grid LayoutsabstractAbstract Collections of objects such as images are often presented visually in a grid because it is a compact representation that lends itself well for search and exploration. Most grid layouts are sorted using very basic criteria, such as date or filename. In this work we present a method to arrange collections of objects respecting an arbitrary distance measure. Pairwise distances are preserved as much as possible, while still producing the specific target arrangement which may be a 2D grid, the surface of a sphere, a hierarchy, or any other shape. We show that our method can be used for infographics, collection exploration, summarization, data visualization, and even for solving problems such as where to seat family members at a wedding. We present a fast algorithm that can work on large collections and quantitatively evaluate how well distances are preserved. Ohad Fried, Stephen DiVerdi, Maciej Halber, Elena Sizikova, Adam Finkelstein |
Comput. Graph. Forum | 2 |
| 2015 | Palette-based photo recoloringabstractImage editing applications offer a wide array of tools for color manipulation. Some of these tools are easy to understand but offer a limited range of expressiveness. Other more powerful tools are time consuming for experts and inscrutable to novices. Researchers have described a variety of more sophisticated methods but these are typically not interactive, which is crucial for creative exploration. This paper introduces a simple, intuitive and interactive tool that allows non-experts to recolor an image by editing a color palette. This system is comprised of several components: a GUI that is easy to learn and understand, an efficient algorithm for creating a color palette from an image, and a novel color transfer algorithm that recolors the image based on a user-modified palette. We evaluate our approach via a user study, showing that it is faster and easier to use than two alternatives, and allows untrained users to achieve results comparable to those of experts using professional software. Huiwen Chang, Ohad Fried, Yiming Liu 0001, Stephen DiVerdi, Adam Finkelstein |
ACM Trans. Graph. | 4 |
| 2015 | A Modular Framework for Digital PaintingabstractWhile there has been tremendous research in the simulation of natural media painting, little academic work has been written to understand how all these contributions interrelate and to use this knowledge to direct future work. In this paper, we survey the set of interesting artistic tools to categorize their effects and motivate a modular framework for digital painting that can reproduce those effects in a loosely coupled way. We use this framework as a lens through which we survey the literature and classify the achievements of previous efforts. We examine our own contributions in the field in more detail, discussing how the framework motivated those results and how it impacted our accomplishments. Finally, we discuss the open challenges that remain for the research community, and how the framework can help to make contributions towards those challenges. Stephen DiVerdi |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2014 | Auto-rectification of user photosabstractThe image auto rectification project at Google aims to create a pleasanter version of user photos by correcting the small, involuntary camera rotations (roll / pitch/ yaw) that often occur in non-professional photographs. Our system takes the image closer to the fronto-parallel view by performing an affine rectification on the image that restores parallelism of lines that are parallel in the fronto-parallel image view. This partially corrects perspective distortions, but falls short of full metric rectification which also restores angles between lines. On the other hand the 2D homography for our rectification can be computed from only two (as opposed to three) estimated vanishing points, allowing us to fire upon many more images. A new RANSAC based approach to vanishing point estimation has been developed. The main strength of our vanishing point detector is that it is line-less, thereby avoiding the hard, binary (line/no-line) upstream decisions that cause traditional algorithm to ignore much supporting evidence and/or admit noisy evidence for vanishing points. A robust RANSAC based technique for detecting horizon lines in an image is also proposed for analyzing correctness of the estimated rectification. We post-multiply our affine rectification homography with a 2D rotation which aligns the closer vanishing point with the image Y axis. Krishnendu Chaudhury, Stephen DiVerdi, Sergey Ioffe |
ICIP | 2 |
| 2013 | Wip chairsabstractNew this year to ISMAR 2013, we are proud to present the Works In Progress (WIP) Program. Augmented Reality is rapidly growing into many new areas, so the WIP is a platform to present the field's latest, emerging results to the larger community before the work has reached its final form. This year, the program includes bread and butter AR technologies such as remote collaboration interfaces, fiducial marker design, and perceptual studies, to even loftier applications like AR interactions aboard the International Space Station. Passive haptics, bare-handed gesture interfaces, and realistic rendering round out the offerings. So come to the WIP sessions to hear about active AR research and find the spark of inspiration! Stephen DiVerdi, Jun Park |
ISMAR | 1 |
| 2013 | Learning part-based templates from large collections of 3D shapesabstractAs large repositories of 3D shape collections continue to grow, understanding the data, especially encoding the inter-model similarity and their variations, is of central importance. For example, many data-driven approaches now rely on access to semantic segmentation information, accurate inter-model point-to-point correspondence, and deformation models that characterize the model collections. Existing approaches, however, are either supervised requiring manual labeling; or employ super-linear matching algorithms and thus are unsuited for analyzing large collections spanning many thousands of models. We propose an automatic algorithm that starts with an initial template model and then jointly optimizes for part segmentation, point-to-point surface correspondence, and a compact deformation model to best explain the input model collection. As output, the algorithm produces a set of probabilistic part-based templates that groups the original models into clusters of models capturing their styles and variations. We evaluate our algorithm on several standard datasets and demonstrate its scalability by analyzing much larger collections of up to thousands of shapes. Vladimir G. Kim, Wilmot Li, Niloy J. Mitra, Siddhartha Chaudhuri, Stephen DiVerdi, Thomas A. Funkhouser |
ACM Trans. Graph. | 5 |
| 2013 | RealBrush: painting with examples of physical mediaabstractConventional digital painting systems rely on procedural rules and physical simulation to render paint strokes. We present an interactive, data-driven painting system that uses scanned images of real natural media to synthesize both new strokes and complex stroke interactions, obviating the need for physical simulation. First, users capture images of real media, including examples of isolated strokes, pairs of overlapping strokes, and smudged strokes. Online, the user inputs a new stroke path, and our system synthesizes its 2D texture appearance with optional smearing or smudging when strokes overlap. We demonstrate high-fidelity paintings that closely resemble the captured media style, and also quantitatively evaluate our synthesis quality via user studies. Jingwan Lu, Connelly Barnes, Stephen DiVerdi, Adam Finkelstein |
ACM Trans. Graph. | 3 |
| 2013 | Painting with Polygons: A Procedural Watercolor EngineabstractExisting natural media painting simulations have produced high-quality results, but have required powerful compute hardware and have been limited to screen resolutions. Digital artists would like to be able to use watercolor-like painting tools, but at print resolutions and on lower end hardware such as laptops or even slates. We present a procedural algorithm for generating watercolor-like dynamic paint behaviors in a lightweight manner. Our goal is not to exactly duplicate watercolor painting, but to create a range of dynamic behaviors that allow users to achieve a similar style of process and result, while at the same time having a unique character of its own. Our stroke representation is vector based, allowing for rendering at arbitrary resolutions, and our procedural pigment advection algorithm is fast enough to support painting on slate devices. We demonstrate our technique in a commercially available slate application used by professional artists. Finally, we present a detailed analysis of the different vector-rendering technologies available. Stephen DiVerdi, Aravind Krishnaswamy, Radomír Mech, Daichi Ito |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2012 | A lightweight, procedural, vector watercolor painting engineabstractExisting natural media painting simulations have produced high quality results, but have required powerful compute hardware and have been limited to screen resolutions. Digital artists would like to be able to use watercolor-like painting tools, but at print resolutions and on lower end hardware such as laptops or even slates. We present a procedural algorithm for generating watercolor-like dynamic paint behaviors in a lightweight manner. Our goal is not to exactly duplicate watercolor painting, but to create a range of dynamic behaviors that allow users to achieve a similar style of process and result, while at the same time having a unique character of its own. Our stroke representation is vector-based, allowing for rendering at arbitrary resolutions, and our procedural pigment advection algorithm is fast enough to support painting on slate devices. We demonstrate our technique in a commercially available slate application used by professional artists. Stephen DiVerdi, Aravind Krishnaswamy, Radomír Mech, Daichi Ito |
I3D | 1 |
| 2012 | Exploring collections of 3D models using fuzzy correspondencesabstractLarge collections of 3D models from the same object class (e.g., chairs, cars, animals) are now commonly available via many public repositories, but exploring the range of shape variations across such collections remains a challenging task. In this work, we present a new exploration interface that allows users to browse collections based on similarities and differences between shapes in user-specified regions of interest (ROIs). To support this interactive system, we introduce a novel analysis method for computing similarity relationships between points on 3D shapes across a collection. We encode the inherent ambiguity in these relationships using fuzzy point correspondences and propose a robust and efficient computational framework that estimates fuzzy correspondences using only a sparse set of pairwise model alignments. We evaluate our analysis method on a range of correspondence benchmarks and report substantial improvements in both speed and accuracy over existing alternatives. In addition, we demonstrate how fuzzy correspondences enable key features in our exploration tool, such as automated view alignment, ROI-based similarity search, and faceted browsing. Vladimir G. Kim, Wilmot Li, Niloy J. Mitra, Stephen DiVerdi, Thomas A. Funkhouser |
ACM Trans. Graph. | 4 |
| 2012 | HelpingHand: example-based stroke stylizationabstractDigital painters commonly use a tablet and stylus to drive software like Adobe Photoshop. A high quality stylus with 6 degrees of freedom (DOFs: 2D position, pressure, 2D tilt, and 1D rotation) coupled to a virtual brush simulation engine allows skilled users to produce expressive strokes in their own style. However, such devices are difficult for novices to control, and many people draw with less expensive (lower DOF) input devices. This paper presents a data-driven approach for synthesizing the 6D hand gesture data for users of low-quality input devices. Offline, we collect a library of strokes with 6D data created by trained artists. Online, given a query stroke as a series of 2D positions, we synthesize the 4D hand pose data at each sample based on samples from the library that locally match the query. This framework optionally can also modify the stroke trajectory to match characteristic shapes in the style of the library. Our algorithm outputs a 6D trajectory that can be fed into any virtual brush stroke engine to make expressive strokes for novices or users of limited hardware. Jingwan Lu, Fisher Yu 0001, Adam Finkelstein, Stephen DiVerdi |
ACM Trans. Graph. | 4 |
| 2010 | Industrial-strength painting with a virtual bristle brushabstractResearch in natural media painting has produced impressive images, but those results have not been adopted by commercial applications to date because of the heavy demands of industrial painting workflows. In this paper, we present a new 3D brush model with associated algorithms for stroke generation and bidirectional paint transfer that is suitable for professional use. Our model can reproduce arbitrary brush tip shapes and can be used to generate raster or vector output, none of which was possible in previous simulations. This is achieved by an efficient formulation of bristle behaviors as strand dynamics in a non-inertial reference frame. To demonstrate the robustness and flexibility of our approach, we have integrated our model into major commercial painting and vector editing applications and given it to professional artists to evaluate. Stephen DiVerdi, Aravind Krishnaswamy, Sunil Hadap |
VRST | 1 |
| 2009 | All around the map: Online spherical panorama construction
Stephen DiVerdi, Jason Wither, Tobias Höllerer |
Comput. Graph. | 1 |
| 2009 | Annotation in outdoor augmented reality
Jason Wither, Stephen DiVerdi, Tobias Höllerer |
Comput. Graph. | 2 |
| 2009 | Mid-air display experiments to create novel user interfacesabstractDisplays are the most visible part of most computer applications. Novel display technologies strongly influence and inspire new forms of computer use and interaction. We are particularly interested in the interplay of novel displays and interaction for ubiquitous computing or ambient media environments, as emerging display technologies may become game-changers in how we define and use computers, possibly changing the context of computing fundamentally. We present some of our experiments and lessons learnt with a new category of displays, the “immaterial” FogScreen. It can be described as a novel media platform, exhibiting some fundamental differences to and advantages over other displays. It also enables novel kinds of user interfaces and experiences. In this paper we give insights about the special properties and strengths of the FogScreen by looking at a set of successfully demonstrated interfaces and applications. We also discuss its future potential for user interface design. Ismo Rakkolainen, Tobias Höllerer, Stephen DiVerdi, Alex Olwal |
Multim. Tools Appl. | 3 |
| 2009 | Depth-Fused 3D Imagery on an Immaterial DisplayabstractWe present an immaterial display that uses a generalized form of depth-fused 3D (DFD) rendering to create unencumbered 3D visuals. To accomplish this result, we demonstrate a DFD display simulator that extends the established depth-fused 3D principle by using screens in arbitrary configurations and from arbitrary viewpoints. The feasibility of the generalized DFD effect is established with a user study using the simulator. Based on these results, we developed a prototype display using one or two immaterial screens to create an unencumbered 3D visual that users can penetrate, examining the potential for direct walk-through and reach-through manipulation of the 3D scene. We evaluate the prototype system in formative and summative user studies and report the tolerance thresholds discovered for both tracking and projector errors. Cha Lee, Stephen DiVerdi, Tobias Höllerer |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2008 | Envisor: Online Environment Map Construction for Mixed RealityabstractOne of the main goals of anywhere augmentation is the development of automatic algorithms for scene acquisition in augmented reality systems. In this paper, we present Envisor, a system for online construction of environment maps in new locations. To accomplish this, Envisor uses vision-based frame to frame and landmark orientation tracking for long-term, drift-free registration. For additional robustness, a gyroscope/compass orientation unit can optionally be used for hybrid tracking. The tracked video is then projected into a cubemap frame by frame. Feedback is presented to the user to help avoid gaps in the cubemap, while any remaining gaps are filled by texture diffusion. The resulting environment map can be used for a variety of applications, including shading of virtual geometry and remote presence. Stephen DiVerdi, Jason Wither, Tobias Höllerer |
VR | 1 |
| 2008 | Heads Up and Camera Down: A Vision-Based Tracking Modality for Mobile Mixed RealityabstractAnywhere Augmentation pursues the goal of lowering the initial investment of time and money necessary to participate in mixed reality work, bridging the gap between researchers in the field and regular computer users. Our paper contributes to this goal by introducing the GroundCam, a cheap tracking modality with no significant setup necessary. By itself, the GroundCam provides high frequency, high resolution relative position information similar to an inertial navigation system, but with significantly less drift. We present the design and implementation of the GroundCam, analyze the impact of several design and run-time factors on tracking accuracy, and consider the implications of extending our GroundCam to different hardware configurations. Motivated by the performance analysis, we developed a hybrid tracker that couples the GroundCam with a wide area tracking modality via a complementary Kalman filter, resulting in a powerful base for indoor and outdoor mobile mixed reality work. To conclude, the performance of the hybrid tracker and its utility within mixed reality applications is discussed. Stephen DiVerdi, Tobias Höllerer |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2007 | Evaluating Display Types for AR Selection and AnnotationabstractThis paper evaluates different display devices for selection or annotation tasks in augmented reality (AR). We compare three different display types - a head mounted display and two hand held displays. The first hand held display is configured as a magic lens where the user sees the augmented space directly behind the display. The second hand held display is configured to be used at waist level (as one would commonly hold a tablet computer) but the view is still of the scene in front of the user. Making a selection or annotation in AR requires two distinct tasks by the user. First, the user must find the real (or virtual) object they want to mark. Second, the user must move a cursor to the object's location. We test and compare our three representative displays with respect to both tasks. We found that using a hand held display in the magic lens configuration was faster for cursor movement than either of the other two displays. There was no significant difference among the displays regarding the amount of time it took users to search for either physical or virtual objects. Jason Wither, Stephen DiVerdi, Tobias Höllerer |
ISMAR | 2 |
| 2007 | GroundCam: A Tracking Modality for Mobile Mixed RealityabstractAnywhere augmentation pursues the goal of lowering the initial investment of time and money necessary to participate in mixed reality work, bridging the gap between researchers in the field and regular computer users. Our paper contributes to this goal by introducing the GroundCam, a cheap tracking modality with no significant setup necessary. By itself, the GroundCam provides high frequency, high resolution relative position information similar to an inertial navigation system, but with significantly less drift. When coupled with a wide area tracking modality via a complementary Kalman filter, the hybrid tracker becomes a powerful base for indoor and outdoor mobile mixed reality work Stephen DiVerdi, Tobias Höllerer |
VR | 1 |
| 2007 | Implicit 3D modeling and tracking for anywhere augmentationabstractThis paper presents an online 3D modeling and tracking methodology that uses aerial photographs for mobile augmented reality. Instead of relying on models which are created in advance, the system generates a 3D model for a real building on the fly by combining frontal and aerial views with the help of an optical sensor, an inertial sensor, a GPS unit and a few mouse clicks. A user’s initial pose is estimated using an aerial photograph, which is retrieved from a database according to the user’s GPS coordinates, and an inertial sensor which measures pitch. To track the user’s position and orientation in real-time, feature-based tracking is carried out based on salient points on the edges and the sides of a building the user is keeping in view. We implemented camera pose estimators using both a least squares and an unscented Kalman filter (UKF) approach. The UKF approach results in more stable and reliable vision-based tracking. We evaluate the speed and accuracy of both approaches, and we demonstrate the usefulness of our computations as important building blocks for an Anywhere Augmentation scenario. Sehwan Kim, Stephen DiVerdi, Jae Sik Chang, Taehyuk Kang, Ronald A. Iltis, Tobias Höllerer |
VRST | 2 |
| 2007 | An immaterial depth-fused 3D displayabstractWe present an immaterial display that uses a generalized form of depth-fused 3D (DFD) rendering to create unencumbered 3D visuals. To accomplish this result, we demonstrate a DFD display simulator that extends the established depth-fused 3D principle by using screens in arbitrary configurations and from arbitrary viewpoints. The performance of the generalized DFD effect is established with a user study using the simulator. Based on these results, we developed a prototype display using two immaterial screens to create an unencumbered 3D visual that users can penetrate, enabling the potential for direct walk-through and reach-through manipulation of the 3D scene. Cha Lee, Stephen DiVerdi, Tobias Höllerer |
VRST | 2 |
| 2006 | 3DTV - Panoramic 3D Model Acquisition and its 3D Visualization on the Interactive FogscreenabstractFuture 3D television critically relies on mechanisms for automatically acquiring and visualizing high quality 3D content of both indoor and outdoor scenes. The envisioned goal is that a photo-realistic 3D real-time rendering from the actual and potentially arbitrary viewpoint of the beholder who is watching 3DTV becomes possible. Such scenes include movie sets in studios, e.g., for talk shows, TV series and blockbuster movies, but also outdoor scenes, e.g., buildings in a neighborhood for a car chase or cultural heritage sites for a documentary. The goal of 3D model acquisition is to provide the 3D background models where potential 3D actors can be embedded. We present both the 3D acquisition and semi-immersive 3D visualization to give an impression how a future 3D television system could be like. Sven Fleck, Florian Busch, Peter Biber, Wolfgang Straßer, Ismo Rakkolainen, Stephen DiVerdi, Tobias Höllerer |
ICIP | 6 |
| 2006 | Using aerial photographs for improved mobile AR annotationabstractWe present a mobile augmented reality system for outdoor annotation of the real world. To reduce user burden, we use aerial photographs in addition to the wearable system's usual data sources (position, orientation, camera and user input). This allows the user to accurately annotate 3D features with only a few simple interactions from a single position by aligning features in both their first-person viewpoint and in the aerial view. We examine three types of aerial photograph features - corners, edges, and regions - that are suitable for a wide variety of useful mobile augmented reality applications, and are easily visible on aerial photographs. By using aerial photographs in combination with wearable augmented reality, we are able to achieve much higher accuracy 3D annotation positions than was previously possible from a single user location. Jason Wither, Stephen DiVerdi, Tobias Höllerer |
ISMAR | 2 |
| 2006 | Image-space Correction of AR Registration Errors Using Graphics Hardwareabstractdirectly on top of physical objects in a video scene. Registration accuracy is a serious problem in these cases since any imprecisions are immediately apparent as virtual and physical edges and features coincide. We present a hardware-accelerated image-based post-processing technique that adjusts rendering of virtual geometry to better match edges present in images of a physical scene, reducing the visual effect of registration errors from both inaccurate tracking and oversimplified modeling. Our algorithm is easily integrable with existing AR applications, having no dependency on the underlying tracking technique. We use the advanced programmable capabilities of modern graphics hardware to achieve high performance without burdening the CPU. Stephen DiVerdi, Tobias Höllerer |
VR | 1 |
| 2006 | An Immaterial, Dual-sided Display System with 3D InteractionabstractWe present an interactive wall-sized immaterial display that introduces a number of interesting possibilities for advanced interface design. The immaterial nature of a thin sheet of fog allows users to penetrate and even walk through the screen, while its dual-sided nature allows for new possibilities in multi-user faceto- face collaboration and pseudo-3D visualization. Alex Olwal, Stephen DiVerdi, Nicola Candussi, Ismo Rakkolainen, Tobias Höllerer |
VR | 2 |
| 2004 | Level of Detail InterfacesabstractWe present the level of detail interface based on the marriage of level of detail geometry and an adaptable user interface. Level of detail interfaces allow applications to paramaterize their display of data and interface widgets with respect to distance from the camera, to best take advantage of diminished screen space in a 3D environment. Stephen DiVerdi, Tobias Höllerer, Richard Schreyer |
ISMAR | 1 |
| 2003 | ARWin-A Desktop Augmented Reality Window ManagerabstractWe present ARWin, a single user 3D augmented reality desktop. We explain our design considerations and system architecture and discuss a variety of applications and interaction techniques designed to take advantage of this new platform. Stephen DiVerdi, Daniel Nurmi, Tobias Höllerer |
ISMAR | 1 |