Cuong Nguyen 0003

dblp:00/6661-3 · DBLP profile ↗
← Back
27ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0001-9234-9960ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 25 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021
YearPublicationVenuePosition
2026 DepthScape: Authoring 2.5D Designs via Depth Estimation, Semantic Understanding, and Geometry Extraction
abstract
2.5D effects, such as occlusion and perspective foreshortening, enhance visual dynamics and realism by introducing 3D depth cues into 2D designs. However, creating these effects remains challenging, as designers must manually infer and author depth relationships—such as relative ordering, occlusion boundaries, and perspective scaling—within 2D representations. We introduce DepthScape, a human–AI collaborative system that facilitates 2.5D effect creation by placing design elements directly into 3D reconstructions. Using monocular depth reconstruction, DepthScape transforms images into 3D scenes, enabling depth-based blending that produces realistic occlusion and perspective foreshortening. To simplify 3D placement, DepthScape leverages a vision-language model to analyze source images and extract key visual components as parametric anchors, which support direct manipulation editing. The system design was iteratively refined through a formative user study with an early prototype. We evaluate DepthScape through a technical study on 100 professional stock images to assess robustness, alongside an expert evaluation confirming design quality, usefulness, and broad application potential, further illustrated through five example scenarios.
Xia Su, Cuong Nguyen 0003, Matheus A. Gadelha, Jon Froehlich
DIS2
2026 Visual Lyrics: Generating Animated Text for Music Lyric Videos with an Augmented Text Editor
abstract
Animated lyric videos transform song lyrics into dynamic visual experiences, offering a powerful medium for artistic expression and audience engagement. However, creating these videos is challenging, requiring expertise in audio, typography, graphic design, and animation, making it inaccessible to novices. To address this challenge, we introduce Visual Lyrics, a proof-of-concept system for generating animated lyric videos controlled with an augmented text editor interface. We examined existing lyric videos to distill a taxonomy and design guidelines, informing the design of Visual Lyrics. Our key insight is a multimodal music analysis pipeline based on the taxonomy and leveraging LLM’s strong natural language understanding and code generation capabilities to synthesize creative and semantically meaningful animations. We collected a dataset of over 300 code-driven creative text animations to serve as inspiration for our LLM-driven pipeline, which we open source. In a user study, Visual Lyrics enabled novices to easily create high-quality animated lyric videos with high ratings of enjoyment, inspiration, and exploration.
David Chuan-En Lin, Cuong Nguyen 0003, Hijung Shin, Nikolas Martelaro
IUI2
2026 Role-Aware Virtual Agents for Navigational Interaction guided by a Multimodal Large Language Model
abstract
We present a role-aware virtual agent navigational interaction that generates consistent, role-aligned movement behaviors. Our approach leverages Multimodal Large Language Models (MLLMs) to interpret multimodal inputs including scene information, user state, and high-level language role instruction, producing discrete navigation decisions and stylized planning path. Our approach enables virtual agents to behave consistently with narrative roles and respond to dynamic actions, such as playing a hide-and-seek taking into account the agent's role and the user's possible intention. Our approach demonstrates how MLLMs can go beyond language-based interaction to support embodied, spatial, and role-aware agent behaviors in immersive environments such as augmented reality.
ChangYang Li, Cuong Nguyen 0003, Lap-Fai Yu
ACM Trans. Graph.3
2024 MemoVis: A GenAI-Powered Tool for Creating Companion Reference Images for 3D Design Feedback
abstract
Providing asynchronous feedback is a critical step in the 3D design workflow. A common approach to providing feedback is to pair textual comments with companion reference images, which helps illustrate the gist of text. Ideally, feedback providers should possess 3D and image editing skills to create reference images that can effectively describe what they have in mind. However, they often lack such skills, so they have to resort to sketches or online images that might not match well with the current 3D design. To address this, we introduce MemoVis , a text editor interface that assists feedback providers in creating reference images with generative AI driven by the feedback comments. First, a novel real-time viewpoint suggestion feature, based on a vision-language foundation model, helps feedback providers anchor a comment with a camera viewpoint. Second, given a camera viewpoint, we introduce three types of image modifiers based on pre-trained 2D generative models to turn a text comment into an updated version of the 3D scene from that viewpoint. We conducted a within-subjects study with \(14\) feedback providers, demonstrating the effectiveness of MemoVis. The quality and explicitness of the companion images were evaluated by another eight participants with prior 3D design experience.
Chen Chen 0070, Cuong Nguyen 0003, Thibault Groueix, Vladimir G. Kim, Nadir Weibel
ACM Trans. Comput. Hum. Interact.2
2023 PointShopAR: Supporting Environmental Design Prototyping Using Point Cloud in Augmented Reality
abstract
We present PointShopAR, a novel tablet-based system for AR environmental design using point clouds as the underlying representation. It integrates point cloud capture and editing in a single AR workflow to help users quickly prototype design ideas in their spatial context. We hypothesize that point clouds are well suited for prototyping, as they can be captured more rapidly than textured meshes and then edited immediately in situ on the capturing device. We based the design of PointShopAR on the practical needs of six architects in a formative study. Our system supports a variety of point cloud editing operations in AR, including selection, transformation, hole filling, drawing, morphing, and animation. We evaluate PointShopAR through a remote study on usability and an in-person study on environmental design support. Participants were able to iterate design rapidly, showing the merits of an integrated capture and editing workflow with point clouds in AR environmental design.
Zeyu Wang 0003, Cuong Nguyen 0003, Paul Asente, Julie Dorsey
CHI2
2023 PaperToPlace: Transforming Instruction Documents into Spatialized and Context-Aware Mixed Reality Experiences
abstract
While paper instructions are a mainstream medium for sharing knowledge, consuming such instructions and translating them into activities can be inefficient due to the lack of connectivity with the physical environment. We propose PaperToPlace, a novel workflow comprising an authoring pipeline, which allows the authors to rapidly transform and spatialize existing paper instructions into an MR experience, and a consumption pipeline, which computationally places each instruction step at an optimal location that is easy to read and does not occlude key interaction areas. Our evaluation of the authoring pipeline with 12 participants demonstrates the usability of our workflow and the effectiveness of using a machine learning based approach to help extract the spatial locations associated with each step. A second within-subjects study with another 12 participants demonstrates the merits of our consumption pipeline to reduce context-switching effort by delivering individual segmented instruction steps and offering hands-free affordances.
Chen Chen 0070, Cuong Nguyen 0003, Jane Hoffswell, Jennifer A. Healey, Trung Bui, Nadir Weibel
UIST2
2023 GestureCanvas: A Programming by Demonstration System for Prototyping Compound Freehand Interaction in VR
abstract
As the use of hand gestures becomes increasingly prevalent in virtual reality (VR) applications, prototyping Compound Freehand Interactions (CFIs) effectively and efficiently has become a critical need in the design process. Compound Freehand Interaction (CFI) is a sequence of freehand interactions where each sub-interaction in the sequence conditions the next. Despite the need for interactive prototypes of CFI in the early design stage, creating them is effortful and remains a challenge for designers since it requires a highly technical workflow that involves programming the recognizers, system responses and conditionals for each sub-interaction. To bridge this gap, we present GestureCanvas, a freehand interaction-based immersive prototyping system that enables a rapid, end-to-end, and code-free workflow for designing, testing, refining, and subsequently deploying CFI by leveraging the three pillars of interaction models: event-driven state machine, trigger-action authoring, and programming by demonstration. The design of GestureCanvas includes three novel design elements — (i) appropriating the multimodal recording of freehand interaction into a CFI authoring workspace called Design Canvas, (ii) semi-automatic identification of the input trigger logic from demonstration to reduce the manual effort of setting up triggers for each sub-interaction, (iii) on the fly testing for independently validating the input conditionals in-situ. We validate the workflow enabled by GestureCanvas through an interview study with professional designers and evaluate its usability through a user study with non-experts. Our work lays the foundation for advancing research on immersive prototyping systems allowing even highly complex gestures to be easily prototyped and tested within VR environments.
Anika Sayara, Emily Lynn Chen, Cuong Nguyen 0003, Robert Xiao, Dongwook Yoon
UIST3
2023 PoseVEC: Authoring Adaptive Pose-aware Effects using Visual Programming and Demonstrations
abstract
Pose-aware visual effects where graphics assets and animations are rendered reactively to the human pose have become increasingly popular, appearing on mobile devices, the web, or even head-mounted displays like AR glasses. Yet, creating such effects still remains difficult for novices. In a traditional video editing workflow, a creator could utilize keyframes to create expressive but non-adaptive results which cannot be reused for other videos. Alternatively, programming-based approaches allow users to develop interactive effects, but are cumbersome for users to quickly express their creative intents. In this work, we propose a lightweight visual programming workflow for authoring adaptive and expressive pose effects. By combining a programming by demonstration paradigm with visual programming, we simplify three key tasks in the authoring process: creating pose triggers, designing animation parameters, and rendering. We evaluated our system with a qualitative user study and a replicated example study, finding that all participants can create effects efficiently.
Cuong Nguyen 0003, Rubaiat Habib Kazi, Lap-Fai Yu
UIST2
2023 WARPY: Sketching Environment-Aware 3D Curves in Mobile Augmented Reality
abstract
Three-dimensional curve drawing in Augmented Reality (AR) enables users to create 3D curves that fit within the real-world scene. It has applications in 3D design, sculpting, and animation. However, the task complexity increases when the desirable path for the curve is obstructed by the physical environment or by what the camera can see. For example, it is difficult to draw a curve that wraps around an object or scales to out-of-reach places. We propose WARPY, an environment-aware 3D curve drawing tool for mobile AR. Our system enables users to draw freeform curves from a distance in AR by combining 2D-to-3D sketch inference with geometric proxies. Geometric Proxies can be obtained via 3D scanning or from a list of pre-defined primitives. WARPY also provides a multi-view mode to enable users to sketch a curve from multiple viewpoints, which is useful if the target curve cannot fit within the camera's field of view. We conducted two user studies and found that WARPY can be a viable tool to help users create complex and large curves in AR.
Rawan Alghofaili, Cuong Nguyen 0003, Vojtech Krs, Nathan Carr 0001, Radomír Mech, Lap-Fai Yu
VR2
2023 Using Online Videos as the Basis for Developing Design Guidelines: A Case Study of AR-Based Assembly Instructions
abstract
Design guidelines serve as an important conceptual tool to guide designers of interactive applications with well-established principles and heuristics. Consulting domain experts is a common way to develop guidelines. However, experts are often not easily accessible, and their time can be expensive. This problem poses challenges in developing comprehensive and practical guidelines. We propose a new guideline development method that uses online public videos as the basis for capturing diverse patterns of design goals and interaction primitives. In a case study focusing on AR-based assembly instructions, we apply our novel Identify-Rationalize pipeline, which distills design patterns from videos featuring AR-based assembly instructions (N=146) into a set of guidelines that cover a wide range of design considerations. The evaluation conducted with 16 AR designers indicated that the pipeline is useful for generating comprehensive guidelines. We conclude by discussing the transferability and practicality of our method.
Niu Chen, Frances Jihae Sin, Laura Mariah Herman, Cuong Nguyen 0003, Ivan Song, Dongwook Yoon
Proc. ACM Hum. Comput. Interact.4
2023 VideoDoodles: Hand-Drawn Animations on Videos with Scene-Aware Canvases
abstract
We present an interactive system to ease the creation of so-called video doodles - videos on which artists insert hand-drawn animations for entertainment or educational purposes. Video doodles are challenging to create because to be convincing, the inserted drawings must appear as if they were part of the captured scene. In particular, the drawings should undergo tracking, perspective deformations and occlusions as they move with respect to the camera and to other objects in the scene - visual effects that are difficult to reproduce with existing 2D video editing software. Our system supports these effects by relying on planar canvases that users position in a 3D scene reconstructed from the video. Furthermore, we present a custom tracking algorithm that allows users to anchor canvases to static or dynamic objects in the scene, such that the canvases move and rotate to follow the position and direction of these objects. When testing our system, novices could create a variety of short animated clips in a dozen of minutes, while professionals praised its speed and ease of use compared to existing tools.
Emilie Yu, Kevin Matzen, Cuong Nguyen 0003, Oliver Wang, Rubaiat Habib Kazi, Adrien Bousseau
ACM Trans. Graph.3
2021 DistanciAR: Authoring Site-Specific Augmented Reality Experiences for Remote Environments
abstract
Most augmented reality (AR) authoring tools only support the author’s current environment, but designers often need to create site-specific experiences for a different environment. We propose DistanciAR, a novel tablet-based workflow for remote AR authoring. Our baseline solution involves three steps. A remote environment is captured by a camera with LiDAR; then, the author creates an AR experience from a different location using AR interactions; finally, a remote viewer consumes the AR content on site. A formative study revealed understanding and navigating the remote space as key challenges with this solution. We improved the authoring interface by adding two novel modes: Dollhouse, which renders a bird’s-eye view, and Peek, which creates photorealistic composite images using captured images. A second study compared this improved system with the baseline, and participants reported that the new modes made it easier to understand and navigate the remote scene.
Zeyu Wang 0003, Cuong Nguyen 0003, Paul Asente, Julie Dorsey
CHI2
2021 Rapido: Prototyping Interactive AR Experiences through Programming by Demonstration
abstract
Programming by Demonstration (PbD) is a well-known technique that allows non-programmers to describe interactivity by performing examples of the expected behavior, but it has not been extensively explored for AR. We present Rapido, a novel early-stage prototyping tool to create fully interactive mobile AR prototypes from non-interactive video prototypes using PbD. In Rapido, designers use a mobile AR device to record a video prototype to capture context, sketch assets, and demonstrate interactions. They can demonstrate touch inputs, animation paths, and rules to, e.g., have a sketch follow the focus area of the device or the user’s world-space touches. Simultaneously, a live website visualizes an editable overview of all the demonstrated examples and infers a state machine of the user flow. Our key contribution is a method that enables designers to turn a video prototype into an executable state machine through PbD. The designer switches between these representations to interactively refine the final interactive prototype. We illustrate the power of Rapido’s approach by prototyping the main interactions of three popular AR mobile applications.
Germán Leiva, Jens Emil Grønbæk, Clemens Nylandsted Klokmose, Cuong Nguyen 0003, Rubaiat Habib Kazi, Paul Asente
UIST4
2020 Pronto: Rapid Augmented Reality Video Prototyping Using Sketches and Enaction
abstract
Designers have limited tools to prototype AR experiences rapidly. Can lightweight, immediate tools let designers prototype dynamic AR interactions while capturing the nuances of a 3D experience? We interviewed three AR experts and identified several recurring issues in AR design: creating and positioning 3D assets, handling the changing user position, and orchestrating multiple animations. We introduce PROJECT PRONTO, a tablet-based video prototyping system that combines 2D video with 3D manipulation. PRONTO supports four intertwined activities: capturing 3D spatial information alongside a video scenario, positioning and sketching 2D drawings in a 3D world, and enacting animations with physical interactions. An observational study with professional designers shows that participants can use PRONTO to prototype diverse AR experiences. All participants performed two tasks: replicating a sample non-trivial AR experience and prototyping their open-ended designs. All participants completed the replication task and found PRONTO easy to use. Most participants found that PRONTO encourages more exploration of designs than their current practices.
Germán Leiva, Cuong Nguyen 0003, Rubaiat Habib Kazi, Paul Asente
CHI2
2020 View-Dependent Effects for 360° Virtual Reality Video
abstract
"View-dependent effects'' have parameters that change with the user's view and are rendered dynamically at runtime. They can be used to simulate physical phenomena such as exposure adaptation, as well as for dramatic purposes such as vignettes. We present a technique for adding view-dependent effects to 360 degree video, by interpolating spatial keyframes across an equirectangular video to control effect parameters during playback. An in-headset authoring tool is used to configure effect parameters and set keyframe positions. We evaluate the utility of view-dependent effects with expert 360 degree filmmakers and the perception of the effects with a general audience. Results show that experts find view-dependent effects desirable for their creative purposes and that these effects can evoke novel experiences in an audience.
Jeremy Hartmann, Stephen DiVerdi, Cuong Nguyen 0003, Daniel Vogel 0001
UIST3
2020 TransceiVR: Bridging Asymmetrical Communication Between VR Users and External Collaborators
abstract
Virtual Reality (VR) users often need to work with other users, who observe them outside of VR using an external display. Communication between them is difficult; the VR user cannot see the external user's gestures, and the external user cannot see VR scene elements outside of the VR user's view. We carried out formative interviews with experts to understand these asymmetrical interactions and identify their goals and challenges. From this, we identify high-level system design goals to facilitate asymmetrical interactions and a corresponding space of implementation approaches based on the level of programmatic access to a VR application. We present TransceiVR, a system that utilizes VR platform APIs to enable asymmetric communication interfaces for third-party applications without requiring source code access. TransceiVR allows external users to explore the VR scene spatially or temporally, to annotate elements in the VR scene at correct depths, and to discuss via a shared static virtual display. An initial co-located user evaluation with 10 pairs shows that our system makes asymmetric collaborations in VR more effective and successful in terms of task time, error rate, and task load index. An informal evaluation with a remote expert gives additional insight on utility of features for real world tasks.
Balasaravanan Thoravi Kumaravel, Cuong Nguyen 0003, Stephen DiVerdi, Björn Hartmann
UIST2
2020 Slicing-Volume: Hybrid 3D/2D Multi-target Selection Technique for Dense Virtual Environments
abstract
3D selection in dense VR environments (e.g., point clouds) is extremely challenging due to occlusion and imprecise mid-air input modalities (e.g., 3D controllers and hand gestures). In this paper, we propose "Slicing-Volume", a hybrid selection technique that enables simultaneous 3D interaction in mid-air, and a 2D pen-and-tablet metaphor in VR. Inspired by well-known slicing plane techniques in data visualization, our technique consists of a 3D volume that encloses target objects in mid-air, which are then projected to a 2D tablet view for precise selection on a tangible physical surface. While slicing techniques and tablets-in-VR have been previously explored, in this paper, we evaluated the potential of this hybrid approach to improve accuracy in highly occluded selection tasks, comparing different multimodal interactions (e.g., Mid-air, Virtual Tablet and Real Tablet). Our results showed that our hybrid technique significantly improved overall accuracy of selection compared to Mid-air selection only, thanks to the added haptic feedback given by the physical tablet surface, rather than the added visualization given by the tablet view.
Roberto A. Montaño-Murillo, Cuong Nguyen 0003, Rubaiat Habib Kazi, Sriram Subramanian, Stephen DiVerdi, Diego Martínez 0001
VR2
2019 TutoriVR: A Video-Based Tutorial System for Design Applications in Virtual Reality
abstract
Virtual Reality painting is a form of 3D-painting done in a Virtual Reality (VR) space. Being a relatively new kind of art form, there is a growing interest within the creative practices community to learn it. Currently, most users learn using community posted 2D-videos on the internet, which are a screencast recording of the painting process by an instructor. While such an approach may suffice for teaching 2D-software tools, these videos by themselves fail in delivering crucial details that required by the user to understand actions in a VR space. We conduct a formative study to identify challenges faced by users in learning to VR-paint using such video-based tutorials. Informed by results of this study, we develop a VR-embedded tutorial system that supplements video tutorials with 3D and contextual aids directly in the user's VR environment. An exploratory evaluation showed users were positive about the system and were able to use the proposed system to recreate painting tasks in VR.
Balasaravanan Thoravi Kumaravel, Cuong Nguyen 0003, Stephen DiVerdi, Björn Hartmann
CHI2
2019 Challenges and Design Considerations for Multimodal Asynchronous Collaboration in VR
abstract
Studies on collaborative virtual environments (CVEs) have suggested capture and later replay of multimodal interactions (e.g., speech, body language, and scene manipulations), which we refer to as multimodal recordings, as an effective medium for time-distributed collaborators to discuss and review 3D content in an immersive, expressive, and asynchronous way. However, there exist gaps of empirical knowledge in understanding how this multimodal asynchronous VR collaboration (MAVRC) context impacts social behaviors in mediated-communication, workspace awareness in cooperative work, and user requirements for authoring and consuming multimedia recording. This study aims to address these gaps by conceptualizing MAVRC as a type of CSCW and by understanding the challenges and design considerations of MAVRC systems. To this end, we conducted an exploratory need-finding study where participants (N = 15) used an experimental MAVRC system to complete a representative spatial task in an asynchronously collaborative setting, involving both consumption and production of multimodal recordings. Qualitative analysis of interview and observation data from the study revealed unique, core design challenges of MAVRC in: (1) coordinating proxemic behaviors between asynchronous collaborators, (2) providing traceability and change awareness across different versions of 3D scenes, (3) accommodating viewpoint control to maintain workspace awareness, and (4) supporting navigation and editing of multimodal recordings. We discuss design implications, ideate on potential design solutions, and conclude the paper with a set of design recommendations for MAVRC systems.
Kevin Chow, Caitlin Coyiuto, Cuong Nguyen 0003, Dongwook Yoon
Proc. ACM Hum. Comput. Interact.3
2018 Depth Conflict Reduction for Stereo VR Video Interfaces
abstract
Applications for viewing and editing 360° video often render user interface (UI) elements on top of the video. For stereoscopic video, in which the perceived depth varies over the image, the perceived depth of the video can conflict with that of the UI elements, creating discomfort and making it hard to shift focus. To address this problem, we explore two new techniques that adjust the UI rendering based on the video content. The first technique dynamically adjusts the perceived depth of the UI to avoid depth conflict, and the second blurs the video in a halo around the UI. We conduct a user study to assess the effectiveness of these techniques in two stereoscopic VR video tasks: video watching with subtitles, and video search.
Cuong Nguyen 0003, Stephen DiVerdi, Aaron Hertzmann, Feng Liu 0015
CHI1
2017 Vremiere: In-Headset Virtual Reality Video Editing
abstract
Creative professionals are creating Virtual Reality (VR) experiences today by capturing spherical videos, but video editing is still done primarily in traditional 2D desktop GUI applications such as Premiere. These interfaces provide limited capabilities for previewing content in a VR headset or for directly manipulating the spherical video in an intuitive way. As a result, editors must alternate between editing on the desktop and previewing in the headset, which is tedious and interrupts the creative process. We demonstrate an application that enables a user to directly edit spherical video while fully immersed in a VR headset. We first interviewed professional VR filmmakers to understand current practice and derived a suitable workflow for in-headset VR video editing. We then developed a prototype system implementing this new workflow. Our system is built upon a familiar timeline design, but is enhanced with custom widgets to enable intuitive editing of spherical video inside the headset. We conducted an expert review study and found that with our prototype, experts were able to edit videos entirely within the headset. Experts also found our interface and widgets useful, providing intuitive controls for their editing needs.
Cuong Nguyen 0003, Stephen DiVerdi, Aaron Hertzmann, Feng Liu 0015
CHI1
2017 CollaVR: Collaborative In-Headset Review for VR Video
abstract
Collaborative review and feedback is an important part of conventional filmmaking and now Virtual Reality (VR) video production as well. However, conventional collaborative review practices do not easily translate to VR video because VR video is normally viewed in a headset, which makes it difficult to align gaze, share context, and take notes. This paper presents CollaVR, an application that enables multiple users to review a VR video together while wearing headsets. We interviewed VR video professionals to distill key considerations in reviewing VR video. Based on these insights, we developed a set of networked tools that enable filmmakers to collaborate and review video in real-time. We conducted a preliminary expert study to solicit feedback from VR video professionals about our system and assess their usage of the system with and without collaboration features.
Cuong Nguyen 0003, Stephen DiVerdi, Aaron Hertzmann, Feng Liu 0015
UIST1
2016 Gaze-based Notetaking for Learning from Lecture Videos
abstract
Taking notes has been shown helpful for learning. This activity, however, is not well supported when learning from watching lecture videos. The conventional video interface does not allow users to quickly locate and annotate important content in the video as notes. Moreover, users sometimes need to manually pause the video while taking notes, which is often distracting. In this paper, we develop a gaze-based system to assist a user in notetaking while watching lecture videos. Our system has two features to support notetaking. First, our system integrates offline video analysis and online gaze analysis to automatically detect and highlight key content from the lecture video for notetaking. Second, our system provides adaptive video control that automatically reduces the video playback speed or pauses it while a user is taking notes to minimize the user's effort in controlling video. Our study shows that our system enables users to take notes more easily and with better quality than the traditional video interface.
Cuong Nguyen 0003, Feng Liu 0015
CHI1
2015 Making Software Tutorial Video Responsive
abstract
Tutorial videos are widely available to help people use software. These videos, however, are viewed by users as captured and offer little direct interaction between users and software. This paper presents a video navigation method that allows users to interact with software tutorial video as if they were using the software. To make the tutorial video responsive, our method records the user interaction events like mouse click and drag during capturing the video. Our method then analyzes, selects, and visualizes these user interaction events at the event locations. When a user directly interacts with an event visualization, our method automatically navigates to the proper video frame to provide the visual feedback as if the software were responding to the user input. Thus, our method provides the experience of interacting with the software through directly manipulating the tutorial video. Our study shows our method can better help users follow tutorial videos to complete tasks than the baseline timeline interface.
Cuong Nguyen 0003, Feng Liu 0015
CHI1
2014 Direct manipulation video navigation on touch screens
abstract
Direct Manipulation Video Navigation (DMVN) systems allow a user to directly drag an object of interest along its motion trajectory and have been shown effective for space-centric video browsing tasks. This paper designs touch-based interface techniques to support DMVN on touchscreen devices. While touch screens can suit DMVN systems naturally and enhance the directness during video navigation, the fat finger problems, such as precise selection and occlusion handling, must be properly addressed. In this paper, we discuss the effect of the fat finger problems on DMVN and develop three touch-based object dragging techniques for DMVN on touch screens, namely Offset Drag, Window Drag, and Drag Anywhere. We conduct user studies to evaluate our techniques as well as two baseline solutions on a smartphone and a desktop touch screen. Our studies show that two of our techniques can support DMVN on touch screen devices well and perform better than the baseline solutions.
Cuong Nguyen 0003, Yuzhen Niu, Feng Liu 0015
Mobile HCI1
2013 Direct manipulation video navigation in 3D
abstract
Direct Manipulation Video Navigation (DMVN) systems allow a user to navigate a video by dragging an object along its motion trajectory. These systems have been shown effective for space-centric video browsing. Their performance, however, is often limited by temporal ambiguities in a video with complex motion, such as recurring motion, self-intersecting motion, and pauses. The ambiguities come from reducing the 3D spatial-temporal motion (x, y, t) to the 2D spatial motion (x, y) in visualizing the motion and dragging the object. In this paper, we present a 3D DMVN system that maps the spatial-temporal motion (x, y, t) to 3D space (x, y, z) by mapping time t to depth z, visualizes the motion and video frame in 3D, and allows to navigate the video by spatial-temporally manipulating the object in 3D. We show that since our 3D DMVN system preserves all the motion information, it resolves the temporal ambiguities and supports intuitive navigation on challenging videos with complex motion.
Cuong Nguyen 0003, Yuzhen Niu, Feng Liu 0015
CHI1
2012 Video summagator: an interface for video summarization and navigation
abstract
This paper presents Video Summagator (VS), a volume-based interface for video summarization and navigation. VS models a video as a space-time cube and visualizes the video cube using real-time volume rendering techniques. VS empowers a user to interactively manipulate the video cube. We show that VS can quickly summarize both the static and dynamic video content by visualizing the space-time information in 3D. We demonstrate that VS enables a user to quickly look into the video cube, understand the content, and navigate to the content of interest.
Cuong Nguyen 0003, Yuzhen Niu, Feng Liu 0015
CHI1