Karthik Ramani

dblp:01/6965 · DBLP profile ↗
← Back
147ranked-venue papers
3as first author
45since 2021 · last 2026
0000-0001-8639-5135ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 71 · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 53 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 26 · 10 since 2021Systems, architecture and hardware · 8 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7Databases, data management, data science and information retrieval · 5Software engineering, systems software and programming languages · 3 · 1 first-author
YearPublicationVenuePosition
2026 SketchConcept: Sketching-based Concept Composition for Product Design using Multimodal Large Language Model
abstract
Sketches are widely used in conceptual design to externalize early ideas and communicate intent. With the rise of generative AI, sketch-to-design workflows have advanced rapidly. However, sketches are limited for organizing component-level structure and intent: parts, functions, and relations are often implicit, making systematic design space exploration difficult. We present SketchConcept, a sketch-to-design system that enables multimodal exploration through sketching and language. It allows designers to sketch out the form, then use voice or text to articulate and refine component functions and structural organization. This enables designers to explore not only satisfying appearances, but also functional and structural alternatives that are essential for design. To support this workflow, SketchConcept introduces a function-to-visual mapping mechanism that connects visual components to functional properties for component-wise iteration. We demonstrate the system through a set of representative use cases and evaluate its efficacy and usability in a two-session user study.
Runlin Duan, Chenfei Zhu, Yuzhao Chen, Dizhi Ma, Jingyu Shi, Yichen Hu, Ziyi Liu 0004, Karthik Ramani
DIS8
2026 AssembleIt: Generating Adaptive On-Demand 3D Animations for Context-Aware Mechanical Assembly Guidance
abstract
Mechanical assembly instructions are commonly delivered through static manuals or fixed-sequence animations, which limit users’ ability to seek clarification, request partial explanations, or adapt guidance to their moment-to-moment needs during physical assembly. We present AssembleIt, an interactive system that generates on-demand 3D assembly animations and verbal explanations directly from natural language user queries, without relying on pre-authored instructional content. AssembleIt automatically derives a part dependency graph from CAD geometry using an Assembly-by-Disassembly strategy and uses this representation to generate query-driven, context-aware animations at runtime, rather than following a single predefined sequence. We evaluate AssembleIt through a controlled user study with 12 participants performing physical assembly tasks using real parts, comparing on-demand, query-driven animations against a static 3D animation baseline. Results indicate that on-demand animation generation supports flexible exploration and targeted clarification during assembly, highlighting the design potential of generative, dependency-driven instructional interfaces for hands-on mechanical tasks.
Mayank Patel 0005, Rahul Jain 0018, Asim Unmesh, Shao-Kang Hsia, Karthik Ramani
DIS5
2026 JustShape: Exploring Co-Speech Gestures for Multimodal LLM-Powered 3D Parametric Modeling
Runlin Duan, Yuzhao Chen, Yichen Hu, Ziyi Liu 0004, Chenfei Zhu, Xiyun Hu, Dizhi Ma, Karthik Ramani
CHI9
2026 ARify: Leveraging Narrated Instructional Videos to Create Augmented Reality Tutorials for Procedural Tasks
abstract
Augmented Reality (AR) tutorials enhance procedural task learning by providing situated, step-by-step guidance. Yet, creating such tutorials requires AR authoring expertise, posing a significant entry barrier. To lower this barrier, we introduce ARify, an authoring system that semi-automatically transforms narrated instructional videos into AR tutorials. To guide system design, we conducted a content analysis of video tutorials and derived a design space of instructional intents, tactics, and AR representations. Building on this, ARify generates AR tutorials by integrating a vision–language model to plan tutorial structures and an AR builder to configure AR representations, and offers interfaces that allow users to refine and customize the results. A numerical study on three machine tasks and a user study with 18 participants showed that ARify achieves promising performance across task types, and allows novices to author effective AR tutorials, validating its effectiveness and usability.
Xiyun Hu, Chenfei Zhu, Shao-Kang Hsia, Dizhi Ma, Rahul Jain 0018, Karthik Ramani
CHI6
2026 AmIWrite: Exploring Scalable One-on-One Handwriting-Based Tutoring for Mathematical Problem-Solving with an LLM-Powered AI Tutor
abstract
Real-time handwriting interactions between tutors and students —where tutors observe individual problem-solving processes, provide personalized annotations, and adapt explanations based on students’ work—are fundamental to effective STEM tutoring. However, scaling such personalized handwriting-based tutoring remains challenging—human tutors cannot be available to every student on demand, and current online platforms often fail to recreate equivalent learning experiences. As an initial step toward tackling this challenge, we present AmIWrite, an LLM-powered AI tutoring system for mathematical problem-solving that provides real-time co-speech handwriting interactions on tablet devices, instantiated here as a case study in linear algebra. We conducted a within-subjects study (N = 40) comparing AmIWrite to a text-based AI tutor on two linear algebra topics. Our case study demonstrates how a multimodal AI tutor can preserve the pedagogical benefits of handwriting-based math tutoring and offer a potential path toward more scalable one-on-one STEM tutoring.
Ziyi Liu 0004, Yuzhao Chen, Runlin Duan, Zhengzhe Zhu, Xiyun Hu, Kylie Peppler, Karthik Ramani
CHI8
2026 AgentCoach: LLM-Based Adaptive Coaching Feedback for Motor Skill Learning
abstract
We present AgentCoach, an LLM-powered system that provides adaptive feedback for motor skill learning from tutorial videos. The system works by extracting key coaching points (CPs) and compiling CP-specific evaluators that map each cue to measurable kinematic parameters. This process allows AgentCoach to connect high-level semantic meaning with low-level postural estimation for accurate, context-aware evaluation. During practice, learners receive concise visual diagnostics of their mistakes paired with prescriptive verbal feedback that adapts based on their performance history. We technically validate the CP extraction and evaluator compilation across a wide range of common sports and exercise videos. A user study confirms the system’s usability and shows the system’s potential effectiveness of its adaptive feedback across multiple skills.
Dizhi Ma, Jiakun Yu, Xiyun Hu, Liang He 0005, Sooyeon Jeong, Karthik Ramani
CHI7
2026 Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
abstract
Generative AI (GenAI) has significantly advanced the ease and flexibility of image creation. However, it remains a challenge to precisely control spatial compositions, including object arrangement and scene conditions. To bridge this gap, we propose Canvas3D, an interactive system leveraging a 3D engine to enable precise spatial manipulation for image generation. Upon user prompt, Canvas3D automatically converts textual descriptions into interactive objects within a 3D engine-driven virtual canvas, empowering direct and precise spatial configuration. These user-defined arrangements generate explicit spatial constraints that guide generative models in accurately reflecting user intentions in the resulting images. We conducted a closed-ended comparative study between Canvas3D and a baseline system, and an open-ended, free-form study to assess overall system usability. The results indicate that Canvas3D outperforms the baseline on spatial control, interactivity, and overall user experience.
Yuzhao Chen, Runlin Duan, Rahul Jain 0018, Yichen Hu, Chenfei Zhu, Jingyu Shi, Karthik Ramani
IUI7
2025 DesignFromX: Empowering Consumer-Driven Design Space Exploration through Feature Composition of Referenced Products
abstract
Generated Design SpaceIteration-1 Iteration-2 Iteration-3Figure 1: Exploring the design space of a desk using DesignFromX.The process begins with the user selecting a component from a reference product image-here, the legs of a wooden chair.The system identifies and suggests design features of the selected component.The user then composes this feature, in this case, the structural form, into a designated part of the desk.Based on the user-defined composition, a Generative AI model generates a new design space for the desk.In subsequent iterations, the user can incorporate additional design features from other reference products to further explore the design space while retaining the features of their previous selections.
Runlin Duan, Chenfei Zhu, Yuzhao Chen, Yichen Hu, Jingyu Shi, Karthik Ramani
Conference on Designing Interactive Systems6
2025 GesPrompt: Leveraging Co-Speech Gestures to Augment LLM-Based Interaction in Virtual Reality
abstract
Large Language Model (LLM)-based copilots have shown great potential in Extended Reality (XR) applications.However, the user faces challenges when describing the 3D environments to the copilots due to the complexity of conveying spatial-temporal information through text or speech alone.To address this, we introduce GesPrompt, a multimodal XR interface that combines co-speech gestures with speech, allowing end-users to communicate more naturally and accurately with LLM-based copilots in XR environments.By incorporating gestures, GesPrompt extracts spatial-temporal reference from co-speech gestures, reducing the need for precise textual prompts and minimizing cognitive load for end-users.Our contributions include (1) a workflow to integrate gesture and speech input in the XR environment, (2) a prototype VR system that implements the workflow, and (3) a user study demonstrating its effectiveness in improving user communication in VR environments.
Xiyun Hu, Dizhi Ma, Fengming He, Zhengzhe Zhu, Shao-Kang Hsia, Chenfei Zhu, Ziyi Liu 0004, Karthik Ramani
Conference on Designing Interactive Systems8
2025 CARING-AI: Towards Authoring Context-aware Augmented Reality INstruction through Generative Artificial Intelligence
Jingyu Shi, Rahul Jain 0018, Seunggeun Chi, Hyungjun Doh, Hyung-Gun Chi, Alexander J. Quinn, Karthik Ramani
CHI7
2025 Transparent Barriers: Natural Language Access Control Policies for XR-Enhanced Everyday Objects
Kentaro Taninaka, Rahul Jain 0018, Jingyu Shi, Kazunori Takashio, Karthik Ramani
CHI5
2025 CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from a Single-View Image
abstract
Recently, generalizable feed-forward methods based on 3D Gaussian Splatting have gained significant attention for their potential to reconstruct 3D scenes using finite resources. These approaches create a 3D radiance field, parameterized by per-pixel 3D Gaussian primitives, from just a few images in a single forward pass. However, unlike multi-view methods that benefit from cross-view correspondences, 3D scene reconstruction with a single-view image remains an underexplored area. In this work, we introduce CATSplat, a novel generalizable transformer-based framework designed to break through the inherent constraints in monocular settings. First, we propose leveraging textual guidance from a visual-language model to complement insufficient information from a single image. By incorporating scene-specific contextual details from text embeddings through cross-attention, we pave the way for context-aware 3D scene reconstruction beyond relying solely on visual cues. Moreover, we advocate utilizing spatial guidance from 3D point features toward comprehensive geometric understanding under single-view settings. With 3D priors, image features can capture rich structural insights for predicting 3D Gaussians without multi-view techniques. Extensive experiments on large-scale datasets demonstrate the state-of-the-art performance of CATSplat in single-view 3D scene reconstruction with high-quality novel view synthesis.
Wonseok Roh, Hwanhee Jung, Jong Wook Kim, Seunggwan Lee, Innfarn Yoo, Andreas Lugmayr, Seunggeun Chi, Karthik Ramani, Sangpil Kim
ICCV8
2025 Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction
abstract
We introduce a novel framework for reconstructing dynamic human-object interactions from monocular video that overcomes challenges associated with occlusions and temporal inconsistencies. Traditional 3D reconstruction methods typically assume static objects or full visibility of dynamic subjects, leading to degraded performance when these assumptions are violated-particularly in scenarios where mutual occlusions occur. To address this, our framework leverages amodal completion to infer the complete structure of partially obscured regions. Unlike conventional approaches that operate on individual frames, our method integrates temporal context, enforcing coherence across video sequences to incrementally refine and stabilize reconstructions. This template-free strategy adapts to varying conditions without relying on predefined models, significantly enhancing the recovery of intricate details in dynamic scenes. We validate our approach using 3D Gaussian Splatting on challenging monocular videos, demonstrating superior precision in handling occlusions and maintaining temporal stability compared to existing techniques.
Hyungjun Doh, Dong In Lee, Seunggeun Chi, Pin-Hao Huang, Kwonjoon Lee, Sangpil Kim, Karthik Ramani
ACM Multimedia7
2025 agentAR: Creating Augmented Reality Applications with Tool-Augmented LLM-based Autonomous Agents
Chenfei Zhu, Shao-Kang Hsia, Xiyun Hu, Ziyi Liu 0004, Jingyu Shi, Karthik Ramani
UIST6
2025 InfoGCN++: Learning Representation by Predicting the Future for Online Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition has made significant advancements recently, with models like InfoGCN showcasing remarkable accuracy. However, these models exhibit a key limitation: they necessitate complete action observation prior to classification, which constrains their applicability in real-time situations such as surveillance and robotic systems. To overcome this barrier, we introduce InfoGCN++, an innovative extension of InfoGCN, explicitly developed for online skeleton-based action recognition. InfoGCN++ augments the abilities of the original InfoGCN model by allowing real-time categorization of action types, independent of the observation sequence's length. It transcends conventional approaches by learning from current and anticipated future movements, thereby creating a more thorough representation of the entire sequence. Our approach to prediction is managed as an extrapolation issue, grounded on observed actions. To enable this, InfoGCN++ incorporates Neural Ordinary Differential Equations, a concept that lets it effectively model the continuous evolution of hidden states. Following rigorous evaluations on three skeleton-based action recognition benchmarks, InfoGCN++ demonstrates exceptional performance in online action recognition. It consistently equals or exceeds existing techniques, highlighting its significant potential to reshape the landscape of real-time action recognition applications. Consequently, this work represents a major leap forward from InfoGCN, pushing the limits of what's possible in online, skeleton-based action recognition.
Seunggeun Chi, Hyung-Gun Chi, Qixing Huang, Karthik Ramani
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Visualizing Causality in Mixed Reality for Manual Task Learning: A Study
abstract
Mixed Reality (MR) is gaining prominence in manual task skill learning due to its in-situ, embodied, and immersive experience. To teach manual tasks, current methodologies break the task into hierarchies (tasks into subtasks) and visualize not only the current subtasks but also the future ones that are causally related. We investigate the impact of visualizing causality within an MR framework on manual task skill learning. We conducted a user study with 48 participants, experimenting with how presenting tasks in hierarchical causality levels (no causality, event-level, interaction-level, and gesture-level causality) affects user comprehension and performance in a complex assembly task. The research finds that displaying all causality levels enhances user understanding and task execution, with a compromise of learning time. Based on the results, we further provide design recommendations and in-depth discussions for future manual task learning systems.
Rahul Jain 0018, Jingyu Shi, Andrew Benton, Moiz Rasheed, Hyungjun Doh, Subramanian Chidambaram, Karthik Ramani
IEEE Trans. Vis. Comput. Graph.7
2024 ClassMeta: Designing Interactive Virtual Classmate to Promote VR Classroom Participation
abstract
Peer influence plays a crucial role in promoting classroom participation, where behaviors from active students can contribute to a collective classroom learning experience. However, the presence of these active students depends on several conditions and is not consistently available across all circumstances. Recently, Large Language Models (LLMs) such as GPT have demonstrated the ability to simulate diverse human behaviors convincingly due to their capacity to generate contextually coherent responses based on their role settings. Inspired by this advancement in technology, we designed ClassMeta, a GPT-4 powered agent to help promote classroom participation by playing the role of an active student. These agents, which are embodied as 3D avatars in virtual reality, interact with actual instructors and students with both spoken language and body gestures. We conducted a comparative study to investigate the potential of ClassMeta for improving the overall learning experience of the class.
Ziyi Liu 0004, Zhengzhe Zhu, Enze Jiang, Xiyun Hu, Kylie Peppler, Karthik Ramani
CHI7
2024 ChatDirector: Enhancing Video Conferencing with Space-Aware Scene Rendering and Speech-Driven Layout Transition
abstract
Remote video conferencing systems (RVCS) are widely adopted in personal and professional communication. However, they often lack the co-presence experience of in-person meetings. This is largely due to the absence of intuitive visual cues and clear spatial relationships among remote participants, which can lead to speech interruptions and loss of attention. This paper presents ChatDirector, a novel RVCS that overcomes these limitations by incorporating space-aware visual presence and speech-aware attention transition assistance. ChatDirector employs a real-time pipeline that converts participants’ RGB video streams into 3D portrait avatars and renders them in a virtual 3D scene. We also contribute a decision tree algorithm that directs the avatar layouts and behaviors based on participants’ speech states. We report on results from a user study (N=16) where we evaluated ChatDirector. The satisfactory algorithm performance and complimentary subject user feedback imply that ChatDirector significantly enhances communication efficacy and user engagement.
Xun Qian, Feitong Tan, Yinda Zhang 0001, Brian Moreno Collins, David Kim 0002, Alex Olwal, Karthik Ramani, Ruofei Du
CHI7
2024 Higher-order Relational Reasoning for Pedestrian Trajectory Prediction
abstract
Social relations have substantial impacts on the potential trajectories of each individual. Modeling these dynamics has been a central solution for more precise and accurate trajectory forecasting. However, previous works ignore the importance of ‘social depth’, meaning the influences flowing from different degrees of social relations. In this work, we propose HighGraph, a graph-based pedestrian relational reasoning method that captures the higherorder dynamics of social interactions. First, we construct a collision-aware relation graph based on the agents' observed trajectories. Upon this graph structure, we build our core module that aggregates the agent features from diverse social distances. As a result, the network is able to model complex social relations, thereby yielding more accurate and socially acceptable trajectories. Our High-Graph is a plug-and-play module that can be easily applied to any current trajectory predictors. Extensive experiments with ETH/UCY and SDD datasets demonstrate that our HighGraph noticeably improves the previous state-of-the-art baselines both quantitatively and qualitatively.
Sungjune Kim, Hyung-Gun Chi, Hyerin Lim, Karthik Ramani, Jinkyu Kim 0001, Sangpil Kim
CVPR4
2024 M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models
Seunggeun Chi, Hyung-Gun Chi, Hengbo Ma, Nakul Agarwal, Faizan Siddiqui 0001, Karthik Ramani, Kwonjoon Lee
ECCV (14)6
2024 Multi-Modal Representation Learning with Tactile Data
abstract
Advancements in embodied language models like PALM-E and RT-2 have significantly enhanced language-conditioned robotic manipulation. However, these advances remain predominantly focused on vision and language, often overlooking the pivotal role of tactile feedback which is advantageous in contact-rich interactions. Our research introduces a novel approach that synergizes tactile information with vision and language. We present the Multi-Modal Wand (MMWand) dataset enriched with linguistic descriptions and tactile data. By integrating tactile feedback, we aim to bridge the divide between human linguistic understanding and robotic sensory interpretation. Our multi-modal representation model is trained on these datasets by employing the multi-modal embedding alignment principle from ImageBind which has shown promising results, emphasizing the potential of tactile data in robotic applications. The validation of our approach in downstream robotics tasks, such as texture-based object classification, cross-modality retrieval, and the dense reward function for visuomotor control, attests to its effectiveness. Our contributions underscore the importance of tactile feedback in multi-modal robotic learning and its potential to enhance robotic tasks. The MMWand dataset is publicly available at https://hyung-gun.me/mmwand/.
Hyung-Gun Chi, Jose A. Barreiros, Jean Mercat, Karthik Ramani, Thomas Kollar
IROS4
2024 Estimating Ego-Body Pose from Doubly Sparse Egocentric Video Data
abstract
We study the problem of estimating the body movements of a camera wearer from egocentric videos. Current methods for ego-body pose estimation rely on temporally dense sensor data, such as IMU measurements from spatially sparse body parts like the head and hands. However, we propose that even temporally sparse observations, such as hand poses captured intermittently from egocentric videos during natural or periodic hand movements, can effectively constrain overall body motion. Naively applying diffusion models to generate full-body pose from head pose and sparse hand pose leads to suboptimal results. To overcome this, we develop a two-stage approach that decomposes the problem into temporal completion and spatial completion. First, our method employs masked autoencoders to impute hand trajectories by leveraging the spatiotemporal correlations between the head pose sequence and intermittent hand poses, providing uncertainty estimates. Subsequently, we employ conditional diffusion models to generate plausible full-body motions based on these temporally dense trajectories of the head and hands, guided by the uncertainty estimates from the imputation. The effectiveness of our methods was rigorously tested and validated through comprehensive experiments conducted on various HMD setup with AMASS and Ego-Exo4D datasets. Project page: https://sgchi.github.io/dsposer
Seunggeun Chi, Pin-Hao Huang, Enna Sachdeva, Hengbo Ma, Karthik Ramani, Kwonjoon Lee
NeurIPS5
2024 avaTTAR: Table Tennis Stroke Training with Embodied and Detached Visualization in Augmented Reality
abstract
Table tennis stroke training is a critical aspect of player development. We designed a new augmented reality (AR) system, avaTTAR, for table tennis stroke training. The system provides both “on-body” (first-person view) and “detached” (third-person view) visual cues, enabling users to visualize target strokes and correct their attempts effectively with this dual perspectives setup. By employing a combination of pose estimation algorithms and IMU sensors, avaTTAR captures and reconstructs the 3D body pose and paddle orientation of users during practice, allowing real-time comparison with expert strokes. Through a user study, we affirm avaTTAR ’s capacity to amplify player experience and training results.
Dizhi Ma, Xiyun Hu, Jingyu Shi, Mayank Patel 0005, Rahul Jain 0018, Ziyi Liu 0004, Zhengzhe Zhu, Karthik Ramani
UIST8
2024 AdapTUI: Adaptation of Geometric-Feature-Based Tangible User Interfaces in Augmented Reality
abstract
With the advents in geometry perception and Augmented Reality (AR), end-users can customize Tangible User Interfaces (TUIs) that control digital assets using intuitive and comfortable interactions with physical geometries (e.g., edges and surfaces). However, it remains challenging to adapt such TUIs in varied physical environments while maintaining the same spatial and ergonomic affordance. We propose AdapTUI, an end-to- end system that enables an end-user to author geometric-based TUIs and automatically adapts the TUIs when the user moves to a new environment. Leveraging a geometry detection module and the spatial awareness of AR, AdapTUI first lets users create custom mappings between geometric features and digital functions. Then, AdapTUI uses an optimization-based adaptation framework, which considers both the geometric variations and human-factor nuances, to dynamically adjust the attachment of the user-authored TUIs. We demonstrate three application scenarios where end-users can utilize TUIs at different locations, including portable car play, efficient AR workstation, and entertainment. We evaluated the effectiveness of the adaptation method as well as the overall usability through a comparison user study (N=12). The satisfactory adaptation of the user-authored TUIs and the positive qualitative feedback demonstrate the effectiveness of our system.
Fengming He, Xiyun Hu, Xun Qian, Zhengzhe Zhu, Karthik Ramani
Proc. ACM Hum. Comput. Interact.5
2023 Ubi Edge: Authoring Edge-Based Opportunistic Tangible User Interfaces in Augmented Reality
abstract
Edges are one of the most ubiquitous geometric features of physical objects. They provide accurate haptic feedback and easy-to-track features for camera systems, making them an ideal basis for Tangible User Interfaces (TUI) in Augmented Reality (AR). We introduce Ubi Edge, an AR authoring tool that allows end-users to customize edges on daily objects as TUI inputs to control varied digital functions. We develop an integrated AR-device and an integrated vision-based detection pipeline that can track 3D edges and detect the touch interaction between fingers and edges. Leveraging the spatial-awareness of AR, users can simply select an edge by sliding fingers along it and then make the edge interactive by connecting it to various digital functions. We demonstrate four use cases including multi-function controllers, smart homes, games, and TUI-based tutorials. We also evaluated and proved our system’s usability through a two-session user study, where qualitative and quantitative results are positive.
Fengming He, Xiyun Hu, Jingyu Shi, Xun Qian, Tianyi Wang 0004, Karthik Ramani
CHI6
2023 InstruMentAR: Auto-Generation of Augmented Reality Tutorials for Operating Digital Instruments Through Recording Embodied Demonstration
abstract
Augmented Reality tutorials, which provide necessary context by directly superimposing visual guidance on the physical referent, represent an effective way of scaffolding complex instrument operations. However, current AR tutorial authoring processes are not seamless as they require users to continuously alternate between operating instruments and interacting with virtual elements. We present InstruMentAR, a system that automatically generates AR tutorials through recording user demonstrations. We design a multimodal approach that fuses gestural information and hand-worn pressure sensor data to detect and register the user’s step-by-step manipulations on the control panel. With this information, the system autonomously generates virtual cues with designated scales to respective locations for each step. Voice recognition and background capture are employed to automate the creation of text and images as AR content. For novice users receiving the authored AR tutorials, we facilitate immediate feedback through haptic modules. We compared InstruMentAR with traditional systems in the user study.
Ziyi Liu 0004, Zhengzhe Zhu, Enze Jiang, Feichi Huang, Ana M. Villanueva, Xun Qian, Tianyi Wang 0004, Karthik Ramani
CHI8
2023 LearnIoTVR: An End-to-End Virtual Reality Environment Providing Authentic Learning Experiences for Internet of Things
abstract
The rapid growth of Internet-of-Things (IoT) applications has generated interest from many industries and a need for graduates with relevant knowledge. An IoT system is comprised of spatially distributed interactions between humans and various interconnected IoT components. These interactions are contextualized within their ambient environment, thus impeding educators from recreating authentic tasks for hands-on IoT learning. We propose LearnIoTVR, an end-to-end virtual reality (VR) learning environment which helps students to acquire IoT knowledge through immersive design, programming, and exploration of real-world environments empowered by IoT (e.g., a smart house). The students start the learning process by installing virtual IoT components we created in different locations inside the VR environment so that the learning will be situated in the same context where the IoT is applied. With our custom-designed 3D block-based language, students can program IoT behaviors directly within VR and get immediate feedback on their programming outcome. In the user study, we evaluated the learning outcomes among students using LearnIoTVR with a pre- and post-test to understand to what extent does engagement in LearnIoTVR lead to gains in learning programming skills and IoT competencies. Additionally, we examined what aspects of LearnIoTVR support usability and learning of programming skills compared to a traditional desktop-based learning environment. The results from these studies were promising. We also acquired insightful user feedback which provides inspiration for further expansions of this system.
Zhengzhe Zhu, Ziyi Liu 0004, Youyou Zhang, Joey Huang, Ana M. Villanueva, Xun Qian, Kylie Peppler, Karthik Ramani
CHI9
2023 AdamsFormer for Spatial Action Localization in the Future
abstract
Predicting future action locations is vital for applications like human-robot collaboration. While some computer vision tasks have made progress in predicting human actions, accurately localizing these actions in future frames remains an area with room for improvement. We introduce a new task called spatial action localization in the future (SALF), which aims to predict action locations in both observed and future frames. SALF is challenging because it requires understanding the underlying physics of video observations to predict future action locations accurately. To address SALF, we use the concept of NeuralODE, which models the latent dynamics of sequential data by solving ordinary differential equations (ODE) with neural networks. We propose a novel architecture, AdamsFormer, which extends observed frame features to future time horizons by modeling continuous temporal dynamics through ODE solving. Specifically, we employ the Adams method, a multi-step approach that efficiently uses information from previous steps without discarding it. Our extensive experiments on UCF101-24 and JHMDB-21 datasets demonstrate that our proposed model outperforms existing long-range temporal modeling methods by a significant margin in terms of frame-mAP.
Hyung-Gun Chi, Kwonjoon Lee, Nakul Agarwal, Yi Xu 0005, Karthik Ramani, Chiho Choi
CVPR5
2023 Pose Relation Transformer Refine Occlusions for Human Pose Estimation
abstract
Accurately estimating the human pose is an essential task for many applications in robotics. However, existing pose estimation methods suffer from poor performance when occlusion occurs. Recent advances in NLP have been very successful in predicting the missing words conditioned on visible words. We draw upon the sentence completion analogy in NLP to guide our model to address occlusions in the pose estimation problem. We propose a novel approach that can mitigate the effect of occlusions motivated by the sentence completion task of NLP. In an analogous manner, we designed our model to reconstruct occluded joints given the visible joints utilizing joint correlations by capturing the implicit joint connectivity through the attention mechanism. In this work, we propose a POse Relation Transformer (PORT) that captures the global context of the pose using self-attention and a local context by aggregating adjacent joint features. To supervise PORT in learning joint correlations, we guide PORT to reconstruct randomly masked joints, which we call Masked Joint Modeling (MJM). PORT trained with MJM adds to existing keypoint detection methods and successfully refines occlusions. Notably, PORT is a model-agnostic plug-and-play module for pose refinement under occlusion that can be plugged into any keypoint detector with substantially low computational costs. We conducted extensive experiments to demonstrate the advantage of PORT mitigating the occlusion on the hand and body pose PORT improves the pose estimation accuracy of existing human pose estimation methods by up to 16% with only 5% of additional parameters. The code is publicly available at https://github.com/stnoah1/PORT.
Hyung-Gun Chi, Seunggeun Chi, Stanley Chan, Karthik Ramani
ICRA4
2023 AircraftVerse: A Large-Scale Multimodal Dataset of Aerial Vehicle Designs
abstract
We present AircraftVerse, a publicly available aerial vehicle design dataset. Aircraft design encompasses different physics domains and, hence, multiple modalities of representation. The evaluation of these designs requires the use of scientific analytical and simulation models ranging from computer-aided design tools for structural and manufacturing analysis, computational fluid dynamics tools for drag and lift computation, battery models for energy estimation, and simulation models for flight control and dynamics. AircraftVerse contains $27{,}714$ diverse air vehicle designs - the largest corpus of designs with this level of complexity. Each design comprises the following artifacts: a symbolic design tree describing topology, propulsion subsystem, battery subsystem, and other design details; a STandard for the Exchange of Product (STEP) model data; a 3D CAD design using a stereolithography (STL) file format; a 3D point cloud for the shape of the design; and evaluation results from high fidelity state-of-the-art physics models that characterize performance metrics such as maximum flight distance and hover-time. We also present baseline surrogate models that use different modalities of design representation to predict design performance metrics, which we provide as part of our dataset release. Finally, we discuss the potential impact of this dataset on the use of learning in aircraft design, and more generally, in the emerging field of deep learning for scientific design. AircraftVerse is accompanied by a datasheet as suggested in the recent literature, and it is released under Creative Commons Attribution-ShareAlike (CC BY-SA) license. The dataset with baseline models are hosted at http://doi.org/10.5281/zenodo.6525446, code at https://github.com/SRI-CSL/AircraftVerse, and the dataset description at https://uavdesignverse.onrender.com/.
Adam D. Cobb, Daniel Elenius, F. Michael Heim, Brian Swenson, Sydney Whittington, James D. Walker, Ted Bapty, Joseph Hite, Karthik Ramani, Christopher McComb, Susmit Jha
NeurIPS10
2023 Ubi-TOUCH: Ubiquitous Tangible Object Utilization through Consistent Hand-object interaction in Augmented Reality
abstract
Utilizing everyday objects as tangible proxies for Augmented Reality (AR) provides users with haptic feedback while interacting with virtual objects. Yet, existing methods focus on the attributes of the objects, constraining the possible proxies and yielding inconsistency in user experience. Therefore, we propose Ubi-TOUCH, an AR system that assists users in seeking a wider range of tangible proxies for AR applications based on the hand-object interaction (HOI) they desire. Given the target interaction with a virtual object, the system scans the users’ vicinity and recommends object proxies with similar interactions. Upon user selection, the system simultaneously tracks and maps users’ physical HOI to the virtual HOI, adaptively optimizing object 6 DoF and the hand gesture to provide consistency between the interactions. We showcase promising use cases of Ubi-TOUCH, such as remote tutorials, AR gaming, and Smart Home control. Finally, we evaluate the performance and usability of Ubi-TOUCH with a user study.
Rahul Jain 0018, Jingyu Shi, Runlin Duan, Zhengzhe Zhu, Xun Qian, Karthik Ramani
UIST6
2023 Simplification of 3D CAD Model in Voxel Form for Mechanical Parts Using Generative Adversarial Networks
Hyunoh Lee, Soonjo Kwon, Karthik Ramani, Hyung-Gun Chi, Duhwan Mun
Comput. Aided Des.4
2023 Advanced modeling method for quantifying cumulative subjective fatigue in mid-air interaction
Ana M. Villanueva, Sujin Jang, Wolfgang Stuerzlinger, Satyajit Ambike, Karthik Ramani
Int. J. Hum. Comput. Stud.5
2022 Towards Modeling of Virtual Reality Welding Simulators to Promote Accessible and Scalable Training
abstract
The US manufacturing industry is currently facing a welding workforce shortage which is largely due to inadequacy of widespread welding training. To address this challenge, we present a Virtual Reality (VR)-based training system aimed at transforming state-of-the-art-welding simulations and in-person instruction into a widely accessible and engaging platform. We applied backward design principles to design a low-cost welding simulator in the form of modularized units through active consulting with welding training experts. Using a minimum viable prototype, we conducted a user study with 24 novices to test the system’s usability. Our findings show (1) greater effectiveness of the system in transferring skills to real-world environments as compared to accessible video-based alternatives and, (2) the visuo-haptic guidance during virtual welding enhances performance and provides a realistic learning experience to users. Using the solution, we expect inexperienced users to achieve competencies faster and be better prepared to enter actual work environments.
Ananya Ipsita, Levi Erickson, Yangzi Dong, Joey Huang, Alexa Bushinski, Sraven Saradhi, Ana M. Villanueva, Kylie Peppler, Thomas Redick, Karthik Ramani
CHI10
2022 ScalAR: Authoring Semantically Adaptive Augmented Reality Experiences in Virtual Reality
abstract
Augmented Reality (AR) experiences tightly associate virtual contents with environmental entities. However, the dissimilarity of different environments limits the adaptive AR content behaviors under large-scale deployment. We propose ScalAR, an integrated workflow enabling designers to author semantically adaptive AR experiences in Virtual Reality (VR). First, potential AR consumers collect local scenes with a semantic understanding technique. ScalAR then synthesizes numerous similar scenes. In VR, a designer authors the AR contents’ semantic associations and validates the design while being immersed in the provided scenes. We adopt a decision-tree-based algorithm to fit the designer’s demonstrations as a semantic adaptation model to deploy the authored AR experience in a physical scene. We further showcase two application scenarios authored by ScalAR and conduct a two-session user study where the quantitative results prove the accuracy of the AR content rendering and the qualitative results show the usability of ScalAR.
Xun Qian, Fengming He, Xiyun Hu, Tianyi Wang 0004, Ananya Ipsita, Karthik Ramani
CHI6
2022 InfoGCN: Representation Learning for Human Skeleton-based Action Recognition
abstract
Human skeleton-based action recognition offers a valuable means to understand the intricacies of human behavior because it can handle the complex relationships between physical constraints and intention. Although several studies have focused on encoding a skeleton, less attention has been paid to embed this information into the latent representations of human action. InfoGCN proposes a learning framework for action recognition combining a novel learning objective and an encoding method. First, we design an information bottleneck-based learning objective to guide the model to learn informative but compact latent representations. To provide discriminative information for classifying action, we introduce attention-based graph convolution that captures the context-dependent intrinsic topology of human action. In addition, we present a multi-modal representation of the skeleton using the relative position of joints, designed to provide complementary spatial information for joints. InfoGcn11Code is available at github.com/stnoahl/infogcn surpasses the known state-of-the-art on multiple skeleton-based action recognition benchmarks with the accuracy of 93.0% on NTU RGB+D 60 cross-subject split, 89.8% on NTU RGB+D 120 cross-subject split, and 97.0% on NW-UCLA.
Hyung-Gun Chi, Myoung Hoon Ha, Seunggeun Chi, Sang Wan Lee, Qixing Huang, Karthik Ramani
CVPR6
2022 EditAR: A Digital Twin Authoring Environment for Creation of AR/VR and Video Instructions from a Single Demonstration
abstract
Augmented/Virtual reality and video-based media play a vital role in the digital learning revolution to train novices in spatial tasks. However, creating content for these different media requires expertise in several fields. We present EditAR, a unified authoring, and editing environment to create content for AR, VR, and video based on a single demonstration. EditAR captures the user’s interaction within an environment and creates a digital twin, enabling users without programming backgrounds to develop content. We conducted formative interviews with both subject and media experts to design the system. The prototype was developed and reviewed by experts. We also performed a user study comparing traditional video creation with 2D video creation from 3D recordings, via a 3D editor, which uses freehand interaction for in-headset editing. Users took 5 times less time to record instructions and preferred EditAR, along with giving significantly higher usability scores.
Subramanian Chidambaram, Sai Swarup Reddy, Matthew Rumple, Ananya Ipsita, Ana M. Villanueva, Thomas Redick, Wolfgang Stuerzlinger, Karthik Ramani
ISMAR8
2022 ARnnotate: An Augmented Reality Interface for Collecting Custom Dataset of 3D Hand-Object Interaction Pose Estimation
abstract
Vision-based 3D pose estimation has substantial potential in hand-object interaction applications and requires user-specified datasets to achieve robust performance. We propose ARnnotate, an Augmented Reality (AR) interface enabling end-users to create custom data using a hand-tracking-capable AR device. Unlike other dataset collection strategies, ARnnotate first guides a user to manipulate a virtual bounding box and records its poses and the user’s hand joint positions as the labels. By leveraging the spatial awareness of AR, the user manipulates the corresponding physical object while following the in-situ AR animation of the bounding box and hand model, while ARnnotate captures the user’s first-person view as the images of the dataset. A 12-participant user study was conducted, and the results proved the system’s usability in terms of the spatial accuracy of the labels, the satisfactory performance of the deep neural networks trained with the data collected by ARnnotate, and the users’ subjective feedback.
Xun Qian, Fengming He, Xiyun Hu, Tianyi Wang 0004, Karthik Ramani
UIST5
2022 MechARspace: An Authoring System Enabling Bidirectional Binding of Augmented Reality with Toys in Real-time
abstract
Augmented Reality (AR), which blends physical and virtual worlds, presents the possibility of enhancing traditional toy design. By leveraging bidirectional virtual-physical interactions between humans and the designed artifact, such AR-enhanced toys can provide more playful and interactive experiences for traditional toys. However, designers are constrained by the complexity and technical difficulties of the current AR content creation processes. We propose MechARspace, an immersive authoring system that supports users to create toy-AR interactions through direct manipulation and visual programming. Based on the elicitation study, we propose a bidirectional interaction model which maps both ways: from the toy inputs to reactions of AR content, and also from the AR content to the toy reactions. This model guides the design of our system which includes a plug-and-play hardware toolkit and an in-situ authoring interface. We present multiple use cases enabled by MechARspace to validate this interaction model. Finally, we evaluate our system with a two-session user study where users first recreated a set of predefined toy-AR interactions and then implemented their own AR-enhanced toy designs.
Zhengzhe Zhu, Ziyi Liu 0004, Tianyi Wang 0004, Youyou Zhang, Xun Qian, Pashin Farsak Raja, Ana M. Villanueva, Karthik Ramani
UIST8
2022 ColabAR: A Toolkit for Remote Collaboration in Tangible Augmented Reality Laboratories
abstract
Current times are accelerating new technologies to provide high-quality education for remote collaboration, as well as hands-on learning. This is particularly important in the case of laboratory-based classes, which play an essential role in STEM education. In this paper, we introduce ColabAR, a toolkit that uses physical proxies to manipulate virtual objects in Tangible Augmented Reality (TAR) laboratories. ColabAR introduces haptic-based customizable interaction techniques to promote remote collaboration between students. Our toolkit provides hardware and software that enable haptic feedback to improve user experience and promote collaboration during learning. Also, we present the architecture of our cloud platform for haptic interaction that supports information sharing between students in a TAR laboratory. We performed two user studies (N=40) to test the effect of our toolkit in enriching local and remote collaborative experiences. Finally, we demonstrated that our TAR laboratory enables students' performance (i.e., lab completion rate, lab scores) to be similar to their performance in an in-person laboratory.
Ana M. Villanueva, Zhengzhe Zhu, Ziyi Liu 0004, Subramanian Chidambaram, Karthik Ramani
Proc. ACM Hum. Comput. Interact.6
2021 ProcessAR: An augmented reality-based tool to create in-situ procedural 2D/3D AR Instructions
abstract
Augmented reality (AR) is an efficient form of delivering spatial information and has great potential for training workers. However, AR is still not widely used for such scenarios due to the technical skills and expertise required to create interactive AR content. We developed ProcessAR, an AR-based system to develop 2D/3D content that captures subject matter expert’s (SMEs) environment-object interactions in situ. The design space for ProcessAR was identified from formative interviews with AR programming experts and SMEs, alongside a comparative design study with SMEs and novice users. To enable smooth workflows, ProcessAR locates and identifies different tools/objects through computer vision within the workspace when the author looks at them. We explored additional features such as embedding 2D videos with detected objects and user-adaptive triggers. A final user evaluation comparing ProcessAR and a baseline AR authoring environment showed that, according to our qualitative questionnaire, users preferred ProcessAR.
Subramanian Chidambaram, Hank Huang, Fengming He, Xun Qian, Ana M. Villanueva, Thomas Redick, Wolfgang Stuerzlinger, Karthik Ramani
Conference on Designing Interactive Systems8
2021 AdapTutAR: An Adaptive Tutoring System for Machine Tasks in Augmented Reality
abstract
Modern manufacturing processes are in a state of flux, as they adapt to increasing demand for flexible and self-configuring production. This poses challenges for training workers to rapidly master new machine operations and processes, i.e. machine tasks. Conventional in-person training is effective but requires time and effort of experts for each worker trained and not scalable. Recorded tutorials, such as video-based or augmented reality (AR), permit more efficient scaling. However, unlike in-person tutoring, existing recorded tutorials lack the ability to adapt to workers’ diverse experiences and learning behaviors. We present AdapTutAR, an adaptive task tutoring system that enables experts to record machine task tutorials via embodied demonstration and train learners with different AR tutoring contents adapting to each user’s characteristics. The adaptation is achieved by continually monitoring learners’ tutorial-following status and adjusting the tutoring content on-the-fly and in-situ. The results of our user study evaluation have demonstrated that our adaptive system is more effective and preferable than the non-adaptive one.
Gaoping Huang, Xun Qian, Tianyi Wang 0004, Fagun Patel, Maitreya Sreeram, Yuanzhi Cao, Karthik Ramani, Alexander J. Quinn
CHI7
2021 RobotAR: An Augmented Reality Compatible Teleconsulting Robotics Toolkit for Augmented Makerspace Experiences
abstract
Distance learning is facing a critical moment finding a balance between high quality education for remote students and engaging them in hands-on learning. This is particularly relevant for project-based classrooms and makerspaces, which typically require extensive trouble-shooting and example demonstrations from instructors. We present RobotAR, a teleconsulting robotics toolkit for creating Augmented Reality (AR) makerspaces. We present the hardware and software for an AR-compatible robot, which behaves as a student’s voice assistant and can be embodied by the instructor for teleconsultation. As a desktop-based teleconsulting agent, the instructor has control of the robot’s joints and position to better focus on areas of interest inside the workspace. Similarly, the instructor has access to the student’s virtual environment and the capability to create AR content to aid the student with problem-solving. We also performed a user study which compares current techniques for distance hands-on learning and an implementation of our toolkit.
Ana M. Villanueva, Ziyi Liu 0004, Zhengzhe Zhu, Joey Huang, Kylie Peppler, Karthik Ramani
CHI7
2021 GesturAR: An Authoring System for Creating Freehand Interactive Augmented Reality Applications
abstract
Freehand gesture is an essential input modality for modern Augmented Reality (AR) user experiences. However, developing AR applications with customized hand interactions remains a challenge for end-users. Therefore, we propose GesturAR, an end-to-end authoring tool that supports users to create in-situ freehand AR applications through embodied demonstration and visual programming. During authoring, users can intuitively demonstrate the customized gesture inputs while referring to the spatial and temporal context. Based on the taxonomy of gestures in AR, we proposed a hand interaction model which maps the gesture inputs to the reactions of the AR contents. Thus, users can author comprehensive freehand applications using trigger-action visual programming and instantly experience the results in AR. Further, we demonstrate multiple application scenarios enabled by GesturAR, such as interactive virtual objects, robots, and avatars, room-level interactive AR spaces, embodied AR presentations, etc. Finally, we evaluate the performance and usability of GesturAR through a user study.
Tianyi Wang 0004, Xun Qian, Fengming He, Xiyun Hu, Yuanzhi Cao, Karthik Ramani
UIST6
2021 Object Synthesis by Learning Part Geometry with Surface and Volumetric Representations
Sangpil Kim, Hyung-Gun Chi, Karthik Ramani
Comput. Aided Des.3
2020 First-Person View Hand Segmentation of Multi-Modal Hand Activity Video Dataset
Sangpil Kim, Hyung-Gun Chi, Xiao Hu 0004, Anirudh Vegesana, Karthik Ramani
BMVC5
2020 An Exploratory Study of Augmented Reality Presence for Tutoring Machine Tasks
abstract
Machine tasks in workshops or factories are often a compound sequence of local, spatial, and body-coordinated human-machine interactions. Prior works have shown the merits of video-based and augmented reality (AR) tutoring systems for local tasks. However, due to the lack of a bodily representation of the tutor, they are not as effective for spatial and body-coordinated interactions. We propose avatars as an additional tutor representation to the existing AR instructions. In order to understand the design space of tutoring presence for machine tasks, we conduct a comparative study with 32 users. We aim to explore the strengths/limitations of the following four tutor options: video, non-avatar-AR, half-body+AR, and full-body+AR. The results show that users prefer the half-body+AR overall, especially for the spatial interactions. They have a preference for the full-body+AR for the body-coordinated interactions and the non-avatar-AR for the local interactions. We further discuss and summarize design recommendations and insights for future machine task tutoring systems.
Yuanzhi Cao, Xun Qian, Tianyi Wang 0004, Rachel Lee, Ke Huo, Karthik Ramani
CHI6
2020 StoryMakAR: Bringing Stories to Life With An Augmented Reality & Physical Prototyping Toolkit for Youth
abstract
Makerspaces can support educational experiences in prototyping for children. Storytelling platforms enable high levels of creativity and expression, but have high barriers of entry. We introduce StoryMakAR, which combines making and storytelling. StoryMakAR is a new AR-IoT system for children that uses block programming, physical prototyping, and event-based storytelling to bring stories to life. We reduce the barriers to entry for youth (Age=14-18) by designing an accessible, plug-and-play system through merging both electro-mechanical devices and virtual characters to create stories. We describe our initial design process, the evolution and workflow of StoryMakAR, and results from multiple single-session workshops with 33 high school students. Our preliminary studies led us to understand what students want to make. We provide evidence of how students both engage and have difficulties with maker-based storytelling. We also discuss the potential for StoryMakAR to be used as a learning environment for classrooms and younger students.
Terrell Glenn, Ananya Ipsita, Caleb Carithers, Kylie Peppler, Karthik Ramani
CHI5
2020 Vipo: Spatial-Visual Programming with Functions for Robot-IoT Workflows
abstract
Mobile robots and IoT (Internet of Things) devices can increase productivity, but only if they can be programmed by workers who understand the domain. This is especially true in manufacturing. Visual programming in the spatial context of the operating environment can enable mental models at a familiar level of abstraction. However, spatial-visual programming is still in its infancy; existing systems lack IoT integration and fundamental constructs, such as functions, that are essential for code reuse, encapsulation, or recursive algorithms. We present Vipo, a spatial-visual programming system for robot-IoT workflows. Vipo was designed with input from managers at six factories using mobile robots. Our user study (n=22) evaluated efficiency, correctness, comprehensibility of spatial-visual programming with functions.
Gaoping Huang, Pawan S. Rao, Meng-Han Wu, Xun Qian, Shimon Y. Nof, Karthik Ramani, Alexander J. Quinn
CHI6
2020 Meta-AR-App: An Authoring Platform for Collaborative Augmented Reality in STEM Classrooms
abstract
Augmented Reality (AR) has become a valuable tool for education and training processes. Meanwhile, cloud-based technologies can foster collaboration and other interaction modalities to enhance learning. We combine the cloud capabilities with AR technologies to present Meta-AR-App, an authoring platform for collaborative AR, which enables authoring between instructors and students. Additionally, we introduce a new application of an established collaboration process, the pull-based development model, to enable sharing and retrieving of AR learning content. We customize this model and create two modalities of interaction for the classroom: local (student to student) and global (instructor to class) pull. Based on observations from our user studies, we organize a four-category classroom model which implements our system: Work, Design, Collaboration, and Technology. Further, our system enables an iterative improvement workflow of the class content and enables synergistic collaboration that empowers students to be active agents in the learning process.
Ana M. Villanueva, Zhengzhe Zhu, Ziyi Liu 0004, Kylie Peppler, Thomas Redick, Karthik Ramani
CHI6
2020 A Large-Scale Annotated Mechanical Components Benchmark for Classification and Retrieval Tasks with Deep Neural Networks
Sangpil Kim, Hyung-Gun Chi, Xiao Hu 0004, Qixing Huang, Karthik Ramani
ECCV (18)5
2020 CAPturAR: An Augmented Reality Tool for Authoring Human-Involved Context-Aware Applications
abstract
Recognition of human behavior plays an important role in context-aware applications. However, it is still a challenge for end-users to build personalized applications that accurately recognize their own activities. Therefore, we present CAPturAR, an in-situ programming tool that supports users to rapidly author context-aware applications by referring to their previous activities. We customize an AR head-mounted device with multiple camera systems that allow for non-intrusive capturing of user's daily activities. During authoring, we reconstruct the captured data in AR with an animated avatar and use virtual icons to represent the surrounding environment. With our visual programming interface, users create human-centered rules for the applications and experience them instantly in AR. We further demonstrate four use cases enabled by CAPturAR. Also, we verify the effectiveness of the AR-HMD and the authoring workflow with a system evaluation using our prototype. Moreover, we conduct a remote user study in an AR simulator to evaluate the usability.
Tianyi Wang 0004, Xun Qian, Fengming He, Xiyun Hu, Ke Huo, Yuanzhi Cao, Karthik Ramani
UIST7
2020 Using social interaction trace data and context to predict collaboration quality and creative fluency in collaborative design learning environments
abstract
Engineering design typically occurs as a collaborative process situated in specific context such as computer-supported environments, however there is limited research examining the dynamics of design collaboration in specific contexts. In this study, drawing from situative learning theory, we developed two analytic lenses to broaden theoretical insights into collaborative design practices in computer-supported environments: (a) the role of spatial and material context, and (b) the role of social interactions. We randomly assigned participants to four conditions varying the material context (paper vs. tablet sketching tools) and spatial environment (private room vs commons area) as they worked collaboratively to generate ideas for a toy design task. We used wearable sociometric badges to automatically and unobtrusively collect social interaction data. Using partial least squares regression, we generated two predictive models for collaboration quality and creative fluency. We found that context matters materially to perceptions of collaboration, where those using collaboration-support tools perceived higher quality collaboration. But context matters spatially to creativity, and those situated in private spaces are more fluent in generating ideas than those in commons areas. We also found that interaction dynamics differ: synchronous interaction is important to quality collaboration, but reciprocal interaction is important to creative fluency. These findings provide important insights into the processual factors in collaborative design in computer-supported environments, and the predictive role of context and conversation dynamics. We discuss the theoretical contributions to computer-supported collaborative design, the methodological contributions of wearable sensor tools, and the practical contributions to structuring computer-supported environments for engineering design practice.
Ninger Zhou, Lorraine G. Kisselburgh, Senthil K. Chandrasegaran, Sriram Karthik Badam, Niklas Elmqvist, Karthik Ramani
Int. J. Hum. Comput. Stud.6
2020 Latent transformations neural network for object view synthesis
Sangpil Kim, Nick Winovich, Hyung-Gun Chi, Guang Lin 0001, Karthik Ramani
Vis. Comput.5
2019 V.Ra: An In-Situ Visual Authoring System for Robot-IoT Task Planning with Augmented Reality
abstract
We present V.Ra, a visual and spatial programming system for robot-IoT task authoring. In V.Ra, programmable mobile robots serve as binding agents to link the stationary IoTs and perform collaborative tasks. We establish an ecosystem that coherently connects the three key elements of robot task planning , the human, robot and IoT, with one single mobile AR device. Users can perform task authoring with the Augmented Reality (AR) handheld interface, then placing the AR device onto the mobile robot directly transfers the task plan in a what-you-do-is-what-robot-does (WYDWRD) manner. The mobile device mediates the interactions between the user, robot, and the IoT oriented tasks, and guides the path planning execution with the embedded simultaneous localization and mapping (SLAM) capability. We demonstrate that V.Ra enables instant, robust and intuitive room-scale navigatory and interactive task authoring through various use cases and preliminary studies.
Yuanzhi Cao, Zhuangying Xu, Wentao Zhong, Ke Huo, Karthik Ramani
Conference on Designing Interactive Systems6
2019 Shape Structuralizer: Design, Fabrication, and User-driven Iterative Refinement of 3D Mesh Models
abstract
Current Computer-Aided Design (CAD) tools lack proper support for guiding novice users towards designs ready for fabrication. We propose Shape Structuralizer (SS), an interactive design support system that repurposes surface models into structural constructions using rods and custom 3D-printed joints. Shape Structuralizer embeds a recommendation system that computationally supports the user during design ideation by providing design suggestions on local refinements of the design. This strategy enables novice users to choose designs that both satisfy stress constraints as well as their personal design intent. The interactive guidance enables users to repurpose existing surface mesh models, analyze them in-situ for stress and displacement constraints, add movable joints to increase functionality, and attach a customized appearance. This also empowers novices to fabricate even complex constructs while ensuring structural soundness. We validate the Shape Structuralizer tool with a qualitative user study where we observed that even novice users were able to generate a large number of structurally safe designs for fabrication.
Subramanian Chidambaram, Venkatraghavan Sundararajan, Niklas Elmqvist, Karthik Ramani
CHI5
2019 Deep Learning 3D Shapes Using Alt-az Anisotropic 2-Sphere Convolution
Min Liu 0018, Fupin Yao, Chiho Choi, Ayan Sinha, Karthik Ramani
ICLR (Poster)5
2019 GhostAR: A Time-space Editor for Embodied Authoring of Human-Robot Collaborative Task with Augmented Reality
abstract
We present GhostAR, a time-space editor for authoring and acting Human-Robot-Collaborative (HRC) tasks in-situ. Our system adopts an embodied authoring approach in Augmented Reality (AR), for spatially editing the actions and programming the robots through demonstrative role-playing. We propose a novel HRC workflow that externalizes user's authoring as demonstrative and editable AR ghost, allowing for spatially situated visual referencing, realistic animated simulation, and collaborative action guidance. We develop a dynamic time warping (DTW) based collaboration model which takes the real-time captured motion as inputs, maps it to the previously authored human actions, and outputs the corresponding robot actions to achieve adaptive collaboration. We emphasize an in-situ authoring and rapid iterations of joint plans without an offline training process. Further, we demonstrate and evaluate the effectiveness of our workflow through HRC use cases and a three-session user study.
Yuanzhi Cao, Tianyi Wang 0004, Xun Qian, Pawan S. Rao, Manav Wadhawan, Ke Huo, Karthik Ramani
UIST7
2018 Plain2Fun: Augmenting Ordinary Objects with Interactive Functions by Auto-Fabricating Surface Painted Circuits
abstract
The growing makers' community demands better supports for designing and fabricating interactive functional objects. Most of the current approaches focus on embedding desired functions within new objects. Instead, we advocate repurposing the existing objects and rapidly authoring interactive functions onto them. We present Plain2Fun, a design and fabrication pipeline enabling users to quickly transform ordinary objects into interactive and functional ones. Plain2Fun allows users to directly design the circuit layouts onto the surfaces of the scanned 3D model of existing objects. Our design tool automatically generates as short as possible circuit paths between any two points while avoiding intersections. Further, we build a digital machine to construct the conductive paths accurately. With a specially designed housing base, users can simply snap the electronic components onto the surfaces and obtain working physical prototypes. Moreover, we evaluate the usability of our system with multiple use cases and a preliminary user study.
Tianyi Wang 0004, Ke Huo, Pratik Chawla, Guiming Chen, Siddharth Banerjee, Karthik Ramani
Conference on Designing Interactive Systems6
2018 Scenariot: Spatially Mapping Smart Things Within Augmented Reality Scenes
abstract
The emerging simultaneous localizing and mapping (SLAM) based tracking technique allows the mobile AR device spatial awareness of the physical world. Still, smart things are not fully supported with the spatial awareness in AR. Therefore, we present Scenariot, a method that enables instant discovery and localization of the surrounding smart things while also spatially registering them with a SLAM based mobile AR system. By exploiting the spatial relationships between mobile AR systems and smart things, Scenariot fosters in-situ interactions with connected devices. We embed Ultra-Wide Band (UWB) RF units into the AR device and the controllers of the smart things, which allows for measuring the distances between them. With a one-time initial calibration, users localize multiple IoT devices and map them within the AR scenes. Through a series of experiments and evaluations, we validate the localization accuracy as well as the performance of the enabled spatial aware interactions. Further, we demonstrate various use cases through Scenariot.
Ke Huo, Yuanzhi Cao, Sang Ho Yoon, Zhuangying Xu, Guiming Chen, Karthik Ramani
CHI6
2018 Ani-Bot: A Modular Robotics System Supporting Creation, Tweaking, and Usage with Mixed-Reality Interactions
abstract
Ani-Bot is a modular robotics system that allows users to control their DIY robots using Mixed-Reality Interaction (MRI). This system takes advantage of MRI to enable users to visually program the robot through the augmented view of a Head-Mounted Display (HMD). In this paper, we first explain the design of the Mixed-Reality (MR) ready modular robotics system, which allows users to instantly perform MRI once they finish assembling the robot. Then, we elaborate the augmentations provided by the MR system in the three primary phases of a construction kit's lifecycle: Creation, Tweaking, and Usage. Finally, we demonstrate Ani-Bot with four application examples and evaluate the system with a two-session user study. The results of our evaluation indicate that Ani-Bot does successfully embed MRI into the lifecycle (Creation, Tweaking, Usage) of DIY robotics and that it does show strong potential for delivering an enhanced user experience.
Yuanzhi Cao, Zhuangying Xu, Terrell Glenn, Ke Huo, Karthik Ramani
TEI5
2018 SynchronizAR: Instant Synchronization for Spontaneous and Spatial Collaborations in Augmented Reality
abstract
We present SynchronizAR, an approach to spatially register multiple SLAM devices together without sharing maps or involving external tracking infrastructures. SynchronizAR employs a distance based indirect registration which resolves the transformations between the separate SLAM coordinate systems. We attach an Ultra-Wide Bandwidth~(UWB) based distance measurements module on each of the mobile AR devices which is capable of self-localization with respect to the environment. As users move on independent paths, we collect the positions of the AR devices in their local frames and the corresponding distance measurements. Based on the registration, we support to create a spontaneous collaborative AR environment to spatially coordinate users' interactions. We run both technical evaluation and user studies to investigate the registration accuracy and the usability towards spatial collaborations. Finally, we demonstrate various collaborative AR experience using SynchronizAR.
Ke Huo, Tianyi Wang 0004, Luis Paredes, Ana M. Villanueva, Yuanzhi Cao, Karthik Ramani
UIST6
2017 3D Object Classification via Spherical Projections
abstract
In this paper, we introduce a new method for classifying 3D objects. Our main idea is to project a 3D object onto a spherical domain centered around its barycenter and develop neural network to classify the spherical projection. We introduce two complementary projections. The first captures depth variations of a 3D object, and the second captures contour-information viewed from different angles. Spherical projections combine key advantages of two main-stream 3D classification methods: image-based and 3D-based. Specifically, spherical projections are locally planar, allowing us to use massive image datasets (e.g, ImageNet) for pre-training. Also spherical projections are similar to voxel-based methods, as they encode complete information of a 3D object in a single neural network capturing dependencies across different views. Our novel network design can fully utilize these advantages. Experimental results on ModelNet40 and ShapeNetCore show that our method is superior to prior methods.
Zhangjie Cao, Qixing Huang, Karthik Ramani
3DV3
2017 Modeling Cumulative Arm Fatigue in Mid-Air Interaction based on Perceived Exertion and Kinetics of Arm Motion
abstract
Quantifying cumulative arm muscle fatigue is a critical factor in understanding, evaluating, and optimizing user experience during prolonged mid-air interaction. A reasonably accurate estimation of fatigue requires an estimate of an individual's strength. However, there is no easy-to-access method to measure individual strength to accommodate inter-individual differences. Furthermore, fatigue is influenced by both psychological and physiological factors, but no current HCI model provides good estimates of cumulative subjective fatigue. We present a new, simple method to estimate the maximum shoulder torque through a mid-air pointing task, which agrees with direct strength measurements. We then introduce a cumulative fatigue model informed by subjective and biomechanical measures. We evaluate the performance of the model in estimating cumulative subjective fatigue in mid-air interaction by performing multiple cross-validations and a comparison with an existing fatigue metric. Finally, we discuss the potential of our approach for real-time evaluation of subjective fatigue as well as future challenges.
Sujin Jang, Wolfgang Stuerzlinger, Satyajit Ambike, Karthik Ramani
CHI4
2017 WireFab: Mix-Dimensional Modeling and Fabrication for 3D Mesh Models
abstract
Many rapid fabrication technologies are directed towards layer wise printing or laser based prototyping. We propose WireFab, a rapid modeling and prototyping system that uses bent metal wires as the structure framework. WireFab approximates both the skeletal articulation and the skin appearance of the corresponding virtual skin meshes, and it allows users to personalize the designs by (1) specifying joint positions and part segmentations, (2) defining joint types and motion ranges to build a wire-based skeletal model, and (3) abstracting the segmented meshes into mixed-dimensional appearance patterns or attachments.
Min Liu 0018, Jing Bai 0004, Yuanzhi Cao, Jeffrey M. Alperovich, Karthik Ramani
CHI6
2017 Co-3Deator: A Team-First Collaborative 3D Design Ideation Tool
abstract
We present Co-3Deator, a sketch-based collaborative 3D modeling system based on the notion of "team-first" ideation tools, where the needs and processes of the entire design team come before that of an individual designer. Co-3Deator includes two specific team-first features: a concept component hierarchy which provides a design representation suitable for multi-level sharing and reusing of design information, and a collaborative design explorer for storing, viewing, and accessing hierarchical design data during collaborative design activities. We conduct two controlled user studies, one with individual designers to elicit the form and functionality of the collaborative design explorer, and the other with design teams to evaluate the utility of the concept component hierarchy and design explorer towards collaborative design ideation. Our results support our rationale for both of the proposed team-first collaboration mechanisms and suggest further ways to streamline collaborative design.
Cecil Piya, Vinayak R. Krishnamurthy, Senthil K. Chandrasegaran, Niklas Elmqvist, Karthik Ramani
CHI5
2017 SurfNet: Generating 3D Shape Surfaces Using Deep Residual Networks
abstract
3D shape models are naturally parameterized using vertices and faces, i.e., composed of polygons forming a surface. However, current 3D learning paradigms for predictive and generative tasks using convolutional neural networks focus on a voxelized representation of the object. Lifting convolution operators from the traditional 2D to 3D results in high computational overhead with little additional benefit as most of the geometry information is contained on the surface boundary. Here we study the problem of directly generating the 3D shape surface of rigid and non-rigid shapes using deep convolutional neural networks. We develop a procedure to create consistent `geometry images' representing the shape surface of a category of 3D objects. We then use this consistent representation for category-specific shape surface generation from a parametric representation or an image by developing novel extensions of deep residual networks for the task of geometry image generation. Our experiments indicate that our network learns a meaningful representation of shape surfaces allowing it to interpolate between shape orientations and poses, invent new shape surfaces and reconstruct 3D shape surfaces from previously unseen images. Our code is available at https://github.com/sinhayan/surfnet.
Ayan Sinha, Asim Unmesh, Qixing Huang, Karthik Ramani
CVPR4
2017 Merging Sketches for Creative Design Exploration: An Evaluation of Physical and Cognitive Operations
Senthil K. Chandrasegaran, Sriram Karthik Badam, Ninger Zhou, Zhenpeng Zhao, Lorraine G. Kisselburgh, Kylie Peppler, Niklas Elmqvist, Karthik Ramani
Graphics Interface8
2017 Learning Hand Articulations by Hallucinating Heat Distribution
abstract
We propose a robust hand pose estimation method by learning hand articulations from depth features and auxiliary modality features. As an additional modality to depth data, we present a function of geometric properties on the surface of the hand described by heat diffusion. The proposed heat distribution descriptor is robust to identify the keypoints on the surface as it incorporates both the local geometry of the hand and global structural representation at multiple time scales. Along this line, we train our heat distribution network to learn the geometrically descriptive representations from the proposed descriptors with the fingertip position labels. Then the hallucination network is guided to mimic the intermediate responses of the heat distribution modality from a paired depth image. We use the resulting geometrically informed responses together with the discriminative depth features estimated from the depth network to regularize the angle parameters in the refinement network. To this end, we conduct extensive evaluations to validate that the proposed framework is powerful as it achieves state-of-the-art performance.
Chiho Choi, Sangpil Kim, Karthik Ramani
ICCV3
2017 Robust Hand Pose Estimation during the Interaction with an Unknown Object
abstract
This paper proposes a robust solution for accurate 3D hand pose estimation in the presence of an external object interacting with hands. Our main insight is that the shape of an object causes a configuration of the hand in the form of a hand grasp. Along this line, we simultaneously train deep neural networks using paired depth images. The object-oriented network learns functional grasps from an object perspective, whereas the hand-oriented network explores the details of hand configurations from a hand perspective. The two networks share intermediate observations produced from different perspectives to create a more informed representation. Our system then collaboratively classifies the grasp types and orientation of the hand and further constrains a pose space using these estimates. Finally, we collectively refine the unknown pose parameters to reconstruct the final hand pose. To this end, we conduct extensive evaluations to validate the efficacy of the proposed collaborative learning approach by comparing it with self-generated baselines and the state-of-the-art method.
Chiho Choi, Sang Ho Yoon, Chin-Ning Chen, Karthik Ramani
ICCV4
2017 Window-Shaping: 3D Design Ideation by Creating on, Borrowing from, and Looking at the Physical World
abstract
We present, Window-Shaping, a tangible mixed-reality (MR) interaction metaphor for design ideation that allows for the direct creation of 3D shapes on and around physical objects. Using the sketch-and-inflate scheme, our metaphor enables quick design of dimensionally consistent and visually coherent 3D models by borrowing visual and dimensional attributes from existing physical objects without the need for 3D reconstruction or fiducial markers. Through a preliminary evaluation of our prototype application we demonstrate the expressiveness provided by our design workflow, the effectiveness of our interaction scheme, and the potential of our metaphor.
Ke Huo, Vinayak R. Krishnamurthy, Karthik Ramani
TEI3
2017 iSoft: A Customizable Soft Sensor with Real-time Continuous Contact and Stretching Sensing
abstract
We present iSoft, a single volume soft sensor capable of sensing real-time continuous contact and unidirectional stretching. We propose a low-cost and an easy way to fabricate such piezoresistive elastomer-based soft sensors for instant interactions. We employ an electrical impedance tomography (EIT) technique to estimate changes of resistance distribution on the sensor caused by fingertip contact. To compensate for the rebound elasticity of the elastomer and achieve real-time continuous contact sensing, we apply a dynamic baseline update for EIT. The baseline updates are triggered by fingertip contact and movement detections. Further, we support unidirectional stretching sensing using a model-based approach which works separately with continuous contact sensing. We also provide a software toolkit for users to design and deploy personalized interfaces with customized sensors. Through a series of experiments and evaluations, we validate the performance of contact and stretching sensing. Through example applications, we show the variety of examples enabled by iSoft.
Sang Ho Yoon, Ke Huo, Guiming Chen, Luis Paredes, Subramanian Chidambaram, Karthik Ramani
UIST7
2017 Integrating Visual Analytics Support for Grounded Theory Practice in Qualitative Text Analysis
abstract
Abstract We present an argument for using visual analytics to aid Grounded Theory methodologies in qualitative data analysis. Grounded theory methods involve the inductive analysis of data to generate novel insights and theoretical constructs. Making sense of unstructured text data is uniquely suited for visual analytics. Using natural language processing techniques such as parts‐of‐speech tagging, retrieving information content, and topic modeling, different parts of the data can be structured and semantically associated, and interactively explored, thereby providing conceptual depth to the guided discovery process. We review grounded theory methods and identify processes that can be enhanced through visual analytic techniques. Next, we develop an interface for qualitative text analysis, and evaluate our design with qualitative research practitioners who analyze texts with and without visual analytics support. The results of our study suggest how visual analytics can be incorporated into qualitative data analysis tools, and the analytic and interpretive benefits that can result.
Senthil K. Chandrasegaran, Sriram Karthik Badam, Lorraine G. Kisselburgh, Karthik Ramani, Niklas Elmqvist
Comput. Graph. Forum4
2017 VizScribe: A visual analytics approach to understand designer behavior
abstract
Design protocol analysis is a technique to understand designers’ cognitive processes by analyzing sequences of observations on their behavior. These observations typically use audio, video, and transcript data in order to gain insights into the designer's behavior and the design process. The recent availability of sophisticated sensing technology has made such data highly multimodal, requiring more flexible protocol analysis tools. To address this need, we present VizScribe, a visual analytics framework that employs multiple coordinated multiple views that enable the viewing of such data from different perspectives. VizScribe allows designers to create, customize, and extend interactive visualizations for design protocol data such as video, transcripts, sketches, sensor data, and user logs. User studies where design researchers used VizScribe for protocol analysis indicated that the linked views and interactive navigation offered by VizScribe afforded the researchers multiple, useful ways to approach and interpret such multimodal data.
Senthil K. Chandrasegaran, Sriram Karthik Badam, Lorraine G. Kisselburgh, Kylie Peppler, Niklas Elmqvist, Karthik Ramani
Int. J. Hum. Comput. Stud.6
2016 CardBoardiZer: Creatively Customize, Articulate and Fold 3D Mesh Models
abstract
Computer-aided design of flat patterns allows designers to prototype foldable 3D objects made of heterogeneous sheets of material. We found origami designs are often characterized by pre-synthesized patterns and automated algorithms. Furthermore, augmenting articulated features to a desired model requires time-consuming synthesis of interconnected joints. This paper presents CardBoardiZer, a rapid cardboard based prototyping platform that allows everyday sculptural 3D models to be easily customized, articulated and folded. We develop a building platform to allow the designer to 1) import a desired 3D shape, 2) customize articulated partitions into planar or volumetric foldable patterns, and 3) define rotational movements between partitions. The system unfolds the model into 2D crease-cut-slot patterns ready for die-cutting and folding. In this paper, we developed interactive algorithms and validated the usability of CardBoardiZer using various 3D models. Furthermore, comparisons between CardBoardiZer and methods of Autodesk® 123D Make, demonstrated significantly shorter time-to-prototype and ease of fabrication.
Luis Paredes, Karthik Ramani
CHI4
2016 DeepHand: Robust Hand Pose Estimation by Completing a Matrix Imputed with Deep Features
abstract
We propose DeepHand to estimate the 3D pose of a hand using depth data from commercial 3D sensors. We discriminatively train convolutional neural networks to output a low dimensional activation feature given a depth map. This activation feature vector is representative of the global or local joint angle parameters of a hand pose. We efficiently identify 'spatial' nearest neighbors to the activation feature, from a database of features corresponding to synthetic depth maps, and store some 'temporal' neighbors from previous frames. Our matrix completion algorithm uses these 'spatio-temporal' activation features and the corresponding known pose parameter values to estimate the unknown pose parameters of the input feature vector. Our database of activation features supplements large viewpoint coverage and our hierarchical estimation of pose parameters is robust to occlusions. We show that our approach compares favorably to state-of-the-art methods while achieving real time performance (≈ 32 FPS) on a standard computer.
Ayan Sinha, Chiho Choi, Karthik Ramani
CVPR3
2016 Deep Learning 3D Shape Surfaces Using Geometry Images
Ayan Sinha, Jing Bai 0004, Karthik Ramani
ECCV (6)3
2016 RealFusion: An Interactive Workflow for Repurposing Real-World Objects towards Early-stage Creative Ideation
Cecil Piya, Vinayak R. Krishnamurthy, Karthik Ramani
Graphics Interface4
2016 Cubimorph: Designing modular interactive devices
abstract
We introduce Cubimorph, a modular interactive device that accommodates touchscreens on each of the six module faces, and that uses a hinge-mounted turntable mechanism to self-reconfigure in the user's hand. Cubimorph contributes toward the vision of programmable matter where interactive devices reconfigure in any shape that can be made out of a chain of cubes in order to fit a myriad of functionalities, e.g. a mobile phone shifting into a console when a user launches a game. We present a design rationale that exposes user requirements to consider when designing homogeneous modular interactive devices. We present our Cubimorph mechanical design, three prototypes demonstrating key aspects (turntable hinges, embedded touchscreens and miniaturization), and an adaptation of the probabilistic roadmap algorithm for the reconfiguration.
Anne Roudaut, Diana Krusteva, Mike McCoy, Abhijit Karnik, Karthik Ramani, Sriram Subramanian
ICRA5
2016 Deconvolving Feedback Loops in Recommender Systems
abstract
Collaborative filtering is a popular technique to infer users' preferences on new content based on the collective information of all users preferences. Recommender systems then use this information to make personalized suggestions to users. When users accept these recommendations it creates a feedback loop in the recommender system, and these loops iteratively influence the collaborative filtering algorithm's predictions over time. We investigate whether it is possible to identify items affected by these feedback loops. We state sufficient assumptions to deconvolve the feedback loops while keeping the inverse solution tractable. We furthermore develop a metric to unravel the recommender system's influence on the entire user-item rating matrix. We use this metric on synthetic and real-world datasets to (1) identify the extent to which the recommender system affects the final rating matrix, (2) rank frequently recommended items, and (3) distinguish whether a user's rated item was recommended or an intrinsic preference. Our results indicate that it is possible to recover the ratings matrix of intrinsic user preferences using a single snapshot of the ratings matrix without any temporal information.
Ayan Sinha, David F. Gleich, Karthik Ramani
NIPS3
2016 MobiSweep: Exploring Spatial Design Ideation Using a Smartphone as a Hand-held Reference Plane
abstract
In this paper, we explore quick 3D shape composition during early-phase spatial design ideation. Our approach is to re-purpose a smartphone as a hand-held reference plane for creating, modifying, and manipulating 3D sweep surfaces. We implemented MobiSweep, a prototype application to explore a new design space of constrained spatial interactions that combine direct orientation control with indirect position control via well-established multi-touch gestures. MobiSweep leverages kinesthetically aware interactions for the creation of a sweep surface without explicit position tracking. The design concepts generated by users, in conjunction with their feedback, demonstrate the potential of such interactions in enabling spatial ideation.
Vinayak R. Krishnamurthy, Devarajan Ramanujan, Cecil Piya, Karthik Ramani
TEI4
2016 TMotion: Embedded 3D Mobile Input using Magnetic Sensing Technique
abstract
We present TMotion, a self-contained 3D input that enables spatial interactions around mobile device using a magnetic sensing technique. We embed a permanent magnet and an inertial measurement unit (IMU) in a stylus. When the stylus moves around the mobile device, we obtain a continuous magnetometer readings. By numerically solving non-linear magnetic field equations with known orientation from IMU, we achieve 3D position tracking with update rate greater than 30Hz. Our experiments evaluated the position tracking accuracy, showing an average error of 4.55mm in the space of 80mm×120mm×100mm. Furthermore, the experiments confirmed the tracking robustness against orientations and dynamic tracings. In task evaluations, we verified the tracking and targeting performance in spatial interactions with users. We demonstrate example applications that highlight TMotion's interaction capability.
Sang Ho Yoon, Ke Huo, Karthik Ramani
TEI3
2016 TRing: Instant and Customizable Interactions with Objects Using an Embedded Magnet and a Finger-Worn Device
abstract
We present TRing, a finger-worn input device which provides instant and customizable interactions. TRing offers a novel method for making plain objects interactive using an embedded magnet and a finger-worn device. With a particle filter integrated magnetic sensing technique, we compute the fingertip's position relative to the embedded magnet. We also offer a magnet placement algorithm that guides the magnet installation location based upon the user's interface customization. By simply inserting or attaching a small magnet, we bring interactivity to both fabricated and existing objects. In our evaluations, TRing shows an average tracking error of 8.6 mm in 3D space and a 2D targeting error of 4.96 mm, which are sufficient for implementing average-sized conventional controls such as buttons and sliders. A user study validates the input performance with TRing on a targeting task (92% accuracy within 45 mm distance) and a cursor control task (91% accuracy for a 10 mm target). Furthermore, we show examples that highlight the interaction capability of our approach.
Sang Ho Yoon, Ke Huo, Karthik Ramani
UIST4
2016 Extracting hand grasp and motion for intent expression in mid-air shape deformation: A concrete and iterative exploration through a virtual pottery application
Vinayak R. Krishnamurthy, Karthik Ramani
Comput. Graph.2
2016 Wearable textile input device with multimodal sensing for eyes-free mobile interaction during daily activities
Sang Ho Yoon, Ke Huo, Karthik Ramani
Pervasive Mob. Comput.3
2016 The Intrinsic Geometric Structure of Protein-Protein Interaction Networks for Protein Interaction Prediction
abstract
Recent developments in high-throughput technologies for measuring protein-protein interaction (PPI) have profoundly advanced our ability to systematically infer protein function and regulation. However, inherently high false positive and false negative rates in measurement have posed great challenges in computational approaches for the prediction of PPI. A good PPI predictor should be 1) resistant to high rate of missing and spurious PPIs, and 2) robust against incompleteness of observed PPI networks. To predict PPI in a network, we developed an intrinsic geometry structure (IGS) for network, which exploits the intrinsic and hidden relationship among proteins in network through a heat diffusion process. In this process, all explicit PPIs participate simultaneously to glue local infinitesimal and noisy experimental interaction data to generate a global macroscopic descriptions about relationships among proteins. The revealed implicit relationship can be interpreted as the probability of two proteins interacting with each other. The revealed relationship is intrinsic and robust against individual, local and explicit protein interactions in the original network. We apply our approach to publicly available PPI network data for the evaluation of the performance of PPI prediction. Experimental results indicate that, under different levels of the missing and spurious PPIs, IGS is able to robustly exploit the intrinsic and hidden relationship for PPI prediction with a higher sensitivity and specificity compared to that of recently proposed methods.
Yi Fang 0006, Mengtian Sun, Guoxian Dai, Karthik Ramani
IEEE ACM Trans. Comput. Biol. Bioinform.4
2016 MotionFlow: Visual Abstraction and Aggregation of Sequential Patterns in Human Motion Tracking Data
abstract
Pattern analysis of human motions, which is useful in many research areas, requires understanding and comparison of different styles of motion patterns. However, working with human motion tracking data to support such analysis poses great challenges. In this paper, we propose MotionFlow, a visual analytics system that provides an effective overview of various motion patterns based on an interactive flow visualization. This visualization formulates a motion sequence as transitions between static poses, and aggregates these sequences into a tree diagram to construct a set of motion patterns. The system also allows the users to directly reflect the context of data and their perception of pose similarities in generating representative pose states. We provide local and global controls over the partition-based clustering process. To support the users in organizing unstructured motion data into pattern groups, we designed a set of interactions that enables searching for similar motion sequences from the data, detailed exploration of data subsets, and creating and modifying the group of motion patterns. To evaluate the usability of MotionFlow, we conducted a user study with six researchers with expertise in gesture-based interaction design. They used MotionFlow to explore and organize unstructured motion tracking data. Results show that the researchers were able to easily learn how to use MotionFlow, and the system effectively supported their pattern analysis activities, including leveraging their perception and domain knowledge.
Sujin Jang, Niklas Elmqvist, Karthik Ramani
IEEE Trans. Vis. Comput. Graph.3
2015 HandiMate: exploring a modular robotics kit for animating crafted toys
abstract
Building from our previous work we explore HandiMate, a robotics kit which enables users to construct and animate their toys using everyday craft materials [32]. The kit contains eight joint modules, a tablet interface and a glove controller. Unlike popular kits, HandiMate does not rely on manufactured parts to construct the toy. Rather this open ended platform engages users to pursue interest driven activities using everyday objects, such as cardboard, construction paper, and spoons. These crafted parts are then fastened together using Velcro to the joint modules and animated using the glove as the controller. In this paper, we discuss the results from two user studies which were designed to understand the affinity of HandiMate among children. The first study reveals that children rated the HandiMate kit as gender-neutral, appealing equally to both female and male students. The second study discusses the benefits of engaging children in engineering design with HandiMate, which has been observed to bring out children's tacit physics-based engineering knowledge and facilitate learning.
Sang Ho Yoon, Ansh Verma, Kylie Peppler, Karthik Ramani
IDC4
2015 Hand grasp and motion for intent expression in mid-air virtual pottery
Vinayak R. Krishnamurthy, Karthik Ramani
Graphics Interface2
2015 A Collaborative Filtering Approach to Real-Time Hand Pose Estimation
abstract
Collaborative filtering aims to predict unknown user ratings in a recommender system by collectively assessing known user preferences. In this paper, we first draw analogies between collaborative filtering and the pose estimation problem. Specifically, we recast the hand pose estimation problem as the cold-start problem for a new user with unknown item ratings in a recommender system. Inspired by fast and accurate matrix factorization techniques for collaborative filtering, we develop a real-time algorithm for estimating the hand pose from RGB-D data of a commercial depth camera. First, we efficiently identify nearest neighbors using local shape descriptors in the RGB-D domain from a library of hand poses with known pose parameter values. We then use this information to evaluate the unknown pose parameters using a joint matrix factorization and completion (JMFC) approach. Our quantitative and qualitative results suggest that our approach is robust to variation in hand configurations while achieving real time performance (≈ 29 FPS) on a standard computer.
Chiho Choi, Ayan Sinha, Joon Hee Choi, Sujin Jang, Karthik Ramani
ICCV5
2015 SOFTii: Soft Tangible Interface for Continuous Control of Virtual Objects with Pressure-based Input
abstract
We present SOFTii, a flexible input system for topography design and continuous control via external force. Our intent is to provide a tactile metaphor for pressure-based surface input. In this study, two prototypes of SOFTii have been fabricated: (a) The first prototype has one pressure surface for topography design with everyday tangible objects, (b) the second prototype, having two force input surfaces, performs as a deformable controller for video games and continuous shape modeling using a SVM algorithm. Both prototypes of SOFTii are constructed by layering Polymethylsiloxane (PDMS), ITO coated PET film, and conductive fabric and foam. The layer configuration allows the capturing of local pressure on the SOFTii surface via distributed electrodes. Here we further discuss the implementation of the device with possible usage scenarios.
Vinh P. Nguyen, Sang Ho Yoon, Ansh Verma, Karthik Ramani
TEI5
2015 HandiMate: Create and Animate using Everyday Objects as Material
abstract
The combination of technological progress and a growing interest in design has promoted the prevalence of DIY (Do It Yourself) and craft activities. We introduce HandiMate, a platform that makes it easier for people without technical expertise to fabricate and animate electro-mechanical systems from everyday objects. Our goal is to encourage creativity, expressiveness and playfulness. The user can assemble his or her hand crafted creations with HandiMate--s joint modules and animate them via gestures. The joint modules are packaged with an actuator, a wireless communication device and a micro-controller. This modularization makes quick electro-mechanical prototyping just a matter of pressing together velcro. Animating these constructions is made intuitive and simple by a glove-based gestural controller. Our study conducted with children and adults demonstrates a high level of usability (system usability score 79.9). It also indicates that creative ideas emerge and are realized in a constructive and iterative manner in less than 90 minutes. This paper describes the design goals, framework, interaction methods, sample creations and evaluations methods.
Jasjeet Singh Seehra, Ansh Verma, Kylie Peppler, Karthik Ramani
TEI4
2015 TIMMi: Finger-worn Textile Input Device with Multimodal Sensing in Mobile Interaction
abstract
We introduce TIMMi, a textile input device for mobile interactions. TIMMi is worn on the index finger to provide a multimodal sensing input metaphor. The prototype is fabricated on a single layer of textile where the conductive silicone rubber is painted and the conductive threads are stitched. The sensing area comprises of three equally spaced dots and a separate wide line. Strain and pressure values are extracted from the line and three dots, respectively via voltage dividers. Regression analysis is performed to model the relationship between sensing values and finger pressure and bending. A multi-level thresholding is applied to capture different levels of finger bending and pressure. A temporal position tracking algorithm is implemented to capture the swipe gesture. In this preliminary study, we demonstrate TIMMi as a finger-worn input device with two applications: controlling music player and interacting with smartglasses.
Sang Ho Yoon, Ke Huo, Vinh P. Nguyen, Karthik Ramani
TEI4
2015 RevoMaker: Enabling Multi-directional and Functionally-embedded 3D printing using a Rotational Cuboidal Platform
abstract
In recent years, 3D printing has gained significant attention from the maker community, academia, and industry to support low-cost and iterative prototyping of designs. Current unidirectional extrusion systems require printing sacrificial material to support printed features such as overhangs. Furthermore, integrating functions such as sensing and actuation into these parts requires additional steps and processes to create "functional enclosures", since design functionality cannot be easily embedded into prototype printing. All of these factors result in relatively high design iteration times. We present "RevoMaker", a self-contained 3D printer that creates direct out-of-the-printer functional prototypes, using less build material and with substantially less reliance on support structures. By modifying a standard low-cost FDM printer with a revolving cuboidal platform and printing partitioned geometries around cuboidal facets, we achieve a multidirectional additive prototyping process to reduce the print and support material use. Our optimization framework considers various orientations and sizes for the cuboidal base. The mechanical, electronic, and sensory components are preassembled on the flattened laser-cut facets and enclosed inside the cuboid when closed. We demonstrate RevoMaker directly printing a variety of customized and fully-functional product prototypes, such as computer mice and toys, thus illustrating the new affordances of 3D printing for functional product design.
Diogo C. Nazzetta, Karthik Ramani, Raymond J. Cipra
UIST4
2015 The status, challenges, and future of additive manufacturing in engineering
Devarajan Ramanujan, Karthik Ramani, Yong Chen 0017, Christopher Williams 0002, Charlie C. L. Wang, Yung C. Shin, Song Zhang 0002, Pablo D. Zavattieri
Comput. Aided Des.4
2015 A gesture-free geometric approach for mid-air expression of design intent in 3D virtual pottery
Vinayak R. Krishnamurthy, Karthik Ramani
Comput. Aided Des.2
2015 Sketcholution: Interaction histories for sketching
Zhenpeng Zhao, William Benjamin, Niklas Elmqvist, Karthik Ramani
Int. J. Hum. Comput. Stud.4
2014 ChiroBot: modularrobotic manipulation via spatial hand gestures
abstract
We introduce ChiroBot, a cyberphysical construction kit that allows users to create custom robots out of craft material, easily assemble the robots using joint modules and control them using hand gestures. These handcrafted robots are assembled using our modules packaged with actuator, wireless communication and controller electronics. These modules eliminate the need for expertise in electronics and enable a plug and play system that directly encourages users to explore by quick prototyping. We designed a glove embedded with sensors to enable the user to control the robots using hand gestures. We present different usage scenarios to demonstrate the system's versatility such as vehicular robot, humanoid puppet, robotic arm, and other combinations. This paper describes the ChiroBot system, interaction methods, few sample creations, and proposes possible "play value".
Jasjeet Singh Seehra, Ansh Verma, Karthik Ramani
IDC3
2014 Tracing and sketching performance using blunt-tipped styli on direct-touch tablets
abstract
Direct-touch tablets are quickly replacing traditional pen-and-paper tools in many applications, but not in case of the designer's sketchbook. In this paper, we explore the tradeoffs inherent in replacing such paper sketchbooks with digital tablets in terms of two major tasks: tracing and free-hand sketching. Given the importance of the pen for sketching, we also study the impact of using a blunt-and-soft-tipped capacitive stylus in tablet settings. We thus conducted experiments to evaluate three sketch media: pen-paper, finger-tablet, and stylus-tablet based on the above tasks. We analyzed the tracing data with respect to speed and accuracy, and the quality of the free-hand sketches through a crowdsourced survey. The pen-paper and stylus-tablet media both performed significantly better than the finger-tablet medium in accuracy, while the pen-paper sketches were significantly rated higher quality compared to both tablet interfaces. A follow-up study comparing the performance of this stylus with a sharp, hard-tip version showed no significant difference in tracing performance, though participants preferred the sharp tip for sketching.
Sriram Karthik Badam, Senthil K. Chandrasegaran, Niklas Elmqvist, Karthik Ramani
AVI4
2014 PuppetX: a framework for gestural interactions with user constructed playthings
abstract
We present PuppetX, a framework for both constructing playthings and playing with them using spatial body and hand gestures. This framework allows users to construct various playthings similar to puppets with modular components representing basic geometric shapes. It is topologically-aware, i.e. depending on its configuration; PuppetX automatically determines its own topological construct. Once the plaything is made the users can interact with them naturally via body and hand gestures as detected by depth-sensing cameras. This gives users the freedom to create playthings using our components and the ability to control them using full body interactions. Our framework creates affordances for a new variety of gestural interactions with physically constructed objects. As its by-product, a virtual 3D model is created, which can be animated as a proxy to the physical construct. Our algorithms can recognize hand and body gestures in various configurations of the playthings. Through our work, we push the boundaries of interaction with user-constructed objects using large gestures involving the whole body or fine gestures involving the fingers. We discuss the results of a study to understand how users interact with the playthings and conclude with a demonstration of the abilities of gestural interactions with PuppetX by exploring a variety of interaction scenarios.
Saikat Gupta, Sujin Jang, Karthik Ramani
AVI3
2014 Juxtapoze: supporting serendipity and creative expression in clipart compositions
abstract
Juxtapoze is a clipart composition workflow that supports creative expression and serendipitous discoveries in the shape domain. We achieve creative expression by supporting a workflow of searching, editing, and composing: the user queries the shape database using strokes, selects the desired search result, and finally modifies the selected image before composing it into the overall drawing. Serendipitous discovery of shapes is facilitated by allowing multiple exploration channels, such as doodles, shape filtering, and relaxed search. Results from a qualitative evaluation show that Juxtapoze makes the process of creating image compositions enjoyable and supports creative expression and serendipity.
William Benjamin, Senthil K. Chandrasegaran, Devarajan Ramanujan, Niklas Elmqvist, S. V. N. Vishwanathan, Karthik Ramani
CHI6
2014 skWiki: a multimedia sketching system for collaborative creativity
abstract
We present skWiki, a web application framework for collaborative creativity in digital multimedia projects, including text, hand-drawn sketches, and photographs. skWiki overcomes common drawbacks of existing wiki software by providing a rich viewer/editor architecture for all media types that is integrated into the web browser itself, thus avoiding dependence on client-side editors. Instead of files, skWiki uses the concept of paths as trajectories of persistent state over time. This model has intrinsic support for collaborative editing, including cloning, branching, and merging paths edited by multiple contributors. We demonstrate skWiki's utility using a qualitative, sketching-based user study.
Zhenpeng Zhao, Sriram Karthik Badam, Senthil K. Chandrasegaran, Deok Gun Park 0001, Niklas Elmqvist, Lorraine G. Kisselburgh, Karthik Ramani
CHI7
2014 BendID: flexible interface for localized deformation recognition
abstract
We present BendID, a bendable input device that recognizes the location, magnitude and direction of its deformation. We use BendID to provide users with a tactile metaphor for pressure based input. The device is constructed by layering an array of indium tin oxide (ITO)-coated PET film electrodes on a Polymethylsiloxane (PDMS) sheet, which is sandwiched between conductive foams. The pressure values that are interpreted from the ITO electrodes are classified using a Support Vector Machine (SVM) algorithm via the Weka library to identify the direction and location of bending. A polynomial regression model is also employed to estimate the overall magnitude of the pressure from the device. A model then maps these variables to a GUI to perform tasks. In this preliminary paper, we demonstrate this device by implementing it as an interface for 3D shape bending and a game controller.
Vinh P. Nguyen, Sang Ho Yoon, Ansh Verma, Karthik Ramani
UbiComp4
2014 Global Voting Model for Protein Function Prediction from Protein-Protein Interaction Networks
Yi Fang 0006, Mengtian Sun, Guoxian Dai, Karthik Ramani
ICIC (3)4
2014 The Intrinsic Geometric Structure of Protein-Protein Interaction Networks for Protein Interaction Prediction
Yi Fang 0006, Mengtian Sun, Guoxian Dai, Karthik Ramani
ICIC (3)4
2014 HexaMorph: A reconfigurable and foldable hexapod robot inspired by origami
abstract
Origami affords the creation of diverse 3D objects through explicit folding processes from 2D sheets of material. Originally as a paper craft from 17th century AD, origami designs reveal the rudimentary characteristics of sheet folding: it is lightweight, inexpensive, compact and combinatorial. In this paper, we present “HexaMorph”, a novel starfish-like hexapod robot designed for modularity, foldability and reconfigurability. Our folding scheme encompasses periodic foldable tetrahedral units, called “Basic Structural Units” (BSU), for constructing a family of closed-loop spatial mechanisms and robotic forms. The proposed hexapod robot is fabricated using single sheets of cardboard. The electronic and battery components for actuation are allowed to be preassembled on the flattened crease-cut pattern and enclosed inside when the tetrahedral modules are folded. The self-deploying characteristic and the mobility of the robot are investigated, and we discuss the motion planning and control strategies for its squirming locomotion. Our design and folding paradigm provides a novel approach for building reconfigurable robots using a range of lightweight foldable sheets.
Ke Huo, Jasjeet Singh Seehra, Karthik Ramani, Raymond J. Cipra
IROS4
2014 Multi-Scale Kernels Using Random Walks
abstract
Abstract We introduce novel multi‐scale kernels using the random walk framework and derive corresponding embeddings and pairwise distances. The fractional moments of the rate of continuous time random walk (equivalently diffusion rate) are used to discover higher order kernels (or similarities) between pair of points. The formulated kernels are isometry, scale and tessellation invariant, can be made globally or locally shape aware and are insensitive to partial objects and noise based on the moment and influence parameters. In addition, the corresponding kernel distances and embeddings are convergent and efficiently computable. We introduce dual Green's mean signatures based on the kernels and discuss the applicability of the multi‐scale distance and embedding. Collectively, we present a unified view of popular embeddings and distance metrics while recovering intuitive probabilistic interpretations on discrete surface meshes.
Ayan Sinha, Karthik Ramani
Comput. Graph. Forum2
2013 The evolution, challenges, and future of knowledge representation in product design systems
Senthil K. Chandrasegaran, Karthik Ramani, Ram D. Sriram, Imre Horváth, Alain Bernard, Ramy F. Harik
Comput. Aided Des.2
2013 Shape-It-Up: Hand gesture based creative expression of 3D shapes using intelligent generalized cylinders
Vinayak R. Krishnamurthy, Sundar Murugappan, Hairong Liu, Karthik Ramani
Comput. Aided Des.4
2012 Center-Shift: An approach towards automatic robust mesh segmentation (ARMS)
abstract
In the area of 3D shape analysis, research in mesh segmentation has always been an important topic, as it is a fundamental low-level task which can be utilized in many applications including computer-aided design, computer animation, biomedical applications and many other fields. We define the automatic robust mesh segmentation (ARMS) method in this paper, which 1) is invariant to isometric transformation, 2) is insensitive to noise and deformation, 3) performs closely to human perception, 4) is efficient in computation, and 5) is minimally dependent on prior knowledge. In this work, we develop a new framework, namely the Center-Shift, which discovers meaningful segments of a 3D object by exploring the intrinsic geometric structure encoded in the biharmonic kernel. Our Center-Shift framework has three main steps: First, we construct a feature space where every vertex on the mesh surface is associated with the corresponding biharmonic kernel density function value. Second, we apply the Center-Shift algorithm for initial segmentation. Third, the initial segmentation result is refined through an efficient iterative process which leads to visually salient segmentation of the shape. The performance of this segmentation method is demonstrated through extensive experiments on various sets of 3D shapes and different types of noise and deformation. The experimental results of 3D shape segmentation have shown better performance of Center-Shift, compared to state-of-the-art segmentation methods.
Mengtian Sun, Yi Fang 0006, Karthik Ramani
CVPR3
2012 Extended multitouch: recovering touch posture and differentiating users using a depth camera
abstract
Multitouch surfaces are becoming prevalent, but most existing technologies are only capable of detecting the user's actual points of contact on the surface and not the identity, posture, and handedness of the user. In this paper, we define the concept of extended multitouch interaction as a richer input modality that includes all of this information. We further present a practical solution to achieve this on tabletop displays based on mounting a single commodity depth camera above a horizontal surface. This will enable us to not only detect when the surface is being touched, but also recover the user's exact finger and hand posture, as well as distinguish between different users and their handedness. We validate our approach using two user studies, and deploy the technique in a scratchpad tool and in a pen + touch sketch tool.
Sundar Murugappan, Vinayak R. Krishnamurthy, Niklas Elmqvist, Karthik Ramani
UIST4
2012 3DMolNavi: A web-based retrieval and navigation tool for flexible molecular shape comparison
abstract
BACKGROUND: Many molecules of interest are flexible and undergo significant shape deformation as part of their function, but most existing methods of molecular shape comparison treat them as rigid shapes, which may lead to incorrect measure of the shape similarity of flexible molecules. Currently, there still is a limited effort in retrieval and navigation for flexible molecular shape comparison, which would improve data retrieval by helping users locate the desirable molecule in a convenient way. RESULTS: To address this issue, we develop a web-based retrieval and navigation tool, named 3DMolNavi, for flexible molecular shape comparison. This tool is based on the histogram of Inner Distance Shape Signature (IDSS) for fast retrieving molecules that are similar to a query molecule, and uses dimensionality reduction to navigate the retrieved results in 2D and 3D spaces. We tested 3DMolNavi in the Database of Macromolecular Movements (MolMovDB) and CATH. Compared to other shape descriptors, it achieves good performance and retrieval results for different classes of flexible molecules. CONCLUSIONS: The advantages of 3DMolNavi, over other existing softwares, are to integrate retrieval for flexible molecular shape comparison and enhance navigation for user's interaction. 3DMolNavi can be accessed via https://engineering.purdue.edu/PRECISE/3dmolnavi/index.html.
Yu-Shen Liu, Meng Wang 0001, Jean-Claude Paul, Karthik Ramani
BMC Bioinform.4
2012 Towards locally and globally shape-aware reverse 3D modeling
Manish Goyal 0001, Sundar Murugappan, Cecil Piya, William Benjamin, Yi Fang 0006, Min Liu 0018, Karthik Ramani
Comput. Aided Des.7
2011 Heat-mapping: A robust approach toward perceptually consistent mesh segmentation
abstract
3D mesh segmentation is a fundamental low-level task with applications in areas as diverse as computer vision, computer-aided design, bio-informatics, and 3D medical imaging. A perceptually consistent mesh segmentation (PCMS), as defined in this paper is one that satisfies 1) in-variance to isometric transformation of the underlying surface, 2) robust to the perturbations of the surface, 3) robustness to numerical noise on the surface, and 4) close conformation to human perception. We exploit the intelligence of the heat as a global structure-aware message on a meshed surface and develop a robust PCMS scheme, called Heat-Mapping based on the heat kernel. There are three main steps in Heat-Mapping. First, the number of the segments is estimated based on the analysis of the behavior of the Laplacian spectrum. Second, the heat center, which is defined as the most representative vertex on each segment, is discovered by a proposed heat center hunting algorithm. Third, a heat center driven segmentation scheme reveals the PCMS with a high consistency towards human perception. Extensive experimental results on various types of models verify the performance of Heat-Mapping with respect to the consistent segmentation of articulated bodies, the topological changes, and various levels of numerical noise.
Yi Fang 0006, Mengtian Sun, Minhyong Kim, Karthik Ramani
CVPR4
2011 sLLE: Spherical locally linear embedding with applications to tomography
abstract
The tomographic reconstruction of a planar object from its projections taken at random unknown view angles is a problem that occurs often in medical imaging. Therefore, there is a need to robustly estimate the view angles given random observations of the projections. The widely used locally linear embedding (LLE) technique provides nonlinear embedding of points on a flat manifold. In our case, the projections belong to a sphere. Therefore, we extend LLE and develop a spherical locally linear embedding (sLLE) algorithm, which is capable of embedding data points on a non-flat spherically constrained manifold. Our algorithm, sLLE, transforms the problem of the angle estimation to a spherically constrained embedding problem. It considers each projection as a high dimensional vector with dimensionality equal to the number of sampling points on the projection. The projections are then embedded onto a sphere, which parametrizes the projections with respect to view angles in a globally consistent manner. The image is reconstructed from parametrized projections through the inverse Radon transform. A number of experiments demonstrate that sLLE is particularly effective for the tomography application we consider. We evaluate its performance in terms of the computational efficiency and noise tolerance, and show that sLLE can be used to shed light on the other constrained applications of LLE.
Yi Fang 0006, Mengtian Sun, S. V. N. Vishwanathan, Karthik Ramani
CVPR4
2011 Ontology-based customer preference modeling for concept generation
Dongxing Cao, Zhanjun Li, Karthik Ramani
Adv. Eng. Informatics3
2011 Editorial for the special issue of information mining and retrieval in design
Ying Liu 0004, Chris A. McMahon, Karthik Ramani, Dirk Schaefer
Adv. Eng. Informatics3
2011 Knowledge-based part similarity measurement utilizing ontology and multi-criteria decision making technique
Duhwan Mun, Karthik Ramani
Adv. Eng. Informatics2
2011 Heat Walk: Robust Salient Segmentation of Non-rigid Shapes
abstract
Abstract Segmenting three dimensional objects using properties of heat diffusion on meshes aim to produce salient results. The few existing algorithms based on heat diffusion do not use the full knowledge that can be gained from heat diffusion and are sensitive to varying kinds of perturbations. Our simple algorithm, Heat Walk, converts the implicit information in the heat kernel to explicit knowledge about the pathways for maximum heat flow capacity. We develop a two stage strategy for segmentation. In the first stage we quickly identify regions which are dominated by heat accumulators by employing a greedy algorithm. The second stage partitions out dissipative regions from the previously discovered accumulative regions by using a KL‐divergence based criterion. The resulting algorithm is both independent of human intervention and fast because of the globally aware directed walk along the maximal heat flow capacity. Extensive experimental evidence shows the method is robust to a variety of noise factors including topological short circuits, surface holes, pose variations, variations in tessellation, missing features, scaling, as well as normal and shot noise. Comparison with the Princeton Segmentation Benchmark (PSB) shows that our method is comparable with state of the art segmentation methods and has additional advantages of being robust and self contained. Based upon theoretical insight the convergence and stability of the Heat Walk is shown.
William Benjamin, Andrew Wood Polk, S. V. N. Vishwanathan, Karthik Ramani
Comput. Graph. Forum4
2011 Computing the Inner Distances of Volumetric Models for Articulated Shape Description with a Visibility Graph
abstract
A new visibility graph-based algorithm is presented for computing the inner distances of a 3D shape represented by a volumetric model. The inner distance is defined as the length of the shortest path between landmark points within the shape. The inner distance is robust to articulation and can reflect the deformation of a shape structure well without an explicit decomposition. Our method is based on the visibility graph approach. To check the visibility between pairwise points, we propose a novel, fast, and robust visibility checking algorithm based on a clustering technique which operates directly on the volumetric model without any surface reconstruction procedure, where an octree is used for accelerating the computation. The inner distance can be used as a replacement for other distance measures to build a more accurate description for complex shapes, especially for those with articulated parts. The binary executable program for the Windows platform is available from https://engineering.purdue.edu/PRECISE/VMID.
Yu-Shen Liu, Karthik Ramani, Min Liu 0018
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Using diffusion distances for flexible molecular shape comparison
abstract
BACKGROUND: Many molecules are flexible and undergo significant shape deformation as part of their function, and yet most existing molecular shape comparison (MSC) methods treat them as rigid bodies, which may lead to incorrect shape recognition. RESULTS: In this paper, we present a new shape descriptor, named Diffusion Distance Shape Descriptor (DDSD), for comparing 3D shapes of flexible molecules. The diffusion distance in our work is considered as an average length of paths connecting two landmark points on the molecular shape in a sense of inner distances. The diffusion distance is robust to flexible shape deformation, in particular to topological changes, and it reflects well the molecular structure and deformation without explicit decomposition. Our DDSD is stored as a histogram which is a probability distribution of diffusion distances between all sample point pairs on the molecular surface. Finally, the problem of flexible MSC is reduced to comparison of DDSD histograms. CONCLUSIONS: We illustrate that DDSD is insensitive to shape deformation of flexible molecules and more effective at capturing molecular structures than traditional shape descriptors. The presented algorithm is robust and does not require any prior knowledge of the flexible regions.
Yu-Shen Liu, Guo-Qin Zheng, Karthik Ramani, William Benjamin
BMC Bioinform.4
2009 StreamRay: a stream filtering architecture for coherent ray tracing
abstract
The wide availability of commodity graphics processors has made real-time graphics an intrinsic component of the human/computer interface. These graphics cores accelerate the z-buffer algorithm and provide a highly interactive experience at a relatively low cost. However, many applications in entertainment, science, and industry require high quality lighting effects such as accurate shadows, reflection, and refraction. These effects can be difficult to achieve with z-buffer algorithms but are straightforward to implement using ray tracing. Although ray tracing is computationally more complex, the algorithm exhibits excellent scaling and parallelism properties. Nevertheless, ray tracing memory access patterns are difficult to predict and the parallelism speedup promise is therefore hard to achieve.
Karthik Ramani, Christiaan P. Gribble, Al Davis
ASPLOS1
2009 On minimal orthographic view covers for polyhedra
abstract
In this paper, we consider the external visibility coverage for polyhedra under the orthographic viewing model. The problem is to compute whether the whole boundary of a polyhedron is visible from a finite set of view directions, and if so, how to compute a minimal set of such view directions. A global visibility map based method is developed to calculate an optimal or near-optimal solution using object space segmentation and viewpoint space sampling. Our method subdivides the concave regions of a polyhedron into a ground set containing two types of segments: concave regions with nonempty global visibility maps and convex polygons in the concave regions whose global visibility maps are empty. The viewpoint space is sampled using a generate-as-required heuristic. The corresponding visibility matrix is computed based on global visibility calculation. Our problem is then modeled as an instance of the classical set-cover problem and solved using a minimal visible set based branch-and-bound algorithm.
Min Liu 0018, Karthik Ramani
Shape Modeling International2
2009 Using least median of squares for structural superposition of flexible proteins
abstract
BACKGROUND: The conventional superposition methods use an ordinary least squares (LS) fit for structural comparison of two different conformations of the same protein. The main problem of the LS fit that it is sensitive to outliers, i.e. large displacements of the original structures superimposed. RESULTS: To overcome this problem, we present a new algorithm to overlap two protein conformations by their atomic coordinates using a robust statistics technique: least median of squares (LMS). In order to effectively approximate the LMS optimization, the forward search technique is utilized. Our algorithm can automatically detect and superimpose the rigid core regions of two conformations with small or large displacements. In contrast, most existing superposition techniques strongly depend on the initial LS estimating for the entire atom sets of proteins. They may fail on structural superposition of two conformations with large displacements. The presented LMS fit can be considered as an alternative and complementary tool for structural superposition. CONCLUSION: The proposed algorithm is robust and does not require any prior knowledge of the flexible regions. Furthermore, we show that the LMS fit can be extended to multiple level superposition between two conformations with several rigid domains. Our fit tool has produced successful superpositions when applied to proteins for which two conformations are known. The binary executable program for Windows platform, tested examples, and database are available from https://engineering.purdue.edu/PRECISE/LMSfit.
Yu-Shen Liu, Yi Fang 0006, Karthik Ramani
BMC Bioinform.3
2009 IDSS: deformation invariant signatures for molecular shape comparison
abstract
BACKGROUND: Many molecules of interest are flexible and undergo significant shape deformation as part of their function, but most existing methods of molecular shape comparison (MSC) treat them as rigid bodies, which may lead to incorrect measure of the shape similarity of flexible molecules. RESULTS: To address the issue we introduce a new shape descriptor, called Inner Distance Shape Signature (IDSS), for describing the 3D shapes of flexible molecules. The inner distance is defined as the length of the shortest path between landmark points within the molecular shape, and it reflects well the molecular structure and deformation without explicit decomposition. Our IDSS is stored as a histogram which is a probability distribution of inner distances between all sample point pairs on the molecular surface. We show that IDSS is insensitive to shape deformation of flexible molecules and more effective at capturing molecular structures than traditional shape descriptors. Our approach reduces the 3D shape comparison problem of flexible molecules to the comparison of IDSS histograms. CONCLUSION: The proposed algorithm is robust and does not require any prior knowledge of the flexible regions. We demonstrate the effectiveness of IDSS within a molecular search engine application for a benchmark containing abundant conformational changes of molecules. Such comparisons in several thousands per second can be carried out. The presented IDSS method can be considered as an alternative and complementary tool for the existing methods for rigid MSC. The binary executable program for Windows platform and database are available from https://engineering.purdue.edu/PRECISE/IDSS.
Yu-Shen Liu, Yi Fang 0006, Karthik Ramani
BMC Bioinform.3
2009 Shape-based clustering for 3D CAD objects: A comparative study of effectiveness
Subramaniam Jayanti, Yagnanarayanan Kalyanaraman, Karthik Ramani
Comput. Aided Des.3
2009 Computing global visibility maps for regions on the boundaries of polyhedra using Minkowski sums
Min Liu 0018, Yu-Shen Liu, Karthik Ramani
Comput. Aided Des.3
2009 Robust principal axes determination for point-based shapes using least median of squares
Yu-Shen Liu, Karthik Ramani
Comput. Aided Des.2
2009 Editorial
Karthik Ramani, Zoltán Rusák
Comput. Aided Des.1
2008 Bio-geometry: challenges, approaches, and future opportunities in proteomics and drug discovery
abstract
Biology has been an experimental science until the recent prominence of Bioinformatics and Computational Biology. With the discovery of the DNA sequence the protein structure determination is now an emerging challenge. The protein structure is closely coupled to the function. Today, with the given increase in available computing power and no-cost storage, the ability to do computational experiments is emerging as a core competence necessary for rapid discovery in the future. The ability to include various complex physics such as electrostatics and hydrophobic interactions in realistic simulations has increased. The discovery of structure of proteins is the next frontier for a number of convergent areas in science. Hence the combination of geometry and physics becomes very critical to do realistic computational experiments.
Ruth Nussinov, Talapady Bhat, Jack Snoeyink, Karthik Ramani
Symposium on Solid and Physical Modeling5
2008 SHape REtrieval contest 2008: CAD models
abstract
This paper presents the summary of all the results of the participants in the event SHREC08 — CAD Model Track
M. Ramanathan 0001, Karthik Ramani
Shape Modeling International2
2008 Combinatorial synthesis approach employing graph networks
Offer Shai, Noel Titus, Karthik Ramani
Adv. Eng. Informatics3
2008 Structure-oriented contour representation and matching for engineering shapes
Suyu Hou, Karthik Ramani
Comput. Aided Des.2
2007 Application driven embedded system design: a face recognition case study
abstract
The key to increasing performance without a commensurate increase in power consumption in modern processors lies in increasing both parallelism and core specialization. Core specialization has been employed in the embedded space and is likely to play an important role in future heterogeneous multi-core architectures as well. In this paper, the face recognition application domain is employed as a case study to showcase an architectural design methodology which generates a specialized core with high performance and very low powercharacteristics. Specifically, we create "ASIC-like" execution flows to sustain the high memory parallelism generated within the core. The price of this benefit is a significant increase in compilation complexity. The crux of the problem is the need to co-schedule the often conflicting constraints of data access, data movement, and computation. A modular compiler approach that employs integer linear programming (ILP) based "interconnect-aware" instruction and data scheduling techniques to solve this problem is then described. The resulting core running the compiled code delivers a 1.65x throughput improvement over a high performance processor (Pentium 4) while simultaneously achieving an 80x energy-delay improvement over an energy-efficient processor (XScale) and performs real-time face recognition at embedded power budgets.
Karthik Ramani, Al Davis
CASES1
2007 Salient critical points for meshes
abstract
A novel method for extracting the salient critical points of meshes, possibly with noise, is presented by combining mesh saliency with Morse theory. In this paper, we use the idea of mesh saliency as a measure of regional importance for meshes. The proposed method defines the salient critical points in a scalar function space using a center-surround filter operator on Gaussian-weighted average of the scalar of vertices. Compared to using a purely geometric measure of shape, such as curvature, our method yields more satisfactory results with the lower number of critical points. We demonstrate the effectiveness of this approach by comparing our results with the results of the conventional approaches in a number of examples. Furthermore, this work has a variety of potential applications. We give a direct application to the hierarchical topological representation for meshes by combining the salient critical points with the Morse-Smale complex.
Yu-Shen Liu, Min Liu 0018, Daisuke Kihara, Karthik Ramani
Symposium on Solid and Physical Modeling4
2007 Computing an exact spherical visibility map for meshed polyhedra
abstract
This paper considers computation of the exact visibility range (or the spherical visibility map) for a closed polyhedron whose boundary is represented as a triangle mesh. For each facet on the mesh, we calculate the set of view directions from which all the points on the facet can be seen from the exterior. The projection of those visible directions onto the unit sphere forms the visibility map for the facet. We show that the exact visibility map is a spherical arrangement of closed 0-cells, 1-cells, and 2-cells embedded on the surface of the unit sphere. Based on a provable method for calculating the potential occlusion regions of a facet, a vector visibility algorithm is developed for computing the exact solution of the spherical visibility map for a facet. Examples are given to illustrate our algorithm.
Min Liu 0018, Karthik Ramani
Symposium on Solid and Physical Modeling2
2007 Anisotropic filtering on normal field and curvature tensor field using optimal estimation theory
abstract
In this paper, we study the problem of mesh denoising for improving the single pass surface estimation on normals and curvature tensors. We focus mainly on the engineering objects represented as dense triangle meshes. In particular, a two run non-linear diffusion algorithm based on optimal estimation theory is proposed to adaptively filter out the undesired discontinuities introduced by noise while preserving the underlying features. We show that the proposed filter can successfully improve the local surface estimates while preserving the desired features in terms of tangential and curvature discontinuities.
Min Liu 0018, Yu-Shen Liu, Karthik Ramani
Shape Modeling International3
2007 Classifier combination for sketch-based 3D part retrieval
Suyu Hou, Karthik Ramani
Comput. Graph.2
2006 Interconnect-Aware Coherence Protocols for Chip Multiprocessors
abstract
Improvements in semiconductor technology have made it possible to include multiple processor cores on a single die. Chip Multi-Processors (CMP) are an attractive choice for future billion transistor architectures due to their low design complexity, high clock frequency, and high throughput. In a typical CMP architecture, the L2 cache is shared by multiple cores and data coherence is maintained among private L1s. Coherence operations entail frequent communication over global on-chip wires. In future technologies, communication between different L1s will have a significant impact on overall processor performance and power consumption. On-chip wires can be designed to have different latency, bandwidth, and energy properties. Likewise, coherence protocol messages have different latency and bandwidth needs. We propose an interconnect composed of wires with varying latency, bandwidth, and energy characteristics, and advocate intelligently mapping coherence operations to the appropriate wires. In this paper, we present a comprehensive list of techniques that allow coherence protocols to exploit a heterogeneous interconnect and evaluate a subset of these techniques to show their performance and power-efficiency potential. Most of the proposed techniques can be implemented with a minimum complexity overhead.
Liqun Cheng, Naveen Muralimanohar, Karthik Ramani, Rajeev Balasubramonian, John B. Carter
ISCA3
2006 Power efficient resource scaling in partitioned architectures through dynamic heterogeneity
abstract
The ever increasing demand for high clock speeds and the desire to exploit abundant transistor budgets have resulted in alarming increases in processor power dissipation. Partitioned (or clustered) architectures have been proposed in recent years to address scalability concerns in future billion-transistor microprocessors. Our analysis shows that increasing processor resources in a clustered architecture results in a linear increase in power consumption, while providing diminishing improvements in single-thread performance. To preserve high performance to power ratios, we claim that the power consumption of additional resources should be in proportion to the performance improvements they yield. Hence, in this paper, we propose the implementation of heterogeneous clusters that have varying delay and power characteristics. A cluster's performance and power characteristic is tuned by scaling its frequency and novel policies dynamically assign frequencies to clusters, while attempting to either meet a fixed power budget or minimize a metric such as Energy /spl times/ Delay/sup 2/ (ED/sup 2/). By increasing resources in a power-efficient manner, we observe an 11% improvement in ED/sup 2/ and a 22.4% average reduction in peak temperature, when compared to a processor with homogeneous units. Our proposed processor model also provides strategies to handle thermal emergencies that have a relatively low impact on performance.
Naveen Muralimanohar, Karthik Ramani, Rajeev Balasubramonian
ISPASS2
2006 Developing an engineering shape benchmark for CAD models
Subramaniam Jayanti, Yagnanarayanan Kalyanaraman, Natraj Iyer, Karthik Ramani
Comput. Aided Des.4
2006 Automatic least-squares projection of points onto point clouds with applications in reverse engineering
Yu-Shen Liu, Jean-Claude Paul, Jun-Hai Yong, Pi-Qiang Yu, Hui Zhang 0013, Jia-Guang Sun 0001, Karthik Ramani
Comput. Aided Des.7
2006 On visual similarity based 2D drawing retrieval
Jiantao Pu, Karthik Ramani
Comput. Aided Des.2
2005 Microarchitectural Wire Management for Performance and Power in Partitioned Architectures
abstract
Future high-performance billion-transistor processors are likely to employ partitioned architectures to achieve high clock speeds, high parallelism, low design complexity, and low power. In such architectures, inter-partition communication over global wires has a significant impact on overall processor performance and power consumption. VLSI techniques allow a variety of wire implementations, but these wire properties have previously never been exposed to the microarchitecture. This paper advocates global wire management at the microarchitecture level and proposes a heterogeneous interconnect that is comprised of wires with varying latency, bandwidth, and energy characteristics. We propose and evaluate microarchitectural techniques that can exploit such a heterogeneous interconnect to improve performance and reduce energy consumption. These techniques include a novel cache pipeline design, the identification of narrow bit-width operands, the classification of non-critical data, and the detection of interconnect load imbalance. For a dynamically scheduled partitioned architecture, our results demonstrate that the proposed innovations result in up to 11% reductions in overall processor ED/sup 2/, compared to a baseline processor that employs a homogeneous interconnect.
Rajeev Balasubramonian, Naveen Muralimanohar, Karthik Ramani, Venkatanand Venkatachalapathy
HPCA3
2005 Three-dimensional shape searching: state-of-the-art review and future trends
Natraj Iyer, Subramaniam Jayanti, Kuiyang Lou, Yagnanarayanan Kalyanaraman, Karthik Ramani
Comput. Aided Des.5
2005 Shape-based searching for product lifecycle applications
Natraj Iyer, Subramaniam Jayanti, Kuiyang Lou, Yagnanarayanan Kalyanaraman, Karthik Ramani
Comput. Aided Des.5
2004 Content-based Three-dimensional Engineering Shape Search
abstract
We discuss the design and implementation of a prototype 3D engineering shape search system. The system incorporates multiple feature vectors, relevance feedback, and query by example and browsing, flexible definition of shape similarity, and efficient execution through multidimensional indexing and clustering. In order to offer more information for a user to determine similarity of 3D engineering shape, a 3D interface that allows users to manipulate shapes is proposed and implemented to present the search results. The system allows users to specify which feature vectors should be used to perform the search. The system is used to conduct extensive experimentation real data to test the effectiveness of various feature vectors for shape - the first such comparison of this type. The test results show that the descending order of the average precision of feature vectors is: principal moments, moment invariants, geometric parameters, and eigenvalues. In addition, a multistep similarity search strategy is proposed and tested to improve the effectiveness of 3D engineering shape search. It is shown that the multistep approach is more effective than the one-shot search approach, when a fixed number of shapes are retrieved.
Kuiyang Lou, Sunil Prabhakar 0001, Karthik Ramani
ICDE3