Shao-Kang Hsia

dblp:307/4928 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0001-7893-5613ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 4 since 2021
YearPublicationVenuePosition
2026 AssembleIt: Generating Adaptive On-Demand 3D Animations for Context-Aware Mechanical Assembly Guidance
abstract
Mechanical assembly instructions are commonly delivered through static manuals or fixed-sequence animations, which limit users’ ability to seek clarification, request partial explanations, or adapt guidance to their moment-to-moment needs during physical assembly. We present AssembleIt, an interactive system that generates on-demand 3D assembly animations and verbal explanations directly from natural language user queries, without relying on pre-authored instructional content. AssembleIt automatically derives a part dependency graph from CAD geometry using an Assembly-by-Disassembly strategy and uses this representation to generate query-driven, context-aware animations at runtime, rather than following a single predefined sequence. We evaluate AssembleIt through a controlled user study with 12 participants performing physical assembly tasks using real parts, comparing on-demand, query-driven animations against a static 3D animation baseline. Results indicate that on-demand animation generation supports flexible exploration and targeted clarification during assembly, highlighting the design potential of generative, dependency-driven instructional interfaces for hands-on mechanical tasks.
Mayank Patel 0005, Rahul Jain 0018, Asim Unmesh, Shao-Kang Hsia, Karthik Ramani
DIS4
2026 ARify: Leveraging Narrated Instructional Videos to Create Augmented Reality Tutorials for Procedural Tasks
abstract
Augmented Reality (AR) tutorials enhance procedural task learning by providing situated, step-by-step guidance. Yet, creating such tutorials requires AR authoring expertise, posing a significant entry barrier. To lower this barrier, we introduce ARify, an authoring system that semi-automatically transforms narrated instructional videos into AR tutorials. To guide system design, we conducted a content analysis of video tutorials and derived a design space of instructional intents, tactics, and AR representations. Building on this, ARify generates AR tutorials by integrating a vision–language model to plan tutorial structures and an AR builder to configure AR representations, and offers interfaces that allow users to refine and customize the results. A numerical study on three machine tasks and a user study with 18 participants showed that ARify achieves promising performance across task types, and allows novices to author effective AR tutorials, validating its effectiveness and usability.
Xiyun Hu, Chenfei Zhu, Shao-Kang Hsia, Dizhi Ma, Rahul Jain 0018, Karthik Ramani
CHI3
2025 GesPrompt: Leveraging Co-Speech Gestures to Augment LLM-Based Interaction in Virtual Reality
abstract
Large Language Model (LLM)-based copilots have shown great potential in Extended Reality (XR) applications.However, the user faces challenges when describing the 3D environments to the copilots due to the complexity of conveying spatial-temporal information through text or speech alone.To address this, we introduce GesPrompt, a multimodal XR interface that combines co-speech gestures with speech, allowing end-users to communicate more naturally and accurately with LLM-based copilots in XR environments.By incorporating gestures, GesPrompt extracts spatial-temporal reference from co-speech gestures, reducing the need for precise textual prompts and minimizing cognitive load for end-users.Our contributions include (1) a workflow to integrate gesture and speech input in the XR environment, (2) a prototype VR system that implements the workflow, and (3) a user study demonstrating its effectiveness in improving user communication in VR environments.
Xiyun Hu, Dizhi Ma, Fengming He, Zhengzhe Zhu, Shao-Kang Hsia, Chenfei Zhu, Ziyi Liu 0004, Karthik Ramani
Conference on Designing Interactive Systems5
2025 agentAR: Creating Augmented Reality Applications with Tool-Augmented LLM-based Autonomous Agents
Chenfei Zhu, Shao-Kang Hsia, Xiyun Hu, Ziyi Liu 0004, Jingyu Shi, Karthik Ramani
UIST2