Dizhi Ma

dblp:277/5928 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0005-1711-4930ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2026 SketchConcept: Sketching-based Concept Composition for Product Design using Multimodal Large Language Model
abstract
Sketches are widely used in conceptual design to externalize early ideas and communicate intent. With the rise of generative AI, sketch-to-design workflows have advanced rapidly. However, sketches are limited for organizing component-level structure and intent: parts, functions, and relations are often implicit, making systematic design space exploration difficult. We present SketchConcept, a sketch-to-design system that enables multimodal exploration through sketching and language. It allows designers to sketch out the form, then use voice or text to articulate and refine component functions and structural organization. This enables designers to explore not only satisfying appearances, but also functional and structural alternatives that are essential for design. To support this workflow, SketchConcept introduces a function-to-visual mapping mechanism that connects visual components to functional properties for component-wise iteration. We demonstrate the system through a set of representative use cases and evaluate its efficacy and usability in a two-session user study.
Runlin Duan, Chenfei Zhu, Yuzhao Chen, Dizhi Ma, Jingyu Shi, Yichen Hu, Ziyi Liu 0004, Karthik Ramani
DIS4
2026 JustShape: Exploring Co-Speech Gestures for Multimodal LLM-Powered 3D Parametric Modeling
Runlin Duan, Yuzhao Chen, Yichen Hu, Ziyi Liu 0004, Chenfei Zhu, Xiyun Hu, Dizhi Ma, Karthik Ramani
CHI7
2026 ARify: Leveraging Narrated Instructional Videos to Create Augmented Reality Tutorials for Procedural Tasks
abstract
Augmented Reality (AR) tutorials enhance procedural task learning by providing situated, step-by-step guidance. Yet, creating such tutorials requires AR authoring expertise, posing a significant entry barrier. To lower this barrier, we introduce ARify, an authoring system that semi-automatically transforms narrated instructional videos into AR tutorials. To guide system design, we conducted a content analysis of video tutorials and derived a design space of instructional intents, tactics, and AR representations. Building on this, ARify generates AR tutorials by integrating a vision–language model to plan tutorial structures and an AR builder to configure AR representations, and offers interfaces that allow users to refine and customize the results. A numerical study on three machine tasks and a user study with 18 participants showed that ARify achieves promising performance across task types, and allows novices to author effective AR tutorials, validating its effectiveness and usability.
Xiyun Hu, Chenfei Zhu, Shao-Kang Hsia, Dizhi Ma, Rahul Jain 0018, Karthik Ramani
CHI4
2026 AgentCoach: LLM-Based Adaptive Coaching Feedback for Motor Skill Learning
abstract
We present AgentCoach, an LLM-powered system that provides adaptive feedback for motor skill learning from tutorial videos. The system works by extracting key coaching points (CPs) and compiling CP-specific evaluators that map each cue to measurable kinematic parameters. This process allows AgentCoach to connect high-level semantic meaning with low-level postural estimation for accurate, context-aware evaluation. During practice, learners receive concise visual diagnostics of their mistakes paired with prescriptive verbal feedback that adapts based on their performance history. We technically validate the CP extraction and evaluator compilation across a wide range of common sports and exercise videos. A user study confirms the system’s usability and shows the system’s potential effectiveness of its adaptive feedback across multiple skills.
Dizhi Ma, Jiakun Yu, Xiyun Hu, Liang He 0005, Sooyeon Jeong, Karthik Ramani
CHI1
2025 GesPrompt: Leveraging Co-Speech Gestures to Augment LLM-Based Interaction in Virtual Reality
abstract
Large Language Model (LLM)-based copilots have shown great potential in Extended Reality (XR) applications.However, the user faces challenges when describing the 3D environments to the copilots due to the complexity of conveying spatial-temporal information through text or speech alone.To address this, we introduce GesPrompt, a multimodal XR interface that combines co-speech gestures with speech, allowing end-users to communicate more naturally and accurately with LLM-based copilots in XR environments.By incorporating gestures, GesPrompt extracts spatial-temporal reference from co-speech gestures, reducing the need for precise textual prompts and minimizing cognitive load for end-users.Our contributions include (1) a workflow to integrate gesture and speech input in the XR environment, (2) a prototype VR system that implements the workflow, and (3) a user study demonstrating its effectiveness in improving user communication in VR environments.
Xiyun Hu, Dizhi Ma, Fengming He, Zhengzhe Zhu, Shao-Kang Hsia, Chenfei Zhu, Ziyi Liu 0004, Karthik Ramani
Conference on Designing Interactive Systems2
2024 avaTTAR: Table Tennis Stroke Training with Embodied and Detached Visualization in Augmented Reality
abstract
Table tennis stroke training is a critical aspect of player development. We designed a new augmented reality (AR) system, avaTTAR, for table tennis stroke training. The system provides both “on-body” (first-person view) and “detached” (third-person view) visual cues, enabling users to visualize target strokes and correct their attempts effectively with this dual perspectives setup. By employing a combination of pose estimation algorithms and IMU sensors, avaTTAR captures and reconstructs the 3D body pose and paddle orientation of users during practice, allowing real-time comparison with expert strokes. Through a user study, we affirm avaTTAR ’s capacity to amplify player experience and training results.
Dizhi Ma, Xiyun Hu, Jingyu Shi, Mayank Patel 0005, Rahul Jain 0018, Ziyi Liu 0004, Zhengzhe Zhu, Karthik Ramani
UIST1
2020 Flexible Spatial and Angular Light Field Super Resolution
abstract
A light field contains information in four dimensions, two spatial and two angular. Representing a light field by sampling it with a fixed number of pixels implies an inherent trade-off between angular resolution and spatial resolution- one apparently fixed at the time of capture. To enable flexible trade-offs in spatial and angular resolution after the fact, in this paper we apply techniques from super resolution in an integrated fashion. Our approach explores the similarity between light field super resolution (LFSR) and single image super resolution (SISR) and proposes a neural network framework that can carry out flexible super resolution tasks. We present concrete instances of the framework for center-view spatial LFSR, full-view spatial LFSR, and combined spatial and angular LFSR. Experiments with synthetic and real-world data sets show the center-view and full-views approaches outperform state-of-the-art spatial LFSR by over 1dB in PSNR and that the combined approach achieves comparable performance to state-of-the-art spatial LFSR algorithms. Visual results for images rendered from the combined approach show improved resolution of detail, without rendering artifacts.
Dizhi Ma, Andrew Lumsdaine, Wenhui Zhou 0001
ICIP1
2020 Fast and Efficient Neural Network for Light Field Disparity Estimation
abstract
As with many imaging tasks, disparity estimation for light fields seems to be well-matched to machine learning approaches. Neural network-based methods can achieve an overall bad pixel rate as low as four percent on the 4D light field benchmark dataset, continued effort to improve accuracy is resulting in diminishing returns. On the other hand, due to the growing importance of mobile and embedded devices, improving the efficiency is emerging as an important problem. In this paper, we improve the efficiency of existing neural net approaches for light field disparity estimation by introducing efficient network blocks, pruning redundant sections of the network and downsampling the resolution of feature vector. To improve performance, we also propose densely sampled epipolar image plane volumes as input. Experiment results show that our approach can achieve similar results compared with state-of-the-art methods while using only one-tenth runtime.
Dizhi Ma, Andrew Lumsdaine
ICPR1