VLDB 2026 Research / reviewers in the wild / expert
Chenfei Zhu
dblp:41/9229
· DBLP profile ↗
9ranked-venue papers
1as first author
7since 2021 · last 2026
0009-0003-3408-2876ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SketchConcept: Sketching-based Concept Composition for Product Design using Multimodal Large Language ModelabstractSketches are widely used in conceptual design to externalize early ideas and communicate intent. With the rise of generative AI, sketch-to-design workflows have advanced rapidly. However, sketches are limited for organizing component-level structure and intent: parts, functions, and relations are often implicit, making systematic design space exploration difficult. We present SketchConcept, a sketch-to-design system that enables multimodal exploration through sketching and language. It allows designers to sketch out the form, then use voice or text to articulate and refine component functions and structural organization. This enables designers to explore not only satisfying appearances, but also functional and structural alternatives that are essential for design. To support this workflow, SketchConcept introduces a function-to-visual mapping mechanism that connects visual components to functional properties for component-wise iteration. We demonstrate the system through a set of representative use cases and evaluate its efficacy and usability in a two-session user study. Runlin Duan, Chenfei Zhu, Yuzhao Chen, Dizhi Ma, Jingyu Shi, Yichen Hu, Ziyi Liu 0004, Karthik Ramani |
DIS | 2 |
| 2026 | JustShape: Exploring Co-Speech Gestures for Multimodal LLM-Powered 3D Parametric Modeling
Runlin Duan, Yuzhao Chen, Yichen Hu, Ziyi Liu 0004, Chenfei Zhu, Xiyun Hu, Dizhi Ma, Karthik Ramani |
CHI | 5 |
| 2026 | ARify: Leveraging Narrated Instructional Videos to Create Augmented Reality Tutorials for Procedural TasksabstractAugmented Reality (AR) tutorials enhance procedural task learning by providing situated, step-by-step guidance. Yet, creating such tutorials requires AR authoring expertise, posing a significant entry barrier. To lower this barrier, we introduce ARify, an authoring system that semi-automatically transforms narrated instructional videos into AR tutorials. To guide system design, we conducted a content analysis of video tutorials and derived a design space of instructional intents, tactics, and AR representations. Building on this, ARify generates AR tutorials by integrating a vision–language model to plan tutorial structures and an AR builder to configure AR representations, and offers interfaces that allow users to refine and customize the results. A numerical study on three machine tasks and a user study with 18 participants showed that ARify achieves promising performance across task types, and allows novices to author effective AR tutorials, validating its effectiveness and usability. Xiyun Hu, Chenfei Zhu, Shao-Kang Hsia, Dizhi Ma, Rahul Jain 0018, Karthik Ramani |
CHI | 2 |
| 2026 | Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual CanvasabstractGenerative AI (GenAI) has significantly advanced the ease and flexibility of image creation. However, it remains a challenge to precisely control spatial compositions, including object arrangement and scene conditions. To bridge this gap, we propose Canvas3D, an interactive system leveraging a 3D engine to enable precise spatial manipulation for image generation. Upon user prompt, Canvas3D automatically converts textual descriptions into interactive objects within a 3D engine-driven virtual canvas, empowering direct and precise spatial configuration. These user-defined arrangements generate explicit spatial constraints that guide generative models in accurately reflecting user intentions in the resulting images. We conducted a closed-ended comparative study between Canvas3D and a baseline system, and an open-ended, free-form study to assess overall system usability. The results indicate that Canvas3D outperforms the baseline on spatial control, interactivity, and overall user experience. Yuzhao Chen, Runlin Duan, Rahul Jain 0018, Yichen Hu, Chenfei Zhu, Jingyu Shi, Karthik Ramani |
IUI | 5 |
| 2025 | DesignFromX: Empowering Consumer-Driven Design Space Exploration through Feature Composition of Referenced ProductsabstractGenerated Design SpaceIteration-1 Iteration-2 Iteration-3Figure 1: Exploring the design space of a desk using DesignFromX.The process begins with the user selecting a component from a reference product image-here, the legs of a wooden chair.The system identifies and suggests design features of the selected component.The user then composes this feature, in this case, the structural form, into a designated part of the desk.Based on the user-defined composition, a Generative AI model generates a new design space for the desk.In subsequent iterations, the user can incorporate additional design features from other reference products to further explore the design space while retaining the features of their previous selections. Runlin Duan, Chenfei Zhu, Yuzhao Chen, Yichen Hu, Jingyu Shi, Karthik Ramani |
Conference on Designing Interactive Systems | 2 |
| 2025 | GesPrompt: Leveraging Co-Speech Gestures to Augment LLM-Based Interaction in Virtual RealityabstractLarge Language Model (LLM)-based copilots have shown great potential in Extended Reality (XR) applications.However, the user faces challenges when describing the 3D environments to the copilots due to the complexity of conveying spatial-temporal information through text or speech alone.To address this, we introduce GesPrompt, a multimodal XR interface that combines co-speech gestures with speech, allowing end-users to communicate more naturally and accurately with LLM-based copilots in XR environments.By incorporating gestures, GesPrompt extracts spatial-temporal reference from co-speech gestures, reducing the need for precise textual prompts and minimizing cognitive load for end-users.Our contributions include (1) a workflow to integrate gesture and speech input in the XR environment, (2) a prototype VR system that implements the workflow, and (3) a user study demonstrating its effectiveness in improving user communication in VR environments. Xiyun Hu, Dizhi Ma, Fengming He, Zhengzhe Zhu, Shao-Kang Hsia, Chenfei Zhu, Ziyi Liu 0004, Karthik Ramani |
Conference on Designing Interactive Systems | 6 |
| 2025 | agentAR: Creating Augmented Reality Applications with Tool-Augmented LLM-based Autonomous Agents
Chenfei Zhu, Shao-Kang Hsia, Xiyun Hu, Ziyi Liu 0004, Jingyu Shi, Karthik Ramani |
UIST | 1 |
| 2013 | Statistical machine translation based text normalization with crowdsourcingabstractIn [1], we have proposed systems for text normalization based on statistical machine translation (SMT) methods which are constructed with the support of Internet users and evaluated those with French texts. Internet users normalize text displayed in a web interface in an annotation process, thereby providing a parallel corpus of normalized and non-normalized text. With this corpus, SMT models are generated to translate non-normalized into normalized text. In this paper, we analyze their efficiency for other languages. Additionally, we embedded the English annotation process for training data in Amazon Mechanical Turk and compare the quality of texts thoroughly annotated in our lab to those annotated by the Turkers. Finally, we investigate how to reduce the user effort by iteratively applying an SMT system to the next sentences to be edited, built from the sentences which have been annotated so far. Tim Schlippe, Chenfei Zhu, Daniel Lemcke, Tanja Schultz |
ICASSP | 2 |
| 2010 | Text normalization based on statistical machine translation and internet user support
Tim Schlippe, Chenfei Zhu, Jan Gebhardt, Tanja Schultz |
INTERSPEECH | 2 |