VLDB 2026 Research / reviewers in the wild / expert
Weinan Shi
dblp:199/3083
· DBLP profile ↗
13ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-1351-9034ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 13 · 1 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TraceRing: Touchpad-like Pointing with a Single IMU Ring through Personalized LearningabstractAchieving touchpad-like pointing with a single IMU ring is highly desirable for portable and wearable interaction, yet challenging due to incomplete motion data and significant user variability. We present TraceRing, a finger-worn IMU system that enables precise two-dimensional cursor control. To address the limitations of generic end-to-end models, we propose a personalized training framework that learns user-specific representations through joint multi-task and contrastive learning, while dynamically selecting the most suitable expert model. This approach enables personalization without requiring per-user fine-tuning, and reduces velocity prediction error by 33.9% over state-of-the-art baselines. Furthermore, a real-time study shows it delivers speed and accuracy far exceeding those of AirMouse (2.26s v.s. 3.01s in average task completion time). These results demonstrate TraceRing as a portable and comfortable alternative for mobile computing and AR interaction applications. Weinan Shi, Zixuan Wang 0018, Suya Wu, Xiyuan Shen, Chengchi Zhou, Chun Yu, Yuanchun Shi |
CHI | 2 |
| 2026 | 3DRing: Enabling Low-Cost 3D Hand Position Tracking by Fusing Inertial and Low-Framerate Optical SensingabstractCurrent mobile hand tracking systems primarily rely on high-framerate (HFR) optical sensors to capture hand positions, resulting in high computational cost and limiting the applicability in end devices. We propose 3DRing, a 3D hand position tracking method that requires only low-framerate (LFR, <10 FPS) optical data and a single IMU ring. It consists of two stages: (1) a Deep Extended Kalman Filter module that predicts high-framerate hand positions from LFR optical measurements and a single IMU; (2) a Reinforcement Learning module that adaptively selects minimal keyframes for calibration, further reducing the average optical framerate. Using only 6.61 FPS optical data, 3DRing achieves an average real-time tracking error of 1.75 cm and an interaction efficiency of 86.0% in a 3D target selection task, compared to the 67 FPS hand tracking system of Meta Quest Pro, demonstrating a strong potential to reduce the reliance on optical data in mobile hand tracking tasks. Zhuojun Li, Chun Yu, Chang Liu 0150, Mingyuan Du, Weinan Shi, Yuanchun Shi |
CHI | 6 |
| 2026 | HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRIabstractLong-range Human-Robot Interaction (HRI) remains underexplored. Within it, Command Source Identification (CSI) — determining who issued a command — is especially challenging due to multi-user and distance-induced sensor ambiguity. We introduce HiSync, an optical-inertial fusion framework that treats hand motion as binding cues by aligning robot-mounted camera optical flow with hand-worn IMU signals. We first elicit a user-defined (N=12) gesture set and collect a multimodal command gesture dataset (N=38) in long-range multi-user HRI scenarios. Next, HiSync extracts frequency-domain hand motion features from both camera and IMU data, and a learned CSINet denoises IMU readings, temporally aligns modalities, and performs distance-aware multi-window fusion to compute cross-modal similarity of subtle, natural gestures, enabling robust CSI. In three-person scenes up to 34 m, HiSync achieves 92.32% CSI accuracy, outperforming the prior SOTA by 48.44%. HiSync is also validated on real-robot deployment. By making CSI reliable and natural, HiSync provides a practical primitive and design guidance for public-space HRI. Chun Yu, Borong Zhuang, Haopeng Jin, Qingyang Wan, Zhuojun Li, Zhoutong Ye, Chang Liu 0150, Weinan Shi, Yuanchun Shi |
CHI | 11 |
| 2025 | Investigating Context-Aware Collaborative Text Entry on Smartphones using Large Language ModelsabstractText entry is a fundamental and ubiquitous task, but users often face challenges such as situational impairments or difficulties in sentence formulation.Motivated by this, we explore the potential of large language models (LLMs) to assist with text entry in realworld contexts.We propose a collaborative smartphone-based text entry system, CATIA, that leverages LLMs to provide text suggestions based on contextual factors, including screen content, time, location, activity, and more.In a 7-day in-the-wild study with 36 participants, the system offered appropriate text suggestions in over 80% of cases.Users exhibited different collaborative behaviors depending on whether they were composing text for interpersonal communication or information services.Additionally, the relevance Yuanchun Shi, Weinan Shi, Meizhu Chen, Yeshuang Zhu, Jinchao Zhang 0001, Chun Yu |
CHI | 4 |
| 2025 | From Operation to Cognition: Automatic Modeling Cognitive Dependencies from User Demonstrations for GUI Task Automation
Yiwen Yin, Chun Yu, Toby Jia-Jun Li, Aamir Khan Jadoon, Sixiang Cheng, Weinan Shi, Mohan Chen 0007, Yuanchun Shi |
CHI | 7 |
| 2025 | InterQuest: A Mixed-Initiative Framework for Dynamic User Interest Modeling in Conversational Search
Yuanxi Wang, Qingyang Wan, Zhuojun Li, Chun Yu, Weinan Shi, Yuanchun Shi |
UIST | 7 |
| 2025 | Prompt2Task: Automating UI Tasks on Smartphones from Textual PromptsabstractUI task automation enables efficient task execution by simulating human interactions with GUIs, without modifying the existing application code. However, its broader adoption is constrained by the need for expertise in both scripting languages and workflow design. To address this challenge, we present Prompt2Task, a system designed to comprehend various task-related textual prompts (e.g., goals, procedures), thereby generating and performing the corresponding automation tasks. Prompt2Task incorporates a suite of intelligent agents that mimic human cognitive functions, specializing in interpreting user intent, managing external information for task generation, and executing operations on smartphones. The agents can learn from user feedback and continuously improve their performance based on the accumulated knowledge. Experimental results indicated a performance jump from a 22.28% success rate in the baseline to 95.24% with Prompt2Task, requiring an average of 0.69 user interventions for each new task. Prompt2Task presents promising applications in fields such as tutorial creation, smart assistance, and customer service. Tian Huang, Chun Yu, Weinan Shi, Zijian Peng, David Yang 0002, Yuanchun Shi |
ACM Trans. Comput. Hum. Interact. | 3 |
| 2024 | ContextCam: Bridging Context Awareness with Creative Human-AI Image Co-CreationabstractThe rapid advancement of AI-generated content (AIGC) promises to transform various aspects of human life significantly. This work particularly focuses on the potential of AIGC to revolutionize image creation, such as photography and self-expression. We introduce ContextCam, a novel human-AI image co-creation system that integrates context awareness with mainstream AIGC technologies like Stable Diffusion. ContextCam provides user’s image creation process with inspiration by extracting relevant contextual data, and leverages Large Language Model-based (LLM) multi-agents to co-create images with the user. A study with 16 participants and 136 scenarios revealed that ContextCam was well-received, showcasing personalized and diverse outputs as well as interesting user behavior patterns. Participants provided positive feedback on their engagement and enjoyment when using ContextCam, and acknowledged its ability to inspire creativity. Xianzhe Fan, Zihan Wu 0002, Chun Yu, Fenggui Rao, Weinan Shi, Teng Tu 0002 |
CHI | 5 |
| 2023 | From Gap to Synergy: Enhancing Contextual Understanding through Human-Machine Collaboration in Personalized SystemsabstractThis paper presents LangAware, a collaborative approach for constructing personalized context for context-aware applications. The need for personalization arises due to significant variations in context between individuals based on scenarios, devices, and preferences. However, there is often a notable gap between humans and machines in the understanding of how contexts are constructed, as observed in trigger-action programming studies such as IFTTT. LangAware enables end-users to participate in establishing contextual rules in-situ using natural language. The system leverages large language models (LLMs) to semantically connect low-level sensor detectors to high-level contexts and provide understandable natural language feedback for effective user involvement. We conducted a user study with 16 participants in real-life settings, which revealed an average success rate of 87.50% for defining contextual rules in a variety of 12 campus scenarios, typically accomplished within just two modifications. Furthermore, users reported a better understanding of the machine’s capabilities by interacting with LangAware. Chun Yu, Lichen Yang, Weinan Shi, Yuanchun Shi |
UIST | 7 |
| 2019 | VIPBoard: Improving Screen-Reader Keyboard for Visually Impaired People with Character-Level Auto CorrectionabstractModern touchscreen keyboards are all powered by the word-level auto-correction ability to handle input errors. Unfortunately, visually impaired users are deprived of such benefit because a screen-reader keyboard offers only character-level input and provides no correction ability. In this paper, we present VIPBoard, a smart keyboard for visually impaired people, which aims at improving the underlying keyboard algorithm without altering the current input interaction. Upon each tap, VIPBoard predicts the probability of each key considering both touch location and language model, and reads the most likely key, which saves the calibration time when the touchdown point misses the target key. Meanwhile, the keyboard layout automatically scales according to users' touch point location, which enables them to select other keys easily. A user study shows that compared with the current keyboard technique, VIPBoard can reduce touch error rate by 63.0% and increase text entry speed by 12.6%. Weinan Shi, Chun Yu, Shuyi Fan, Xin Yi 0001, Xiaojun Bi 0001, Yuanchun Shi |
CHI | 1 |
| 2018 | Lip-Interact: Improving Mobile Device Interaction with Silent Speech CommandsabstractWe present Lip-Interact, an interaction technique that allows users to issue commands on their smartphone through silent speech. Lip-Interact repurposes the front camera to capture the user's mouth movements and recognize the issued commands with an end-to-end deep learning model. Our system supports 44 commands for accessing both system-level functionalities (launching apps, changing system settings, and handling pop-up windows) and application-level functionalities (integrated operations for two apps). We verify the feasibility of Lip-Interact with three user experiments: evaluating the recognition accuracy, comparing with touch on input efficiency, and comparing with voiced commands with regards to personal privacy and social norms. We demonstrate that Lip-Interact can help users access functionality efficiently in one step, enable one-handed input when the other hand is occupied, and assist touch to make interactions more fluent. Ke Sun 0003, Chun Yu, Weinan Shi, Yuanchun Shi |
UIST | 3 |
| 2017 | Word Clarity as a Metric in Sampling Keyboard Test SetsabstractTest sets play an essential role in evaluating text entry techniques. In this paper, we argue that in addition to the widely adopted metric of bigram representativeness and memorability, word clarity should also be considered as a metric when creating test sets from the target dataset. Word clarity quantifies the extent to which a word is likely to confuse with other words on a keyboard. We formally define word clarity, derive equations calculating it, and both theoretically and empirically show that word clarity has a significant effect on text entry performance: it can yield up to 26.4% difference in error rate, and 25% difference in input speed. We later propose a Pareto optimization method for sampling test sets with different sizes, which optimizes the word clarity and bigram representativeness, and memorability of the test set. The obtained test sets are published on the Internet. Xin Yi 0001, Chun Yu, Weinan Shi, Xiaojun Bi 0001, Yuanchun Shi |
CHI | 3 |
| 2017 | Is it too small?: Investigating the performances and preferences of users when typing on tiny QWERTY keyboards
Xin Yi 0001, Chun Yu, Weinan Shi, Yuanchun Shi |
Int. J. Hum. Comput. Stud. | 3 |