Cheng Zhang 0022

dblp:82/6384-22 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0002-5079-5927ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 18 · 1 first-author · 15 since 2021Computer networks · 3 · 3 since 2021
YearPublicationVenuePosition
2026 WatchHand: Enabling Continuous Hand Pose Tracking On Off-the-Shelf Smartwatches
abstract
Tracking hand poses on wrist-wearables enables rich, expressive interactions, yet remains unavailable on commercial smartwatches, as prior implementations rely on external sensors or custom hardware, limiting their real-world applicability. To address this, we present WatchHand, the first continuous 3D hand pose tracking system implemented on off-the-shelf smartwatches using only their built-in speaker and microphone. WatchHand emits inaudible frequency-modulated continuous waves and captures their reflections from the hand. These acoustic signals are processed by a deep-learning model that estimates 3D hand poses for 20 finger joints. We evaluate WatchHand across diverse real-world conditions—multiple smartwatch models, wearing-hands, body postures, noise conditions, pose-variation protocols—and achieve a mean per-joint position error of 7.87 mm in cross-session tests with device remounting. Although performance drops for unseen users or gestures, the model adapts effectively with lightweight fine-tuning on small amounts of data. Overall, WatchHand lowers the barrier to smartwatch-based hand tracking by eliminating additional hardware while enabling robust, always-available interactions on millions of existing devices.
Chi-Jung Lee, Hohurn Jung, Tianhong Catherine Yu, Ian Oakley, Cheng Zhang 0022
CHI7
2026 μTouch: Enabling Accurate, Lightweight Self-Touch Sensing with Passive Magnets
abstract
Self-touch gestures (e.g., nuanced facial touches and subtle finger scratches) provide rich insights into human behaviors, from hygiene practices to health monitoring. However, existing approaches fall short in detecting such micro gestures due to their diverse movement patterns.This paper presents μTouch, a novel magnetic sensing platform for self-touch gesture recognition. μTouch features (1) a compact hardware design with low-power magnetometers and magnetic silicon, (2) a lightweight semi-supervised framework requiring minimal user data, and (3) an ambient field detection module to mitigate environmental interference. We evaluated μTouch in two representative applications in user studies with 11 and 12 participants. μTouch only requires three-second fine-tuning data for each gesture — new users need less than one minute before starting to use the system. μTouch can distinguish eight different face-touching behaviors with an average accuracy of 93.41%, and reliably detect body-scratch behaviors with an average accuracy of 94.63%. μTouch demonstrates accurate and robust sensing performance even after a month, showcasing its potential as a practical tool for hygiene monitoring and dermatological health applications.
Ke Li 0013, Jike Wang, Cheng Zhang 0022, Alanson Sample, Dongyao Chen
PerCom5
2025 SpellRing: Recognizing Continuous Fingerspelling in American Sign Language using a Ring
abstract
Fingerspelling is a critical part of American Sign Language (ASL) recognition and has become an accessible optional text entry method for Deaf and Hard of Hearing (DHH) individuals. In this paper, we introduce SpellRing, a single smart ring worn on the thumb that recognizes words continuously fingerspelled in ASL. SpellRing uses active acoustic sensing (via a microphone and speaker) and an inertial measurement unit (IMU) to track handshape and movement, which are processed through a deep learning algorithm using Connectionist Temporal Classification (CTC) loss. We evaluated the system with 20 ASL signers (13 fluent and 7 learners), using the MacKenzie-Soukoref Phrase Set of 1,164 words and 100 phrases. Offline evaluation yielded top-1 and top-5 word recognition accuracies of 82.45% (9.67%) and 92.42% (5.70%), respectively. In real-time, the system achieved a word error rate (WER) of 0.099 (0.039) on the phrases. Based on these results, we discuss key lessons and design implications for future minimally obtrusive ASL recognition wearables.
Hyunchul Lim, Nam Anh Dang, Dylan Lee, Tianhong Catherine Yu, Jane Lu, Franklin Mingzhe Li, Yiqi Jin, Yan Ma 0006, Xiaojun Bi 0001, François Guimbretière, Cheng Zhang 0022
CHI11
2025 Exploring the Impact of Emotional Voice Integration in Sign-to-Speech Translators for Deaf-to-Hearing Communication
abstract
Emotional voice communication plays a crucial role in effective daily interactions. Deaf and Hard of Hearing (DHH) individuals, who often have limited use of voice, rely on facial expressions to supplement sign language and convey emotions. However, in American Sign Language (ASL), facial expressions serve not only emotional purposes but also function as linguistic markers that can alter the meaning of signs. This dual role can often confuse non-signers when interpreting a signer's emotional state. In this paper, we present studies that: (1) confirm the challenges non-signers face when interpreting emotions from facial expressions in ASL communication, and (2) demonstrate how integrating emotional voice into translation systems can enhance hearing individuals' understanding of a signer's emotional intent. An online survey with 45 hearing participants (non-ASL signers) revealed frequent misinterpretations of signers' emotions when emotional and linguistic facial expressions were used simultaneously. The findings show that incorporating emotional voice into translation systems significantly improves emotion recognition by 32%. Additionally, follow-up survey with 48 DHH participants highlights design considerations for implementing emotional voice features, emphasizing the importance of emotional voice integration to bridge communication gaps between DHH and hearing communities.
Hyunchul Lim, Minghan Gao, Franklin Mingzhe Li, Nam Anh Dang, Ianip Sit, Michelle M. Olson, Cheng Zhang 0022
Proc. ACM Hum. Comput. Interact.7
2024 EchoWrist: Continuous Hand Pose Tracking and Hand-Object Interaction Recognition Using Low-Power Active Acoustic Sensing On a Wristband
abstract
Our hands serve as a fundamental means of interaction with the world around us. Therefore, understanding hand poses and interaction contexts is critical for human-computer interaction (HCI). We present EchoWrist, a low-power wristband that continuously estimates 3D hand poses and recognizes hand-object interactions using active acoustic sensing. EchoWrist is equipped with two speakers emitting inaudible sound waves toward the hand. These sound waves interact with the hand and its surroundings through reflections and diffractions, carrying rich information about the hand’s shape and the objects it interacts with. The information captured by the two microphones goes through a deep learning inference system that recovers hand poses and identifies various everyday hand activities. Results from the two 12-participant user studies show that EchoWrist is effective and efficient at tracking 3D hand poses and recognizing hand-object interactions. Operating at 57.9 mW, EchoWrist can continuously reconstruct 20 3D hand joints with MJEDE of 4.81 mm and recognize 12 naturalistic hand-object interactions with 97.6% accuracy.
Chi-Jung Lee, Devansh Agarwal, Tianhong Catherine Yu, Vipin Gunda, Oliver Lopez, James Kim, Sicheng Yin, Boao Dong, Ke Li 0013, Mose Sakashita, François Guimbretière, Cheng Zhang 0022
CHI13
2024 EyeEcho: Continuous and Low-power Facial Expression Tracking on Glasses
abstract
In this paper, we introduce EyeEcho, a minimally-obtrusive acoustic sensing system designed to enable glasses to continuously monitor facial expressions. It utilizes two pairs of speakers and microphones mounted on glasses, to emit encoded inaudible acoustic signals directed towards the face, capturing subtle skin deformations associated with facial expressions. The reflected signals are processed through a customized machine-learning pipeline to estimate full facial movements. EyeEcho samples at 83.3 Hz with a relatively low power consumption of 167mW. Our user study involving 12 participants demonstrates that, with just four minutes of training data, EyeEcho achieves highly accurate tracking performance across different real-world scenarios, including sitting, walking, and after remounting the devices. Additionally, a semi-in-the-wild study involving 10 participants further validates EyeEcho’s performance in naturalistic scenarios while participants engage in various daily activities. Finally, we showcase EyeEcho’s potential to be deployed on a commercial-off-the-shelf (COTS) smartphone, offering real-time facial expression tracking.
Ke Li 0013, Boao Chen, Mose Sakashita, François Guimbretière, Cheng Zhang 0022
CHI7
2024 Beyond-Voice: Towards Continuous 3D Hand Pose Tracking on Commercial Home Assistant Devices
abstract
The surging popularity of home assistants and their voice user interface (VUI) have made them an ideal central control hub for smart home devices. However, current form factors heavily rely on VUI, which poses accessibility and usability issues; some latest ones are equipped with additional cameras and displays, which are costly and raise privacy concerns. These concerns jointly motivate Beyond-Voice, a novel high-fidelity acoustic sensing system that allows commodity home assistant devices to track and reconstruct hand poses continuously. It transforms the home assistant into an active sonar system using its existing onboard microphones and speakers. We feed a high-resolution range profile to the deep learning model that can analyze the motions of multiple body parts and predict the 3D positions of 21 finger joints, bringing the granularity for acoustic hand tracking to the next level. It operates across different environments and users without the need for personalized training data. A user study with 11 participants in 3 different environments shows that Beyond-Voice can track joints with an average mean absolute error of 16.47mm without any training data provided by the testing subject.
Yin Li 0008, Rohan Reddy, Cheng Zhang 0022, Rajalakshmi Nandakumar
IPSN3
2024 Poster Abstract: Beyond-Voice - Towards Continuous 3D Hand Pose Tracking on Commercial Home Assistant Devices
abstract
The surging popularity of home assistants and their voice user interface (VUI) have made them an ideal central control hub for smart home devices. However, current form factors heavily rely on VUI, which poses accessibility and usability issues; some latest ones are equipped with additional cameras and displays, which are costly and raise privacy concerns. These concerns jointly motivate Beyond-Voice, a novel high-fidelity acoustic sensing system that allows commodity home assistant devices to track and reconstruct hand poses continuously. It transforms the device into an active sonar system using its existing onboard microphones and speakers. By feeding a high-resolution range profile to the deep learning model, we can localize 21 finger joints in 3D, bringing the granularity for acoustic hand tracking to the next level. A user study with 11 participants in 3 different environments shows that Beyond-Voice can track joints with an average mean absolute error of 16.47mm for unseen environments and users.
Yin Li 0008, Rohan Reddy, Cheng Zhang 0022, Rajalakshmi Nandakumar
IPSN3
2024 GazeTrak: Exploring Acoustic-based Eye Tracking on a Glass Frame
abstract
In this paper, we present GazeTrak, the first acoustic-based eye tracking system on glasses. Our system only needs one speaker and four microphones attached to each side of the glasses. These acoustic sensors capture the formations of the eyeballs and the surrounding areas by emitting encoded inaudible sound towards eyeballs and receiving the reflected signals. These reflected signals are further processed to calculate the echo profiles, which are fed to a customized deep learning pipeline to continuously infer the gaze position. In a user study with 20 participants, GazeTrak achieves an accuracy of 3.6° within the same remounting session and 4.9° across different sessions with a refreshing rate of 83.3 Hz and a power signature of 287.9 mW. Furthermore, we report the performance of our gaze tracking system fully implemented on an MCU with a low-power CNN accelerator (MAX78002). In this configuration, the system runs at up to 83.3 Hz and has a total power signature of 95.4 mW with a 30 Hz FPS.
Ke Li 0013, Boao Chen, Sicheng Yin, Saif Mahmud, Qikang Liang, François Guimbretière, Cheng Zhang 0022
MobiCom9
2024 SeamPose: Repurposing Seams as Capacitive Sensors in a Shirt for Upper-Body Pose Tracking
abstract
Seams are areas of overlapping fabric formed by stitching two or more pieces of fabric together in the cut-and-sew apparel manufacturing process. In SeamPose, we repurposed seams as capacitive sensors in a shirt for continuous upper-body pose estimation. Compared to previous all-textile motion-capturing garments that place the electrodes on the clothing surface, our solution leverages existing seams inside of a shirt by machine-sewing insulated conductive threads over the seams. The unique invisibilities and placements of the seams afford the sensing shirt to look and wear similarly as a conventional shirt while providing exciting pose-tracking capabilities. To validate this approach, we implemented a proof-of-concept untethered shirt with 8 capacitive sensing seams. With a 12-participant user study, our customized deep-learning pipeline accurately estimates the relative (to the pelvis) upper-body 3D joint positions with a mean per joint position error (MPJPE) of 6.0 cm. SeamPose represents a step towards unobtrusive integration of smart clothing for everyday pose estimation.
Tianhong Catherine Yu, Manru Mary Zhang, Peter He 0002, Chi-Jung Lee, Cassidy Cheesman, Saif Mahmud, François Guimbretière, Cheng Zhang 0022
UIST9
2023 ReMotion: Supporting Remote Collaboration in Open Space with Automatic Robotic Embodiment
abstract
Design activities, such as brainstorming or critique, often take place in open spaces combining whiteboards and tables to present artefacts. In co-located settings, peripheral awareness enables participants to understand each other’s locus of attention with ease. However, these spatial cues are mostly lost while using videoconferencing tools. Telepresence robots could bring back a sense of presence, but controlling them is distracting. To address this problem, we present ReMotion, a fully automatic robotic proxy designed to explore a new way of supporting non-collocated open-space design activities. ReMotion combines a commodity body tracker (Kinect) to capture a user’s location and orientation over a wide area with a minimally invasive wearable system (NeckFace) to capture facial expressions. Due to its omnidirectional platform, ReMotion embodiment can render a wide range of body movements. A formative evaluation indicated that our system enhances the sharing of attention and the sense of co-presence enabling seamless movement-in-space during a design review task.
Mose Sakashita, Michael Russo, Cheng Zhang 0022, Malte F. Jung, François Guimbretière
CHI6
2023 EchoSpeech: Continuous Silent Speech Recognition on Minimally-obtrusive Eyewear Powered by Acoustic Sensing
abstract
We present EchoSpeech, a minimally-obtrusive silent speech interface (SSI) powered by low-power active acoustic sensing. EchoSpeech uses speakers and microphones mounted on a glass-frame and emits inaudible sound waves towards the skin. By analyzing echos from multiple paths, EchoSpeech captures subtle skin deformations caused by silent utterances and uses them to infer silent speech. With a user study of 12 participants, we demonstrate that EchoSpeech can recognize 31 isolated commands and 3-6 figure connected digits with 4.5% (std 3.5%) and 6.1% (std 4.2%) Word Error Rate (WER), respectively. We further evaluated EchoSpeech under scenarios including walking and noise injection to test its robustness. We then demonstrated using EchoSpeech in demo applications in real-time operating at 73.3mW, where the real-time pipeline was implemented on a smartphone with only 1-6 minutes of training data. We believe that EchoSpeech takes a solid step towards minimally-obtrusive wearable SSI for real-life deployment.
Ke Li 0013, Yihong Hao, Zhengnan Lai, François Guimbretière, Cheng Zhang 0022
CHI7
2023 D-Touch: Recognizing and Predicting Fine-grained Hand-face Touching Activities Using a Neck-mounted Wearable
abstract
This paper presents D-Touch, a neck-mounted wearable sensing system that can recognize and predict how a hand touches the face. It uses a neck-mounted infrared camera (IR), which takes pictures of the head from the neck. These IR camera images are processed and used to train a deep-learning model to recognize and predict touch time and positions. The study showed D-Touch distinguished 17 Facial related Activity (FrA), including 11 face touch positions and 6 other activities, with over 92.1% accuracy and predict the hand-touching T-zone from other FrA activities with an accuracy of 82.12% within 150 ms after the hand appeared in the camera. A study with 10 participants conducted in their homes without any constraints on participants showed that D-Touch can predict the hand-touching T-zone from other FrA activities with an accuracy of 72.3% within 150 ms after the camera saw the hand. Based on the study results, we further discuss the opportunities and challenges of deploying D-Touch in real-world scenarios.
Hyunchul Lim, Samhita Pendyal, Jeyeon Jo, Cheng Zhang 0022
IUI5
2022 Understanding How People with Visual Impairments Take Selfies: Experiences and Challenges
abstract
Selfies are a pervasive form of communication in social media. While there has been some work on systems that guide people with visual impairments (PVI) in taking photos, nearly all has focused on using the camera on the back of the device. We do not know whether and how PVI take selfies. The aim of our work is to understand (1) PVI selfie-taking experiences and challenges, (2) what information do PVI need when taking selfies, and (3) what modalities do PVI prefer (e.g., tactile, verbal, or non-verbal audio) to support selfie-taking. To address this gap, we conducted interviews with 10 PVI. Our findings show that current selfie-taking applications do not provide enough assistance to meet the needs of PVI. We contribute design guidelines that researchers and designers can implement for creating accessible selfie-taking applications.
Ricardo E. Gonzalez, Paul Vermette, Cheng Zhang 0022, Keith Vertanen, Shiri Azenkot
ASSETS4
2022 EatingTrak: Detecting Fine-grained Eating Moments in the Wild Using a Wrist-mounted IMU
abstract
In this paper, we present EatingTrak, an AI-powered sensing system using a wrist-mounted inertial measurement unit (IMU) to recognize eating moments in a near-free-living semi-wild setup. It significantly improves the SOTA in time resolution using similar hardware on identifying eating moments, from over five minutes to three seconds. Different from prior work which directly learns from raw IMU data, it proposes intelligent algorithms which can estimate the arm posture in 3D in the wild and then learns the detailed eating moments from the series of estimated arm postures. To evaluate the system, we collected eating activity data from 9 participants in semi-wild scenarios for over 113 hours. Results showed that it was able to recognize eating moments at three time-resolutions: 3 seconds and 15 minutes with F-1 scores of 73.7% and 83.8%, respectively. EatingTrak would introduce new opportunities in sensing detailed eating behavior information requiring high time resolution, such as eating frequency, snack-taking, on-site behavior intervention. We also discuss the opportunities and challenges in deploying EatingTrak on commodity devices at scale.
Jihai Zhang 0002, Nitish Gade, Seyun Kim, Junchi Yan, Cheng Zhang 0022
Proc. ACM Hum. Comput. Interact.7
2021 TeethTap: Recognizing Discrete Teeth Gestures Using Motion and Acoustic Sensing on an Earpiece
abstract
Teeth gestures become an alternative input modality for different situations and accessibility purposes. In this paper, we present TeethTap, a novel eyes-free and hands-free input technique, which can recognize up to 13 discrete teeth tapping gestures. TeethTap adopts a wearable 3D printed earpiece with an IMU sensor and a contact microphone behind both ears, which works in tandem to detect jaw movement and sound data, respectively. TeethTap uses a support vector machine to classify gestures from noise by fusing acoustic and motion data, and implements K-Nearest-Neighbor (KNN) with a Dynamic Time Warping (DTW) distance measurement using motion data for gesture classification. A user study with 11 participants demonstrated that TeethTap could recognize 13 gestures with a real-time classification accuracy of 90.9% in a laboratory environment. We further uncovered the accuracy differences on different teeth gestures when having sensors on single vs. both sides. Moreover, we explored the activation gesture under real-world environments, including eating, speaking, walking and jumping. Based on our findings, we further discussed potential applications and practical challenges of integrating TeethTap into future devices.
Wei Sun 0050, Franklin Mingzhe Li, Benjamin Steeper, Songlin Xu, Feng Tian 0001, Cheng Zhang 0022
IUI6
2021 ThumbTrak: Recognizing Micro-finger Poses Using a Ring with Proximity Sensing
abstract
ThumbTrak is a novel wearable input device that recognizes 12 micro-finger poses in real-time. Poses are characterized by the thumb touching each of the 12 phalanges on the hand. It uses a thumb-ring, built with a flexible printed circuit board, which hosts nine proximity sensors. Each sensor measures the distance from the thumb to various parts of the palm or other fingers. ThumbTrak uses a support-vector-machine (SVM) model to classify finger poses based on distance measurements in real-time. A user study with ten participants showed that ThumbTrak could recognize 12 micro finger poses with an average accuracy of 93.6%. We also discuss potential opportunities and challenges in applying ThumbTrak in real-world applications.
Wei Sun 0050, Franklin Mingzhe Li, Congshu Huang, Zhenyu Lei 0005, Benjamin Steeper, Songyun Tao, Feng Tian 0001, Cheng Zhang 0022
MobileHCI8
2021 HandyTrak: Recognizing the Holding Hand on a Commodity Smartphone from Body Silhouette Images
abstract
Understanding which hand a user holds a smartphone with can help improve the mobile interaction experience. For instance, the layout of the user interface (UI) can be adapted to the holding hand. In this paper, we present HandyTrak, an AI-powered software system that recognizes the holding hand on a commodity smartphone using body silhouette images captured by the front-facing camera. The silhouette images are processed and sent to a customized user-dependent deep learning model (CNN) to infer how the user holds the smartphone (left, right, or both hands). We evaluated our system on each participant’s smartphone at five possible front camera positions in a user study with ten participants under two hand positions (in the middle and skewed) and three common usage cases (standing, sitting, and resting against a desk). The results showed that HandyTrak was able to continuously recognize the holding hand with an average accuracy of 89.03% (SD: 8.98%) at a 2 Hz sampling rate. We also discuss the challenges and opportunities to deploy HandyTrak on different commodity smartphones and potential applications in real-world scenarios.
Hyunchul Lim, Jessica Tweneboah, Cheng Zhang 0022
UIST4
2020 C-Face: Continuously Reconstructing Facial Expressions by Deep Learning Contours of the Face with Ear-mounted Miniature Cameras
abstract
C-Face (Contour-Face) is an ear-mounted wearable sensing technology that uses two miniature cameras to continuously reconstruct facial expressions by deep learning contours of the face. When facial muscles move, the contours of the face change from the point of view of the ear-mounted cameras. These subtle changes are fed into a deep learning model which continuously outputs 42 facial feature points representing the shapes and positions of the mouth, eyes and eyebrows. To evaluate C-Face, we embedded our technology into headphones and earphones. We conducted a user study with nine participants. In this study, we compared the output of our system to the feature points outputted by a state of the art computer vision library (Dlib) from a font facing camera. We found that the mean error of all 42 feature points was 0.77 mm for earphones and 0.74 mm for headphones. The mean error for 20 major feature points capturing the most active areas of the face was 1.43 mm for earphones and 1.39 mm for headphones. The ability to continuously reconstruct facial expressions introduces new opportunities in a variety of applications. As a demonstration, we implemented and evaluated C-Face for two applications: facial expression detection (outputting emojis) and silent speech recognition. We further discuss the opportunities and challenges of deploying C-Face in real-world applications.
Tuochao Chen, Benjamin Steeper, Kinan Alsheikh, Songyun Tao, François Guimbretière, Cheng Zhang 0022
UIST6
2011 T-Maze: a tangible programming tool for children
abstract
This paper presents a tangible programming tool 'T-Maze' for children aged 5 to 9. Children could use T-Maze to create their own maze maps and complete some maze escaping tasks by the tangible programming blocks and sensors. T-Maze uses a camera to, in real-time, catch the programming sequence of the wooden blocks' arrangement, which will be used to analyze the semantic correctness and enable the children to receive feedbacks immediately. And children could join in the game by controlling the sensors during program's running. A user study shows that T-Maze is an interesting programming approach for children and easy to learn and use.
Danli Wang, Cheng Zhang 0022, Hongan Wang
IDC2
2011 CoolMag: a tangible interaction tool to customize instruments for children in music education
abstract
In this paper, we describe CoolMag, a tangible interaction tool to enable children to create different instruments collaboratively in music education. With CoolMag, children could learn the basic playing methods of different instruments. It also has the potential to inspire children's creativity, because children could adopt objects in daily life (broom, cup, pen etc.) as the carrier of their novel instruments whose appearance may differ from the traditional one.
Cheng Zhang 0022, Danli Wang, Feng Tian 0001, Hongan Wang
UbiComp1