Yuanchun Shi

dblp:08/5313 · DBLP profile ↗
← Back
195ranked-venue papers
3as first author
77since 2021 · last 2026
0000-0003-2273-6927ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 123 · 2 first-author · 62 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 16 · 8 since 2021Systems, architecture and hardware · 15Applied, interdisciplinary, general and emerging computing · 11 · 2 since 2021Computer networks · 4Software engineering, systems software and programming languages · 4Databases, data management, data science and information retrieval · 4Security and privacy · 3
YearPublicationVenuePosition
2026 SituFont: A Just-in-Time Adaptive Intervention Interface for Enhancing Mobile Readability in Situational Visual Impairments
abstract
Situational visual impairments (SVIs) hinder mobile readability, causing discomfort and limiting information access. Building on prior work in adaptive typography and accessibility, this paper presents SituFont, a context-aware and human-in-the-loop adaptive typography adjustment approach that enhances smartphone mobile readability by dynamically adjusting font parameters based on real-time contextual changes. Using smartphone sensors and a human-in-the-loop approach, SituFont personalizes text presentation to accommodate personal factors (e.g., fatigue, distraction) and environmental conditions (e.g., lighting, motion, location). To inform its design, we conducted formative interviews (N=15) to identify key SVI factors and controlled experiments (N=18) to quantify their impact on optimal text parameters. A comparative user study (N=12) across eight simulated SVI scenarios demonstrated SituFont’s effectiveness in improving smartphone mobile readability in terms of improved efficiency and reduced workload compared with a non-trivial manual adjustment baseline.
Jingruo Chen, Kexin Nie, Mingshan Zhang, Chun Yu, Zhiqi Gao, Kun Yue, Yuanchun Shi
CHI7
2026 TraceRing: Touchpad-like Pointing with a Single IMU Ring through Personalized Learning
abstract
Achieving touchpad-like pointing with a single IMU ring is highly desirable for portable and wearable interaction, yet challenging due to incomplete motion data and significant user variability. We present TraceRing, a finger-worn IMU system that enables precise two-dimensional cursor control. To address the limitations of generic end-to-end models, we propose a personalized training framework that learns user-specific representations through joint multi-task and contrastive learning, while dynamically selecting the most suitable expert model. This approach enables personalization without requiring per-user fine-tuning, and reduces velocity prediction error by 33.9% over state-of-the-art baselines. Furthermore, a real-time study shows it delivers speed and accuracy far exceeding those of AirMouse (2.26s v.s. 3.01s in average task completion time). These results demonstrate TraceRing as a portable and comfortable alternative for mobile computing and AR interaction applications.
Weinan Shi, Zixuan Wang 0018, Suya Wu, Xiyuan Shen, Chengchi Zhou, Chun Yu, Yuanchun Shi
CHI8
2026 Does Personalized Nudging Wear Off? A Longitudinal Study of AI Self-Modeling for Behavioral Engagement
abstract
Sustaining the effectiveness of behavior change technologies remains a key challenge. AI self-modeling, which generates personalized portrayals of one’s ideal self, has shown promise for motivating behavior change, yet prior work largely examines short-term effects. We present one of the first longitudinal evaluations of AI self-modeling in fitness engagement through a two-stage empirical study. A 1-week, three-arm experiment (visual self-modeling (VSM), auditory self-modeling (ASM), Control; N=28) revealed that VSM drove initial performance gains, while ASM showed no significant effects. A subsequent 4-week study (VSM vs. Control; N=31) demonstrated that VSM sustained higher performance levels but exhibited diminishing improvement rates after two weeks. Interviews uncovered a catalyst effect that fostered early motivation through clear, attainable goals, followed by habituation and internalization which stabilized performance. These findings highlight the temporal dynamics of personalized nudging and inform the design of behavior change technologies for long-term engagement.
Yuzhou Du, Jiahuan Ding, Yuanchun Shi, Yuntao Wang 0001
CHI5
2026 3DRing: Enabling Low-Cost 3D Hand Position Tracking by Fusing Inertial and Low-Framerate Optical Sensing
abstract
Current mobile hand tracking systems primarily rely on high-framerate (HFR) optical sensors to capture hand positions, resulting in high computational cost and limiting the applicability in end devices. We propose 3DRing, a 3D hand position tracking method that requires only low-framerate (LFR, <10 FPS) optical data and a single IMU ring. It consists of two stages: (1) a Deep Extended Kalman Filter module that predicts high-framerate hand positions from LFR optical measurements and a single IMU; (2) a Reinforcement Learning module that adaptively selects minimal keyframes for calibration, further reducing the average optical framerate. Using only 6.61 FPS optical data, 3DRing achieves an average real-time tracking error of 1.75 cm and an interaction efficiency of 86.0% in a 3D target selection task, compared to the 67 FPS hand tracking system of Meta Quest Pro, demonstrating a strong potential to reduce the reliance on optical data in mobile hand tracking tasks.
Zhuojun Li, Chun Yu, Chang Liu 0150, Mingyuan Du, Weinan Shi, Yuanchun Shi
CHI7
2026 Routine Computing: A Systematic Review of Sensing Daily Life Dimensions Towards Human-Centered Goals
abstract
Human routines structure daily life, yet remain challenging for computational systems to understand. This paper presents the first systematic review of routine computing, a previously implicit but increasingly recognized field that focuses on computationally sensing and modeling human behaviors. It synthesizes 203 studies published up to August 2025. The paper presents a new taxonomy of the literature, focusing on temporal structures, behavioral interactions, cognitive aspects, and how variability and deviations are addressed. The common goals of routine computing extend across four major application domains, including accessibility care, the promotion of healthy habits, adaptive and context-aware support, and large-scale population insights. Persistent challenges that limit the design of truly human-centered systems are identified, including the gap between low-level activity recognition and high-level intent, the tension between personalization and generalization, unresolved privacy concerns, and data-related limitations. By consolidating these findings, this paper provides a foundational framework for HCI researchers, outlining principles for designing ethical, adaptive, and human-centered routine-aware systems.
Borislav Pavlov, Yuntao Wang 0001, Yuanchun Shi
CHI5
2026 ActivitySeeker: Towards Collaborative Personalized Human Activity Discovery and Recognition on Smartphones
abstract
Smartphones provide an attractive yet challenging platform for human activity recognition (HAR). They are ubiquitous, but also limit the input of HAR systems to a single IMU. These systems are also challenged by the inherent diversity of human activities and varying phone placement on the user’s body. This results in traditional smartphone HAR systems having limited personalization potential or imposing a high user burden. We propose ActivitySeeker, a personalized smartphone HAR system that combines self-supervised activity discovery and low-burden user interaction to collaboratively label IMU data and adapt HAR models to individual users on-device through transfer learning. We evaluated ActivitySeeker through simulated online learning and in-the-wild user experiments, where it discovered 95.5% of personal activity types and achieved high recognition accuracy (93.3%) while maintaining a positive user experience. Leveraging the synergy between user and smartphone, ActivitySeeker opens up new possibilities for HAR-based applications like fitness, health and personalized recommendation.
Zhoutong Ye, Yanwen Huang, Chun Yu, Yuntao Wang 0001, Yuqi Luo, Yuanchun Shi
CHI6
2026 GazeCoT: Unleashing Social Intelligence in Multimodal LLMs With Gaze-Informed Chain-of-Thought Reasoning
abstract
Social intelligence is vital for effective human-AI interaction. While LLMs demonstrate strong text-based social intelligence, the vision modality remains challenging due to the presence of non-verbal social cues. For example, gaze is the primary conveyor of social attention, yet it cannot be accurately perceived and understood by multimodal LLMs (MLLMs). Therefore, we propose GazeCoT, a pipeline using gaze estimation models to provide MLLMs with the attention of people in images or videos. The gaze information is provided as visual and text prompts compiled into a structured context to support MLLM social reasoning. Benchmark evaluation confirms that GazeCoT enhances MLLMs’ social intelligence by improving gaze perception. A user study in a challenging application involving parent-child interactions demonstrates that GazeCoT improves perceived explainability and trustworthiness by aligning MLLM social perception and social reasoning with human norms. We hope that GazeCoT, a versatile plug-and-play pipeline, can enable socially aware, MLLM-based HCI applications.
Zhoutong Ye, Xutong Wang, Ruiwen Zhang, Qinwei Li, Chun Yu, Yuanchun Shi
CHI8
2026 Proactive AI as a Catalyst for Creativity? Balancing Human Agency and AI Contribution in Collaborative Story Writing
abstract
Large Language Models (LLMs) hold promise in supporting creative writing, yet the role of proactive AI in collaborative writing remains underexplored due to concerns around human agency and disruption. To investigate effective strategies for proactive AI support, we conducted a Wizard-of-Oz study simulating two suggestion styles: intrusive suggestions (next-sentence completions) and non-intrusive suggestions (exploratory proposals), where participants completed two story outlining tasks under each style, receiving real-time proactive suggestions from a human wizard acting as the AI. Both quantitative and qualitative results show that proactive AI can enhance creativity and accelerate writing. However, we observed a trade-off between AI involvement and perceived human agency. This trade-off was moderated by how strongly AI stimulated users–greater inspiration led to stronger perceived agency even under high AI involvement. Based on wizards’ behavior, we offer guidance on suggestion style and timing to better balance creativity and agency for future proactive AI writing systems.
Yiwen Yin, Mingze Wu, Ruijie Huang, Chun Yu, Yuanchun Shi
CHI7
2026 Enabling Adaptive Cardio-Respiratory Biofeedback Training on Ubiquitous Hand-Worn Devices
abstract
We introduce an adaptive cardio-respiratory biofeedback system implemented on ubiquitous hand-worn devices such as smart watches and rings, enabling accessible and real-time physiological training outside clinical settings. Users place a hand on their abdomen to promote embodied awareness of breathing rhythms, while PPG and IMU sensors continuously capture cardio-respiratory signals. Unlike conventional open-loop biofeedback that delivers fixed breathing guidance irrespective of user response, our system employs a closed-loop adaptation: real-time physiological signals adjust breathing cues to optimize cardio-respiratory coupling, ensuring personalized training trajectories. This shift from static to adaptive guidance markedly improves user engagement and training efficacy. A user performance evaluation study further showed that adaptive biofeedback significantly boosts HRV, prolongs high-HRV states, and enhances user experience, demonstrating clear advantages over non-adaptive methods. Together, these findings position adaptive, hand-worn biofeedback as a promising approach for ubiquitous, user-centered mental health interventions.
Ruotong Yu, Xintong Wu, Lily Sheng, Yuntao Wang 0001, Yuanchun Shi
CHI6
2026 HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI
abstract
Long-range Human-Robot Interaction (HRI) remains underexplored. Within it, Command Source Identification (CSI) — determining who issued a command — is especially challenging due to multi-user and distance-induced sensor ambiguity. We introduce HiSync, an optical-inertial fusion framework that treats hand motion as binding cues by aligning robot-mounted camera optical flow with hand-worn IMU signals. We first elicit a user-defined (N=12) gesture set and collect a multimodal command gesture dataset (N=38) in long-range multi-user HRI scenarios. Next, HiSync extracts frequency-domain hand motion features from both camera and IMU data, and a learned CSINet denoises IMU readings, temporally aligns modalities, and performs distance-aware multi-window fusion to compute cross-modal similarity of subtle, natural gestures, enabling robust CSI. In three-person scenes up to 34 m, HiSync achieves 92.32% CSI accuracy, outperforming the prior SOTA by 48.44%. HiSync is also validated on real-robot deployment. By making CSI reliable and natural, HiSync provides a practical primitive and design guidance for public-space HRI.
Chun Yu, Borong Zhuang, Haopeng Jin, Qingyang Wan, Zhuojun Li, Zhoutong Ye, Chang Liu 0150, Weinan Shi, Yuanchun Shi
CHI12
2026 OpenCD: Empowering Diagnosis of Children's Mathematical Cognition through Open-ended Multimodal Tasks
abstract
Assessing children’s cognitive development in early mathematics is vital for effective teaching. Compared to closed-ended questions, which may fail to capture nuanced developmental spectrum, open-ended elicitation tasks (e.g., asking students to manipulate objects or draw to represent numbers) serve as a promising approach to reveal deeper cognitive processes. However, their diverse and unstructured nature makes systematic analysis challenging for teachers. We present OpenCD, a teacher-facing system that automatically analyzes multimodal student responses to capture individualized insights. Based on Evidence-Centered Design, it combines Vision-Language Models (VLMs) and expert models to generate interactive diagnostic graphs and reports with traceability back to behavioral evidence. In our two-part evaluation, a validation study found 90.3% of the system’s diagnoses “completely reasonable,” and a user study showed that OpenCD reduced teachers’ analysis burden and enhanced their insights into student thinking. Our work contributes to scalable process-based assessment for mathematical literacy.
Chun Yu, Minzheng Song, Binglin Liu, Jianyang Liu, Xutong Wang, Jie Cai 0003, Yuanchun Shi
CHI10
2026 Beyond a single perspective: A multi-agent debate framework for affective computing
Yijie Pan, Yuanchun Shi, Chun Yu, Xiangzeng Kong, Naian Xiao
Pattern Recognit.2
2025 Non-Contact Health Monitoring During Daily Personal Care Routines
abstract
Remote photoplethysmography (rPPG) enables noncontact, continuous monitoring of physiological signals and offers a practical alternative to traditional health sensing methods. Although rPPG is promising for daily health monitoring, its application in long-term personal care scenarios-such as mirrorfacing routines in high-altitude environments-remains challenging due to ambient lighting variations, frequent occlusions from hand movements, and dynamic facial postures. To address these challenges, we present the Long-term Altitude Daily Health (LADH) dataset, the first long-term rPPG dataset containing 240 synchronized RGB and infrared (IR) facial videos from 21 participants across five common personal care scenarios, along with ground-truth PPG, respiration, and blood oxygen signals. Our experiments demonstrate that combining RGB and IR video inputs improves the accuracy and robustness of non-contact physiological monitoring, achieving a mean absolute error (MAE) of 4.99 BPM in heart rate estimation. Furthermore, we find that multi-task learning enhances performance across multiple physiological indicators simultaneously. Dataset and code are open at https://github.com/McJackTang/FusionVitals.
Xulin Ma, Jiankai Tang, Zhang Jiang, Songqin Cheng, Yuanchun Shi, Xin Liu 0034, Daniel McDuff, Yuntao Wang 0001
BSN5
2025 Modeling the Impact of Visual Stimuli on Redirection Noticeability with Gaze Behavior in Virtual Reality
Zhipeng Li 0001, Yishu Ji, Ruijia Chen, Yuntao Wang 0001, Yuanchun Shi, Yukang Yan
CHI6
2025 The Odyssey Journey: Top-Tier Medical Resource Seeking for Specialized Disorder in China
abstract
It is pivotal for patients to receive accurate health information, diagnoses, and timely treatments. However, in China, the significant imbalanced doctor-to-patient ratio intensifies the information and power asymmetries in doctor-patient relationships. Health information-seeking, which enables patients to collect information from sources beyond doctors, is a potential approach to mitigate these asymmetries. While HCI research predominantly focuses on common chronic conditions, our study focuses on specialized disorders, which are often familiar to specialists but not to general practitioners and the public. With Hemifacial Spasm (HFS) as an example, we aim to understand patients' health information and top-tier1 medical resource seeking journeys in China. Through interviews with three neurosurgeons and 12 HFS patients from rural and urban areas, and applying Actor-Network Theory, we provide empirical insights into the roles, interactions, and workflows of various actors in the health information-seeking network. We also identified five strategies patients adopted to mitigate asymmetries and access top-tier medical resources, illustrating these strategies as subnetworks within the broader health information-seeking network and outlining their advantages and challenges. © 2025 Copyright held by the owner/author(s).
Ka I Chan, Siying Hu, Yuntao Wang 0001, Xuhai Xu, Zhicong Lu, Yuanchun Shi
CHI6
2025 Investigating Context-Aware Collaborative Text Entry on Smartphones using Large Language Models
abstract
Text entry is a fundamental and ubiquitous task, but users often face challenges such as situational impairments or difficulties in sentence formulation.Motivated by this, we explore the potential of large language models (LLMs) to assist with text entry in realworld contexts.We propose a collaborative smartphone-based text entry system, CATIA, that leverages LLMs to provide text suggestions based on contextual factors, including screen content, time, location, activity, and more.In a 7-day in-the-wild study with 36 participants, the system offered appropriate text suggestions in over 80% of cases.Users exhibited different collaborative behaviors depending on whether they were composing text for interpersonal communication or information services.Additionally, the relevance
Yuanchun Shi, Weinan Shi, Meizhu Chen, Yeshuang Zhu, Jinchao Zhang 0001, Chun Yu
CHI2
2025 Unknown Word Detection for English as a Second Language (ESL) Learners using Gaze and Pre-trained Language Models
Jiexin Ding, Bowen Zhao 0004, Yuntao Wang 0001, Xinyun Liu, Ishan Chatterjee, Yuanchun Shi
CHI7
2025 Palmpad: Enabling Real-Time Index-to-Palm Touch Interaction with a Single RGB Camera
Yuanchun Shi, Chi Hsia, Chun Yu
CHI3
2025 WritingRing: Enabling Natural Handwriting Input with a Single IMU Ring
Zixuan Wang 0018, Chun Yu, Xiyuan Shen, Yuanchun Shi
CHI6
2025 Enhancing Smartphone Eye Tracking with Cursor-Based Interactive Implicit Calibration
Chang Liu 0150, Chun Yu, Yingtian Shi, Yuanchun Shi
CHI8
2025 VAction: A Lightweight and Integrated VR Training System for Authentic Film-Shooting Experience
Che Qu, Minjing Yu, Chao Zhou 0012, Yuntao Wang 0001, Yu-Hui Wen, Yuanchun Shi, Yong-Jin Liu 0001
CHI7
2025 From Operation to Cognition: Automatic Modeling Cognitive Dependencies from User Demonstrations for GUI Task Automation
Yiwen Yin, Chun Yu, Toby Jia-Jun Li, Aamir Khan Jadoon, Sixiang Cheng, Weinan Shi, Mohan Chen 0007, Yuanchun Shi
CHI9
2025 AutoPBL: An LLM-powered Platform to Guide and Support Individual Learners Through Self Project-based Learning
Zhoutong Ye, Wenxuan Tang, Chun Yu, Yuanchun Shi
CHI6
2025 GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-Based VLM Agent Training
abstract
Reinforcement learning with verifiable outcome rewards (RLVR) has effectively scaled up chain-of-thought (CoT) reasoning in large language models (LLMs). Yet, its efficacy in training vision-language model (VLM) agents for goal-directed action reasoning in visual environments is less established. This work investigates this problem through extensive experiments on complex card games, such as 24 points, and embodied tasks from ALFWorld. We find that when rewards are based solely on action outcomes, RL fails to incentivize CoT reasoning in VLMs, instead leading to a phenomenon we termed thought collapse, characterized by a rapid loss of diversity in the agent's thoughts, state-irrelevant and incomplete reasoning, and subsequent invalid actions, resulting in negative rewards. To counteract thought collapse, we highlight the necessity of process guidance and propose an automated corrector that evaluates and refines the agent's reasoning at each RL step. This simple and scalable GTR (Guided Thought Reinforcement) framework trains reasoning and action simultaneously without the need for dense, per-step human labeling. Our experiments demonstrate that GTR significantly enhances the performance and generalization of the LLaVA-7b model across various visual environments, achieving 3-5 times higher task success rates compared to SoTA models with notably smaller model sizes.
Junliang Xing, Yuanchun Shi, Zongqing Lu 0002, Deheng Ye
ICCV4
2025 BodyGen: Advancing Towards Efficient Embodiment Co-Design
abstract
Embodiment co-design aims to optimize a robot's morphology and control policy simultaneously. While prior work has demonstrated its potential for generating environment-adaptive robots, this field still faces persistent challenges in optimization efficiency due to the (i) combinatorial nature of morphological search spaces and (ii) intricate dependencies between morphology and control. We prove that the ineffective morphology representation and unbalanced reward signals between the design and control stages are key obstacles to efficiency. To advance towards efficient embodiment co-design, we propose **BodyGen**, which utilizes (1) topology-aware self-attention for both design and control, enabling efficient morphology representation with lightweight model sizes; (2) a temporal credit assignment mechanism that ensures balanced reward signals for optimization. With our findings, BodyGen achieves an average **60.03%** performance improvement against state-of-the-art baselines. We provide codes and more results on the website: https://genesisorigin.github.io.
Haofei Lu, Junliang Xing, Jianshu Li, Yuanchun Shi
ICLR7
2025 InterQuest: A Mixed-Initiative Framework for Dynamic User Interest Modeling in Conversational Search
Yuanxi Wang, Qingyang Wan, Zhuojun Li, Chun Yu, Weinan Shi, Yuanchun Shi
UIST8
2025 Understanding Users' Perceptions and Expectations toward a Social Balloon Robot via an Exploratory Study
Tianyi Xia, Manqiu Liao, Yuan Gao 0024, Chun Yu, Yuntao Wang 0001, Yuanchun Shi
UIST12
2025 Asymmetric interaction preference induces cooperation in human-agent hybrid game
Danyang Jia, Xiangfeng Dai, Junliang Xing, Pin Tao, Yuanchun Shi, Zhen Wang 0004
Sci. China Inf. Sci.5
2025 EchoMind: Supporting Real-time Complex Problem Discussions through Human-AI Collaborative Facilitation
abstract
Teams often engage in group discussions to leverage collective intelligence when solving complex problems. However, in real-time discussions, such as face-to-face meetings, participants frequently struggle with managing diverse perspectives and structuring content, which can lead to unproductive outcomes like forgetfulness and off-topic conversations. Through a formative study, we explores a human-AI collaborative facilitation approach, where AI assists in establishing a shared knowledge framework to provide a guiding foundation. We present EchoMind, a system that visualizes discussion knowledge through real-time issue mapping. EchoMind empowers participants to maintain focus on specific issues, review key ideas or thoughts, and collaboratively expand the discussion. The system leverages large language models (LLMs) to dynamically organize dialogues into nodes based on the current context recorded on the map. Our user study with four teams (N=16) reveals that EchoMind helps clarify discussion objectives, trace knowledge pathways, and enhance overall productivity. We also discuss the design implications for human-AI collaborative facilitation and the potential of shared knowledge visualization to transform group dynamics in future collaborations.
Chun Yu, Meizhu Chen, Yipeng Xu, Yuanchun Shi
Proc. ACM Hum. Comput. Interact.6
2025 Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts
abstract
UI task automation enables efficient task execution by simulating human interactions with GUIs, without modifying the existing application code. However, its broader adoption is constrained by the need for expertise in both scripting languages and workflow design. To address this challenge, we present Prompt2Task, a system designed to comprehend various task-related textual prompts (e.g., goals, procedures), thereby generating and performing the corresponding automation tasks. Prompt2Task incorporates a suite of intelligent agents that mimic human cognitive functions, specializing in interpreting user intent, managing external information for task generation, and executing operations on smartphones. The agents can learn from user feedback and continuously improve their performance based on the accumulated knowledge. Experimental results indicated a performance jump from a 22.28% success rate in the baseline to 95.24% with Prompt2Task, requiring an average of 0.69 user interventions for each new task. Prompt2Task presents promising applications in fields such as tutorial creation, smart assistance, and customer service.
Tian Huang, Chun Yu, Weinan Shi, Zijian Peng, David Yang 0002, Yuanchun Shi
ACM Trans. Comput. Hum. Interact.7
2025 AmplitudeArrow: On-the-Go AR Menu Selection Using Consecutive Simple Head Gestures and Amplitude Visualization
abstract
Heads-up computing aims to provide synergistic digital assistance that minimally interferes with users' on-the-go daily activities. Currently, the input modalities of heads-up computing are mainly voice and finger gestures. In this work, we propose and evaluate the AmplitudeArrow (AA) technique designed for on-the-go AR menu selection to demonstrate that consecutive simple head gestures can also be an effective input modality for heads-up computing. Specifically, AA arranges menu icons into one/two row(s). To select a target icon, the user first makes their head yaw to pre-select the target icon or the column containing it and then makes their head pitch to make the arrow in the target icon expand until the arrow covers the target icon completely, i.e., the pitch amplitude surpasses the selection confirmation threshold. User studies indicated that AA demonstrated robust resistance to walking-caused head perturbation and external factors such as other people/obstacles, delivering high accuracy (error rate $< $< 5$\%$%) and fast speed ($< $< 1.5s per selection) when there were no more than six icon columns (twelve icons) distributed horizontally and evenly in a menu area with a horizontal visual angle of $43^{\circ }$43∘.
Yang Tian 0008, Yukang Yan, Shengdong Zhao 0001, Xiaojuan Ma, Yuanchun Shi
IEEE Trans. Vis. Comput. Graph.6
2024 Time2Stop: Adaptive and Explainable Human-AI Loop for Smartphone Overuse Intervention
abstract
Despite a rich history of investigating smartphone overuse intervention techniques, AI-based just-in-time adaptive intervention (JITAI) methods for overuse reduction are lacking. We develop Time2Stop, an intelligent, adaptive, and explainable JITAI system that leverages machine learning to identify optimal intervention timings, introduces interventions with transparent AI explanations, and collects user feedback to establish a human-AI loop and adapt the intervention model over time. We conducted an 8-week field experiment (N=71) to evaluate the effectiveness of both the adaptation and explanation aspects of Time2Stop. Our results indicate that our adaptive models significantly outperform the baseline methods on intervention accuracy (>32.8% relatively) and receptivity (>8.0%). In addition, incorporating explanations further enhances the effectiveness by 53.8% and 11.4% on accuracy and receptivity, respectively. Moreover, Time2Stop significantly reduces overuse, decreasing app visit frequency by 7.0 ∼ 8.9%. Our subjective data also echoed these quantitative measures. Participants preferred the adaptive interventions and rated the system highly on intervention time accuracy, effectiveness, and level of trust. We envision our work can inspire future research on JITAI systems with a human-AI loop to evolve with users.
Adiba Orzikulova, Zhipeng Li 0001, Yukang Yan, Yuntao Wang 0001, Yuanchun Shi, Marzyeh Ghassemi, Sung-Ju Lee 0001, Anind K. Dey, Xuhai Xu
CHI6
2024 MouseRing: Always-available Touchpad Interaction with IMU Rings
abstract
Tracking fine-grained finger movements with IMUs for continuous 2D-cursor control poses significant challenges due to limited sensing capabilities. Our findings suggest that finger-motion patterns and the inherent structure of joints provide beneficial physical knowledge, which lead us to enhance motion perception accuracy by integrating physical priors into ML models. We propose MouseRing, a novel ring-shaped IMU device that enables continuous finger-sliding on unmodified physical surfaces like a touchpad. A motion dataset was created using infrared cameras, touchpads, and IMUs. We then identified several useful physical constraints, such as joint co-planarity, rigid constraints, and velocity consistency. These principles help refine the finger-tracking predictions from an RNN model. By incorporating touch state detection as a cursor movement switch, we achieved precise cursor control. In a Fitts’ Law study, MouseRing demonstrated input efficiency comparable to touchpads. In real-world applications, MouseRing ensured robust, efficient input and good usability across various surfaces and body postures.
Xiyuan Shen, Chun Yu, Xutong Wang, Haozhan Chen, Yuanchun Shi
CHI6
2024 PepperPose: Full-Body Pose Estimation with a Companion Robot
abstract
Accurate full-body pose estimation across diverse actions in a user-friendly and location-agnostic manner paves the way for interactive applications in realms like sports, fitness, and healthcare. This task becomes challenging in real-world scenarios due to factors like the user’s dynamic positioning, the diversity of actions, and the varying acceptability of the pose-capturing system. In this context, we present PepperPose, a novel companion robot system tailored for optimized pose estimation. Unlike traditional methods, PepperPose actively tracks the user and refines its viewpoint, facilitating enhanced pose accuracy across different locations and actions. This allows users to enjoy a seamless action-sensing experience. Our evaluation, involving 30 participants undertaking daily functioning and exercise actions in a home-like space, underscores the robot’s promising capabilities. Moreover, we demonstrate the opportunities that PepperPose presents for human-robot interaction, its current limitations, and future developments.
Lingxiao Zhong, Chun Yu, Yuntao Wang 0001, Yuan Gao 0024, Tin Lun Lam, Yuanchun Shi
CHI9
2024 MindShift: Leveraging Large Language Models for Mental-States-Based Problematic Smartphone Use Intervention
abstract
Problematic smartphone use negatively affects physical and mental health. Despite the wide range of prior research, existing persuasive techniques are not flexible enough to provide dynamic persuasion content based on users’ physical contexts and mental states. We first conducted a Wizard-of-Oz study (N=12) and an interview study (N=10) to summarize the mental states behind problematic smartphone use: boredom, stress, and inertia. This informs our design of four persuasion strategies: understanding, comforting, evoking, and scaffolding habits. We leveraged large language models (LLMs) to enable the automatic and dynamic generation of effective persuasion content. We developed MindShift, a novel LLM-powered problematic smartphone use intervention technique. MindShift takes users’ in-the-moment app usage behaviors, physical contexts, mental states, goals & habits as input, and generates personalized and dynamic persuasive content with appropriate persuasion strategies. We conducted a 5-week field experiment (N=25) to compare MindShift with its simplified version (remove mental states) and baseline techniques (fixed reminder). The results show that MindShift improves intervention acceptance rates by 4.7-22.5% and reduces smartphone usage duration by 7.4-9.8%. Moreover, users have a significant drop in smartphone addiction scale scores and a rise in self-efficacy scale scores. Our study sheds light on the potential of leveraging LLMs for context-aware persuasion in other behavior change domains.
Ruolan Wu, Chun Yu, Xiaole Pan, Yujia Liu 0004, Ningning Zhang, Yuhan Wang 0015, Qiaolei Jiang, Xuhai Xu, Yuanchun Shi
CHI12
2024 MEFusion: Unsupervised Mutual Enhancement for Multimodal Image Fusion
abstract
Image fusion aims to extract valuable information from each modality to create a fused image. Currently, state-of-the-art image fusion approaches tend to initially decompose each modality into distinct yet complementary features, and transfer beneficial information through carefully hand-crafted or learned fusion rules to the target. Nevertheless, previous approaches treat each modality in isolation before fusion, potentially under-utilising the complementary information available across modalities. In this word, we introduce a novel method called MEFusion that pioneers cross-modality mutual enhancement before feature decomposition. By harnessing the individual strengths of each modality, MEFusion elevates the overall quality and comprehensiveness of the fusion outcome. To facilitate a bidirectional enhancement for each feature across modalities, we have designed a pluggable co-attention mechanism that seamlessly integrates into a lightweight dual-path transformer. Furthermore, to enrich the details of each modality, we propose an unsupervised cross-modality mutual enhancement loss, which overcomes the limitations of requiring paired training data for enhancement tasks. Extensive experiments conducted on several benchmark datasets demonstrate the superiority of our proposed MEFusion method in terms of traditional fusion metrics and perceptual quality improvement of fused images.
Yushe Cao, Siwen Jiao, Penghao Sun, Baoyun Peng, Dian-xi Shi, Yuanchun Shi
ECAI6
2024 PoseAugment: Generative Human Pose Data Augmentation with Physical Plausibility for IMU-Based Motion Capture
Zhuojun Li, Chun Yu, Yuanchun Shi
ECCV (32)4
2024 PAE: Reinforcement Learning from External Knowledge for Efficient Exploration
abstract
Human intelligence is adept at absorbing valuable insights from external knowledge. This capability is equally crucial for artificial intelligence. In contrast, classical reinforcement learning agents lack such capabilities and often resort to extensive trial and error to explore the environment. This paper introduces $\textbf{PAE}$: $\textbf{P}$lanner-$\textbf{A}$ctor-$\textbf{E}$valuator, a novel framework for teaching agents to $\textit{learn to absorb external knowledge}$. PAE integrates the Planner's knowledge-state alignment mechanism, the Actor's mutual information skill control, and the Evaluator's adaptive intrinsic exploration reward to achieve 1) effective cross-modal information fusion, 2) enhanced linkage between knowledge and state, and 3) hierarchical mastery of complex tasks. Comprehensive experiments across 11 challenging tasks from the BabyAI and MiniHack environment suites demonstrate PAE's superior exploration efficiency with good interpretability.
Haofei Lu, Junliang Xing, Renye Yan, Yaozhong Gan, Yuanchun Shi
ICLR7
2024 DreamCatcher: A Wearer-aware Multi-modal Sleep Event Dataset Based on Earables in Non-restrictive Environments
abstract
Poor quality sleep can be characterized by the occurrence of events ranging from body movement to breathing impairment. Widely available earbuds equipped with sensors (also known as earables) can be combined with a sleep event detection algorithm to offer a convenient alternative to laborious clinical tests for individuals suffering from sleep disorders. Although various solutions utilizing such devices have been proposed to detect sleep events, they ignore the fact that individuals often share sleeping spaces with roommates or couples. To address this issue, we introduce DreamCatcher, the first publicly available dataset for wearer-aware sleep event algorithm development on earables. DreamCatcher encompasses eight distinct sleep events, including synchronous dual-channel audio and motion data collected from 12 pairs (24 participants) totaling 210 hours (420 hour.person) with fine-grained label. We tested multiple benchmark models on three tasks related to sleep event detection, demonstrating the usability and unique challenge of DreamCatcher. We hope that the proposed DreamCatcher can inspire other researchers to further explore efficient wearer-aware human vocal activity sensing on earables. DreamCatcher is publicly available at https://github.com/thuhci/DreamCatcher.
Xiyuxing Zhang, Ruotong Yu, Yuntao Wang 0001, Kenneth Christofferson, Jingru Zhang 0005, Alexander Mariakakis, Yuanchun Shi
NeurIPS8
2024 ContextMate: a context-aware smart agent for efficient data analysis
Aamir Khan Jadoon, Chun Yu, Yuanchun Shi
CCF Trans. Pervasive Comput. Interact.3
2024 HCI Research and Innovation in China: A 10-Year Perspective
abstract
In the past years, human computer interaction (HCI) research and innovation have developed substantially, leading to a number of fruitful research topics. In this paper, we surveyed the HCI research and innovation in China from a 10-year perspective. We analyzed the popular research methodology and topics among Chinese researchers, including human modeling, user interface techniques, context awareness, user acceptance and performance, user experience design, human-AI interaction, HCI applications and social influences. We also conducted a bibliography analysis on the published papers in top-tier conferences and journals, which revealed a significant rising trend, and a generally broad distribution of research types. Moreover, we described typical applications and the industry influence of the research outcomes. We concluded with implications and reflections for HCI researchers across the world and shared the future research trends envisioned by Chinese researchers.
Yuanchun Shi, Xin Yi 0001, Yuntao Wang 0001, Yukang Yan, Zhimin Cheng, Pengye Zhu, Yongjuan Li, Yanci Liu, Weixuan Zhou, Diya Zhao
Int. J. Hum. Comput. Interact.1
2024 Freedom of choice disrupts cyclic dominance but maintains cooperation in voluntary prisoner's dilemma game
Danyang Jia, Chen Shen 0006, Xiangfeng Dai, Xinyu Wang 0022, Junliang Xing, Pin Tao, Yuanchun Shi, Zhen Wang 0004
Knowl. Based Syst.7
2024 Previs-Real:Interactive virtual previsualization system for news shooting rehearsal and evaluation
abstract
In the demanding field of live news broadcasting, the intricate studio production procedures and tight schedules pose significant challenges for physical rehearsals by cameramen. This paper explores the design and implementation of a lightweight virtual news previsualization system, leveraging virtual production technology and interaction design methods to address the lack of fidelity in presentations and manipulations, and the quantitative feedback of rehearsal effects in previous virtual approaches. Our system, Previs-Real, is informed by user investigation with professional cameramen and studio technicians, and adheres to principles of high fidelity, accurate replication of actual hardware operations, and real-time feedback on rehearsal results. The system's software and hardware development are implemented based on Unreal Engine and accompanying toolsets, incorporating cutting-edge modeling and camera calibration methods. We validated Previs-Real through a user study, demonstrating superior performance in previsualization shooting tasks using the virtual system compared to traditional camera setups. The findings, supported by both objective performance metrics and subjective responses, underline Previs-Real's effectiveness and potential in transforming news broadcasting rehearsals. Previs-Real eliminates the requirement for complex equipment interconnections and team coordination inherent in a physical studio by implementing methodologies complying the above principles, objectively resulting in a lightweight design of applicable version of virtual news previsualization system. It offers a novel solution to the challenges in news studio previsualization by focusing on key operational features rather than full environment replication. This design approach is equally effective in the process of designing lightweight systems in other fields.
Che Qu, Tongchen Zhao, Cheng Wa Wong, Chi Deng, Yuhui Wen, Yuanchun Shi
Virtual Real. Intell. Hardw.10
2023 Modeling the Trade-off of Privacy Preservation and Activity Recognition on Low-Resolution Images
abstract
A computer vision system using low-resolution image sensors can provide intelligent services (e.g., activity recognition) but preserve unnecessary visual privacy information from the hardware level. However, preserving visual privacy and enabling accurate machine recognition have adversarial needs on image resolution. Modeling the trade-off of privacy preservation and machine recognition performance can guide future privacy-preserving computer vision systems using low-resolution image sensors. In this paper, using the at-home activity of daily livings (ADLs) as the scenario, we first obtained the most important visual privacy features through a user survey. Then we quantified and analyzed the effects of image resolution on human and machine recognition performance in activity recognition and privacy awareness tasks. We also investigated how modern image super-resolution techniques influence these effects. Based on the results, we proposed a method for modeling the trade-off of privacy preservation and activity recognition on low-resolution images.
Yuntao Wang 0001, Zirui Cheng, Xin Yi 0001, Yan Kong, Xuhai Xu, Yukang Yan, Chun Yu, Shwetak N. Patel, Yuanchun Shi
CHI10
2023 Squeez'In: Private Authentication on Smartphones based on Squeezing Gestures
abstract
In this paper, we proposed Squeez’In, a technique on smartphones that enabled private authentication by holding and squeezing the phone with a unique pattern. We first explored the design space of practical squeezing gestures for authentication by analyzing the participants’ self-designed gestures and squeezing behavior. Results showed that varying-length gestures with two levels of touch pressure and duration were the most natural and unambiguous. We then implemented Squeez’In on an off-the-shelf capacitive sensing smartphone, and employed an SVM-GBDT model for recognizing gestures and user-specific behavioral patterns, achieving 99.3% accuracy and 0.93 F1-score when tested on 21 users. A following 14-day study validated the memorability and long-term stability of Squeez’In. During usability evaluation, compared with gesture and pin code, Squeez’In achieved significantly faster authentication speed and higher user preference in terms of privacy and security.
Xin Yi 0001, Louisa Shi, Fengyan Han, Yan Kong, Hewu Li, Yuanchun Shi
CHI8
2023 Enabling Voice-Accompanying Hand-to-Face Gesture Recognition with Cross-Device Sensing
abstract
Gestures performed accompanying the voice are essential for voice interaction to convey complementary semantics for interaction purposes such as wake-up state and input modality. In this paper, we investigated voice-accompanying hand-to-face (VAHF) gestures for voice interaction. We targeted on hand-to-face gestures because such gestures relate closely to speech and yield significant acoustic features (e.g., impeding voice propagation). We conducted a user study to explore the design space of VAHF gestures, where we first gathered candidate gestures and then applied a structural analysis to them in different dimensions (e.g., contact position and type), outputting a total of 8 VAHF gestures with good usability and least confusion. To facilitate VAHF gesture recognition, we proposed a novel cross-device sensing method that leverages heterogeneous channels (vocal, ultrasound, and IMU) of data from commodity devices (earbuds, watches, and rings). Our recognition model achieved an accuracy of 97.3% for recognizing 3 gestures and 91.5% for recognizing 8 gestures (excluding the "empty" gesture), proving the high applicability. Quantitative analysis also shed light on the recognition capability of each sensor channel and their different combinations. In the end, we illustrated the feasible use cases and their design principles to demonstrate the applicability of our system in various scenarios.
Zisu Li, Yuntao Wang 0001, Chun Yu, Yukang Yan, Mingming Fan 0001, Yuanchun Shi
CHI8
2023 ResType: Invisible and Adaptive Tablet Keyboard Leveraging Resting Fingers
abstract
Text entry on tablet touchscreens is a basic need nowadays. Tablet keyboards require visual attention for users to locate keys, thus not supporting efficient touch typing. They also take up a large proportion of screen space, which affects the access to information. To solve these problems, we propose ResType, an adaptive and invisible keyboard on three-state touch surfaces (e.g. tablets with unintentional touch prevention). ResType allows users to rest their hands on it and automatically adapts the keyboard to the resting fingers. Thus, users do not need visual attention to locate keys, which supports touch typing. We quantitatively explored users’ resting finger patterns on ResType, based on which we proposed an augmented Bayesian decoding algorithm for ResType, with 96.3% top-1 and 99.0% top-3 accuracies. After a 5-day evaluation, ResType achieved 41.26 WPM, outperforming normal tablet keyboards by 13.5% and reaching 86.7% of physical keyboards. It solves the occlusion problem while maintaining comparable typing speed with current methods on visible tablet keyboards.
Zhuojun Li, Chun Yu, Yizheng Gu, Yuanchun Shi
CHI4
2023 A Human-Computer Collaborative Editing Tool for Conceptual Diagrams
abstract
Editing (e.g., editing conceptual diagrams) is a typical office task that requires numerous tedious GUI operations, resulting in poor interaction efficiency and user experience, especially on mobile devices. In this paper, we present a new type of human-computer collaborative editing tool (CET) that enables accurate and efficient editing with little interaction effort. CET divides the task into two parts, and the human and the computer focus on their respective specialties: the human describes high-level editing goals with multimodal commands, while the computer calculates, recommends, and performs detailed operations. We conducted a formative study (N = 16) to determine the concrete task division and implemented the tool on Android devices for the specific tasks of editing concept diagrams. The user study (N = 24 + 20) showed that it increased diagram editing speed by 32.75% compared with existing state-of-the-art commercial tools and led to better editing results and user experience.
Lihang Pan, Chun Yu, Yuanchun Shi
CHI4
2023 Selecting Real-World Objects via User-Perspective Phone Occlusion
abstract
Perceiving the region of interest (ROI) and target object by smartphones from the user’s first-person perspective can enable diverse spatial interactions. In this paper, we propose a novel ROI input method and a target selecting method for smartphones by utilizing the user-perspective phone occlusion. This concept of turning the phone into real-world physical cursor benefits from the proprioception, gets rid of the constraint of camera preview, and allows users to rapidly and accurately select the target object. Meanwhile, our method can provide a resizable and rotatable rectangular ROI to disambiguate dense targets. We implemented the prototype system by positioning the user’s iris with the front camera and estimating the rectangular area blocked by the phone with the rear camera simultaneously, followed by a target prediction algorithm with the distance-weighted Jaccard index. We analyzed the behavioral models of using our method and evaluated our prototype system’s pointing accuracy and usability. Results showed that our method is well-accepted by the users for its convenience, accuracy, and efficiency.
Chun Yu, Jiachen Yao, Yueting Weng, Yukang Yan, Yuanchun Shi
CHI8
2023 SmartRecorder: An IMU-based Video Tutorial Creation by Demonstration System for Smartphone Interaction Tasks
abstract
This work focuses on an active topic in the HCI community, namely tutorial creation by demonstration. We present a novel tool named SmartRecorder that facilitates people, without video editing skills, creating video tutorials for smartphone interaction tasks. As automatic interaction trace extraction is a key component to tutorial generation, we seek to tackle the challenges of automatically extracting user interaction traces on smartphones from screencasts. Uniquely, with respect to prior research in this field, we combine computer vision techniques with IMU-based sensing algorithms, and the technical evaluation results show the importance of smartphone IMU data in improving system performance. With the extracted key information of each step, SmartRecorder generates instructional content initially and provides tutorial creators with a tutorial refinement editor designed based on a high recall (99.38%) of key steps to revise the initial instructional content. Finally, SmartRecorder generates video tutorials based on refined instructional content. The results of the user study demonstrate that SmartRecorder allows non-experts to create smartphone usage video tutorials with less time and higher satisfaction from recipients.
Xiaozhu Hu, Yanwen Huang, Bo Liu 0091, Ruolan Wu, Yongquan Hu, Aaron J. Quigley, Mingming Fan 0001, Chun Yu, Yuanchun Shi
IUI9
2023 From Gap to Synergy: Enhancing Contextual Understanding through Human-Machine Collaboration in Personalized Systems
abstract
This paper presents LangAware, a collaborative approach for constructing personalized context for context-aware applications. The need for personalization arises due to significant variations in context between individuals based on scenarios, devices, and preferences. However, there is often a notable gap between humans and machines in the understanding of how contexts are constructed, as observed in trigger-action programming studies such as IFTTT. LangAware enables end-users to participate in establishing contextual rules in-situ using natural language. The system leverages large language models (LLMs) to semantically connect low-level sensor detectors to high-level contexts and provide understandable natural language feedback for effective user involvement. We conducted a user study with 16 participants in real-life settings, which revealed an average success rate of 87.50% for defining contextual rules in a variety of 12 campus scenarios, typically accomplished within just two modifications. Furthermore, users reported a better understanding of the machine’s capabilities by interacting with LangAware.
Chun Yu, Lichen Yang, Weinan Shi, Yuanchun Shi
UIST8
2023 Reprogrammable Digital Metamaterials for Interactive Devices
abstract
We present digital mechanical metamaterials that enable multiple computation loops and reprogrammable logic functions, making a significant step towards passive yet interactive devices. Our materials consist of many cells that transmit signals using an embedded bistable spring. When triggered, the bistable spring displaces and triggers the next cell. We integrate a recharging mechanism to recharge the bistable springs, enabling multiple computation rounds. Between the iterations, we enable reprogramming the logic functions after fabrication. We demonstrate that such materials can trigger a simple controlled actuation anywhere in the material to change the local shape, texture, stiffness, and display. This enables large-scale interactive and functional materials with no or a small number of external actuators. We showcase the capabilities of our system with various examples: a haptic floor with tunable stiffness for different VR scenarios, a display with easy-to-reconfigure messages after fabrication, or a tactile notification integrated into users’ desktops.
Yu Jiang 0010, Shobhit Aggarwal, Zhipeng Li 0001, Yuanchun Shi, Alexandra Ion
UIST4
2023 ShadowTouch: Enabling Free-Form Touch-Based Hand-to-Surface Interaction with Wrist-Mounted Illuminant by Shadow Projection
abstract
We present ShadowTouch, a novel sensing method to recognize the subtle hand-to-surface touch state for independent fingers based on optical auxiliary. ShadowTouch mounts a forward-facing light source on the user’s wrist to construct shadows on the surface in front of the fingers when the corresponding fingers are close to the surface. With such an optical design, the subtle vertical movements of near-surface fingers are magnified and turned to shadow features cast on the surface, which are recognizable for computer vision algorithms. To efficiently recognize the touch state of each finger, we devised a two-stage CNN-based algorithm that first extracted all the fingertip regions from each frame and then classified the touch state of each region from the cropped consecutive frames. Evaluations showed our touch state detection algorithm achieved a recognition accuracy of 99.1% and an F-1 score of 96.8% in the leave-one-out cross-user evaluation setting. We further outlined the hand-to-surface interaction space enabled by ShadowTouch’s sensing capability from the aspects of touch-based interaction, stroke-based interaction, and out-of-surface information and developed four application prototypes to showcase ShadowTouch’s interaction potential. The usability evaluation study showed the advantages of ShadowTouch over threshold-based techniques in aspects of lower mental demand, lower effort, lower frustration, more willing to use, easier to use, better integrity, and higher confidence.
Xutong Wang, Zisu Li, Chi Hsia, Mingming Fan 0001, Chun Yu, Yuanchun Shi
UIST7
2023 ConeSpeech: Exploring Directional Speech Interaction for Multi-Person Remote Communication in Virtual Reality
abstract
Remote communication is essential for efficient collaboration among people at different locations. We present ConeSpeech, a virtual reality (VR) based multi-user remote communication technique, which enables users to selectively speak to target listeners without distracting bystanders. With ConeSpeech, the user looks at the target listener and only in a cone-shaped area in the direction can the listeners hear the speech. This manner alleviates the disturbance to and avoids overhearing from surrounding irrelevant people. Three featured functions are supported, directional speech delivery, size-adjustable delivery range, and multiple delivery areas, to facilitate speaking to more than one listener and to listeners spatially mixed up with bystanders. We conducted a user study to determine the modality to control the cone-shaped delivery area. Then we implemented the technique and evaluated its performance in three typical multi-user communication tasks by comparing it to two baseline methods. Results show that ConeSpeech balanced the convenience and flexibility of voice communication.
Yukang Yan, Haohua Liu, Yingtian Shi, Ruici Guo, Zisu Li, Xuhai Xu, Chun Yu, Yuntao Wang 0001, Yuanchun Shi
IEEE Trans. Vis. Comput. Graph.10
2022 FaceOri: Tracking Head Position and Orientation Using Ultrasonic Ranging on Earphones
abstract
Face orientation can often indicate users’ intended interaction target. In this paper, we propose FaceOri, a novel face tracking technique based on acoustic ranging using earphones. FaceOri can leverage the speaker on a commodity device to emit an ultrasonic chirp, which is picked up by the set of microphones on the user’s earphone, and then processed to calculate the distance from each microphone to the device. These measurements are used to derive the user’s face orientation and distance with respect to the device. We conduct a ground truth comparison and user study to evaluate FaceOri’s performance. The results show that the system can determine whether the user orients to the device at a 93.5% accuracy within a 1.5 meters range. Furthermore, FaceOri can continuously track user’s head orientation with a median absolute error of 10.9 mm in the distance, 3.7° in yaw, and 5.8° in pitch. FaceOri can allow for convenient hands-free control of devices and produce more intelligent context-aware interactions.
Yuntao Wang 0001, Jiexin Ding, Ishan Chatterjee, Farshid Salemi Parizi, Yuzhou Zhuang, Yukang Yan, Shwetak N. Patel, Yuanchun Shi
CHI8
2022 Automatically Generating and Improving Voice Command Interface from Operation Sequences on Smartphones
abstract
Using voice commands to automate smartphone tasks (e.g., making a video call) can effectively augment the interactivity of numerous mobile apps. However, creating voice command interfaces requires a tremendous amount of effort in labeling and compiling the graphical user interface (GUI) and the utterance data. In this paper, we propose AutoVCI, a novel approach to automatically generate voice command interface (VCI) from smartphone operation sequences. The generated voice command interface has two distinct features. First, it automatically maps a voice command to GUI operations and fills in parameters accordingly, leveraging the GUI data instead of corpus or hand-written rules. Second, it launches a complementary Q&A dialogue to confirm the intention in case of ambiguity. In addition, the generated voice command interface can learn and evolve from user interactions. It accumulates the history command understanding results to annotate the user’s input and improve its semantic understanding ability. We implemented this approach on Android devices and conducted a two-phase user study with 16 and 67 participants in each phase. Experimental results of the study demonstrated the practical feasibility of AutoVCI.
Lihang Pan, Chun Yu, Jiahui Li 0010, Tian Huang, Xiaojun Bi 0001, Yuanchun Shi
CHI6
2022 TypeOut: Leveraging Just-in-Time Self-Affirmation for Smartphone Overuse Reduction
abstract
Smartphone overuse is related to a variety of issues such as lack of sleep and anxiety. We explore the application of Self-Affirmation Theory on smartphone overuse intervention in a just-in-time manner. We present TypeOut, a just-in-time intervention technique that integrates two components: an in-situ typing-based unlock process to improve user engagement, and self-affirmation-based typing content to enhance effectiveness. We hypothesize that the integration of typing and self-affirmation content can better reduce smartphone overuse. We conducted a 10-week within-subject field experiment (N=54) and compared TypeOut against two baselines: one only showing the self-affirmation content (a common notification-based intervention), and one only requiring typing non-semantic content (a state-of-the-art method). TypeOut reduces app usage by over 50%, and both app opening frequency and usage duration by over 25%, all significantly outperforming baselines. TypeOut can potentially be used in other domains where an intervention may benefit from integrating self-affirmation exercises with an engaging just-in-time mechanism.
Xuhai Xu, Tianyuan Zou, Yanzhang Li, Ruolin Wang, Tianyi Yuan, Yuntao Wang 0001, Yuanchun Shi, Jennifer Mankoff, Anind K. Dey
CHI8
2022 DEEP: 3D Gaze Pointing in Virtual Reality Leveraging Eyelid Movement
abstract
Gaze-based target suffers from low input precision and target occlusion. In this paper, we explored to leverage the continuous eyelid movement to support high-efficient and occlusion-robust dwell-based gaze pointing in virtual reality. We first conducted two user studies to examine the users’ eyelid movement pattern both in unintentional and intentional conditions. The results proved the feasibility of leveraging intentional eyelid movement that was distinguishable with natural movements for input. We also tested the participants’ dwelling pattern for targets with different sizes and locations. Based on these results, we propose DEEP, a novel technique that enables the users to see through occlusions by controlling the aperture angle of their eyelids and dwell to select the targets with the help of a probabilistic input prediction model. Evaluation results showed that DEEP with dynamic depth and location selection incorporation significantly outperformed its static variants, as well as a naive dwelling baseline technique. Even for 100% occluded targets, it could achieve an average selection speed of 2.5s with an error rate of 2.3%.
Xin Yi 0001, Leping Qiu, Wenjing Tang, Yehan Fan, Hewu Li, Yuanchun Shi
UIST6
2022 Color-to-Depth Mappings as Depth Cues in Virtual Reality
abstract
Despite significant improvements to Virtual Reality (VR) technologies, most VR displays are fixed focus and depth perception is still a key issue that limits the user experience and the interaction performance. To supplement humans’ inherent depth cues (e.g., retinal blur, motion parallax), we investigate users’ perceptual mappings of distance to virtual objects’ appearance to generate visual cues aimed to enhance depth perception. As a first step, we explore color-to-depth mappings for virtual objects so that their appearance differs in saturation and value to reflect their distance. Through a series of controlled experiments, we elicit and analyze users’ strategies of mapping a virtual object’s hue, saturation, value and a combination of saturation and value to its depth. Based on the collected data, we implement a computational model that generates color-to-depth mappings fulfilling adjustable requirements on confusion probability, number of depth levels, and consistent saturation/value changing tendency. We demonstrate the effectiveness of color-to-depth mappings in a 3D sketching task, showing that compared to single-colored targets and strokes, with our mappings, the users were more confident in the accuracy without extra cognitive load and reduced the perceived depth error by 60.8%. We also implement four VR applications and demonstrate how our color cues can benefit the user experience and interaction performance in VR.
Zhipeng Li 0001, Yikai Cui, Tianze Zhou, Yu Jiang 0010, Yuntao Wang 0001, Yukang Yan, Michael Nebeling, Yuanchun Shi
UIST8
2022 GazeDock: Gaze-Only Menu Selection in Virtual Reality using Auto-Triggering Peripheral Menu
abstract
Gaze-only input techniques in VR face the challenge of avoiding false triggering due to continuous eye tracking while maintaining interaction performance. In this paper, we proposed GazeDock, a technique for enabling fast and robust gaze-based menu selection in VR. GazeDock features a view-fixed peripheral menu layout that automatically triggers appearing and selection when the user’s gaze approaches and leaves the menu zone, thus facilitating interaction speed and minimizing the false triggering rate. We built a dataset of 12 participants’ natural gaze movements in typical VR applications. By analyzing their gaze movement patterns, we designed the menu UI personalization and optimized selection detection algorithm of GazeDock. We also examined users’ gaze selection precision for targets on the peripheral menu and found that 4–8 menu items yield the highest throughput when considering both speed and accuracy. Finally, we validated the usability of GazeDock in a VR navigation game that contains both scene exploration and menu selection. Results showed that GazeDock achieved an average selection time of 471ms and a false triggering rate of 3.6%. And it received higher user preference ratings compared with dwell-based and pursuit-based techniques.
Xin Yi 0001, Yiqin Lu, Ziyin Cai, Zihan Wu 0002, Yuntao Wang 0001, Yuanchun Shi
VR6
2022 Easily-add battery-free wireless sensors to everyday objects: system implementation and usability study
Tengxiang Zhang, Zi Qian, Hsuan-Wei Fan, Jie Ren 0017, Yuntao Wang 0001, Yuanchun Shi
CCF Trans. Pervasive Comput. Interact.6
2022 Investigating user-defined flipping gestures for dual-display phones
Zhican Yang, Chun Yu, Xin Chen 0083, Jingjia Luo, Yuanchun Shi
Int. J. Hum. Comput. Stud.5
2022 Design and Evaluation of Window Management Operations in AR Headset + Smartphone Interface
abstract
Combining the use of an AR headset and a smartphone can provide wider display and precise touch input simultaneously; it can redefine the way we use applications today. Unfortunately, users are deprived of such benefit because of the independence of two devices. There lacks a kind of intuitive and direct interactions across them. In this paper, we conduct a formative study to understand the window management requirements and interaction preferences of using an AR headset and a smartphone simultaneously and report the insights we gained. Also, we introduce an example vocabulary of window management operations in AR headset + smartphone interface. It allows users to manipulate windows in virtual space and shift windows between devices efficiently and seamlessly.
Jie Ren 0017, Chun Yu, Yueting Weng, Chengchi Zhou, Yuanchun Shi
Virtual Real. Intell. Hardw.5
2022 Intelligent interaction in mixed reality
abstract
With the rise of the concept of the metaverse, mixed reality continues to receive keen attention from all over the world.Academia and industry worldwide are constantly innovating the key technologies of mixed reality software and hardware.For a long time, the field of mixed reality mainly focused on display technology, but with the increase of mixed reality applications, people gradually found that the lack of interactive technology and solutions has become a bottleneck restricting the development of mixed reality technology.On the mixed reality platform represented by AR/VR helmets, people cannot completely get rid of the remote control to complete human-computer interaction, which greatly limits the development of applications on it.How to enable users to exchange information with computers based on the natural
Yuanchun Shi, Chun Yu
Virtual Real. Intell. Hardw.1
2021 Understanding the Design Space of Mouth Microgestures
abstract
As wearable devices move toward the face (i.e. smart earbuds, glasses), there is an increasing need to facilitate intuitive interactions with these devices. Current sensing techniques can already detect many mouth-based gestures; however, users’ preferences of these gestures are not fully understood. In this paper, we investigate the design space and usability of mouth-based microgestures. We first conducted brainstorming sessions (N=16) and compiled an extensive set of 86 user-defined gestures. Then, with an online survey (N=50), we assessed the physical and mental demand of our gesture set and identified a subset of 14 gestures that can be performed easily and naturally. Finally, we conducted a remote Wizard-of-Oz usability study (N=11) mapping gestures to various daily smartphone operations under a sitting and walking context. From these studies, we develop a taxonomy for mouth gestures, finalize a practical gesture set for common applications, and provide design guidelines for future mouth-based gesture interactions.
Xuhai Xu, Richard Li 0002, Yuanchun Shi, Shwetak N. Patel, Yuntao Wang 0001
Conference on Designing Interactive Systems4
2021 Facilitating Text Entry on Smartphones with QWERTY Keyboard for Users with Parkinson's Disease
abstract
QWERTY is the primary smartphone text input keyboard configuration. However, insertion and substitution errors caused by hand tremors, often experienced by users with Parkinson’s disease, can severely affect typing efficiency and user experience. In this paper, we investigated Parkinson’s users’ typing behavior on smartphones. In particular, we identified and compared the typing characteristics generated by users with and without Parkinson’s symptoms. We then proposed an elastic probabilistic model for input prediction. By incorporating both spatial and temporal features, this model generalized the classical statistical decoding algorithm to correct insertion, substitution and omission errors, while maintaining direct physical interpretation. User study results confirmed that the proposed algorithm outperformed baseline techniques: users reached 22.8 WPM typing speed with a significantly lower error rate and higher user-perceived performance and preference. We concluded that our method could effectively improve the text entry experience on smartphones for users with Parkinson’s disease.
Yuntao Wang 0001, Ao Yu, Xin Yi 0001, Yuanwei Zhang, Ishan Chatterjee, Shwetak N. Patel, Yuanchun Shi
CHI7
2021 LightWrite: Teach Handwriting to The Visually Impaired with A Smartphone
abstract
Learning to write is challenging for blind and low vision (BLV) people because of the lack of visual feedback. Regardless of the drastic advancement of digital technology, handwriting is still an essential part of daily life. Although tools designed for teaching BLV to write exist, many are expensive and require the help of sighted teachers. We propose LightWrite, a low-cost, easy-to-access smartphone application that uses voice-based descriptive instruction and feedback to teach BLV users to write English lowercase letters and Arabian digits in a specifically designed font. A two-stage study with 15 BLV users with little prior writing knowledge shows that LightWrite can successfully teach users to learn handwriting characters in an average of 1.09 minutes for each letter. After initial training and 20-minute daily practice for 5 days, participants were able to write an average of 19.9 out of 26 letters that are recognizable by sighted raters.
Zihan Wu 0002, Chun Yu, Xuhai Xu, Tianyuan Zou, Ruolin Wang, Yuanchun Shi
CHI7
2021 PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback
abstract
Second language (L2) English learners often find it difficult to improve their pronunciations due to the lack of expressive and personalized corrective feedback. In this paper, we present Pronunciation Teacher (PTeacher), a Computer-Aided Pronunciation Training (CAPT) system that provides personalized exaggerated audio-visual corrective feedback for mispronunciations. Though the effectiveness of exaggerated feedback has been demonstrated, it is still unclear how to define the appropriate degrees of exaggeration when interacting with individual learners. To fill in this gap, we interview 100 L2 English learners and 22 professional native teachers to understand their needs and experiences. Three critical metrics are proposed for both learners and teachers to identify the best exaggeration levels in both audio and visual modalities. Additionally, we incorporate the personalized dynamic feedback mechanism given the English proficiency of learners. Based on the obtained insights, a comprehensive interactive pronunciation training course is designed to help L2 learners rectify mispronunciations in a more perceptible, understandable, and discriminative manner. Extensive user studies demonstrate that our system significantly promotes the learners’ learning efficiency.
Yaohua Bu, Hang Zhou 0009, Jia Jia 0001, Shengqi Chen 0001, Dachuan Shi, Haozhe Wu, Kun Li 0003, Zhiyong Wu 0001, Yuanchun Shi, Xiaobo Lu, Ziwei Liu 0002
CHI13
2021 Auth+Track: Enabling Authentication Free Interaction on Smartphone by Continuous User Tracking
abstract
We propose Auth+Track, a novel authentication model that aims to reduce redundant authentication in everyday smartphone usage. By sparse authentication and continuous tracking of the user’s status, Auth+Track eliminates the “gap” authentication between fragmented sessions and enables “Authentication Free when User is Around”. To instantiate the Auth+Track model, we present PanoTrack, a prototype that integrates body and near field hand information for user tracking. We install a fisheye camera on the top of the phone to achieve a panoramic vision that can capture both user’s body and on-screen hands. Based on the captured video stream, we develop an algorithm to extract 1) features for user tracking, including body keypoints and their temporal and spatial association, near field hand status, and 2) features for user identity assignment. The results of our user studies validate the feasibility of PanoTrack and demonstrate that Auth+Track not only improves the authentication efficiency but also enhances user experiences with better usability.
Chun Yu, Xiaoying Wei, Xuhai Xu, Yongquan Hu, Yuntao Wang 0001, Yuanchun Shi
CHI7
2021 Tactile Compass: Enabling Visually Impaired People to Follow a Path with Continuous Directional Feedback
abstract
Accurate and effective directional feedback is crucial for an electronic traveling aid device that guides visually impaired people in walking through paths. This paper presents Tactile Compass, a hand-held device that provides continuous directional feedback with a rotatable needle pointing toward the planned direction. We conducted two lab studies to evaluate the effectiveness of the feedback solution. Results showed that, using Tactile Compass, participants could reach the target direction in place with a mean deviation of 3.03° and could smoothly navigate along paths of 60cm width, with a mean deviation from the centerline of 12.1cm. Subjective feedback showed that Tactile Compass was easy to learn and use.
Guanhong Liu, Tianyu Yu 0001, Chun Yu, Haiqing Xu 0001, Shuchang Xu, Ciyuan Yang, Haipeng Mi, Yuanchun Shi
CHI9
2021 ProxiMic: Convenient Voice Activation via Close-to-Mic Speech Detected by a Single Microphone
abstract
Wake-up-free techniques (e.g., Raise-to-Speak) are important for improving the voice input experience. We present ProxiMic, a close-to-mic (within 5 cm) speech sensing technique using only one microphone. With ProxiMic, a user keeps a microphone-embedded device close to the mouth and speaks directly to the device without wake-up phrases or button presses. To detect close-to-mic speech, we use the feature from pop noise observed when a user speaks and blows air onto the microphone. Sound input is first passed through a low-pass adaptive threshold filter, then analyzed by a CNN which detects subtle close-to-mic features (mainly pop noise). Our two-stage algorithm can achieve 94.1% activation recall, 12.3 False Accepts per Week per User (FAWU) with 68 KB memory size, which can run at 352 fps on the smartphone. The user study shows that ProxiMic is efficient, user-friendly, and practical.
Chun Yu, Zhaoheng Li, Mingyuan Zhong 0001, Yukang Yan, Yuanchun Shi
CHI6
2021 FaceSight: Enabling Hand-to-Face Gesture Interaction on AR Glasses with a Downward-Facing Camera Vision
abstract
We present FaceSight, a computer vision-based hand-to-face gesture sensing technique for AR glasses. FaceSight fixes an infrared camera onto the bridge of AR glasses to provide extra sensing capability of the lower face and hand behaviors. We obtained 21 hand-to-face gestures and demonstrated the potential interaction benefits through five AR applications. We designed and implemented an algorithm pipeline that segments facial regions, detects hand-face contact (f1 score: 98.36%), and trains convolutional neural network (CNN) models to classify the hand-to-face gestures. The input features include gesture recognition, nose deformation estimation, and continuous fingertip movement. Our algorithm achieves classification accuracy of all gestures at 83.06%, proved by the data of 10 users. Due to the compact form factor and rich gestures, we recognize FaceSight as a practical solution to augment input capability of AR glasses in the future.
Yueting Weng, Chun Yu, Yingtian Shi, Yuhang Zhao 0001, Yukang Yan, Yuanchun Shi
CHI6
2021 HulaMove: Using Commodity IMU for Waist Interaction
abstract
We present HulaMove, a novel interaction technique that leverages the movement of the waist as a new eyes-free and hands-free input method for both the physical world and the virtual world. We first conducted a user study (N=12) to understand users’ ability to control their waist. We found that users could easily discriminate eight shifting directions and two rotating orientations, and quickly confirm actions by returning to the original position (quick return). We developed a design space with eight gestures for waist interaction based on the results and implemented an IMU-based real-time system. Using a hierarchical machine learning model, our system could recognize waist gestures at an accuracy of 97.5%. Finally, we conducted a second user study (N=12) for usability testing in both real-world scenarios and virtual reality settings. Our usability study indicated that HulaMove significantly reduced interaction time by 41.8% compared to a touch screen method, and greatly improved users’ sense of presence in the virtual world. This novel technique provides an additional input method when users’ eyes or hands are busy, accelerates users’ daily operations, and augments their immersive experience in the virtual world.
Xuhai Xu, Tianyi Yuan, Liang He 0005, Xin Liu 0034, Yukang Yan, Yuntao Wang 0001, Yuanchun Shi, Jennifer Mankoff, Anind K. Dey
CHI8
2021 SemanticAdapt: Optimization-based Adaptation of Mixed Reality Layouts Leveraging Virtual-Physical Semantic Connections
abstract
We present an optimization-based approach that automatically adapts Mixed Reality (MR) interfaces to different physical environments. Current MR layouts, including the position and scale of virtual interface elements, need to be manually adapted by users whenever they move between environments, and whenever they switch tasks. This process is tedious and time consuming, and arguably needs to be automated for MR systems to be beneficial for end users. We contribute an approach that formulates this challenge as a combinatorial optimization problem and automatically decides the placement of virtual interface elements in new environments. To achieve this, we exploit the semantic association between the virtual interface elements and physical objects in an environment. Our optimization furthermore considers the utility of elements for users’ current task, layout factors, and spatio-temporal consistency to previous layouts. All those factors are combined in a single linear program, which is used to adapt the layout of MR interfaces in real time. We demonstrate a set of application scenarios, showcasing the versatility and applicability of our approach. Finally, we show that compared to a naive adaptive baseline approach that does not take semantic associations into account, our approach decreased the number of manual interface adaptations by 33%.
Yi Fei Cheng 0001, Yukang Yan, Xin Yi 0001, Yuanchun Shi, David Lindlbauer
UIST4
2021 Just Speak It: Minimize Cognitive Load for Eyes-Free Text Editing with a Smart Voice Assistant
abstract
Entering text precisely by voice, users might encounter colloquial inserts, inappropriate wording, and recognition errors, which brings difficulties to voice editing. Users need to locate the errors and then correct them. In eyes-free scenarios, this select-modify mode brings a cognitive burden and a risk of error. This paper introduces neural networks and pre-trained models to understand users’ revision intention based on semantics, reducing the need for the information from users’ statements. We present two strategies. One is to remove the colloquial inserts automatically. The other is to allow users to edit by just speaking out the target words without having to say the context and the incorrect text. Accordingly, our approach can predict whether to insert or replace, the incorrect text to replace, and the position to insert. We implement these strategies in SmartEdit, an eyes-free voice input agent controlled with earphone buttons. The evaluation shows that our techniques reduce the cognitive load and decrease the average failure rate by 54.1% compared to descriptive command or re-speaking.
Jiayue Fan, Chenning Xu, Chun Yu, Yuanchun Shi
UIST4
2021 TypeBoard: Identifying Unintentional Touch on Pressure-Sensitive Touchscreen Keyboards
abstract
Text input is essential in tablet computer interaction. However, tablet software keyboards face the problem of misrecognizing unintentional touch, which affects efficiency and usability [29, 49]. In this paper, we proposed TypeBoard, a pressure-sensitive touchscreen keyboard that prevents unintentional touches. The TypeBoard allows users to rest their fingers on the touchscreen, which changes the user behavior: on average, users generate 40.83 unintentional touches every 100 keystrokes. The TypeBoard prevents unintentional touch with an accuracy of 98.88%. A typing study showed that the TypeBoard reduced fatigue (p < 0.005) and typing errors (p < 0.01), and improved the touchscreen keyboard’ typing speed by 11.78% (p < 0.005). As users could touch the screen without triggering responses, we added tactile landmarks on the TypeBoard, allowing users to locate the keys by the sense of touch. This feature further improves the typing speed, outperforming the ordinary tablet keyboard by 21.19% (p < 0.001). Results show that pressure-sensitive touchscreen keyboards can prevent unintentional touch, improving usability from many aspects, such as avoiding fatigue, reducing errors, and mediating touch typing on tablets.
Yizheng Gu, Chun Yu, Xuanzhong Chen, Zhuojun Li, Yuanchun Shi
UIST5
2021 ReflecTrack: Enabling 3D Acoustic Position Tracking Using Commodity Dual-Microphone Smartphones
abstract
3D position tracking on smartphones has the potential to unlock a variety of novel applications, but has not been made widely available due to limitations in smartphone sensors. In this paper, we propose ReflecTrack, a novel 3D acoustic position tracking method for commodity dual-microphone smartphones. A ubiquitous speaker (e.g., smartwatch or earbud) generates inaudible Frequency Modulated Continuous Wave (FMCW) acoustic signals that are picked up by both smartphone microphones. To enable 3D tracking with two microphones, we introduce a reflective surface that can be easily found in everyday objects near the smartphone. Thus, the microphones can receive sound from the speaker and echoes from the surface for FMCW-based acoustic ranging. To simultaneously estimate the distances from the direct and reflective paths, we propose the echo-aware FMCW technique with a new signal pattern and target detection process. Our user study shows that ReflecTrack achieves a median error of 28.4 mm in the 60cm × 60cm × 60cm space and 22.1 mm in the 30cm × 30cm × 30cm space for 3D positioning. We demonstrate the easy accessibility of ReflecTrack using everyday surfaces and objects with several typical applications of 3D position tracking, including 3D input for smartphones, fine-grained gesture recognition, and motion tracking in smartphone-based VR systems.
Yuzhou Zhuang, Yuntao Wang 0001, Yukang Yan, Xuhai Xu, Yuanchun Shi
UIST5
2020 MoveVR: Enabling Multiform Force Feedback in Virtual Reality using Household Cleaning Robot
abstract
Haptic feedback can significantly enhance the realism and immersiveness of virtual reality (VR) systems. In this paper, we propose MoveVR, a technique that enables realistic, multiform force feedback in VR leveraging commonplace cleaning robots. MoveVR can generate tension, resistance, impact and material rigidity force feedback with multiple levels of force intensity and directions. This is achieved by changing the robot's moving speed, rotation, position as well as the carried proxies. We demonstrated the feasibility and effectiveness of MoveVR through interactive VR gaming. In our quantitative and qualitative evaluation studies, participants found that MoveVR provides more realistic and enjoyable user experience when compared to commercially available haptic solutions such as vibrotactile haptic systems.
Yuntao Wang 0001, Zichao (Tyson) Chen, Hanchuan Li, Zhengyi Cao, Huiyi Luo, Tengxiang Zhang, Ke Ou, John Raiti, Chun Yu, Shwetak N. Patel, Yuanchun Shi
CHI11
2020 EarBuddy: Enabling On-Face Interaction via Wireless Earbuds
abstract
Past research regarding on-body interaction typically requires custom sensors, limiting their scalability and generalizability. We propose EarBuddy, a real-time system that leverages the microphone in commercial wireless earbuds to detect tapping and sliding gestures near the face and ears. We develop a design space to generate 27 valid gestures and conducted a user study (N=16) to select the eight gestures that were optimal for both human preference and microphone detectability. We collected a dataset on those eight gestures (N=20) and trained deep learning models for gesture detection and classification. Our optimized classifier achieved an accuracy of 95.3%. Finally, we conducted a user study (N=12) to evaluate EarBuddy's usability. Our results show that EarBuddy can facilitate novel interaction and that users feel very positively about the system. EarBuddy provides a new eyes-free, socially acceptable input method that is compatible with commercial wireless earbuds and has the potential for scalability and generalizability
Xuhai Xu, Haitian Shi, Xin Yi 0001, Wenjia Liu, Yukang Yan, Yuanchun Shi, Alexander Mariakakis, Jennifer Mankoff, Anind K. Dey
CHI6
2020 FrownOnError: Interrupting Responses from Smart Speakers by Facial Expressions
abstract
In the conversations with smart speakers, misunderstandings of users' requests lead to erroneous responses. We propose FrownOnError, a novel interaction technique that enables users to interrupt the responses by intentional but natural facial expressions. This method leverages the human nature that the facial expression changes when we receive unexpected responses. We conducted a first user study (N=12) to understand users' intuitive reactions to the correct and incorrect responses. Our results reveal the significant difference in the frequency of occurrence and intensity of users' facial expressions between two conditions, and frowning and raising eyebrows are intuitive to perform and easy to control. Our second user study (N=16) evaluated the user experience and interruption efficiency of FrownOnError and the third user study (N=12) explored suitable conversation recovery strategies after the interruptions. Our results show that FrownOnError can be accurately detected (precision: 97.4%, recall: 97.6%), provides the most timely interruption compared to the baseline methods of wake-up word and button press, and is rated as most intuitive and easiest to be performed by users.
Yukang Yan, Chun Yu, Wengrui Zheng, Ruining Tang, Xuhai Xu, Yuanchun Shi
CHI6
2020 PalmBoard: Leveraging Implicit Touch Pressure in Statistical Decoding for Indirect Text Entry
abstract
We investigated how to incorporate implicit touch pressure, finger pressure applied to a touch surface during typing, to improve text entry performance via statistical decoding. We focused on one-handed touch-typing on indirect interface as an example scenario. We first collected typing data on a pressure-sensitive touchpad, and analyzed users' typing behavior such as touch point distribution, key-to-finger mappings, and pressure images. Our investigation revealed distinct pressure patterns for different keys. Based on the findings, we performed a series of simulations to iteratively optimize the statistical decoding algorithm. Our investigation led to a Markov-Bayesian decoder incorporating pressure image data into decoding. It improved the top-1 accuracy from 53% to 74% over a naive Bayesian decoder. We then implemented PalmBoard, a text entry method that implemented the Markov-Bayesian decoder and effectively supported one-handed touch-typing on indirect interfaces. A user study showed participants achieved an average speed of 32.8 WPM with 0.6% error rate. Expert typists could achieve 40.2 WPM with 30 minutes of practice. Overall, our investigation showed that incorporating implicit touch pressure is effective in improving text entry decoding.
Xin Yi 0001, Chen Wang 0049, Xiaojun Bi 0001, Yuanchun Shi
CHI4
2020 Cross-VAE: Towards Disentangling Expression from Identity For Human Faces
abstract
Facial expression and identity are two independent yet intertwined components for representing a face. For facial expression recognition, identity can contaminate the training procedure by providing tangled but irrelevant information. In this paper, we propose to learn clearly disentangled and discriminative features that are invariant of identities for expression recognition. However, such disentanglement normally requires annotations of both expression and identity on one large dataset, which is often unavailable. Our solution is to extend conditional VAE to a crossed version named Cross-VAE, which is able to use partially labeled data to disentangle expression from identity. We emphasis the following novel characteristics of our Cross-VAE: (1) It is based on an independent assumption that the two latent representations' distributions are orthogonal. This ensures both encoded representations to be disentangled and expressive. (2) It utilizes a symmetric training procedure where the output of each encoder is fed as the condition of the other. Thus two partially labeled sets can be jointly used. Extensive experiments show that our proposed method is capable of encoding expressive and disentangled features for facial expression. Compared with the baseline methods, our model shows an improvement of 3.56% on average in terms of accuracy.
Haozhe Wu, Jia Jia 0001, Lingxi Xie, Guo-Jun Qi, Yuanchun Shi, Qi Tian 0001
ICASSP5
2020 Investigating Bubble Mechanism for Ray-Casting to Improve 3D Target Acquisition in Virtual Reality
abstract
Ray-casting, i.e., a ray cast from a hand-held controller to select targets, is widely used in 3D environments. Inspired by the bubble cursor [12] which dynamically resizes its selection range on 2D surfaces, we investigate a bubble mechanism for ray-casting in virtual reality. Bubble mechanism identifies the target nearest to the ray, with which users do not have to accurately shoot through the target. We first design the criterion of selection and the visual feedback of the bubble. We then conduct two experiments to evaluate ray-casting techniques with bubble mechanism in both simple and complicated 3D target acquisition tasks. Results show the bubble mechanism significantly improves ray-casting on both performance and preference, and our Bubble Ray technique with angular distance definition is competitive compared with other target acquisition techniques. We also discuss potential improvements to show more practical implementations of ray-casting with bubble mechanism.
Yiqin Lu, Chun Yu, Yuanchun Shi
VR3
2019 "I Bought This for Me to Look More Ordinary": A Study of Blind People Doing Online Shopping
abstract
Online shopping, by reducing the needs for traveling, has become an essential part of lives for people with visual impairments. However, in HCI, research on online shopping for them has only been limited to the analysis of accessibility and usability issues. To develop a broader and better understanding of how visually impaired people shop online and design accordingly, we conducted a qualitative study with twenty blind people. Our study highlighted that blind people's desire of being treated as ordinary had significantly shaped their online shopping practices: very attentive to the visual appearance of the goods even they themselves could not see and taking great pain to find and learn what commodities are visually appropriate for them. This paper reports how their trying to appear ordinary is manifested in online shopping and suggests design implications to support these practices.
Guanhong Liu, Xianghua Ding, Chun Yu, Lan Gao 0001, Xingyu Chi, Yuanchun Shi
CHI6
2019 Typing on Split Keyboards with Peripheral Vision
abstract
Split keyboards are widely used on hand-held touchscreen devices (e.g., tablets). However, typing on a split keyboard often requires eye movement and attention switching between two halves of the keyboard, which slows users down and increases fatigue. We explore peripheral typing, a superior typing mode in which a user focuses her visual attention on the output text and keeps the split keyboard in peripheral vision. Our investigation showed that peripheral typing reduced attention switching, enhanced user experience and increased overall performance (27 WPM, 28% faster) over the typical eyes-on typing mode. This typing mode can be well supported by accounting the typing behavior in statistical decoding. Based on our study results, we have designed GlanceType, a text entry system that supported both peripheral and eyes-on typing modes for real typing scenario. Our evaluation showed that peripheral typing not only well co-existed with the existing eyes-on typing, but also substantially improved the text entry performance. Overall, peripheral typing is a promising typing mode and supporting it would significantly improve the text entry performance on a split keyboard.
Yiqin Lu, Chun Yu, Shuyi Fan, Xiaojun Bi 0001, Yuanchun Shi
CHI5
2019 VIPBoard: Improving Screen-Reader Keyboard for Visually Impaired People with Character-Level Auto Correction
abstract
Modern touchscreen keyboards are all powered by the word-level auto-correction ability to handle input errors. Unfortunately, visually impaired users are deprived of such benefit because a screen-reader keyboard offers only character-level input and provides no correction ability. In this paper, we present VIPBoard, a smart keyboard for visually impaired people, which aims at improving the underlying keyboard algorithm without altering the current input interaction. Upon each tap, VIPBoard predicts the probability of each key considering both touch location and language model, and reads the most likely key, which saves the calibration time when the touchdown point misses the target key. Meanwhile, the keyboard layout automatically scales according to users' touch point location, which enables them to select other keys easily. A user study shows that compared with the current keyboard technique, VIPBoard can reduce touch error rate by 63.0% and increase text entry speed by 12.6%.
Weinan Shi, Chun Yu, Shuyi Fan, Xin Yi 0001, Xiaojun Bi 0001, Yuanchun Shi
CHI8
2019 EarTouch: Facilitating Smartphone Use for Visually Impaired People in Mobile and Public Scenarios
abstract
Interacting with a smartphone using touch input and speech output is challenging for visually impaired people in mobile and public scenarios, where only one hand may be available for input (e.g., while holding a cane) and using the loudspeaker for speech output is constrained by environmental noise, privacy, and social concerns. To address these issues, we propose EarTouch, a one-handed interaction technique that allows the users to interact with a smartphone using the ear to perform gestures on the touchscreen. Users hold the phone to their ears and listen to speech output from the ear speaker privately. We report how the technique was designed, implemented, and evaluated through a series of studies. Results show that EarTouch is easy, efficient, fun and socially acceptable to use.
Ruolin Wang, Chun Yu, Xing-Dong Yang, Yuanchun Shi
CHI5
2019 HandSee: Enabling Full Hand Interaction on Smartphone with Front Camera-based Stereo Vision
abstract
We present HandSee, a novel sensing technique that can capture the state and movement of the user's hands touching or gripping a smartphone. We place a right angle prism mirror on the front camera to achieve a stereo vision of the scene above the touchscreen surface. We develop a pipeline to extract the depth image of hands from a monocular RGB image, which consists of three components: a stereo matching algorithm to estimate the pixel-wise depth of the scene, a CNN-based online calibration algorithm to detect hand skin, and a merging algorithm that outputs the depth image of the hands. Building on the output, a substantial set of valuable interaction information, such as fingers' 3D location, gripping posture, and finger identity can be recognized concurrently. Due to this unique sensing ability, HandSee enables a variety of novel interaction techniques and expands the design space for full hand interaction on smartphones.
Chun Yu, Xiaoying Wei, Shubh Vachher, Yueting Weng, Yizheng Gu, Yuanchun Shi
CHI8
2019 Accurate and Low-Latency Sensing of Touch Contact on Any Surface with Finger-Worn IMU Sensor
abstract
Head-mounted Mixed Reality (MR) systems enable touch in­teraction on any physical surface. However, optical methods (i.e., with cameras on the headset) have difficulty in determin­ing the touch contact accurately. We show that a finger ring with Inertial Measurement Unit (IMU) can substantially im­prove the accuracy of contact sensing from 84.74% to 98.61% (f1 score), with a low latency of 10 ms. We tested different ring wearing positions and tapping postures (e.g., with different fingers and parts). Results show that an IMU-based ring worn on the proximal phalanx of the index finger can accurately sense touch contact of most usable tapping postures. Partici­pants preferred wearing a ring for better user experience. Our approach can be used in combination with the optical touch sensing to provide robust and low-latency contact detection.
Yizheng Gu, Chun Yu, Zhipeng Li 0001, Shuchang Xu, Xiaoying Wei, Yuanchun Shi
UIST7
2019 The dynamic grouping keyboard: a general keyboard optimization approach for users with motor impairment
Yizheng Gu, Chun Yu, Yuanchun Shi
CCF Trans. Pervasive Comput. Interact.3
2019 Photoplethysmogram-based Cognitive Load Assessment Using Multi-Feature Fusion Model
abstract
Cognitive load assessment is crucial for user studies and human--computer interaction designs. As a noninvasive and easy-to-use category of measures, current photoplethysmogram- (PPG) based assessment methods rely on single or small-scale predefined features to recognize responses induced by people’s cognitive load, which are not stable in assessment accuracy. In this study, we propose a machine-learning method by using 46 kinds of PPG features together to improve the measurement accuracy for cognitive load. We test the method on 16 participants through the classical n-back tasks (0-back, 1-back, and 2-back). The accuracy of the machine-learning method in differentiating different levels of cognitive loads induced by task difficulties can reach 100% in 0-back vs. 2-back tasks, which outperformed the traditional HRV-based and single-PPG-feature-based methods by 12--55%. When using “leave-one-participant-out” subject-independent cross validation, 87.5% binary classification accuracy was reached, which is at the state-of-the-art level. The proposed method can also support real-time cognitive load assessment by beat-to-beat classifications with better performance than the traditional single-feature-based real-time evaluation method.
Xiao Zhang 0008, Yongqiang Lyu 0001, Tong Qu, Pengfei Qiu, Xiaomin Luo, Shunjie Fan, Yuanchun Shi
ACM Trans. Appl. Percept.8
2019 Exploring Low-Occlusion Qwerty Soft Keyboard Using Spatial Landmarks
abstract
The Qwerty soft keyboard is widely used on mobile devices. However, keyboards often consume a large portion of the touchscreen space, occluding the application view on the smartphone and requiring a separate input interface on the smartwatch. Such space consumption can affect the user experience of accessing information and the overall performance of text input. In order to free up the screen real estate, this article explores the concept of Sparse Keyboard and proposes two new ways of presenting the Qwerty soft keyboard. The idea is to use users’ spatial memory and the reference effect of spatial landmarks on the graphical interface. Our final design K3-SGK displays only three keys while L5-EYOCN displays only five line segments instead of the entire Qwerty layout. To achieve this, we employ a user-centered computational design method: first study the reference effect of a single landmark key (line segment) from empirical data, then make assumptions to generalize the effect to multiple landmarks, and finally optimize the best designs. To make the text entry function more complete, we also design and implement gestural interactions for editing operations and non-alphabetical characters’ input. User evaluation shows that participants can quickly learn how to type with K3-SGK and L5-EYOCN . After five 15-phrase typing sessions, participants achieve 88.1%--92.8% of the full Qwerty keyboard in terms of words per minute on the smartphone and 98.4%--99.1% on the smartwatch. The differences on character and word error rate between our keyboard designs and the full Qwerty keyboard are not significant. The results of out-of-vocabulary words input are also promising. In addition, participants can quickly recall the typing skills and maintain the input performance even after a few days. User feedbacks in real application contexts show that with the low occlusion keyboard, users can acquire more information and perform less scrolling on the smartphone and achieve a higher input efficiency on the smartwatch with a more fluent input experience.
Ke Sun 0003, Chun Yu, Yuanchun Shi
ACM Trans. Comput. Hum. Interact.3
2019 Gesture-based target acquisition in virtual and augmented reality
abstract
Background Gesture is a basic interaction channel that is frequently used by humans to communicate in daily life. In this paper, we explore to use gesture-based approaches for target acquisition in virtual and augmented reality. A typical process of gesture-based target acquisition is: when a user intends to acquire a target, she performs a gesture with her hands, head or other parts of the body, the computer senses and recognizes the gesture and infers the most possible target. Methods We build mental model and behavior model of the user to study two key parts of the interaction process. Mental model describes how user thinks up a gesture for acquiring a target, and can be the intuitive mapping between gestures and targets. Behavior model describes how user moves the body parts to perform the gestures, and the relationship between the gesture that user intends to perform and signals that computer senses. Results In this paper, we present and discuss three pieces of research that focus on the mental model and behavior model of gesture-based target acquisition in VR and AR. Conclusions We show that leveraging these two models, interaction experience and performance can be improved in VR and AR environments.
Yukang Yan, Xin Yi 0001, Chun Yu, Yuanchun Shi
Virtual Real. Intell. Hardw.4
2018 Eyes-Free Target Acquisition in Interaction Space around the Body for Virtual Reality
abstract
Eyes-free target acquisition is a basic and important human ability to interact with the surrounding physical world, relying on the sense of space and proprioception. In this research, we leverage this ability to improve interaction in virtual reality (VR), by allowing users to acquire a virtual object without looking at it. We expect this eyes-free approach can effectively reduce head movements and focus changes, so as to speed up the interaction and alleviate fatigue and VR sickness. We conduct three lab studies to progressively investigate the feasibility and usability of eyes-free target acquisition in VR. Results show that, compared with the eyes-engaged manner, the eyes-free approach is significantly faster, provides satisfying accuracy, and introduces less fatigue and sickness; Most participants (13/16) prefer this approach. We also measure the accuracy of motion control and evaluate subjective experience of users when acquiring targets at different locations around the body. Based on the results, we make suggestions on designing appropriate target layout and discuss several design issues for eyes-free target acquisition in VR.
Yukang Yan, Chun Yu, Xiaojuan Ma, Yuanchun Shi
CHI6
2018 VirtualGrasp: Leveraging Experience of Interacting with Physical Objects to Facilitate Digital Object Retrieval
abstract
We propose VirtualGrasp, a novel gestural approach to retrieve virtual objects in virtual reality. Using VirtualGrasp, a user retrieves an object by performing a barehanded gesture as if grasping its physical counterpart. The object-gesture mapping under this metaphor is of high intuitiveness, which enables users to easily discover, remember the gestures to retrieve the objects. We conducted three user studies to demonstrate the feasibility and effectiveness of the approach. Progressively, we investigated the consensus of the object-gesture mapping across users, the expressivity of grasping gestures, and the learnability and performance of the approach. Results showed that users achieved high agreement on the mapping, with an average agreement score [35] of 0.68 (SD=0.27). Without exposure to the gestures, users successfully retrieved 76% objects with VirtualGrasp. A week after learning the mapping, they could recall the gestures for 93% objects.
Yukang Yan, Chun Yu, Xiaojuan Ma, Xin Yi 0001, Ke Sun 0003, Yuanchun Shi
CHI6
2018 ForceBoard: Subtle Text Entry Leveraging Pressure
abstract
We present ForceBoard, a pressure-based input technique that enables text entry by subtle finger motion. To enter text, users apply pressure to control a multi-letter-wide sliding cursor on a one-dimensional keyboard with alphabetical ordering, and confirm the selection with a quick release. We examined the error model of pressure control for successive and error-tolerant input, which was incorporated into a Bayesian algorithm to infer user input. A user study showed that, after a 10-minute training, the average text entry rate reached 4.2 wpm (Words Per Minute) for character-level input, and 11.0 wpm for word-level input. Users reported that ForceBoard was easy to learn and interesting to use. These results demonstrated the feasibility of applying pressure as the main channel for text entry. We conclude by discussing the limitation, as well as the potential of ForceBoard to support interaction with constraints from form factor, social concern and physical environments.
Mingyuan Zhong 0001, Chun Yu, Xuhai Xu, Yuanchun Shi
CHI5
2018 Lip-Interact: Improving Mobile Device Interaction with Silent Speech Commands
abstract
We present Lip-Interact, an interaction technique that allows users to issue commands on their smartphone through silent speech. Lip-Interact repurposes the front camera to capture the user's mouth movements and recognize the issued commands with an end-to-end deep learning model. Our system supports 44 commands for accessing both system-level functionalities (launching apps, changing system settings, and handling pop-up windows) and application-level functionalities (integrated operations for two apps). We verify the feasibility of Lip-Interact with three user experiments: evaluating the recognition accuracy, comparing with touch on input efficiency, and comparing with voiced commands with regards to personal privacy and social norms. We demonstrate that Lip-Interact can help users access functionality efficiently in one step, enable one-handed input when the other hand is occupied, and assist touch to make interactions more fluent.
Ke Sun 0003, Chun Yu, Weinan Shi, Yuanchun Shi
UIST5
2018 Interpreting User Input Intention in Natural Human Computer Interaction
abstract
Human Computer Interaction (HCI)1 is about information exchange between human and computers. Interaction between users and computers occurs at the User Interface (UI). Now, computers become pervasive, they are embedded in everyday things and UIs are the main value-added competitive advantages. UIs should be more natural for users. NUI (natural user interface) expands forms beyond formal input devices like the mouse and keyboard to more and more natural forms of interaction such as touch, speech, gestures, handwriting, and vision. Unlike speech, handwriting and vision, which have been researched for decades and put into practical use recently, touch and gestures are interaction tasks related, and yet lack of study. This talk will introduce methods of modelling user input action based on data with the random noise for fast touch input and natural gestures.
Yuanchun Shi
UMAP1
2018 Evaluating Photoplethysmogram as a Real-Time Cognitive Load Assessment during Game Playing
abstract
Accurate evaluation for user experience during computer game playing is very important in optimizing game design to improve the gaming experience. When evaluating user experiences, the concept of cognitive load is crucial for dynamically responding to game players’ mental status. In our former studies, the photoplethysmogram (PPG)-based Stress-Induced Vascular Response Index (sVRI) shows better sensitivity and reliability in measuring cognitive loads compared with heart-rate variation, blood pressure, and galvanic skin response. In this study, we use memory matrixes and Pop Cap’s tower defense game Plants vs. Zombies as cognitive tasks and use sVRI to assess players’ cognitive load dynamically during the real-time computer games. Evaluations on cognitive tasks verified the usability of sVRI in comparison with other indexes derived from PPG, such as heartbeat interval, area under curve, digital pulse amplitude, reflection index, and inflection point area ratio. Our findings indicate the potential of sVRI for assessing game players’ mental workload in real time.
Xiao Zhang 0008, Yongqiang Lyu 0001, Ziyue Hu, Yuanchun Shi
Int. J. Hum. Comput. Interact.5
2018 Real-Time Movie-Induced Discrete Emotion Recognition from EEG Signals
abstract
Recognition of a human's continuous emotional states in real time plays an important role in machine emotional intelligence and human-machine interaction. Existing real-time emotion recognition systems use stimuli with low ecological validity (e.g., picture, sound) to elicit emotions and to recognise only valence and arousal. To overcome these limitations, in this paper, we construct a standardised database of 16 emotional film clips that were selected from over one thousand film excerpts. Based on emotional categories that are induced by these film clips, we propose a real-time movie-induced emotion recognition system for identifying an individual's emotional states through the analysis of brain waves. Thirty participants took part in this study and watched 16 standardised film clips that characterise real-life emotional experiences and target seven discrete emotions and neutrality. Our system uses a 2-s window and a 50 percent overlap between two consecutive windows to segment the EEG signals. Emotional states, including not only the valence and arousal dimensions but also similar discrete emotions in the valence-arousal coordinate space, are predicted in each window. Our real-time system achieves an overall accuracy of 92.26 percent in recognising high-arousal and valenced emotions from neutrality and 86.63 percent in recognising positive from negative emotions. Moreover, our system classifies three positive emotions (joy, amusement, tenderness) with an average of 86.43 percent accuracy and four negative emotions (anger, disgust, fear, sadness) with an average of 65.09 percent accuracy. These results demonstrate the advantage over the existing state-of-the-art real-time emotion recognition systems from EEG signals in terms of classification accuracy and the ability to recognise similar discrete emotions that are close in the valence-arousal coordinate space.
Yong-Jin Liu 0001, Minjing Yu, Guozhen Zhao, Jinjing Song, Yan Ge 0007, Yuanchun Shi
IEEE Trans. Affect. Comput.6
2018 Non-Invasive Measurement of Cognitive Load and Stress Based on the Reflected Stress-Induced Vascular Response Index
abstract
Measuring cognitive load and stress is crucial for ubiquitous human--computer interaction applications to dynamically understand and respond to the mental status of users, such as in smart healthcare, smart driving, and robotics. Various quantitative methods have been employed for this purpose, such as physiological and behavioral methods. However, the sensitivity, reliability, and usability are not satisfactory in many of the current methods, so they are not ideal for ubiquitous applications. In this study, we employed a reflected photoplethysmogram-based stress-induced vascular response index, i.e., the reflected sVRI (sVRI-r), to non-invasively measure the cognitive load and stress. This method has high usability as well as good sensitivity and reliability compared with the previously proposed transmitted sVRI (sVRI-t). We developed the basic methodology and detailed algorithm framework to validate the sVRI-r measurements, and it was implemented by employing two light sources, i.e., infrared light and green light. Compared with the simultaneously recorded blood pressure, heart rate variation, and sVRI-t, our findings demonstrated the greater potential of the sVRI-r for use as a sensitive, reliable, and usable parameter, as well as suggesting its potential integration with ubiquitous touch interactions for dynamic cognition and stress-sensing scenarios.
Yongqiang Lyu 0001, Xiao Zhang 0008, Xiaomin Luo, Ziyue Hu, Yuanchun Shi
ACM Trans. Appl. Percept.6
2018 Real-Time Assessment of the Cross-Task Mental Workload Using Physiological Measures During Anomaly Detection
abstract
The ability to detect anomalies in perceived stimuli is critical to a broad range of practical and applied activities involving human operators. In this paper, we propose a real-time physiological-based system to assess the cross-task mental workload during anomaly detection. Forty participants were recruited to detect anomalous images from a set of different distracting images (Task I) and abnormal activities from surveillance videos (Task II). In Task I, the task difficulty levels were manipulated by changing the number of anomalies/distracting stimuli (15, 21, 28, or 36) with and without time constraints (i.e., 4 × 2 = 8 task difficulty levels). Physiological and behavioral data from four task difficulty levels were divided into four categories according to subjective ratings of the mental workload. The support vector machine (SVM) classifiers were trained on these data to predict the mental workload categories of: 1) the same four task difficulty levels (within level); and 2) the other four task difficulty levels in Task I (cross level). Within-level classifications (with an average of 95.29%) were more accurate than cross-level classifications (average of 72.2%), which were much more accurate than random level classifications (25%). In Task II, the same participants monitored one, two, or four video clips simultaneously in accordance with three task difficulty levels. The same physiological signals were processed for real-time recognition of a participant's mental workload after he or she completed each activity detection task. The three-class SVM classifiers were trained on physiological data from Task I to predict the mental workload categories of the Task II (cross task), achieving an overall classification accuracy of 53.83%, compared to a 33.33% accuracy at random. These results are discussed in terms of their implications for developing situation-aware recognition systems of the mental workload and adaptive human-computer interaction platforms.
Guozhen Zhao, Yong-Jin Liu 0001, Yuanchun Shi
IEEE Trans. Hum. Mach. Syst.3
2017 Float: One-Handed and Touch-Free Target Selection on Smartwatches
abstract
Touch interaction on smartwatches suffers from the awkwardness of having to use two hands and the "fat finger" problem. We present Float, a wrist-to-finger input approach that enables one-handed and touch-free target selection on smartwatches with high efficiency and precision using only commercially-available built-in sensors. With Float, a user tilts the wrist to point and performs an in-air finger tap to click. To realize Float, we first explore the appropriate motion space for wrist tilt and determine the clicking action (finger tap) through a user-elicitation study. We combine the photoplethysmogram (PPG) signal with accelerometer and gyroscope to detect finger taps with a recall of 97.9% and a false discovery rate of 0.4%. Experiments show that using just one hand, Float allows users to acquire targets with size ranging from 2mm to 10mm in less than 2s to 1s, meanwhile achieve much higher accuracy than direct touch in both stationary (>98.9%) and walking (>71.5%) contexts.
Ke Sun 0003, Yuntao Wang 0001, Chun Yu, Yukang Yan, Hongyi Wen, Yuanchun Shi
CHI6
2017 Word Clarity as a Metric in Sampling Keyboard Test Sets
abstract
Test sets play an essential role in evaluating text entry techniques. In this paper, we argue that in addition to the widely adopted metric of bigram representativeness and memorability, word clarity should also be considered as a metric when creating test sets from the target dataset. Word clarity quantifies the extent to which a word is likely to confuse with other words on a keyboard. We formally define word clarity, derive equations calculating it, and both theoretically and empirically show that word clarity has a significant effect on text entry performance: it can yield up to 26.4% difference in error rate, and 25% difference in input speed. We later propose a Pareto optimization method for sampling test sets with different sizes, which optimizes the word clarity and bigram representativeness, and memorability of the test set. The obtained test sets are published on the Internet.
Xin Yi 0001, Chun Yu, Weinan Shi, Xiaojun Bi 0001, Yuanchun Shi
CHI5
2017 COMPASS: Rotational Keyboard on Non-Touch Smartwatches
abstract
Entering text is very challenging on smartwatches, especially on non-touch smartwatches where virtual keyboards are unavailable. In this paper, we designed and implemented COMPASS, a non-touch bezel-based text entry technique. COMPASS positions multiple cursors on a circular keyboard, with the location of each cursor dynamically optimized during typing to minimize rotational distance. To enter text, a user rotates the bezel to select keys with any nearby cursors. The design of COMPASS was justified by an iterative design process and user studies. Our evaluation showed that participants achieved a pick-up speed around 10 WPM and reached 12.5 WPM after 90-minute practice. COMPASS allows users to enter text on non-touch smartwatches, and also serves as an alternative for entering text on touch smartwatches when touch is unavailable (e.g., wearing gloves).
Xin Yi 0001, Chun Yu, Weijie Xu, Xiaojun Bi 0001, Yuanchun Shi
CHI5
2017 Tap, Dwell or Gesture?: Exploring Head-Based Text Entry Techniques for HMDs
abstract
Despite the increasing popularity of head mounted displays (HMDs), development of efficient text entry methods on these devices has remained under explored. In this paper, we investigate the feasibility of head-based text entry for HMDs, by which, the user controls a pointer on a virtual keyboard using head rotation. Specifically, we investigate three techniques: TapType, DwellType, and GestureType. Users of TapType select a letter by pointing to it and tapping a button. Users of DwellType select a letter by pointing to it and dwelling over it for a period of time. Users of GestureType perform word-level input using a gesture typing style. Two lab studies were conducted. In the first study, users typed 10.59 WPM, 15.58 WPM, and 19.04 WPM with DwellType, TapType, and GestureType, respectively. Users subjectively felt that all three of the techniques were easy to learn and considered the induced fatigue to be acceptable. In the second study, we further investigated GestureType. We improved its gesture-word recognition algorithm by incorporating the head movement pattern obtained from the first study. This resulted in users reaching 24.73 WPM after 60 minutes of training. Based on these results, we argue that head-based text entry is feasible and practical on HMDs, and deserves more attention.
Chun Yu, Yizheng Gu, Zhican Yang, Xin Yi 0001, Hengliang Luo, Yuanchun Shi
CHI6
2017 ViVo: Video-Augmented Dictionary for Vocabulary Learning
abstract
Research on Computer-Assisted Language Learning (CALL) has shown that the use of multimedia materials such as images and videos can facilitate interpretation and memorization of new words and phrases by providing richer cues than text alone. We present ViVo, a novel video-augmented dictionary that provides an inexpensive, convenient, and scalable way to exploit huge online video resources for vocabulary learning. ViVo automatically generates short video clips from existing movies with the target word highlighted in the subtitles. In particular, we apply a word sense disambiguation algorithm to identify the appropriate movie scenes with adequate contextual information for learning. We analyze the challenges and feasibility of this approach and describe our interaction design. A user study showed that learners were able to retain nearly 30% more new words with ViVo than with a standard bilingual dictionary days after learning. They preferred our video-augmented dictionary for its benefits in memorization and enjoyable learning experience.
Yeshuang Zhu, Yuntao Wang 0001, Chun Yu, Shaoyun Shi, Yankai Zhang, Shuang He, Peijun Zhao, Xiaojuan Ma, Yuanchun Shi
CHI9
2017 CEPT: Collaborative Editing Tool for Non-Native Authors
abstract
Due to language deficiencies, individual non-native speakers (NNS) face many difficulties while writing. In this paper, we propose to build a collaborative editing system that aims to facilitate the sharing of language knowledge among non-native co-authors, with the ultimate goal of improving writing quality. We describe CEPT, which allows individual co-authors to generate their own revisions as well as incorporating edits from others to achieve mutual inspiration. The main technical challenge is how to aggregate edits of multiple co-authors and present them in an easy-to-understand way. After iterative design, CEPT highlights three novel features: 1) cross-version sentence mapping for edit tracking, 2) summarization of edits from multiple co-authors, and 3) a collaborative editing interface that enables co-authors to examine, comment on, and borrow edits of others. A preliminary lab study showed that CEPT could significantly improve both the language quality and collaboration experience of NNS writers, due to its efficacy for sharing language knowledge.
Yeshuang Zhu, Shichao Yue, Chun Yu, Yuanchun Shi
CSCW4
2017 pbSE: Phase-Based Symbolic Execution
abstract
The study of software bugs has long been a key area in software security. Dynamic symbolic execution, in exploring the program's execution paths, finds bugs by analyzing all potential dangerous operations. Due to its high coverage and abilities to generate effective testcases, dynamic symbolic execution has attracted wide attention in the research community. However, the success of dynamic symbolic execution is limited due to complex program logic and its difficulty to handle large symbolic data. In our experiments we found that phase-related features of a program often prevents dynamic symbolic execution from exploring deep paths. On the basis of this discovery, we proposed a novel symbolic execution technology guided by program phase characteristics. Compared to KLEE, the most well-known symbolic execution approach, our method is capable of covering more code and discovering more bugs. We designed and implemented pbSE system, which was used to test several commonly used tools and libraries in Linux. Our results showed that pbSE on average covers code twice as much as what KLEE does, and we discovered 21 previously unknown vulnerabilities by using pbSE, out of which 7 are assigned CVE IDs.
Qixue Xiao, Yu Chen 0004, Chengang Wu, Kang Li 0001, Junjie Mao, Shize Guo, Yuanchun Shi
DSN7
2017 BitID: Easily Add Battery-Free Wireless Sensors to Everyday Objects
abstract
Radio-Frequency Identification (RFID) systems are becoming increasingly used within smart environments. In this paper, we propose BitID, a passive Ultra-High Frequency (UHF) RFID based sensing technique that can easily be made using off-the-shelf tags. BitID can be added to everyday objects to enable sensing and control capabilities. With a simple shorting mechanism, BitID is able to differentiate between two states of the object to which it is attached (for example, whether a door is open or closed). We explain the working principle of BitID, and demonstrate how to build and apply it to target objects. We also show that by using a three- layered system architecture, BitID can be used for various applications, including event detection, energy monitoring, fitness tracking, human behavior tracking and control.
Tengxiang Zhang, Nicholas Becker, Yuntao Wang 0001, Yuanchun Shi
SMARTCOMP5
2017 Is it too small?: Investigating the performances and preferences of users when typing on tiny QWERTY keyboards
Xin Yi 0001, Chun Yu, Weinan Shi, Yuanchun Shi
Int. J. Hum. Comput. Stud.4
2016 Scalable Kernel TCP Design and Implementation for Short-Lived Connections
abstract
With the rapid growth of network bandwidth, increases in CPU cores on a single machine, and application API models demanding more short-lived connections, a scalable TCP stack is performance-critical. Although many clean-state designs have been proposed, production environments still call for a bottom-up parallel TCP stack design that is backward-compatible with existing applications.
Yu Chen 0004, Junjie Mao, Jiaquan He, Wei Xu 0005, Yuanchun Shi
ASPLOS7
2016 RID: Finding Reference Count Bugs with Inconsistent Path Pair Checking
abstract
Reference counts are widely used in OS kernels for resource management. However, reference counts are not trivial to be used correctly in large scale programs because it is left to developers to make sure that an increment to a reference count is always paired with a decrement. This paper proposes inconsistent path pair checking, a novel technique that can statically discover bugs related to reference counts without knowing how reference counts should be changed in a function. A prototype called RID is implemented and evaluations show that RID can discover more than 80 bugs which were confirmed by the developers in the latest Linux kernel. The results also show that RID tends to reveal bugs caused by developers' misunderstanding on API specifications or error conditions that are not handled properly.
Junjie Mao, Yu Chen 0004, Qixue Xiao, Yuanchun Shi
ASPLOS4
2016 One-Dimensional Handwriting: Inputting Letters and Words on Smart Glasses
abstract
We present 1D Handwriting, a unistroke gesture technique enabling text entry on a one-dimensional interface. The challenge is to map two-dimensional handwriting to a reduced one-dimensional space, while achieving a balance between memorability and performance efficiency. After an iterative design, we finally derive a set of ambiguous two-length unistroke gestures, each mapping to 1-4 letters. To input words, we design a Bayesian algorithm that takes into account the probability of gestures and the language model. To input letters, we design a pause gesture allowing users to switch into letter selection mode seamlessly. Users studies show that 1D Handwriting significantly outperforms a selection-based technique (a variation of 1Line Keyboard) for both letter input (4.67 WPM vs. 4.20 WPM) and word input (9.72 WPM vs. 8.10 WPM). With extensive training, text entry rate can reach 19.6 WPM. Users' subjective feedback indicates 1D Handwriting is easy to learn and efficient to use. Moreover, it has several potential applications for other one-dimensional constrained interfaces.
Chun Yu, Ke Sun 0003, Mingyuan Zhong 0001, Peijun Zhao, Yuanchun Shi
CHI6
2016 Investigating Effects of Post-Selection Feedback for Acquiring Ultra-Small Targets on Touchscreen
abstract
In this paper, we investigate the effects of post-selection feedback for acquiring ultra-small (2-4mm) targets on touchscreens. Post-selection feedback shows the contact point on touchscreen after a user lifts his/her fingers to increase users' awareness of touching. Three experiments are conducted progressively using a single crosshair target, two reciprocally acquired targets and 2D random targets. Results show that in average post-selection feedback can reduce touch error rates by 78.4%, with a compromise of target acquisition time no more than 10%. In addition, we investigate participants' adjustment behavior based on correlation between successive trials. We conclude that the benefit of post-selection feedback is the outcome of both improved understanding about finger/point mapping and the dynamic adjustment of finger movement enabled by the visualization of the touch point.
Chun Yu, Hongyi Wen, Xiaojun Bi 0001, Yuanchun Shi
CHI5
2015 Measuring Photoplethysmogram-Based Stress-Induced Vascular Response Index to Assess Cognitive Load and Stress
abstract
Quantitative assessment for cognitive load and mental stress is very important in optimizing human-computer system designs to improve performance and efficiency. Traditional physiological measures, such as heart rate variation (HRV), blood pressure and electrodermal activity (EDA), are widely used but still have limitations in sensitivity, reliability and usability. In this study, we propose a novel photoplethysmogram-based stress induced vascular index (sVRI) to measure cognitive load and stress. We also provide the basic methodology and detailed algorithm framework. We employed a classic experiment with three levels of task difficulty and three stages of testing period to verify the new measure. Compared with the blood pressure, heart rate and HRV components recorded simultaneously, the sVRI reached the same level of significance on the effect of task difficulty/period as the most significant other measure. Our findings showed sVRI's potential as a sensitive, reliable and usable parameter.
Yongqiang Lyu 0001, Xiaomin Luo, Chun Yu, Congcong Miao, Yuanchun Shi, Ken-ichi Kameyama
CHI7
2015 Mitigating Code-Reuse Attacks on CISC Architectures in a Hardware Approach
Zhijiao Zhang, Ya-Shuai Lü, Yu Chen 0004, Yongqiang Lyu 0001, Yuanchun Shi
SEC5
2015 ATK: Enabling Ten-Finger Freehand Typing in Air Based on 3D Hand Tracking Data
abstract
Ten-finger freehand mid-air typing is a potential solution for post-desktop interaction. However, the absence of tactile feedback as well as the inability to accurately distinguish tapping finger or target keys exists as the major challenge for mid-air typing. In this paper, we present ATK, a novel interaction technique that enables freehand ten-finger typing in the air based on 3D hand tracking data. Our hypothesis is that expert typists are able to transfer their typing ability from physical keyboards to mid-air typing. We followed an iterative approach in designing ATK. We first empirically investigated users' mid-air typing behavior, and examined fingertip kinematics during tapping, correlated movement among fingers and 3D distribution of tapping endpoints. Based on the findings, we proposed a probabilistic tap detection algorithm, and augmented Goodman's input correction model to account for the ambiguity in distinguishing tapping finger. We finally evaluated the performance of ATK with a 4-block study. Participants typed 23.0 WPM with an uncorrected word-level error rate of 0.3% in the first block, and later achieved 29.2 WPM in the last block without sacrificing accuracy.
Xin Yi 0001, Chun Yu, Mingrui Ray Zhang, Sida Gao, Ke Sun 0003, Yuanchun Shi
UIST6
2015 RegionalSliding: Facilitating small target selection with marking menu for one-handed thumb use on touchscreen-based mobile devices
Wenchang Xu, Chun Yu, Jie Liu 0027, Yuanchun Shi
Pervasive Mob. Comput.4
2015 Requester-Based Spin Lock: A Scalable and Energy Efficient Locking Scheme on Multicore Systems
abstract
In response to the increasing ubiquity of multicore processors, applications are usually designed or deployed to make each core busy. Unfortunately, lock contention within operating systems can limit the scalability of multicore systems so severely that an increase in the number of cores can actually lead to reduced performance (i.e., scalability collapse). Existing lock implementations have disadvantages in scalability, power consumption, and energy efficiency. In this paper, we observe that the number of tasks requesting a lock has a significant correlation with the occurrence of scalability collapse. Based on this observation, a lock implementation that allows tasks waiting for a lock to either spin or enter a power-saving state based on the number of requesters is proposed. Our lock protocol is called requester-based lock and is implemented in the Linux kernel to replace its default spin lock. Based on the results of a sensitivity analysis, we find that the best policy, in practice, for a task waiting for a lock to be granted is to enter the power-saving state immediately after noticing the lock cannot be acquired. Our requester-based lock scheme is evaluated using intensive benchmarking on AMD 32-core and Intel 40-core systems. Experimental results suggest that our lock avoids scalability collapse completely for most applications and shows better scalability, power consumption, and energy efficiency than previous work. Besides, the requester-based lock is extensible, which means using together with other kinds of spin locks can provide better scalability and energy efficiency.
Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi
IEEE Trans. Computers4
2015 LockSim: An Event-Driven Simulator for Modeling Spin Lock Contention
abstract
Spin lock contention in operating systems can limit scalability on multicore systems so significantly that an increase in the number of cores actually leads to reduced speedup (i.e., scalability collapse). Modeling spin lock contention is an effective way to understand the scalability collapse phenomenon and explore collapse avoidance schemes. However, previous spin lock models have disadvantages in accuracy and efficiency. To overcome these drawbacks, this paper proposes LockSim, an event-driven simulator which models both the sequential execution in lock-protected codes (i.e., critical sections) and shared hardware resource contention caused by the cache coherence protocol. Our simulator is verified against real-world workloads with different degrees of spin lock contention. Experimental results suggest that LockSim can reproduce the scalability collapse phenomenon with better accuracy than previous work. Besides, several metrics are also used to characterize this phenomenon and collapse avoidance methods are investigated.
Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi
IEEE Trans. Parallel Distributed Syst.4
2014 Running Multiple Androids on One ARM Platform
Zhijiao Zhang, Lei Zhang 0060, Yu Chen 0004, Yuanchun Shi
ACISP4
2014 FOCUS: enhancing children's engagement in reading by using contextual BCI training sessions
abstract
Reading is an important aspect of a child's development. Reading outcome is heavily dependent on the level of engagement while reading. In this paper, we present FOCUS, an EEG-augmented reading system which monitors a child's engagement level in real time, and provides contextual BCI training sessions to improve a child's reading engagement. A laboratory experiment was conducted to assess the validity of the system. Results showed that FOCUS could significantly improve engagement in terms of both EEG-based measurement and teachers' subjective measure on the reading outcome.
Chun Yu, Yuntao Wang 0001, Yuhang Zhao 0001, Chou Mo, Jie Liu 0027, Lie Zhang, Yuanchun Shi
CHI9
2014 QOOK: enhancing information revisitation for active reading with a paper book
abstract
Revisiting information on previously accessed pages is a common activity during active reading. Both physical and digital books have their own benefits in supporting such activity according to their manipulation natures. In this paper, we introduce QOOK, a paper-book based interactive reading system, which integrates the advanced technology of digital books with the affordances of physical books to facilitate people's information revisiting process. The design goals of QOOK are derived from the literature survey and our field study on physical and digital books respectively. QOOK allows page flipping just like on a real book and enables people to use electronic functions such as keyword searching, highlighting and bookmarking. A user study is conducted and the study results demonstrate that QOOK brings faster information revisiting and better reading experience to readers.
Yuhang Zhao 0001, Yongqiang Qin, Taoshuai Zhang, Yuanchun Shi
TEI6
2014 Mitigating Resource Contention on Multicore Systems via Scheduling
abstract
This paper addresses the resource contention issue caused by the sharing of the last level caches by introducing a novel contention-aware scheduler. To accurately determine a task's resource requirements, i.e. the unique input of our scheduler, we develop a methodology to select the best heuristic metric from five candidates to represent a task's resource requirements. Based on the heuristic of each task acquired by exploiting the performance monitor unit, our scheduler co-schedules tasks with complementary resource requirements by combining scheduling order adjustments with task-to-core reassignments. The proposed scheduler has been implemented in the completely fair scheduler, rotating staircase deadline scheduler and O(1) schedulers. Using eight workloads constructed from nine NASA advanced supercomputing serial benchmarks on an Intel dual-core platform, the execution time of an individual task is reduced by up to 21%, system scalability and performance of a workload are improved by up to 13%, and the full potential of the contention-aware scheduling can be achieved if the time slice length and the period of executing each task once are short enough. In addition, our proposal exhibits benefits in the reduction of the execution time fluctuation of individual tasks due to an enforcement of reasonable usage of shared resources. Finally, we demonstrate an expected performance improvement on an Intel eight-core platform in order to suggest the broad applicability of our protocol.
Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi
Comput. J.4
2014 Towards scalability collapse behavior on multicores
abstract
SUMMARY Multicore processor systems have become mainstream. To release the full potential of multiple cores, applications are programmed to be parallel to keep every core busy. Unfortunately, lock contention within operating systems can limit the scalability so seriously that use of more cores leads to reduced throughput (scalability collapse). To understand and characterize the collapse behavior easily, a discrete‐event simulation model, which considers both the sequential execution of critical sections and the overhead of hardware resource contention, is designed and implemented. By the use of the model, we observe that the percentage of time used to wait for locks and the number of tasks requesting for a lock have a significant correlation with the occurrence of scalability collapse. On the basis of these observations, two new techniques (lock contention aware scheduler and requester‐based adaptive lock) are proposed to remove the scalability collapse on multicores. The proposed methods are implemented in the Linux kernel 2.6.29.4 and evaluated on an AMD 32‐core system to verify their effectiveness. By using micro‐benchmarks and macro‐benchmarks, we find that these methods can remove scalability collapse totally for four of five workloads exhibiting the collapse behavior. For one workload that does not suffer scalability collapse, these proposed methods only introduce negligible overhead. Copyright © 2012 John Wiley & Sons, Ltd.
Yan Cui 0002, Yu Chen 0004, Yuanchun Shi
Concurr. Comput. Pract. Exp.3
2013 Facilitating parallel web browsing through multiple-page view
abstract
Parallel web browsing describes the behavior where users visit web pages in multiple concurrent threads. Qualitative studies have observed this activity being performed with multiple browser windows or tabs. However, these solutions are not satisfying since a large amount of time is wasted on switch among windows and tabs. In this paper, we propose the multiple-page view to facilitate parallel web browsing. Specifically, we provide users with the experience of visiting multiple web pages in one browser window and tab with extensions of prevalent desktop web browsers. Through user study and survey, we found that 2-4 pages within the window size were preferred for multiple-page view in spite of the diverse screen sizes and resolutions. Analytical results of logs from the user study also showed an improvement of 26.3% in users' efficiency of performing parallel web browsing tasks, compared to traditional browsing with multiple windows or tabs.
Wenchang Xu, Chun Yu, Songmin Zhao, Jie Liu 0027, Yuanchun Shi
CHI5
2013 Understanding performance of eyes-free, absolute position control on touchable mobile phones
abstract
Many eyes-free interaction techniques have been proposed for touchscreens, but few researches have studied human's eyes-free pointing ability with mobile phones. In this paper, we investigate the single-handed thumb performance of eyes-free, absolute position control on mobile touch screens. Both 1D and 2D experiments were conducted. We explored the effects of target size and location on eyes-free touch patterns and accuracy. Our findings show that variance of touch points per target will converge as target size decreases. The centroid of touch points per target tends to be offset to the left of target center along horizontal direction, and shift toward screen center along vertical direction. Average accuracy drops from 99.6% of 2×2 layout to 85.0% of 4×4 layout, and average per target varies depending on the location of target. Our findings and design implications provide a foundation for future researches based on eyes-free, absolute position control using thumb on mobile devices.
Yuntao Wang 0001, Chun Yu, Jie Liu 0027, Yuanchun Shi
Mobile HCI4
2013 Implicit bookmarking: Improving support for revisitation in within-document reading tasks
Chun Yu, Ravin Balakrishnan, Ken Hinckley, Tomer Moscovich, Yuanchun Shi
Int. J. Hum. Comput. Stud.5
2013 Lock-contention-aware scheduler: A scalable and energy-efficient method for addressing scalability collapse on multicore systems
abstract
In response to the increasing ubiquity of multicore processors, there has been widespread development of multithreaded applications that strive to realize their full potential. Unfortunately, lock contention within operating systems can limit the scalability of multicore systems so severely that an increase in the number of cores can actually lead to reduced performance (i.e., scalability collapse). Existing efforts of solving scalability collapse mainly focus on making critical sections of kernel code fine-grained or designing new synchronization primitives. However, these methods have disadvantages in scalability or energy efficiency. In this article, we observe that the percentage of lock-waiting time over the total execution time for a lock intensive task has a significant correlation with the occurrence of scalability collapse. Based on this observation, a lock-contention-aware scheduler is proposed. Specifically, each task in the scheduler monitors its percentage of lock waiting time continuously. If the percentage exceeds a predefined threshold, this task is considered as lock intensive and migrated to a Special Set of Cores (i.e., SSC). In this way, the number of concurrently running lock-intensive tasks is limited to the number of cores in the SSC, and therefore, the degree of lock contention is controlled. A central challenge of using this scheme is how many cores should be allocated in the SSC to handle lock-intensive tasks. In our scheduler, the optimal number of cores is determined online by the model-driven search. The proposed scheduler is implemented in the recent Linux kernel and evaluated using micro- and macrobenchmarks on AMD and Intel 32-core systems. Experimental results suggest that our proposal is able to remove scalability collapse completely and sustains the maximal throughput of the spin-lock-based system for most applications. Furthermore, the percentage of lock-waiting time can be reduced by up to 84%. When compared with scalability collapse reduction methods such as requester-based locking scheme and sleeping-based synchronization primitives, our scheme exhibits significant advantages in scalability, power consumption, and energy efficiency.
Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi
ACM Trans. Archit. Code Optim.4
2012 Clustering web pages to facilitate revisitation on mobile devices
abstract
Due to small screens, inaccuracy of input and other limitations of mobile devices, revisitation of Web pages in mobile browsers takes more time than that in desktop browsers. In this paper, we propose a novel approach to facilitate revisitation. We designed AutoWeb, a system that clusters opened Web pages into different topics based on their contents. Users can quickly find a desired opened Web page by narrowing down the searching scope to a group of Web pages that share the same topic. Clustering accuracy is evaluated to be 92.4% and computing resource consumption was proved to be acceptable. A user study was conducted to explore user experience and how much AutoWeb facilitates revisitation. Results showed that AutoWeb could save up a significant time for revisitation and participants rated the system highly.
Jie Liu 0027, Chun Yu, Wenchang Xu, Yuanchun Shi
IUI4
2012 Reducing Scalability Collapse via Requester-Based Locking on Multicore Systems
abstract
In response to the increasing ubiquity of multicore processors, there has been widespread development of multithreaded applications that strive to realize their full potential. Unfortunately, lock contention within operating systems can limit the scalability of multicore systems so severely that an increase in the number of cores can actually lead to reduced performance (i.e. scalability collapse). Existing lock implementations have disadvantages in scalability, resource utilization and energy efficiency. In this work, we observe that the number of tasks requesting a lock has a significant correlation with the occurrence of scalability collapse. Based on this observation, we propose a novel lock implementation that allows tasks blocked on a lock to either spin or maintain a power-saving state according to the number of lock requesters. We call our lock implementation protocol a requester-based lock and implement it in the Linux kernel to replace its default spin lock. Based on the results of an analysis, we find that the best policy for a task waiting for a lock to become free is to enter the power saving state immediately after noticing that the lock cannot be acquired. Our lock-requester based lock scheme is evaluated using micro- and macro-benchmarks on AMD 32-core and Intel 40-core systems. Experimental results indicate our lock scheme removes scalability collapse completely for most applications. Furthermore, our method shows better scalability and energy efficiency than mutex locks and adaptive locks.
Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi
MASCOTS4
2012 AutoWeb: automatic classification of mobile web pages for revisitation
abstract
Revisitation in mobile Web browsers takes more time than that in desktop browsers due to the limitations of mobile phones. In this paper, we propose AutoWeb, a novel approach to speed up revisitation in mobile Web browsing. In AutoWeb, opened Web pages are automatically classified into different groups based on their contents. Users can more quickly revisit an opened Web page by narrowing down search scope into a group of pages that share the same topic. We evaluated the classification accuracy and the accuracy is 92.4%. Three experiments were conducted to investigate revisitation performance in three specific tasks. Results show AutoWeb can save significant time for revisitation by 29.5%, especially for long time Web browsing, and that it improves overall mobile Web revisitation experience. We also compare automatic classification with other revisitation methods.
Jie Liu 0027, Wenchang Xu, Yuanchun Shi
Mobile HCI3
2012 Digging unintentional displacement for one-handed thumb use on touchscreen-based mobile devices
abstract
There is usually an unaware screen distance between initial contact and final lift-off when users tap on touchscreen-based mobile devices with their fingers, which may affect users' target selection accuracy, gesture performance, etc. In this paper, we summarize such case as unintentional displacement and give its models under both static and dynamic scenarios. We then conducted two user studies to understand unintentional displacement for the widely-adopted one-handed thumb use on touchscreen-based mobile devices under both scenarios respectively. Our findings shed light on the following four questions: 1) what are the factors that affect unintentional displacement; 2) what is the distance range of the displacement; 3) how is the distance varying over time; 4) how are the unintentional points distributed around the initial contact point. These results not only explain certain touch inaccuracy, but also provide important reference for optimization and future design of UI components, gestures, input techniques, etc.
Wenchang Xu, Jie Liu 0027, Chun Yu, Yuanchun Shi
Mobile HCI4
2012 A Scalable Distributed Architecture for Intelligent Vision System
abstract
The complexity of intelligent computer vision systems demands novel system architectures that are capable of integrating various computer vision algorithms into a working system with high scalability. The real-time applications of human-centered computing are based on multiple cameras in current systems, which require a transparent distributed architecture. This paper presents an application-oriented service share model for the generalization of vision processing. Based on the model, a vision system architecture is presented that can readily integrate computer vision processing and make application modules share services and exchange messages transparently. The architecture provides a standard interface for loading various modules and a mechanism for modules to acquire inputs and publish processing results that can be used as inputs by others. Using this architecture, a system can load specific applications without considering the common low-layer data processing. We have implemented a prototype vision system based on the proposed architecture. The latency performance and 3-D track function were tested with the prototype system. The architecture is scalable and open, so it will be useful for supporting the development of an intelligent vision system, as well as a distributed sensor system.
Guojian Wang, Linmi Tao, Huijun Di, Xiyong Ye, Yuanchun Shi
IEEE Trans. Ind. Informatics5
2012 How Much to Share: A Repeated Game Model for Peer-to-Peer Streaming under Service Differentiation Incentives
abstract
In this paper, we propose a service differentiation incentive for P2P streaming system, according to peers' instant contributions. Also, a repeated game model is designed to analyze how much the peers should contribute in each round under this incentive. Simulations show that satisfying streaming quality is achieved in the Nash Equilibrium state.
Xin Xiao 0005, Qian Zhang 0001, Yuanchun Shi
IEEE Trans. Parallel Distributed Syst.3
2011 Experience on Comparison of Operating Systems Scalability on the Multi-core Architecture
abstract
Multi-core processor architectures have become ubiquitous in today's computing platforms, especially in parallel computing installations, with their power and cost advantages. While the technology trend continues towards having hundreds of cores on a chip in the foreseeable future, an urgent question posed to system designers as well as application users is whether applications can receive sufficient support on today's operating systems for them to scale to many cores. To this end, people need to understand the strengths and weaknesses on their support on scalability and to identify major bottlenecks limiting the scalability, if any. As open-source operating systems are of particular interests in the research and industry communities, in this paper we choose three operating systems (Linux, Solaris and FreeBSD) to systematically evaluate and compare their scalability by using a set of highly-focused microbenchmarks for broad and detailed understanding their scalability on an AMD 32-core system. We use system profiling tools and analyze kernel source codes to find out the root cause of each observed scalability bottleneck. Our results reveal that there is no single operating system among the three standing out on all system aspects, though some system(s) can prevail on some of the system aspects. For example, Linux outperforms Solaris and FreeBSD significantly for file-descriptor and process-intensive operations. For applications with intensive sockets creation and deletion operations, Solaris leads FreeBSD, which scales better than Linux. With the help of performance tools and source code instrumentation and analysis, we find that synchronization primitives protecting shared data structures in the kernels are the major bottleneck limiting system scalability. Empowered by the knowledge obtained through targeted experiments and analysis on a small-scale system, we are able to project the scalability of an application on any of the investigated operating systems running on a system of a larger number of cores.
Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi
CLUSTER4
2011 Surprise Grabber: a co-located tangible social game using phone hand gesture
abstract
Social network games (SNGs) are among the most popular games recently. Different from the asynchronous and online based SNGs, we present Surprise Grabber to see how tangible gesture interface could benefit the synchronous co-located social game. In Surprise Grabber, users control a virtual grabber's moving in 3D game to catch the gifts by using their camera phone. An efficient code running on the phone detects hand motion, delivers results to Serve PC and provides feedbacks in real time. Distinguished from online SNGs, all players stand together in front of a public display. The main results of the pilot user studies showed that: 1) Gesture interface was easy to catch up and made the game more immersive; 2) Occasionally inaccuracy in hand motion detection made the game more competitive instead of frustrating players; 3) Players were eager to share their experience by talking while playing; 4) Players' performances were obviously influenced by the collocated playing social atmosphere; 5) In most cases, players' performances became better or worse at the same time instead of random results.
Mingming Fan 0001, Xin Li 0215, Yuanchun Shi, Hao Wang 0021
CSCW5
2011 A rotation based method for detecting on-body positions of mobile devices
abstract
We present a novel rotation based method for detecting where a mobile device is worn on a user's body that utilizes the fusion of the data from accelerometer and gyroscope. Detecting the position of a mobile device could improve the performance of on-body sensor based human activity recognition and the adaptability of many mobile applications. In our method, the radius and angular velocity for a position is calculated based on the data read from the sensors integrated in a mobile device. We have evaluated our method with an experiment to detect four commonly used positions: breast pocket, trouser pocket, hip pocket and hand.
Yuanchun Shi, Jie Liu 0027
UbiComp2
2011 Air finger: enabling multi-scale navigation by finger height above the surface
abstract
We present Air Finger, a novel technique that enables controlling CD ratio by finger height above the touch sur-face for multi-scale navigation tasks. Extending previous research on virtual touch, Air Finger divides the space above surface into two layers and associates the high, medium and low CD ratios to the touch surface, the lower air and the higher air respectively. Users can fluidly switch between the three navigation scales by lifting and pressing the finger. Air Finger enables multi-scale navigation control using one hand
Chun Yu, Yuanchun Shi
UbiComp4
2011 Smart home on smart phone
abstract
Mobile phone with high accessibility and usability is regarded as the ideal interface for the users to monitor and control the approaching smart home environment. Moreover, networking technologies and protocols have been advanced enough to support a universal monitoring and controlling interface on smart phones. This paper presents HouseGenie, an interactive, direct manipulation application on mobile, which supports a range of basic home monitoring and controlling functionalities as a replacement of individual remotes of smart home appliances. HouseGenie also addresses several common requirements that may be behind the vision, such as scenario, short-delay alarm, area restriction and so on. We demonstrate that HouseGenie not only provides intuitive presentations and interactions for smart home management, but also improves user experience comparing to present solutions.
Yue Suo, Wenchang Xu, Chun Yu, Yuhang Zhao 0001, Yuanchun Shi
UbiComp7
2010 Structured laser pointer: enabling wrist-rolling movements as a new interactive dimension
abstract
In this paper, we re-visit the issue of multi-point laser pointer interaction from a wrist-rolling perspective. Firstly, we proposed SLP---Structured Laser Pointer, and detects a laser pointer's rotation along its emitting axis. SLP adds the wrist-rolling gestures as a new interactive dimension to the conventional laser pointer interaction approach. We asked a group of users to perform certain tasks using SLP, and derived from test results a set of criteria to distinguish between incidental and intentional SLP rolling, and then the experimental results also approved the high accuracy and acceptable speed as well as throughput of such rolling interaction.
Yongqiang Qin, Yuanchun Shi, Chun Yu
AVI2
2010 HyMTO: The Hybrid Mesh/Tree Overlay for Large Scale Multimedia Interactive Applications over the Internet
abstract
In the large scale multimedia interactive applications over the Internet, data can be categorized into the streaming data (e.g. the audio and the video streaming) and the sporadic data (e.g. the annotation in a whiteboard). The mesh-pull overlay fits for the streaming data, for it is resilient in heterogeneous network, while the tree-push overlay is more suitable for the sporadic data, for its easy-to-manage structure can solve the problem of packet loss and disorder. To address the transmission of both types of data, a Hybrid Mesh/Tree Overlay (HyMTO) is proposed in this paper, where the sporadic data and the streaming data are transmitted through different overlays with a novel synchronous scheme. HyMTO can provide the resilient, reliable and ordered transmission for both types of data respectively. Simulations are presented to evaluate the effectiveness and efficiency of HyMTO and a practical e-conference system based on HyMTO is implemented and deployed to demonstrate its practicability and reasonability.
Yuanchun Shi, Xin Xiao 0005, Jianhua Shen
ICC2
2010 A Discrete Event Simulation Model for Understanding Kernel Lock Thrashing on Multi-core Architectures
abstract
Multi-core architectures have become mainstream. Trends suggest that the number of cores integrated on a single chip will increase continuously. However, lock contention in operating systems can limit the parallel scalability on multi-cores so significantly that the speedup decreases with the increasing number of cores (thrashing). Although the phenomenon can be easily reproduced experimentally, most existing lock models are not able to do so. To overcome this challenge, this paper develops a discrete event simulation model which has the capability of capturing both the sequential execution in critical sections and the contention for shared hardware resources. The model is evaluated using a series of typical parameter configurations which can represent different degrees of lock contention. Experimental results suggest that the thrashing phenomenon can be observed when the model parameters are selected properly. To further understand this phenomenon, statistics such as the percentage of time spent waiting for locks and the number of cores waiting for a lock are exploited to characterize the lock thrashing. In addition, the model sensitivity to changes in memory latency and hardware architectures are also examined. Finally, we use this model to compare three methods which are proposed for preventing the lock thrashing.
Yan Cui 0002, Weiyi Wu, Yingxin Wang, Xufeng Guo, Yu Chen 0004, Yuanchun Shi
ICPADS6
2010 A Scheduling Method for Avoiding Kernel Lock Thrashing on Multi-cores
abstract
Multi-core architectures have been adopted in various computing environments. Predictions based on Moore's Law state that thousands of cores can be integrated on a single chip within 10 years. To achieve better performance and scalability on multi-cores, applications should be multi-threaded, and therefore threads assigned on different cores can execute concurrently. However, lock contention in kernels can affect the scalability so significantly that the speedup decreases with the increasing number of cores (thrashing). Existing efforts to address this problem mainly focus on deferring lock thrashing, and therefore these techniques cannot prevent thrashing fundamentally. In this paper, we propose to use lock-aware scheduling to avoid thrashing. Our method detects thrashing on a per-thread basis and migrates contended threads to a smaller set of cores. The optimal number of cores is determined by maximizing the proposed normalized throughput model of migrated threads. The proposed method is implemented in Linux 2.6.29.4 and evaluated on a 32-core system. Experimental results on a series of lock-intensive micro- and macro-benchmarks show the effectiveness: for 3 of 5 workloads exhibiting thrashing behaviour, lock-aware scheduling can detect the speedup decrease accurately and sustain the maximal speedup, for the remaining 2 workloads, the performance can be improved greatly although the maximal speedup is not sustained, for 1 workload which does not suffer thrashing, the method introduces negligible runtime overhead.
Yan Cui 0002, Weida Zhang, Yu Chen 0004, Yuanchun Shi
ICPADS4
2010 Scaling OLTP applications on commodity multi-core platforms
abstract
Multi-core processor architectures can have significant performance advantage over traditional single core designs, which are limited by power and processor complexity. Predictions based on Moore's Law state that a processor chip may accommodate thousands of cores in 5–10 years. Can software scale with the number of cores and achieve the performance potential?
Yan Cui 0002, Yu Chen 0004, Yuanchun Shi
ISPASS3
2010 Scalability comparison of commodity operating systems on multi-cores
abstract
In this paper, we evaluate and compare the parallel scalability of three commodity operating systems (Linux, Solaris and FreeBSD) on an AMD 32-core platform. Measurements of microbenchmarks and a real-life application reveal that no operating system scales totally better than another for microbenchmarks; for the real-life application, Linux and Solaris are competitive in scalability and perform better than FreeBSD. Related kernel source analysis and performance data suggest that synchronization primitives protecting the shared data structure in kernels are the root cause of the poor scalability on multi-cores.
Yan Cui 0002, Yu Chen 0004, Yuanchun Shi, Qingbo Wu 0003
ISPASS3
2010 Reinventing Lock Modeling for Multi-Core Systems
abstract
Multi-core architectures have become mainstream. Trends suggest that the number of cores integrated on a single chip will continue to increase. However, lock contention in applications or kernels can degrade the scalability so significantly that the speedup decreases with the increasing number of cores (thrashing). Although the phenomenon can be easily reproduced on real multi-core platforms, existing lock models are not able to do so. To overcome the disadvantage, this paper proposes an analysis model which has the capability of capturing both the sequential execution of critical sections and the overhead of lock implementation. Numerical results indicate that thrashing can be observed by using the proposed model. Furthermore, this model can also be exploited to compare different mechanisms designed for avoiding the lock thrashing.
Yan Cui 0002, Weiyi Wu, Yingxin Wang, Xufeng Guo, Yu Chen 0004, Yuanchun Shi
MASCOTS6
2010 A Low-Cost Ubiquitous Family Healthcare Framework
Yongqiang Lyu 0001, Lei Zhang 0060, Yu Chen 0004, Yingjie Ren, Weikang Yang, Yuanchun Shi
UIC7
2010 The satellite cursor: achieving MAGIC pointing without gaze tracking using multiple cursors
abstract
We present the satellite cursor - a novel technique that uses multiple cursors to improve pointing performance by reducing input movement. The satellite cursor associates every target with a separate cursor in its vicinity for pointing, which realizes the MAGIC (manual and gaze input cascade) pointing method without gaze tracking. We discuss the problem of visual clutter caused by multiple cursors and propose several designs to mitigate it. Two controlled experiments were conducted to evaluate satellite cursor performance in a simple reciprocal pointing task and a complex task with multiple targets of varying layout densities. Results show the satellite cursor can save significant mouse movement and consequently pointing time, especially for sparse target layouts, and that satellite cursor performance can be accurately modeled by Fitts' Law.
Chun Yu, Yuanchun Shi, Ravin Balakrishnan, Xiangliang Meng, Yue Suo, Mingming Fan 0001, Yongqiang Qin
UIST2
2010 A Policy-Driven Service Composition Method for Adaptation in Pervasive Computing Environment
abstract
Service composition allows distributed application, such as multimedia application, to be composed from atomic service units and to adapt dynamically to users' requirements and environment conditions in pervasive computing system. It augments the adaptation action space for the application of pervasive computing. According to the multidimensional QoS (Quality of Service) requirement of pervasive computing system, we proposed a comprehensive service composition method to enhance the capability of application adaptation. First, according to a hierarchy policy model and a policy specification language, strengthened by event calculus, service discovery policy action integrating the situation of user, application, environment and resource can be triggered. Secondly, the proposed physical space model can support the location-aware service discovery and explicit range query to improve the efficiency of the query. To the end, an adaptation policy evaluation model is utilized to maximize an evaluation criterion–quality of satisfaction of users and environment by optimizing the optional service selection and the composition path. Through experiment and discussion of the algorithm, the paper further illustrates the great potential advantage of the solution to service composition.
Baopeng Zhang, Yuanchun Shi, Xin Xiao 0005
Comput. J.2
2010 Toward Systematical Data Scheduling for Layered Streaming in Peer-to-Peer Networks: Can We Go Farther?
abstract
Layered streaming in P2P networks has become a hot topic recently. However, the "layered" feature makes the data scheduling quite different from that for nonlayered streaming, and it hasn't been systematically studied yet. In this paper, first, according to the unique characteristics caused by layered coding, we present four objectives that should be addressed by scheduling: throughput, layer delivery ratio, useless packets ratio, and subscription jitter prevention; then a three-stage scheduling approach LayerP2P is designed to request data, where the min-cost flow model, probability decision mechanism, and multiwindow remedy mechanism are used in Free Stage, Decision Stage, and Remedy Stage, respectively, to collaboratively achieve the above objectives. With the basic version of LayerP2P and corresponding experiment results achieved in our previous work, in this paper, more efforts are put on its mechanism details and analysis to its unique features; besides, to further guarantee the performance under sharp bandwidth variation, we propose the enhanced approach by improving the Decision Stage strategy. Extensive experiments by simulation and real network implementation indicate that it outperforms other schemes. LayerP2P has also been deployed in PDEPS Project in China, which is expected to be the first practical layered streaming system for education in P2P networks.
Xin Xiao 0005, Yuanchun Shi, Qian Zhang 0001, Jianhua Shen
IEEE Trans. Parallel Distributed Syst.2
2009 CFS Optimizations to KVM Threads on Multi-Core Environment
abstract
Multi-core architecture provides more on-chip parallelism and powerful computational capability. It helps virtualization achieve scalable performance. KVM (kernel based virtual machine) is different from other virtualization solutions which can make use of the Linux kernel components such as completely fair scheduler (CFS). However, CFS treats the KVM threads as normal tasks without considering about their unique features such as thread allocation mechanism and lock inside guest virtual machine, which may harm the KVM virtualization performance. In this paper, we analyze a phenomenon that some guest multi-threaded applications have very low performance when scheduled by CFS. As a solution to this problem, we introduce two kinds of optimizations in CFS: (1) configuration optimizations (2) lock optimizations. Our contributions are: (1) implement 5 original and 2 newest proposed optimizations in the newest Linux kernel. (2) Classify and compare them, a brief analysis is also given. They are all very simple and general to other virtual machine monitors such as Xen and schedulers as O(1). The performance of our CFS optimizations to KVM threads is measured by running some well-known benchmarks in two guest virtual machines on an 8-core server which models the real world applications. The results indicate our scheduling optimizations can improve the overall system performance. This paper can provide useful advices to KVM developers and virtualization data center administrators.
Yisu Zhou, Yan Cui 0002, Yu Chen 0004, Yuanchun Shi, Qingbo Wu 0003
ICPADS6
2009 LayerP2P: A New Data Scheduling Approach for Layered Streaming in Heterogeneous Networks
abstract
Although layered streaming in heterogeneous peer-to-peer networks has drawn great interest in recent years, there's still a lack of systematical studies on its data scheduling issue. In this paper, we propose a new scheduling approach for layered video streaming, called LayerP2P. The key idea and main contributions of LayerP2P come in two-fold: 1) According to the characteristics caused by layered coding, we propose four objectives that should be achieved by data scheduling: high throughput, high layer delivery ratio, low useless packets ratio, and low subscription jitter; 2) We design a 3-stage scheduling mechanism to request absent blocks, where the min-cost flow model, probability decision mechanism and multi-window remedy mechanism are employed in Free Stage, Decision Stage and Remedy Stage, respectively. Each stage has different scheduling objective while collaborates with each other, to achieve the above four objectives. Experimental results indicate that our approach outperforms other schemes in simulation environment. Besides, LayerP2P is implemented in the PDEPS Project in China, which is expected to be the first practical layered streaming system for education in peer-to-peer networks.
Xin Xiao 0005, Yuanchun Shi, Qian Zhang 0001
INFOCOM2
2009 Human pose estimation from corrupted silhouettes using a sub-manifold voting strategy in latent variable space
Chunfeng Shen, Xueyin Lin, Yuanchun Shi
Pattern Recognit. Lett.3
2009 Learning in an Ambient Intelligent World: Enabling Technologies and Practices
abstract
The rapid evolution of information and communication technology opens a wide spectrum of opportunities to change our surroundings into an Ambient Intelligent (AmI) world. AmI is a vision of future information society, where people are surrounded by a digital environment that is sensitive to their needs, personalized to their requirements, anticipatory of their behavior, and responsive to their presence. It emphasizes on greater user friendliness, user empowerment, and more effective service support, with an aim to bring information and communication technology to everyone, every home, every business, and every school, thus improving the quality of human life. AmI unprecedentedly enhances learning experiences by endowing the users with the opportunities of learning in context, a breakthrough from the traditional education settings. In this survey paper, we examine some major characteristics of an AmI learning environment. To deliver a feasible and effective solution to ambient learning, we overview a few latest developed enabling technologies in context awareness and interactive learning. Associated practices are meanwhile reported. We also describe our experience in designing and implementing a smart class prototype, which allows teachers to simultaneously instruct both local and remote students in a context-aware and natural way.
Lizhu Zhou, Yuanchun Shi
IEEE Trans. Knowl. Data Eng.4
2009 Open Smart Classroom: Extensible and Scalable Learning System in Smart Space Using Web Service Technology
abstract
Real-time interactive virtual classroom with teleeducation experience is an important approach in distance learning. However, most current systems fail to meet new challenges in extensibility and scalability, which mainly lie with three issues. First, an open system architecture is required to better support the integration of increasing human-computer interfaces and personal mobile devices in the classroom. Second, the learning system should facilitate opening its interfaces, which will help easy deployment that copes with different circumstances and allows other learning systems to talk to each other. Third, problems emerge on binding existing systems of classrooms together in different places or even different countries such as tackling systems intercommunication and distant intercultural learning in different languages. To address these issues, we build a prototype application called Open Smart Classroom built on our software infrastructure based on the multiagent system architecture using Web Service technology in Smart Space. Besides the evaluation of the extensibility and scalability of the system, an experiment connecting two Open Smart Classrooms deployed in different countries is also undertaken, which demonstrates the influence of these new features on the educational effect. Interesting and optimistic results obtained show a significant research prospect for developing future distant learning systems.
Yue Suo, Naoki Miyata, Hiroki Morikawa, Toru Ishida 0001, Yuanchun Shi
IEEE Trans. Knowl. Data Eng.5
2009 Finding User's Interest Blocks using Significant Implicit Evidence for Web Browsing on Small Screen Devices
Peifeng Xiang, Yuanchun Shi
World Wide Web3
2008 NALP: Navigating Assistant for Large Display Presentation Using Laser Pointer
abstract
In this paper, we present NALP (navigating assistant using laser pointer), a novel interaction technique for large display presentation control. NALP is based on laser dot detection, track recognition and space segmentation, allowing users to manipulate the presentation software (such as PowerPoint ) directly and freely using any kind of common laser pointer. Demanding users only to sweep the pointer to and fro on display, this technique to a large extent solved the problems caused by hand jitter, latency and detection errors existing in other navigation assistant systems. Report on evaluation experiments shows that when delivering an electronic and interactive presentation with large displays, NALP can effectively meet the need and is preferred over other laser pointer interactive methods.
Yuanchun Shi, Boliang Chen
ACHI2
2008 UMDD: User Model Driven Software Development
abstract
The existing software engineering seldom considers software usability, and human-computer interaction (HCI) techniques which can improve the software usability cannot guarantee development efficiency. Recently, more and more stakeholders and users begin to regard the usability as an important software requirement. In order to bring respective advantages of software engineering and HCI techniques into full play to improve software usability and development efficiency, the paper presents user model driven software development, which integrates HCI techniques into software development method by eliciting user model under the participation of user, HCI designer and software engineer. Applications show that the method can be effectively applied to small software development team, and can raise software usability and development efficiency.
Xiaochun Wang, Yuanchun Shi
EUC (1)2
2008 OCals: A Novel Overlay Construction Approach for Layered Streaming
abstract
Layered streaming in overlay networks has drawn great interests since not only can it accomodate large scales of clients but also it handles client heterogeneities. However, to our knowledge, there's still a lack of overlay construction (i.e. neighbor selection) approach suited for layered streaming, because i) In existing works neighbors are selected only based on their network conditions. However, a neighbor with good network condition may not be able to provide sufficient layers (e.g. a neighbor in the same LAN), ii) Previous works usually select "good" neighbors for the new node, ignoring that the joining of the new node could also be utilized to improve the performance of existing nodes. In this paper, OCals - a two-stage QoS aware overlay construction approach for layered streaming is proposed. The main contribution of OCals is that i) when selecting neighbor, it considers existing nodes' network conditions and their providing layers as a whole; ii) it guarantees the QoS for the new node as well as improves the QoS for existing nodes so that with the joining of new nodes, the performance of the overlay could be consecutively improved; iii) it's easy to implement and low time cost. Experiments demonstrate that compared with two other approaches: SCAMP (a pure random neighbor selection method) and Narada (a QoS aware method), the throughput and average packet delay of the layered streaming on top of the overlay constructed by OCals can be remarkably improved. Besides, the time spent on joining and recovery is very short.
Xin Xiao 0005, Yuanchun Shi, Baopeng Zhang
ICC2
2008 On optimal scheduling for layered video streaming in heterogeneous peer-to-peer networks
abstract
Layered video streaming in peer-to-peer networks has drawn great interests since not only it can accommodate a large number of users but also it handles heterogeneities of client networks. However, to our knowledge, there's still a lack of systematical study on the data scheduling (i.e. requesting and relaying data) for layered streaming, and previous works in this area just focus on maximizing the throughput and/or minimizing the packet delay. In this paper, firstly, according to the characteristics caused by layered coding, we propose the four objectives that should be achieved by data scheduling for layered streaming: high throughput, high layer delivery ratio, low useless packets ratio and low subscription jitter; then, we use a 3-stage scheduling approach to request missed blocks, where each stage has different scheduling objective but collaborate with each other. The min-cost network flow model, probability decision mechanism and multi-window remedy mechanism are employed in Free Stage, Decision Stage and Remedy Stage, respectively, to achieve the above four goals. Extensive experimental results indicate that our approach outperforms other schemes in both throughput and layer delivery ratio. Besides, the useless packets number and subscription jitters are kept low.
Xin Xiao 0005, Yuanchun Shi
ACM Multimedia2
2008 UCam: direct manipulation using handheld camera for 3d gesture interaction
abstract
This paper presents UCam a novel approach in 3D Gesture Interaction based on handheld camera movement. UCam reflects hand's movement and directly maps it to the movement of 3D object based on visual tracking of feature-like points on incoming frames. Only one button is needed to differentiate rotation and translation. The advantages of this technique lie in the popularity and low cost of handheld cameras, low requirement and no need of adjustment of background and easy to use for beginners. To evaluate UCam, it is compared with mouse in some 3D controlling tasks. The results show that UCam is more flexible and easier to use and master in most cases. Even for complicated tasks, UCam has comparable performance as mouse.
Yuanchun Shi, Mingming Fan 0001
ACM Multimedia2
2008 Combining User Profiles and Situation Contexts for Spontaneous Service Provision in Smart Assistive Environments
Weijun Qin, Daqing Zhang 0001, Yuanchun Shi, Kejun Du
UIC3
2008 Hand's 3D movement detection with one handheld camera
abstract
This paper presents a scheme to create a real-time and reliable method for recognizing vision-based hand's 3D movement and to use the movement parameters for controlling 3D objects. The algorithm for 3D movement detection is totally based on analyzing feature points from the only camera in user's hand. As the algorithm is based on frames captured from one camera in untrained environment, it's difficult to distinguish similar movements on optical flow images, especially between shifting and rotating. A novel differentiation algorithm by voting from some weak classifiers is used. The algorithm provides a method of direct mapping user's hand movement to object control. We design an application of controlling a virtual 3D cube's movement and estimate the accuracy of the algorithm. And the experients' result presents that the 3D movement detection algorithm is efficient and robust enough for real-time interaction.
Mingming Fan 0001, Yuanchun Shi
VRST3
2007 An Experimental Study on Cheating and Anti-Cheating in Gossip-Based Protocol
abstract
The Internet has witnessed a rapid growth in deployment of gossip-based protocol in many multicast applications. In a typical gossip-based protocol, each node independently exchanges data with its neighbors, acting as dual roles of receiver and sender to facilitate scalability and resilience. However, most of previous work in this literature seldom considered cheating issue of end users, which is also very important in face of the fact that the mutual cooperation inherently determines overall system performance. In this paper, we mainly investigate the dishonest behaviors in decentralized gossip-based protocol through extensive experimental study. Our original contributions come in two-fold: In the first part of cheating study, we analytically discuss two typical cheating strategies, that is, intentionally increasing subscription requests and untruthfully calculating forwarding probability, and further evaluate their negative impacts. The results indicate that more attention should be paid on defending cheating behaviors in gossip-based protocol. In the second part of anti-cheating study, we propose a simple receiver-driven measurement mechanism, which evaluates individual forwarding traffic from the perspective of receivers and thus identifies cheating nodes with high incoming/outgoing ratio. The experiments under various conditions show that it performs quite well in case of serious cheating and achieves considerable performance in other cases.
Yuanchun Shi, Shiqiang Yang, Yuzhuo Zhong
ICC3
2007 Web Page Segmentation Based on Gestalt Theory
abstract
Automatic Web page segmentation is the basis to adaptive Web browsing on mobile devices. It breaks a large page into smaller blocks, in which contents with coherent semantics are keeping together. Then, various adaptations like single column and thumbnail view can be developed. However, page segmentation remains a challenging task, and its poor result directly yields a frustrating user experience. As human usually understand the Web page well, in this paper, we start from Gestalt theory, a psychological theory that can explain human's visual perceptive processes. Four basic laws, proximity, similarity, closure, and simplicity, are drawn from Gestalt theory and then implemented in a program to simulate how human understand the layout of Web pages. The experiments show that this method outperforms existing methods.
Peifeng Xiang, Yuanchun Shi
ICME3
2007 An Investigation and a Preventing Strategy for the Redundant Packets in P2P Networks with Push Method
abstract
The push method for data transmission in peer-to-peer system has drawn great interest, since it can efficiently reduce the accumulated latency observed at user nodes. However, it can cause redundant packets more easily than the pull method, because the nodes using pull method can completely control the data request process while it's not the case for push method. In this paper, we systematically study the distribution of the received redundant packets in push mode, and model it as a heterogeneous Poisson arrival. Both the theoretical analysis and the simulation show that the redundant packets caused by push method can be remarkable if the end-to-end delays among peers are not small enough. Furthermore, a redundant packets preventing strategy is proposed, and the experiment on PlanetLab demonstrates that it can effectively reduce the amount of received redundant packets.
Xin Xiao 0005, Yuanchun Shi
ICME2
2007 A Peer-to-Peer Semantic-Based Service Discovery Method for Pervasive Computing Environment
Baopeng Zhang, Yuanchun Shi, Xin Xiao 0005
UIC2
2006 Fusion of Texture Variation and On-Line Color Sampling for Moving Object Detection Under Varying Chromatic Illumination
Chunfeng Shen, Xueyin Lin, Yuanchun Shi
ACCV (1)3
2006 HyMoNet: a peer-to-peer hybrid multicast overlay network for efficient live media streaming
abstract
This paper presents HyMoNet, a peer-to-peer hybrid multicast overlay network for efficient live media streaming. The system comprises two parts: service infrastructure and user space. Service infrastructure contains the data stream servers manually deployed in the backbone network by the service provider. In user space, all users are dynamically classified into several tree structures, in which data is dispatched successively. With the user space management algorithm, the system can detect IP multicast service intelligently and allow users to join and to leave dynamically. NAT is an obstacle for widespread of P2P system, so a NAT pass strategy is designed to overcome this problem. The system guarantees the reliable data transmission which is friendly to TCP data flow. HyMoNet is now deployed in Internet providing real-time TV program streaming for the public.
Bin Chang, Yuanchun Shi
AINA (1)2
2006 Direct pointer: direct manipulation for large-display interaction using handheld cameras
abstract
This paper describes the design and evaluation of a technique, Direct Pointer, that enables users to interact intuitively with large displays using cameras equipped on handheld devices, such as mobile phones and personal digital assistant (PDA). In contrast to many existing interaction methods that attempt to address the same problem, ours offers direct manipulation of the pointer position with continuous visual feedback. The primary advantage of this technique is that it only requires equipment that is readily available: an electronic display, a handheld digital camera, and a connection between the two. No special visual markers in the display content are needed, nor are fixed cameras pointing at the display. We evaluated the performance of Direct Pointer as an interaction product, showing that it performs as well as comparable techniques that require more sophisticated equipment.
Eyal Ofek, Neema Moraveji, Yuanchun Shi
CHI4
2006 A Novel Approach for Sharing White Board Between PC and PDAs with Multi-users
Xin Xiao 0005, Yuanchun Shi, Weisheng He
EUC2
2006 Impact of Node Cheating on Gossip-Based Protocol
Yuanchun Shi, Bin Chang
EUC2
2006 Cicada: A Highly-Precise Easy-Embedded and Omni-Directional Indoor Location Sensing System
Hongliang Gu, Yuanchun Shi, Yu Chen 0004, Bibo Wang, Wenfeng Jiang
GPC2
2006 CAMPS: A Middleware for Providing Context-Aware Services for Smart Space
Weijun Qin, Yue Suo, Yuanchun Shi
GPC3
2006 Ermdclime: Enabling Real-time Multimedia Discussion for Collaborative Learning in Mobile Environment
abstract
With the aid of mobile technology, it's possible to carry out collaborative learning (CL) in any places at any time. And the most essential and effective way to achieve this goal is through real-time multimedia discussion. In this paper, a multi-agent system, ErmdClime, is designed and implemented to enable real-time multimedia discussion for collaborative learning in mobile environment. The realtime bi-directional audio/video interaction and shared white board are integrated seamlessly into the system for learners with mobile devices. Besides, this system enables both local and remote learners to control the play of PowerPoint slides by voice command when discussing and annotate on the shared white board. Context-aware technology is used to switch real video and white board display on mobile devices. User study demonstrates that learners feel quite natural and highly efficient when they are engaged in the CL with mobile devices in our system
Xin Xiao 0005, Yuanchun Shi
ICME2
2006 Degree Pre-Reserved Hierarchical Tree for Multimedia Multicast
abstract
Overlay multicast tree is widely used to support large-scale real-time multimedia applications. The scalability and the robustness are two key issues in the overlay structure design. In this paper, we propose a degree pre-reserved hierarchical tree for multimedia multicast, called DTree. It organizes the overlay tree in two hierarchies, and combines application-level multicast with IP multicast to achieve a low delay data delivery. A degree pre-reserved mechanism is designed in DTree, which greatly shortens the time to resume the data flow, and achieve a fast recovery. Our simulation results show that DTree has a better delivery performance, compared with other overlay tree. It is quite responsive to the changes in the tree, and almost 3 times faster than other recovery strategies in some cases
Yuanchun Shi, Bin Chang
ICME2
2006 Recovering semantic relations from web pages based on visual cues
abstract
Recovering semantic relations between different parts of web pages are of great importance for multi-platform web interface development, as they make it possible to re-distribute interaction objects and change the structure of interfaces while preserving the semantics of the UI. Important semantic relations include topic, order, hierarchy, etc. This paper presents a visual cues based approach, which is tag-tree structure independent, to automatically detect such kind of semantic relations in web pages. Comparing with other existing techniques, such as DOM-based methods, this approach mostly depends on interfaces' perceptible visual information that is more reliable. The preliminary evaluation on complex web sites shows promising results. We believe further exploration is worth taken.
Peifeng Xiang, Yuanchun Shi
IUI2
2006 Drag and Drop by Laser Pointer: Seamless Interaction with Multiple Large Displays
Yuanchun Shi, Jichun Chen
UIC2
2006 Effective Page Segmentation Combining Pattern Analysis and Visual Separators for Browsing on Small Screens
abstract
Page segmentation plays a key role in browsing on small screens. It breaks a large page into smaller segments according to their semantic relationships. Then, various approaches such as single column adaptation and thumbnail view with zooming links can be implemented based on these page segments. However, for current flexible Web pages, segmentation remains a challenging task. This paper proposes an effective automatic segmentation method which combining pattern analysis and visual separators. The basic idea is that a page's semantic structure is largely reflected by repeated continuous patterns and visual separators, which coincides with human's visual perception. The proposed method works in three steps: generating a refined tag tree from the DOM tree, recognizing and merging inexact patterns recursively, and segmenting the others by visual separators. Our experimental results show that the proposed method outperforms existing methods, especially for pages automatically generated from templates
Peifeng Xiang, Yuanchun Shi
Web Intelligence3
2006 Moving object tracking under varying illumination conditions
Chunfeng Shen, Xueyin Lin, Yuanchun Shi
Pattern Recognit. Lett.3
2005 A Personalized Agents Platform Design and Implementation for Personalized Education
Apple W. P. Fok, Xin Xiao 0005, Yuanchun Shi, Horace Ho-Shing Ip
ICCE3
2005 uPen: laser-based, personalized, multi-user interaction on large displays
abstract
We present the uPen, a laser pointer combined with a contact-pushed switch, three press buttons and a wireless communication module. This novel interaction device allows users to interact on large displays at a distance or directly on the surface with full-function of mouse. Onboard software enable the uPen system to identify different users and provide personalized services to them, such as associating users with corresponding privileges, giving access to each participant's private content (e.g., home pages, personal calendars). Additionally, with our two-step association method, the uPen system has the ability to distinguish strokes of different uPens working simultaneously and support multi-user simultaneous interaction. A prototype system has been implemented in our Smart Classroom [1]. And user studies show the benefit of using it.
Xiaojun Bi 0001, Yuanchun Shi, Peifeng Xiang
ACM Multimedia2
2005 Guest Editor's Introduction
Wenyin Liu, Yuanchun Shi, Hai Zhuge
World Wide Web2
2004 Rich Metadata Searches Using the JXTA Content Manager Service
abstract
With the development of networking technologies and the advent of the peer-to-peer computing paradigm, distributed file-sharing systems like Gnutella are becoming prevalent over time. JXTA, an interoperable and platform independent peer-to-peer computing infrastructure, has been adopted in an increasing number of network applications providing file-sharing and cooperative service. In this paper, we propose the metadata search layer which serves as an enhancement to the CMS (content manager service), a JXTA-based file-sharing service. Through the metadata search layer, more precise search of resources could be conducted as opposed to the inflexible keyword-based search mechanism employed in most of current file-sharing applications in use. A compact general-purpose query language is proposed to facilitate the use of standard metadata schema and bridge the gap between users and metadata descriptions of resources. We exemplify the advantage of the metadata-based search mechanism over the traditional keyword search with a proof-of-concept application.
Yuanchun Shi
AINA (1)2
2004 Context-Aware Computing During Seamless Transfer Based on Random Set Theory for Active Space
Yuanchun Shi, Enyi Chen, Guangyou Xu, Baopeng Zhang
EUC2
2004 Use Web Usage Mining to Assist Online E-Learning Assessment
abstract
In this paper we present a set of models related to online learning activities. We also present our approach to assess student behavior through Web log mining.
Yuanchun Shi
ICALT3
2004 TORM: a hybrid multicast infrastructure for interactive distance learning
abstract
Multicast has been intensively studied over the past decade. Existing multicast services come in two-flavors, IP multicast and application level multicast, depending on which layer the service is implemented on. After investigating the pros and cons of both approaches, as well as the characteristics of interactive distance learning, we find it an imperative to combine these two multicast approaches. We present TORM (totally ordered reliable multicast), a hybrid multicast infrastructure for interactive distance learning. Apart from blending IP multicast and application level multicast, TORM builds and maintains a tree structure upon which transport functionalities are implemented, including a differentiated error recovery service, flow and congestion control and global ordering.
Yi Che, Runting Shi, Yuanchun Shi
ICME3
2003 Search and Delivery of Standardized Learning Resources Based on SOAP Messaging and Native XML Databases
Yuanchun Shi
Dublin Core Conference3
2003 A Conformance Test Suite of Localized LOM Model
abstract
Since the approval of IEEE LOM draft standard and the advance of network-driven learning technology, a large number of resource database constructors, content developers and learning management system vendors are likely to describe their relevant learning resources using LOM model. To ensure the process of carrying out correct implementation of LOM, there is need for a conformance test suite of LOM model. Some design and implementation issues related to LOM XML binding, parsing of metadata instance as well as carrying out conformance testing procedure are to be discussed. We hope that the initiative could elicit some opinions with respect to conformance testing of learning technology standards.
Yuanchun Shi
ICALT2
2002 A Learning Resource Metadata Management System Based on LOM Specification
abstract
The rapid increase of learning resources makes it difficult to search, manage and reuse. Using metadata is an efficacious way to solve this problem. With consistent descriptions of the characteristics of learning resources, searching becomes more specific and accurate, management becomes more simple and uniform, and sharing becomes more efficient and in-depth. The Learning Object Metadata Schema developed by IEEE P1484.12 is one of the most promising metadata approaches for describing learning resources, on which we developed a learning resource metadata management system (LRMMS). The system provides a platform for users to register, browse, search and evaluate learning resources. It is a decentralized framework with several metadata servers to provide services. The system supports distributed queries and the evaluation loop of learning resources. Users can search resources from different points of view, especially educational needs. We also made the interface as user-friendly as possible.
Zhongnan Shen, Yuanchun Shi, Guangyou Xu
CSCWD2
2002 Smart Platform - A Software Infrastructure for Smart Space (SISS)
abstract
A software infrastructure is fundamental to a Smart Space. Previously proposed software infrastructures for Smart Space (SISS) did not sufficiently address the issue of performance and usability. A new solution, Smart Platform, which is focused on improving these aspects of a SISS, is presented in this paper. To optimize its intermodule communication performance, the stream-oriented communication is distinguished from the message-oriented ones, and a corresponding hybrid communication scheme is proposed. To improve the usability, a featured loose coupling structure, a straightforward Publish-and-Subscribe coordination model as well as a set of user-friendly deployment and development tools are developed. Besides, Smart Platform is intended as an open and generic SISS available for other research groups. To this end, XML-based message syntax and the open wire-protocol based architecture are adopted to make sharing research efforts more easily.
Weikai Xie, Yuanchun Shi, Guangyou Xu, Yanhua Mao
ICMI2
2001 Supporting Group Awareness in Collaborative Design
abstract
Different from multi-user database management systems, CSCW systems must support group awareness explicitly, namely that participants should perceive the presence of each other during the process of cooperation, which is essential to effective collaboration. The authors firstly analyze and compare several means of supporting sensibility, then investigate the model of collaborative design and suggest achieving group awareness by product data which forms the foundation of a relationship among designers. The information model and the shared workspace organization of CECAD (Collaborative Environment for Computer-Aided Design, a prototype system developed by the authors) are presented and its group awareness supporting mechanism is described in succession.
Yuanchun Shi, Guangyou Xu
CSCWD2
2000 A Pragmatic Semantic Reliable Multicast Architecture for Distant Learning
Yuanchun Shi, Guangyou Xu
ICMI2