Lik-Hang Lee

dblp:203/8146 · also Lik Hang Lee · DBLP profile ↗
← Back
59ranked-venue papers
10as first author
52since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 31 · 9 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 23 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Who Gets Left Out of Digital Banking in Later Life? Barriers and Opportunities in Hong Kong's Silver Population
abstract
As digital banking increasingly replaces face-to-face financial services, older adults face growing challenges in navigating self-service and mobile platforms. This issue is particularly salient in Hong Kong, where a highly digitalized yet fragmented multi-channel banking ecosystem combines branches, ATMs, mobile apps, and phone banking. While prior research has identified general barriers such as usability and trust, less is known about how banking practices, challenges, and support needs differ across stages of later life. We address this gap through a mixed-methods study in Hong Kong, combining an in-person survey with 151 adults aged 60+ and semi-structured interviews with older adults and frontline bank staff. Participants were grouped into young-old (60–69), old-old (70–79), and oldest-old (80+) cohorts. Our findings reveal clear age-related patterns: young-old adults actively use ATMs and digital banking but report strong psychological concerns; old-old adults rely on hybrid channel use and face increasing knowledge-related barriers; and oldest-old adults depend primarily on physical branches due to compounded physical and cognitive limitations. We conclude with age-specific design implications for more inclusive digital banking systems.
Clarence Chi S. Cheung, Lulin Chen, Qiongyan Chen, Luchen Li, Pan Hui 0001, Lik-Hang Lee, Mingming Fan 0001
DIS7
2026 When Generative Artificial Intelligence Meets Extended Reality: A Systematic Review
abstract
With the continuous advancement of technology, the application of generative artificial intelligence (AI) in various fields is gradually demonstrating great potential, particularly when combined with Extended Reality (XR), creating unprecedented possibilities. This survey article systematically reviews the applications of generative AI in XR, covering as much relevant literature as possible from 2023 to 2025. The application areas of generative AI in XR and its key technology implementations are summarised through PRISMA screening and analysis of the final 26 articles. The survey highlights existing articles from the last three years related to how XR utilises generative AI, providing insights into current trends and research gaps. We also explore potential opportunities for future research to further empower XR through generative AI, providing guidance and information for future generative XR research.
Xinyu Ning, Yan Zhuo, Chan-In Sio, Lik-Hang Lee
Int. J. Hum. Comput. Interact.5
2026 Conflict Resolution Strategies for Co-Manipulation of Virtual Objects Under Non-Disjoint Conditions
abstract
Virtual Reality (VR) co-manipulation enables multiple users to collaboratively interact with shared virtual objects. However, existing research treats objects as monolithic entities, overlooking scenarios where users need to manipulate different sub-components simultaneously. This work addresses conflict resolution when users select overlapping vertices (non-disjoint sets) during co-manipulation. We present a comprehensive framework comprising preventive strategies (Object-level and Action-level Restrictions) and reactive strategies (computational conflict resolution). Through two user studies with 76 participants (38 pairs), we evaluated these approaches in collaborative wireframe editing tasks. Study 1 identified Averaging as the optimal computational method, balancing task efficiency with user experience. Study 2 highlighted that Action-level Restriction, which permits overlapping selections but restricts concurrent identical operations, achieved better performance compared to exclusive object locking. Reactive strategies using averaging provided smooth collaboration for experienced users, while second-user priority enabled quick corrections. Our findings indicate that optimal strategy selection depends on task requirements, user expertise, and collaboration patterns. Based on the findings, we provide design implications for developing VR collaboration systems that support flexible sub-components manipulation while maintaining collaborative awareness and minimizing conflicts.
Xuanru Cheng, Rongkai Shi, Lei Chen 0088, Jingyao Zheng, Hai-Ning Liang, Lik-Hang Lee
IEEE Trans. Vis. Comput. Graph.7
2026 When Effort Becomes Visible: Facet-Level Shifts in Evaluation and Workload During VR Teamwork
abstract
Surfacing peers' workload and speed can reshape collaboration in VR, yet prior work often conflates social cues with environmental load. We isolate the social channel with TRACE-VR, which holds geometry, physics, rules, and timing constant while independently manipulating effort identifiability (traceable vs. anonymous) and peer effort (pace) (high vs. low). In a 2×2 within-subjects study $(n=32)$, each participant worked with nine scripted co-actors for three minutes, transporting 16 crates along a self-chosen path between fixed pickup/drop points; co-actors followed fixed pre-authored routes. We measured intrinsic motivation, social-evaluative load, NASA-TLX, and behavioral/process outcomes. Our findings show that identifiability and higher peer effort (pace) each modestly increased completion rates, with asymmetric effects suggesting partial substitution between accountability cues and normative pace. Perceived monitoring and temporal demand depended on their combination-identifiability raised monitoring at low pace and amplified time pressure at high pace-while composite TLX and finishers' times remained comparable across conditions. Thus, these cues chiefly determine who finishes rather than how fast finishers move. Motivation did not uniformly increase under higher pace. We frame identifiability and pace as partially substitutable levers for shaping social-evaluative experience and facet-level motivation, and outline tempo-aware, accountability-aware guidance for collaborative VR.
Zheng Wei 0003, Hyeonmin Lee, Junxiang Liao, Hao Li 0102, Hayoung Oh 0002, Lik-Hang Lee, Wai Tong, Pan Hui 0001
IEEE Trans. Vis. Comput. Graph.8
2026 Non-Urgent Messages Do Not Jump into My Headset Suddenly! Adaptive Notification Design in Mixed Reality
abstract
Mixed reality (MR) notification systems currently display all messages in fixed central locations regardless of urgency, leading to unnecessary interruptions and cognitive overload. Drawing from previous MR/Virtual Reality (VR) notification design work and calm technology principles, we developed an adaptive notification system that adjusts spatial placement based on urgency levels: non-urgent notifications appear as peripheral icons accessible via head movement, moderately urgent messages anchor to the user's hand, and very urgent notifications transition progressively from peripheral to central view. Through a within-subjects study (N=18), we evaluated our adaptive system against the default centralised approach. Results demonstrate that the adaptive system significantly reduces mental workload $(p=0.041)$, temporal workload $(p=0.008)$, and frustration $(p=0.004)$ while maintaining comparable notification awareness. Logistic regression analysis reveals that users prefer the adaptive system even with classification errors, provided the combined misclassification rate (disruptiveness + omission errors) remains below a determinable threshold. Our findings establish the first empirical evidence that urgency-based spatial notification distribution effectively addresses core MR usability challenges, offering practical design guidelines for immersive notification systems that balance user attention management with information accessibility.
Jingyao Zheng, Sven Mayer, Lik-Hang Lee
IEEE Trans. Vis. Comput. Graph.4
2026 MetaCineMoji: Visualizing film set communication in an interactive interface for collaboration in virtual LED production
abstract
LED-VP is now mainstream in high-end studios, yet the capital cost of real-time volume stages renders training with this technology prohibitively expensive. Most film institutions simply cannot afford to build one LED-VP set for their students. This paper explores an alternative solution to the training of LED-VP by developing a virtual counterpart. However, significant challenges arise in simulating a virtual environment where students can operate in an LED-VP. Our study proposes the virtual environment of LED-VP, a virtual reality Collaborative Learning (VRCL) system. We developed an interface featuring visual symbols, namely MetaCineMoji , for film operation to facilitate smooth communication and learning processes in LED-VP workflows. MetaCineMoji demonstrates the feasibility of multi-person collaborative learning in a virtual filming studio, translating film set communication to visual symbols for lighting design, scene construction, and collaborative work with key stakeholders, e.g., directors, cinematographers, and gaffers. We explore the impact of the film-operation interfaces containing visual symbols on social interaction factors within the virtual studio. We conducted evaluations with 24 participants. Our findings show that our system equipped with the film-operation interface, enabled by visual symbols, significantly enhances social interaction among learners and results in significantly higher learning outcomes than systems without the interface. • MetaCineMoji VR system enables collaborative LED production training, significantly enhancing team communication and learning outcomes. • Visual symbols in MetaCineMoji streamline film set operations, facilitating efficient role-based interactions among directors, cinematographers, and gaffers. • Participants using MetaCineMoji reported improved workflow understanding, and strengthened team presence compared to non-symbolic systems.
Zheng Wei 0003, Shan Jin 0002, Wai Tong, Pan Hui 0001, Lik-Hang Lee
Vis. Informatics5
2025 DysVis: A User-Centred Data Visualization System for Dyslexia Pre-screening
abstract
Dyslexia is a common neurobiological learning disorder significantly impacting reading, writing, and spelling worldwide. Early identification and intervention are essential, but most pre-screening tools focus on Latin languages, leaving Chinese-speaking students underserved. To address this gap, we conduct semi-structured interviews with special education (special-ed) teachers to gather their needs for dyslexia pre-screening tailored to Chinese contexts. Using their insights, we have developed DysVis, a user-centered data visualization system that combines handwriting analysis, body movement keypoint conversion, and a comprehensive visualization interface. DysVis provides teachers with multi-level visualizations, such as performance overviews, task analyses, handwriting observations, and behavioural insights, enabling them to identify the root causes of learning difficulties. Our evaluations, including case studies, a user study, and expert interviews, demonstrate that DysVis is user-friendly and effective in quickly identifying at-risk students, ultimately enhancing learning outcomes for Chinese-speaking students with dyslexia.
Ka Yan Fung, Lik-Hang Lee, Linping Yuan, Kwong Chiu Fung, Kuen Fung Sin, Tze-Leung Rick Lui, Huamin Qu, Shenghui Song 0001
CHI2
2025 MIRAGE: Multimodal Intention Recognition and Admittance-Guided Enhancement in VR-Based Multi-Object Teleoperation
abstract
Effective human-robot interaction (HRI) in multi-object teleoperation tasks faces significant challenges due to perceptual ambiguities in virtual reality (VR) environments and the limitations of single-modality intention recognition. This paper proposes a shared control framework that combines a virtual admittance (VA) model with a Multimodal-CNN-based Human Intention Perception Network (MMIPN) to enhance teleoperation performance and user experience. The VA model employs artificial potential fields to guide operators toward target objects by adjusting admittance force and optimizing motion trajectories. MMIPN processes multimodal inputs-gaze movement, robot motions, and environmental context-to estimate human grasping intentions, helping overcome depth perception challenges in VR. Our user study evaluated four conditions across two factors, and the results showed that MMIPN significantly improved grasp success rates, while the VA model enhanced movement efficiency by reducing path lengths. Gaze data emerged as the most crucial input modality. These findings demonstrate the effectiveness of combining multimodal cues with implicit guidance in VR-based teleoperation, providing a robust solution for multi-object grasping tasks and enabling more natural interactions across various applications in the future.
Chi Sun, Abhishek Kumar 0011, Chengbin Cui, Lik-Hang Lee
ISMAR5
2025 Exploring Gaze Dynamics in Vr Film Education: Gender, Avatar, and the Shift Between Male and Female Perspectives
abstract
In virtual reality (VR) education, especially in creative fields like film production, avatar design and narrative style extend beyond appearance and aesthetics. This study explores how the interaction between avatar gender, the dominant narrative actor's gender, and the learner's gender influences film production learning in VR, focusing on gaze dynamics and gender perspectives. Using a$2 \times 2 \times 2$experimental design, 48 participants operated avatars of different genders and interacted with male or female-dominant narratives. The results show that the consistency between the avatar and gender affects presence, and learners' control over the avatar is also influenced by gender matching. Learners using avatars of the opposite gender reported stronger control, suggesting gender incongruity prompted more focus on the avatar. Additionally, female participants with female avatars were more likely to adopt a “female gaze,” favoring soft lighting and emotional shots, while male participants with male avatars were more likely to adopt a “male gaze,” choosing dynamic shots and high contrast. When male participants used female avatars, they favored “female gaze,” while female participants with male avatars focused on “male gaze”. These findings advance our understanding of how avatar design and narrative style in VRbased education influence creativity and the cultivation of gender perspectives, and they offer insights for developing more inclusive and diverse VR teaching tools going forward.
Zheng Wei 0003, Jia Sun 0011, Junxiang Liao, Lik-Hang Lee, Pan Hui 0001, Huamin Qu, Wai Tong
ISMAR4
2025 PersoNo: Personalised Notification Urgency Classifier in Mixed Reality
abstract
Mixed Reality (MR) is increasingly integrated into daily life, providing enhanced capabilities across various domains. However, users face growing notification streams that disrupt their immersive experience. We present PersoNo, a personalised notification urgency classifier for MR that intelligently classifies notifications based on individual user preferences. Through a user study ($\mathrm{N}=18$), we created the first MR notification dataset containing both selflabelled and interaction-based data across activities with varying cognitive demands. Our thematic analysis revealed that, unlike in mobiles, the activity context is equally important as the content and the sender in determining notification urgency in MR. Leveraging these insights, we developed PersoNo using large language models that analyse users' replying behaviour patterns. Our multi-agent approach achieved 81.5% accuracy and significantly reduced false negative rates (0.381) compared to baseline models. PersoNo has the potential not only to reduce unnecessary interruptions but also to offer users understanding and control of the system, adhering to Human-Centered Artificial Intelligence design principles.
Jingyao Zheng, Haodi Weng, Chengbin Cui, Sven Mayer, Chi-lok Tai, Lik-Hang Lee
ISMAR7
2025 Perceived User Reachability in Mobile UIs Using Data Analytics and Machine Learning
abstract
One-handed interactions on smartphone interfaces offer a prominent feature of highly mobile inputs. Thus, the design factor of user reachability is essential to realizing the incentive. However, the sole consideration of physical characteristics, such as hand size and thumb length, does not fully reflect the users’ perceived choices of hand poses and the corresponding inertia. We first conducted a 6-week questionnaire-based study of UI rating tasks and collected 62,156 responses reflecting user preferences for 3000 clustered UIs. Our analysis of the responses shows that user perceptions of smartphone UI components are divergent from their physical ability of thumb reaches; e.g. they can reach an icon with a thumb reach, but they prefer alternative hand poses. Accordingly, we propose a machine learning model, i.e. XGBoost (XGB), to predict the user’s choices of hand poses, with a reasonable prediction accuracy of 64% that can be regarded as a practical preliminary evaluation tool. With illustrative examples, our model can offer auxiliary information in the assessment of perceived user reachability with one-handed interaction on smartphone interfaces, which paves a path toward a computational understanding of UI designs, and such findings can be further extended to 2D UIs in 3D worlds.
Lik-Hang Lee, Yui-Pan Yau, Pan Hui 0001
Int. J. Hum. Comput. Interact.1
2025 EmojiChat: Toward Designing Emoji-Driven Social Interaction in VR Museums
abstract
Museums have traditionally been places of learning, evolving into spaces that also facilitate socializing. This shift is particularly evident in virtual reality (VR) museums, which have become popular venues for activities like friend gatherings. However, education has long established the social norm that museums need to “maintaining silence.” Even in virtual environments, this can influence visitor behavior. This perception often prevents people from using verbal communication, leading them to prefer quieter forms of interaction. In museums, including the VR museums that now largely replicate the layout of physical museums, this preference may be reinforced by specific features, such as spaciousness, quietness, and block-based layouts, which may create visual obstructions, restricting interaction modes dependent on shared view, such as gesture interaction. These limitations necessitate the introduction of an additional interaction mode. Emojis, with their capacity for rapid message exchange and adjustable positioning, emerge as a suitable interaction mode in this context. Thus, we introduce EmojiChat, an innovative VR museum experience designed to respect the social norms of traditionally keeping quiet while promoting natural interaction. We first design and iterate a customized emoji set for the museum context through semi-structured interviews, participatory design, and an online survey. Then, this emoji set is integrated into a VR museum to facilitate interaction between visitors. Finally, we conduct a comparative study to evaluate the performance of EmojiChat. Our results show that the integration of emojis can improve communication enjoyment and efficiency. Additionally, we identify usage patterns for interaction modes and the advantages offered by emojis. We also identify several challenges that point toward future directions for enhancing emoji integration and facilitating social interaction.
Luyao Shen, Lik-Hang Lee, Mingming Fan 0001, Pan Hui 0001
Int. J. Hum. Comput. Interact.4
2025 "I Can't Even Recall What I Bought": How Design Influences Impulsive Buying in Douyin Live Sales
abstract
Impulsive buying tendencies exist on Douyin, the most popular Chinese social media platform, primarily due to the users’ exposure to live sales events. This study delved into examining impulsive buying behaviour, specifically triggered by external stimuli, through the lens of the Stimulus-Organism-Response framework model. Thus, our study implemented a tailored Douyin client, namely Douyin X, that contains four interventions: Visualizing the wallet’s balances, enhancing payment friction, Prolonging the duration of purchase decision-making, and imposing browsing time limits and usage statistics. Our user study with 20 participants implies individuals’ impulsive buying due to external stimuli, combined with the proliferation of impulsive buying-promoting designs, which has led to the excessive prevalence of impulsive buying in live e-commerce. Our research offers a comprehensive framework that can effectively mitigate the likelihood of impulsive buying behaviours, specifically from a design-oriented standpoint.
Zheng Wei 0003, Lik-Hang Lee, Wai Tong, Chaozhe Zhang, Huamin Qu, Pan Hui 0001
Int. J. Hum. Comput. Interact.2
2025 ChatSync: Large-Language-Model-Enabled Spatial-Temporal Knowledge Reasoning for Production Logistics Synchronization
abstract
With increasing pressure from customized demands, discrete manufacturing systems face challenges due to fluctuating resource requirements. These challenges hinder the synchronization of production logistics (PL), which is essential for coordinating resources and ensuring smooth production. Poor synchronization will result in resources waiting on each other, leading to delays and idle time. Accordingly, this paper proposes ChatSync, a framework leveraging large language model (LLM) and spatial-temporal knowledge reasoning to optimize resource allocation, delivery, and monitoring in industrial applications, particularly within the Industrial Internet of Things (IIoT) environment. First, the resource spatial-temporal graph (RSTG) is constructed by integrating real-time IIoT data and expert operational experience, enhancing the knowledge base of LLM through cross-domain knowledge fusion. Second, graph-based reasoning optimization is presented, incorporating spatial-temporal, contextual, and relational reasoning mechanisms, enabling LLM to achieve credible and responsible analysis and decision-making. Third, the PL-oriented ChatSync framework with knowledge and reasoning engines is proposed, supporting chat-based interactions for resilient resource allocation, personalized suggestion, and precise traceability. A case study in air conditioning manufacturing demonstrates that ChatSync outperforms existing benchmark methods in various PL phases, achieving a delivery punctuality rate of 91.2%.
Zhiheng Zhao, Chen Yang 0011, Sihan Huang, Lik-Hang Lee, George Q. Huang
IEEE Internet Things J.5
2025 MetaRoundWorm: A Virtual Reality Escape Room Game for Learning the Lifecycle and Immune Response to Parasitic Infections
abstract
Promoting public health is challenging owing to its abstract nature, and individuals may be apprehensive about confronting it. Recently, there has been an increasing interest in using the metaverse and gamification as novel educational techniques to improve learning experiences related to the immune system. Thus, we present MetaRoundWorm, an immersive virtual reality (VR) escape room game designed to enhance the understanding of parasitic infections and host immune responses through interactive, gamified learning. The application simulates the lifecycle of Ascaris lumbricoides and corresponding immunological mechanisms across anatomically accurate environments within the human body. Integrating serious game mechanics with embodied learning principles, MetaRoundWorm offers players a task-driven experience combining exploration, puzzle-solving, and immune system simulation. To evaluate the educational efficacy and user engagement, we conducted a controlled study comparing MetaRoundWorm against a traditional approach, i.e., interactive slides. Results indicate that MetaRoundWorm significantly improves immediate learning outcomes, cognitive engagement, and emotional experience, while maintaining knowledge retention over time. Our findings suggest that immersive VR gamification holds promise as an effective pedagogical tool for communicating complex biomedical concepts and advancing digital health education.
Xuanru Cheng, Chi-lok Tai, Lik-Hang Lee
IEEE Trans. Vis. Comput. Graph.4
2025 TeamPortal: Exploring Virtual Reality Collaboration Through Shared and Manipulating Parallel Views
abstract
Virtual Reality (VR) offers a unique collaborative experience, with parallel views playing a pivotal role in Collaborative Virtual Environments by supporting the transfer and delivery of items. Sharing and manipulating partners' views provides users with a broader perspective that helps them identify the targets and partner actions. We proposed TeamPortal accordingly and conducted two user studies with 72 participants (36 pairs) to investigate the potential benefits of interactive, shared perspectives in VR collaboration. Our first study compared ShaView and TeamPortal against a baseline in a collaborative task that encompassed a series of searching and manipulation tasks. The results show that TeamPortal significantly reduced movement and increased collaborative efficiency and social presence in complex tasks. Following the results, the second study evaluated three variants: TeamPortal+, SnapTeamPortal+, and DropTeamPortal+. The results show that both SnapTeamPortal+ and DropTeamPortal+ improved task efficiency and willingness to further adopt these technologies, though SnapTeamPortal+ reduced co-presence. Based on the findings, we proposed three design implications to inform the development of future VR collaboration systems.
Luyao Shen, Lei Chen 0088, Mingming Fan 0001, Lik-Hang Lee
IEEE Trans. Vis. Comput. Graph.5
2025 Illuminating the Scene: How Virtual Environments and Learning Modes Shape Film Lighting Mastery in Virtual Reality
abstract
In virtual reality (VR) education, particularly in creative fields like film production, the role of different virtual environments in shaping learning outcomes remains underexplored. This study investigates how three distinct environments-baseline, a dynamic beach setting, and a familiar office space-affect students' ability to learn film lighting techniques and whether team-based learning offers advantages over individual learning. We conducted a 3×2 factorial experiment with 36 participants to examine the effects of these environments on learning performance. Our results show for individual learners, the dynamic and potentially distracting beach environment increased frustration and effort but also heightened their sense of engagement and perceived performance. In contrast, team-based learning in familiar environments like the office significantly reduced frustration and fostered collaboration, leading to improved performance. Interestingly, team-based learning excelled in the baseline environment, whereas individual learners performed better in more challenging settings like the beach. These findings provide practical insights into optimizing virtual environments to enhance both individual and collaborative learning in VR education.
Zheng Wei 0003, Jia Sun 0011, Junxiang Liao, Lik-Hang Lee, Chan-In Sio, Pan Hui 0001, Huamin Qu, Wai Tong
IEEE Trans. Vis. Comput. Graph.4
2025 Transforming cinematography lighting education in the metaverse
abstract
Lighting education is a foundational component of cinematography education. However, many art schools do not have expensive soundstages for traditional cinematography lessons. Migrating physical setups to virtual experiences is a potential solution driven by metaverse initiatives. Yet there is still a lack of knowledge on the design of a VR system for teaching cinematography. We first analyzed the educational needs for cinematography lighting education by conducting interviews with six cinematography professionals from academia and industry. Accordingly, we presented Art Mirror, a VR soundstage for teachers and students to emulate cinematography lighting in virtual scenarios. We evaluated Art Mirror from the aspects of usability, realism, presence, sense of agency, and collaboration. Sixteen participants were invited to take a cinematography lighting course and assess the design elements of Art Mirror. Our results demonstrate that Art Mirror is usable and useful for cinematography lighting education, which sheds light on the design of VR cinematography education.
Wai Tong, Zheng Wei 0003, Meng Xia 0002, Lik-Hang Lee, Huamin Qu
Vis. Informatics5
2024 DreamScene: 3D Gaussian-Based Text-to-3D Scene Generation via Formation Pattern Sampling
Haoran Li 0020, Haolin Shi, Yong Liao 0003, Lin Wang 0025, Lik-Hang Lee, Peng Yuan Zhou
ECCV (74)7
2024 A Study of Partisan News Sharing in the Russian Invasion of Ukraine
abstract
Since the Russian invasion of Ukraine, a large volume of biased and partisan news has been spread via social media platforms. As this may lead to wider societal issues, we argue that understanding how partisan news sharing impacts users' communication is crucial for better governance of online communities. In this paper, we perform a measurement study of partisan news sharing. We aim to characterize the role of such sharing in influencing users' communications. Our analysis covers an eight-month dataset across six Reddit communities related to the Russian invasion. We first perform an analysis of the temporal evolution of partisan news sharing. We confirm that the invasion stimulates discussion in the observed communities, accompanied by an increased volume of partisan news sharing. Next, we characterize users' response to such sharing. We observe that partisan bias plays a role in narrowing its propagation. More biased media is less likely to be spread across multiple subreddits. However, we find that partisan news sharing attracts more users to engage in the discussion, by generating more comments. We then built a predictive model to identify users likely to spread partisan news. The prediction is challenging though, with 61.57% accuracy on average. Our centrality analysis on the commenting network further indicates that the users who disseminate partisan news possess lower network influence in comparison to those who propagate neutral news.
Ehsan ul Haq, Gareth Tyson, Lik-Hang Lee, Yuyang Wang 0002, Pan Hui 0001
ICWSM4
2024 Towards Trustworthy MetaShopping: Studying Manipulative Audiovisual Designs in Virtual-Physical Commercial Platforms
abstract
E-commerce has emerged as a significant endeavour in which technological advancements influence the shopping experience. Simultaneously, the metaverse is the next breakthrough to transform multimedia engagement. However, under such situations, deceiving designs aimed at deceiving users into making desired choices might be more successful. This paper proposes the design space of manipulative techniques in e-commerce applications for the metaverse. We construct our arguments by evaluating user interaction with manipulative design in metaverse shopping experiences, followed by a survey among users to understand the effect of counteracting manipulative e-commerce scenarios. Our findings can understanding of design guidelines according to metaverse e-commerce experiences and the possibility of opportunities to improve user awareness of manipulative experiences. © 2024 ACM.
Esmée Henrieke Anne de Haas, Lik-Hang Lee, Yiming Huang 0007, Carlos Bermejo 0001, Pan Hui 0001
ACM Multimedia2
2024 MetaDragonBoat: Exploring Paddling Techniques of Virtual Dragon Boating in a Metaverse Campus
abstract
The preservation of cultural heritage, as mandated by the United Nations Sustainable Development Goals (SDGs), is integral to sustainable urban development. This paper focuses on the Dragon Boat Festival, a prominent event in Chinese cultural heritage, and proposes leveraging Virtual Reality (VR), to enhance its preservation and accessibility. Traditionally, participation in the festival's dragon boat races was limited to elite athletes, excluding broader demographics. Our proposed solution, named MetaDragonBoat, enables virtual participation in dragon boat racing, offering immersive experiences that replicate physical exertion through a cultural journey. Thus, we build a digital twin of a university campus located in a region with a rich dragon boat racing tradition. Coupled with three paddling techniques that are enabled by either commercial controllers or physical paddle controllers with haptic feedback, diversified users can engage in realistic rowing experiences. Our results demonstrate that by integrating resistance into the paddle controls, users could simulate the physical effort of dragon boat racing, promoting a deeper understanding and appreciation of this cultural heritage.
Wei He 0028, Xiang Li 0101, Shengtian Xu, Yuzheng Chen, Chan-In Sio, Ge Lin 0001, Lik-Hang Lee
ACM Multimedia7
2024 Hearing the Moment with MetaEcho! From Physical to Virtual in Synchronized Sound Recording
abstract
In film education, high expenses and limited space significantly challenge teaching synchronized sound recording (SSR). Traditional methods, which emphasize theory with limited practical experience, often fail to bridge the gap between theoretical understanding and practical application. As such, we introduce MetaEcho, an educational virtual reality leveraging the presence theory for teaching SSR. MetaEcho provides realistic simulations of various recording equipment and facilitates communication between learners and instructors, offering an immersive learning experience that closely mirrors actual practices. An evaluation with 24 students demonstrated that MetaEcho surpasses the traditional method in presence, collaboration, usability, realism, comprehensibility, and creativity. Three experts also commented on the benefits of MetaEcho and the opportunities for promoting SSR education in the metaverse era.
Zheng Wei 0003, Yuzheng Chen, Wai Tong, Huamin Qu, Lik-Hang Lee
ACM Multimedia7
2024 Create-to-learn Paradigm: A Proxy Visual Storytelling Tool (PVST) for Stimulating Children's Story Sense and Structure
abstract
Storytelling is vital to children’s development by nurturing creative thinking, effective communication, and self-expression. Many tools have been created to support children’s creativity. Unfortunately, the existing tools do not adequately integrate visual elements with storytelling, limiting children’s imaginative potential. This study addresses the gap by introducing a proxy visual storytelling tool (PVST) that employs a character-based approach (i.e., proxy character assembling) to enhance children’s creativity and storytelling skills. Through a comparative study using Kurt Vonnegut’s “The Shape of Stories" theory, the PVST was evaluated. The results from a pilot test show that the PVST can increase children’s sense of agency and engagement in the storytelling learning process. Additionally, it can stimulate children’s creative imagination, improve their storytelling abilities, and enable them to construct more fluent and articulate narratives. The findings highlight the importance of incorporating visual storytelling elements in enhancing children’s creativity and storytelling skills, ultimately fostering a more engaging and enriching learning experience.
Ka Yan Fung, Lik-Hang Lee, Huamin Qu, Yuelu Li, Shenghui Song 0001, David Kei-Man Yip
VINCI2
2024 Jump Cut Effects in Cinematic Virtual Reality: Editing with the 30-degree Rule and 180-degree Rule
abstract
Virtual reality (VR) is an immersive medium that offers users a unique opportunity to experience a digital environment realistically. As the demand for VR content continues to grow, the importance of effective VR editing techniques becomes increasingly apparent. This paper is a pioneering work investigating the effects of jump cuts on the viewer’s sense of presence, viewing experience, and edit quality in cinematic VR. Specifically, this work focuses on using the 30-degree and 180-degree rules in VR editing to minimize the adverse effects of jump cuts. We conducted a user study with thirteen participants, who watched nine different VR edits and completed a survey for each edited video. Our results indicate that employing the 30-degree and 180-degree rules in VR editing can significantly improve the sense of presence, viewing experience, and edit quality while mitigating the negative effects of jump cuts. We provide valuable insights for VR content creators and editors to achieve more effective and immersive VR experiences.
Lik-Hang Lee, Yuyang Wang 0002, Shan Jin 0002, Danlu Fei, Pan Hui 0001
VR2
2024 APT-Pipe: A Prompt-Tuning Tool for Social Data Annotation using ChatGPT
abstract
Recent research has highlighted the potential of LLMs, like ChatGPT, for performing label annotation on social computing data. However, it is already well known that performance hinges on the quality of the input prompts. To address this, there has been a flurry of research into prompt tuning --- techniques and guidelines that attempt to improve the quality of prompts. Yet these largely rely on manual effort and prior knowledge of the dataset being annotated. To address this limitation, we propose APT-Pipe, an automated prompt-tuning pipeline. APT-Pipe aims to automatically tune prompts to enhance ChatGPT's text classification performance on any given dataset. We implement APT-Pipe and test it across twelve distinct text classification datasets. We find that prompts tuned by APT-Pipe help ChatGPT achieve higher weighted F1-score on nine out of twelve experimented datasets, with an improvement of 7.01% on average. We further highlight APT-Pipe's flexibility as a framework by showing how it can be extended to support additional tuning mechanisms.
Zhizhuo Yin, Gareth Tyson, Ehsan ul Haq, Lik-Hang Lee, Pan Hui 0001
WWW5
2024 The Dark Side of Augmented Reality: Exploring Manipulative Designs in AR
abstract
Augmented Reality (AR) applications are becoming more mainstream, with successful examples in the mobile environment like Pokemon GO. Current malicious techniques can exploit these environments’ immersive and mixed nature (physical-virtual) to trick users into providing more personal information, i.e., dark patterns. Dark patterns are deceiving techniques (e.g., interface tricks) designed to influence individuals’ behavioural decisions. However, there are few studies regarding dark patterns’ potential issues in AR environments. In this work, using scenario construction to build our prototypes, we investigate the potential future approaches that dark patterns can have. We use VR mockups in our user study to analyze the effects of dark patterns in AR. Our study indicates that dark patterns are effective in immersive scenarios, and the use of novel techniques, such as “haptic grabbing” to draw participants’ attention, can influence their movements. Finally, we discuss the impact of such malicious techniques and what techniques can mitigate them.
Lik-Hang Lee, Carlos Bermejo 0001, Pan Hui 0001
Int. J. Hum. Comput. Interact.2
2024 Danger, Nuisance, Disregard: Analyzing User-Generated Videos for Augmented Reality Gameplay on Hand-held Devices
abstract
Augmented Reality (AR) has been largely deployed on smartphones in recent years. AR gaming, featured with geo-reference, is anchored to our real-world environments. Nevertheless, designing such AR user interaction in-the-wild is under-explored. Therefore, we examined 242 YouTube videos regarding AR gameplay, primarily Pokémon GO, (1) to reveal personal and social threats during gameplay, and (2) to identify how such threats connect to the user and spatial contexts. Our video analysis generalises the threats, including but not limited to hitting bystanders, falling off a cliff, car crashes, crowds and blockages, and stampedes. Through video analysis, we connect the threats with the user and spatial contexts. Subsequently, we propose several design tactics to deliver safe and nuisance-free AR usages in physical environments. Finally, we suggest strategies to consider large-scale deployment of AR at the levels of individuals and local communities, such as an auditable and accountable framework for user safety and social acceptability.
Lik-Hang Lee
Proc. ACM Hum. Comput. Interact.1
2023 Understanding Characteristics of Catalyst Users in the WallStreetBets Community
abstract
WallStreetBets (WSB), a Reddit community, impacted stock markets during the 2021 GameStop Short Squeeze. We examine the content and user properties that influence engagement in WSB. Despite WSB's association with emojis and informal terms, engagement among community members depends on more than surface-level factors. Although emojis are commonly used, they are not as effective at fostering interactions among users. Community members engage more with posts that have longer and topic-specific text. Simply producing a high volume of posts is not enough to attract an audience. Consistent topical focus, reciprocal interactions, and previous authorship of catalyst posts influence engagement. WSB posts, regardless of length, generally remain relevant to the community's theme of stock trading. Our findings provide insights into WSB engagement patterns and can be useful for downstream research, such as financial predictive tasks using WSB data.
Ehsan ul Haq, Haodi Weng, Gareth Tyson, Lik-Hang Lee, Reza Hadi Mogavi, Tristan Braud, Pan Hui 0001
ASONAM6
2023 Designing Loving-Kindness Meditation in Virtual Reality for Long-Distance Romantic Relationships
abstract
Loving-kindness meditation (LKM) is used in clinical psychology for couples' relationship therapy, but physical isolation can make the relationship more strained and inaccessible to LKM. Virtual reality (VR) can provide immersive LKM activities for long-distance couples. However, no suitable commercial VR applications for couples exist to engage in LKM activities of long-distance. This paper organized a series of workshops with couples to build a prototype of a couple-preferred LKM app. Through analysis of participants' design works and semi-structured interviews, we derived design considerations for such VR apps and created a prototype for couples to experience. We conducted a study with couples to understand their experiences of performing LKM using the VR prototype and a traditional video conferencing tool. Results show that LKM session utilizing both tools has a positive effect on the intimate relationship and the VR prototype is a more preferable tool for long-term use. We believe our experience can inform future researchers.
Xiaoyu Mo, Lik-Hang Lee, Xiaoying Wei, Xiaofu Jin, Mingming Fan 0001, Pan Hui 0001
ACM Multimedia3
2023 Feeling Present! From Physical to Virtual Cinematography Lighting Education with Metashadow
abstract
The high cost and limited availability of soundstages for cinematography lighting education pose significant challenges for art institutions. Traditional teaching methods, combining basic lighting equipment operation with slide lectures, often yield unsatisfactory results, hindering students' mastery of cinematography lighting techniques. Therefore, we propose Metashadow, a virtual reality (VR) cinematography lighting education system demonstrating the feasibility of learning in a virtual soundstage. Based on the presence theory, Metashadow features high-fidelity lighting devices that enable users to adjust multiple parameters, providing a quantifiable learning approach. We evaluated Metashadow with 24 participants and found that it provides better learning outcomes than traditional teaching methods regarding presence, collaboration, usability, realism, creativity, and flexibility. Six experts also praised the Metashadow's expressiveness and its learning outcomes. Our study demonstrates the potential of VR technology to enhance cinematography lighting education while imposing a smaller cost burden and space requirement.
Zheng Wei 0003, Lik-Hang Lee, Wai Tong, Huamin Qu, Pan Hui 0001
ACM Multimedia3
2023 Demo: Real-Time WebXR Edge-based Object Detection for AR
abstract
Web-based extended reality (WebXR) can enable lightweight, easy-to-access, and cross-platform augmented reality (AR) experiences. Context awareness is one key feature of AR. Supporting this in browser-based WebXR applications is challenging as typical object detection algorithms are too computationally demanding to be run in-browser, leading to slow response times and decreased battery life. In this demo, we show a WebXR AR application that uses a technique of WebRTC-based video streaming to obtain a usable video stream on an edge server to perform object detection.
Jacky Cao, Kit-Yung Lam, Lik-Hang Lee
MobiSys3
2023 Development of an immersive simulator for improving student chemistry learning efficiency
abstract
Virtual reality (VR) technology has been used for educational purposes in different learning contents during teaching and training. VR could improve users’ learning efficiency and motivation to study abstract concepts. This work designed a VR environment for chemistry education to support computer-mediated hands-on exercises, including Self-propagating high-temperature synthesis (SHS) and Electrode sheet fabrication (ESF). In our evaluation with 39 participants who wore heart beat measurement wearables, we compared the students’ performances in hands-on chemistry tasks, either with or without score-keeping and time-sensitive conditions. Accordingly, we designed questionnaires reflecting sixteen qualitative aspects (e.g., content, perspicuity, and interaction) and perceived user workloads. The experimental results indicate participants’ preferences and attitudes in terms of efficiency and sense of safety. 94.87% of participants reported that the learning simulator could improve learning efficiency, and 92.31% of the participants indicated that it can improve their sense of safety. The results of the data analysis show that the different learning scenarios we simulated have positive significance. Our findings shed light on the quality and learning performance of operational skills for chemistry education.
Shan Jin 0002, Yuyang Wang 0002, Lik-Hang Lee, Pan Hui 0001
VINCI3
2023 MyoKey: Inertial Motion Sensing and Gesture-Based QWERTY Keyboard for Extended Realities
abstract
Usability challenges and social acceptance of textual input in a context of extended realities (XR) motivate the research of novel input modalities. We investigate the fusion of inertial measurement unit (IMU) control and surface electromyography (sEMG) gesture recognition applied to text entry using a QWERTY-layout virtual keyboard. We design, implement, and evaluate the proposed multi-modal solution named MyoKey. The user can select characters with a combination of arm movements and hand gestures. MyoKey employs a lightweight convolutional neural network classifier that can be deployed on a mobile device with insignificant inference time. We demonstrate the practicality of interruption-free text entry with MyoKey, by recruiting 12 participants and by testing three sets of grasp micro-gestures in three scenarios: empty hand text input, tripod grasp (e.g., pen), and a cylindrical grasp (e.g., umbrella). With MyoKey, users achieve an average text entry rate of 9.33 words per minute (WPM), 8.76 WPM, and 8.35 WPM for the freehand, tripod grasp, and cylindrical grasp conditions, respectively.
Kirill A. Shatilov, Young D. Kwon, Lik-Hang Lee, Dimitris Chatzopoulos, Pan Hui 0001
IEEE Trans. Mob. Comput.3
2022 Designing a Game for Pre-Screening Students with Specific Learning Disabilities in Chinese
abstract
Most students with specific learning disabilities (SLDs) have difficulties in reading and writing. The SLDs pre-screening is crucial because the golden period for therapy is before six years old. However, many students in Hong Kong receive SLDs assessments after the golden period. Also, the SLDs pre-screening is challenging, especially in a language with the logographic script but without prominent sound-script correspondence (e.g., Chinese, Japanese). To make pre-screening SLDs in Chinese more effective and efficient, we designed a new comprehensive pre-screening game for SLDs in Chinese (i.e., dyslexia, dysgraphia, and dyspraxia). Notably, we designed a Chinese morphological awareness puzzle that challenges students to recognize different words made up with the first character that is identical and the second character that is different, such as樹枝 (literally means tree branch),樹幹 (literally means tree truck),樹葉 (literally means tree leaves), and樹根 (literally means tree root). We experimented with students, which showed that our game can effectively pre-screen students with SLDs in Chinese. Our work contributes an approach to quick SLDs in Chinese pre-screening, potentially useful for other logographic languages (e.g., Japanese).
Ka Yan Fung, Kuen Fung Sin, Zikai Wen, Lik-Hang Lee, Shenghui Song 0001, Huamin Qu
ASSETS4
2022 Exploring Mental Health Communications among Instagram Coaches
abstract
There has been a significant expansion in the use of online social networks (OSNs) to support people experiencing mental health issues. This paper studies the role of Instagram influencers who specialize in coaching people with mental health issues. Using a dataset of 97k posts, we characterize such users' linguistic and behavioural features. We explore how these observations impact audience engagement (as measured by likes). We show that the support provided by these accounts varies based on their self-declared professional identities. For instance, Instagram accounts that declare themselves as Authors offer less support than accounts that label themselves as a Coach. We show that increasing information support in general communication positively affects user engagement. However, the effect of vocabulary on engagement is not consistent across the Instagram account types. Our findings shed light on this understudied topic and guide how mental health practitioners can improve outreach.
Ehsan ul Haq, Lik-Hang Lee, Gareth Tyson, Reza Hadi Mogavi, Tristan Braud, Pan Hui 0001
ASONAM2
2022 Causal Analysis on the Anchor Store Effect in a Location-based Social Network
abstract
A particular phenomenon of interest in Retail Eco-nomics is the spillover effect of anchor stores (specific stores with a reputable brand) to non-anchor stores in terms of customer traffic. Prior works in this area rely on small and survey-based datasets that are often confidential or expensive to collect on a large scale. Also, very few works study the underlying causal mechanisms between factors that underpin the spillover effect. In this work, we analyze the causal relationship between anchor stores and customer traffic to non-anchor stores and employ a propensity score matching framework to investigate this effect more efficiently. First of all, to demonstrate the effect, we leverage open and mobile data from London Datastore and Location-Based Social Networks (LBSNs) such as Foursquare. We then perform a large-scale empirical analysis of customer visit patterns from anchor stores to non-anchor stores (e.g., non-chain restaurants) located in the Greater London area as a case study. By studying over 600 neighbourhoods in the Greater London area, we find that anchor stores cause a 14.2-26.5% increase in customer traffic for the non-anchor stores reinforcing the established economic theory Moreover, we evaluate the efficiency of our methodology by studying the confounder balance, dose difference and performance of the matching framework on synthetic data. Through this work, we point decision-makers in the retail industry to a more systematic approach to estimate the anchor store effect and pave the way for further research to discover more complex causal relationships underlying this effect with open data.
Anish K. Vallapuram, Young D. Kwon, Lik-Hang Lee, Fengli Xu, Pan Hui 0001
ASONAM3
2022 Beyond the Blue Sky of Multimodal Interaction: A Centennial Vision of Interplanetary Virtual Spaces in Turn-based Metaverse
abstract
Human habitation across multiple planets requires communication and social connection between planets. When the infrastructure of a deep space network becomes mature, immersive cyberspace, known as the Metaverse, can exchange diversified user data and host multitudinous virtual worlds. Nevertheless, such immersive cyberspace unavoidably encounters latency in minutes, and thus operates in a turn-taking manner. This Blue Sky paper illustrates a vision of an interplanetary Metaverse that connects Earthian and Martian users in a turn-based Metaverse. Accordingly, we briefly discuss several grand challenges to catalyze research initiatives for the ‘Digital Big Bang’ on Mars.
Lik-Hang Lee, Carlos Bermejo 0001, Ahmad Yousef Alhilal, Tristan Braud, Simo Hosio, Esmée Henrieke Anne de Haas, Pan Hui 0001
ICMI1
2022 Decentralized, not Dehumanized in the Metaverse: Bringing Utility to NFTs through Multimodal Interaction
abstract
User Interaction for NFTs (Non-fungible Tokens) is gaining increasing attention. Although NFTs have been traditionally single-use and monolithic, recent applications aim to connect multimodal interaction with human behavior. This paper reviews the related technological approaches and business practices in NFT art. We highlight that multimodal interaction is a currently under-studied issue in mainstream NFT art, and conjecture that multimodal interaction is a crucial enabler for decentralization in the NFT community. We present a continuum theory and propose a framework combining a bottom-up approach with AI multimodal process. Through this framework, we put forward integrating human behavior data into generative NFT units, as "multimodal interactive NFT." Our work displays the possibilities of NFTs in the art world, beyond the traditional 2D and 3D static content.
Anqi Wang 0003, Ze Gao 0003, Lik-Hang Lee, Tristan Braud, Pan Hui 0001
ICMI3
2022 Towards Reproducible Evaluations for Flying Drone Controllers in Virtual Environments
abstract
Research attention on natural user interfaces (NUIs) for drone flights are rising. Nevertheless, NUIs are highly diversified, and primarily evaluated by different physical environments leading to hard-to-compare performance between such solutions. We propose a virtual environment, namely VRFlightSim, enabling comparative evaluations with enriched drone flight details to address this issue. We first replicated a state-of-the-art (SOTA) interface and designed two tasks (crossing and pointing) in our virtual environment. Then, two user studies with 13 participants demonstrate the necessity of VRFlightSim and further highlight the potential of open-data interface designs.
Yiming Huang 0007, Yui-Pan Yau, Pan Hui 0001, Lik-Hang Lee
IROS5
2022 PassWalk: Spatial Authentication Leveraging Lateral Shift and Gaze on Mobile Headsets
abstract
Secure and usable user authentication on mobile headsets is a challenging problem. The miniature-sized touchpad on such devices becomes a hurdle to user interactions that impact usability. However, the most common authentication methods, i.e., the standard QWERTY virtual keyboard or mid-air inputs to enter passwords are highly vulnerable to shoulder surfing attacks. In this paper, we present PassWalk, a keyboard-less authentication system leveraging multi-modal inputs on mobile headsets. PassWalk demonstrates the feasibility of user authentication driven by the user's gaze and lateral shifts (i.e., footsteps) simultaneously. The keyboard-less authentication interface in PassWalk enables users to accomplish highly mobile inputs of graphical passwords, containing digital overlays and physical objects. We conduct an evaluation with 22 recruited participants (15 legitimate users and 7 attackers). Our results show that PassWalk provides high security (only 1.1% observation attacks were successful) with a mean authentication time of 8.028s, which outperforms the commercial method of using the QWERTY virtual keyboard (21.5% successful attacks) and a research prototype LookUnLock (5.5% successful attacks). Additionally, PassWalk entails a significantly smaller workload on the user than the current commercial methods.
Abhishek Kumar 0011, Lik-Hang Lee, Jagmohan Chauhan, Xiang Su 0001, Mohammad Ashraful Hoque, Susanna Pirttikangas, Sasu Tarkoma, Pan Hui 0001
ACM Multimedia2
2022 Human-Avatar Interaction in Metaverse: Framework for Full-Body Interaction
abstract
The metaverse is a network of shared virtual environments where people can interact synchronously through their avatars. To enable this, it is necessary to accurately capture and recreate (physical) human motion. This is used to render avatars correctly, reflecting the motion of their corresponding users. In large-scale environments this must be done in real-time. This paper proposes a human-avatar framework with full-body motion capture. Its goal is to deliver high-accuracy capture with low computational and network overheads. It relies on a lightweight Octree data structure to record and transmit motion to other users. We conduct a user study with 22 participants and perform a preliminary evaluation of its scalability. Our user study shows that Octree with Inverse Kinematic achieves the best trade-off, achieving low delay and high accuracy. Our proposed solution delivers the lowest delay, with an average of 67ms in an environment of 8 concurrent users. It attains a 55.7% improvement over the prior techniques.
Kit-Yung Lam, Ahmad Yousef Alhilal, Lik-Hang Lee, Gareth Tyson, Pan Hui 0001
MMAsia4
2022 3DeformR: freehand 3D model editing in virtual environments considering head movements on mobile headsets
abstract
3D objects are the primary media in virtual reality environments in immersive cyberspace, also known as the Metaverse. Users, through editing such objects, can communicate with other individuals on mobile headsets. Knowing that the tangible controllers cause the burden to carry such addendum devices, the body-centric interaction techniques, such as hand gestures, get rid of such burdens. However, object editing with hand gestures is usually overlooked. Accordingly, we propose and implement a palm-based virtual embodiment for hand gestural model editing, namely 3DeformR. We employ three optimized hand gestures on bi-harmonic deformation algorithms that enable selecting and editing 3D models in fine granularity. Our evaluation with nine participants considers three interaction techniques (two-handed tangible controller (OMC), a naive implementation of hand gestures (SH), and 3DeformR. Two experimental tasks of planar and spherical objects imply that 3DeformR outperforms SH, in terms of task completion time (~51%) and required actions (~17%). Also, our participants with 3DeformR make significantly better performance than the commercial standard (OMC) - saved task time (~43%) and actions (~3%). Remarkably, the edited objects by 3DeformR show no discernible difference from those with tangible controllers characterised by accurate and responsive detection.
Kit-Yung Lam, Lik-Hang Lee, Pan Hui 0001
MMSys2
2022 Screenshots, Symbols, and Personal Thoughts: The Role of Instagram for Social Activism
abstract
In this paper, we highlight the use of Instagram for social activism, taking 2019 Hong Kong protests as a case study. Instagram focuses on image content and provides users with few features to share or repost, limiting information propagation. Nevertheless, users who are politically active offline also share their activism on Instagram. We first evaluate the effect of protests on social media activity for protesters and non-protesters over two significant protests. Protesters’ exposure to protest-related posts is much higher than non-protesters, and their network activity follows the protest schedule. They are also much more active on posts related to the protest that they participate in than the other protest. We then analyze the images posted by the users. Users predominantly use symbols related to protests and share personal thoughts on its primary actors. Users primarily share content to raise their network’s awareness, and the content choice is directly affected by Instagram’s intrinsic interaction modalities.
Ehsan ul Haq, Tristan Braud, Yui-Pan Yau, Lik-Hang Lee, Franziska B. Keller, Pan Hui 0001
WWW4
2022 EdgeXAR: A 6-DoF Camera Multi-target Interaction Framework for MAR with User-friendly Latency Compensation
abstract
The computational capabilities of recent mobile devices enable the processing of natural features for Augmented Reality (AR), but the scalability is still limited by the devices' computation power and available resources. In this paper, we propose EdgeXAR, a mobile AR framework that utilizes the advantages of edge computing through task offloading to support flexible camera-based AR interaction. We propose a hybrid tracking system for mobile devices that provides lightweight tracking with 6 Degrees of Freedom and hides the offloading latency from users' perception. A practical, reliable and unreliable communication mechanism is used to achieve fast response and consistency of crucial information. We also propose a multi-object image retrieval pipeline that executes fast and accurate image recognition tasks on the cloud and edge servers. Extensive experiments are carried out to evaluate the performance of EdgeXAR by building mobile AR apps upon it. Regarding the Quality of Experience (QoE), the mobile AR apps powered by EdgeXAR framework run on average at the speed of 30 frames per second with precise tracking of only 1-2 pixel errors and accurate image recognition of at least 97% accuracy. As compared to Vuforia, one of the leading commercial AR frameworks, EdgeXAR transmits 87% less data while providing a stable 30FPS performance and reducing the offloading latency by 50 to 70% depending on the transmission medium. Our work facilitates the large-scale deployment of AR as the next generation of ubiquitous interfaces.
Sikun Lin, Farshid Hassani Bijarbooneh, Hao Fei Cheng, Tristan Braud, Peng Yuan Zhou, Lik-Hang Lee, Pan Hui 0001
Proc. ACM Hum. Comput. Interact.7
2022 AICP: Augmented Informative Cooperative Perception
abstract
Connected vehicles, whether equipped with advanced driver-assistance systems or fully autonomous, require human driver supervision and are currently constrained to visual information in their line-of-sight. A cooperative perception system among vehicles increases their situational awareness by extending their perception range. Existing solutions focus on improving perspective transformation and fast information collection. However, such solutions fail to filter out large amounts of less relevant data and thus impose significant network and computation load. Moreover, presenting all this less relevant data can overwhelm the driver and thus actually hinder them. To address such issues, we present Augmented Informative Cooperative Perception (AICP), the first fast-filtering system which optimizes the informativeness of shared data at vehicles to improve the fused presentation. To this end, an informativeness maximization problem is presented for vehicles to select a subset of data to display to their drivers. Specifically, we propose (i) a dedicated system design with custom data structure and lightweight routing protocol for convenient data encapsulation, fast interpretation and transmission, and (ii) a comprehensive problem formulation and efficient fitness-based sorting algorithm to select the most valuable data to display at the application layer. We implement a proof-of-concept prototype of AICP with a bandwidth-hungry, latency-constrained real-life augmented reality application. The prototype adds only 12.6 milliseconds of latency to a current informativeness-unaware system. Next, we test the networking performance of AICP at scale and show that AICP effectively filters out less relevant packets and decreases the channel busy time.
Peng Yuan Zhou, Pranvera Kortoçi, Yui-Pan Yau, Benjamin Finley, Xiujun Wang, Tristan Braud, Lik-Hang Lee, Sasu Tarkoma, Jussi Kangasharju, Pan Hui 0001
IEEE Trans. Intell. Transp. Syst.7
2021 PARA: Privacy Management and Control in Emerging IoT Ecosystems using Augmented Reality
abstract
The ubiquity of smart devices, combined with a lack of information about data garnered by them, make privacy a significant challenge for adopting smart devices. Ensuring users can safeguard their privacy without compromising the devices’ functionality requires effective yet intuitive ways to manage personal privacy preferences. Current solutions for privacy management are severely lacking as they are ineffective in making users aware of potential privacy risks or how to mitigate them and as they offer limited support for interaction. As our first contribution, we develop a novel AR privacy management interface (PARA) that uses AR visualization to contextualize data disclosure and improve user’s perceptions of privacy threats. Besides offering support for enhancing user’s privacy perceptions, our interface supports privacy control on compatible devices through privacy-enhancing technologies. As our second contribution, we systematically study factors affecting privacy perceptions and privacy control for two device classes (smart camera and smart speaker) through a user study with N = 32 participants. Our results show that PARA’s contextualization and visualization of privacy disclosure strongly affect the participants’ privacy perceptions. For privacy control, we demonstrate that our prototype improves the participant’s capability to identify risks and provides an effective and easy-to-use mechanism for controlling privacy disclosure, in contrast to existing state-of-the-art privacy management interfaces.
Carlos Bermejo 0001, Lik-Hang Lee, Petteri Nurmi, Pan Hui 0001
ICMI2
2021 Theophany: Multimodal Speech Augmentation in Instantaneous Privacy Channels
abstract
Many factors affect speech intelligibility in face-to-face conversations. These factors lead conversation participants to speak louder and more distinctively, exposing the content to potential eavesdroppers. To address these issues, we introduce Theophany, a privacy-preserving framework for augmenting speech. Theophany establishes ad-hoc social networks between conversation participants to exchange contextual information, improving speech intelligibility in real-time. At the core of Theophany, we develop the first privacy perception model that assesses the privacy risk of a face-to-face conversation based on its topic, location, and participants. This framework allows to develop any privacy-preserving application for face-to-face conversation. We implement the framework within a prototype system that augments the speaker's speech with real-life subtitles to overcome the loss of contextual cues brought by mask-wearing and social distancing during the COVID-19 pandemic. We evaluate Theophany through a user survey and a user study on 53 and 17 participants, respectively. Theophany's privacy predictions match the participants' privacy preferences with an accuracy of 71.26%. Users considered Theophany to be useful to protect their privacy (3.88/5), easy to use (4.71/5), and enjoyable to use (4.24/5). We also raise the question of demographic and individual differences in the design of privacy-preserving solutions.
Abhishek Kumar 0011, Tristan Braud, Lik-Hang Lee, Pan Hui 0001
ACM Multimedia3
2021 A2W: Context-Aware Recommendation System for Mobile Augmented Reality Web Browser
abstract
Augmented Reality (AR) offers new capabilities for blurring the boundaries between physical reality and digital media. However, the capabilities of integrating web contents and AR remain underexplored. This paper presents an AR web browser with an integrated context-aware AR-to-Web content recommendation service named as A2W browser, to provide continuously user-centric web browsing experiences driven by AR headsets. We implement the A2W browser on an AR headset as our demonstration application, demonstrating the features and performance of A2W framework. The A2W browser visualizes the AR-driven web contents to the user, which is suggested by the content-based filtering model in our recommendation system. In our experiments, 20 participants with the adaptive UIs and recommendation system in A2W browser achieve up to 30.69% time saving compared to smartphone conditions. Accordingly, A2W-supported web browsing on workstations facilitates the recommended information leading to 41.67% faster reaches to the target information than typical web browsing.
Kit-Yung Lam, Lik-Hang Lee, Pan Hui 0001
ACM Multimedia2
2021 Exploring Button Designs for Mid-air Interaction in Virtual Reality: A Hexa-metric Evaluation of Key Representations and Multi-modal Cues
abstract
The continued advancement in user interfaces comes to the era of virtual reality that requires a better understanding of how users will interact with 3D buttons in mid-air. Although virtual reality owns high levels of expressiveness and demonstrates the ability to simulate the daily objects in the physical environment, the most fundamental issue of designing virtual buttons is surprisingly ignored. To this end, this paper presents four variants of virtual buttons, considering two design dimensions of key representations and multi-modal cues (audio, visual, haptic). We conduct two multi-metric assessments to evaluate the four virtual variants and the baselines of physical variants. Our results indicate that the 3D-lookalike buttons help users with more refined and subtle mid-air interactions (i.e. lesser press depth) when haptic cues are available; while the users with 2D-lookalike buttons unintuitively achieve better keystroke performance than the 3D counterparts. We summarize the findings, and accordingly, suggest the design choices of virtual reality buttons among the two proposed design dimensions.
Carlos Bermejo 0001, Lik-Hang Lee, Paul Chojecki, David Przewozny, Pan Hui 0001
Proc. ACM Hum. Comput. Interact.2
2021 Press-n-Paste: Copy-and-Paste Operations with Pressure-sensitive Caret Navigation for Miniaturized Surface in Mobile Augmented Reality
abstract
Copy-and-paste operations are the most popular features on computing devices such as desktop computers, smartphones and tablets. However, the copy-and-paste operations are not sufficiently addressed on the Augmented Reality (AR) smartglasses designated for real-time interaction with texts in physical environments. This paper proposes two system solutions, namely Granularity Scrolling (GS) and Two Ends (TE), for the copy-and-paste operations on AR smartglasses. By leveraging a thumb-size button on a touch-sensitive and pressure-sensitive surface, both the multi-step solutions can capture the target texts through indirect manipulation and subsequently enables the copy-and-paste operations. Based on the system solutions, we implemented an experimental prototype named Press-n-Paste (PnP). After the eight-session evaluation capturing 1,296 copy-and-paste operations, 18 participants with GS and TE achieve the peak performance of 17,574 ms and 13,951 ms per copy-and-paste operation, with 93.21% and 98.15% accuracy rates respectively, which are as good as the commercial solutions using direct manipulation on touchscreen devices. The user footprints also show that PnP has a distinctive feature of miniaturized interaction area within 12.65 mm * 14.48 mm. PnP not only proves the feasibility of copy-and-paste operations with the flexibility of various granularities on AR smartglasses, but also gives significant implications to the design space of pressure widgets as well as the input design on smart wearables.
Lik-Hang Lee, Yui-Pan Yau, Pan Hui 0001, Susanna Pirttikangas
Proc. ACM Hum. Comput. Interact.1
2021 Emerging ExG-based NUI Inputs in Extended Realities: A Bottom-up Survey
abstract
Incremental and quantitative improvements of two-way interactions with e x tended realities (XR) are contributing toward a qualitative leap into a state of XR ecosystems being efficient, user-friendly, and widely adopted. However, there are multiple barriers on the way toward the omnipresence of XR; among them are the following: computational and power limitations of portable hardware, social acceptance of novel interaction protocols, and usability and efficiency of interfaces. In this article, we overview and analyse novel natural user interfaces based on sensing electrical bio-signals that can be leveraged to tackle the challenges of XR input interactions. Electroencephalography-based brain-machine interfaces that enable thought-only hands-free interaction, myoelectric input methods that track body gestures employing electromyography, and gaze-tracking electrooculography input interfaces are the examples of electrical bio-signal sensing technologies united under a collective concept of ExG. ExG signal acquisition modalities provide a way to interact with computing systems using natural intuitive actions enriching interactions with XR. This survey will provide a bottom-up overview starting from (i) underlying biological aspects and signal acquisition techniques, (ii) ExG hardware solutions, (iii) ExG-enabled applications, (iv) discussion on social acceptance of such applications and technologies, as well as (v) research challenges, application directions, and open problems; evidencing the benefits that ExG-based Natural User Interfaces inputs can introduce to the area of XR.
Kirill A. Shatilov, Dimitris Chatzopoulos, Lik-Hang Lee, Pan Hui 0001
ACM Trans. Interact. Intell. Syst.3
2020 Force9: Force-assisted Miniature Keyboard on Smart Wearables
abstract
Smartwatches and other wearables are characterized by small-scale touchscreens that complicate the interaction with content. In this paper, we present Force9, the first optimized miniature keyboard leveraging force-sensitive touchscreens on wrist-worn computers. Force9 enables character selection in an ambiguous layout by analyzing the trade-off between interaction space and the easiness of force-assisted interaction. We argue that dividing the screen's pressure range into three contiguous force levels is sufficient to differentiate characters for fast and accurate text input. Our pilot study captures and calibrates the ability of users to perform force-assisted touches on miniature-sized keys on touchscreen devices. We then optimize the keyboard layout considering the goodness of character pairs (with regards to the selected English corpus) under the force-based configuration and the users? familiarity with the QWERTY layout. We finally evaluate the performance of the trimetric optimized Force9 layout, and achieve an average of 10.18 WPM by the end of the final session. Compared to the other state-of-the-art approaches, Force9 allows for single-gesture character selection without addendum sensors.
Lik-Hang Lee, Ngo Yan Yeung, Tristan Braud, Tong Li 0013, Xiang Su 0001, Pan Hui 0001
ICMI1
2020 UbiPoint: towards non-intrusive mid-air interaction for hardware constrained smart glasses
abstract
Throughout the past decade, numerous interaction techniques have been designed for mobile and wearable devices. Among these devices, smartglasses mostly rely on hardware interfaces such as touchpad and buttons, which are often cumbersome and counterintuitive to use. Furthermore, smartglasses feature cheap and low-power hardware preventing the use of advanced pointing techniques. To overcome these issues, we introduce UbiPoint, a freehand mid-air interaction technique. UbiPoint uses the monocular camera embedded in smartglasses to detect the user's hand without relying on gloves, markers, or sensors, enabling intuitive and non-intrusive interaction. We introduce a computationally fast and light-weight algorithm for fingertip detection, which is especially suited for the limited hardware specifications and the short battery life of smartglasses. UbiPoint processes pictures at a rate of 20 frames per second with high detection accuracy - no more than 6 pixels deviation. Our evaluation shows that UbiPoint, as a mid-air non-intrusive interface, delivers a better experience for users and smart glasses interactions, with users completing typical tasks 1.82 times faster than when using the original hardware.
Lik-Hang Lee, Tristan Braud, Farshid Hassani Bijarbooneh, Pan Hui 0001
MMSys1
2020 One-thumb Text Acquisition on Force-assisted Miniature Interfaces for Mobile Headsets
abstract
Touchscreen interfaces are shrinking and even dis-appearing on mobile headsets. The existing approaches for text acquisition on mobile headsets, for instance, speech commands and hand gestures, are cumbersome and coarse. In this paper, we show the feasibility of interaction on a miniature area as small as 12 * 13 mm2that offers an input alternative on small form-factor devices such as smartwatches, smart rings, or the spectacles frames of mobile headsets. To this end, we propose and implement two interaction approaches, namely FRS and DupleFR, for acquiring textual contents on mobile headsets. Both approaches leverage force-assisted interaction on a miniature-size interface. They enable the user to acquire textual content with various granularities such as characters, words, sentences, paragraphs, and the entire text. After 8 sessions, 22 participants with FRS and DupleFR achieve the peak performance of respectively 11.455 and 10.611 seconds per textual acquisition with accuracy rates of 91.41% and 94.95%. Although FRS and DupleFR as indirect manipulations are disadvantageous, they are at least 37.06% faster than the commercial standards designated to direct manipulation on touchscreens.
Lik-Hang Lee, Yui-Pan Yau, Tristan Braud, Xiang Su 0001, Pan Hui 0001
PerCom1
2020 From seen to unseen: Designing keyboard-less interfaces for text entry on the constrained screen real estate of Augmented Reality headsets
Lik-Hang Lee, Tristan Braud, Kit-Yung Lam, Yui-Pan Yau, Pan Hui 0001
Pervasive Mob. Comput.1
2019 TiPoint: detecting fingertip for mid-air interaction on computational resource constrained smartglasses
abstract
Smartglasses mostly rely on hardware interfaces such as touch-pad and buttons, which are often cumbersome and counter-intuitive to use. Furthermore, smartglasses feature cheap and low-power hardware preventing the use of advanced pointing techniques. To overcome these issues, we introduce TiPoint, a freehand mid-air interaction technique. TiPoint uses the monocular camera embedded in smartglasses to detect the user's hand, enabling intuitive and non-intrusive interaction. We introduce a light-weight algorithm for fingertip detection, which is especially suited for the limited hardware specifications and the short battery life time of smartglasses. Our evaluation shows that TiPoint as a mid-air non-intrusive interface delivers a better experience for users and smart glasses interactions, with users completing typical tasks 1.82 times faster than when using the original hardware.
Lik-Hang Lee, Tristan Braud, Farshid Hassani Bijarbooneh, Pan Hui 0001
UbiComp1
2019 M2A: A Framework for Visualizing Information from Mobile Web to Mobile Augmented Reality
abstract
Mobile Augmented Reality (MAR) drastically changes our approach to computing and user interaction. Web browsing, in particular, is impractical on AR devices as current web design principles do not account for three-dimensional display and navigation of virtual content. In this paper, we propose Mobile to AR (M2A), the first framework for designing web pages for AR devices. M2A exploits the visual context to display more content while enabling users to locate relevant data intuitively with minimal modifications to the website. To evaluate the principles behind the framework, we implement a demonstration application in AR and conduct two user-focused experiments. Our experimental study reveals that participants with M2A are 5 times faster to find information on a web page compared to a smartphone, and 2 times faster than a traditional AR web browser. Furthermore, users consider navigation on M2A websites to be significantly more intuitive and easy to use compared to their desktop and mobile counterparts.
Kit-Yung Lam, Lik-Hang Lee, Tristan Braud, Pan Hui 0001
PerCom2
2019 HIBEY: Hide the Keyboard in Augmented Reality
abstract
Text input is a very challenging task in Augmented Reality (AR). On non-touch AR headsets, virtual keyboards are counter-intuitive and character keys are hard to locate inside the constrained screen real estate. In this paper, we present the design, implementation and evaluation of HIBEY, a text input system for smartglasses. HIBEY provides a fast, reliable, affordable, and easy-to-use text entry solution through vision-based freehand interactions. Supported by a probabilistic spatial model and a language model, a three-level holographic environment enables users to apply fast and continuous hand gesture to pick characters and predictive words in a keyboardless interface. Through the pilot study and a thorough evaluations lasting 8 days, we show that HIBEY leads to a mean text entry rate of 9.95 word per minute (WPM) with 96.06% accuracy, which is comparable to other state-of-the-art approaches. After 8 days, participants can achieve an average of 13.19 WPM. In addition, HIBEY only occupies 13.14% of the screen real estate at the edge region, which is 62.80% smaller than the default keyboard layout on Microsoft Hololens.
Lik-Hang Lee, Kit-Yung Lam, Yui-Pan Yau, Tristan Braud, Pan Hui 0001
PerCom1