Michael Xuelin Huang

dblp:86/11412 · DBLP profile ↗
← Back
28ranked-venue papers
11as first author
5since 2021 · last 2025
0000-0001-5695-2869ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 19 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Tap&Say: Touch Location-Informed Large Language Model for Multimodal Text Correction on Smartphones
abstract
layer that integrates the tap location into the LLM's attention mechanism, enabling it to utilize the tap location for text correction. We fine-tuned the touch location-informed LLM on synthetic touch locations and correction commands, achieving significantly higher correction accuracy than the state-of-the-art method VT [45]. A 16-person user study demonstrated that Tap&Say outperforms VT [45] with 16.4% shorter task completion time and 47.5% fewer keyboard clicks and is preferred by users.
Maozheng Zhao, Michael Xuelin Huang, Nathan G. Huang, Shanqing Cai, Henry Huang, Michael G. Huang, Shumin Zhai, I. V. Ramakrishnan, Xiaojun Bi 0001
CHI2
2024 Rambler: Supporting Writing With Speech via LLM-Assisted Gist Manipulation
abstract
Dictation enables efficient text input on mobile devices. However, writing with speech can produce disfluent, wordy, and incoherent text and thus requires heavy post-processing. This paper presents Rambler, an LLM-powered graphical user interface that supports gist-level manipulation of dictated text with two main sets of functions: gist extraction and macro revision. Gist extraction generates keywords and summaries as anchors to support the review and interaction with spoken text. LLM-assisted macro revisions allow users to respeak, split, merge, and transform dictated text without specifying precise editing locations. Together they pave the way for interactive dictation and revision that help close gaps between spontaneously spoken words and well-structured writing. In a comparative study with 12 participants performing verbal composition tasks, Rambler outperformed the baseline of a speech-to-text editor + ChatGPT, as it better facilitates iterative revisions with enhanced user control over the content while supporting surprisingly diverse user strategies.
Susan Lin, Jeremy Warner, J. D. Zamfirescu-Pereira, Matthew G. Lee, Sauhard Jain, Shanqing Cai, Piyawat Lertvittayakumjorn, Michael Xuelin Huang, Shumin Zhai, Björn Hartmann, Can Liu 0003
CHI8
2023 Real-time flashover prediction model for multi-compartment building structures using attention based recurrent neural networks
Wai Cheong Tam, Eugene Yujun Fu, Richard Peacock, Paul A. Reneke, Grace Ngai, Hong Va Leong, Thomas Cleary, Michael Xuelin Huang
Expert Syst. Appl.9
2022 A spatial temporal graph neural network model for predicting flashover in arbitrary building floorplans
Wai Cheong Tam, Eugene Yujun Fu, Michael Xuelin Huang
Eng. Appl. Artif. Intell.6
2021 TapNet: The Design, Training, Implementation, and Applications of a Multi-Task Learning CNN for Off-Screen Mobile Input
abstract
To make off-screen interaction without specialized hardware practical, we investigate using deep learning methods to process the common built-in IMU sensor (accelerometers and gyroscopes) on mobile phones into a useful set of one-handed interaction events. We present the design, training, implementation and applications of TapNet, a multi-task network that detects tapping on the smartphone. With phone form factor as auxiliary information, TapNet can jointly learn from data across devices and simultaneously recognize multiple tap properties, including tap direction and tap location. We developed two datasets consisting of over 135K training samples, 38K testing samples, and 32 participants in total. Experimental evaluation demonstrated the effectiveness of the TapNet design and its significant improvement over the state of the art. Along with the datasets, codebase1, and extensive experiments, TapNet establishes a new technical foundation for off-screen mobile input.
Michael Xuelin Huang, Yang Li 0058, Nazneen Nazneen, Alexander Chao, Shumin Zhai
CHI1
2020 Designing and Evaluating Head-based Pointing on Smartphones for People with Motor Impairments
abstract
Head-based pointing is an alternative input method for people with motor impairments to access computing devices. This paper proposes a calibration-free head-tracking input mechanism for mobile devices that makes use of the front-facing camera that is standard on most devices. To evaluate our design, we performed two Fitts’ Law studies. First, a comparison study of our method with an existing head-based pointing solution, Eva Facial Mouse, with subjects without motor impairments. Second, we conducted what we believe is the first Fitts’ Law study using a mobile head tracker with subjects with motor impairments. We extend prior studies with a greater range of index of difficulties (IDs) [1.62, 5.20] bits and achieved promising throughput (average 0.61 bps with motor impairments and 0.90 bps without). We found that users’ throughput was 0.95 bps on average in our most difficult task (IDs: 5.20 bits), which involved selecting a target half the size of the Android recommendation for a touch target after moving nearly the full height of the screen. This suggests the system is capable of fine precision tasks. We summarize our observations and the lessons from our user studies into a set of design guidelines for head-based pointing systems.
Muratcan Cicek, Ankit Dave, Wenxin Feng 0001, Michael Xuelin Huang, Julia Katherine Haines, Jeffrey Nichols 0001
ASSETS4
2020 WATouCH: Enabling Direct Input on Non-touchscreen Using Smartwatch's Photoplethysmogram and IMU Sensor Fusion
abstract
Interacting with non-touchscreens such as TV or public displays can be difficult and inefficient. We propose WATouCH, a novel method that localizes a smartwatch on a display and allows direct input by turning the smartwatch into a tangible controller. This low-cost solution leverages sensor fusion of the built-in inertial measurement unit (IMU) and photoplethysmogram (PPG) sensor on a smartwatch that is used for heart rate monitoring. Specifically, WATouCH tracks the smartwatch movement using IMU data and corrects its location error caused by drift using the PPG responses to a dynamic visual pattern on the display. We conducted a user study on two tasks -- a point and click and line tracing task -- to evaluate the system usability and user performance. Evaluation results suggested that our sensor fusion mechanism effectively confined IMU-based localization error, achieved encouraging targeting and tracing precision, was well received by the participants, and thus opens up new opportunities for interaction.
Hui-Shyong Yeo, Wenxin Feng 0001, Michael Xuelin Huang
CHI3
2019 SacCalib: reducing calibration distortion for stationary eye trackers using saccadic eye movements
abstract
Recent methods to automatically calibrate stationary eye trackers were shown to effectively reduce inherent calibration distortion. However, these methods require additional information, such as mouse clicks or on-screen content. We propose the first method that only requires users' eye movements to reduce calibration distortion in the background while users naturally look at an interface. Our method exploits that calibration distortion makes straight saccade trajectories appear curved between the saccadic start and end points. We show that this curving effect is systematic and the result of distorted gaze projection plane. To mitigate calibration distortion, our method undistorts this plane by straightening saccade trajectories using image warping. We show that this approach improves over the common six-point calibration and is promising for reducing distortion. As such, it provides a non-intrusive solution to alleviating accuracy decrease of eye tracker during long-term use.
Michael Xuelin Huang, Andreas Bulling
ETRA1
2019 Reducing calibration drift in mobile eye trackers by exploiting mobile phone usage
abstract
Automatic saliency-based recalibration is promising for addressing calibration drift in mobile eye trackers but existing bottom-up saliency methods neglect user's goal-directed visual attention in natural behaviour. By inspecting real-life recordings of egocentric eye tracker cameras, we reveal that users are likely to look at their phones once these appear in view. We propose two novel automatic recalibration methods that exploit mobile phone usage: The first builds saliency maps using the phone location in the egocentric view to identify likely gaze locations. The second uses the occurrence of touch events to recalibrate the eye tracker, thereby enabling privacy-preserving recalibration. Through in-depth evaluations on a recent mobile eye tracking dataset (N=17, 65 hours) we show that our approaches outperform a state-of-the-art saliency approach for automatic recalibration. As such, our approach improves mobile eye tracking and gaze-based interaction, particularly for long-term use.
Philipp Müller 0001, Daniel Buschek, Michael Xuelin Huang, Andreas Bulling
ETRA3
2019 Privacy-aware eye tracking using differential privacy
abstract
With eye tracking being increasingly integrated into virtual and augmented reality (VR/AR) head-mounted displays, preserving users' privacy is an ever more important, yet under-explored, topic in the eye tracking community. We report a large-scale online survey (N=124) on privacy aspects of eye tracking that provides the first comprehensive account of with whom, for which services, and to what extent users are willing to share their gaze data. Using these insights, we design a privacy-aware VR interface that uses differential privacy, which we evaluate on a new 20-participant dataset for two privacy sensitive tasks: We show that our method can prevent user re-identification and protect gender information while maintaining high performance for gaze-based document type classification. Our results highlight the privacy challenges particular to gaze data and demonstrate that differential privacy is a potential means to address them. Thus, this paper lays important foundations for future research on privacy-aware gaze interfaces.
Julian Steil, Inken Hagestedt, Michael Xuelin Huang, Andreas Bulling
ETRA3
2019 Moment-to-Moment Detection of Internal Thought during Video Viewing from Eye Vergence Behavior
abstract
Internal thought refers to the process of directing attention away from a primary visual task to internal cognitive processing. It is pervasive and closely related to primary task performance. As such, automatic detection of internal thought has significant potential for user modeling in human-computer interaction and multimedia applications. Despite the close link between the eyes and the human mind, only few studies have investigated vergence behavior during internal thought and none has studied moment-to-moment detection of internal thought from gaze. While prior studies relied on long-term data analysis and required a large number of gaze characteristics, we describe a novel method that is user-independent, computationally light-weight and only requires eye vergence information readily available from binocular eye trackers. We further propose a novel paradigm to obtain ground truth internal thought annotations by exploiting human blur perception. We evaluated our method during natural viewing of lecture videos and achieved a 12.1% improvement over the state of the art. These results demonstrate the effectiveness and robustness of vergence-based detection of internal thought and, as such, open new research directions for attention-aware interfaces.
Michael Xuelin Huang, Grace Ngai, Hong Va Leong, Andreas Bulling
ACM Multimedia1
2019 Towards High-Frequency SSVEP-Based Target Discrimination with an Extended Alphanumeric Keyboard
abstract
Despite significant advances in using Steady-State Visually Evoked Potentials (SSVEP) for on-screen target discrimination, existing methods either require intrusive, low-frequency visual stimulation or only support a small number of targets. We propose SSVEPNet: a convolutional long short-term memory (LSTM) recurrent neural network for high-frequency stimulation (≥30Hz) using a large number of visual targets. We evaluate our method for discriminating between 43 targets on an extended alphanumeric virtual keyboard and compare three different frequency assignment strategies. Our experimental results show that SSVEPNet significantly outperforms state-of-the-art correlation-based methods and convolutional neural networks. As such, our work opens up an exciting new direction of research towards a new class of unobtrusive and highly expressive SSVEP-based interfaces for text entry and beyond.
Sahar Abdelnabi, Michael Xuelin Huang, Andreas Bulling
SMC2
2018 Training Person-Specific Gaze Estimators from User Interactions with Multiple Devices
abstract
Learning-based gaze estimation has significant potential to enable attentive user interfaces and gaze-based interaction on the billions of camera-equipped handheld devices and ambient displays. While training accurate person- and device-independent gaze estimators remains challenging, person-specific training is feasible but requires tedious data collection for each target device. To address these limitations, we present the first method to train person-specific gaze estimators across multiple devices. At the core of our method is a single convolutional neural network with shared feature extraction layers and device-specific branches that we train from face images and corresponding on-screen gaze locations. Detailed evaluations on a new dataset of interactions with five common devices (mobile phone, tablet, laptop, desktop computer, smart TV) and three common applications (mobile game, text editing, media center) demonstrate the significant potential of cross-device training. We further explore training with gaze locations derived from natural interactions, such as mouse or touch input.
Xucong Zhang, Michael Xuelin Huang, Yusuke Sugano, Andreas Bulling
CHI2
2018 Robust eye contact detection in natural multi-person interactions using gaze and speaking behaviour
abstract
Eye contact is one of the most important non-verbal social cues and fundamental to human interactions. However, detecting eye contact without specialised eye tracking equipment poses significant challenges, particularly for multiple people in real-world settings. We present a novel method to robustly detect eye contact in natural three- and four-person interactions using off-the-shelf ambient cameras. Our method exploits that, during conversations, people tend to look at the person who is currently speaking. Harnessing the correlation between people's gaze and speaking behaviour therefore allows our method to automatically acquire training data during deployment and adaptively train eye contact detectors for each target user. We empirically evaluate the performance of our method on a recent dataset of natural group interactions and demonstrate that it achieves a relative improvement over the state-of-the-art method of more than 60%, and also improves over a head pose based baseline.
Philipp Müller 0001, Michael Xuelin Huang, Xucong Zhang, Andreas Bulling
ETRA2
2018 Fixation detection for head-mounted eye tracking based on visual similarity of gaze targets
abstract
Fixations are widely analysed in human vision, gaze-based interaction, and experimental psychology research. However, robust fixation detection in mobile settings is profoundly challenging given the prevalence of user and gaze target motion. These movements feign a shift in gaze estimates in the frame of reference defined by the eye tracker's scene camera. To address this challenge, we present a novel fixation detection method for head-mounted eye trackers. Our method exploits that, independent of user or gaze target motion, target appearance remains about the same during a fixation. It extracts image information from small regions around the current gaze position and analyses the appearance similarity of these gaze patches across video frames to detect fixations. We evaluate our method using fine-grained fixation annotations on a five-participant indoor dataset (MPIIEgoFixation) with more than 2,300 fixations in total. Our method outperforms commonly used velocity- and dispersion-based algorithms, which highlights its significant potential to analyse scene image information for eye movement detection.
Julian Steil, Michael Xuelin Huang, Andreas Bulling
ETRA2
2018 Every Little Movement Has a Meaning of Its Own: Using Past Mouse Movements to Predict the Next Interaction
abstract
User experience could be enhanced if the computer could understand human interaction intention. For instance, it could react to intercept and prevent interaction errors. This paper presents an approach to predicting users intention in interaction tasks based on past mouse movements. We adopt a long short-term memory (LSTM) model to predict the users» intention via their next mouse click interaction, upon being trained with past mouse interaction behaviors. To evaluate, we consider two scenarios in daily computer usage: a more structured crowdsourcing annotation task and a more free-form, open-ended web search task. Our results indicate that we could predict the next interaction event with reasonable accuracy. We also conducted a pilot study to investigate the possibility of applying our model for non-intentional mouse click detection. We believe that our findings would be beneficial towards the development of better intelligent agents.
Tiffany C. K. Kwok, Eugene Yujun Fu, Erin You Wu, Michael Xuelin Huang, Grace Ngai, Hong Va Leong
IUI4
2018 Detecting Low Rapport During Natural Interactions in Small Groups from Non-Verbal Behaviour
abstract
Rapport, the close and harmonious relationship in which interaction partners are "in sync" with each other, was shown to result in smoother social interactions, improved collaboration, and improved interpersonal outcomes. In this work, we are first to investigate automatic prediction of low rapport during natural interactions within small groups. This task is challenging given that rapport only manifests in subtle non-verbal signals that are, in addition, subject to influences of group dynamics as well as inter-personal idiosyncrasies. We record videos of unscripted discussions of three to four people using a multi-view camera system and microphones. We analyse a rich set of non-verbal signals for rapport detection, namely facial expressions, hand motion, gaze, speaker turns, and speech prosody. Using facial features, we can detect low rapport with an average precision of 0.7 (chance level at 0.25), while incorporating prior knowledge of participants' personalities can even achieve early prediction without a drop in performance. We further provide a detailed analysis of different feature sets and the amount of information contained in different temporal segments of the interactions.
Philipp Müller 0001, Michael Xuelin Huang, Andreas Bulling
IUI2
2018 Cross-Species Learning: A Low-Cost Approach to Learning Human Fight from Animal Fight
abstract
Detecting human fight behavior from videos is important in social signal processing, especially in the context of surveillance. However, the uncommon occurrence of real human fight events generally restricts the data collection for fight detection in machine learning, and thus hampers the performance of contemporary data-driven approaches. To address this challenge, we present a novel cross-species learning method with a set of low-computational cost motion features for fight detection. It effectively circumvents the problem of limited human fight data for data-demaining approaches. Our method exploits the intrinsic commonality between human and animal fights, such as the physical acceleration of moving body parts. It also leverages an ensemble learning mechanism to adapt useful knowledge from similar source subsets across species. Our evaluation results demonstrate the effectiveness of the proposed feature representation for cross-species adaptation. We believe that cross-species learning is not only a promising solution to the data constraint issue, but it also sheds lights on the studies of other human mental and social behaviors in cross-disciplinary research.
Eugene Yujun Fu, Michael Xuelin Huang, Hong Va Leong, Grace Ngai
ACM Multimedia2
2018 Quick Bootstrapping of a Personalized Gaze Model from Real-Use Interactions
abstract
Understanding human visual attention is essential for understanding human cognition, which in turn benefits human--computer interaction. Recent work has demonstrated a Personalized, Auto-Calibrating Eye-tracking (PACE) system, which makes it possible to achieve accurate gaze estimation using only an off-the-shelf webcam by identifying and collecting data implicitly from user interaction events. However, this method is constrained by the need for large amounts of well-annotated data. We thus present fast-PACE, an adaptation to PACE that exploits knowledge from existing data from different users to accelerate the learning speed of the personalized model. The result is an adaptive, data-driven approach that continuously “learns” its user and recalibrates, adapts, and improves with additional usage by a user. Experimental evaluations of fast-PACE demonstrate its competitive accuracy in iris localization, validity of alignment identification between gaze and interactions, and effectiveness of gaze transfer. In general, fast-PACE achieves an initial visual error of 3.98 degrees and then steadily improves to 2.52 degrees given incremental interaction-informed data. Our performance is comparable to state-of-the-art, but without the need for explicit training or calibration. Our technique addresses the data quality and quantity problems. It therefore has the potential to enable comprehensive gaze-aware applications in the wild.
Michael Xuelin Huang, Grace Ngai, Hong Va Leong
ACM Trans. Intell. Syst. Technol.1
2018 Fast-PADMA: Rapidly Adapting Facial Affect Model From Similar Individuals
abstract
A user-specific model generally performs better in facial affect recognition. Existing solutions, however, have usability issues since the annotation can be long and tedious for the end users (e.g., consumers). We address this critical issue by presenting a more user-friendly user-adaptive model to make the personalized approach more practical. This paper proposes a novel user-adaptive model, which we have called fast-Personal Affect Detection with Minimal Annotation (Fast-PADMA). Fast-PADMA integrates data from multiple source subjects with a small amount of data from the target subject. Collecting this target subject data is feasible since fast-PADMA requires only one self-reported affect annotation per facial video segment. To alleviate overfitting in this context of limited individual training data, we propose an efficient bootstrapping technique, which strengthens the contribution of multiple similar source subjects. Specifically, we employ an ensemble classifier to construct pretrained weak generic classifiers from data of multiple source subjects, which is weighted according to the available data from the target user. The result is a model that does not require expensive computation, such as distribution dissimilarity calculation or model retraining. We evaluate our method with in-depth experimental evaluations on five publicly available facial datasets, with results that compare favorably with the state-of-the-art performance on classifying pain, arousal, and valence. Our findings show that fast-PADMA is effective at rapidly constructing a user-adaptive model that outperforms both its generic and user-specific counterparts. This efficient technique has the potential to significantly improve user-adaptive facial affect recognition for personal use and, therefore, enable comprehensive affect-aware applications.
Michael Xuelin Huang, Grace Ngai, Hong Va Leong, Kien A. Hua
IEEE Trans. Multim.1
2017 Are you stressed? Your eyes and the mouse can tell
abstract
Stress is a fact of daily life. Stress can also deteriorate human's attention and memory, which, when a user is engaged in interactive applications, will negatively affect the user experience and downgrade the delivered performance. Traditional stress inference is mainly based on user physical features like Blood Volume Pulse, Galvanic Skin Response, often captured via devices that intrude on the user space. In contrast, this paper proposes a non-intrusive approach that exploits the consistency of users' behavioral patterns when interacting with a user interface, specifically, in terms of eye gaze and mouse movement. The relationship between the stress experienced by the user and his/her eye gaze and gaze-mouse coordination patterns are investigated. We show that both eye gaze and gaze-mouse coordination patterns can be exploited to distinguish whether a user is under stress. We also discover that a user's eye gaze behavior patterns are more consistent when he/she is under stress. This understanding of how a user's behavior differs under stress could be useful in the development of effective adaptive systems that can maximize user potential.
Jun Wang 0136, Michael Xuelin Huang, Grace Ngai, Hong Va Leong
ACII2
2017 ScreenGlint: Practical, In-situ Gaze Estimation on Smartphones
abstract
Gaze estimation has widespread applications. However, little work has explored gaze estimation on smartphones, even though they are fast becoming ubiquitous. This paper presents ScreenGlint, a novel approach which exploits the glint (reflection) of the screen on the user's cornea for gaze estimation, using only the image captured by the front-facing camera. We first conduct a user study on common postures during smartphone use. We then design an experiment to evaluate the accuracy of ScreenGlint under varying face-to-screen distances. An in-depth evaluation involving multiple users is conducted and the impact of head pose variations is investigated. ScreenGlint achieves an overall angular error of 2.44º without head pose variations, and 2.94º with head pose variations. Our technique compares favorably to state-of-the-art research works, indicating that the glint of the screen is an effective and practical cue to gaze estimation on the smartphone platform. We believe that this work can open up new possibilities for practical and ubiquitous gaze-aware applications.
Michael Xuelin Huang, Grace Ngai, Hong Va Leong
CHI1
2016 Building a Personalized, Auto-Calibrating Eye Tracker from User Interactions
abstract
We present PACE, a Personalized, Automatically Calibrating Eye-tracking system that identifies and collects data unobtrusively from user interaction events on standard computing systems without the need for specialized equipment. PACE relies on eye/facial analysis of webcam data based on a set of robust geometric gaze features and a two-layer data validation mechanism to identify good training samples from daily interaction data. The design of the system is founded on an in-depth investigation of the relationship between gaze patterns and interaction cues, and takes into consideration user preferences and habits. The result is an adaptive, data-driven approach that continuously recalibrates, adapts and improves with additional use. Quantitative evaluation on 31 subjects across different interaction behaviors shows that training instances identified by the PACE data collection have higher gaze point-interaction cue consistency than those identified by conventional approaches. An in-situ study using real-life tasks on a diverse set of interactive applications demonstrates that the PACE gaze estimation achieves an average error of 2.56º, which is comparable to state-of-the-art, but without the need for explicit training or calibration. This demonstrates the effectiveness of both the gaze estimation method and the corresponding data collection mechanism.
Michael Xuelin Huang, Tiffany C. K. Kwok, Grace Ngai, Stephen Chi-fai Chan, Hong Va Leong
CHI1
2016 StressClick: Sensing Stress from Gaze-Click Patterns
abstract
Stress sensing is valuable in many applications, including online learning crowdsourcing and other daily human-computer interactions. Traditional affective computing techniques investigate affect inference based on different individual modalities, such as facial expression, vocal tones, and physiological signals or the aggregation of signals of these independent modalities, without explicitly exploiting their inter-connections. In contrast, this paper focuses on exploring the impact of mental stress on the coordination between two human nervous systems, the somatic and autonomic nervous systems. Specifically, we present the analysis of the subtle but indicative pattern of human gaze behaviors surrounding a mouse-click event, i.e. the gaze-click pattern. Our evaluation shows that mental stress affects the gaze-click pattern, and this influence has largely been ignored in previous work. This paper, therefore, further proposes a non-intrusive approach to inferring human stress level based on the gaze-click pattern, using only data collected from the common computer webcam and mouse. We conducted a human study on solving math questions under different stress levels to explore the validity of stress recognition based on this coordination pattern. Experimental results show the effectiveness of our technique and the generalizability of the proposed features for user-independent modeling. Our results suggest that it may be possible to detect stress non-intrusively in the wild, without the need for specialized equipment.
Michael Xuelin Huang, Grace Ngai, Hong Va Leong
ACM Multimedia1
2016 Identifying User-Specific Facial Affects from Spontaneous Expressions with Minimal Annotation
abstract
This paper presents Personalized Affect Detection with Minimal Annotation (PADMA), a user-dependent approach for identifying affective states from spontaneous facial expressions without the need for expert annotation. The conventional approach relies on the use of key frames in recorded affect sequences and requires an expert observer to identify and annotate the frames. It is susceptible to user variability and accommodating individual differences is difficult. The alternative is a user-dependent approach, but it would be prohibitively expensive to collect and annotate data for each user. PADMA uses a novel Association-based Multiple Instance Learning (AMIL) method, which learns a personal facial affect model through expression frequency analysis, and does not need expert input or frame-based annotation. PADMA involves a training/calibration phase in which the user watches short video segments and reports the affect that best describes his/her overall feeling throughout the segment. The most indicative facial gestures are identified and extracted from the facial response video, and the association between gesture and affect labels is determined by the distribution of the gesture over all reported affects. Hence both the geometric deformation and distribution of key facial gestures are specially adapted for each user. We show results that demonstrate the feasibility, effectiveness and extensibility of our approach.
Michael Xuelin Huang, Grace Ngai, Kien A. Hua, Stephen Chi-fai Chan, Hong Va Leong
IEEE Trans. Affect. Comput.1
2015 Emotar: Communicating Feelings through Video Sharing
abstract
Affect exchange is essential for healthy physical and social development [7], and friends and family communicate their emotions to each other instinctively. In particular, watching movies has always been a popular mode of socialization and video sharing is increasingly viewed as an effective way to facilitate communication of feelings and affects, even when the parties are not in the same location. We present an asynchronous video-sharing platform that uses Emotars to facilitate affect sharing in order to create and enhance the sense of togetherness through the experience of asynchronous movie watching. We investigate its potential impact and benefits, including a better viewing experience, supporting relationships, and strengthening engagement, connectedness and emotion awareness among individuals.
Tiffany C. K. Kwok, Michael Xuelin Huang, Wai Cheong Tam, Grace Ngai
IUI2
2014 Building a Self-Learning Eye Gaze Model from User Interaction Data
abstract
Most eye gaze estimation systems rely on explicit calibration, which is inconvenient to the user, limits the amount of possible training data and consequently the performance. Since there is likely a strong correlation between gaze and interaction cues, such as cursor and caret locations, a supervised learning algorithm can learn the complex mapping between gaze features and the gaze point by training on incremental data collected implicitly from normal computer interactions. We develop a set of robust geometric gaze features and a corresponding data validation mechanism that identifies good training data from noisy interaction-informed data collected in real-use scenarios. Based on a study of gaze movement patterns, we apply behavior-informed validation to extract gaze features that correspond with the interaction cue, and data-driven validation provides another level of crosschecking using previous good data. Experimental evaluation shows that the proposed method achieves an average error of 4.06º, and demonstrates the effectiveness of the proposed gaze estimation method and corresponding validation mechanism.
Michael Xuelin Huang, Tiffany C. K. Kwok, Grace Ngai, Hong Va Leong, Stephen Chi-fai Chan
ACM Multimedia1
2012 MelodicBrush: a novel system for cross-modal digital art creation linking calligraphy and music
abstract
MelodicBrush is a novel system that connects two ancient art forms: Chinese ink-brush calligraphy and Chinese music. Our system uses vision-based techniques to create a digitized ink-brush calligraphic writing surface with enhanced interaction functionalities. The music generation combines cross-modal stroke-note mapping and statistical language modeling techniques into a hybrid model that generates music as a real-time, auditory response and feedback to the user's calligraphic strokes.
Michael Xuelin Huang, Will W. W. Tang, Kenneth W. K. Lo, Chi Kin Lau, Grace Ngai, Stephen Chi-fai Chan
Conference on Designing Interactive Systems1