Eugene Yujun Fu

dblp:151/8780 · DBLP profile ↗
← Back
27ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0003-1048-1904ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Prior-Informed Normative EEG Scoring for Alzheimer's Detection: Framework and Cross-Paradigm Boundary Analysis
Harihar Rengan, Ali-Mansur Valiyev, Ryan Borui Zhu, Chi-Wai Do, Eugene Yujun Fu
COMPSAC5
2025 Is Your Autonomous Vehicle Safe? Understanding the Threat of Electromagnetic Signal Injection Attacks on Traffic Scene Perception
abstract
Autonomous vehicles rely on camera-based perception systems to comprehend their driving environment and make crucial decisions, thereby ensuring vehicles to steer safely. However, a significant threat known as Electromagnetic Signal Injection Attacks (ESIA) can distort the images captured by these cameras, leading to incorrect AI decisions and potentially compromising the safety of autonomous vehicles. Despite the serious implications of ESIA, there is limited understanding of its impacts on the robustness of AI models across various and complex driving scenarios. To address this gap, our research analyzes the performance of different models under ESIA, revealing their vulnerabilities to the attacks. Moreover, due to the challenges in obtaining real-world attack data, we develop a novel ESIA simulation method and generate a simulated attack dataset for different driving scenarios. Our research provides a comprehensive simulation and evaluation framework, aiming to enhance the development of more robust AI models and secure intelligent systems, ultimately contributing to the advancement of safer and more reliable technology across various fields.
Wenhao Liao, Sineng Yan, Youqian Zhang, Xinwei Zhai, Eugene Yujun Fu
AAAI6
2025 Combating Phone Scams with LLM-based Detection: Where Do We Stand? (Student Abstract)
abstract
Phone scams pose a significant threat to individuals and communities, causing substantial financial losses and emotional distress. Despite ongoing efforts to combat these scams, scammers continue to adapt and refine their tactics, making it imperative to explore innovative countermeasures. This research explores the potential of large language models (LLMs) to provide detection of fraudulent phone calls. By analyzing the conversational dynamics between scammers and victims, LLM-based detectors can identify potential scams as they occur, offering immediate protection to users. While such approaches demonstrate promising results, we also acknowledge the challenges of biased datasets, relatively low recall, and hallucinations that must be addressed for further advancement in this field.
Zitong Shen, Kangzhong Wang, Youqian Zhang, Grace Ngai, Eugene Yujun Fu
AAAI5
2025 One Size Fits All? A Modular Adaptive Sanitization Kit (MASK) for Customizable Privacy-Preserving Phone Scam Detection
abstract
Phone scams remain a pervasive threat to both personal safety and financial security worldwide. Recent advances in large language models (LLMs) have demonstrated strong potential in detecting fraudulent behavior by analyzing transcribed phone conversations. However, these capabilities introduce notable privacy risks, as such conversations frequently contain sensitive personal information that may be exposed to third-party service providers during processing. In this work, we explore how to harness LLMs for phone scam detection while preserving user privacy. We propose MASK (Modular Adaptive Sanitization Kit), a trainable and extensible framework that enables dynamic privacy adjustment based on individual preferences. MASK provides a pluggable architecture that accommodates diverse sanitization methods-from traditional keyword-based techniques for high-privacy users to sophisticated neural approaches for those prioritizing accuracy. We also discuss potential modeling approaches and loss function designs for future development, enabling the creation of truly personalized, privacy-aware LLM-based detection systems that balance user trust and detection effectiveness, even beyond phone scam context.
Kangzhong Wang, Zitong Shen, Youqian Zhang, MK Michael Cheung, Xiapu Luo, Grace Ngai, Eugene Yujun Fu
ACM Multimedia7
2025 Anti-ESIA: Analyzing and Mitigating Impacts of Electromagnetic Signal Injection Attacks on Image Sensing
Denglin Kang, Youqian Zhang, Wai Cheong Tam, Xiapu Luo, Eugene Yujun Fu
MoMM5
2025 TMAN: A temporal multimodal attention network for backchannel detection
Kangzhong Wang, Xinwei Zhai, MK Michael Cheung, Eugene Yujun Fu, Peter Q. Chen, Grace Ngai, Hong Va Leong
Neurocomputing4
2025 Personalized e-learning resource recommendation using multimodal-enhanced collaborative filtering
Xinwei Zhai, Luwen Liang, Kangzhong Wang, Fengchun Pei, Eugene Yujun Fu
Knowl. Based Syst.6
2024 Understanding Impacts of Electromagnetic Signal Injection Attacks on Object Detection
abstract
Object detection can localize and identify objects in images, and it is extensively employed in critical multimedia applications such as security surveillance and autonomous driving. Despite the success of existing object detection models, they are often evaluated in ideal scenarios where captured images guarantee the accurate and complete representation of the detecting scenes. However, images captured by image sensors may be affected by different factors in real applications, including cyber-physical attacks. In particular, attackers can exploit hardware properties within the systems to inject electromagnetic interference so as to manipulate the images. Such attacks can cause noisy or incomplete information about the captured scene, leading to incorrect detection results, potentially granting attackers malicious control over critical functions of the systems. This paper presents a research work that comprehensively quantifies and analyzes the impacts of such attacks on state-of-the-art object detection models in practice. It also sheds light on the underlying reasons for the incorrect detection outcomes.
Youqian Zhang, Eugene Yujun Fu, Qinhong Jiang, Chen Yan 0001, Sze-Yiu Chau, Grace Ngai, Hong Va Leong, Xiapu Luo, Wenyuan Xu 0001
ICME3
2023 Unveiling Subtle Cues: Backchannel Detection Using Temporal Multimodal Attention Networks
abstract
Automatic detection of backchannel has great potential to enhance artificial mediators, which indicate listeners' attention and agreement in human communication. It is often expressed by subtle non-verbal cues that occur briefly and sparsely. Focusing on identifying and locating these subtle cues (i.e., their occurrence moment and the involved body parts), this paper proposes a novel approach for backchannel detection. In particular, our model utilizes temporal- and modality-attention modules to determine and lead the model to pay more attention to both the indicative moment and the accompanying body parts at that specific time. It achieves an accuracy of 68.6% on the testing set in MultiMediate'23 backchannel detection challenge, outperforming the counterparts. Furthermore, we conducted an ablation study to thoroughly understand the contributions of our model. This study underscores the effectiveness of our selection of modality inputs and the importance of the two attention modules in our model.
Kangzhong Wang, MK Michael Cheung, Youqian Zhang, Peter Q. Chen, Eugene Yujun Fu, Grace Ngai
ACM Multimedia6
2023 MultiMediate 2023: Engagement Level Detection using Audio and Video Features
abstract
Real-time engagement estimation holds significant potential across various research areas, particularly in the realm of human-computer interaction. It empowers artificial agents to dynamically adjust their responses based on user engagement levels, fostering more intuitive and immersive interactions. Despite the strides in automating real-time engagement estimation, the task remains challenging in real-world settings, especially when handling multi-modal human social signals. Capitalizing on human body and audio signals, this paper explores the appropriate feature representations of different modalities and effective modelling of dual conversations. This results in a novel and efficient multi-modal engagement detection model.We thoroughly evaluated our method in the MultiMediate'23 grand challenge. It performs consistently, with a notable improvement over the baseline model. Specifically, while the baseline achieves a concordance correlation coefficient (CCC) of 0.59, our approach yields a CCC of 0.70, suggesting its promising efficacy in real-life engagement detection.
Kangzhong Wang, Peter Q. Chen, MK Michael Cheung, Youqian Zhang, Eugene Yujun Fu, Grace Ngai
ACM Multimedia6
2023 Is your mouse attracted by your eyes: Non-intrusive stress detection in off-the-shelf desktop environments
Jun Wang 0136, Eugene Yujun Fu, Grace Ngai, Hong Va Leong
Eng. Appl. Artif. Intell.3
2023 Real-time flashover prediction model for multi-compartment building structures using attention based recurrent neural networks
Wai Cheong Tam, Eugene Yujun Fu, Richard Peacock, Paul A. Reneke, Grace Ngai, Hong Va Leong, Thomas Cleary, Michael Xuelin Huang
Expert Syst. Appl.2
2022 Identifying Key Learning Factors in Service-Leaning Programs Using Machine Learning
abstract
As an impactful experiential learning pedagogy in higher education, service-learning (SL) can enhance students' academic learning and their sense of community and social responsibility by involving them in comprehensive community services. Much extant literature has justified the positive impacts of SL. However, the lack of quantitative analysis on identifying significant learning and course factors that strongly impact students' SL outcomes limits SL's further enhancement and adaptive development. This paper proposes to use machine learning approaches for modeling and identifying key learning factors in SL. We collect and study a large-scale dataset, including students' feedback on learning factors related to the different student experiences, course elements, and self-perceived learning outcomes. Machine learning algorithms are applied to model the various learning factors, contributing to effective classification models that predict students' learning outcomes using their evaluation on the learning factors. The most predictive model is then selected to identify a key set of important variables most indicative to students' SL outcomes. Our experiment results show that learning factors related to study challenges and interactions have significant positive impacts on students' learning gains. We believe that this paper will benefit future studies in this field.
Kangzhong Wang, Eugene Yujun Fu, Grace Ngai, Hong Va Leong
COMPSAC2
2022 A spatial temporal graph neural network model for predicting flashover in arbitrary building floorplans
Wai Cheong Tam, Eugene Yujun Fu, Michael Xuelin Huang
Eng. Appl. Artif. Intell.2
2022 Investigating Differences in Gaze and Typing Behavior Across Writing Genres
abstract
Writing is one of the most common activities undertaken on a computer, and the activity of writing has been widely studied. Given that writing is an intensively cognitive process, it makes sense that the type of writing that is being produced would have an effect on the writer’s gaze and typing behaviors. However, only a few studies have explored this relationship. In this paper, we study the gaze-typing behaviors, specifically, the coordination between eye gaze and typing dynamics, of writers who are producing original articles in different genres: reminiscent, logical and creative. Our study focuses on Chinese typing, particularly via the Pinyin input method, which generates text via a two step method, and requires additional cognitive processes compared to typing in phonographic languages such as English. Our study involves 46 native Chinese speakers of varying ages from children to elderly. Our method deploys statistics- and sequence-based features to infer the mental state of the author during the writing process. The statistics-based features focus on modeling the overall gaze-typing behaviors during the process and the sequence-based features focus on the transition of the gaze-typing behaviors as the piece of writing progresses. Using a linear support-vector machine, we achieve an overall accuracy over 88% for the article-genre detection by using a leave-one-subject-out cross-validation evaluation.
Jun Wang 0136, Eugene Yujun Fu, Grace Ngai, Hong Va Leong
Int. J. Hum. Comput. Interact.2
2021 Predicting Flashover Occurrence using Surrogate Temperature Data
abstract
Fire fighter fatalities and injuries in the U.S. remain too high and fire fighting too hazardous. Until now, fire fighters rely only on their experience to avoid life-threatening fire events, such as flashover. In this paper, we describe the development of a flashover prediction model which can be used to warn fire fighters before flashover occurs. Specifically, we consider the use of a fire simulation program to generate a set of synthetic data and an attention-based bidirectional long short-term memory to learn the complex relationships between temperature signals and flashover conditions. We first validate the fire simulation program with temperature measurements obtained from full-scale fire experiments. Then, we generate a set of synthetic temperature data which account for the realis-tic fire and vent opening conditions in a multi-compartment structure. Results show that our proposed method achieves promising performance for prediction of flashover even when temperature data is completely lost in the room of fire origin. It is believed that the flashover prediction model can facilitate the transformation of fire fighting tactics from traditional experience-based decision marking to data-driven decision marking and reduce fire fighter deaths and injuries.
Eugene Yujun Fu, Wai Cheong Tam, Jun Wang 0136, Richard Peacock, Paul A. Reneke, Grace Ngai, Hong Va Leong, Thomas Cleary
AAAI1
2021 Using Motion Histories for Eye Contact Detection in Multiperson Group Conversations
abstract
Eye contact detection in group conversations is the key to developing artificial mediators that can understand and interact with a group. In this paper, we propose to model a group's appearances and behavioral features to perform eye contact detection for each participant in the conversation. Specifically, we extract the participants' appearance features at the detection moment, and extract the participants' behavioral features based on their motion history image, which is encoded with the participants' body movements within a small time window before the detection moment. In order to attain powerful representative features from these images, we propose to train a Convolutional Neural Network (CNN) to model them. A set of relevant features are obtained from the network, which achieves an accuracy of 0.60 on the validation set in the eye contact detection challenge in ACM MM 2021. Furthermore, our experimental results also demonstrate that making use of both participants' appearance and behavior features can lead to higher accuracy at eye detection than only using one of them.
Eugene Yujun Fu, Michael W. Ngai
ACM Multimedia1
2020 Exploiting Active Learning in Novel Refractive Error Detection with Smartphones
abstract
Refractive errors, such as myopia and astigmatism, can lead to severe visual impairment if not detected and corrected in time. Traditional methods of refractive error diagnosis rely on well-trained optometrists operating expensive and importable devices, constraining the vision screening process. Advance in smartphone camera has enabled novel low-cost ubiquitous vision screening to detect refractive error or ametropia through eye image processing, based on the principle of photorefraction. However, contemporary smartphone-based methods rely heavily on hand-crafted features and sufficiency of well-labeled data. To address these challenges, this paper exploits active learning methods with a set of Convolutional Neural Network features encoding information of human eyes from pre-trained gaze estimation model. This enables more effective training on refractive error detection models with less labeled data. Our experimental results demonstrate the encouraging effectiveness of our active learning approach. The new set of features is able to attain screening accuracy of more than 80% with mean absolute error less than 0.66, meeting the expectation of optometrists for 0.5 to 1. The proposed active learning also requires significantly fewer training samples of 18% in achieving satisfactory performance.
Eugene Yujun Fu, Zhongqi Yang, Hong Va Leong, Grace Ngai, Chi-Wai Do, Lily Chan
ACM Multimedia1
2020 Screening for refractive error with low-quality smartphone images
abstract
Uncorrected refractive errors can lead to permanent debilitating eye conditions if not corrected in a timely manner. Contemporary diagnostic methods rely on the professional acumen of optometrists and the use of expensive devices, which may not be easily accessible to all. According to the optical principle of photorefraction, refractive error can be estimated based on a relative pupil and crescent size of an eye image taken by a camera from a specified working distance. A low-cost approach would be to leverage smartphones with cameras for this purpose. However, the poor image quality generated from basic smartphones poses a challenge for the current approach as they often fail to accurately distinguish the crescent from the iris. We propose a novel method to detect and accurately measure the iris and crescent from smartphone photos. Based on this method, we further propose a set of features for machine learning to build our refractive error estimation model. The performance of our models are evaluated in an in-depth experiment.
Zhongqi Yang, Eugene Yujun Fu, Grace Ngai, Hong Va Leong, Chi-Wai Do, Lily Chan
MoMM2
2019 Investigating Differences in Gaze and Typing Behavior Across Age Groups and Writing Genres
abstract
Typing is one of the most common activities that are undertaken on a computer. It would therefore be interesting to investigate whether it is possible to deduce characteristics of the user, such as their age or the type of the document that they are writing, just simply from typing dynamics. In this paper, we study the coordination between eye gaze and typing dynamics, or the gaze-typing behavior, of subjects who are producing original text. We focus upon the differences between different age groups (children vs elderly seniors) and different genres of writing (reminiscent, logical and creative). Using machine-learning, we achieve an accuracy of 93.5% for age detection and 61.1% for the article-category detection, using a leave-one-subject-out cross-validation evaluation, which is 44% and 28% higher than baselines.
Jun Wang 0136, Eugene Yujun Fu, Grace Ngai, Hong Va Leong
COMPSAC (1)2
2019 Your Body Signals Expose Your Fall
abstract
Fall is a common cause of severe injuries that may lead to irreversible body damage and even death. A real-time fall monitoring system can reveal a fall in time for timely medical aid to a victim. This is particularly important in the context of mobile healthcare. Fall detection with most contemporary wearable devices relied solely on acceleration signals, often not flexible and robust enough. In this paper, we propose to deploy body signals in a multi-modality approach. Besides the common acceleration signals, we also make use of physiological signals returned by wearable devices for multiple modalities. Fall detection would not fail easily even if some acceleration signals become ineffective. Our experiment results indicate that we are able to attain an accuracy of more than 96%. An in-depth evaluation demonstrates that physiological signals can contribute in distinguishing falls from actions generating similar acceleration signals, such as jumps, sit-downs and walking-downstairs.
Eugene Yujun Fu, Cheuk Yin Wong, Katie T. Y. Lau, Hong Va Leong, Grace Ngai
iiWAS1
2019 Activity Recognition and Stress Detection via Wristband
abstract
Advancement of micro-electromechanical systems enables easy daily activity and physiological data collection with a smart wristband and smartphone. Making use of those signals in various intelligent algorithm can contribute much to trending m-health applications. The ability of continuously monitoring physical activities and stress level can help users to better track their health condition. In this study, we propose to recognize different physical activities and detect long lasting stress level based on the 3-axis acceleration signals and physiological signals. We are able to achieve accuracy of around 97% for physical activities recognition and more than 80% for stress detection. We also discover that physiological signals alone cannot distinguish well between the high intensity activities and the stress condition.
Johnny Chun Yiu Wong, Jun Wang 0136, Eugene Yujun Fu, Hong Va Leong, Grace Ngai
MoMM3
2018 Every Little Movement Has a Meaning of Its Own: Using Past Mouse Movements to Predict the Next Interaction
abstract
User experience could be enhanced if the computer could understand human interaction intention. For instance, it could react to intercept and prevent interaction errors. This paper presents an approach to predicting users intention in interaction tasks based on past mouse movements. We adopt a long short-term memory (LSTM) model to predict the users» intention via their next mouse click interaction, upon being trained with past mouse interaction behaviors. To evaluate, we consider two scenarios in daily computer usage: a more structured crowdsourcing annotation task and a more free-form, open-ended web search task. Our results indicate that we could predict the next interaction event with reasonable accuracy. We also conducted a pilot study to investigate the possibility of applying our model for non-intentional mouse click detection. We believe that our findings would be beneficial towards the development of better intelligent agents.
Tiffany C. K. Kwok, Eugene Yujun Fu, Erin You Wu, Michael Xuelin Huang, Grace Ngai, Hong Va Leong
IUI2
2018 Cross-Species Learning: A Low-Cost Approach to Learning Human Fight from Animal Fight
abstract
Detecting human fight behavior from videos is important in social signal processing, especially in the context of surveillance. However, the uncommon occurrence of real human fight events generally restricts the data collection for fight detection in machine learning, and thus hampers the performance of contemporary data-driven approaches. To address this challenge, we present a novel cross-species learning method with a set of low-computational cost motion features for fight detection. It effectively circumvents the problem of limited human fight data for data-demaining approaches. Our method exploits the intrinsic commonality between human and animal fights, such as the physical acceleration of moving body parts. It also leverages an ensemble learning mechanism to adapt useful knowledge from similar source subsets across species. Our evaluation results demonstrate the effectiveness of the proposed feature representation for cross-species adaptation. We believe that cross-species learning is not only a promising solution to the data constraint issue, but it also sheds lights on the studies of other human mental and social behaviors in cross-disciplinary research.
Eugene Yujun Fu, Michael Xuelin Huang, Hong Va Leong, Grace Ngai
ACM Multimedia1
2017 Your Mouse Reveals Your Next Activity: Towards Predicting User Intention from Mouse Interaction
abstract
This paper presents an investigation into user intention prediction in two common web-based tasks: crowdsourcing annotation and web search, based on human-mouse interaction information. User experience is gaining importance within the research area of human-centered computing, and is particularly useful for complex, multi-step tasks. To enhance user experience, the computer should be intelligent enough to be able to predict the user intention. For instance, an intelligent agent might be able to anticipate when the user is about to press a button, and helpfully enlarge or highlight it in advance. In this paper, we propose two prediction models on user intention: a classical model that considers only historical mouse activity sequence, and a multimodal model that utilizes mouse interaction signals as well as features extracted from mouse trajectory and clicking events. We evaluate our models and find that they achieve reasonable accuracy. Our preliminary results indicate that we can dynamically learn a multimodal model that can effectively predict a user's next activity from historical activity sequence and mouse interaction signals.
Eugene Yujun Fu, Tiffany C. K. Kwok, Erin You Wu, Hong Va Leong, Grace Ngai, Stephen Chi-fai Chan
COMPSAC (1)1
2016 Automatic Fight Detection in Surveillance Videos
Eugene Yujun Fu, Hong Va Leong, Grace Ngai, Stephen Chi-fai Chan
MoMM1
2015 Automatic Fight Detection Based on Motion Analysis
abstract
Social signal processing is becoming an important topic in affective computing. In this paper, we focus on an important social interaction in real life, namely, fighting. Fight detection will be useful in public transportation, prisons, bars, or even sport. A robust mechanism in detecting fights from a video will be extremely useful, especially in applications relevant to surveillance systems. Recent research works focus on extracting visual features from high resolution video, leading to computationally expensive systems. In this paper, we propose an approach to detect fights in a natural and robust way based on motion analysis, which is not only intuitive, but also robust. Experimental results show that we can accurately detect fight activities in different video surveillance settings.
Eugene Yujun Fu, Hong Va Leong, Grace Ngai, Stephen Chi-fai Chan
ISM1