Beibin Li

dblp:176/9123 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0002-4426-3449ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 10 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
abstract
Yun He, Wenzhe Li, Hejia Zhang, Songlin Li, Karishma Mandyam, Sopan Khosla, Yuanhao Xiong, Nanshu Wang, Xiaoliang Peng, Beibin Li, Shengjie Bi, Shishir G Patil, Qi Qi, Shengyu Feng, Julian Katz-Samuels, Richard Yuanzhe Pang, Sujan Kumar Gonugondla, Hunter Lang, Yue Yu, Yundi Qian, Maryam Fazel-Zarandi, Licheng Yu, Amine Benhalloum, Hany Hassan Awadalla, Manaal Faruqui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Karishma Mandyam, Sopan Khosla, Yuanhao Xiong, Nanshu Wang, Xiaoliang Peng, Beibin Li, Shengjie Bi, Shishir G. Patil, Shengyu Feng, Julian Katz-Samuels, Richard Yuanzhe Pang, Sujan K. Gonugondla, Hunter Lang, Yue Yu 0009, Yundi Qian, Maryam Fazel-Zarandi, Licheng Yu, Amine Benhalloum, Hany Hassan, Manaal Faruqui
ACL (1)10
2025 Genetic Associations of Blink Behavior in Infants and Toddlers: A Computer Vision Approach
Yeji Bae, Beibin Li, Kelsey Jackson Dommer, Megan Reninger, Katie Hegerberg, Arya Ajwani, Terje Falck-Ytter, Frédérick Shic
ETRA2
2025 PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology
abstract
Diagnosing diseases through histopathology whole slide images (WSIs) is fundamental in modern pathology but is challenged by the gigapixel scale and complexity of WSIs. Trained histopathologists overcome this challenge by navigating the WSI, looking for relevant patches, taking notes, and compiling them to produce a final holistic diagnostic. Traditional AI approaches, such as multiple instance learning and transformer-based models, fail short of such a holistic, iterative, multi-scale diagnostic procedure, limiting their adoption in the real-world. We introduce PathFinder, a multi-modal, multi-agent framework that emulates the decision-making process of expert pathologists. PathFinder integrates four AI agents, the Triage Agent, Navigation Agent, Description Agent, and Diagnosis Agent, that collaboratively navigate WSIs, gather evidence, and provide comprehensive diagnoses with natural language explanations. The Triage Agent classifies the WSI as benign or risky; if risky, the Navigation and Description Agents iteratively focus on significant regions, generating importance maps and descriptive insights of sampled patches. Finally, the Diagnosis Agent synthesizes the findings to determine the patient's diagnostic classification. Our Experiments show that PathFinder outperforms state-of-the-art methods in skin melanoma diagnosis by 8% while offering inherent explainability through natural language descriptions of diagnostically relevant patches. Qualitative analysis by pathologists shows that the Description Agent's outputs are of high quality and comparable to GPT-4o. PathFinder is also the first AI-based system to surpass the average performance of pathologists in this challenging melanoma classification task by 9%, setting a new record for efficient, accurate, and interpretable AI-assisted diagnostics in pathology. Data, code and models available at https://pathfinder-dx.github.io/
Fatemeh Ghezloo, Mehmet Saygin Seyfioglu, Rustin Soraki, Wisdom Oluchi Ikezogwo, Beibin Li, Tejoram Vivekanandan, Joann G. Elmore, Ranjay Krishna, Linda G. Shapiro
ICCV5
2025 Towards Foundation Models for Mixed Integer Linear Programming
abstract
Mixed Integer Linear Programming (MILP) is essential for modeling complex decision-making problems but faces challenges in computational tractability and interpretability. Current deep learning approaches for MILP focus on specific problem classes and do not generalize to unseen classes. To address this shortcoming, we take a foundation model training approach, where we train a single deep learning model on a diverse set of MILP problems to generalize across problem classes. As existing datasets for MILP lack diversity and volume, we introduce MILP-Evolve, a novel LLM-based evolutionary framework that is capable of generating a large set of diverse MILP classes with an unlimited amount of instances. We study our methodology on three key learning tasks that capture diverse aspects of MILP: (1) integrality gap prediction, (2) learning to branch, and (3) a new task of aligning MILP instances with natural language descriptions. Our empirical results show that models trained on the data generated by MILP-Evolve achieve significant improvements on unseen problems, including MIPLIB benchmarks. Our work highlights the potential of moving towards a foundation model approach for MILP that can generalize to a broad range of MILP problem classes. Our code and data are publicly available at https://github.com/microsoft/OptiGuide.
Janardhan Kulkarni, Ishai Menache, Cathy Wu 0002, Beibin Li
ICLR5
2025 Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems
abstract
Failure attribution in LLM multi-agent systems—identifying the agent and step responsible for task failures—provides crucial clues for systems debugging but remains underexplored and labor-intensive. In this paper, we propose and formulate a new research area: automated failure attribution for LLM multi-agent systems. To support this initiative, we introduce the Who&When dataset, comprising extensive failure logs from 127 LLM multi-agent systems with fine-grained annotations linking failures to specific agents and decisive error steps. Using the Who&When, we develop and evaluate three automated failure attribution methods, summarizing their corresponding pros and cons. The best method achieves 53.5% accuracy in identifying failure-responsible agents but only 14.2% in pinpointing failure steps, with some methods performing below random. Even SOTA reasoning models, such as OpenAI o1 and DeepSeek R1, fail to achieve practical usability. These results highlight the task’s complexity and the need for further research in this area. Code and dataset are available in https://github.com/mingyin1/Agents_Failure_Attribution.
Ming Yin 0009, Jieyu Zhang 0001, Zhiguang Han, Jingyang Zhang, Beibin Li, Chi Wang 0001, Huazheng Wang, Yiran Chen 0001, Qingyun Wu
ICML7
2024 Reflect-RL: Two-Player Online RL Fine-Tuning for LMs
abstract
Environment Observation You are in the middle of a room.You see a cabinet 18, a cabinet 3, a ... Possible Actions 1. Close microwave 2. Put tomato in microwave 3. Go to cabinet ... Reflect Agent ReflectionTaking the tomato to the microwave is correct.Next step is to place the tomato ...
Runlong Zhou, Simon S. Du, Beibin Li
ACL (1)3
2024 Towards Safer Heuristics With XPlain
abstract
Many problems that cloud operators solve are computationally expensive, and operators often use heuristic algorithms (that are faster and scale better than optimal) to solve them more efficiently. Heuristic analyzers enable operators to find when and by how much their heuristics underperform. However, these tools do not provide enough detail for operators to mitigate the heuristic's impact in practice: they only discover a single input instance that causes the heuristic to underperform (and not the full set) and they do not explain why.
Pantea Karimi, Solal Pirelli, Siva Kesava Reddy K., Ryan Beckett, Santiago Segarra, Beibin Li, Pooria Namyar, Behnaz Arzani
HotNets6
2023 Comparing Attention to Biological Motion in Autism across Age Groups Using Eye-Tracking
abstract
This study tracked eye movement in children with and without autism spectrum disorder (ASD) watching emotional biological and non-biological motion point-light-displays (PLDs). Older children with ASD focused on extremities while older typically developing (TD) children looked at figure’s heads, whereas not evident in the younger groups. These results suggest developmental advances in social-information biases in TD children not evident in children with ASD, together with atypical and potentially adaptive increases in attentional biases towards local motion cues with age in ASD. Potential avenues for future computational and methodological analyses are discussed.
Michal Hochhauser, Kelsey Jackson Dommer, Adham Atyabi, Beibin Li, Yeojin A. Ahn, Madeline Aubertine, Minah Kim, Sarah Corrigan, Kevin A. Pelphrey, Frédérick Shic
ETRA4
2023 Kerveros: Efficient and Scalable Cloud Admission Control
Sultan Mahmud Sajal, Luke Marshall, Beibin Li, Shandan Zhou, Abhisek Pan, Konstantina Mellou, Deepak Narayanan, Timothy Zhu, David Dion, Thomas Moscibroda, Ishai Menache
OSDI3
2023 VSGD-Net: Virtual Staining Guided Melanocyte Detection on Histopathological Images
abstract
Detection of melanocytes serves as a critical prerequisite in assessing melanocytic growth patterns when diagnosing melanoma and its precursor lesions on skin biopsy specimens. However, this detection is challenging due to the visual similarity of melanocytes to other cells in routine Hematoxylin and Eosin (H&E) stained images, leading to the failure of current nuclei detection methods. Stains such as Sox10 can mark melanocytes, but they require an additional step and expense and thus are not regularly used in clinical practice. To address these limitations, we introduce VSGD-Net, a novel detection network that learns melanocyte identification through virtual staining from H&E to Sox10. The method takes only routine H&E images during inference, resulting in a promising approach to support pathologists in the diagnosis of melanoma. To the best of our knowledge, this is the first study that investigates the detection problem using image synthesis features between two distinct pathology stainings. Extensive experimental results show that our proposed model outperforms state-of-the-art nuclei detection methods for melanocyte detection. The source code and pre-trained model are available at: https://github.com/kechunl/VSGD-Net.
Kechun Liu, Beibin Li, Caitlin J. May, Oliver Chang, Stevan Knezevich, Lisa M. Reisch, Joann G. Elmore, Linda G. Shapiro
WACV2
2023 Stratification of Children with Autism Spectrum Disorder Through Fusion of Temporal Information in Eye-gaze Scan-Paths
abstract
Background: Looking pattern differences are shown to separate individuals with Autism Spectrum Disorder (ASD) and Typically Developing (TD) controls. Recent studies have shown that, in children with ASD, these patterns change with intellectual and social impairments, suggesting that patterns of social attention provide indices of clinically meaningful variation in ASD. Method: We conducted a naturalistic study of children with ASD (n = 55) and typical development (TD, n = 32). A battery of eye-tracking video stimuli was used in the study, including Activity Monitoring (AM), Social Referencing (SR), Theory of Mind (ToM), and Dyadic Bid (DB) tasks. This work reports on the feasibility of spatial and spatiotemporal scanpaths generated from eye-gaze patterns of these paradigms in stratifying ASD and TD groups. Algorithm: This article presents an approach for automatically identifying clinically meaningful information contained within the raw eye-tracking data of children with ASD and TD. The proposed mechanism utilizes combinations of eye-gaze scan-paths (spatial information), fused with temporal information and pupil velocity data and Convolutional Neural Network (CNN) for stratification of diagnosis (ASD or TD). Results: Spatial eye-gaze representations in the form of scanpaths in stratifying ASD and TD (ASD vs. TD: DNN: 74.4%) are feasible. These spatial eye-gaze features, e.g., scan-paths, are shown to be sensitive to factors mediating heterogeneity in ASD: age (ASD: 2–4 y/old vs. 10–17 y/old CNN: 80.5%), gender (Male vs. Female ASD: DNN: 78.0%) and the mixture of age and gender (5–9 y/old Male vs. 5–9 y/old Female ASD: DNN:98.8%). Limiting scan-path representations temporally increased variance in stratification performance, attesting to the importance of the temporal dimension of eye-gaze data. Spatio-Temporal scan-paths that incorporate velocity of eye movement in their images of eye-gaze are shown to outperform other feature representation methods achieving classification accuracy of 80.25%. Conclusion: The results indicate the feasibility of scan-path images to stratify ASD and TD diagnosis in children of varying ages and gender. Infusion of temporal information and velocity data improves the classification performance of our deep learning models. Such novel velocity fused spatio-temporal scan-path features are shown to be able to capture eye gaze patterns that reflect age, gender, and the mixed effect of age and gender, factors that are associated with heterogeneity in ASD and difficulty in identifying robust biomarkers for ASD.
Adham Atyabi, Frédérick Shic, Jiajun Jiang, Claire E. Foster, Erin Barney, Minah Kim, Beibin Li, Pamela Ventola, Chung-Hao Chen
ACM Trans. Knowl. Discov. Data7
2022 Calibration Error Prediction: Ensuring High-Quality Mobile Eye-Tracking
abstract
Gaze calibration is common in traditional infrared oculographic eye tracking. However, it is not well studied in visible-light mobile/remote eye tracking. We developed a lightweight real-time gaze error estimator and analyzed calibration errors from two perspectives: facial feature-based and Monte Carlo-based. Both methods correlated with gaze estimation errors, but the Monte Carlo method associated more strongly. Facial feature associations with gaze error were interpretable, relating movements of the face to the visibility of the eye. We highlight the degradation of gaze estimation quality in a sample of children with autism spectrum disorder (as compared to typical adults), and note that calibration methods may improve Euclidean error by 10%.
Beibin Li, James C. Snider, Quan Wang 0003, Sachin Mehta, Claire E. Foster, Erin Barney, Linda G. Shapiro, Pamela Ventola, Frédérick Shic
ETRA1
2022 Warper: Efficiently Adapting Learned Cardinality Estimators to Data and Workload Drifts
abstract
Recent learned cardinality estimation (CE) models are vulnerable when query predicates or the underlying datasets drift from what the models were trained upon. We propose a system Warper that accelerates model adaptation to drifts; Warper generates additional queries when limited examples are available from the new workload and carefully picks which queries to use to update the CE model. We show that Warper can be used to adapt different CE models including ones that support queries over single tables and join expressions. Experiments with different drifts suggest that Warper has a small computational cost and adapts much faster compared to state-of-the-art solutions. We also show that faster model adaptation improves query performance by shortening the period for which imperfect query plans are picked by a query optimizer due to incorrect cardinality estimates.
Beibin Li, Yao Lu 0028, Srikanth Kandula
SIGMOD Conference1
2021 Learning Oculomotor Behaviors from Scanpath
abstract
Identifying oculomotor behaviors relevant for eye-tracking applications is a critical but often challenging task. Aiming to automatically learn and extract knowledge from existing eye-tracking data, we develop a novel method that creates rich representations of oculomotor scanpaths to facilitate the learning of downstream tasks. The proposed stimulus-agnostic Oculomotor Behavior Framework (OBF) model learns human oculomotor behaviors from unsupervised and semi-supervised tasks, including reconstruction, predictive coding, fixation identification, and contrastive learning tasks. The resultant pre-trained OBF model can be used in a variety of applications. Our pre-trained model outperforms baseline approaches and traditional scanpath methods in autism spectrum disorder and viewed-stimulus classification tasks. Ablation experiments further show our proposed method could achieve even better results with larger model sizes and more diverse eye-tracking training datasets, supporting the model’s potential for future eye-tracking applications. Open source code: http://github.com/BeibinLi/OBF.
Beibin Li, Nicholas Nuechterlein, Erin Barney, Claire E. Foster, Minah Kim, Monique Mahony, Adham Atyabi, Quan Wang 0003, Pamela Ventola, Linda G. Shapiro, Frédérick Shic
ICMI1
2020 Selection of Eye-Tracking Stimuli for Prediction by Sparsely Grouped Input Variables for Neural Networks: towards Biomarker Refinement for Autism
abstract
Eye tracking has become a powerful tool in the study of autism spectrum disorder (ASD). Current, large-scale efforts aim to identify specific eye-tracking stimuli to be used as biomarkers for ASD, with the intention of informing the diagnostic process, monitoring therapeutic response, predicting outcomes, or identifying subgroups with the spectrum. However, there are hundreds of candidate experimental paradigms, each of which contains dozens or even hundreds of individual stimuli. Each stimuli is associated with an array of potential derived outcome variables, thus the number of variables to consider can be enormous. Standard variable selection techniques are not applicable to this problem, because selection must be done at the level of stimuli and not individual variables. In other words, this is a grouped variable selection problem. In this work, we apply lasso, group lasso, and a new technique, Sparsely Grouped Input Variables for Neural Network (SGIN), to select experimental stimuli for group discrimination and regression with clinical variables. Using a dataset obtained from children with and without ASD who were administered a battery containing 109 different stimuli presentations involving 9647 features, we are able to retain strong group separation even with only 11 out of the 109 stimuli. This work sets the stage for concerted techniques designed around engines to iteratively refine and define next-generation biomarkers using eye tracking for psychiatric conditions. http://github.com/beibinli/SGIN
Beibin Li, Erin Barney, Caitlin Hudac, Nicholas Nuechterlein, Pamela Ventola, Linda G. Shapiro, Frédérick Shic
ETRA1
2020 Classifying Breast Histopathology Images with a Ductal Instance-Oriented Pipeline
abstract
In this study, we propose the Ductal Instance-Oriented Pipeline (DIOP) that contains a duct-level instance segmentation model, a tissue-level semantic segmentation model, and three-levels of features for diagnostic classification. Based on recent advancements in instance segmentation and the Mask RCNN model, our duct-level segmenter tries to identify each ductal individual inside a microscopic image; then, it extracts tissue-level information from the identified ductal instances. Leveraging three levels of information obtained from these ductal instances and also the histopathology image, the proposed DIOP outperforms previous approaches (both feature-based and CNN-based) in all diagnostic tasks; for the four-way classification task, the DIOP achieves comparable performance to general pathologists in this unique dataset. The proposed DIOP only takes a few seconds to run in the inference time, which could be used interactively on most modern computers. More clinical explorations are needed to study the robustness and generalizability of this system in the future.
Beibin Li, Ezgi Mercan, Sachin Mehta, Stevan Knezevich, Corey W. Arnold, Donald L. Weaver, Joann G. Elmore, Linda G. Shapiro
ICPR1
2020 Leveraging Unlabeled Data for Glioma Molecular Subtype and Survival Prediction
abstract
In this paper, we address two long-standing radio-genomic challenges in glioma subtype and survival prediction: (1) how to leverage large amounts of unlabeled magnetic resonance (MR) imaging data and (2) how to unite MR data and genomic data. We propose a novel application of multi-task learning (MTL) that leverages unlabeled MR data by jointly learning an auxiliary tumor segmentation task with glioma subtype prediction and that can learn from patients with and without genomic data. We analyze multi-parametric MR data from 542 patients in the combined training, validation, and testing sets of the 2018 Multimodal Brain Tumor Segmentation Challenge and somatic copy number alteration (SCNA) data from 1090 patients in The Cancer Genome Atlas' (TCGA) lower-grade glioma and glioblastoma projects. Our MTL model significantly outperforms comparable classification models trained only on labeled MR data for both IDH1/2 mutation and 1p/19q co-deletion subtype prediction tasks. We also show that embeddings produced by our MTL models improve survival predictions beyond MR or SCNA on their own. Our code is available at https://github.com/nknuecht/glioma_mtl.
Nicholas Nuechterlein, Beibin Li, Mehmet Saygin Seyfioglu, Sachin Mehta, Patrick J. Cimino, Linda G. Shapiro
ICPR2
2019 A Facial Affect Analysis System for Autism Spectrum Disorder
abstract
In this paper, we introduce an end-to-end machine learning-based system for classifying autism spectrum disorder (ASD) using facial attributes such as expressions, action units, arousal, and valence. Our system classifies ASD using representations of different facial attributes from convolutional neural networks, which are trained on images in the wild. Our experimental results show that different facial attributes used in our system are statistically significant and improve sensitivity, specificity, and F1 score of ASD classification by a large margin. In particular, the addition of different facial attributes improves the performance of ASD classification by about 7% which achieves a F1 score of 76%.
Beibin Li, Sachin Mehta, Deepali Aneja, Claire E. Foster, Pamela Ventola, Frédérick Shic, Linda G. Shapiro
ICIP1
2018 Social Influences on Executive Functioning in Autism: Design of a Mobile Gaming Platform
abstract
Most studies of executive function (EF) in Autism Spectrum Disorder (ASD) focus on cognitive information processing, emphasizing less the social interaction deficits core to ASD. We designed a mobile game that uses social and nonsocial stimuli to assess children's EF skills. The game comprised three components involving different EF skills: cognitive flexibility (shifting/inference), inhibitory control, and short-term memory. By recruiting 65 children with and without ASD to play the mobile game, we investigated the potential of such platforms for capturing important phenotypic characteristics of individuals with autism. Results highlighted between-diagnostic-group differences in playing patterns with children with ASD showing broad patterns of EF deficits, but with relative strengths in nonsocial short-term memory, and preserved response to emotional inhibition cues. We showed the system could predict IQ, an important target for clinical treatment, towards the goal of developing platforms to act as long-term, efficient, and effective behavioral biomarkers for ASD.
Beibin Li, Adham Atyabi, Minah Kim, Erin Barney, Amy Yeo-jin Ahn, Yawen Luo, Madeline Aubertine, Sarah Corrigan, Tanya St. John, Quan Wang 0003, Marilena Mademtzi, Mary Best, Frédérick Shic
CHI1
2017 An exploratory analysis targeting diagnostic classification of AAC app usage patterns
abstract
Augmentative and Alternative Communication (AAC) apps are apps that enable non-speech communicative forms. One class of AAC apps are speech-generating devices (SGDs), where icons/pictures are tapped to produce spoken words. These apps are widely used to support communication and language learning for individuals with disabilities such as autism spectrum disorder (ASD). Given that these apps are used in everyday scenarios, they can generate massive streams of data, providing a wealth of information regarding individual usage patterns and for developing usage model profiles. However, the utility and potential of these streams of data has been little explored from a data mining perspective. The objective of this study is to evaluate several feature representations of usage patterns, coupled with data mining and data modelling techniques, for identifying differences in AAC usage patterns between users with and without ASD. The study is conducted using data streams aggregated from an AAC app called FreeSpeech, specifically designed for individuals with learning disabilities and ASD. Several feature representations for modeling usage profiles based on temporal, behavioral and frequency of usage, are investigated. The potential of each usage representation is assessed using a collection of well-known and well-established learning methods such as support vector machine and ensemble learning. While, in general, prediction performance was only slightly above chance in most representations, results from unsupervised class labeling experiments showed promising results regarding the potential of stationary keypress usage representations with bootstrapped ensembles for separating ASD from non-ASD users.
Adham Atyabi, Beibin Li, Amy Yeo-jin Ahn, Minah Kim, Erin Barney, Frédérick Shic
IJCNN2
2016 Modified DBSCAN algorithm on oculomotor fixation identification
abstract
This paper modifies the DBSCAN algorithm to identify fixations and saccades. This method combines advantages from dispersion-based algorithms, such as resilience to noise and intuitive fixational structure, and from velocity-based algorithms, such as the ability to deal appropriately with smooth pursuit (SP) movements.
Beibin Li, Quan Wang 0003, Erin Barney, Logan Hart, Carla A. Wall, Katarzyna Chawarska, Irati Saez de Urabain, Timothy J. Smith, Frédérick Shic
ETRA1
2016 Optimality of the distance dispersion fixation identification algorithm
abstract
Researchers use fixation identification algorithms to parse eye movement trajectories into a series of fixations and saccades, simplifying analyses and providing measures which may relate to cognition. The Distance Dispersion (I-DD) a widely-used elementary fixation identification algorithm. Yet the "optimality" properties of its most popular greedy implementation have not been described. This paper: (1) asks how "optimal" should be defined, and advances maximizing total fixation time and minimizing number of clusters as a definition; (2) asks whether the greedy implementation of I-DD is optimal, and shows that it is when no fixations are rejected for being too short; and (3) we show that when fixation time rejection criterion are enabled, the greedy algorithm is not optimal. We propose an O(n2) algorithm which is.
Beibin Li, Quan Wang 0003, Laura Boccanfuso, Frédérick Shic
ETRA1
2016 Thermographic eye tracking
abstract
Far infrared thermography, which can be used to detect thermal radiation emitted by humans, has been used to detect physical disease, physiological changes relating to emotion, and polygraph testing, but has not been used for eye tracking. However, because the surface temperature of the cornea is colder than the limbus, it is theoretically possible to track corneal movements through thermal imaging. To explore the feasibility of thermal eye tracking, we invited 10 adults and tracked their corneal movements with passive thermal imaging at 60 Hz. We combined shape models of eyes with intensity threshold to segment the cornea from other parts of the eye in thermal images. We used an animation sequence as a calibration target for 5 point calibration/validation 5 times. Our results were compared to simultaneously collected data using an SR EyeLink eye tracker at 500 Hz, demonstrating the feasibility of eye tracking with thermal images. Blinking and breathing frequencies, which reflect the psychophysical status of the participants, were also robustly detected during thermal eye tracking.
Quan Wang 0003, Laura Boccanfuso, Beibin Li, Amy Yeo-jin Ahn, Claire E. Foster, Margaret P. Orr, Brian Scassellati, Frédérick Shic
ETRA3
2016 A thermal emotion classifier for improved human-robot interaction
abstract
In their expanding role as tutors, home and healthcare assistants, robots must effectively interact with individuals of varying ability and temperament. Indeed, deploying robots in long-term social engagements will almost certainly require robots to reliably detect and adapt to changes in the demeanor of social partners to promote trust and more productive collaboration. However, the recognition of emotional state typically relies on the interpretation of very subtle cues, often varying from one person to the next. In addition, while facial expressions, body posture and features of speech have been used to detect affective changes, the robustness of these measures is often hindered by cultural and age differences. Recently, infrared thermography has shown promise in detecting guilt, fear and stress, indicating that it may be a viable sensing modality for improved human-robot interaction. In this study, we evaluated the efficacy of using a far infrared (FIR) camera for detecting robot-elicited affective response compared to video-elicited affective response by tracking thermal changes in five areas of the face. Further, we analyzed localized changes in the face to assess whether thermal and electrodermal responses to emotions elicited by traditional video techniques and by robots are similar. Finally, we performed principal component analysis to reduce the dimensionality of data and evaluated the performance using machine learning techniques for classifying thermal data by emotion state, resulting in a thermal classifier with a performance accuracy of 77.5%.
Laura Boccanfuso, Quan Wang 0003, Iolanda Leite, Beibin Li, Colette Torres, Lisa Chen, Nicole Salomons, Claire E. Foster, Erin Barney, Amy Yeo-jin Ahn, Brian Scassellati, Frédérick Shic
RO-MAN4