VLDB 2026 Research / reviewers in the wild / expert
Mihai Burzo
dblp:120/4061 · also Mihai G. Burzo
· DBLP profile ↗
24ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-3968-7343ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 8 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Personalized Multimodal Efficient Detection of Human Circadian States
Kapotaksha Das, Mihai Burzo, Mohamed Abouelenien |
ICPR (7) | 2 |
| 2026 | More Than Meets the Ear: Multimodal Driver Alertness Detection Leveraging LLMs and Synthetic Speech
Salem Sharak, Kapotaksha Das, Mihai Burzo, Mohamed Abouelenien |
ICPR (13) | 3 |
| 2023 | Towards Modeling Human Biorhythms Using Non-Invasive ApproachesabstractBiorhythms in humans hold vital importance that is directly related to a person’s health and well-being. This paper introduces a novel dataset for modeling biorhythms while focusing primarily on using an ingestible pill to monitor core body temperature and four physiological signals. We use two different settings to label our data to develop different scenarios of energization and enervation. We propose an approach that could potentially act as a reliable alternative to the usage of core temperature pills and genome sequencing while avoiding their invasiveness and complexities. We also explore the effect of the data size on the overall performance. Our findings indicate that our proposed approach is reliable and feasible, and presents future promise in using less invasive approaches to monitor and predict the circadian state of humans. Kapotaksha Das, John Elson, Mohamed Abouelenien, Ali Hassani 0007, Mihai Burzo, Clay Maranville |
BIBM | 5 |
| 2023 | Non-Contact Based Modeling of EnervationabstractSignificant research is currently carried out with a focus on autonomous vehicles; research is starting to focus on areas such as the modeling of occupant states and behavioral elements. This paper contributes to this line of research by developing a pipeline that extracts physiological signals from thermal imagery and modeling occupant enervation using a fully non-contact based approach. These signals are obtained via a multimodal dataset of 36 subjects across multiple channels, including the thermal and physiological modalities. Moreover, we provide a comparative analysis of non-contact and contact based channels to model the enervation state of individuals. Our analysis indicates that non-contact physiological signals extracted from thermal imagery can reach and exceed the performance of contact-based physiological signals. In addition, modeling of enervation is possible using said non-contact physiological signals and thermal features, with an accuracy of up to 70% in identifying energized and enervated occupant states. Our findings provide a novel approach for future research and opens the possibility for integration of unrestrictive sensors in future automobiles. Kais Riani, Salem Sharak, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea, John Elson, Clay Maranville, Kwaku O. Prakah-Asante, Waqas Manzoor |
FG | 4 |
| 2022 | In-the-Wild Video Question AnsweringabstractExisting video understanding datasets mostly focus on human interactions, with little attention being paid to the “in the wild” settings, where the videos are recorded outdoors. We propose WILDQA, a video understanding dataset of videos recorded in outside settings. In addition to video question answering (Video QA), we also introduce the new task of identifying visual support for a given question and answer (Video Evidence Selection). Through evaluations using a wide range of baseline models, we show that WILDQA poses new challenges to the vision and language research communities. The dataset is available at https://lit.eecs.umich.edu/wildqa/. Santiago Castro, Naihao Deng, Pingxuan Huang, Mihai Burzo, Rada Mihalcea |
COLING | 4 |
| 2022 | Contact Versus Noncontact Detection of Driver's DrowsinessabstractWith an estimated number of injuries in the millions, accidents caused due to drowsy driving remain a significant source of financial costs and loss of life. Accurate detection of driver’s drowsiness could provide a clear avenue towards eliminating a great majority of the associated accidents and losses. Existing research on the subject could be defined as either contact-based or noncontact-based alertness detection. This paper utilizes a novel multimodal driver’s alertness dataset consisting of 45 subjects via seven recorded channels, including four contact-based and three noncontact-based channels, to investigate the performance of said modalities in detecting driver’s drowsiness as well as provide a novel comparison between the results of multiple contact and noncontact methods. Our results highlight the viability of noncontact methods to detect driver’s drowsiness as an implementable technology in automobiles. Salem Sharak, Kapotaksha Das, Kais Riani, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea |
ICPR | 5 |
| 2022 | Multimodal Deception Detection Using Real-Life Trial DataabstractHearings of witnesses and defendants play a crucial role when reaching court trial decisions. Given the high-stakes nature of trial outcomes, developing computational models that assist the decision-making process is an important research venue. In this article, we address the identification of deception in real-life trial data. We use a dataset consisting of videos collected from public court trials. We explore the use of verbal and non-verbal modalities to build a multimodal deception detection system that aims to discriminate between truthful and deceptive statements provided by defendants and witnesses. In particular, three complementary modalities (visual, acoustic and linguistic) are evaluated for the classification of deception at the subject level. The final classifier is obtained by combining the three modalities via score-level classification, achieving 83.05 percent accuracy in subject-level deceit detection. To place our results in perspective, we present a human deception detection study where we evaluate the human capability of detecting deception using different modalities and compare the results to the developed system. The results show that our system outperforms the average non-expert human capability of identifying deceit. Mehmet Umut Sen, Verónica Pérez-Rosas, Berrin A. Yanikoglu, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea |
IEEE Trans. Affect. Comput. | 5 |
| 2022 | Detection and Recognition of Driver Distraction Using Multimodal SignalsabstractDistracted driving is a leading cause of accidents worldwide. The tasks of distraction detection and recognition have been traditionally addressed as computer vision problems. However, distracted behaviors are not always expressed in a visually observable way. In this work, we introduce a novel multimodal dataset of distracted driver behaviors, consisting of data collected using twelve information channels coming from visual, acoustic, near-infrared, thermal, physiological and linguistic modalities. The data were collected from 45 subjects while being exposed to four different distractions (three cognitive and one physical). For the purposes of this paper, we performed experiments with visual, physiological, and thermal information to explore potential of multimodal modeling for distraction recognition. In addition, we analyze the value of different modalities by identifying specific visual, physiological, and thermal groups of features that contribute the most to distraction characterization. Our results highlight the advantage of multimodal representations and reveal valuable insights for the role played by the three modalities on identifying different types of driving distractions. Kapotaksha Das, Michalis Papakostas, Kais Riani, Andrew Brian Gasiorowski, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea |
ACM Trans. Interact. Intell. Syst. | 6 |
| 2021 | Towards Classifying Human Circadian Rhythm Using Multiple ModalitiesabstractAutonomous vehicles represent one of the most active technologies currently being developed, with research areas addressing, among others, the modeling of the states and behavioral elements of the occupants. This paper contributes to this line of research by studying the circadian rhythm of individuals using a novel multimodal dataset of 36 subjects consisting of five information channels. These channels include visual, thermal, physiological, linguistic, and background data. Moreover, we propose a framework to explore whether the circadian rhythm can be modeled without continuous monitoring and investigate the hypothesis that multimodal features have a greater propensity for improved performance using data points specific to certain times during the day. Our analysis shows that multimodal fusion can lead to an accuracy of up to 77% on identifying energized and enervated states of the participants. Our findings highlight the validity of our hypothesis and present a novel approach for future research. Kais Riani, Salem Sharak, Kapotaksha Das, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea, John Elson, Clay Maranville, Kwaku O. Prakah-Asante, Waqas Manzoor |
ACII | 5 |
| 2021 | Multimodal Detection of Drivers Drowsiness and DistractionabstractConsidering the ever-growing presence of automobiles around the world, ensuring the safety of those on and near roadways is of great importance. From the causes of accidents, drowsiness and distractedness are among the most consequential. In this paper, we use a multimodal dataset consisting of 11 recorded channels over 45 subjects to model driver’s drowsiness and distraction. Our work puts forward the application of this dataset by using segmented windows as features, resulting in four main contributions. We explore the performance of each individual modality and specify which signals and features have a better capability of detecting drowsiness and different kinds of distractions. In addition, we analyze the effects of early fusion on the classification of the driver’s state using multiple physiological and thermal channels. Finally, we use cascaded late fusion and test three voting strategies to evaluate the performance of our proposed approach. Our results confirm the effectiveness of utilizing a multimodal approach in detecting both drowsiness and distraction as two separate factors influencing the driver and provide guidelines on which signals are appropriate for detecting different driver’s states. Kapotaksha Das, Salem Sharak, Kais Riani, Mohamed Abouelenien, Mihai Burzo, Michalis Papakostas |
ICMI | 5 |
| 2021 | Understanding Driving Distractions: A Multimodal Analysis on Distraction CharacterizationabstractDistracted driving is a leading cause of accidents worldwide. The tasks of distraction detection and recognition have been traditionally addressed as computer vision problems. However, distracted behaviors are not always expressed in a visually observable way. In this work, we introduce a novel multimodal dataset of distracted driver behaviors, consisting of data collected using twelve information channels coming from visual, acoustic, near-infrared, thermal, physiological and linguistic modalities. The data were collected from 45 subjects while being exposed to four different distractions (three cognitive and one physical). For the purposes of this paper, we experiment with visual and physiological information and explore the potential of multimodal modeling for distraction recognition. In addition, we analyze the value of different modalities by identifying specific visual and physiological groups of features that contribute the most to distraction characterization. Our results highlight the advantage of multimodal representations and reveal valuable insights for the role played by the two modalities on identifying different types of driving distractions. Michalis Papakostas, Kais Riani, Andrew Brian Gasiorowski, Mohamed Abouelenien, Rada Mihalcea, Mihai Burzo |
IUI | 7 |
| 2021 | MUSER: MUltimodal Stress detection using Emotion Recognition as an Auxiliary TaskabstractYiqun Yao, Michalis Papakostas, Mihai Burzo, Mohamed Abouelenien, Rada Mihalcea. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yiqun Yao, Michalis Papakostas, Mihai Burzo, Mohamed Abouelenien, Rada Mihalcea |
NAACL-HLT | 3 |
| 2020 | MORSE: MultimOdal sentiment analysis for Real-life SEttingsabstractMultimodal sentiment analysis aims to detect and classify sentiment expressed in multimodal data. Research to date has focused on datasets with a large number of training samples, manual transcriptions, and nearly-balanced sentiment labels. However, data collection in real settings often leads to small datasets with noisy transcriptions and imbalanced label distributions, which are therefore significantly more challenging than in controlled settings. In this work, we introduce MORSE, a domain-specific dataset for MultimOdal sentiment analysis in Real-life SEttings. The dataset consists of 2,787 video clips extracted from 49 interviews with panelists in a product usage study, with each clip annotated for positive, negative, or neutral sentiment. The characteristics of MORSE include noisy transcriptions from raw videos, naturally imbalanced label distribution, and scarcity of minority labels. To address the challenging real-life settings in MORSE, we propose a novel two-step fine-tuning method for multimodal sentiment classification using transfer learning and the Transformer model architecture; our method starts with a pre-trained language model and one step of fine-tuning on the language modality, followed by the second step of joint fine-tuning that incorporates the visual and audio modalities. Experimental results show that while MORSE is challenging for various baseline models such as SVM and Transformer, our two-step fine-tuning method is able to capture the dataset characteristics and effectively address the challenges. Our method outperforms related work that uses both single and multiple modalities in the same transfer learning settings. Yiqun Yao, Verónica Pérez-Rosas, Mohamed Abouelenien, Mihai Burzo |
ICMI | 4 |
| 2020 | MuSE: a Multimodal Dataset of Stressed EmotionabstractEndowing automated agents with the ability to provide support, entertainment and interaction with human beings requires sensing of the users’ affective state. These affective states are impacted by a combination of emotion inducers, current psychological state, and various conversational factors. Although emotion classification in both singular and dyadic settings is an established area, the effects of these additional factors on the production and perception of emotion is understudied. This paper presents a new dataset, Multimodal Stressed Emotion (MuSE), to study the multimodal interplay between the presence of stress and expressions of affect. We describe the data collection protocol, the possible areas of use, and the annotations for the emotional content of the recordings. The paper also presents several baselines to measure the performance of multimodal features for emotion and stress classification. Mimansa Jaiswal, Cristian-Paul Bara, Yuanhang Luo, Mihai Burzo, Rada Mihalcea, Emily Mower Provost |
LREC | 4 |
| 2019 | Muse-ing on the Impact of Utterance Ordering on Crowdsourced Emotion AnnotationsabstractEmotion recognition algorithms rely on data annotated with high quality labels. However, emotion expression and perception are inherently subjective. There is generally not a single annotation that can be unambiguously declared "correct." As a result, annotations are colored by the manner in which they were collected. In this paper, we conduct crowdsourcing experiments to investigate this impact on both the annotations themselves and on the performance of these algorithms. We focus on one critical question: the effect of context. We present a new emotion dataset, Multimodal Stressed Emotion (MuSE), and annotate the dataset using two conditions: randomized, in which annotators are presented with clips in random order, and contextualized, in which annotators are presented with clips in order. We find that contextual labeling schemes result in annotations that are more similar to a speaker's own self-reported labels and that labels generated from randomized schemes are most easily predictable by automated systems. Mimansa Jaiswal, Zakaria Aldeneh, Cristian-Paul Bara, Yuanhang Luo, Mihai Burzo, Rada Mihalcea, Emily Mower Provost |
ICASSP | 5 |
| 2017 | Multimodal gender detectionabstractAutomatic gender classification is receiving increasing attention in the computer interaction community as the need for personalized, reliable, and ethical systems arises. To date, most gender classification systems have been evaluated on textual and audiovisual sources. This work explores the possibility of enhancing such systems with physiological cues obtained from thermography and physiological sensor readings. Using a multimodal dataset consisting of audiovisual, thermal, and physiological recordings of males and females, we extract features from five different modalities, namely acoustic, linguistic, visual, thermal, and physiological. We then conduct a set of experiments where we explore the gender prediction task using single and combined modalities. Experimental results suggest that physiological and thermal information can be used to recognize gender at reasonable accuracy levels, which are comparable to the accuracy of current gender prediction systems. Furthermore, we show that the use of non-contact physiological measurements, such as thermography readings, can enhance current systems that are based on audio or visual input. This can be particularly useful for scenarios where non-contact approaches are preferred, i.e., when data is captured under noisy audiovisual conditions or when video or speech data are not available due to ethical considerations. Mohamed Abouelenien, Verónica Pérez-Rosas, Rada Mihalcea, Mihai Burzo |
ICMI | 4 |
| 2017 | Detecting Deceptive Behavior via Integration of Discriminative Features From Multiple ModalitiesabstractDeception detection has received an increasing amount of attention in recent years, due to the significant growth of digital media, as well as increased ethical and security concerns. Earlier approaches to deception detection were mainly focused on law enforcement applications and relied on polygraph tests, which had proved to falsely accuse the innocent and free the guilty in multiple cases. In this paper, we explore a multimodal deception detection approach that relies on a novel data set of 149 multimodal recordings, and integrates multiple physiological, linguistic, and thermal features. We test the system on different domains, to measure its effectiveness and determine its limitations. We also perform feature analysis using a decision tree model, to gain insights into the features that are most effective in detecting deceit. Our experimental results indicate that our multimodal approach is a promising step toward creating a feasible, non-invasive, and fully automated deception detection system. Mohamed Abouelenien, Verónica Pérez-Rosas, Rada Mihalcea, Mihai Burzo |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2015 | Verbal and Nonverbal Clues for Real-life Deception DetectionabstractDeception detection has been receiving an increasing amount of attention from the computational linguistics, speech, and multimodal processing communities. One of the major challenges encountered in this task is the availability of data, and most of the research work to date has been conducted on acted or artificially collected data. The generated deception models are thus lacking real-world evidence. In this paper, we explore the use of multimodal real-life data for the task of deception detection. We develop a new deception dataset consisting of videos from reallife scenarios, and build deception tools relying on verbal and nonverbal features. We achieve classification accuracies in the range of 77-82% when using a model that extracts and fuses features from the linguistic and visual modalities. We show that these results outperform the human capability of identifying deceit. Verónica Pérez-Rosas, Mohamed Abouelenien, Rada Mihalcea, C. J. Linton, Mihai Burzo |
EMNLP | 6 |
| 2015 | Deception Detection using Real-life Trial DataabstractHearings of witnesses and defendants play a crucial role when reaching court trial decisions. Given the high-stake nature of trial outcomes, implementing accurate and effective computational methods to evaluate the honesty of court testimonies can offer valuable support during the decision making process. In this paper, we address the identification of deception in real-life trial data. We introduce a novel dataset consisting of videos collected from public court trials. We explore the use of verbal and non-verbal modalities to build a multimodal deception detection system that aims to discriminate between truthful and deceptive statements provided by defendants and witnesses. We achieve classification accuracies in the range of 60-75% when using a model that extracts and fuses features from the linguistic and gesture modalities. In addition, we present a human deception detection study where we evaluate the human capability of detecting deception in trial hearings. The results show that our system outperforms the human capability of identifying deceit. Verónica Pérez-Rosas, Mohamed Abouelenien, Rada Mihalcea, Mihai Burzo |
ICMI | 4 |
| 2014 | Deception detection using a multimodal approachabstractIn this paper we address the automatic identification of deceit by using a multimodal approach. We collect deceptive and truthful responses using a multimodal setting where we acquire data using a microphone, a thermal camera, as well as physiological sensors. Among all available modalities, we focus on three modalities namely, language use, physiological response, and thermal sensing. To our knowledge, this is the first work to integrate these specific modalities to detect deceit. Several experiments are carried out in which we first select representative features for each modality, and then we analyze joint models that integrate several modalities. The experimental results show that the combination of features from different modalities significantly improves the detection of deceptive behaviors as compared to the use of one modality at a time. Moreover, the use of non-contact modalities proved to be comparable with and sometimes better than existing contact-based methods. The proposed method increases the efficiency of detecting deceit by avoiding human involvement in an attempt to move towards a completely automated non-invasive deception detection process. Mohamed Abouelenien, Verónica Pérez-Rosas, Rada Mihalcea, Mihai Burzo |
ICMI | 4 |
| 2014 | A Multimodal Dataset for Deception Detection
Verónica Pérez-Rosas, Rada Mihalcea, Alexis Narvaez, Mihai Burzo |
LREC | 4 |
| 2013 | Automatic detection of deceit in verbal communicationabstractThis paper presents experiments in building a classifier for the automatic detection of deceit. Using a dataset of deceptive videos, we run several comparative evaluations focusing on the verbal component of these videos, with the goal of understanding the difference in deceit detection when using manual versus automatic transcriptions, as well as the difference between spoken and written lies. We show that using only the linguistic component of the deceptive videos, we can detect deception with accuracies in the range of 52-73%. Rada Mihalcea, Verónica Pérez-Rosas, Mihai Burzo |
ICMI | 3 |
| 2012 | Towards sensing the influence of visual narratives on human affectabstractIn this paper, we explore a multimodal approach to sensing affective state during exposure to visual narratives. Using four different modalities, consisting of visual facial behaviors, thermal imaging, heart rate measurements, and verbal descriptions, we show that we can effectively predict changes in human affect. Our experiments show that these modalities complement each other, and illustrate the role played by each of the four modalities in detecting human affect. Mihai Burzo, Daniel McDuff, Rada Mihalcea, Louis-Philippe Morency, Alexis Narvaez, Verónica Pérez-Rosas |
ICMI | 1 |
| 2012 | Towards multimodal deception detection - step 1: building a collection of deceptive videosabstractIn this paper, we introduce a novel crowdsourced dataset of deceptive videos. We describe the collection process and the characteristics of the dataset, and we validate it through initial experiments in the recognition of deceptive language. The collection, consisting of 140 truthful and deceptive videos, will enable future experiments in multimodal deceptive detection. Rada Mihalcea, Mihai Burzo |
ICMI | 2 |