Ekta Sood

dblp:225/6412 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0002-6267-7151ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 InteRead: An Eye Tracking Dataset of Interrupted Reading
abstract
Eye movements during reading offer a window into cognitive processes and language comprehension, but the scarcity of reading data with interruptions – which learners frequently encounter in their everyday learning environments – hampers advances in the development of intelligent learning technologies. We introduce InteRead – a novel 50-participant dataset of gaze data recorded during self-paced reading of real-world text. InteRead further offers fine-grained annotations of interruptions interspersed throughout the text as well as resumption lags incurred by these interruptions. Interruptions were triggered automatically once readers reached predefined target words. We validate our dataset by reporting interdisciplinary analyses on different measures of gaze behavior. In line with prior research, our analyses show that the interruptions as well as word length and word frequency effects significantly impact eye movements during reading. We also explore individual differences within our dataset, shedding light on the potential for tailored educational solutions. InteRead is accessible from our datasets web-page: https://www.ife.uni-stuttgart.de/en/llis/research/datasets/.
Francesca Zermiani, Prajit Dhar, Ekta Sood, Fabian Kögel, Andreas Bulling, Maria Wirzberger
LREC/COLING3
2024 GEARS: Generalizable Multi-Purpose Embeddings for Gaze and Hand Data in VR Interactions
abstract
Machine learning models using users’ gaze and hand data to encode user interaction behavior in VR are often tailored to a single task and sensor set, limiting their applicability in settings with constrained compute resources. We propose GEARS, a new paradigm that learns a shared feature extraction mechanism across multiple tasks and sensor sets to encode gaze and hand tracking data of users VR behavior into multi-purpose embeddings. GEARS leverages a contrastive learning framework to learn these embeddings, which we then use to train linear models to predict task labels. We evaluated our paradigm across four VR datasets with eye tracking that comprise different sensor sets and task goals. The performance of GEARS was comparable to results from models trained for a single task with data of a single sensor set. Our research advocates a shift from using sensor set and task specific models towards using one shared feature extraction mechanism to encode users’ interaction behavior in VR.
Philipp Hallgarten, Naveen Sendhilnathan, Ting Zhang 0013, Ekta Sood, Tanya R. Jonker
UMAP4
2023 Impact of Privacy Protection Methods of Lifelogs on Remembered Memories
abstract
Lifelogging is traditionally used for memory augmentation. However, recent research shows that users’ trust in the completeness and accuracy of lifelogs might skew their memories. Privacy-protection alterations such as body blurring and content deletion are commonly applied to photos to circumvent capturing sensitive information. However, their impact on how users remember memories remain unclear. To this end, we conduct a white-hat memory attack and report on an iterative experiment (N=21) to compare the impact of viewing 1) unaltered lifelogs, 2) blurred lifelogs, and 3) a subset of the lifelogs after deleting private ones, on confidently remembering memories. Findings indicate that all the privacy methods impact memories’ quality similarly and that users tend to change their answers in recognition more than recall scenarios. Results also show that users have high confidence in their remembered content across all privacy methods. Our work raises awareness about the mindful designing of technological interventions.
Passant El Agroudy, Mohamed Khamis, Florian Mathis, Diana Irmscher, Ekta Sood, Andreas Bulling, Albrecht Schmidt 0001
CHI5
2023 Improving neural saliency prediction with a cognitive model of human visual attention
Ekta Sood, Lei Shi 0032, Matteo Bortoletto, Yao Wang 0018, Philipp Müller 0001, Andreas Bulling
CogSci1
2022 Gaze-enhanced Crossmodal Embeddings for Emotion Recognition
abstract
Emotional expressions are inherently multimodal -- integrating facial behavior, speech, and gaze -- but their automatic recognition is often limited to a single modality, e.g. speech during a phone call. While previous work proposed crossmodal emotion embeddings to improve monomodal recognition performance, despite its importance, an explicit representation of gaze was not included. We propose a new approach to emotion recognition that incorporates an explicit representation of gaze in a crossmodal emotion embedding framework. We show that our method outperforms the previous state of the art for both audio-only and video-only emotion classification on the popular One-Minute Gradual Emotion Recognition dataset. Furthermore, we report extensive ablation experiments and provide detailed insights into the performance of different state-of-the-art gaze representations and integration strategies. Our results not only underline the importance of gaze for emotion recognition but also demonstrate a practical and highly effective approach to leveraging gaze information for this task.
Ahmed Abdou, Ekta Sood, Philipp Müller 0001, Andreas Bulling
Proc. ACM Hum. Comput. Interact.2
2021 VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
abstract
We present VQA-MHUG -a novel 49participant dataset of multimodal human gaze on both images and questions during visual question answering (VQA) collected using a high-speed eye tracker.We use our dataset to analyze the similarity between human and neural attentive strategies learned by five state-ofthe-art VQA models: Modular Co-Attention Network (MCAN) with either grid or region features, Pythia, Bilinear Attention Network (BAN), and the Multimodal Factorized Bilinear Pooling Network (MFB).While prior work has focused on studying the image modality, our analyses show -for the first time -that for all models, higher correlation with human attention on text is a significant predictor of VQA performance.This finding points at a potential for improving VQA performance and, at the same time, calls for further research on neural text attention mechanisms and their integration into architectures for vision and language tasks, including but potentially also beyond VQA.
Ekta Sood, Fabian Kögel, Florian Strohm, Prajit Dhar, Andreas Bulling
CoNLL1
2021 Neural Photofit: Gaze-based Mental Image Reconstruction
abstract
We propose a novel method that leverages human fixations to visually decode the image a person has in mind into a photofit (facial composite). Our method combines three neural networks: An encoder, a scoring network, and a decoder. The encoder extracts image features and predicts a neural activation map for each face looked at by a human observer. A neural scoring network compares the human and neural attention and predicts a relevance score for each extracted image feature. Finally, image features are aggregated into a single feature vector as a linear combination of all features weighted by relevance which a decoder de-codes into the final photofit. We train the neural scoring network on a novel dataset containing gaze data of 19 participants looking at collages of synthetic faces. We show that our method significantly outperforms a mean baseline predictor and report on a human study that shows that we can decode photofits that are visually plausible and close to the observer’s mental image.
Florian Strohm, Ekta Sood, Sven Mayer, Philipp Müller 0001, Mihai Bâce, Andreas Bulling
ICCV2
2020 Interpreting Attention Models with Human Visual Attention in Machine Reading Comprehension
abstract
While neural networks with attention mechanisms have achieved superior performance on many natural language processing tasks, it remains unclear to which extent learned attention resembles human visual attention.In this paper, we propose a new method that leverages eye-tracking data to investigate the relationship between human visual attention and neural attention in machine reading comprehension.To this end, we introduce a novel 23 participant eye tracking dataset -MQA-RC, in which participants read movie plots and answered pre-defined questions.We compare state of the art networks based on long shortterm memory (LSTM), convolutional neural models (CNN) and XLNet Transformer architectures.We find that higher similarity to human attention and performance significantly correlates to the LSTM and CNN models.However, we show this relationship does not hold true for the XLNet models -despite the fact that the XLNet performs best on this challenging task.Our results suggest that different architectures seem to learn rather different neural attention strategies and similarity of neural to human attention does not guarantee best performance.
Ekta Sood, Simon Tannert, Diego Frassinelli, Andreas Bulling, Ngoc Thang Vu
CoNLL1
2020 Anticipating Averted Gaze in Dyadic Interactions
abstract
We present the first method to anticipate averted gaze in natural dyadic interactions. The task of anticipating averted gaze, i.e. that a person will not make eye contact in the near future, remains unsolved despite its importance for human social encounters as well as a number of applications, including human-robot interaction or conversational agents. Our multimodal method is based on a long short-term memory (LSTM) network that analyses non-verbal facial cues and speaking behaviour. We empirically evaluate our method for different future time horizons on a novel dataset of 121 YouTube videos of dyadic video conferences (74 hours in total). We investigate person-specific and person-independent performance and demonstrate that our method clearly outperforms baselines in both settings. As such, our work sheds light on the tight interplay between eye contact and other non-verbal signals and underlines the potential of computational modelling and anticipation of averted gaze for interactive applications.
Philipp Müller 0001, Ekta Sood, Andreas Bulling
ETRA2
2020 Improving Natural Language Processing Tasks with Human Gaze-Guided Neural Attention
abstract
A lack of corpora has so far limited advances in integrating human gaze data as a supervisory signal in neural attention mechanisms for natural language processing (NLP). We propose a novel hybrid text saliency model (TSM) that, for the first time, combines a cognitive model of reading with explicit human gaze supervision in a single machine learning framework. On four different corpora we demonstrate that our hybrid TSM duration predictions are highly correlated with human gaze ground truth. We further propose a novel joint modeling approach to integrate TSM predictions into the attention layer of a network designed for a specific upstream NLP task without the need for any task-specific human gaze data. We demonstrate that our joint model outperforms the state of the art in paraphrase generation on the Quora Question Pairs corpus by more than 10% in BLEU-4 and achieves state of the art performance for sentence compression on the challenging Google Sentence Compression corpus. As such, our work introduces a practical approach for bridging between data-driven and cognitive models and demonstrates a new way to integrate human gaze-guided neural attention into NLP tasks.
Ekta Sood, Simon Tannert, Philipp Müller 0001, Andreas Bulling
NeurIPS1
2018 Comparing Attention-Based Convolutional and Recurrent Neural Networks: Success and Limitations in Machine Reading Comprehension
abstract
We propose a machine reading comprehension model based on the compare-aggregate framework with two-staged attention that achieves state-of-the-art results on the MovieQA question answering dataset.To investigate the limitations of our model as well as the behavioral difference between convolutional and recurrent neural networks, we generate adversarial examples to confuse the model and compare to human performance.Furthermore, we assess the generalizability of our model by analyzing its differences to human inference, drawing upon insights from cognitive science.
Matthias Blohm, Glorianna Jagfeld, Ekta Sood, Ngoc Thang Vu
CoNLL3