VLDB 2026 Research / reviewers in the wild / expert
David R. Reich
dblp:321/1783 · also David Robert Reich
· DBLP profile ↗
22ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0002-3524-3788ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 15 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fixation Sequences as Time Series: A Topological Approach to Dyslexia DetectionabstractPersistent homology, a method from topological data analysis, extracts robust, multi-scale features from data. It produces stable representations of time series by applying varying thresholds to their values (a process known as a filtration). We develop novel filtrations for time series and introduce topological methods for the analysis of eye-tracking data, by interpreting fixation sequences as time series, and constructing “hybrid models” that combine topological features with traditional statistical features. We empirically evaluate our method by applying it to the task of dyslexia detection from eye-tracking-while-reading data using the Copenhagen Corpus, which contains scanpaths from dyslexic and non-dyslexic L1 and L2 readers. Our hybrid models outperform existing approaches that rely solely on traditional features, showing that persistent homology captures complementary information encoded in fixation sequences. The strength of these topological features is further underscored by their achieving performance comparable to established baseline methods. Importantly, our proposed filtrations outperform existing ones. Marius Huber, David R. Reich, Lena A. Jäger |
ETRA | 2 |
| 2025 | CoLAGaze: A Corpus of Eye Movements for Linguistic AcceptabilityabstractWe present CoLAGaze, the first broad-coverage eye-tracking-while-reading corpus on grammatical and ungrammatical sentences sourced from CoLA — a Natural Language Processing (NLP) benchmark for evaluating the grammatical knowledge of language models (LMs). CoLAGaze provides eye-tracking data from native English speakers in different formats including the raw eye-tracking signal, gaze event data, and reading measures computed at the character, word, and sentence levels alongside comprehensive meta-data and data quality documentation. CoLAGaze enables psycholinguistic research on the processing of diverse (un)grammatical structures, allows the training of generative models of eye-movements-in-reading capable of generalizing to ungrammatical stimuli, facilitates the alignment of LMs to human language processing, and supports gaze-augmented NLP applications for grammatical error detection. CoLAGaze and the preprocessing code, is available at OSF and GitHub. We have also integrated it into the pymovements Python package. Anna Bondar, David R. Reich, Lena A. Jäger |
ETRA | 2 |
| 2025 | Neural Additive Models Uncover Predictive Gaze Features in Reading
Deborah N. Jakobi, David R. Reich, Paul Prasse, Lena A. Jäger |
ETRA | 2 |
| 2025 | MultiplEYE: Creating a multilingual eye-tracking-while-reading corpusabstractContains fulltext : 326363.pdf (Publisher’s version ) (Open Access) Deborah N. Jakobi, Maja Stegenwallner-Schütz, Nora Hollenstein, Cui Ding, Ramune Kaspere, Ana Matic Skoric, Eva Pavlinusic Vilus, Stefan Frank, Marie-Luise Müller, Kristine M. Jensen de López, Nik Kharlamov, Hanne B. Søndergaard Knudsen, Yevgeni Berzak, Ella Lion, Irina A. Sekerina, Cengiz Acartürk, Mohd Faizan Ansari, Katarzyna Harezlak, Pawel Kasprowski, Ana Bautista, Lisa Beinborn, Anna Bondar, Antonia Boznou, Leah Bradshaw, Jana Mara Hofmann, Thyra Krosness, Not Battesta Soliva, Anila Çepani, Kristina Cergol, Ana Dosen, Marijan Palmovic, Adelina Çerpja, Dalí Chirino, Jan Chromý, Vera Demberg, Iza Skrjanec, Nazik Dinçtopal Deniz, Inmaculada Fajardo, Mariola Giménez-Salvador, Xavier Mínguez-López, Maros Filip, Zigmunds Freibergs, Jessica Gomes, Andreia Janeiro, Paula Luegi, João Veríssimo, Sasho Gramatikov, Jana Hasenäcker, Alba Haveriku, Nelda Kote, Muhammad Mohsin Kamal, Hanna Kedzierska, Dorota Klimek-Jankowska, Sara Kosutar, Daniel Krakowczyk, Izabela Krejtz, Marta Lockiewicz, Kaidi Lõo, Jurgita Motiejuniene, Jamal Abdul Nasir, Johanne Sofie Krog Nedergård, Aysegül Özkan, Mikulás Preininger, Loredana Punga, David R. Reich, Chiara Tschirner, Spela Rot, Andreas Säuberli, Jordi Solé i Casals, Ekaterina Strati, Igor Svoboda, Evis Trandafili, Spyridoula Varlokosta, Mila Dimitrova-Vulchanova, Lena A. Jäger |
ETRA | 65 |
| 2025 | The More the Merrier: Boost Your Dataset Visibility and Discover Eye-Tracking Datasets with pymovements
Daniel Krakowczyk, David R. Reich, Andreas Säuberli, Iza Skrjanec, Isabelle Caroline Rose Cretton, Deborah N. Jakobi, Sergiu Nisioi, Paul Prasse, Lena A. Jäger |
ETRA | 2 |
| 2025 | Detection of Alcohol Inebriation from Eye Movements using Remote and Wearable Eye TrackersabstractThis OSF contains the data for the paper 'Detection of Alcohol Inebriation from Eye Movements using Remote and Wearable Eye Trackers'. The raw data can be found in the folder raw_data (each zip file contains recorded data up- /downsampled to 1,000 Hz as csv-files). - Each csv file contains the recording (remote and wearable) for one subject for one PVT trial. - Each csv file contains the following columns: trial_id: trial-id for current recording block_id: block-id for current recording x_pix_eyelink: x-pixel coordinates using eyelink remote eye-tracker y_pix_eyelink: y-pixel coordinates using eyelink remote eye-tracker eyelink_timestamp: timestamp or recording in ms x_pix_pupilcore_interpolated: x-pixel coordinates using pupil-core eye-tracker upsampled to 1,000 Hz y_pix_pupilcore_interpolated: y-pixel coordinates using pupil-core eye-tracker upsampled to 1,000 Hz pupil_size_eyelink: pupil-size of pupil using eyelink remote eye-tracker target_distance: distance to eyelink remote eye-tracker (screen) in mm pupil_size_pupilcore_interpolated: pupil-size of pupil pupil-core eye-tracker upsampled to 1,000 Hz pupil_confidence_interpolated: pupil detection confidence of pupil pupil-core eye-tracker upsampled to 1,000 Hz time_to_prev_bac: elapsed time from previous BAC testing in ms time_to_next_bac: remaining time for next BAC testing in ms prev_bac: previous BAC concentration next_bac: next BAC concentration For more details see: https://github.com/aeye-lab/etra-potsdam-binge-pvt Paul Prasse, David R. Reich, Jakob Chwastek, Silvia Makowski, Lena A. Jäger, Tobias Scheffer |
ETRA | 2 |
| 2025 | Proxy-Based Pre-Training for Eye-Tracking Applications
David R. Reich, Cui Ding, Lena S. Bolliger, Patrick Haller 0001, Paul Prasse, Lena A. Jäger |
ETRA | 1 |
| 2025 | EyeBench: Predictive Modeling from Eye Movements in ReadingabstractWe present EyeBench, the first benchmark designed to evaluate machine learning models that decode cognitive and linguistic information from eye movements during reading. EyeBench offers an accessible entry point to the challenging and underexplored domain of modeling eye tracking data paired with text, aiming to foster innovation at the intersection of multimodal AI and cognitive science. The benchmark provides a standardized evaluation framework for predictive models, covering a diverse set of datasets and tasks, ranging from assessment of reading comprehension to detection of developmental dyslexia. Progress on the EyeBench challenge will pave the way for both practical real-world applications, such as adaptive user interfaces and personalized education, and scientific advances in understanding human language processing. The benchmark is released as an open-source software package which includes data downloading and harmonization scripts, baselines and state-of-the-art models, as well as evaluation code, publicly available at https://github.com/EyeBench/eyebench. Omer Shubi, David R. Reich, Keren Gruteke Klein, Yuval Angel, Paul Prasse, Lena A. Jäger, Yevgeni Berzak |
NeurIPS | 2 |
| 2025 | ScanDL 2.0: A Generative Model of Eye Movements in Reading Synthesizing Scanpaths and Fixation DurationsabstractEye movements in reading have become a vital tool for investigating the cognitive mechanisms involved in language processing. They are not only used within psycholinguistics but have also been leveraged within the field of NLP to improve the performance of language models on downstream tasks. However, the scarcity of real eye-tracking data and its limited generalizability at inference time present challenges for data-driven approaches. In response, synthetic scanpaths have emerged as a promising alternative. Despite advances, however, existing machine learning-based methods, including the state-of-the-art ScanDL [9], fail to incorporate fixation durations into the generated scanpaths, which are crucial for a complete representation of reading behavior. We therefore propose a novel model, denoted ScanDL 2.0, which synthesizes both fixation locations and durations. It sets a new benchmark in generating human-like synthetic scanpaths, demonstrating superior performance across various evaluation settings. Furthermore, psycholinguistic analyses confirm its ability to emulate key phenomena in human reading. Our code as well as pre-trained model weights are available via https://github.com/DiLi-Lab/ScanDL-2.0. Lena S. Bolliger, David R. Reich, Lena A. Jäger |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2025 | Evaluating Gaze Event Detection Algorithms: Impacts on Machine Learning-based Classification and Psycholinguistic Statistical ModelingabstractEye movements offer valuable, non-invasive insights into cognitive processes and are widely used in both psycholinguistic research and machine-learning applications, such as assessing reading comprehension and cognitive load. These applications typically rely on fixations and saccades detected through gaze event algorithms, which may be either proprietary or open-source. The impact of different gaze event detection algorithms on subsequent analysis is underexplored and often overlooked. This study investigates how two threshold-based algorithms, I-DT and I-VT, influence both machine-learning classification tasks and psycholinguistic statistical modeling. Using diverse datasets-including stationary, remote, and VR eye-tracking data across multiple sampling frequencies-our findings show significant differences in downstream performance. For ML tasks, I-DT generally outperforms I-VT, with I-VT being highly sensitive to threshold choices. In psycholinguistic analysis, results confirm established findings only when thresholds align with established fixation metrics, emphasizing the importance of appropriate threshold selection for meaningful analysis. Our code is publicly available: https://github.com/aeye-lab/eye-movement-preprocessing. David R. Reich, Paul Prasse, Lena A. Jäger |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2024 | Reading Does Not Equal Reading: Comparing, Simulating and Exploiting Reading Behavior across PopulationsabstractEye-tracking-while-reading corpora play a crucial role in the study of human language processing, and, more recently, have been leveraged for cognitively enhancing neural language models. A critical limitation of existing corpora is that they often lack diversity, comprising primarily native speakers. In this study, we expand the eye-tracking-while-reading dataset CopCo, which initially included only Danish L1 readers with and without dyslexia, by incorporating a new dataset of L2 readers with diverse L1 backgrounds. Thus, the extended CopCo corpus constitutes the first eye-tracking-while-reading dataset encompassing neurotypical L1 and L1 readers with dyslexia as well as L2 readers, all reading the same materials. We first provide extensive descriptive statistics of the extended CopCo corpus. Second, we investigate how different degrees of diversity of the training data affect a state-of-the-art generative model of eye movements in reading. Finally, we use this scanpath generation model for gaze-augmented language modeling and investigate the impact of diversity in the training data on the model’s performance on a range of NLP downstream tasks. The code can be found here: https://github.com/norahollenstein/copco-processing. David R. Reich, Shuwen Deng, Marina Björnsdóttir, Lena A. Jäger, Nora Hollenstein |
LREC/COLING | 1 |
| 2024 | Reverse-Engineering the ReaderabstractNumerous previous studies have sought to determine to what extent language models, pretrained on natural language text, can serve as useful models of human cognition.In this paper, we are interested in the opposite question: whether we can directly optimize a language model to be a useful cognitive model by aligning it to human psychometric data.To achieve this, we introduce a novel alignment technique in which we fine-tune a language model to implicitly optimize the parameters of a linear regressor that directly predicts humans' reading times of in-context linguistic units, e.g., phonemes, morphemes, or words, using surprisal estimates derived from the language model.Using words as a test case, we evaluate our technique across multiple model sizes and datasets and find that it improves language models' psychometric predictive power.However, we find an inverse relationship between psychometric power and a model's performance on downstream NLP tasks as well as its perplexity on held-out test data.While this latter trend has been observed before (Oh et al., 2022;Shain et al., 2024), we are the first to induce it by manipulating a model's alignment to psychometric data. Samuel Kiegeland, Ethan Wilcox, Afra Amini, David R. Reich, Ryan Cotterell |
EMNLP | 4 |
| 2024 | Improving cognitive-state analysis from eye gaze with synthetic eye-movement dataabstractEye movements can be used to analyze a viewer’s cognitive capacities or mental state. Neural networks that process the raw eye-tracking signal can outperform methods that operate on scan paths preprocessed into fixations and saccades. However, the scarcity of such data poses a major challenge. We therefore develop SP-EyeGAN, a neural network that generates synthetic raw eye-tracking data. SP-EyeGAN consists of Generative Adversarial Networks; it produces a sequence of gaze angles indistinguishable from human ocular micro- and macro-movements. We explore the use of these synthetic eye movements for pre-training neural networks using contrastive learning. We find that pre-training on synthetic data does not help for biometric identification, while results are inconclusive for the detection of ADHD and gender classification. However, for the eye movement-based assessment of higher-level cognitive skills such general reading comprehension, text comprehension, and the distinction of native from non-native readers, pre-training on synthetic eye-gaze data improves the models’ performance and even advances the state-of-the-art for reading comprehension. The SP-EyeGAN model, pre-trained on GazeBase, along with the code for developing your own raw eye-tracking machine learning model with contrastive learning, is available at https://github.com/aeye-lab/sp-eyegan. Paul Prasse, David R. Reich, Silvia Makowski, Tobias Scheffer, Lena A. Jäger |
Comput. Graph. | 2 |
| 2023 | ScanDL: A Diffusion Model for Generating Synthetic Scanpaths on TextsabstractEye movements in reading play a crucial role in psycholinguistic research studying the cognitive mechanisms underlying human language processing.More recently, the tight coupling between eye movements and cognition has also been leveraged for language-related machine learning tasks such as the interpretability, enhancement, and pre-training of language models, as well as the inference of reader-and text-specific properties.However, scarcity of eye movement data and its unavailability at application time poses a major challenge for this line of research.Initially, this problem was tackled by resorting to cognitive models for synthesizing eye movement data.However, for the sole purpose of generating humanlike scanpaths, purely data-driven machinelearning-based methods have proven to be more suitable.Following recent advances in adapting diffusion processes to discrete data, we propose SCANDL, a novel discrete sequence-tosequence diffusion model that generates synthetic scanpaths on texts.By leveraging pretrained word representations and jointly embedding both the stimulus text and the fixation sequence, our model captures multi-modal interactions between the two inputs.We evaluate SCANDL within-and across-dataset and demonstrate that it significantly outperforms state-of-the-art scanpath generation methods.Finally, we provide an extensive psycholinguistic analysis that underlines the model's ability to exhibit human-like reading behavior.Our implementation is made available at https://github.com/DiLi-Lab/ScanDL. Lena S. Bolliger, David R. Reich, Patrick Haller 0001, Deborah N. Jakobi, Paul Prasse, Lena A. Jäger |
EMNLP | 2 |
| 2023 | Pre-Trained Language Models Augmented with Synthetic Scanpaths for Natural Language UnderstandingabstractHuman gaze data offer cognitive information that reflects natural language comprehension.Indeed, augmenting language models with human scanpaths has proven beneficial for a range of NLP tasks, including language understanding.However, the applicability of this approach is hampered because the abundance of text corpora is contrasted by a scarcity of gaze data.Although models for the generation of humanlike scanpaths during reading have been developed, the potential of synthetic gaze data across NLP tasks remains largely unexplored.We develop a model that integrates synthetic scanpath generation with a scanpath-augmented language model, eliminating the need for human gaze data.Since the model's error gradient can be propagated throughout all parts of the model, the scanpath generator can be fine-tuned to downstream tasks.We find that the proposed model not only outperforms the underlying language model, but achieves a performance that is comparable to a language model augmented with real human gaze data.Our code is publicly available.1 Shuwen Deng, Paul Prasse, David R. Reich, Tobias Scheffer, Lena A. Jäger |
EMNLP | 3 |
| 2023 | Bridging the Gap: Gaze Events as Interpretable Concepts to Explain Deep Neural Sequence ModelsabstractRecent work in XAI for eye tracking data has evaluated the suitability of feature attribution methods to explain the output of deep neural sequence models for the task of oculomotric biometric identification. These methods provide saliency maps to highlight important input features of a specific eye gaze sequence. However, to date, its localization analysis has been lacking a quantitative approach across entire datasets. In this work, we employ established gaze event detection algorithms for fixations and saccades and quantitatively evaluate the impact of these events by determining their concept influence. Input features that belong to saccades are shown to be substantially more important than features that belong to fixations. By dissecting saccade events into sub-events, we are able to show that gaze samples that are close to the saccadic peak velocity are most influential. We further investigate the effect of event properties like saccadic amplitude or fixational dispersion on the resulting concept influence. Daniel Krakowczyk, Paul Prasse, David R. Reich, Sebastian Lapuschkin, Tobias Scheffer, Lena A. Jäger |
ETRA | 3 |
| 2023 | pymovements: A Python Package for Eye Movement Data ProcessingabstractWe introduce pymovements: a Python package for analyzing eye-tracking data that follows best practices in software development, including rigorous testing and adherence to coding standards. The package provides functionality for key processes along the entire preprocessing pipeline. This includes parsing of eye tracker data files, transforming positional data into velocity data, detecting gaze events like saccades and fixations, computing event properties like saccade amplitude and fixational dispersion and visualizing data and results with several types of plotting methods. Moreover, pymovements also provides an easily accessible interface for downloading and processing publicly available datasets. Additionally, we emphasize how rigorous testing in scientific software packages is critical to the reproducibility and transparency of research, enabling other researchers to verify and build upon previous findings. Daniel Krakowczyk, David R. Reich, Jakob Chwastek, Deborah N. Jakobi, Paul Prasse, Assunta Süss, Oleksii Turuta, Pawel Kasprowski, Lena A. Jäger |
ETRA | 2 |
| 2023 | SP-EyeGAN: Generating Synthetic Eye Movement Data with Generative Adversarial NetworksabstractNeural networks that process the raw eye-tracking signal can outperform traditional methods that operate on scanpaths preprocessed into fixations and saccades. However, the scarcity of such data poses a major challenge. We, therefore, present SP-EyeGAN, a neural network that generates synthetic raw eye-tracking data. SP-EyeGAN consists of Generative Adversarial Networks; it produces a sequence of gaze angles indistinguishable from human micro- and macro-movements. We demonstrate how the generated synthetic data can be used to pre-train a model using contrastive learning. This model is fine-tuned on labeled human data for the task of interest. We show that for the task of predicting reading comprehension from eye movements, this approach outperforms the previous state-of-the-art. Paul Prasse, David R. Reich, Silvia Makowski, Seoyoung Ahn, Tobias Scheffer, Lena A. Jäger |
ETRA | 2 |
| 2023 | Eyettention: An Attention-based Dual-Sequence Model for Predicting Human Scanpaths during ReadingabstractEye movements during reading offer insights into both the reader's cognitive processes and the characteristics of the text that is being read. Hence, the analysis of scanpaths in reading have attracted increasing attention across fields, ranging from cognitive science over linguistics to computer science. In particular, eye-tracking-while-reading data has been argued to bear the potential to make machine-learning-based language models exhibit a more human-like linguistic behavior. However, one of the main challenges in modeling human scanpaths in reading is their dual-sequence nature: the words are ordered following the grammatical rules of the language, whereas the fixations are chronologically ordered. As humans do not strictly read from left-to-right, but rather skip or refixate words and regress to previous words, the alignment of the linguistic and the temporal sequence is non-trivial. In this paper, we develop Eyettention, the first dual-sequence model that simultaneously processes the sequence of words and the chronological sequence of fixations. The alignment of the two sequences is achieved by a cross-sequence attention mechanism. We show that Eyettention outperforms state-of-the-art models in predicting scanpaths. We provide an extensive within- and across-data set evaluation on different languages. An ablation study and qualitative analysis support an in-depth understanding of the model's behavior. Shuwen Deng, David R. Reich, Paul Prasse, Patrick Haller 0001, Tobias Scheffer, Lena A. Jäger |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2022 | Fairness in Oculomotoric Biometric IdentificationabstractGaze patterns are known to be highly individual, and therefore eye movements can serve as a biometric characteristic. We explore aspects of the fairness of biometric identification based on gaze patterns. We find that while oculomotoric identification does not favor any particular gender and does not significantly favor by age range, it is unfair with respect to ethnicity. Moreover, fairness concerning ethnicity cannot be achieved by balancing the training data for the best-performing model. Paul Prasse, David R. Reich, Silvia Makowski, Lena A. Jäger, Tobias Scheffer |
ETRA | 2 |
| 2022 | Inferring Native and Non-Native Human Reading Comprehension and Subjective Text Difficulty from Scanpaths in ReadingabstractEye movements in reading are known to reflect cognitive processes involved in reading comprehension at all linguistic levels, from the sub-lexical to the discourse level. This means that reading comprehension and other properties of the text and/or the reader should be possible to infer from eye movements. Consequently, we develop the first neural sequence architecture for this type of tasks which models scan paths in reading and incorporates lexical, semantic and other linguistic features of the stimulus text. Our proposed model outperforms state-of-the-art models in various tasks. These include inferring reading comprehension or text difficulty, and assessing whether the reader is a native speaker of the text’s language. We further conduct an ablation study to investigate the impact of each component of our proposed neural network on its performance. David R. Reich, Paul Prasse, Chiara Tschirner, Patrick Haller 0001, Frank Goldhammer, Lena A. Jäger |
ETRA | 1 |
| 2022 | Detection of ADHD Based on Eye Movements During Natural Viewing
Shuwen Deng, Paul Prasse, David R. Reich, Sabine Dziemian, Maja Stegenwallner-Schütz, Daniel Krakowczyk, Silvia Makowski, Nicolas Langer, Tobias Scheffer, Lena A. Jäger |
ECML/PKDD (6) | 3 |