EDBT 2026 Demo / reviewers in the wild / expert
Sounak Mondal
dblp:207/4785
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0009-9802-4652ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Image recognition and object detection · 52% 3D vision · 15% Face, body and person analysis · 15% | |
| Human-computer interaction and pervasive computing
5 papers |
Wearable and physiological sensing · 66% Usability and user experience research · 17% Human-AI interaction · 12% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
saliency prediction |
1.8 | 3 | 2025 | Few-shot Personalized Scanpath Prediction · CVPR 2025 Unifying Top-Down and Bottom-Up Scanpath Prediction Using Transformers · CVPR 2024 Target-Absent Human Attention · ECCV (4) 2022 |
Wearable and physiological sensing › gaze prediction
scanpath prediction |
1.5 | 2 | 2025 | Few-shot Personalized Scanpath Prediction · CVPR 2025 Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human Attention · CVPR 2023 |
Computer vision › Image recognition and object detection
visual search |
0.9 | 1 | 2025 | Gaze-Language Alignment for Zero-Shot Prediction of Visual Search Targets from Human Gaze Scanpaths · ICCV 2025 |
Wearable and physiological sensing › eye tracking
gaze-based interaction |
0.9 | 1 | 2025 | Gaze-Language Alignment for Zero-Shot Prediction of Visual Search Targets from Human Gaze Scanpaths · ICCV 2025 |
Usability and user experience research
user modeling |
0.9 | 1 | 2025 | Few-shot Personalized Scanpath Prediction · CVPR 2025 |
Computer vision › Face, body and person analysis › gaze analysis
gaze following |
0.8 | 1 | 2024 | Diffusion-Refined VQA Annotations for Semi-supervised Gaze Following · ECCV (39) 2024 |
Computer vision › Video understanding and tracking
gaze prediction |
0.8 | 1 | 2024 | Look Hear: Gaze Prediction for Speech-Directed Human Attention · ECCV (42) 2024 |
Computer vision › 3D vision › biological vision modeling
scanpath prediction |
0.8 | 1 | 2024 | Unifying Top-Down and Bottom-Up Scanpath Prediction Using Transformers · CVPR 2024 |
Wearable and physiological sensing
gaze prediction |
0.7 | 1 | 2023 | Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human Attention · CVPR 2023 |
Human-AI interaction
human attention |
0.6 | 1 | 2022 | Target-Absent Human Attention · ECCV (4) 2022 |
Wearable and physiological sensing
eye tracking |
0.3 | 1 | 2025 | Gaze-Language Alignment for Zero-Shot Prediction of Visual Search Targets from Human Gaze Scanpaths · ICCV 2025 |
Machine learning › Generative modeling
diffusion model |
0.2 | 1 | 2024 | Diffusion-Refined VQA Annotations for Semi-supervised Gaze Following · ECCV (39) 2024 |
Methods — techniques the papers use, named apart from their topics
vision-language alignment · 1.7subject embedding network · 1.7gaze scanpath modeling · 1.7few-shot learning · 1.7multimodal learning · 1.5transformer · 1.4semi-supervised learning · 0.8foveated retina · 0.8diffusion model · 0.8dense heatmap prediction · 0.8natural language model · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Few-shot Personalized Scanpath PredictionabstractA personalized model for scanpath prediction provides insights into the visual preferences and attention patterns of individual subjects. However, existing methods for training scanpath prediction models are data-intensive and cannot be effectively personalized to new individuals with only a few available examples. In this paper, we propose few-shot personalized scanpath prediction task (FS-PSP) and a novel method to address it, which aims to predict scan-paths for an unseen subject using minimal support data of that subject’s scanpath behavior. The key to our method’s adaptability is the Subject-Embedding Network (SE-Net), specifically designed to capture unique, individualized representations for each subject’s scanpaths. SE-Net generates subject embeddings that effectively distinguish between subjects while minimizing variability among scanpaths from the same individual. The personalized scanpath prediction model is then conditioned on these subject embeddings to produce accurate, personalized results. Experiments on multiple eye-tracking datasets demonstrate that our method excels in FS-PSP settings and does not require any fine-tuning steps at test time. Code is available at: https://github.com/cvlab-stonybrook/few-shot-scanpath Ruoyu Xue, Sounak Mondal, Hieu Le 0001, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras |
CVPR | 3 |
| 2025 | Gaze-Language Alignment for Zero-Shot Prediction of Visual Search Targets from Human Gaze Scanpaths
Sounak Mondal, Naveen Sendhilnathan, Ting Zhang 0013, Michael Proulx, Michael L. Iuzzolino, Tanya R. Jonker |
ICCV | 1 |
| 2024 | Unifying Top-Down and Bottom-Up Scanpath Prediction Using TransformersabstractMost models of visual attention aim at predicting either top-down or bottom-up control, as studied using different visual search and free-viewing tasks. In this paper we propose the Human Attention Transformer (HAT), a single model that predicts both forms of attention control. HAT uses a novel transformer-based architecture and a simplified foveated retina that collectively create a spatio-temporal awareness akin to the dynamic visual working memory of humans. HAT not only establishes a new state-of-the-art in predicting the scanpath of fixations made during target-present and target-absent visual search and "taskless" free viewing, but also makes human gaze behavior interpretable. Unlike previous methods that rely on a coarse grid of fixation cells and experience information loss due to fixation discretization, HAT features a sequential dense prediction architecture and outputs a dense heatmap for each fixation, thus avoiding discretizing fixations. HAT sets a new standard in computational attention, which emphasizes effectiveness, generality, and interpretability. HAT's demonstrated scope and applicability will likely inspire the development of new attention models that can better predict human behavior in various attention-demanding scenarios. Code is available at https://github.com/cvlab-stonybrook/HAT. Zhibo Yang 0002, Sounak Mondal, Seoyoung Ahn, Ruoyu Xue, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras |
CVPR | 2 |
| 2024 | Diffusion-Refined VQA Annotations for Semi-supervised Gaze Following
Qiaomu Miao, Alexandros Graikos, Sounak Mondal, Minh Hoai, Dimitris Samaras |
ECCV (39) | 4 |
| 2024 | Look Hear: Gaze Prediction for Speech-Directed Human Attention
Sounak Mondal, Seoyoung Ahn, Zhibo Yang 0002, Niranjan Balasubramanian, Dimitris Samaras, Gregory J. Zelinsky, Minh Hoai |
ECCV (42) | 1 |
| 2023 | Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human AttentionabstractPredicting human gaze is important in Human-Computer Interaction (HCI). However, to practically serve HCI applications, gaze prediction models must be scalable, fast, and accurate in their spatial and temporal gaze predictions. Recent scanpath prediction models focus on goaldirected attention (search). Such models are limited in their application due to a common approach relying on trained target detectors for all possible objects, and the availability of human gaze data for their training (both not scalable). In response, we pose a new task called ZeroGaze, a new variant of zero-shot learning where gaze is predicted for never-before-searched objects, and we develop a novel model, Gazeformer; to solve the ZeroGaze problem. In contrast to existing methods using object detector modules, Gazeformer encodes the target using a natural language model, thus leveraging semantic similarities in scanpath prediction. We use a transformer-based encoder-decoder architecture because transformers are particularly useful for generating contextual representations. Gazeformer surpasses other models by a large margin (19%-70%) on the ZeroGaze setting. It also outperforms existing target-detection models on standard gaze prediction for both target-present and target-absent search tasks. In addition to its improved performance, Gazeformer is more than five times faster than the state-of-the-art target-present visual search model. Code can be found at https://github.com/cvlab-stonybrook/Gazeformer/ Sounak Mondal, Zhibo Yang 0002, Seoyoung Ahn, Dimitris Samaras, Gregory J. Zelinsky, Minh Hoai |
CVPR | 1 |
| 2022 | Target-Absent Human Attention
Zhibo Yang 0002, Sounak Mondal, Seoyoung Ahn, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras |
ECCV (4) | 2 |