Jakob Suchan

dblp:120/2899 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
5since 2021 · last 2026
0000-0002-1356-3297ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 How Do Naturalistic Visuo-Auditory Cues Guide Human Attention? Insights from Systematic Explorations in Visual Perception of Embodied Multimodal Interaction
abstract
Studies in visual cognition highlight the importance of visual, spatial, and auditory cues in influencing human attention. Such cues often tend to be indicative of actions or events, thereby serving as predictive indicators in both passive observation as well as in interactive engagement. Our research focuses on visual attention in passive observation, particularly examining the manner in which visual, spatial, and auditory cues—henceforth visuo-auditory (shorthand) cues—influence attention on everyday multimodal interaction. We systematically develop a visuo-auditory event model for investigating visual attention in naturalistic embodied settings. Rooted in this event model, we explore the influence of five select visuo-auditory cues—namely, speaking, gaze, relative motion, hand action, and visibility—on visual attention. Our analysis utilizes eye-tracking data from 90 participants observing 27 carefully designed naturalistic event scenarios and correlating their attentional metrics with the select visuo-auditory cues in the backdrop of the developed event model. Findings reveal strong associations between attention and both intra-modal (irrespective of other cues) and cross-modal (combined with other cues) cueing effects, thereby highlighting the nuanced interplay amongst the cues influencing attentional patterns. We develop a systematic and generalized method for analyzing interactions and behavioral parameters, thereby characterizing the impact of visuo-auditory cues on attentional dynamics. Our methodology, combined with the obtained insights into the attentional cueing effects, provides an analytical framework explicating the manner in which everyday (interactive) events directly drive attention under naturalistic conditions. This facilitates not only the precise modeling of behavior and attention allocation but also offers a high-level “experimental lens” for examining interactions in relation to behavioral parameters. Taken together, our methodological and behavioral findings are well-positioned to benefit multiple fields, particularly by advancing human-centered design across diverse application domains. Lastly, towards promoting open-science and for wider dissemination, the complete experimental basis of this research—e.g., event-scenarios, high-quality annotated data, dataset supplementary—has been documented together with instructions on how to use and access experimental data ( https://codesign-lab.org/cognitive-vision/multimodal-cues ).
Vipul Nair, Mehul Bhatt, Jakob Suchan, Erik Billing, Paul Hemeren
ACM Trans. Appl. Percept.3
2025 ASP-Driven Visual Commonsense: A General Framework for Reasoning About Embodied Interaction in the Wild
abstract
We present a general framework for declaratively grounded visual commonsense (reasoning) about embodied interaction in naturalistic, in-the-wild settings relevant to a range of AI application domains. The core computational capabilities of the framework pertaining visual commonsense are driven by a robust neurosymbolic architecture primarily consisting of: (1) answer set programming based modelling of foundational aspects pertaining spatio-temporal dynamics, encompassing space, time, events, action, motion; (2) modularly integrated visual computing techniques constituting the neural substrate linking quantitative perceptual features serving as low-level counterparts to high-level semantic characterisations of (inter)active visual commonsense. Practically, we also present a first open-release of the developed framework with the aim to promote independent extensions and real-world applied KRR. The release comprises: (a) demonstrated case-studies in domains such as autonomous driving, psychology and media studies; (b) systematic evaluation mechanisms for community benchmarking; and (c) supporting material such as tutorials and datasets.
Jakob Suchan, Mehul Bhatt, Julius Monsen
KR1
2025 Probabilistic Answer Set Programming Driven Ranking of Dynamic Space-Time Belief Models
Julius Monsen, Jakob Suchan, Mehul Bhatt
RuleML+RR2
2022 Attentional synchrony in films: A window to visuospatial characterization of events
abstract
The study of event perception emphasizes the importance of visuospatial attributes in everyday human activities and how they influence event segmentation, prediction and retrieval. Attending to these visuospatial attributes is the first step toward event understanding, and therefore correlating attentional measures to such attributes would help to further our understanding of event comprehension. In this study, we focus on attentional synchrony amongst other attentional measures and analyze select film scenes through the lens of a visuospatial event model. Here we present the first results of an in-depth multimodal (such as head-turn, hand-action etc.) visuospatial analysis of 10 movie scenes correlated with visual attention (eye-tracking 32 participants per scene). With the results, we tease apart event segments of high and low attentional synchrony and describe the distribution of attention in relation to the visuospatial features. This analysis gives us an indirect measure of attentional saliency for a scene with a particular visuospatial complexity, ultimately directing the attentional selection of the observers in a given context.
Vipul Nair, Jakob Suchan, Mehul Bhatt, Paul Hemeren
SAP2
2021 Commonsense visual sensemaking for autonomous driving - On generalised neurosymbolic online abduction integrating vision and semantics
abstract
We demonstrate the need and potential of systematically integrated vision and semantics solutions for visual sensemaking in the backdrop of autonomous driving. A general neurosymbolic method for online visual sensemaking using answer set programming (ASP) is systematically formalised and fully implemented. The method integrates state of the art in visual computing, and is developed as a modular framework that is generally usable within hybrid architectures for realtime perception and control. We evaluate and demonstrate with community established benchmarks KITTIMOD, MOT-2017, and MOT-2020. As use-case, we focus on the significance of human-centred visual sensemaking —e.g., involving semantic representation and explainability, question-answering, commonsense interpolation— in safety-critical autonomous driving situations. The developed neurosymbolic framework is domain-independent, with the case of autonomous driving designed to serve as an exemplar for online visual sensemaking in diverse cognitive interaction settings in the backdrop of select human-centred AI technology design considerations.
Jakob Suchan, Mehul Bhatt, Srikrishna Varadarajan
Artif. Intell.1
2020 Cognitive Vision and Perception
abstract
Semantic interpretation of dynamic visuospatial imagery calls for a general and systematic integration of methods in knowledge representation and computer vision. Towards this, we highlight research articulating & developing deep semantics, characterised by the existence of declarative models –e.g., pertaining space and motion– and corresponding formalisation and reasoning methods supporting capabilities such as semantic question-answering, relational visuospatial learning, and (non-monotonic) visuospatial explanation. We position a working model for deep semantics by highlighting select recent / closely related works from IJCAI [8, 4], AAAI [10], ILP [7], and ACS [9]. We posit that human-centred, explainable visual sensemaking necessitates both high-level semantics and low-level visual computing, with the highlighted works providing a model for systematic, modular integration of diverse multifaceted techniques developed in AI, ML, and Computer Vision.
Mehul Bhatt, Jakob Suchan
ECAI2
2020 Driven by Commonsense
Jakob Suchan, Mehul Bhatt, Srikrishna Varadarajan
ECAI1
2019 Out of Sight But Not Out of Mind: An Answer Set Programming Based Online Abduction Framework for Visual Sensemaking in Autonomous Driving
abstract
We demonstrate the need and potential of systematically integrated vision and semantics solutions for visual sensemaking (in the backdrop of autonomous driving). A general method for online visual sensemaking using answer set programming is systematically formalised and fully implemented. The method integrates state of the art in visual computing, and is developed as a modular framework usable within hybrid architectures for perception & control. We evaluate and demo with community established benchmarks KITTIMOD and MOT. As use-case, we focus on the significance of human-centred visual sensemaking ---e.g., semantic representation and explainability, question-answering, commonsense interpolation--- in safety-critical autonomous driving situations.
Jakob Suchan, Mehul Bhatt, Srikrishna Varadarajan
IJCAI1
2018 Visual Explanation by High-Level Abduction: On Answer-Set Programming Driven Reasoning About Moving Objects
abstract
We propose a hybrid architecture for systematically computing robust visual explanation(s) encompassing hypothesis formation, belief revision, and default reasoning with video data. The architecture consists of two tightly integrated synergistic components: (1) (functional) answer set programming based abductive reasoning with space-time tracklets as native entities; and (2) a visual processing pipeline for detection based object tracking and motion analysis. We present the formal framework, its general implementation as a (declarative) method in answer set programming, and an example application and evaluation based on two diverse video datasets: the MOTChallenge benchmark developed by the vision community, and a recently developed Movie Dataset.
Jakob Suchan, Mehul Bhatt, Przemyslaw Andrzej Walega, Carl P. L. Schultz
AAAI1
2016 Artificial Intelligence for Predictive and Evidence Based Architecture Design
abstract
The evidence-based analysis of people's navigation and wayfinding behaviour in large-scale built-up environments (e.g., hospitals, airports) encompasses the measurement and qualitative analysis of a range of aspects including people's visual perception in new and familiar surroundings, their decision-making procedures and intentions, the affordances of the environment itself, etc. In our research on large-scale evidence-based qualitative analysis of wayfinding behaviour, we construe visual perception and navigation in built-up environments as a dynamic narrative construction process of movement and exploration driven by situation-dependent goals, guided by visual aids such as signage and landmarks, and influenced by environmental (e.g., presence of other people, time of day, lighting) and personal (e.g., age, physical attributes) factors. We employ a range of sensors for measuring the embodied visuo-locomotive experience of building users: eye-tracking, egocentric gaze analysis, external camera based visual analysis to interpret fine-grained behaviour (e.g., stopping, looking around, interacting with other people), and also manual observations made by human experimenters. Observations are processed, analysed, and integrated in a holistic model of the visuo-locomotive narrative experience at the individual and group level. Our model also combines embodied visual perception analysis with analysis of the structure and layout of the environment (e.g., topology, routes, isovists) computed from available 3D models of the building. In this framework, abstract regions like the visibility space, regions of attention, eye movement clusters, are treated as first class visuo-spatial and iconic objects that can be used for interpreting the visual experience of subjects in a high-level qualitative manner. The final integrated analysis of the wayfinding experience is such that it can even be presented in a virtual reality environment thereby providing an immersive experience (e.g., using tools such as the Oculus Rift) of the qualitative analysis for single participants, as well as for a combined analysis of large group. This capability is especially important for experiments in post-occupancy analysis of building performance. Our construction of indoor wayfinding experience as a form of moving image analysis centralizes the role and influence of perceptual visuo-spatial characteristics and morphological features of the built environment into the discourse on wayfinding research. We will demonstrate the impact of this work with several case-studies, particularly focussing on a large-scale experiment conducted at the New Parkland Hospital in Dallas Texas, USA.
Mehul Bhatt, Jakob Suchan, Carl P. L. Schultz, Vasiliki Kondyli, Saurabh Goyal
AAAI2
2016 Embodied visuo-locomotive experience analysis: immersive reality based summarisation of experiments in environment-behaviour studies
abstract
Evidence-based design (EBD) for architecture involves the study of post-occupancy behaviour of building users with the aim to provide an empirical basis for improving building performance [Hamilton and Watkins 2009]. Within EBD, the high-level, qualitative analysis of the embodied visuo-locomotive experience of representative groups of building users (e.g., children, senior citizens, individuals facing physical challenges) constitutes a foundational approach for understanding the impact of architectural design decisions, and functional building performance from the viewpoint of areas such as environmental psychology, wayfinding research, human visual perception studies, spatial cognition, and the built environment [Bhatt and Schultz 2016].
Mehul Bhatt, Jakob Suchan, Vasiliki Kondyli, Carl P. L. Schultz
SAP2
2016 The perception of symmetry in the moving image: multi-level computational analysis of cinematographic scene structure and its visual reception
abstract
This research is driven by visuo-spatial perception focussed cognitive film studies, where the key emphasis is on the systematic study and generation of evidence that can characterise and establish correlates between principles for the synthesis of the moving image, and its cognitive (e.g., embodied visuo-auditory, emotional) recipient effects on observers [Suchan and Bhatt 2016b; Suchan and Bhatt 2016a]. Within this context, we focus on the case of "symmetry" in the cinematographic structure of the moving image, and propose a multi-level model of interpreting symmetric patterns therefrom. This provides the foundation for integrating scene analysis with the analysis of its visuo-spatial perception based on eye-tracking data. This is achieved by the integration of: computational semantic interpretation of the scene [Suchan and Bhatt 2016b] ---involving scene objects (people, objects in the scene), cinematographic aids (camera movement, shot types, cuts and scene structure)--- and perceptual artefacts (fixations, saccades, scan-path, areas of attention).
Jakob Suchan, Mehul Bhatt, Stella X. Yu
SAP1
2016 Robust Natural Language Processing - Combining Reasoning, Cognitive Semantics, and Construction Grammar for Spatial Language
Michael Spranger, Jakob Suchan, Mehul Bhatt
IJCAI2
2016 Semantic Question-Answering with Video and Eye-Tracking Data: AI Foundations for Human Visual Perception Driven Cognitive Film Studies
Jakob Suchan, Mehul Bhatt
IJCAI1
2016 The geometry of a scene: On deep semantics for visual perception driven cognitive film, studies
abstract
We present a general computational narrative model encompassing primitives of space, time, and motion from the viewpoint of deep knowledge representation and reasoning about visuo-spatial dynamics, and (eye-tracking based) visual perception of the moving image. The declarative model, implemented within constraint logic programming, integrates knowledge-based qualitative reasoning (e.g., about object / character placement, scene structure) with state of the art computer vision methods for detecting, tracking, and recognition of people, objects, and cinematographic devices such as cuts, shot types, types of camera movement. A key feature is that primitives of the theory - things, time, space and motion predicates, actions and events, perceptual objects (e.g., eye-tracking / gaze points, regions of attention etc) - are available as first-class objects with deep semantics suited for inference and query from the viewpoint of analytical Q&A or studies in visual perception. We present the formal framework and its implementation in the context of a large-scale experiment concerned with analysis of visual perception and reception of the moving image in the context of cognitive film studies.
Jakob Suchan, Mehul Bhatt
WACV1
2014 Grounding Dynamic Spatial Relations for Embodied (Robot) Interaction
Michael Spranger, Jakob Suchan, Mehul Bhatt, Manfred Eppe
PRICAI2