Constantin A. Rothkopf

dblp:71/5555 · DBLP profile ↗
← Back
32ranked-venue papers
2as first author
19since 2021 · last 2025
0000-0002-5636-0801ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 1 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Physical reasoning during motor learning aids people in transferring mass, but not motor control mappings
Fabian Tatai, Dominik Ürüm, Maria K. Eckstein, Constantin A. Rothkopf
CogSci4
2025 Inverse decision-making using neural amortized Bayesian actors
abstract
Bayesian observer and actor models have provided normative explanations for many behavioral phenomena in perception, sensorimotor control, and other areas of cognitive science and neuroscience. They attribute behavioral variability and biases to interpretable entities such as perceptual and motor uncertainty, prior beliefs, and behavioral costs. However, when extending these models to more naturalistic tasks with continuous actions, solving the Bayesian decision-making problem is often analytically intractable. Inverse decision-making, i.e. performing inference over the parameters of such models given behavioral data, is computationally even more difficult. Therefore, researchers typically constrain their models to easily tractable components, such as Gaussian distributions or quadratic cost functions, or resort to numerical approximations. To overcome these limitations, we amortize the Bayesian actor using a neural network trained on a wide range of parameter settings in an unsupervised fashion. Using the pre-trained neural network enables performing efficient gradient-based Bayesian inference of the Bayesian actor model's parameters. We show on synthetic data that the inferred posterior distributions are in close alignment with those obtained using analytical solutions where they exist. Where no analytical solution is available, we recover posterior distributions close to the ground truth. We then show how our method allows for principled model comparison and how it can be used to disentangle factors that may lead to unidentifiabilities between priors and costs. Finally, we apply our method to empirical data from three sensorimotor tasks and compare model fits with different cost functions to show that it can explain individuals' behavioral patterns.
Dominik Straub, Tobias F. Niehues, Jan Peters 0001, Constantin A. Rothkopf
ICLR4
2025 Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?
abstract
Recently, newly developed Vision-Language Models (VLMs), such as OpenAI’s o1, have emerged, seemingly demonstrating advanced reasoning capabilities across text and image modalities. However, the depth of these advances in language-guided perception and abstract reasoning remains underexplored, and it is unclear whether these models can truly live up to their ambitious promises. To assess the progress and identify shortcomings, we enter the wonderland of Bongard problems, a set of classic visual reasoning puzzles that require human-like abilities of pattern recognition and abstract reasoning. With our extensive evaluation setup, we show that while VLMs occasionally succeed in identifying discriminative concepts and solving some of the problems, they frequently falter. Surprisingly, even elementary concepts that may seem trivial to humans, such as simple spirals, pose significant challenges. Moreover, when explicitly asked to recognize ground truth concepts, they continue to falter, suggesting not only a lack of understanding of these elementary visual concepts but also an inability to generalize to unseen concepts. We compare the results of VLMs to human performance and observe that a significant gap remains between human visual reasoning capabilities and machine cognition.
Antonia Wüst, Tim Woydt, Lukas Helff, Inga Ibs, Wolfgang Stammer, Devendra Singh Dhami, Constantin A. Rothkopf, Kristian Kersting
ICML7
2025 Personalized Adaptive Magnification in Gaze-Based Interaction
abstract
Magnification has become the standard approach to tackle accuracy problems in eye-tracking based interaction systems, but suffers many problems itself. Magnification can be inefficient and oftentimes occludes contents in direct proximity to the magnifier, which is especially problematic when dealing with inaccuracies. While some of these problems can be tackled with special magnifying shapes, like hybrid linear-fisheye lenses, most approaches are rather limited one-size-fits-all approaches agnostic to user characteristics and content semantics. This limits the ability of magnification approaches to help users in scenarios where they encounter a below average accuracy or interact with targets that are especially small or crammed together. This paper presents Adaptive Magnification: a novel approach that is personalized to user-specific parameters and also context-aware by utilizing and mapping the screen content's semantic information to choose a magnification factor and size to optimally support interaction. The approach was evaluated against the use of no magnification and the use of hybrid linearfisheye magnification in a user study ($\mathrm{n}=20$). Task effectivity was evaluated by measuring the misclick count, while task efficiency wasn't measured as the approach is not focused on speed increase but on functionality and usability improvements. The user study revealed a 95.7% lower misclick count when compared to the use of no magnification and 13.1% fewer misclicks compared to hybrid linear-fisheye magnification.
Florian Eggenkemper, Jana Swerew, Teresa Rehers, Manuel Hanhoff, Constantin A. Rothkopf, Robert Mertens 0002
ISM5
2025 What do you know? Bayesian knowledge inference for navigating agents
abstract
Human behavior is characterized by continuous learning to reduce uncertainties about the world in pursuit of goals. When trying to understand such behavior from observations, it is essential to account for this adaptive nature and reason about the uncertainties that may have led to seemingly suboptimal decisions. Nevertheless, most inverse approaches to sequential decision-making focus on inferring cost functions underlying stationary behavior or are limited to low-dimensional tasks. In this paper, we address this gap by considering the problem of inferring an agent's knowledge or awareness about the environment based on a given trajectory. We assume that the agent aims to reach a goal in an environment they only partially know, and integrates new information into their plan as they act. We propose a Bayesian approach to infer their latent knowledge state, leveraging an approximate navigation model that optimistically incorporates partial information while accounting for uncertainty. By combining sample-based Bayesian inference with dynamic graph algorithms, we achieve an efficient method for computing posterior beliefs about the agent's knowledge. Empirical validation using simulated behavioral data and human data from an online experiment demonstrates that our model effectively captures human navigation under uncertainty and reveals interpretable insights into their environmental knowledge.
Matthias Schultheis, Jana-Sophie Schönfeld, Constantin A. Rothkopf, Heinz Koeppl
NeurIPS3
2025 Adaptation optimizes sensory encoding for future stimuli
abstract
Sensory neurons continually adapt their response characteristics according to recent stimulus history. However, it is unclear how such a reactive process can benefit the organism. Here, we test the hypothesis that adaptation actually acts proactively in the sense that it optimally adjusts sensory encoding for future stimuli. We first quantified human subjects' ability to discriminate visual orientation under different adaptation conditions. Using an information theoretic analysis, we found that adaptation leads to a reallocation of coding resources such that encoding accuracy peaks at the mean orientation of the adaptor while total coding capacity remains constant. We then asked whether this characteristic change in encoding accuracy is predicted by the temporal statistics of natural visual input. Analyzing the retinal input of freely behaving human subjects showed that the distribution of local visual orientations in the retinal input stream indeed peaks at the mean orientation of the preceding input history (i.e., the adaptor). We further tested our hypothesis by analyzing the internal sensory representations of a recurrent neural network trained to predict the next frame of natural scene videos (PredNet). Simulating our human adaptation experiment with PredNet, we found that the network exhibited the same change in encoding accuracy as observed in human subjects. Taken together, our results suggest that adaptation-induced changes in encoding accuracy prepare the visual system for future stimuli.
Jiang Mao, Constantin A. Rothkopf, Alan A. Stocker
PLoS Comput. Biol.2
2024 Task Diversity and Human Decision-Making: A Taxonomic View
Inga Ibs, Claire Ott, Constantin A. Rothkopf, Frank Jäkel
CogSci3
2024 If it looks like online control, it is probably model-based control
Dominik Straub, Constantin A. Rothkopf
CogSci2
2024 What Matters for Active Texture Recognition With Vision-Based Tactile Sensors
abstract
This paper explores active sensing strategies that employ vision-based tactile sensors for robotic perception and classification of fabric textures. We formalize the active sampling problem in the context of tactile fabric recognition and provide an implementation of information-theoretic exploration strategies based on minimizing predictive entropy and variance of probabilistic models. Through ablation studies and human experiments, we investigate which components are crucial for quick and reliable texture recognition. Along with the active sampling strategies, we evaluate neural network architectures, representations of uncertainty, influence of data augmentation, and dataset variability. By evaluating our method on a previously published Active Clothing Perception Dataset and on a real robotic system, we establish that the choice of the active exploration strategy has only a minor influence on the recognition accuracy, whereas data augmentation and dropout rate play a significantly larger role. In a comparison study, while humans achieve 66.9% recognition accuracy, our best approach reaches 90.0% in under 5 touches, highlighting that vision-based tactile sensors are highly effective for fabric texture recognition.
Alina Böhm, Boris Belousov, Alap Kshirsagar, Lisa Pui Yee Lin, Katja Doerschner, Knut Drewing, Constantin A. Rothkopf, Jan Peters 0001
ICRA8
2023 Finding your Way Out: Planning Strategies in Human Maze-Solving Behavior
Florian Kadner, Hannah Willkomm, Inga Ibs, Constantin A. Rothkopf
CogSci4
2023 People use Newtonian physics in intuitive sensorimotor decisions under risk
Fabian Tatai, Dominik Straub, Constantin A. Rothkopf
CogSci3
2023 Learning Individualized Automatic Content Magnification in Gaze-based Interaction
abstract
The precision of modern commercial off-the-shelf eye trackers has reached a level sufficient for developing gaze-based applications. In many but not all applications, accuracy even allows for replacing a computer mouse with gaze-bazed pointing. The Multi-Modal Interaction Concept for Efficient input (M2ice) tackles accuracy problems with an on-demand hybrid fisheye magnifier. This paper introduces an image-analysis-based approach that identifies areas on the screen where to automatically activate magnification for improved interaction. It combines the separate actions magnifying and clicking into one seamless action. The approach works by combining OpenCV filters for detection of clickable elements on the screen with a local machine learning algorithm predicting whether magnification is needed based on size and position of screen elements. A user study (n = 28) showed a significant speed increase of 18.50 percent (t(27)=-3.95, p=.0002 at α = .05) with automatic magnification compared to separate shortcuts for clicking and magnification.
Florian Eggenkemper, Lars Kölker, Mike Valente, Constantin A. Rothkopf, Robert Mertens 0002
ISM4
2023 Probabilistic inverse optimal control for non-linear partially observable systems disentangles perceptual uncertainty and behavioral costs
abstract
Inverse optimal control can be used to characterize behavior in sequential decision-making tasks. Most existing work, however, is limited to fully observable or linear systems, or requires the action signals to be known. Here, we introduce a probabilistic approach to inverse optimal control for partially observable stochastic non-linear systems with unobserved action signals, which unifies previous approaches to inverse optimal control with maximum causal entropy formulations. Using an explicit model of the noise characteristics of the sensory and motor systems of the agent in conjunction with local linearization techniques, we derive an approximate likelihood function for the model parameters, which can be computed within a single forward pass. We present quantitative evaluations on stochastic and partially observable versions of two classic control tasks and two human behavioral tasks. Importantly, we show that our method can disentangle perceptual factors and behavioral costs despite the fact that epistemic and pragmatic actions are intertwined in sequential decision-making under uncertainty, such as in active sensing and active learning. The proposed method has broad applicability, ranging from imitation learning to sensorimotor neuroscience.
Dominik Straub, Matthias Schultheis, Heinz Koeppl, Constantin A. Rothkopf
NeurIPS4
2023 What Can I Help You With: Towards Task-Independent Detection of Intentions for Interaction in a Human-Robot Environment
abstract
Assistive robots interacting with people promise to increase quality of life and productivity in households, caregiving, or industry settings. Importantly, the quality of such interactions crucially depends on the intuitive ease and reliability of humans being able to request the robot’s assistance. Thus, the ability to detect a human’s Intention for Interaction (IFI) is beneficial for human-robot interaction across multiple application domains. However, existing works that detect IFIs often focus on single tasks, contexts, or interactions or limit their data collection to invariability in human positions. In contrast, here we aim for a more task-independent IFI detection. We record natural human behavior in an experimental setup with a two-armed robot that includes different tasks and interactions, and different positions and orientations of the human towards the robot. We collected audio and RGB-D data from 21 human subjects in the proposed experimental setup resulting in overall 405 IFIs. Using head orientation, shoulder orientation, distance, speech activity recognition, and hotword detection as features, we trained multimodal probabilistic classifiers. We compare feature fusion and decision fusion using the Bayesian fusion method Independent Opinion Pool. The resulting multimodal classifiers can detect task-independent IFIs from natural human behavior with an F1 score of up to 0.81. Overall, we show that good IFI detection can be achieved by modularly combining individual classifiers probabilistically.
Susanne Trick, Vilja Lott, Lisa Kempf, Constantin A. Rothkopf, Dorothea Koert
RO-MAN4
2023 Improving saliency models' predictions of the next fixation with humans' intrinsic cost of gaze shifts
abstract
The human prioritization of image regions can be modeled in a time invariant fashion with saliency maps or sequentially with scanpath models. However, while both types of models have steadily improved on several benchmarks and datasets, there is still a considerable gap in predicting human gaze. Here, we leverage two recent developments to reduce this gap: theoretical analyses establishing a principled framework for predicting the next gaze target and the empirical measurement of the human cost for gaze switches independently of image content. We introduce an algorithm in the framework of sequential decision making, which converts any static saliency map into a sequence of dynamic history-dependent value maps, which are recomputed after each gaze shift. These maps are based on 1) a saliency map provided by an arbitrary saliency model, 2) the recently measured human cost function quantifying preferences in magnitude and direction of eye movements, and 3) a sequential exploration bonus, which changes with each subsequent gaze shift. The parameters of the spatial extent and temporal decay of this exploration bonus are estimated from human gaze data. The relative contributions of these three components were optimized on the MIT1003 dataset for the NSS score and are sufficient to significantly outperform predictions of the next gaze target on NSS and AUC scores for five state of the art saliency models on three image data sets.
Florian Kadner, Tobias Thomas, David Hoppe, Constantin A. Rothkopf
WACV4
2022 Bayesian Classifier Fusion with an Explicit Model of Correlation
abstract
Combining the outputs of multiple classifiers or experts into a single probabilistic classification is a fundamental task in machine learning with broad applications from classifier fusion to expert opinion pooling. Here we present a hierarchical Bayesian model of probabilistic classifier fusion based on a new correlated Dirichlet distribution. This distribution explicitly models positive correlations between marginally Dirichlet-distributed random vectors thereby allowing explicit modeling of correlations between base classifiers or experts. The proposed model naturally accommodates the classic Independent Opinion Pool and other independent fusion algorithms as special cases. It is evaluated by uncertainty reduction and correctness of fusion on synthetic and real-world data sets. We show that a change in performance of the fused classifier due to uncertainty reduction can be Bayes optimal even for highly correlated base classifiers.
Susanne Trick, Constantin A. Rothkopf
AISTATS2
2022 Reinforcement Learning with Non-Exponential Discounting
abstract
Commonly in reinforcement learning (RL), rewards are discounted over time using an exponential function to model time preference, thereby bounding the expected long-term reward. In contrast, in economics and psychology, it has been shown that humans often adopt a hyperbolic discounting scheme, which is optimal when a specific task termination time distribution is assumed. In this work, we propose a theory for continuous-time model-based reinforcement learning generalized to arbitrary discount functions. This formulation covers the case in which there is a non-exponential random termination time. We derive a Hamilton–Jacobi–Bellman (HJB) equation characterizing the optimal policy and describe how it can be solved using a collocation method, which uses deep learning for function approximation. Further, we show how the inverse RL problem can be approached, in which one tries to recover properties of the discount function given decision data. We validate the applicability of our proposed approach on two simulated problems. Our approach opens the way for the analysis of human discounting in sequential decision-making tasks.
Matthias Schultheis, Constantin A. Rothkopf, Heinz Koeppl
NeurIPS2
2021 AdaptiFont: Increasing Individuals' Reading Speed with a Generative Font Model and Bayesian Optimization
abstract
Digital text has become one of the primary ways of exchanging knowledge, but text needs to be rendered to a screen to be read. We present AdaptiFont, a human-in-the-loop system that is aimed at interactively increasing readability of text displayed on a monitor. To this end, we first learn a generative font space with non-negative matrix factorization from a set of classic fonts. In this space we generate new true-type-fonts through active learning, render texts with the new font, and measure individual users’ reading speed. Bayesian optimization sequentially generates new fonts on the fly to progressively increase individuals’ reading speed. The results of a user study show that this adaptive font generation system finds regions in the font space corresponding to high reading speeds, that these fonts significantly increase participants’ reading speed, and that the found fonts are significantly different across individual readers.
Florian Kadner, Yannik Keller, Constantin A. Rothkopf
CHI3
2021 Inverse Optimal Control Adapted to the Noise Characteristics of the Human Sensorimotor System
abstract
Computational level explanations based on optimal feedback control with signal-dependent noise have been able to account for a vast array of phenomena in human sensorimotor behavior. However, commonly a cost function needs to be assumed for a task and the optimality of human behavior is evaluated by comparing observed and predicted trajectories. Here, we introduce inverse optimal control with signal-dependent noise, which allows inferring the cost function from observed behavior. To do so, we formalize the problem as a partially observable Markov decision process and distinguish between the agent’s and the experimenter’s inference problems. Specifically, we derive a probabilistic formulation of the evolution of states and belief states and an approximation to the propagation equation in the linear-quadratic Gaussian problem with signal-dependent noise. We extend the model to the case of partial observability of state variables from the point of view of the experimenter. We show the feasibility of the approach through validation on synthetic data and application to experimental data. Our approach enables recovering the costs and benefits implicit in human sequential sensorimotor behavior, thereby reconciling normative and descriptive approaches in a computational framework.
Matthias Schultheis, Dominik Straub, Constantin A. Rothkopf
NeurIPS3
2020 Alfie: An Interactive Robot with Moral Compass
abstract
This work introduces Alfie, an interactive robot that is capable of answering moral (deontological) questions of a user. The interaction of Alfie is designed in a way in which the user can offer an alternative answer when the user disagrees with the given answer so that Alfie can learn from its interactions. Alfie's answers are based on a sentence embedding model that uses state-of-the-art language models, e.g. Universal Sentence Encoder and BERT. Alfie is implemented on a Furhat Robot, which provides a customizable user interface to design a social robot.
Cigdem Turan, Patrick Schramowski, Constantin A. Rothkopf, Kristian Kersting
ICMI3
2020 Intuitive physical reasoning about objects' masses transfers to a visuomotor decision task consistent with Newtonian physics
abstract
While interacting with objects during every-day activities, e.g. when sliding a glass on a counter top, people obtain constant feedback whether they are acting in accordance with physical laws. However, classical research on intuitive physics has revealed that people's judgements systematically deviate from predictions of Newtonian physics. Recent research has explained at least some of these deviations not as consequence of misconceptions about physics but instead as the consequence of the probabilistic interaction between inevitable perceptual uncertainties and prior beliefs. How intuitive physical reasoning relates to visuomotor actions is much less known. Here, we present an experiment in which participants had to slide pucks under the influence of naturalistic friction in a simulated virtual environment. The puck was controlled by the duration of a button press, which needed to be scaled linearly with the puck's mass and with the square-root of initial distance to reach a target. Over four phases of the experiment, uncertainties were manipulated by altering the availability of sensory feedback and providing different degrees of knowledge about the physical properties of pucks. A hierarchical Bayesian model of the visuomotor interaction task incorporating perceptual uncertainty and press-time variability found substantial evidence that subjects adjusted their button-presses so that the sliding was in accordance with Newtonian physics. After observing collisions between pucks, which were analyzed with a hierarchical Bayesian model of the perceptual observation task, subjects transferred the relative masses inferred perceptually to adjust subsequent sliding actions. Crucial in the modeling was the inclusion of a cost function, which quantitatively captures participants' implicit sensitivity to errors due to their motor variability. Taken together, in the present experiment we find evidence that our participants transferred their intuitive physical reasoning to a subsequent visuomotor control task consistent with Newtonian physics and weighed potential outcomes with a cost functions based on their knowledge about their own variability.
Nils Neupärtl, Fabian Tatai, Constantin A. Rothkopf
PLoS Comput. Biol.3
2019 Semantics Derived Automatically from Language Corpora Contain Human-like Moral Choices
abstract
Allowing machines to choose whether to kill humans would be devastating for world peace and security. But how do we equip machines with the ability to learn ethical or even moral choices? Here, we show that applying machine learning to human texts can extract deontological ethical reasoning about "right" and "wrong" conduct. We create a template list of prompts and responses, which include questions, such as "Should I kill people?", "Should I murder people?", etc. with answer templates of "Yes/no, I should (not)." The model's bias score is now the difference between the model's score of the positive response ("Yes, I should'') and that of the negative response ("No, I should not"). For a given choice overall, the model's bias score is the sum of the bias scores for all question/answer templates with that choice. We ran different choices through this analysis using a Universal Sentence Encoder. Our results indicate that text corpora contain recoverable and accurate imprints of our social, ethical and even moral choices. Our method holds promise for extracting, quantifying and comparing sources of moral choices in culture, including technology.
Sophie F. Jentzsch, Patrick Schramowski, Constantin A. Rothkopf, Kristian Kersting
AIES3
2019 Actor-Critic Instance Segmentation
abstract
Most approaches to visual scene analysis have emphasised parallel processing of the image elements. However, one area in which the sequential nature of vision is apparent, is that of segmenting multiple, potentially similar and partially occluded objects in a scene. In this work, we revisit the recurrent formulation of this challenging problem in the context of reinforcement learning. Motivated by the limitations of the global max-matching assignment of the ground-truth segments to the recurrent states, we develop an actor-critic approach in which the actor recurrently predicts one instance mask at a time and utilises the gradient from a concurrently trained critic network. We formulate the state, action, and the reward such as to let the critic model long-term effects of the current prediction and in- corporate this information into the gradient signal. Furthermore, to enable effective exploration in the inherently high-dimensional action space of instance masks, we learn a compact representation using a conditional variational auto-encoder. We show that our actor-critic model consistently provides accuracy benefits over the recurrent baseline on standard instance segmentation benchmarks.
Nikita Araslanov, Constantin A. Rothkopf, Stefan Roth 0001
CVPR2
2019 Multimodal Uncertainty Reduction for Intention Recognition in Human-Robot Interaction
abstract
Assistive robots can potentially improve the quality of life and personal independence of elderly people by supporting everyday life activities. To guarantee a safe and intuitive interaction between human and robot, human intentions need to be recognized automatically. As humans communicate their intentions multimodally, the use of multiple modalities for intention recognition may not just increase the robustness against failure of individual modalities but especially reduce the uncertainty about the intention to be recognized. This is desirable as particularly in direct interaction between robots and potentially vulnerable humans a minimal uncertainty about the situation as well as knowledge about this actual uncertainty is necessary. Thus, in contrast to existing methods, in this work a new approach for multimodal intention recognition is introduced that focuses on uncertainty reduction through classifier fusion. For the four considered modalities speech, gestures, gaze directions and scene objects individual intention classifiers are trained, all of which output a probability distribution over all possible intentions. By combining these output distributions using the Bayesian method Independent Opinion Pool [1] the uncertainty about the intention to be recognized can be decreased. The approach is evaluated in a collaborative human-robot interaction task with a 7-DoF robot arm. The results show that fused classifiers, which combine multiple modalities, outperform the respective individual base classifiers with respect to increased accuracy, robustness, and reduced uncertainty.
Susanne Trick, Dorothea Koert, Jan Peters 0001, Constantin A. Rothkopf
IROS4
2018 Modeling sensory-motor decisions in natural behavior
abstract
Although a standard reinforcement learning model can capture many aspects of reward-seeking behaviors, it may not be practical for modeling human natural behaviors because of the richness of dynamic environments and limitations in cognitive resources. We propose a modular reinforcement learning model that addresses these factors. Based on this model, a modular inverse reinforcement learning algorithm is developed to estimate both the rewards and discount factors from human behavioral data, which allows predictions of human navigation behaviors in virtual reality with high accuracy across different subjects and with different tasks. Complex human navigation trajectories in novel environments can be reproduced by an artificial agent that is based on the modular model. This model provides a strategy for estimating the subjective value of actions and how they influence sensory-motor decisions in natural behavior.
Matthew H. Tong, Yuchen Cui, Constantin A. Rothkopf, Dana H. Ballard, Mary M. Hayhoe
PLoS Comput. Biol.5
2017 I See What You See: Inferring Sensor and Policy Models of Human Real-World Motor Behavior
abstract
Human motor behavior is naturally guided by sensing the environment. To predict such sensori-motor behavior, it is necessary to model what is sensed and how actions are chosen based on the obtained sensory measurements. Although several models of human sensing haven been proposed, rarely data of the assumed sensory measurements is available. This makes statistical estimation of sensor models problematic. To overcome this issue, we propose an abstract structural estimation approach building on the ideas of Herman et al.'s Simultaneous Estimation of Rewards and Dynamics (SERD). Assuming optimal fusion of sensory information and rational choice of actions the proposed method allows to infer sensor models even in absence of data of the sensory measurements. To the best of our knowledge, this work presents the first general approach for joint inference of sensor and policy models. Furthermore, we consider its concrete implementation in the important class of sensor scheduling linear quadratic Gaussian problems. Finally, the effectiveness of the approach is demonstrated for prediction of the behavior of automobile drivers. Specifically, we model the glance and steering behavior of driving in the presence of visually demanding secondary tasks. The results show, that prediction benefits from the inference of sensor models. This is the case, especially, if also information is considered, that is contained in gaze switching behavior.
Felix Schmitt 0001, Hans-Joachim Bieg, Michael Herman, Constantin A. Rothkopf
AAAI4
2017 Adversarially Tuned Scene Generation
abstract
Generalization performance of trained computer vision (CV) systems that use computer graphics (CG) generated data is not yet effective due to the concept of domain-shift between virtual and real data. Although simulated data augmented with a few real-world samples has been shown to mitigate domain shift and improve transferability of trained models, guiding or bootstrapping the virtual data generation with the distributions learnt from target real world domain is desired, especially in the fields where annotating even few real images is laborious (such as semantic labeling, optical flow, and intrinsic images etc.). In order to address this problem in an unsupervised manner, our work combines recent advances in CG, which aims at generating stochastic scene layouts using large collections of 3D object models, and generative adversarial training, which aims at training generative models by measuring discrepancy between generated and real data in terms of their separability in the space of a deep discriminatively-trained classifier. Our method uses iterative estimation of the posterior density of prior distributions for a generative graphical model. This is done within a rejection sampling framework. Initially, we assume uniform distributions as priors over parameters of a scene described by a generative graphical model. As iterations proceed the uniform prior distributions are updated sequentially to distributions that are closer to the unknown distributions of target data. We demonstrate the utility of adversarially tuned scene generation on two real world benchmark datasets (CityScapes and CamVid) for traffic scene semantic labeling with a deep convolutional net (DeepLab). We obtained performance improvements by 2.28 and 3.14 points on the IoU metric between the DeepLab models trained on simulated sets prepared from the scene generation models before and after tuning to CityScapes and CamVid respectively.
V. S. R. Veeravasarapu, Constantin A. Rothkopf, Visvanathan Ramesh
CVPR2
2017 Model-Driven Simulations for Computer Vision
abstract
There is a growing interest to utilize Computer Graphics (CG) renderings to generate large scale annotated data in order to train machine learning systems, such as Deep convolutional neural networks, for Computer Vision (CV). However, there has been a long debate on the usefulness of CG generated data for tuning CV systems (even from the 1980's). Especially, the impact of modeling errors and computational rendering approximations, due to choices in the rendering pipeline, on trained CV systems generalization performance is still not clear. In this paper, we take a case study in traffic scenario to empirically analyze the performance degradation when CV systems trained with virtual data are transferred to real data. We: a) discuss a generative model coupled with 3D CAD shapes for scene instance synthesis and, b) explore system performance tradeoffs due to the choice of rendering engine (e.g. Lambertian shader (LS), ray-tracing (RT), and Monte-carlo path tracing (MCPT)) and their respective parameters. DeepLab, that performs semantic segmentation, is chosen as the CV system being evaluated. In our case study, involving traffic scenes, when the CV system is trained with CG data samples (that use MCPT or RT) and augmented with only 10% of real-world training data from CityScapes dataset, the performance levels achieved are comparable to that of training DeepLab with the complete CityScapes dataset. Use of samples from LS degraded the performance of DeepLab by 20%. Physics-based MCPT rendering improved the performance by 6% but at the cost of more than 3 times the rendering time.
V. S. R. Veeravasarapu, Constantin A. Rothkopf, Visvanathan Ramesh
WACV2
2017 A model of human motor sequence learning explains facilitation and interference effects based on spike-timing dependent plasticity
abstract
The ability to learn sequential behaviors is a fundamental property of our brains. Yet a long stream of studies including recent experiments investigating motor sequence learning in adult human subjects have produced a number of puzzling and seemingly contradictory results. In particular, when subjects have to learn multiple action sequences, learning is sometimes impaired by proactive and retroactive interference effects. In other situations, however, learning is accelerated as reflected in facilitation and transfer effects. At present it is unclear what the underlying neural mechanism are that give rise to these diverse findings. Here we show that a recently developed recurrent neural network model readily reproduces this diverse set of findings. The self-organizing recurrent neural network (SORN) model is a network of recurrently connected threshold units that combines a simplified form of spike-timing dependent plasticity (STDP) with homeostatic plasticity mechanisms ensuring network stability, namely intrinsic plasticity (IP) and synaptic normalization (SN). When trained on sequence learning tasks modeled after recent experiments we find that it reproduces the full range of interference, facilitation, and transfer effects. We show how these effects are rooted in the network's changing internal representation of the different sequences across learning and how they depend on an interaction of training schedule and task similarity. Furthermore, since learning in the model is based on fundamental neuronal plasticity mechanisms, the model reveals how these plasticity mechanisms are ultimately responsible for the network's sequence learning abilities. In particular, we find that all three plasticity mechanisms are essential for the network to learn effective internal models of the different training sequences. This ability to form effective internal models is also the basis for the observed interference and facilitation effects. This suggests that STDP, IP, and SN may be the driving forces behind our ability to learn complex action sequences.
Quan Wang 0003, Constantin A. Rothkopf, Jochen Triesch
PLoS Comput. Biol.2
2016 Catching heuristics are optimal control policies
abstract
Two seemingly contradictory theories attempt to explain how humans move to intercept an airborne ball. One theory posits that humans predict the ball trajectory to optimally plan future actions; the other claims that, instead of performing such complicated computations, humans employ heuristics to reactively choose appropriate actions based on immediate visual feedback. In this paper, we show that interception strategies appearing to be heuristics can be understood as computational solutions to the optimal control problem faced by a ball-catching agent acting under uncertainty. Modeling catching as a continuous partially observable Markov decision process and employing stochastic optimal control theory, we discover that the four main heuristics described in the literature are optimal solutions if the catcher has sufficient time to continuously visually track the ball. Specifically, by varying model parameters such as noise, time to ground contact, and perceptual latency, we show that different strategies arise under different circumstances. The catcher's policy switches between generating reactive and predictive behavior based on the ratio of system to observation noise and the ratio between reaction time and task duration. Thus, we provide a rational account of human ball-catching behavior and a unifying explanation for seemingly contradictory theories of target interception on the basis of stochastic optimal control.
Boris Belousov, Gerhard Neumann, Constantin A. Rothkopf, Jan Peters 0001
NIPS3
2011 Preference Elicitation and Inverse Reinforcement Learning
Constantin A. Rothkopf, Christos Dimitrakakis
ECML/PKDD (3)1
2004 Head movement estimation for wearable eye tracker
abstract
In the study of eye movements in natural tasks, where subjects are able to freely move in their environment, it is desirable to capture a video of the surroundings of the subject not limited to a small field of view as obtained by the scene camera of an eye tracker. Moreover, recovering the head movements could give additional information about the type of eye movement that was carried out, the overall gaze change in world coordinates, and insight into high-order perceptual strategies. Algorithms for the classification of eye movements in such natural tasks could also benefit form the additional head movement data.We propose to use an omnidirectional vision sensor consisting of a small CCD video camera and a hyperbolic mirror. The camera is mounted on an ASL eye tracker and records an image sequence at 60 Hz. Several algorithms for the extraction of rotational motion from this image sequence were implemented and compared in their performance against the measurements of a Fasttrack magnetic tracking system. Using data from the eye tracker together with the data obtained by the omnidirectional image sensor, a new algorithm for the classification of different types of eye movements based on a Hidden-Markov-Model was developed.
Constantin A. Rothkopf, Jeff B. Pelz
ETRA1