EDBT 2026 Demo / reviewers in the wild / expert
Daniel Sonntag
dblp:83/5858
· DBLP profile ↗
82ranked-venue papers
15as first author
38since 2021 · last 2026
0000-0002-8857-8709ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 41 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 35 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reinforce Trustworthiness in Multimodal Emotional Support SystemabstractIn today’s world, emotional support is increasingly essential, yet it remains challenging for both those seeking help and those offering it. Multimodal approaches to emotional support show great promise by integrating diverse data sources to provide empathetic, contextually relevant responses, fostering more effective interactions. However, current methods have notable limitations, often relying solely on text or converting other data types into text, or providing emotion recognition only, thus overlooking the full potential of multimodal inputs. Moreover, many studies prioritize response generation without accurately identifying critical emotional support elements or ensuring the reliability of outputs. To overcome these issues, we introduce MULTIMOOD, a new framework that (i) leverages multimodal embeddings from video, audio, and text to predict emotional components and to produce responses responses aligned with professional therapeutic standards. To improve trustworthiness, we (ii) incorporate novel psychological criteria and apply Reinforcement Learning (RL) to optimize large language models (LLMs) for consistent adherence to these standards. We also (iii) analyze several advanced LLMs to assess their multimodal emotional support capabilities. Experimental results show that MultiMood achieves state-of-the-art on MESC and DFEW datasets while RL-driven trustworthiness improvements are validated through human and LLM evaluations, demonstrating its superior capability in applying a multimodal framework in this domain. Huy M. Le, Tien Dat Nguyen, Ngan T. T. Vo, Tuan D. Q. Nguyen, Nguyen Binh Le, Duy M. H. Nguyen, Daniel Sonntag, Lizi Liao, Binh T. Nguyen 0001 |
AAAI | 7 |
| 2026 | Do You (Dis)agree With Me? Modelling Implicit User Disagreement in Human-AI Interaction Using Gaze DataabstractThe widespread use of generative AI has led to increased focus on human–AI interaction. However, AI systems can generate unexpected outputs, leading to disagreement or human–AI conflict. This paper focuses on modelling user disagreement using machine learning (ML) by observing users’ implicit viewing behaviour. We conducted a controlled study with 30 participants evaluating captions from a simulated ML image-captioning system. Participants indicated agreement or disagreement with each caption while we recorded their gaze and facial-expression data, which we used to predict (dis)agreement. We show that unimodal gaze-based personalised modelling (0.684 average balanced accuracy) outperforms generalised modelling (0.570), whereas multimodal approaches did not improve performance. Our exploratory post hoc gaze-based analysis highlights the importance of feature selection and temporal dynamics, which help guide system design and future work. We release the dataset to support reproducibility and further work. Due to the nature of this research, we also discuss the potential ethical and privacy implications of continuous passive gaze and facial monitoring. Abdulrahman Mohamed Selim, Omair Shahzad Bhatti, Amr Gomaa, Michael Barz, Daniel Sonntag |
CHI | 5 |
| 2026 | Aligning Instruction-Tuned LLMs for Event Extraction with Multi-objective Reinforcement Learning
Omar Adjali, Siting Liang, Omair Shahzad Bhatti, Daniel Sonntag |
ECIR (2) | 4 |
| 2026 | EyeGestureLogin: Spontaneous Hands‑Free Gaze‑Based Lock Pattern Authentication for Public DisplaysabstractAs public displays become more ubiquitous, common authentication methods face security (e.g., shoulder surfing) and sanitary risks. Gaze-based systems offer a promising alternative, but their adoption is often limited by required calibration and the Midas Touch problem. We present EyeGestureLogin, a touch-free, knowledge-based method that uses gaze to enter lock patterns on a familiar 3 × 3 layout. EyeGestureLogin supports spontaneous walk-up use by replacing explicit calibration with an implicit, single-point offset estimation, and mitigates the Midas Touch using a short dwell-time trigger for fully hands-free input. The interface captures patterns by processing AOI-based fixations; we additionally evaluate an offline saccade-based detector. In a controlled study (n = 24), we achieved 91.59% accuracy (91.88% offline) with a mean entry time of 4.36 s. These results suggest that EyeGestureLogin enables fast and accurate hands-free authentication for public displays, motivating further evaluation under real-world deployment conditions. Omair Shahzad Bhatti, Abdulrahman Mohamed Selim, László Kopácsi, Maximilian Biwersi, Michael Barz, Daniel Sonntag |
ETRA | 6 |
| 2026 | Train the Spire: An ML-Driven Single Player GWAP for Image AnnotationabstractTraining image classification models requires large labelled datasets, which is particularly challenging in specialised domains where manual expert annotation remains the default, such as eye tracking. To address this challenge, we present Train the Spire, a single-player game with a purpose (GWAP) that embeds image annotation within turn-based card game mechanics for crowdsourcing data annotations. The system uses a few-shot deep learning classifier to validate player-generated labels, providing immediate feedback through rewards and penalties. The game incorporates elements, such as progression systems, a companion agent for system transparency, and balanced difficulty, to maintain player engagement while ensuring annotation accuracy. In this paper, we present the system design, implementation details, and evaluation study design for comparing Train the Spire against a baseline annotation tool using the VISUS mobile eye-tracking dataset in an online user study measuring effectiveness, usability, and enjoyment. Keno Nanninga, Abdulrahman Mohamed Selim, Sara-Jane Bittner, Pascal Lessel, Michael Barz, Daniel Sonntag |
ETRA | 6 |
| 2026 | OpenGazeLab: An Interactive Toolkit for Gaze Analysis and Event DetectionabstractEvent detection (e.g., fixations and saccades) is a prerequisite for many eye-tracking analyses. However, existing solutions often require coding experience, rely on closed-source vendor tools, or focus on narrow paradigms (e.g., reading). Additionally, most are designed for stationary eye tracking, whereas head-mounted recordings are affected by head/scene motion, making event detection more difficult. Therefore, we present OpenGazeLab, a publicly available, browser-based toolkit that unifies event extraction, parameter configuration, and data inspection for both stationary and head-mounted eye tracking. OpenGazeLab implements I-DT and I-VT, and extends them for head-mounted data using scene-motion compensation and adaptive thresholds. The toolkit also provides a timeline-based visualisation that overlays gaze and detected events on the stimulus image or scene video, enabling quick visual verification. OpenGazeLab is implemented using widely used, well-maintained Python frameworks to support reproducible, easy-to-adopt workflows, and we plan to extend it with additional event classes (e.g., smooth pursuit) and alternative detectors. Khue Minh Pham, Abdulrahman Mohamed Selim, Omair Shahzad Bhatti, László Kopácsi, Michael Barz, Daniel Sonntag |
ETRA | 6 |
| 2025 | Towards Interpretable Radiology Report Generation via Concept Bottlenecks Using a Multi-agentic RAG
Hasan Md Tusfiqur Alam, Devansh Srivastav, Md Abdul Kadir, Daniel Sonntag |
ECIR (3) | 4 |
| 2025 | On Zero-Initialized Attention: Optimal Prompt and Gating Factor EstimationabstractLLaMA-Adapter has recently emerged as an efficient fine-tuning technique for LLaMA models, leveraging zero-initialized attention to stabilize training and enhance performance. However, despite its empirical success, the theoretical foundations of zero-initialized attention remain largely unexplored. In this paper, we provide a rigorous theoretical analysis, establishing a connection between zero-initialized attention and mixture-of-expert models. We prove that both linear and non-linear prompts, along with gating functions, can be optimally estimated, with non-linear prompts offering greater flexibility for future applications. Empirically, we validate our findings on the open LLM benchmarks, demonstrating that non-linear prompts outperform linear ones. Notably, even with limited training data, both prompt types consistently surpass vanilla attention, highlighting the robustness and adaptability of zero-initialized attention. Nghiem Tuong Diep, Minh Le, Duy M. H. Nguyen, Daniel Sonntag, Mathias Niepert, Nhat Ho |
ICML | 6 |
| 2025 | Eliciting Multimodal Approaches for Machine Learning-Assisted Photobook Creation
Sara-Jane Bittner, Michael Barz, Daniel Sonntag |
INTERACT (2) | 3 |
| 2025 | Gaze-Based Menu Navigation in Virtual Reality: A Comparative Study of Layouts and Interaction TechniquesabstractAbstract Integrating eye-tracking technologies in Extended Reality (XR) headsets has enabled intuitive, hands-free system interaction, such as gaze-based menu navigation. However, there is a lack of comprehensive comparisons and consensus in the literature on the optimal use of gaze-based menu navigation. This paper presents a comparative analysis of gaze-based menu navigation in virtual environments, focusing on two common menu layouts: pie and list menus, with three interaction methods: gaze-based dwell, controller-based, and a multimodal approach combining gaze and controller inputs. We conducted a 19-participant within-subject study, measuring task completion time, error rate, usability, and user preference for each condition. The results indicate that while the pie layout was statistically faster and less erroneous than the list layout, novice users tend to favour list layouts. Furthermore, we found that users preferred the multimodal interaction method, despite its lower task completion times and higher error rates compared to controller-based navigation. Based on our findings, we offer design guidelines and recommendations for implementing gaze-based menu systems. László Kopácsi, Albert Klimenko, Abdulrahman Mohamed Selim, Michael Barz, Daniel Sonntag |
INTERACT (1) | 5 |
| 2025 | Towards Trustable Intelligent Clinical Decision Support Systems: A User Study with OphthalmologistsabstractIntegrating Artificial Intelligence (AI) into Clinical Decision Support Systems (CDSS) presents significant opportunities for improving healthcare delivery, particularly in fields like ophthalmology.This paper explores the usability and trustworthiness of an AI-driven CDSS designed to assist ophthalmologists in treating diabetic retinopathy and age-related macular degeneration.Therefore, we created a CDSS and evaluated its impact on efficiency, informedness, and user experience through task-based semi-structured interviews and questionnaires with 11 ophthalmologists.The usability of the CDSS was rated highly, with a SUS of 81.75.Additionally, results show that participants felt like the CDSS would improve their efficiency and informedness with one major aspect being integrating Electronic Health Records (EHR) and Optical Coherence Tomography (OCT) data into a single interface.Additionally, we explored aspects of the trustworthiness of AI components, specifically OCT segmentation, treatment recommendation, and visual acuity forecasting.Through thematic analysis, we identified key factors influencing trustworthiness and clinical adoption.Results show that a larger degree of abstraction from input to output of a model correlates with decreased trust.From our findings, we propose three guidelines for designing trustworthy CDSS. Robert Andreas Leist, Hans-Jürgen Profitlich, Tim Hunsicker, Daniel Sonntag |
IUI | 4 |
| 2025 | ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language ModelsabstractState-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach can lead to weak vision-language alignment, making these models overly dependent on costly instruction-following data. To address this, we introduce ExGra-Med, a novel multi-graph alignment framework that jointly aligns images, instruction responses, and extended captions in the latent space, advancing semantic grounding and cross-modal coherence. To scale to large LLMs (e.g., LLaMa-7B), we develop an efficient end-to-end training scheme using black-box gradient estimation, enabling fast and scalable optimization. Empirically, ExGra-Med matches LLaVA-Med’s performance using just 10\% of pre-training data, achieving a 20.13\% gain on VQA-RAD and approaching full-data performance. It also outperforms strong baselines like BioMedGPT and RadFM on visual chatbot and zero-shot classification tasks, demonstrating its promise for efficient, high-quality vision-language integration in medical AI. Duy M. H. Nguyen, Nghiem Tuong Diep, Hoang Bao Le, Tai D. Nguyen, Anh-Tien Nguyen, TrungTin Nguyen, Nhat Ho, Pengtao Xie, Roger Wattenhofer, Daniel Sonntag, James Zou 0001, Mathias Niepert |
NeurIPS | 11 |
| 2025 | Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance SamplingabstractRecently, Direct Alignment Algorithms (DAAs) such as Direct Preference Optimization (DPO) have emerged as alternatives to the standard Reinforcement Learning from Human Feedback (RLHF) for aligning large language models (LLMs) with human values.
Surprisingly, while DAAs do not use a separate proxy reward model as in RLHF, their performance can still deteriorate over the course of training -- an over-optimization phenomenon found in RLHF where the learning policy exploits the overfitting to inaccuracies of the reward model to achieve high rewards.
One attributed source of over-optimization in DAAs is the under-constrained nature of their offline optimization, which can gradually shift probability mass toward non-preferred responses not presented in the preference dataset. This paper proposes a novel importance-sampling approach to mitigate the distribution shift problem of offline DAAs.
This approach, called (IS-DAAs), multiplies the DAA objective with an importance ratio that accounts for the reference policy distribution. IS-DAAs additionally avoid the high variance issue associated with importance sampling by clipping the importance ratio to a maximum value. Our extensive experiments demonstrate that IS-DAAs can effectively mitigate over-optimization, especially under low regularization strength, and achieve better performance than other methods designed to address this problem. Phuc Minh Nguyen, Ngoc-Hieu Nguyen, Duy M. H. Nguyen, Anji Liu, An Mai, Binh T. Nguyen 0001, Daniel Sonntag, Khoa D. Doan |
NeurIPS | 7 |
| 2025 | How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?abstractRecent advances in 3D point cloud transformers have led to state-of-the-art results in tasks such as semantic segmentation and reconstruction. However, these models typically rely on dense token representations, incurring high computational and memory costs during training and inference. In this work, we present the finding that tokens are remarkably redundant, leading to substantial inefficiency. We introduce \textbf{GitMerge3D}, a \textbf{g}lobally \textbf{i}nformed graph \textbf{t}oken \textbf{merging} method that can reduce the token count by up to 90–95\% while maintaining competitive performance. This finding challenges the prevailing assumption that more tokens inherently yield better performance and highlights that many current models are over-tokenized and under-optimized for scalability. We validate our method across multiple 3D vision tasks and show consistent improvements in computational efficiency. This work is the first to assess redundancy in large-scale 3D transformer models, providing insights into the development of more efficient 3D foundation architectures. Our code and checkpoints are publicly available at \href{https://gitmerge3d.github.io/}{https://gitmerge3d.github.io}. Duy M. H. Nguyen, Hoai-Chau Tran, Michael Barz, Khoa D. Doan, Roger Wattenhofer, Ngo Anh Vien, Mathias Niepert, Daniel Sonntag, Paul Swoboda |
NeurIPS | 9 |
| 2024 | Dude: Dual Distribution-Aware Context Prompt Learning For Large Vision-Language Model
Duy M. H. Nguyen, An T. Le 0001, Trung Quoc Nguyen, Nghiem Tuong Diep, Tai Nguyen 0008, Duy Duong-Tran, Jan Peters 0001, Li Shen 0001, Mathias Niepert, Daniel Sonntag |
ACML | 10 |
| 2024 | Speech Imagery BCI Training Using Game with a PurposeabstractGames are used in multiple fields of brain–computer interface (BCI) research and applications to improve participants’ engagement and enjoyment during electroencephalogram (EEG) data collection. However, despite potential benefits, no current studies have reported on implemented games for Speech Imagery BCI. Imagined speech is speech produced without audible sounds or active movement of the articulatory muscles. Collecting imagined speech EEG data is a time-consuming, mentally exhausting, and cumbersome process, which requires participants to read words off a computer screen and produce them as imagined speech. To improve this process for study participants, we implemented a maze-like game where a participant navigated a virtual robot capable of performing five actions that represented our words of interest while we recorded their EEG data. The study setup was evaluated with 15 participants. Based on their feedback, the game improved their engagement and enjoyment while resulting in a 69.10% average classification accuracy using a random forest classifier. Abdulrahman Mohamed Selim, Maurice Rekrut, Michael Barz, Daniel Sonntag |
AVI | 4 |
| 2024 | HumanEYEze 2024: Workshop on Eye Tracking for Multimodal Human-Centric ComputingabstractThe HumanEYEze 2024 workshop aims to explore the role of eye tracking in developing human-centered multimodal AI systems. Over the past two decades, eye tracking has evolved from a diagnostic tool to an important input modality for real-time interactive systems, driven by advancements in hardware that have improved its affordability, availability, and performance. Initially used in specialized applications, eye tracking now significantly impacts research on gaze-based multimodal interaction. Recently, eye-based user and context modeling has emerged, utilizing eye movements to provide rich insights into user behavior and interaction contexts. The workshop aims to bring together researchers from eye tracking, multimodal human-computer interaction, and AI. It aims to enhance understanding of integrating eye tracking into multimodal human-centered computing. The expected outcomes include fostering collaborations and promoting knowledge exchange. Michael Barz, Roman Bednarik, Andreas Bulling, Cristina Conati, Daniel Sonntag |
ICMI | 5 |
| 2024 | Perceived Text Relevance Estimation Using Scanpaths and GNNsabstractA scanpath is an important concept in eye tracking that represents a person’s eye movements in a graph-like structure. Passive gaze-based interfaces, in which users do not consciously interact using their eyes, typically interpret users’ scanpaths to enable adaptive and personalised interaction. Despite the benefits of graph neural networks (GNNs) in graph processing, this technology has not been considered for that purpose. An example application is perceived relevance estimation, which still suffers from low classification performance. In this work, we investigate how and whether GNNs can be used to analyse scanpaths for readers’ perceived relevance estimation using the gazeRE dataset. This dataset contains eye tracking data from 24 participants, who rated the relevance of 12 short and 12 long documents in relation to a given query. The relevance was assigned either to an entire short document or to each paragraph within a long document, which allowed us to investigate two different GNN tasks. For comparison, we reproduced the gazeRE baseline using Random Forest and Support Vector classifiers, and an additional Convolutional Neural Network (CNN) from the literature. All models were evaluated using leave-users-out cross-validation. For short documents, the GNNs surpassed the baseline methods, with certain experiments showing an absolute balanced accuracy improvement of 7.6% and 14.3% over the CNN and gazeRE baselines, respectively. However, similar improvements were not observed in long documents. This work investigates and discusses the future potential of using GNNs as a scanpath analysis method for passive gaze-based applications, such as implicit relevance estimation. Abdulrahman Mohamed Selim, Omair Shahzad Bhatti, Michael Barz, Daniel Sonntag |
ICMI | 4 |
| 2024 | Structure-Aware E(3)-Invariant Molecular Conformer Aggregation NetworksabstractA molecule’s 2D representation consists of its atoms, their attributes, and the molecule’s covalent bonds. A 3D (geometric) representation of a molecule is called a conformer and consists of its atom types and Cartesian coordinates. Every conformer has a potential energy, and the lower this energy, the more likely it occurs in nature. Most existing machine learning methods for molecular property prediction consider either 2D molecular graphs or 3D conformer structure representations in isolation. Inspired by recent work on using ensembles of conformers in conjunction with 2D graph representations, we propose E(3)-invariant molecular conformer aggregation networks. The method integrates a molecule’s 2D representation with that of multiple of its conformers. Contrary to prior work, we propose a novel 2D–3D aggregation mechanism based on a differentiable solver for the Fused Gromov-Wasserstein Barycenter problem and the use of an efficient conformer generation method based on distance geometry. We show that the proposed aggregation mechanism is E(3) invariant and propose an efficient GPU implementation. Moreover, we demonstrate that the aggregation mechanism helps to significantly outperform state-of-the-art molecule property prediction methods on established datasets. Duy M. H. Nguyen, Nina Lukashina, Tai Nguyen 0008, An T. Le 0001, TrungTin Nguyen, Nhat Ho, Jan Peters 0001, Daniel Sonntag, Viktor Zaverkin, Mathias Niepert |
ICML | 8 |
| 2024 | Demo: Enhancing Wildlife Acoustic Data Annotation Efficiency through Transfer and Active Learning
Hannes Kath, Patricia P. Serafini, Ivan Braga Campos, Thiago S. Gouvêa, Daniel Sonntag |
IJCAI | 5 |
| 2024 | Accelerating Transformers with Spectrum-Preserving Token MergingabstractIncreasing the throughput of the Transformer architecture, a foundational component used in numerous state-of-the-art models for vision and language tasks (e.g., GPT, LLaVa), is an important problem in machine learning. One recent and effective strategy is to merge token representations within Transformer models, aiming to reduce computational and memory requirements while maintaining accuracy. Prior work has proposed algorithms based on Bipartite Soft Matching (BSM), which divides tokens into distinct sets and merges the top $k$ similar tokens. However, these methods have significant drawbacks, such as sensitivity to token-splitting strategies and damage to informative tokens in later layers. This paper presents a novel paradigm called PiToMe, which prioritizes the preservation of informative tokens using an additional metric termed the \textit{energy score}. This score identifies large clusters of similar tokens as high-energy, indicating potential candidates for merging, while smaller (unique and isolated) clusters are considered as low-energy and preserved. Experimental findings demonstrate that PiToMe saved from 40-60\% FLOPs of the base models while exhibiting superior off-the-shelf performance on image classification (0.5\% average performance drop of ViT-MAEH compared to 2.6\% as baselines), image-text retrieval (0.3\% average performance drop of Clip on Flick30k compared to 4.5\% as others), and analogously in visual questions answering with LLaVa-7B. Furthermore, PiToMe is theoretically shown to preserve intrinsic spectral properties to the original token space under mild conditions. Chau Tran, Duy M. H. Nguyen, Duy Nguyen 0003, TrungTin Nguyen, T. Hoang Ngan Le, Pengtao Xie, Daniel Sonntag, James Zou 0001, Mathias Niepert |
NeurIPS | 7 |
| 2024 | The MASTER XR Platform for Robotics Training in ManufacturingabstractThe MASTER project introduces an open Extended Reality (XR) platform designed to enhance human-robot collaboration and train workers in robotics within manufacturing settings. It includes modules for creating safe workspaces, intuitive robot programming, and user-friendly human-robot interactions (HRI), including eye-tracking technologies. The development of the platform is supported by two open calls targeting technical SMEs and educational institutes to enhance and test its functionalities. By employing the learning-by-doing methodology and integrating effective teaching principles, the MASTER platform aims to provide a comprehensive learning environment, preparing students and professionals for the complexities of flexible and collaborative manufacturing settings. László Kopácsi, Panagiotis Karagiannis, Sotiris Makris, Johan Kildal, Andoni Rivera-Pinto, Judit Ruiz de Munain, Jesús Rosel, Maria Madarieta, Nikolaos Tseregkounis, Konstantina Salagianni, Panagiotis Aivaliotis, Michael Barz, Daniel Sonntag |
VRST | 13 |
| 2024 | GazeLock: Gaze- and Lock Pattern-Based AuthenticationabstractPassword entry is common authentication approach in Extended Reality (XR) applications for its simplicity and familiarity, but it faces challenges in public and dynamic environments due to its cumbersome nature and susceptibility to observation attacks. Manual password input can be disruptive and prone to theft through shoulder surfing or surveillance. While alternative knowledge-based approaches exist, they often require complex physical gestures and are impractical for frequent public use. We present GazeLock, an eye-tracking and lock pattern-based authentication method. This method aims to provide an easy-to-learn and efficient alternative by leveraging familiar lock patterns operated through gaze. It ensures resilience to external observation, as physical interaction is unnecessary and eyes are obscured by the headset. Its hands-free, discreet nature makes it suitable for secure public use. We demonstrate this method by simulating the unlocking of a smart lock via an XR headset, showcasing its potential applications and benefits in real-world scenarios. László Kopácsi, Tobias Sebastian Schneider, Chiara Karr, Michael Barz, Daniel Sonntag |
VRST | 5 |
| 2024 | A References Architecture for Human Cyber Physical Systems, Part II: Fundamental Design Principles for Human-CPS InteractionabstractAs automation increases qualitatively and quantitatively in safety-critical human cyber-physical systems, it is becoming more and more challenging to increase the probability or ensure that human operators still perceive key artifacts and comprehend their roles in the system. In the companion paper, we proposed an abstract reference architecture capable of expressing all classes of system-level interactions in human cyber-physical systems. Here we demonstrate how this reference architecture supports the analysis of levels of communication between agents and helps to identify the potential for misunderstandings and misconceptions. We then develop a metamodel for safe human machine interaction. Therefore, we ask what type of information exchange must be supported on what level so that humans and systems can cooperate as a team, what is the criticality of exchanged information, what are timing requirements for such interactions, and how can we communicate highly critical information in a limited time frame in spite of the many sources of a distorted perception. We highlight shared stumbling blocks and illustrate shared design principles, which rest on established ontologies specific to particular application classes. In order to overcome the partial opacity of internal states of agents, we anticipate a key role of virtual twins of both human and technical cooperation partners for designing a suitable communication. Klaus Bengler, Werner Damm, Andreas Lüdtke, Jochem W. Rieger, Benedikt Austel, Bianca Biebl, Martin Fränzle, Willem Hagemann, Moritz Held, David Hess, Klas Ihme, Severin Kacianka, Alyssa J. Kerscher, Forrest Laine, Sebastian Lehnhoff, Alexander Pretschner, Astrid Rakow, Daniel Sonntag, Janos Sztipanovits, Maike Schwammberger, Mark Schweda, Anirudh Unni, Eric M. S. P. Veith |
ACM Trans. Cyber Phys. Syst. | 18 |
| 2024 | A Reference Architecture of Human Cyber-Physical Systems - Part III: Semantic FoundationsabstractThe design and analysis of multi-agent human cyber-physical systems in safety-critical or industry-critical domains calls for an adequate semantic foundation capable of exhaustively and rigorously describing all emergent effects in the joint dynamic behavior of the agents that are relevant to their safety and well-behavior. We present such a semantic foundation. This framework extends beyond previous approaches by extending the agent-local dynamic state beyond state components under direct control of the agent and belief about other agents (as previously suggested for understanding cooperative as well as rational behavior) to agent-local evidence and belief about the overall cooperative, competitive, or coopetitive game structure. We argue that this extension is necessary for rigorously analyzing systems of human cyber-physical systems because humans are known to employ cognitive replacement models of system dynamics that are both non-stationary and potentially incongruent. These replacement models induce visible and potentially harmful effects on their joint emergent behavior and the interaction with cyber-physical system components. Werner Damm, Martin Fränzle, Alyssa J. Kerscher, Forrest Laine, Klaus Bengler, Bianca Biebl, Willem Hagemann, Moritz Held, David Hess, Klas Ihme, Severin Kacianka, Sebastian Lehnhoff, Andreas Lüdtke, Alexander Pretschner, Astrid Rakow, Jochem W. Rieger, Daniel Sonntag, Janos Sztipanovits, Maike Schwammberger, Mark Schweda, Alexander Trende, Anirudh Unni, Eric M. S. P. Veith |
ACM Trans. Cyber Phys. Syst. | 17 |
| 2024 | A Reference Architecture of Human Cyber-Physical Systems - Part I: Fundamental ConceptsabstractWe propose a reference architecture of safety-critical or industry-critical human cyber-physical systems (CPSs) capable of expressing essential classes of system-level interactions between CPS and humans relevant for the societal acceptance of such systems. To reach this quality gate, the expressivity of the model must go beyond classical viewpoints such as operational, functional, and architectural views and views used for safety and security analysis. The model does so by incorporating elements of such systems for mutual introspections in situational awareness, capabilities, and intentions to enable a synergetic, trusted relation in the interaction of humans and CPSs, which we see as a prerequisite for their societal acceptance. The reference architecture is represented as a metamodel incorporating conceptual and behavioral semantic aspects. We illustrate the key concepts of the metamodel with examples from cooperative autonomous driving, the operating room of the future, cockpit-tower interaction, and crisis management. Werner Damm, David Hess, Mark Schweda, Janos Sztipanovits, Klaus Bengler, Bianca Biebl, Martin Fränzle, Willem Hagemann, Moritz Held, Klas Ihme, Severin Kacianka, Alyssa J. Kerscher, Sebastian Lehnhoff, Andreas Lüdtke, Alexander Pretschner, Astrid Rakow, Jochem W. Rieger, Daniel Sonntag, Maike Schwammberger, Benedikt Austel, Anirudh Unni, Eric M. S. P. Veith |
ACM Trans. Cyber Phys. Syst. | 18 |
| 2023 | Joint Self-Supervised Image-Volume Representation Learning with Intra-inter Contrastive ClusteringabstractCollecting large-scale medical datasets with fully annotated samples for training of deep networks is prohibitively expensive, especially for 3D volume data. Recent breakthroughs in self-supervised learning (SSL) offer the ability to overcome the lack of labeled training samples by learning feature representations from unlabeled data. However, most current SSL techniques in the medical field have been designed for either 2D images or 3D volumes. In practice, this restricts the capability to fully leverage unlabeled data from numerous sources, which may include both 2D and 3D data. Additionally, the use of these pre-trained networks is constrained to downstream tasks with compatible data dimensions. In this paper, we propose a novel framework for unsupervised joint learning on 2D and 3D data modalities. Given a set of 2D images or 2D slices extracted from 3D volumes, we construct an SSL task based on a 2D contrastive clustering problem for distinct classes. The 3D volumes are exploited by computing vectored embedding at each slice and then assembling a holistic feature through deformable self-attention mechanisms in Transformer, allowing incorporating long-range dependencies between slices inside 3D volumes. These holistic features are further utilized to define a novel 3D clustering agreement-based SSL task and masking embedding prediction inspired by pre-trained language models. Experiments on downstream tasks, such as 3D brain segmentation, lung nodule detection, 3D heart structures segmentation, and abnormal chest X-ray detection, demonstrate the effectiveness of our joint 2D and 3D SSL approach. We improve plain 2D Deep-ClusterV2 and SwAV by a significant margin and also surpass various modern 2D and 3D SSL approaches. Duy M. H. Nguyen, Truong Thanh Nhat Mai, Tri Cao, Binh T. Nguyen 0001, Nhat Ho, Paul Swoboda, Shadi Albarqouni, Pengtao Xie, Daniel Sonntag |
AAAI | 10 |
| 2023 | Interactive Machine Learning Solutions for Acoustic Monitoring of Animal Wildlife in Biosphere ReservesabstractBiodiversity loss is taking place at accelerated rates globally, and a business-as-usual trajectory will lead to missing internationally established conservation goals. Biosphere reserves are sites designed to be of global significance in terms of both the biodiversity within them and their potential for sustainable development, and are therefore ideal places for the development of local solutions to global challenges. While the protection of biodiversity is a primary goal of biosphere reserves, adequate information on the state and trends of biodiversity remains a critical gap for adaptive management in biosphere reserves. Passive acoustic monitoring (PAM) is an increasingly popular method for continued, reproducible, scalable, and cost-effective monitoring of animal wildlife. PAM adoption is on the rise, but its data management and analysis requirements pose a barrier for adoption for most agencies tasked with monitoring biodiversity. As an interdisciplinary team of machine learning scientists and ecologists experienced with PAM and working at biosphere reserves in marine and terrestrial ecosystems on three different continents, we report on the co-development of interactive machine learning tools for semi-automated assessment of animal wildlife. Thiago S. Gouvêa, Hannes Kath, Ilira Troshani, Bengt Lüers, Patricia P. Serafini, Ivan Braga Campos, André S. Afonso, Sergio M. F. M. Leandro, Lourens Swanepoel, Nicholas Theron, Anthony M. Swemmer, Daniel Sonntag |
IJCAI | 12 |
| 2023 | A Human-in-the-Loop Tool for Annotating Passive Acoustic Monitoring DatasetsabstractDeep learning methods are well suited for data analysis in several domains, but application is often limited by technical entry barriers and the availability of large annotated datasets. We present an interactive machine learning tool for annotating passive acoustic monitoring datasets created for wildlife monitoring, which are time-consuming and costly to annotate manually. The tool, designed as a web application, consists of an interactive user interface implementing a human-in-the-loop workflow. Class label annotations provided manually as bounding boxes drawn over a spectrogram are consumed by a deep generative model (DGM) that learns a low-dimensional representation of the input data, as well as the available class labels. The learned low-dimensional representation is displayed as an interactive interface element, where new bounding boxes can be efficiently generated by the user with lasso-selection; alternatively, the DGM can propose new, automatically generated bounding boxes on demand. The user can accept, edit, or reject annotations suggested by the model, thus owning final judgement. Generated annotations can be used to fine-tune the underlying model, thus closing the loop. Investigations of the prediction accuracy and first empirical experiments show promising results on an artificial data set, laying the ground for application to a real life scenario. Hannes Kath, Thiago S. Gouvêa, Daniel Sonntag |
IJCAI | 3 |
| 2023 | EdgeAL: An Edge Estimation Based Active Learning Approach for OCT Segmentation
Md Abdul Kadir, Hasan Md Tusfiqur Alam, Daniel Sonntag |
MICCAI (2) | 3 |
| 2023 | LVM-Med: Learning Large-Scale Self-Supervised Vision Models for Medical Imaging via Second-order Graph MatchingabstractObtaining large pre-trained models that can be fine-tuned to new tasks with limited annotated samples has remained an open challenge for medical imaging data. While pre-trained networks on ImageNet and vision-language foundation models trained on web-scale data are the prevailing approaches, their effectiveness on medical tasks is limited due to the significant domain shift between natural and medical images. To bridge this gap, we introduce LVM-Med, the first family of deep networks trained on large-scale medical datasets. We have collected approximately 1.3 million medical images from 55 publicly available datasets, covering a large number of organs and modalities such as CT, MRI, X-ray, and Ultrasound. We benchmark several state-of-the-art self-supervised algorithms on this dataset and propose a novel self-supervised contrastive learning algorithm using a graph-matching formulation. The proposed approach makes three contributions: (i) it integrates prior pair-wise image similarity metrics based on local and global information; (ii) it captures the structural constraints of feature embeddings through a loss function constructed through a combinatorial graph-matching objective, and (iii) it can be trained efficiently end-to-end using modern gradient-estimation techniques for black-box solvers. We thoroughly evaluate the proposed LVM-Med on 15 downstream medical tasks ranging from segmentation and classification to object detection, and both for the in and out-of-distribution settings. LVM-Med empirically outperforms a number of state-of-the-art supervised, self-supervised, and foundation models. For challenging tasks such as Brain Tumor Classification or Diabetic Retinopathy Grading, LVM-Med improves previous vision-language models trained on 1 billion masks by 6-7% while using only a ResNet-50. Duy M. H. Nguyen, Nghiem Tuong Diep, Tan Ngoc Pham, Tri Cao, Binh T. Nguyen 0001, Paul Swoboda, Nhat Ho, Shadi Albarqouni, Pengtao Xie, Daniel Sonntag, Mathias Niepert |
NeurIPS | 11 |
| 2022 | Interactive Assessment Tool for Gaze-based Machine Learning Models in Information RetrievalabstractEye movements were shown to be an effective source of implicit relevance feedback in information retrieval tasks. They can be used to, e.g., estimate the relevance of read documents and expand search queries using machine learning. In this paper, we present the Reading Model Assessment tool (ReMA), an interactive tool for assessing gaze-based relevance estimation models. Our tool allows experimenters to easily browse recorded trials, compare the model output to a ground truth, and visualize gaze-based features at the token- and paragraph-level that serve as model input. Our goal is to facilitate the understanding of the relation between eye movements and the human relevance estimation process, to understand the strengths and weaknesses of a model at hand, and, eventually, to enable researchers to build more effective models. Pablo Valdunciel, Omair Shahzad Bhatti, Michael Barz, Daniel Sonntag |
CHIIR | 4 |
| 2022 | LMGP: Lifted Multicut Meets Geometry Projections for Multi-Camera Multi-Object TrackingabstractMulti-Camera Multi-Object Tracking is currently drawing attention in the computer vision field due to its superior performance in real-world applications such as video surveillance with crowded scenes or in wide spaces. In this work, we propose a mathematically elegant multi-camera multiple object tracking approach based on a spatial-temporal lifted multicut formulation. Our model utilizes state-of-the-art tracklets produced by single-camera trackers as proposals. As these tracklets may contain ID-Switch errors, we refine them through a novel pre-clustering obtained from 3D geometry projections. As a result, we derive a better tracking graph without ID switches and more precise affinity costs for the data association phase. Tracklets are then matched to multi-camera trajectories by solving a global lifted multicut formulation that incorporates short and long-range temporal interactions on tracklets located in the same camera as well as inter-camera ones. Experimental results on the WildTrack dataset yield near-perfect performance, outperforming state-of-the-art trackers on Campus while being on par on the PETS-09 dataset. We will release our implementations at this link https://github.com/nhmduy/LMGP. Duy M. H. Nguyen, Roberto Henschel, Bodo Rosenhahn, Daniel Sonntag, Paul Swoboda |
CVPR | 4 |
| 2022 | TATL: Task agnostic transfer learning for skin attributes detection
Duy M. H. Nguyen, Thu T. Nguyen, Huong Vu, Quang Pham, Duy Nguyen 0003, Binh T. Nguyen 0001, Daniel Sonntag |
Medical Image Anal. | 7 |
| 2021 | Anomaly Detection for Skin Lesion Images Using Replicator Neural Networks
Fabrizio Nunnari, Hasan Md Tusfiqur Alam, Daniel Sonntag |
CD-MAKE | 3 |
| 2021 | On the Overlap Between Grad-CAM Saliency Maps and Explainable Visual Features in Skin Cancer Images
Fabrizio Nunnari, Md Abdul Kadir, Daniel Sonntag |
CD-MAKE | 3 |
| 2021 | Explainable Automatic Evaluation of the Trail Making Test for Dementia ScreeningabstractThe Trail Making Test (TMT) is a frequently used neuropsychological test for assessing cognitive performance. The subject connects a sequence of numbered nodes by using a pen on normal paper. We present an automatic cognitive assessment tool that analyzes samples of the TMT which we record using a digital pen. This enables us to analyze digital pen features that are difficult or impossible to evaluate manually. Our system automatically measures several pen features, including the completion time which is the main performance indicator used by clinicians to score the TMT in practice. In addition, our system provides a structured report of the analysis of the test, for example indicating missed or erroneously connected nodes, thereby offering more objective, transparent and explainable results to the clinician. We evaluate our system with 40 elderly subjects from a geriatrics daycare clinic of a large hospital. Alexander Prange, Michael Barz, Anika Heimann-Steinert, Daniel Sonntag |
CHI | 4 |
| 2021 | Assessing Cognitive Test Performance Using Automatic Digital Pen Features AnalysisabstractMost cognitive assessments, for dementia screening for example, are conducted with a pen on normal paper. We record these tests with a digital pen as part of a new interactive cognitive assessment tool with automatic analysis of pen input. The clinician can, first, observe the sketching process in real-time on a mobile tablet, e.g., in telemedicine settings or to follow Covid-19 distancing regulations. Second, the results of an automatic test analysis are presented to the clinician in real-time, thereby reducing manual scoring effort and producing objective reports. The presented research describes the architecture of our cognitive assessment tool and examines how accurately different machine learning (ML) models can automatically score cognitive tests, without a semantic content analysis. Our system uses a set of more than 170 pen features, calculated directly from the raw digital pen signal. We evaluate our system with 40 subjects from a geriatrics daycare clinic. Using standard ML techniques our feature set outperforms previous approaches on the cognitive tests we consider, i.e., the Clock Drawing, the Rey-Osterrieth Complex Figure, and the Trail Making Test, by automatically scoring tests with up to 82% accuracy in a binary classification task. Alexander Prange, Daniel Sonntag |
UMAP | 2 |
| 2020 | A Study on the Fusion of Pixels and Patient Metadata in CNN-Based Classification of Skin Lesion Images
Fabrizio Nunnari, Chirag Bhuvaneshwara, Abraham Obinwanne Ezema, Daniel Sonntag |
CD-MAKE | 4 |
| 2020 | Visual Search Target Inference in Natural Interaction Settings with Machine LearningabstractVisual search is a perceptual task in which humans aim at identifying a search target object such as a traffic sign among other objects. Search target inference subsumes computational methods for predicting this target by tracking and analyzing overt behavioral cues of that person, e.g., the human gaze and fixated visual stimuli. We present a generic approach to inferring search targets in natural scenes by predicting the class of the surrounding image segment. Our method encodes visual search sequences as histograms of fixated segment classes determined by SegNet, a deep learning image segmentation model for natural scenes. We compare our sequence encoding and model training (SVM) to a recent baseline from the literature for predicting the target segment. Also, we use a new search target inference dataset. The results show that, first, our new segmentation-based sequence encoding outperforms the method from the literature, and second, that it enables target inference in natural settings. Michael Barz, Sven Stauden, Daniel Sonntag |
ETRA | 3 |
| 2020 | Digital Pen Features Predict Task Difficulty and User Performance of Cognitive TestsabstractDigital pen signals were shown to be predictive for cognitive states, cognitive load and emotion in educational settings. We investigate whether low-level pen-based features can predict the difficulty of tasks in a cognitive test and the learner's performance in these tasks, which is inherently related to cognitive load, without a semantic content analysis. We record data for tasks of varying difficulty in a controlled study with children from elementary school. We include two versions of the Trail Making Test (TMT) and six drawing patterns from the Snijders-Oomen Non-verbal intelligence test (SON) as tasks that feature increasing levels of difficulty. We examine how accurately we can predict the task difficulty and the user performance as a measure for cognitive load using support vector machines and gradient boosted decision trees with different feature selection strategies. The results show that our correlation-based feature selection is beneficial for model training, in particular when samples from TMT and SON are concatenated for joint modelling of difficulty and time. Our findings open up opportunities for technology-enhanced adaptive learning. Michael Barz, Kristin Altmeyer, Sarah Malone, Luisa Lauer, Daniel Sonntag |
UMAP | 5 |
| 2019 | Modeling Cognitive Status through Automatic Scoring of a Digital Version of the Clock Drawing TestabstractThe Clock Drawing Test is used as a cognitive assessment tool in geriatrics to detect signs of dementia or to model the progress of stroke recovery. The result is scored manually by a trained professional. We implement the Mendez scoring scheme and create a hierarchy of error categories that model the test characteristics of the clock drawing test, based on a set of impaired clock examples provided by a geriatrics clinic. Using a digital pen we recorded 120 clock samples for evaluating the automatic scoring system, with a total of 2400 error samples distributed over the 20 error classes of the Mendez scoring scheme. Error classes are scored automatically using a handwriting and gesture recognition framework. Results show that we provide a clinically relevant cognitive model for each subject. In addition, we heavily reduce the time spent on manual scoring. We compare manual scoring results with results produced by our automated system. Alexander Prange, Daniel Sonntag |
UMAP | 2 |
| 2019 | An architecture of open-source tools to combine textual information extraction, faceted search and information visualisation
Daniel Sonntag, Hans-Jürgen Profitlich |
Artif. Intell. Medicine | 1 |
| 2018 | Towards a Multimodal Multisensory Cognitive Assessment FrameworkabstractTraditionally, neurocognitive testing is done using pen and paper, which is both expensive and time consuming and often leads to a biased outcome. In this paper, we present an approach towards selecting and digitizing existing cognitive tests and supporting the assessment of cognitive impairments through automated evaluation of different input modalities recorded during the assessments. Our multimodal multisensory framework currently records and analyzes handwriting input captured using a digital pen and electrodermal activity captured by the BITalino sensor board. Using artificial intelligence methods, we aim at analyzing the multisensory data in order to support objective assessments of cognitive impairments. In this work, we describe the current state of our framework and outline future research objectives. Mira Niemann, Alexander Prange, Daniel Sonntag |
CBMS | 3 |
| 2018 | Error-aware gaze-based interfaces for robust mobile gaze interactionabstractGaze estimation error can severely hamper usability and performance of mobile gaze-based interfaces given that the error varies constantly for different interaction positions. In this work, we explore error-aware gaze-based interfaces that estimate and adapt to gaze estimation error on-the-fly. We implement a sample error-aware user interface for gaze-based selection and different error compensation methods: a naïve approach that increases component size directly proportional to the absolute error, a recent model by Feit et al. that is based on the two-dimensional error distribution, and a novel predictive model that shifts gaze by a directional error estimate. We evaluate these models in a 12-participant user study and show that our predictive model significantly outperforms the others in terms of selection rate, particularly for small gaze targets. These results underline both the feasibility and potential of next generation error-aware gaze-based user interfaces. Michael Barz, Florian Daiber, Daniel Sonntag, Andreas Bulling |
ETRA | 3 |
| 2017 | Automatic Extraction of Breast Cancer Information from Clinical ReportsabstractThe majority of clinical data is only available in unstructured text documents. Thus, their automated usage in data-based clinical application scenarios, like quality assurance and clinical decision support by treatment suggestions, is hindered because it requires high manual annotation efforts. In this work, we introduce a system for the automated processing of clinical reports of mamma carcinoma patients that allows for the automatic extraction and seamless processing of relevant textual features. Its underlying information extraction pipeline employs a rule-based grammar approach that is integrated with semantic technologies to determine the relevant information from the patient record. The accuracy of the system, developed with nine thousand clinical documents, reaches accuracy levels of 90% for lymph node status and 69% for the structurally most complex feature, the hormone status. Claudia Bretschneider, Sonja Zillner, Matthias Hammon, Paul Gass, Daniel Sonntag |
CBMS | 5 |
| 2017 | A Digital Pen Based Tool for Instant Digitisation and Digitalisation of Biopsy ProtocolsabstractIn order to improve medical processes in nephrology, we present an application that allows doctors to create biopsy protocols by using a digital pen on a tablet. The biopsy protocol app is seamlessly integrated into the existing infrastructure at the hospital (see figure 1). Compared to other reporting tools, we provide (1) real-time hand-writing/gesture recognition and real-time feedback on the recognition results on the screen; (2) a real-time digitisation into structured data and PDF documents; and (3) the mapping of the transcribed contents into concepts of the Banff classification. Our approach combines the benefits of paper with the automatic digitisation and digitalisation of hand-written user input. A fully digital and mobile approach should empower nephrologists to produce high quality data more effectively and in real-time so that it can be directly used in hospital processes. Alexander Prange, Danilo Schmidt, Daniel Sonntag |
CBMS | 3 |
| 2017 | Integrated Decision Support by Combining Textual Information Extraction, Facetted Search and Information VisualisationabstractThis work focusses on our integration steps of complex and partly unstructured medical data into a clinical research database with subsequent decision support. Our main application is an integrated facetted search tool, followed by information visualisation based on automatic information extraction results from textual documents. We describe the details of our technical architecture (open-source tools), to be replicated at other universities, research institutes, or hospitals. Our exemplary use case is nephrology, where we try to answer questions about the temporal characteristics of sequences and gain significant insight from the data for cohort selection. We report on this case study, illustrating how the application can be used by a clinician and which questions can be answered. Daniel Sonntag, Hans-Jürgen Profitlich |
CBMS | 1 |
| 2017 | Speech-based Medical Decision Support in VR using a Deep Neural Network (Demonstration)abstractWe present a speech dialogue system that facilitates medical decision support for doctors in a virtual reality (VR) application. The therapy prediction is based on a recurrent neural network model that incorporates the examination history of patients. A central supervised patient database provides input to our predictive model and allows us, first, to add new examination reports by a pen-based mobile application on-the-fly, and second, to get therapy prediction results in real-time. This demo includes a visualisation of patient records, radiology image data, and the therapy prediction results in VR. Alexander Prange, Michael Barz, Daniel Sonntag |
IJCAI | 3 |
| 2017 | A Multimodal Dialogue System for Medical Decision Support inside Virtual RealityabstractWe present a multimodal dialogue system that allows doctors to interact with a medical decision support system in virtual reality (VR).We integrate an interactive visualization of patient records and radiology image data, as well as therapy predictions.Therapy predictions are computed in realtime using a deep learning model. Alexander Prange, Margarita Chikobava, Peter Poller, Michael Barz, Daniel Sonntag |
SIGDIAL Conference | 5 |
| 2016 | Digital PI-RADS: Smartphone Sketches for Instant Knowledge Acquisition in Prostate Cancer DetectionabstractIn order to improve reporting practices for the detection of prostate cancer, we present an application that allows urologists to create structured reports by using a digital pen on a smartphone. In this domain, printed documents cannot be easily replaced by computer systems because they contain free-form sketches and textual annotations, and the acceptance of traditional PC reporting tools is rather low among the doctors. Our approach provides an instant knowledge acquisition system by automatically interpreting the written strokes, texts, and sketches. We have incorporated this structured reporting system for MRI of the prostate (PI-RADS). Our system imposes only minimal overhead on traditional form-filling processes and provides for a direct, ontology-based structuring of the user input for semantic search and retrieval applications. Alexander Prange, Daniel Sonntag |
CBMS | 2 |
| 2016 | A model-based approach to qualified process automation for anomaly detection and treatmentabstractModern machineries are becoming complex cyber-physical systems with increasingly intelligent support for process automation. For the dependability and performance, a combination of measures for fault avoidance, robust architecture, and runtime anomaly handling is necessary. These in turn call for a formalization of knowledge across different system lifecycle stages and a provision of novel methods and tools for qualified system synthesis and effective risk management. This paper presents a model-based approach to qualified process automation for the operation and maintenance of production systems. The contribution is centered on the formalizations of a wide range of system concerns, and thereby a consolidation of the rationale behind the design of run-time process logic in BPMN2.0. In particular, the approach allows an integration of formal system descriptions, FTA and FEMA based anomaly analysis, and executable process models for effective anomaly detection and treatment. The approach adopts mature modeling methods and tools through EAST-ADL. In this paper, a prototype tool-chain with MetaEdit+ Domain-Specific Modeling (DSM) Workbench, HiP-HOPS Tool and Camunda BPM Platform is also presented. Dejiu Chen, Dmitri Valeri Panfilenko, Mahmood Reza Khabbazi, Daniel Sonntag |
ETFA | 4 |
| 2016 | BPMN for knowledge acquisition and anomaly handling in CPS for smart factoriesabstractIn this paper, we describe a real-time knowledge acquisition and anomaly handling architecture in a maintenance scenario of an industrial cyber-physical production environment. We use the Business Process Model and Notation (BPMN) for modeling and controlling of the maintenance procedure. Automatic handwriting and pen gesture recognition is combined with a networked smart pen that is used on semantically structured paper forms. Here, we discuss our architecture with BPMN-modeled workflows for real-time data processing. All detected anomalies are automatically associated to the corresponding knowledge sources such as the process model, structure model, business process model and user models; linked Semantic Media-Wiki (SMW) pages are built up accordingly. Our architecture provides for a seamless integration of real-time knowledge sources by smart pen technology and real-time CPS processing by a BPMN engine handling the execution and coordination of the modeled interactions. Dmitri Valeri Panfilenko, Peter Poller, Daniel Sonntag, Sonja Zillner |
ETFA | 3 |
| 2016 | Human Gaze and Focus-of-Attention in Dual Reality Human-Robot CollaborationabstractHuman gaze is an important indicator of the direction of visual focus-of-attention. This information can be very useful in human-robot interaction scenarios. This paper describes a research prototype that utilizes the user's visual attention in a collaborative dual reality environment. We add an additional dimension to existing human-robot interaction scenarios and describe a human-robot collaboration scenario in which the involved human participants are in two different physical locations. One user has the same physical actions space as the robot, the second user is monitoring the setup and thereby collaborating through a virtual reality system. The proposed research prototype monitors the user's visual attention in both real and virtual environments. The prototype also provides information in both the virtual and real environment which results in a dual reality collaboration scenario. As a result, new human-robot interactions brought about by Industrie 4.0, with novel forms of collaborative factory work, can be constructed. Mohammad Mehdi Moniri, Fabio Espinosa, Dieter Merkel, Daniel Sonntag |
Intelligent Environments | 4 |
| 2015 | Interaction and Humans in Internet of Things
Markku Turunen, Daniel Sonntag, Klaus-Peter Engelbrecht, Thomas Olsson 0002, Dirk Schnelle-Walka, Andrés Lucero |
INTERACT (4) | 2 |
| 2015 | Tutorial 3: Intelligent User InterfacesabstractIUI - Intelligent User Interfaces: will introduce you to the design and implementation of Intelligent User Interfaces (IUIs). IUIs aim to incorporate intelligent automated capabilities in human computer interaction, where the net impact is a human-computer interaction that improves performance or usability in critical ways. It also involves designing and implementing an artificial intelligence (AI) component that effectively leverages human skills and capabilities, so that human performance with an application excels. IUIs embody capabilities that have traditionally been associated more strongly with humans than with computers: how to perceive, interpret, learn, use language, reason, plan, and decide. Daniel Sonntag |
ISMAR | 1 |
| 2015 | Halo Content: Context-aware Viewspace Management for Non-invasive Augmented RealityabstractIn mobile augmented reality, text and content placed in a user's immediate field of view through a head worn display can interfere with day to day activities. In particular, messages, notifications, or navigation instructions overlaid in the central field of view can become a barrier to effective face-to-face meetings and everyday conversation. Many text and view management methods attempt to improve text viewability, but fail to provide a non-invasive personal experience for the user. Jason Orlosky, Kiyoshi Kiyokawa, Takumi Toyama, Daniel Sonntag |
IUI | 4 |
| 2015 | Attention Engagement and Cognitive State Analysis for Augmented Reality Text Display FunctionsabstractHuman eye gaze has recently been used as an effective input interface for wearable displays. In this paper, we propose a gaze-based interaction framework for optical see-through displays. The proposed system can automatically judge whether a user is engaged with virtual content in the display or focused on the real environment and can determine his or her cognitive state. With these analytic capacities, we implement several proactive system functions including adaptive brightness, scrolling, messaging, notification, and highlighting, which would otherwise require manual interaction. The goal is to manage the relationship between virtual and real, creating a more cohesive and seamless experience for the user. We conduct user experiments including attention engagement and cognitive state analysis, such as reading detection and gaze position estimation in a wearable display towards the design of augmented reality text display applications. The results from the experiments show robustness of the attention engagement and cognitive state analysis methods. A majority of the experiment participants (8/12) stated the proactive system functions are beneficial. Takumi Toyama, Daniel Sonntag, Jason Orlosky, Kiyoshi Kiyokawa |
IUI | 2 |
| 2015 | ModulAR: Eye-Controlled Vision Augmentations for Head Mounted DisplaysabstractIn the last few years, the advancement of head mounted display technology and optics has opened up many new possibilities for the field of Augmented Reality. However, many commercial and prototype systems often have a single display modality, fixed field of view, or inflexible form factor. In this paper, we introduce Modular Augmented Reality (ModulAR), a hardware and software framework designed to improve flexibility and hands-free control of video see-through augmented reality displays and augmentative functionality. To accomplish this goal, we introduce the use of integrated eye tracking for on-demand control of vision augmentations such as optical zoom or field of view expansion. Physical modification of the device's configuration can be accomplished on the fly using interchangeable camera-lens modules that provide different types of vision enhancements. We implement and test functionality for several primary configurations using telescopic and fisheye camera-lens systems, though many other customizations are possible. We also implement a number of eye-based interactions in order to engage and control the vision augmentations in real time, and explore different methods for merging streams of augmented vision into the user's normal field of view. In a series of experiments, we conduct an in depth analysis of visual acuity and head and eye movement during search and recognition tasks. Results show that methods with larger field of view that utilize binary on/off and gradual zoom mechanisms outperform snapshot and sub-windowed methods and that type of eye engagement has little effect on performance. Jason Orlosky, Takumi Toyama, Kiyoshi Kiyokawa, Daniel Sonntag |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2014 | A natural interface for multi-focal plane head mounted displays using 3D gazeabstractIn mobile augmented reality (AR), it is important to develop interfaces for wearable displays that not only reduce distraction, but that can be used quickly and in a natural manner. In this paper, we propose a focal-plane based interaction approach with several advantages over traditional methods designed for head mounted displays (HMDs) with only one focal plane. Using a novel prototype that combines a monoscopic multi-focal plane HMD and eye tracker, we facilitate interaction with virtual elements such as text or buttons by measuring eye convergence on objects at different depths. This can prevent virtual information from being unnecessarily overlaid onto real world objects that are at a different range, but in the same line of sight. We then use our prototype in a series of experiments testing the feasibility of interaction. Despite only being presented with monocular depth cues, users have the ability to correctly select virtual icons in near, mid, and far planes in 98.6% of cases. Takumi Toyama, Daniel Sonntag, Jason Orlosky, Kiyoshi Kiyokawa |
AVI | 2 |
| 2014 | The Medical Cyber-physical Systems Activity at EIT: A Look under the HoodabstractIn this paper, we describe how we combine active and passive user input modes in clinical environments for knowledge discovery and knowledge acquisition towards decision support in clinical environments. Active input modes include digital pens, smartphones, and automatic handwriting recognition for a direct digitalisation of patient data. Passive input modes include sensors of the clinical environment and or mobile smartphones. This combination for knowledge acquisition and decision support (while using machine learning techniques) has not yet been explored in clinical environments and is of specific interest because it combines previously unconnected information sources for individualised treatments. The innovative aspect is a holistic view on individual patients based on ontologies, terminologies, and textual patient records whereby individual active and passive real-time patient data can be taken into account for improving clinical decision support. Daniel Sonntag, Sonja Zillner, Samarjit Chakraborty, András Lörincz, Esko Strömmer, Luciano Serafini |
CBMS | 1 |
| 2014 | A mixed reality head-mounted text translation system using eye gaze inputabstractEfficient text recognition has recently been a challenge for augmented reality systems. In this paper, we propose a system with the ability to provide translations to the user in real-time. We use eye gaze for more intuitive and efficient input for ubiquitous text reading and translation in head mounted displays (HMDs). The eyes can be used to indicate regions of interest in text documents and activate optical-character-recognition (OCR) and translation functions. Visual feedback and navigation help in the interaction process, and text snippets with translations from Japanese to English text snippets, are presented in a see-through HMD. We focus on travelers who go to Japan and need to read signs and propose two different gaze gestures for activating the OCR text reading and translation function. We evaluate which type of gesture suits our OCR scenario best. We also show that our gaze-based OCR method on the extracted gaze regions provide faster access times to information than traditional OCR approaches. Other benefits include that visual feedback of the extracted text region can be given in real-time, the Japanese to English translation can be presented in real-time, and the augmentation of the synchronized and calibrated HMD in this mixed reality application are presented at exact locations in the augmented user view to allow for dynamic text translation management in head-up display systems. Takumi Toyama, Daniel Sonntag, Andreas Dengel 0001, Takahiro Matsuda 0001, Masakazu Iwamura, Koichi Kise |
IUI | 2 |
| 2013 | Integrating Digital Pens in Breast Imaging for Instant Knowledge AcquisitionabstractFuture radiology practices assume that the radiology reports should be uniform, comprehensive, and easily managed. This means that reports must be “readable” to humans and machines alike. In order to improve reporting practices in breast imaging, we allow the radiologist to write structured reports with a special pen on paper with an invisible dot pattern. In this way, we provide a knowledge acquisition system for printed mammography patient forms for the combined work with printed and digital documents. In this domain, printed documents cannot be easily replaced by computer systems because they contain free-form sketches and textual annotations, and the acceptance of traditional PC reporting tools is rather low among the doctors. This is due to the fact that current electronic reporting systems significantly add to the amount of time it takes to complete the reports. We describe our real-time digital paper application and focus on the use case study of our deployed application. We think that our results motivate the design and implementation of intuitive pen-based user interfaces for the medical reporting process and similar knowledge work domains. Our system imposes only minimal overhead on traditional form-filling processes and provides for a direct, ontology-based structuring of the user input for semantic search and retrieval applications, as well as other applied artificial intelligence scenarios which involve manual form-based data acquisition. Daniel Sonntag, Matthias Hammon, Alexander Cavallaro |
IAAI | 1 |
| 2012 | Clinical Trial and Disease Search with Ad Hoc Interactive Ontology Alignments
Daniel Sonntag, Jochen Setz, Maha Ahmed Baker, Sonja Zillner |
ESWC | 1 |
| 2012 | RadSpeech's mobile dialogue system for radiologistsabstractWith RadSpeech, we aim to build the next generation of intelligent, scalable, and user-friendly semantic search interfaces for the medical imaging domain, based on semantic technologies. Ontology-based knowledge representation is used not only for the image contents, but also for the complex natural language understanding and dialogue management process. This demo shows a speech-based annotation system for radiology images and focuses on a new and effective way to annotate medical image regions with a specific medical, structured, diagnosis while using speech and pointing gestures on the go. Daniel Sonntag, Christian Husodo-Schulz, Christian Reuschling, Luis Galárraga |
IUI | 1 |
| 2011 | Aligning medical ontologies by axiomatic models, corpus linguistic syntactic rules and context informationabstractWe investigate formal semantics, as well as corpus linguistics and context based rules for ontology alignment in the medical domain. Semantic image retrieval should provide the basis for help in clinical decision support and computer aided diagnosis. Medical image and data retrieval for anatomy, diseases, or other patient-centric information require a comprehensive mapping of medical ontologies. We enhanced previous approaches of ontology matching for supporting collaboration by incorporating domain-specific context information of the application domain. The evaluation shows that axiomatic models in combination with syntactic rules and context information are very effective in terms of precision, recall, and F1 measure. Sonja Zillner, Daniel Sonntag |
CBMS | 2 |
| 2011 | Digital pen in mammography patient formsabstractWe present a digital pen based interface for clinical radiology reports in the field of mammography. It is of utmost importance in future radiology practices that the radiology reports be uniform, comprehensive, and easily managed. This means that reports must be "readable" to humans and machines alike. In order to improve reporting practices in mammography, we allow the radiologist to write structured reports with a special pen on paper with an invisible dot pattern. A handwriting software takes care of the interpretation of the written report which is transferred into an ontological representation. In addition, a gesture recogniser allows radiologists to encircle predefined annotation suggestions which turns out to be the most beneficial feature. The radiologist can (1) provide the image and image region annotations mapped to a FMA, RadLex, or ICD10 code, (2) provide free text entries, and (3) correct/select annotations while using multiple gestures on the forms and sketch regions. The resulting, automatically generated PDF report is then stored in a semantic backend system for further use and contains all transcribed annotations as well as all free form sketches. Daniel Sonntag, Marcus Liwicki |
ICMI | 1 |
| 2011 | Towards learned feedback for enhancing trust in information seeking dialogue for radiologistsabstractDialogue-based Question Answering (QA) in the context of information seeking applications is a highly complex user interaction task. QA systems normally include various natural language processing components (i.e., components for question classification and information extraction) and information retrieval components. This paper presents a new approach to equip a multimodal QA system for radiologists with some form of self-knowledge about the expected dialogue processing behaviour and the results themselves. The learned models are used to provide feedback of the QA process, i.e., what the system is doing and delivers as results. The resulting automatic feedback behaviour should enhance the user's trust in the system. To this end, examples of the learned feedback are provided in the context of the generation of system-initiative dialogue feedback to a radiologist's questions. Daniel Sonntag |
IUI | 1 |
| 2011 | Interactive paper for radiology findingsabstractThis paper presents a pen-based interface for clinical radiologists. It is of utmost importance in future radiology practices that the radiology reports be uniform, comprehensive, and easily managed. This means that reports must be "readable" to humans and machines alike. In order to improve reporting practices, we allow the radiologist to write structured reports with a special pen on normal paper. A handwriting recognition and interpretation software takes care of the interpretation of the written report which is transferred into an ontological representation. The resulting report is then stored in a semantic backend system for further use. We will focus on the pen-based interface and new interaction possibilities with gestures in this scenario. Daniel Sonntag, Marcus Liwicki |
IUI | 1 |
| 2011 | Mobile thumb interaction and speechabstractThe design of future spoken and multimodal interfaces should not only compare voice with other modalities of interaction. Instead, it should take screen-based smartphones as a basis and add new, gesture-based anthropocentric interaction forms to it. A potential speech-based interaction can then be smoothly integrated. We focus on the thumb's role while carrying a handheld device in the left or right hand. Mikhail Blinov, Matthieu Deru, Daniel Sonntag |
Mobile HCI | 3 |
| 2010 | Representing the International Classification of Diseases Version 10 in OWL
Manuel Möller, Michael Sintek, Ralf Biedert, Patrick Ernst, Andreas Dengel 0001, Daniel Sonntag |
KEOD | 6 |
| 2010 | A Spatio-anatomical Medical Ontology and Automatic Plausibility Checks
Manuel Möller, Daniel Sonntag, Patrick Ernst |
IC3K | 2 |
| 2010 | Modeling the International Classification of Diseases (ICD-10) in OWL
Manuel Möller, Daniel Sonntag, Patrick Ernst |
IC3K | 2 |
| 2010 | Prototyping Semantic Dialogue Systems for RadiologistsabstractIn the future, speech-based semantic image retrieval and annotation of medical images should provide the basis for clinical decision support and help in computer aided diagnosis. In this paper, we describe today's clinical workflow and interaction requirements and present a semantic dialogue system installation for radiologists. Our research focus is on the interaction design in combination with the implementation of our prototype system for patient image search and image annotation while using a speech based dialogue shell and a big touchscreen in the radiology environment. Ontology modeling provides the backbone for knowledge representation in the dialogue shell and the specific medical application domain. Daniel Sonntag, Manuel Möller |
Intelligent Environments | 1 |
| 2010 | A multimodal dialogue mashup for medical image semanticsabstractThis paper presents a multimodal dialogue mashup where different users are involved in the use of different user interfaces for the annotation and retrieval of medical images. Our solution is a mashup that integrates a multimodal interface for speech-based annotation of medical images and dialogue-based image retrieval with a semantic image annotation tool for manual annotations on a desktop computer. A remote RDF repository connects the annotation and querying task into a common framework and serves as the semantic backend system for the advanced multimodal dialogue a radiologist can use. Daniel Sonntag, Manuel Möller |
IUI | 1 |
| 2010 | Speech Grammars for Textual Entailment Patterns in Multimodal Question Answering
Daniel Sonntag, Bogdan Sacaleanu |
LREC | 1 |
| 2009 | Introspection and Adaptable Model Integration for Dialogue-based Question Answering
Daniel Sonntag |
IJCAI | 1 |
| 2009 | New business to business interaction: shake your iPhone and speak to itabstractWe present a new multimodal interaction sequence for a mobile multimodal Business-to-Business interaction system. A mobile client application on the iPhone supports users in accessing an online service marketplace and allows business experts to intuitively search and browse for services using natural language speech and gestures while on the go. For this purpose, we utilize an ontology-based multimodal dialogue platform as well as an integrated trainable gesture recognizer. Daniel Porta, Daniel Sonntag, Robert Neßelrath |
Mobile HCI | 2 |
| 2008 | Semiotic-based Ontology Evaluation Tool (S-OntoEval)
Renata Queiroz Dividino, Massimo Romanelli, Daniel Sonntag |
LREC | 3 |
| 2006 | A Multimodal Result Ontology for Integrated Semantic Web Dialogue Applications
Daniel Sonntag, Massimo Romanelli |
LREC | 1 |
| 2005 | A look under the hood: design and development of the first SmartWeb system demonstratorabstractExperience shows that decisions in the early phases of the development of a multimodal system prevail throughout the life-cycle of a project. The distributed architecture and the requirement for robust multimodal interaction in our project SmartWeb resulted in an approach that uses and extends W3C standards like EMMA and RDFS. These standards for the interface structure and content allowed us to integrate available tools and techniques. However, the requirements in our system called for various extensions, e.g., to introduce result feedback tags for an extended version of EMMA. The interconnection framework depends on a commercial telephone voice dialog system platform for the dialog-centric components while the information access processes are linked using web service technology. Also in the area of this underlying infrastructure, enhancements and extensions were necessary. The first demonstration system is operable now and will be presented at the Football World Cup 2006 in Germany. Norbert Reithinger, Simon Bergweiler, Ralf Engel, Gerd Herzog, Norbert Pfleger, Massimo Romanelli, Daniel Sonntag |
ICMI | 7 |
| 2005 | An integration framework for a mobile multimodal dialogue system accessing the semantic webabstractAdvanced intelligent multimodal interface systems usually comprise many sub-systems. For the integration of already existing software components in the SMARTWEB 1 system we developed an integration framework, the IHUB. It allows us to reuse already existing components for interpretation and processing of multimodal user interactions. The framework facilitates the integration of the user in the interpretation loop by controlling the message flow in the system which is important in our domain, the multimodal access to the Semantic Web. A technical evaluation of the framework shows the efficient routing of messages to make real-time interactive editing of semantic queries possible. 1. Norbert Reithinger, Daniel Sonntag |
INTERSPEECH | 2 |