Sebastian Simon

dblp:191/6817 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Designing Advanced Interfaces for Learning and Teaching
abstract
The goal of this workshop is to present an overview of ongoing research on Human–Computer Interaction (HCI) for teaching and learning. The workshop focuses on how emerging interaction paradigms, such as adaptive interfaces, multimodal interaction, and game-based learning, can support richer, more engaging, and more inclusive Technology-Enhanced Learning (TEL) experiences. The discussion will be structured around three central research directions that remain insufficiently explored: facilitating the design and evaluation of TEL systems that promote effective learning and quality of experience through innovative HCI approaches; adapting TEL systems to the complexity of educational environments; and modeling learners and users in TEL systems to account for their diversity, their preferences, and constraints. The ultimate goal of the workshop is to bring together researchers interested in discussing and converging on the future challenges and opportunities of the field.
Audrey Serna, Antonio Bucchiarone, Iza Marfisi, Sebastian Simon, Arnaud Prouzeau, David Bertolo, Élise Lavoué, Alina Glushkova
AVI4
2026 Disagreement as Data: Reasoning Trace Analytics in Multi-Agent Systems
abstract
Learning analytics researchers often analyze qualitative student data such as coded annotations or interview transcripts to understand learning processes. With the rise of generative AI, fully automated and human–AI workflows have emerged as promising methods for analysis. However, methodological standards to guide such workflows remain limited. In this study, we propose that reasoning traces generated by large language model (LLM) agents, especially within multi-agent systems, constitute a novel and rich form of process data to enhance interpretive practices in qualitative coding. We apply cosine similarity to LLM reasoning traces to systematically detect, quantify, and interpret disagreements among agents, reframing disagreement as a meaningful analytic signal. Analyzing nearly 10,000 instances of agent pairs coding human tutoring dialog segments, we show that LLM agents’ semantic reasoning similarity robustly differentiates consensus from disagreement and correlates with human coding reliability. Qualitative analysis guided by this metric reveals nuanced instructional sub-functions within codes and opportunities for conceptual codebook refinement. By integrating quantitative similarity metrics with qualitative review, our method bears potential to improve and accelerate the process of establishing inter-rater reliability during coding by surfacing interpretive ambiguity, especially when LLMs collaborate with humans. We discuss how reasoning-trace disagreements represent a valuable new class of analytic signals advancing methodological rigor and interpretive depth in educational research.
Elham Tajik, Conrad Borchers, Bahar Shahrokhian, Sebastian Simon, Ali Keramati, Sonika Pal, Sreecharan Sankaranarayanan
LAK4
2025 Comparing a Human's and a Multi-Agent System's Thematic Analysis: Assessing Qualitative Coding Consistency
Sebastian Simon, Sreecharan Sankaranarayanan, Elham Tajik, Conrad Borchers, Bahar Shahrokhian, Francesco Balzan, Sebastian Strauss, Sree Aurovindh Viswanathan, Amine Hatun Atas, Mia Carapina, Berkan Celik
AIED (3)1
2025 Themes of Building LLM-Based Applications for Production: A Practitioner's View
abstract
Background: Large language models (LLMs) have become a paramount interest of researchers and practitioners alike, yet a comprehensive overview of key considerations for those developing LLM-based systems is lacking. This study addresses this gap by collecting and mapping the topics practitioners discuss online, offering practical insights into where priorities lie in developing LLM-based applications. Method: We collected 189 videos from 2022 to 2024 by practitioners actively developing such systems and discussing various aspects they encounter during development and deployment of LLMs in production. We analyzed the transcripts using BERTopic, then manually sorted and merged the generated topics into themes, leading to a total of 20 topics in 8 themes. Results: The most prevalent topics fall within the theme Design & Architecture, with a strong focus on retrieval-augmented generation (RAG) systems. Other frequently discussed topics include model capabilities and enhancement techniques (e.g., finetuning, prompt engineering), infrastructure and tooling, and risks and ethical challenges. Implications: Our results highlight current discussions and challenges in deploying LLMs in production. This way, we provide a systematic overview of key aspects practitioners should be aware of when developing LLM-based applications. We further highlight topics of interest for academics where further research is needed.
Alina Mailach, Sebastian Simon, Johannes Dorn, Norbert Siegmund
CAIN2
2025 On Automating Configuration Dependency Validation via Retrieval-Augmented Generation
abstract
Configuration dependencies arise when multiple technologies in a software system require coordinated settings for correct interplay. Existing approaches for detecting such dependencies often yield high false-positive rates, require additional validation mechanisms, and are typically limited to specific projects or technologies. Recent work that incorporates large language models (LLMs) for dependency validation still suffers from inaccuracies due to project- and technology-specific variations, as well as from missing contextual information.In this work, we propose to use retrieval-augmented generation (RAG) systems for configuration dependency validation, which allows to incorporate additional project- and technology-specific context information. Specifically, we evaluate whether RAG can improve LLM-based validation of configuration dependencies and what contextual information are needed to overcome the static knowledge base of LLMs. To this end, we conducted a large empirical study on validating configuration dependencies using RAG. Our evaluation shows that vanilla LLMs already demonstrate solid validation abilities, while RAG has only marginal or even negative effects on the validation performance of the models. By incorporating tailored contextual information into the RAG system–derived from a qualitative analysis of validation failures–we achieve significantly more accurate validation results across all models, with an average precision of 0.84 and recall of 0.70, representing improvements of 35% and 133% over vanilla LLMs, respectively. In addition, these results offer two important insights: Simplistic RAG systems may not benefit from additional information if it is not tailored to the task at hand, and it is often unclear upfront what kind of information yields improved performance.
Sebastian Simon, Alina Mailach, Johannes Dorn, Norbert Siegmund
ASE1
2024 Peephole Technology for Mobile Collaborative Learning: An In- Classroom Exploratory Study
abstract
International audience
Sebastian Simon, Iza Marfisi, Sébastien George
CSEDU (1)1
2024 Exploring the Role of the Portable Stimulus Standard in Enhancing Security Property Verification
abstract
The security of the hardware design has been a critical concern in the semiconductor industry in the recent years. Modern integrated circuit designs of the security-critical application contain numerous sensitive data that needs to be protected against adversary attacks. An Intellectual Property (IP) block such as security controllers is integrated into the System on Chip (SoC) to protect the sensitive asset and its integrity with certain measures that implement cryptographic algorithms. Verification of security properties of this type of IP block is especially important and crucial to ensure its trustworthiness. This paper presents an approach to enhance security property verification of the design using Portable Test and Stimulus Standard (PSS) along with existing testbench. PSS defines a specification to generate a single representation of stimulus and test scenarios, which is usable by a various user across different level of integration. We have modelled functionality and security property of the Security Controller IP at abstract level using PSS. Then, PSS model has facilitated to generate a set of concrete test cases in Specman/e for various scenarios. The generated test cases are executed in verification environment and analyzed the results that verify security properties.
Jaimini Nagar, Thorsten Dworzak, Sebastian Simon, Ulrich Heinkel, Djones Lettnin
VLSI-SoC3
2023 Exploring Hyperparameter Usage and Tuning in Machine Learning Research
abstract
The success of machine learning (ML) models depends on careful experimentation and optimization of their hyperparameters. Tuning can affect the reliability and accuracy of a trained model and is the subject of ongoing research. However, little is known on whether and how hyperparameters are used and optimized in research practice. This lack of knowledge not only limits the adoption of best practices for tuning in research, but also affects the reproducibility of published results. Our research systematically analyzes the use and tuning of hyperparameters in ML publications. For this, we analyze 2000 code repositories and their associated research papers from Papers with Code. We compare the use and tuning of hyperparameters of three widely used ML libraries: scikit-learn, TensorFlow, and PyTorch. Our results show that the most of the available hyperparameters remain untouched, and those that have been changed use constant values. In particular, there is a significant difference between tuning hyperparameters and the reporting about it in the corresponding research papers. Our results suggest that there is a need for improved research and reporting practices when using ML methods to improve the reproducibility of published results.
Sebastian Simon, Nikolay Kolyada, Christopher Akiki, Martin Potthast, Benno Stein 0001, Norbert Siegmund
CAIN1
2023 Obfuscation Padding Schemes that Minimize Rényi Min-Entropy for Privacy
Sebastian Simon, Cezara Petrui, Carlos Antonio Pinzón, Catuscia Palamidessi
ISPEC1
2023 CfgNet: A Framework for Tracking Equality-Based Configuration Dependencies Across a Software Project
abstract
Modern software development incorporates various technologies, such as containerization, CI/CD pipelines, and build tools, which have to be jointly configured to enable building, testing, deployment, and execution of software systems. The vast configuration space spans several different configuration artifacts with their own syntax and semantics, encoding hundreds of configuration options and their values. The interplay of these technologies requires some level of coordination, which is realized by matching configurations. That is, configuration options and their according values may depend on other options and values from entirely different technologies and artifacts. This creates non-obvious configuration dependencies that are hard to track. The missing awareness and overview of such configuration dependencies across diverse configuration artifacts, tools, and frameworks can lead to dependency conflicts and severe configuration errors. We proposeCFGNET, a framework that models the configuration landscape of a software project as a configuration network in an extensible and artifact-independent way. This way, we enable the early detection of possible dependency violations and proactively prevent misconfigurations during software development and maintenance. In a literature study, we found that the most common form of dependencies is the equality of values of different options. Based on this result, we developed an equality-based linker to determine dependent options across different artifacts. To demonstrate the extensibility of our framework, we also implemented nine plugins for popular technologies, such as Maven and Docker. To evaluate our approach, we injected and violated five real-world configuration dependencies extracted from Stack Overflow, which we support with our technology plugins, in five subject systems.CFGNETfound all injected dependency violations and four additional ones already present in these systems. Moreover, we appliedCFGNETto the commit history of 50 repositories selected from GitHub and found dependency conflicts in about two thirds of these repositories. We manually inspected 883 conflicts, with about 89 % true positives, demonstrating the need to reliably track cross-technology configuration dependencies and prevent their misconfiguration.
Sebastian Simon, Nicolai Ruckel, Norbert Siegmund
IEEE Trans. Software Eng.1
2022 Towards an Authoring Tool to Help Teachers Create Mobile Collaborative Learning Games for Field Trips
Iza Marfisi, Aurélie Laine, Pierre Laforcade, Sébastien George, Sebastian Simon, Madeth May, Moez Zammit, Ludovic Blin
EC-TEL5
2022 A Conceptual Framework for Creating Mobile Collaboration Tools
Sebastian Simon, Iza Marfisi, Sébastien George
EC-TEL1
2019 Symbolic QED Pre-silicon Verification for Automotive Microcontroller Cores: Industrial Case Study
abstract
We present an industrial case study that demonstrates the practicality and effectiveness of Symbolic Quick Error Detection (Symbolic QED) in detecting logic design flaws (logic bugs) during pre-silicon verification. Our study focuses on several microcontroller core designs (~1,800 flip-flops, ~70,000 logic gates) that have been extensively verified using an industrial verification flow and used for various commercial automotive products. The results of our study are as follows: 1. Symbolic QED detected all logic bugs in the designs that were detected by the industrial verification flow (which includes various flavors of simulation-based verification and formal verification). 2. Symbolic QED detected additional logic bugs that were not recorded as detected by the industrial verification flow. (These additional bugs were also perhaps detected by the industrial verification flow.)3.Symbolic QED enables significant design productivity improvements: (a) 8X improved (i.e., reduced) verification effort for a new design (8 person-weeks for Symbolic QED vs. 17 person-months using the industrial verification flow). (b) 60X improved verification effort for subsequent designs (2 person-days for Symbolic QED vs. 4-7 person-months using the industrial verification flow). (c) Quick bug detection (runtime of 20 seconds or less), together with short counterexamples (10 or fewer instructions) for quick debug, using Symbolic QED.
Eshan Singh, Keerthikumara Devarajegowda, Sebastian Simon, Ralf Schnieder, Karthik Ganesan 0001, Mohammad Rahmani Fadiheh, Dominik Stoffel, Wolfgang Kunz, Clark W. Barrett, Wolfgang Ecker, Subhasish Mitra
DATE3
2017 Coverage-driven mixed-signal verification of smart power ICs in a UVM environment
abstract
The complexity of integrated circuits is continuously increasing, leading to a growing demand for methodologies that offer comprehensive mixed-signal verification concepts. However, compared to the highly automated verification methodologies in the digital domain, pre-silicon verification in the analog domain usually implies a substantial amount of manual work and computational effort. In order to meet the rising challenges, various attempts were made to extend well-established approaches from the field of digital verification to also enable systematic mixed-signal verification. However, no methodology could be identified that meets our requirements for high reusability and maintainability, tool independence as well as capabilities for functional coverage collection. For this reason, we propose a mixed-signal verification methodology that covers the aforementioned as well as additional aspects required for a successful coverage closure. The presented concept is applied to a smart power application to demonstrate its potential and outline the gained benefits.
Sebastian Simon, Deeksha Bhat, Alexander W. Rath, Jérôme Kirscher, Linus Maurer
ETS1
2016 Automatically comparing analog behavior using Earth Mover's Distance
abstract
Evaluating the outcome of analog simulations is a common, mostly manually carried out task in the pre-silicon verification process of mixed-signal ICs. Its non-automated nature makes it an error-prone and time-consuming procedure. For this very reason, we introduce a novel approach for performing this evaluation automatically resulting in significantly reduced turnaround times as well as a considerably increased reliability of verification results. The presented concept is motivated by an algorithm that is used in optical pattern recognition and is called Earth Mover's Distance. Furthermore, we compare our approach with already existing algorithms, namely Fréchet Distance and Pearson Coefficient, in order to analyze its capability. Finally, we present a case study in which we prove the algorithm by applying it to the results of a mixed-signal simulation at chip-level demonstrating the efficiency of our approach.
Alexander W. Rath, Sebastian Simon, Volkan Esen, Wolfgang Ecker
VLSI-SoC2