Vera Schmitt

dblp:295/6533 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0002-9735-6956ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Uncovering Temporal Framing in the News
abstract
Tarek Mahmoud, Veronika Solopova, Premtim Sahitaj, Ariana Sahitaj, Max Upravitelev, Mervat Abassy, Hana Fatima Shaikh, Neda Foroutan, Vera Schmitt, Preslav Nakov. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tarek Mahmoud, Veronika Solopova, Premtim Sahitaj, Ariana Sahitaj, Max Upravitelev, Mervat Abassy, Hana Fatima Shaikh, Neda Foroutan, Vera Schmitt, Preslav Nakov
ACL (1)9
2026 From Weights to Activations: Is Steering the Next Frontier of Adaptation?
abstract
Simon Ostermann, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich, Sebastian Lapuschkin, Wojciech Samek, Vera Schmitt. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Simon Ostermann 0002, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich, Sebastian Lapuschkin, Wojciech Samek, Vera Schmitt
ACL (1)7
2026 Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation
abstract
Qianli Wang, Van Bach Nguyen, Yihong Liu, Fedor Splitt, Nils Feldhus, Christin Seifert, Hinrich Schuetze, Sebastian Möller, Vera Schmitt. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Van Bach Nguyen, Yihong Liu 0001, Fedor Splitt, Nils Feldhus, Christin Seifert, Hinrich Schütze, Sebastian Möller 0001, Vera Schmitt
ACL (1)9
2026 Extending Information Bottleneck Attribution to Video Sequences for Deepfake Detection
Veronika Solopova, Lucas Schmidt, Vera Schmitt, Dorothea Kolossa
IDA3
2026 From Articles to Premises: Building PrimeFacts, an Extraction Methodology and Resource for Fact-Checking Evidence
abstract
Fact-checking articles encode rich supporting evidence and reasoning, yet this evidence remains largely inaccessible to automated verification systems due to unstructured presentation. We introduce PrimeFacts, a methodology and resource for extracting fine-grained evidence from full fact-checking articles. We compile 13,106 PolitiFact articles with claims, verdicts, and all referenced sources, and we identify 49,718 in-article hyperlinks as natural anchors to pinpoint key evidence. Our framework leverages large language models (LLMs) to rewrite these anchor sentences into stand-alone, context-independent premises and investigates the extraction of additional implicit evidence. In evaluations on cross-article evidence retrieval and claim verification, the extracted premises substantially improve performance. Decontextualized evidence yields higher retrievability, achieving up to a 30 percent relative gain in Mean Reciprocal Rank over verbatim sentences, and using the evidence for verdict prediction raises Macro-F1 by 10-20 points over the baseline. These gains are consistent across different verdict granularities (2-class vs. 5-class) and model architectures. A qualitative analysis indicates that the decontextualized premises remain faithful to the original sources. Our work highlights the promise of reusing fact-checkers' evidence for automation and provides a large-scale resource of structured evidence from real-world fact-checks.
Premtim Sahitaj, Jawan Kolanowski, Ariana Sahitaj, Veronika Solopova, Max Upravitelev, Daniel Röder, Iffat Maab, Junichi Yamagishi, Sebastian Möller 0001, Vera Schmitt
LREC10
2026 The 5th ACM International Workshop on Multimedia AI against Disinformation (MAD'26)
abstract
Verifying the authenticity of media has become an increasingly challenging task. Rapid advances in AI-generated content, spanning modalities like text, images, video, audio have significantly blurred the line between genuine and synthetic information. Nowadays, powerful foundation models can easily be leveraged to create, amplify and disseminate information at scale, enabling disinformation campaigns, defamation, or impersonation. This results in the erosion of trust in online information, which poses a great threat to society. The MAD’26 workshop seeks to address this problem by bringing together researchers and practitioners from diverse disciplines, united by the goal of combating disinformation through AI-driven approaches. Now in its fifth edition, the workshop aims to cultivate a collaborative environment that encourages the exchange of ideas, methodologies, and practical experiences. The workshop focuses on key research directions, including the detection of AI-generated and manipulated content, the analysis of disinformation propagation, and the examination of its broader societal impact.
Dan-Cristian Stanciu, Symeon Papadopoulos, Giorgos Kordopatis-Zilos, Bogdan Ionescu, Adrian Popescu 0001, Roberto Caldelli, Milica Gerhardt, Vera Schmitt
ICMR8
2025 Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem
abstract
Natural language explanations (NLEs) are vital for elucidating the reasoning behind large language model (LLM) decisions. Many techniques have been developed to generate NLEs using LLMs. However, like humans, LLMs might not always produce optimal NLEs on first attempt. Inspired by human learning processes, we introduce Cross-Refine, which employs role modeling by deploying two LLMs as generator and critic, respectively. The generator outputs a first NLE and then refines this initial explanation using feedback and suggestions provided by the critic. Cross-Refine does not require any supervised training data or additional training. We validate Cross-Refine across three NLP tasks using three state-of-the-art open-source LLMs through automatic and human evaluation. We select Self-Refine (Madaan et al., 2023) as the baseline, which only utilizes self-feedback to refine the explanations. Our findings from automatic evaluation and a user study indicate that Cross-Refine outperforms Self-Refine. Meanwhile, Cross-Refine can perform effectively with less powerful LLMs, whereas Self-Refine only yields strong results with ChatGPT. Additionally, we conduct an ablation study to assess the importance of feedback and suggestions. Both of them play an important role in refining explanations. We further evaluate Cross-Refine on a bilingual dataset in English and German.
Tatiana Anikina, Nils Feldhus, Simon Ostermann 0002, Sebastian Möller 0001, Vera Schmitt
COLING6
2025 Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
abstract
Counterfactual examples are widely employed to enhance the performance and robustness of large language models (LLMs) through counterfactual data augmentation (CDA). However, the selection of the judge model used to evaluate label flipping, the primary metric for assessing the validity of generated counterfactuals for CDA, yields inconsistent results. To decipher this, we define four types of relationships between the counterfactual generator and judge models: being the same model, belonging to the same model family, being independent models, and having an distillation relationship. Through extensive experiments involving two state-of-the-art LLM-based methods, three datasets, four generator models, and 15 judge models, complemented by a user study (n = 90), we demonstrate that judge models with an independent, non-fine-tuned relationship to the generator model provide the most reliable label flipping evaluations. Relationships between the generator and judge models, which are closely aligned with the user study for CDA, result in better model performance and robustness. Nevertheless, we find that the gap between the most effective judge models and the results obtained from the user study remains considerably large. This suggests that a fully automated pipeline for CDA may be inadequate and requires human intervention.
Van Bach Nguyen, Nils Feldhus, Luis Felipe Villa-Arenas, Christin Seifert, Sebastian Möller 0001, Vera Schmitt
INLG7
2025 MAD'25: 4th ACM International Workshop on Multimedia AI against Disinformation
abstract
2148
Dan-Cristian Stanciu, Bogdan Ionescu, Symeon Papadopoulos, Giorgos Kordopatis-Zilos, Adrian Popescu 0001, Roberto Caldelli, Milica Gerhardt, Vera Schmitt
ICMR8
2024 From Construction to Application: Advancing Argument Mining with the Large-Scale KIALOPRIME Dataset
abstract
In this study, we introduce KIALOPRIME, a novel large-scale dataset comprising 5,687 argument discussion graphs with a total of 1,088,801 of supporting, attacking, and neutral argument relations, derived from the structured debates of the online discussion platform Kialo.com. This dataset facilitates in-depth analysis of argument structures and the dynamics of discourse, serving as a substantial resource for computational argumentation research. We explore argument inference through traditional sequence classification and a modern generative reasoning based approach, employing an open-source mixture of experts LLM to interpret and enrich each argument pair with high-quality synthetic elaborations about the argumentative interaction. We achieve baseline results of F1 .899 and .840 within discussions and F1 .908 and .840 across discussions for the argument relation and elaboration classification models, respectively. While the elaboration-based model scores slightly lower on the classification task, we highlight areas of improvement to better capture the hidden complexities of argumentative text. These initial findings are promising as they not only establish robust benchmarks for future studies but also demonstrate the potential for using generative reasoning to provide a more insightful analysis of argument relations.
Premtim Sahitaj, Ramon Ruiz-Dolz, Ariana Sahitaj, Ata Nizamoglu, Vera Schmitt, Salar Mohtaj, Sebastian Möller 0001
COMMA5
2024 Improving a Pupillometry Signal Through Video Luminance Modulation
abstract
Objectively measuring affective states remains a recurring challenge in psychology and user experience research. A promising proxy for perceived arousal is the momentary assessment of the pupil diameter. However, the pupillometric signal is highly susceptible to variations in stimulus luminance, which can substantially confound the results. We introduce a method to improve the accuracy of pupillometric signals as an arousal predictor by eliminating frame-by-frame luminance variations in visual stimuli (i.e. scaling the brightness value for all pixels such that each frame has the same average luminance). This process was termed "equalization". We tested our approach in a study with 31 participants between the ages of 21 and 64. Our study focused on two metrics: (1) the perceived quality of the stimuli should not be affected by the process, and (2) the predictive power of the pupillometric signal for the perceived arousal should be greater for the equalized videos than for the non-equalized ones. We used a within-subjects design, where each participant rated their affective response and five distinct quality dimensions of four videos (two in the equalized and two in the non-equalized condition). Our results indicate that the equalization process does improve the predictive power of the pupillometric signal by a substantial amount. Statistical analyses also show that the participants rated all quality dimensions equivalently for an equalized and non-equalized variant of the same stimulus, with only small differences that are limited to a subset of the perceptual dimensions. While potentially limiting the efficacy of the process in some scenarios, the strengthened explanatory power of the pupillometric signal leaves room for a wide range of possible applications.
Leon Schreiber, Wafaa Wardah, Vera Schmitt, Sebastian Möller 0001, Robert P. Spang
QoMEX3
2024 Disentangling User States in QoE: Situation-Dependent and Independent Factors
abstract
While current Quality of Experience (QoE) formation models recognize the impact of user states on perception, they often overlook the subjective nuances and individual variations in these experiences. We discuss the integration of both situation-dependent and independent user states, show how dependent states are influenced by varying multimedia content quality, and reflect back into the QoE assessment. To address this gap, we conducted a comprehensive within-subjects lab study (N=92) in a video-telephony setting, employing variables representative of both dependent and independent states, such as affective states, social relationship, sympathy, and bodily needs. Dependent states were assessed per trial and independent states before or after the experiment. Our study investigates two key questions: How does video-telephony call quality influence user-dependent states, and how do both dependent and independent states collectively impact QoE ratings? Our findings reveal a substantial influence of call quality on user emotional states, underscoring the importance of considering these factors in QoE assessments. Moreover, a structural equation model comparison favored the dependent+independent state structure model, highlighting the significant impact of both independent and preceding dependent states on QoE ratings. This research advances QoE models by incorporating a more nuanced mental model of user states, leading to more personalized and accurate multimedia assessment tools, potentially enhancing users’ degree of delight or annoyance and engagement.
Robert P. Spang, Maximilian Warsinke, Vera Schmitt, Luis Felipe Villa-Arenas, Navid Ashrafi, Sebastian Möller 0001
QoMEX3
2023 Protect and Extend - Using GANs for Synthetic Data Generation of Time-Series Medical Records
abstract
Preservation of private user data is of paramount importance for high Quality of Experience (QoE) and acceptability, particularly with services treating sensitive data, such as IT-based health services. Whereas anonymization techniques were shown to be prone to data re-identification, synthetic data generation has gradually replaced anonymization since it is relatively less time and resource-consuming and more robust to data leakage. Generative Adversarial Networks (GANs) have been used for generating synthetic datasets, especially GAN frameworks adhering to the differential privacy phenomena. This research compares state-of-the-art GAN-based models for synthetic data generation to generate time-series synthetic medical records of dementia patients which can be distributed without privacy concerns. Predictive modeling, autocorrelation, and distribution analysis are used to assess the Quality of Generating (QoG) of the generated data. The privacy preservation of the respective models is assessed by applying membership inference attacks to determine potential data leakage risks. Our experiments indicate the superiority of the privacy-preserving GAN (PPGAN) model over other models regarding privacy preservation while maintaining an acceptable level of QoG. The presented results can support better data protection for medical use cases in the future.
Navid Ashrafi, Vera Schmitt, Robert P. Spang, Sebastian Möller 0001, Jan-Niklas Voigt-Antons
QoMEX2
2023 Comparing Simulated and Real Conversations for QoE Assessments: Insights from ARKit-Based Facial Configuration Analyses
abstract
This manuscript investigates the suitability of video-based conversation simulations for studying human reactions to quality degradations, focusing on facial configurations as a proxy for QoE. We analyze data from two distinct studies: a video-simulated video-telephony scenario using the storytime dataset, where participants passively watched videos, and a second study involving real conversations between participants. In both studies, facial features were continuously recorded using Apple's iOS ARKit API. We identify a factor structure of facial features that significantly relates to participants' QoE ratings in the first study and validate its robustness by replicating it in the second, independent study. Our findings suggest statistically significant estimations of QoE ratings across both paradigms, demonstrating the suitability of passive conversation simulations for studying human reactions to quality degradation. We assess the value of the proposed approach at its present stage and conclude that it can be a valuable tool when used in conjunction with other methods, as its predictive capabilities are still not robust enough to rely solely on this analysis technique.
Robert P. Spang, Wafaa Wardah, Vera Schmitt, Sebastian Möller 0001
QoMEX3
2023 Unraveling the Hangry Rater: Non-linear Effects of Hunger on Multimedia Quality Perception
abstract
The subjective quality of experience (QoE) in multimedia contexts is influenced by various factors, including individual differences among raters and experimental setups. While the latter has been extensively studied, the former remains relatively unexplored. This paper investigates the impact of hangriness - a mental state of irritability and frustration caused by hunger - on QoE ratings. In our analysis, hangriness appears to be prevalent in a specific time interval, where individuals have not consumed any food between five and eleven hours. Our analysis, comprising ratings from 100 participants, reveals a significant, non-linear effect of hangriness on QoE ratings, specifically for multimedia stimuli with subpar quality. Participants in the hangry state rated such stimuli significantly worse compared to those who had eaten recently or abstained from food for more than eleven hours. Interestingly, this effect was not observed for high-quality multimedia content. Our findings highlight the importance of considering individual differences, such as hangriness, in QoE research, as they can significantly impact subjective ratings. Further research is needed to corroborate these results and explore other factors that may influence QoE ratings. This work contributes to a better understanding of individual variability in multimedia quality perception and provides insights for designing more reliable QoE assessment methods.
Robert P. Spang, Wafaa Wardah, Vera Schmitt, Sebastian Möller 0001
QoMEX3
2023 What is Your Location Privacy Worth? Monetary Valuation of Different Location Types and Privacy Influencing Factors
abstract
Nowadays, many apps use location data to estimate the user's behavior for targeted advertising, predicting significant locations, personal preferences, state of health, and sports activities. Users of location-based services are often left with no other choice than to accept or reject location tracking when they want to use various applications. Especially, users with higher privacy concerns may reduce the frequency of location tracking by turning it off in the settings. However, most users are unaware that many applications installed on their phones are continuously tracking them. Therefore, this study attempts to answer how (obviously) being tracked over one-week influences a user's privacy concerns. The study was implemented using an iOS app, which participants could install on their smartphones. Moreover, over one week, the participants were requested to answer daily mini-questionnaires about how much they would be willing to pay for the protection of their location information on a monthly basis and how much money they were willing to accept in exchange for their location information. Hereby, the context was an important criterion to determine how the monetary values vary among different location types for, among others, home location, work location, and meeting family and friends. The participants (N=51) interacted with the app on a daily basis by filling out various daily mini-surveys based on their significant locations visited. The results show a significant difference between the monetary valuating of willingness to pay and to accept for all location types except work location and sharing scenarios contributing to further empirical evidence for the endowment effect. The obvious fact of continuously being tracked did not increase the privacy concern of participants.
Vera Schmitt, Zhenni Li, Maija Poikela, Robert P. Spang, Sebastian Möller 0001
WISEC1
2022 Android Permission Manager, Visual Cues, and their Effect on Privacy Awareness and Privacy Literacy
abstract
Android applications request specific permissions from users during the installations to perform required functionalities by accessing system resources and personal information. Usually, users must approve the permissions requested by applications (apps) during the installation process and before the apps can collect privacy- or security-relevant information. However, recent studies have shown that users are overwhelmed with the information provided in privacy policies and do not understand permission requests and which functionalities are necessary for certain applications. Hereby, the collection of personal information remains mostly hidden, as the task of verifying to which information different apps have access to can be very complicated. Therefore, it is necessary to develop frameworks and apps that enable the user to perform informed decisions about apps’ run-time permission access to facilitate the control over sensitive information collected by various apps on smartphones. In this work, we conducted an online study with 70 participants who interacted with a mockup app that enables advanced control over permission requests. The selected permissions are based on the apps’ run-time permission access patterns and explanations, and commonly known visual cues are used to facilitate the user’s understanding and privacy-conscious decision making. Furthermore, the effects of perceived control over information sharing and privacy awareness are examined in combination with the permission manager mockup app to investigate if increased control over information sharing increases general privacy awareness.
Vera Schmitt, Maija Poikela, Sebastian Möller 0001
ARES1