Georg Groh

dblp:09/5335 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-5942-2297ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 1 first-author · 13 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 9 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LLMs are the Ideal Candidate for Mixed-Initiative Game Design Pillar Workflows
abstract
Game Design Pillars are natural language artifacts commonly used in game development to communicate a project’s core vision and ensure a coherent player experience. Their linguistic nature aligns well with the strengths of Large Language Models (LLMs), which excel at generating and interpreting natural language, making them promising candidates for supporting mixed-initiative pillar workflows. In this study, we introduce a formal definition of game design pillars, present an initial prototype—SPINE—and investigate the utility of LLMs in the creation and decision-making processes associated with pillar-driven workflows. We begin with a pre-study comparing gemini-2.0-flash and GPT-4o-mini. Results show that Gemini is better suited to our tasks due to greater output variety and consistency. We then conduct a case study by deploying the tool at a local game jam. Findings indicate positive reception and clear value in integrating SPINE into early-stage development. Finally, we interview four experts, demonstrating the tool and allowing them to experiment with it in a controlled environment. While individual perspectives vary, the overall perception is encouraging and supports our intuition: LLMs can meaningfully contribute to game design pillar workflows. These early findings highlight the potential of formalizing pillar-driven design as a research space and point toward several promising avenues for future work.
Julian Geheeb, Marvin Julian Schwarz, Daniel Dyrda, Georg Groh
FDG4
2026 MTDiag: A Multi-Turn Diagnostic Dataset Towards Clinically Meaningful LLM Evaluation
abstract
Clinical diagnosis is fundamentally interactive and incremental, yet the dominant paradigm for evaluating Large Language Models (LLMs) in medicine remains static QA benchmarks or template-based dialogues. These benchmarks say little about whether a model can serve as a diagnostic agent in a dynamic clinical encounter, with LLMs showing significant accuracy and reliability degradation in multi-turn settings. To address this issue, we present MTDiag, a large multi-turn diagnostic dialogue dataset constructed from three heterogeneous sources: DDXPlus, MIMIC-IV, and published case reports (AJCR), covering common ED presentations as well as long-tail rare and atypical conditions. All cases are normalized into a canonical PatientVector schema anchored in the most comprehensive and widely-adopted medical knowledge bases (UMLS concept identifiers, with ICD-10 diagnosis codes). We release the PatientVector schema, a UserLM-8B-based utterance-generation pipeline, and the physician-validated dataset that converts structured clinical evidence into natural-language utterances. Importantly, we introduce and motivate clinical knowledge-grounded metrics for evaluating LLMs as diagnostic agents, beyond diagnostic accuracy, for the task of multi-turn differential diagnosis.
Pia Chouayfati, Alexander Fichtl, Miriam Anschütz, George Doumat, Georg Groh
SIGDIAL5
2025 Cross-lingual Text Classification Transfer: The Case of Ukrainian
abstract
Despite the extensive amount of labeled datasets in the NLP text classification field, the persistent imbalance in data availability across various languages remains evident. To support further fair development of NLP models, exploring the possibilities of effective knowledge transfer to new languages is crucial. Ukrainian, in particular, stands as a language that still can benefit from the continued refinement of cross-lingual methodologies. Due to our knowledge, there is a tremendous lack of Ukrainian corpora for typical text classification tasks, i.e., different types of style, or harmful speech, or texts relationships. However, the amount of resources required for such corpora collection from scratch is understandable. In this work, we leverage the state-of-the-art advances in NLP, exploring cross-lingual knowledge transfer methods avoiding manual data curation: large multilingual encoders and translation systems, LLMs, and language adapters. We test the approaches on three text classification tasks—toxicity classification, formality classification, and natural language inference (NLI)—providing the “recipe” for the optimal setups for each task.
Daryna Dementieva, Valeriia Khylenko, Georg Groh
COLING3
2025 German4All - A Dataset and Model for Readability-Controlled Paraphrasing in German
abstract
The ability to paraphrase texts across different complexity levels is essential for creating accessible texts that can be tailored toward diverse reader groups. Thus, we introduce German4All, the first large-scale German dataset of aligned readability-controlled, paragraph-level paraphrases. It spans five readability levels and comprises over 25,000 samples. The dataset is automatically synthesized using GPT-4 and rigorously evaluated through both human and LLM-based judgments. Using German4All, we train an open-source, readability-controlled paraphrasing model that achieves state-of-the-art performance in German text simplification, enabling more nuanced and reader-specific adaptations. We open-source both the dataset and the model to encourage further research on multi-level paraphrasing.
Miriam Anschütz, Thanh Mai Pham, Eslam Nasrallah, Maximilian Müller, Cristian-George Craciun, Georg Groh
INLG6
2024 Adapter-Based Approaches to Knowledge-Enhanced Language Models: A Survey
abstract
Knowledge-enhanced language models (KELMs) have emerged as promising tools to bridge the gap between large-scale language models and domain-specific knowledge. KELMs can achieve higher factual accuracy and mitigate hallucinations by leveraging knowledge graphs (KGs). They are frequently combined with adapter modules to reduce the computational load and risk of catastrophic forgetting. In this paper, we conduct a systematic literature review (SLR) on adapter-based approaches to KELMs. We provide a structured overview of existing methodologies in the field through quantitative and qualitative analysis and explore the strengths and potential shortcomings of individual approaches. We show that general knowledge and domain-specific approaches have been frequently explored along with various adapter architectures and downstream tasks. We particularly focused on the popular biomedical domain, where we provided an insightful performance comparison of existing KELMs. We outline the main trends and propose promising future directions.
Alexander Fichtl, Juraj Vladika, Georg Groh
KEOD3
2024 Exploring Hallucinations in Task-oriented Dialogue Systems with Narrow Domains
Yan Pan 0018, Davide Cadamuro, Georg Groh
PACLIC3
2024 FELIX: Automatic and Interpretable Feature Engineering Using LLMs
Simon Malberg, Edoardo Mosca, Georg Groh
ECML/PKDD (4)3
2023 PARL: A Dialog System Framework with Prompts as Actions for Reinforcement Learning
Tao Xiang 0003, Yangzhe Li, Monika Wintergerst, Ana Pecini, Dominika Mlynarczyk, Georg Groh
ICAART (3)6
2023 This is not correct! Negation-aware Evaluation of Language Generation Systems
abstract
Large language models underestimate the impact of negations on how much they change the meaning of a sentence.Therefore, learned evaluation metrics based on these models are insensitive to negations.In this paper, we propose NegBLEURT, a negation-aware version of the BLEURT evaluation metric.For that, we designed a rule-based sentence negation tool and used it to create the CANNOT negation evaluation dataset.Based on this dataset, we fine-tuned a sentence transformer and an evaluation metric to improve their negation sensitivity.Evaluating these models on existing benchmarks shows that our fine-tuned models outperform existing metrics on the negated sentences by far while preserving their base models' performances on other perturbations.
Miriam Anschütz, Diego Miguel Lozano, Georg Groh
INLG3
2023 Data-Augmented Task-Oriented Dialogue Response Generation with Domain Adaptation
Yan Pan 0018, Davide Cadamuro, Georg Groh
PACLIC3
2022 "That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial Attacks
abstract
Adversarial attacks are a major challenge faced by current machine learning research.These purposely crafted inputs fool even the most advanced models, precluding their deployment in safety-critical applications.Extensive research in computer vision has been carried to develop reliable defense strategies.However, the same issue remains less explored in natural language processing.Our work presents a model-agnostic detector of adversarial text examples.The approach identifies patterns in the logits of the target classifier when perturbing the input text.The proposed detector improves the current state-ofthe-art performance in recognizing adversarial inputs and exhibits strong generalization capabilities across different NLP models, datasets, and word-level attacks.
Edoardo Mosca, Shreyash Agarwal, Javier Rando-Ramirez, Georg Groh
ACL (1)4
2022 SHAP-Based Explanation Methods: A Review for NLP Interpretability
abstract
Model explanations are crucial for the transparent, safe, and trustworthy deployment of machine learning models. The SHapley Additive exPlanations (SHAP) framework is considered by many to be a gold standard for local explanations thanks to its solid theoretical background and general applicability. In the years following its publication, several variants appeared in the literature—presenting adaptations in the core assumptions and target applications. In this work, we review all relevant SHAP-based interpretability approaches available to date and provide instructive examples as well as recommendations regarding their applicability to NLP use cases.
Edoardo Mosca, Ferenc Szigeti, Stella Tragianni, Daniel Gallagher, Georg Groh
COLING5
2022 Introducing an Abusive Language Classification Framework for Telegram to Investigate the German Hater Community
Maximilian Wich, Adrian Gorniak, Tobias Eder, Daniel Bartmann, Burak Enes Çakici, Georg Groh
ICWSM6
2022 User Satisfaction Modeling with Domain Adaptation in Task-oriented Dialogue Systems
abstract
User Satisfaction Estimation (USE) is crucial in helping measure the quality of a task-oriented dialogue system.However, the complex nature of implicit responses poses challenges in detecting user satisfaction, and most datasets are limited in size or not available to the public due to user privacy policies.Unlike task-oriented dialogue, large-scale annotated chitchat with emotion labels is publicly available.Therefore, we present a novel user satisfaction model with domain adaptation (USMDA) to utilize this chitchat.We adopt a dialogue Transformer encoder to capture contextual features from the dialogue.And we reduce domain discrepancy to learn dialogue-related invariant features.Moreover, USMDA jointly learns satisfaction signals in the chitchat context with user satisfaction estimation, and user actions in task-oriented dialogue with dialogue action recognition.Experimental results on two benchmarks show that our proposed framework for the USE task outperforms existing unsupervised domain adaptation methods.To the best of our knowledge, this is the first work to study user satisfaction estimation with unsupervised domain adaptation from chitchat to task-oriented dialogue.
Yan Pan 0018, Bernhard Pflugfelder, Georg Groh
SIGDIAL4
2022 Effects and challenges of using a nutrition assistance system: results of a long-term mixed-method study
abstract
Abstract Healthy nutrition contributes to preventing non-communicable and diet-related diseases. Recommender systems, as an integral part of mHealth technologies, address this task by supporting users with healthy food recommendations. However, knowledge about the effects of the long-term provision of health-aware recommendations in real-life situations is limited. This study investigates the impact of a mobile, personalized recommender system named Nutrilize. Our system offers automated personalized visual feedback and recommendations based on individual dietary behaviour, phenotype, and preferences. By using quantitative and qualitative measures of 34 participants during a study of 2–3 months, we provide a deeper understanding of how our nutrition application affects the users’ physique, nutrition behaviour, system interactions and system perception. Our results show that Nutrilize positively affects nutritional behaviour (conditional $$R^2=.342$$ R 2 = . 342 ) measured by the optimal intake of each nutrient. The analysis of different application features shows that reflective visual feedback has a more substantial impact on healthy behaviour than the recommender (conditional $$R^2=.354$$ R 2 = . 354 ). We further identify system limitations influencing this result, such as a lack of diversity, mistrust in healthiness and personalization, real-life contexts, and personal user characteristics with a qualitative analysis of semi-structured in-depth interviews. Finally, we discuss general knowledge acquired on the design of personalized mobile nutrition recommendations by identifying important factors, such as the users’ acceptance of the recommender’s taste, health, and personalization.
Hanna Hauptmann, Nadja Leipold, Mira Madenach, Monika Wintergerst, Martin Lurz, Georg Groh, Markus Böhm 0001, Kurt Gedrich, Helmut Krcmar
User Model. User Adapt. Interact.6
2021 DYME: A Dynamic Metric for Dialog Modeling Learned from Human Conversations
Florian von Unold, Monika Wintergerst, Lenz Belzner, Georg Groh
ICONIP (5)4
2021 Explainable Abusive Language Classification Leveraging User and Network Data
Maximilian Wich, Edoardo Mosca, Adrian Gorniak, Johannes Hingerl, Georg Groh
ECML/PKDD (5)5
2020 Evaluation Metrics for Headline Generation Using Deep Pre-Trained Embeddings
abstract
With the explosive growth in textual data, it is becoming increasingly important to summarize text automatically. Recently, generative language models have shown promise in abstractive text summarization tasks. Since these models rephrase text and thus use similar but different words as found in the summarized text, existing metrics such as ROUGE that use n-gram overlap may not be optimal. Therefore we evaluate two embedding-based evaluation metrics that are applicable to abstractive summarization: Fr ́echet embedding distance, which has been introduced recently, and angular embedding similarity, which is our proposed metric. To demonstrate the utility of both metrics, we analyze the headline generation capacity of two state-of-the-art language models: GPT-2 and ULMFiT. In particular, our proposed metric shows close relation with human judgments in our experiments and has overall better correlations with them. To provide reproducibility, the source code plus human assessments of our experiments is available on GitHub.
Abdul Moeed, Gerhard Hagerer, Georg Groh
LREC4
2020 An Evaluation of Progressive Neural Networksfor Transfer Learning in Natural Language Processing
abstract
A major challenge in modern neural networks is the utilization of previous knowledge for new tasks in an effective manner, otherwise known as transfer learning. Fine-tuning, the most widely used method for achieving this, suffers from catastrophic forgetting. The problem is often exacerbated in natural language processing (NLP). In this work, we assess progressive neural networks (PNNs) as an alternative to fine-tuning. The evaluation is based on common NLP tasks such as sequence labeling and text classification. By gauging PNNs across a range of architectures, datasets, and tasks, we observe improvements over the baselines throughout all experiments.
Abdul Moeed, Gerhard Hagerer, Sumit Dugar, Sarthak Gupta, Mainak Ghosh, Hannah Danner, Oliver Mitevski, Andreas Nawroth, Georg Groh
LREC9
2019 Detection of topical influence in social networks via granger-causal inference: a Twitter case study
abstract
With the ever-increasing importance of computer-mediated communication in our everyday life, understanding the effects of social influence in online social networks has become a necessity. In this work, we argue that cascade models of information diffusion do not adequately capture attitude change, which we consider to be an essential element of social influence. To address this concern, we propose a topical model of social influence and attempt to establish a connection between influence and Granger-causal effects on a theoretical and empirical level. While our analysis of a social media dataset finds effects that are consistent with our model of social influence, evidence suggests that these effects can be attributed largely to external confounders. The dominance of external influencers, including mass media, over peer influence raises new questions about the correspondence between objectively measurable information diffusion and social influence as perceived by human observers.
Jan Hauffa, Wolfgang Bräu, Georg Groh
ASONAM3
2016 Conversational context helps improve mobile notification management
abstract
We explore if and how identifying the character of face-to-face conversations can help manage notifications on smartphones so that they become less disruptive. We show that the social dimensions depth/importance and formality/goal orientation of a conversation are strong indicators of receptiveness. Furthermore, we find that there are types of conversation, e.g. small talk, in which individuals are even more receptive to notifications than in situations without any verbal social interaction at all. This refutes the assumption currently found in the literature that the occurrence of a conversation is a strong predictor of unavailability. We demonstrate a system that tracks conversations in which the user is engaged and that analyzes speech in terms of embedded affective and social cues. Eventually, we find that information of either kind, derived from audio, improves the accuracy of personal notification preference models substantially.
Florian Schulze, Georg Groh
MobileHCI2
2015 Appropriateness of Search Engines, Social Networks, and Directly Approaching Friends to Satisfy Information Needs
abstract
One form of social search is to integrate one's social network in the search process by querying friends, leading to more subjective but also highly individualized answers. Previous studies analyzed users' social search behavior using (broadcasted) status messages on social networking platforms to communicate information needs (Status Message Question Asking, SMQA) and revealed a limited willingness of information seekers to use SMQA when comparing it to traditional search engines. We describe the results of a survey with 112 participants and show that directly approaching well chosen friends is considered more attractive and is associated with higher expectations in terms of response quality than SMQA. Our findings suggest that users anticipate quality improvements gained from forwarding queries especially for certain content types of information needs and that response time is an important factor.
Christoph Fuchs 0001, Georg Groh
ASONAM2
2012 Spatio-temporal small worlds for decentralized information retrieval in social networking
abstract
We discuss foundations and options for alternative, agent-based information retrieval (IR) approaches in Social Networking (SN). In addition to usual semantic contexts, these approaches make use of long-term social and spatio-temporal contexts according to Human IR heuristics. Using a large Twitter dataset, we investigate foundations for these approaches and especially the question in how far spatio-temporal contexts can act as a conceptual bracket implicating social and semantic cohesion, giving rise to the concept of Spatio-Temporal Small Worlds.
Georg Groh, Florian Straub, Benjamin Koster
SIGSPATIAL/GIS1
2011 Characterizing Social Relations Via NLP-Based Sentiment Analysis
Georg Groh, Jan Hauffa
ICWSM1
2010 Group Management in P2P Networks
abstract
Groups are both, a social phenomenon inherent to the human nature and a widely used structure in distributed systems. Recently, the Web 2.0 trend has begun to unite both aspects. Social networks provide their users with simple means to create, modify, join, and leave groups dynamically. Popular applications such as chat rooms and multi-player online games require social group structures. The same holds true for many e-science applications. But despite their growing importance, the structure of social groups has not yet been studied with respect to peer-to-peer or cloud computing systems. In this paper, we analyze the requirement structure that the different types of social groups induce in distributed systems. In particular, we focus on fully decentralized peer-to-peer systems. First, we formalize the requirements using results from field studies in social networks. Then we classify the various group management models by considering their differences in group creation and management. We also discuss scalability, maintainability, security issues and privacy guarantees of the different group models.
Benedikt Elser, Georg Groh, Thomas Fuhrmann
ICCCN2
2010 On the Impact of Chat Communication on Computer-Supported Idea Generation Processes
Florian Forster, Marc René Frieß, Michele Brocco, Georg Groh
ICCC4
2009 Team recommendation in open innovation networks
abstract
Open Innovation has become an important new paradigm for incorporating external knowledge and sources in the innovation process of organizations. Besides other discussed arguments the resulting large size of innovator networks suggests that algorithmic approaches for team recommendation may be needed in that scenario. The current work identifies the related difficulties and thoroughly investigates aspects entities for the problem of team recommendation. Based on that, we develop a meta model which allows to instantiate and integrate most of the vast number of the existing socio-/psychological models on optimal team composition. This meta model is necessary for operationalizing our intended team recommendation approach.
Michele Brocco, Georg Groh
RecSys2
2007 Recommendations in taste related domains: collaborative filtering vs. social filtering
abstract
We investigate how social networks can be used in recommendation generation in taste related domains. Social Filtering (using social networks for neighborhood generation) is compared to Collaborative Filtering with respect to prediction accuracy in the domain of rating clubs. After reviewing background and related work, we present an extensive empirical study where over thousand participants from a social networking community where asked to provide ratings for clubs in Munich. We then compare a typical traditional CF-approach to a social recommender / social filtering approach where friends from the underlying social network are used as rating neighborhood and analyze the experiments statistically. Surprisingly, the social filtering approach outperforms the CF approach in all variants of the experiment. The implications of the experiment for professional and private-life collaborative environments and services where recommendations play a role are discussed. We conclude with future perspectives on social recommender systems, especially in upcoming mobile environments.
Georg Groh, Christian Ehmig
GROUP1
2007 Groups and Group-Instantiations in Mobile Communities - Detection, Modeling and Applications
Georg Groh
ICWSM1