EDBT 2026 Demo / reviewers in the wild / expert
Florian Matthes
dblp:m/FlorianMatthes
· DBLP profile ↗
101ranked-venue papers
8as first author
48since 2021 · last 2026
0000-0002-6667-5452ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 23 since 2021Software engineering, systems software and programming languages · 26 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 16 · 4 first-author · 4 since 2021Security and privacy · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Text2Tabular - Reconstructing Tabular Research Data from Scientific PublicationsabstractThe increasing reliance on data-driven research underscores the need for accessible datasets, particularly in the medical domain.However, raw datasets are frequently unavailable due to privacy constraints and ethical considerations, which complicates reproducibility, meta-analyses, and large-scale data-driven research.Text2Tabular addresses this challenge by reconstructing research datasets from scientific publications using advanced natural language processing and statistical modeling.Our key contributions include: (1) a unified framework combining Large Language Model driven information extraction with copula-based distribution modeling, (2) novel integration of statistical test results as distribution constraints through constrained Markov Chain Monte Carlo refinement, and (3) an own comprehensive benchmark comprising real scientific publications with corresponding raw datasets for evaluating our literaturebased data reconstruction.Evaluation on both benchmark datasets and our curated collection demonstrates strong performance in Trainon-Synthetic-Test-on-Real (TSTR) evaluations, alongside accurate replication of descriptive statistics showing that Text2Tabular preserves the statistical properties and multivariate relationships of the original datasets.Text2Tabular facilitates scientific progress by enabling immediate access to realistic, domain-specific synthetic data, thus, improving data accessibility, and mitigating data scarcity in fields with limited real-world data. Jonas Gottal, Florian Matthes |
ACL (1) | 2 |
| 2026 | Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
Anum Afzal, Yuki Saito 0001, Hiroya Takamura, Katsuhito Sudoh, Shinnosuke Takamichi, Graham Neubig, Florian Matthes, Tatsuya Ishigaki |
LREC | 7 |
| 2025 | MedSEBA: Synthesizing Evidence-Based Answers Grounded in Evolving Medical LiteratureabstractIn the digital age, people often turn to the Internet in search of medical advice and recommendations. With the increasing volume of online content, it has become difficult to distinguish reliable sources from misleading information. Similarly, millions of medical studies are published every year, making it challenging for researchers to keep track of the latest scientific findings. These evolving studies can reach differing conclusions, which is not reflected in traditional search tools. To address these challenges, we introduce MedSEBA, an interactive AI-powered system for synthesizing evidence-based answers to medical questions. It utilizes the power of Large Language Models to generate coherent and expressive answers, but grounds them in trustworthy medical studies dynamically retrieved from the research database PubMed. The answers consist of key points and arguments, which can be traced back to respective studies. Notably, the platform also provides an overview of the extent to which the most relevant studies support or refute the given medical claim, and a visualization of how the research consensus evolved through time. Our user study revealed that medical experts and lay users find the system usable and helpful, and the provided answers trustworthy and informative. This makes the system well-suited for both everyday health questions and advanced research insights. Juraj Vladika, Florian Matthes |
CIKM | 2 |
| 2025 | Spend Your Budget Wisely: Towards an Intelligent Distribution of the Privacy Budget in Differentially Private Text RewritingabstractThe task of Differentially Private Text Rewriting is a class of text privatization techniques in which (sensitive) input textual documents are rewritten under Differential Privacy (DP) guarantees. The motivation behind such methods is to hide both explicit and implicit identifiers that could be contained in text, while still retaining the semantic meaning of the original text, thus preserving utility. Recent years have seen an uptick in research output in this field, offering a diverse array of word-, sentence-, and document-level DP rewriting methods. Common to these methods is the selection of a privacy budget (i.e., the ε parameter), which governs the degree to which a text is privatized. One major limitation of previous works, stemming directly from the unique structure of language itself, is the lack of consideration of where the privacy budget should be allocated, as not all aspects of language, and therefore text, are equally sensitive or personal. In this work, we are the first to address this shortcoming, asking the question of how a given privacy budget can be intelligently and sensibly distributed amongst a target document. We construct and evaluate a toolkit of linguistics- and NLP-based methods used to allocate a privacy budget to constituent tokens in a text document. In a series of privacy and utility experiments, we empirically demonstrate that given the same privacy budget, intelligent distribution leads to higher privacy levels and more positive trade-offs than a naive distribution of epsilon. Our work highlights the intricacies of text privatization with DP, and furthermore, it calls for further work on finding more efficient ways to maximize the privatization benefits offered by DP in text rewriting. Stephen Meisenbacher, Chaeeun Joy Lee, Florian Matthes |
CODASPY | 3 |
| 2025 | DA-Pred: Performance Prediction for Text Summarization under Domain-Shift and Instruct-TuningabstractLarge Language Models (LLMs) often don't perform as expected under Domain Shift or after Instruct-tuning.A reliable indicator of LLM performance in these settings could assist in decision-making.We present a method that uses the known performance in high-resource domains and fine-tuning settings to predict performance in low-resource domains or base models, respectively.In our paper, we formulate the task of performance prediction, construct a dataset for it, and train regression models to predict the said change in performance.Our proposed methodology is lightweight and, in practice, can help researchers & practitioners decide if resources should be allocated for data labeling and LLM Instruct-tuning. Anum Afzal, Florian Matthes, Alexander R. Fabbri |
EMNLP | 2 |
| 2025 | Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy GuaranteesabstractMany works at the intersection of Differential Privacy (DP) in Natural Language Processing aim to protect privacy by transforming texts under DP guarantees.This can be performed in a variety of ways, from word perturbations to full document rewriting, and most often under local DP.Here, an input text must be made indistinguishable from any other potential text, within some bound governed by the privacy parameter ε.Such a guarantee is quite demanding, and recent works show that privatizing texts under local DP can only be done reasonably under very high ε values.Addressing this challenge, we introduce DP-ST, which leverages semantic triples for neighborhoodaware private document generation under local DP guarantees.Through the evaluation of our method, we demonstrate the effectiveness of the divide-and-conquer paradigm, particularly when limiting the DP notion (and privacy guarantees) to that of a privatization neighborhood.When combined with LLM post-processing, our method allows for coherent text generation even at lower ε values, while still balancing privacy and utility.These findings highlight the importance of coherence in achieving balanced privatization outputs at reasonable ε levels. Stephen Meisenbacher, Maulik Chevli, Florian Matthes |
EMNLP | 3 |
| 2025 | Lexical Substitution is not Synonym Substitution: On the Importance of Producing Contextually Relevant Word Substitutes
Juraj Vladika, Stephen Meisenbacher, Florian Matthes |
ICAART (3) | 3 |
| 2025 | LLMs for Legal Subsumption in German Employment ContractsabstractLegal work, characterized by its text-heavy and resource-intensive nature, presents unique challenges and opportunities for NLP research. While data-driven approaches have advanced the field, their lack of interpretability and trustworthiness limits their applicability in dynamic legal environments. To address these issues, we collaborated with legal experts to extend an existing dataset and explored the use of Large Language Models (LLMs) and in-context learning to evaluate the legality of clauses in German employment contracts. Our work evaluates the ability of different LLMs to classify clauses as "valid," "unfair," or "void" under three legal context variants: no legal context, full-text sources of laws and court rulings, and distilled versions of these (referred to as examination guidelines). Results show that full-text sources moderately improve performance, while examination guidelines significantly enhance recall for void clauses and weighted F1-Score, reaching 80%. Despite these advancements, LLMs’ performance when using full-text sources remains substantially below that of human lawyers. We contribute an extended dataset, including examination guidelines, referenced legal sources, and corresponding annotations, alongside our code and all log files. Our findings highlight the potential of LLMs to assist lawyers in contract legality review while also underscoring the limitations of the methods presented. Oliver Wardas, Florian Matthes |
ICAIL | 2 |
| 2025 | "We are not Future-ready": Understanding AI Privacy Risks and Existing Mitigation Strategies from the Perspective of AI Developers in Europe
Alexandra Klymenko, Stephen Meisenbacher, Patrick Gage Kelley, Sai Teja Peddinti, Kurt Thomas, Florian Matthes |
SOUPS | 6 |
| 2025 | Knowledge Sharing and Coordination in Large-Scale Agile Software Development - A Systematic Literature Review and an Interview StudyabstractAbstract Due to their benefits regarding resilience to change and adaptability, agile methods are widely adopted in the software development industry. Despite being intended for small co-located teams, organizations have started to scale agile methods, e.g., applying them in multi-team projects. Consequently, effective and efficient knowledge sharing and coordination become more complex, and agile intra-team practices are not sufficient anymore. While several case studies investigate knowledge sharing and coordination in the scaled agile context, an overview across multiple organizations is missing. To fill this gap, we combined the results of a literature review of 69 studies and an interview study to gain an overview of how coordination and knowledge sharing are conducted in scaled agile organizations and what factors hinder or facilitate their effectiveness and efficiency. Our findings show that organizations implement various mechanisms, including roles (e.g., Product Owners), meetings (e.g., ad-hoc exchange), tools and artifacts (e.g., chats), and other structures (e.g., Communities of Practice). Organizations struggle particularly with misalignment among teams and units, insufficient communication, and distribution across different working locations. Key enabling factors are documentation efforts, as well as both formal and informal exchanges. Context is crucial and should be regularly reassessed as needs evolve. Also, strong organizational support is needed to facilitate coordination and knowledge sharing. Franziska Tobisch, Florian Matthes |
XP | 2 |
| 2025 | Fostering New Work Practices Through a Community of Practice A Case Study in a Large-Scale Software Development OrganizationabstractAbstract New Work and agile methodologies share a common foundation in their aim to foster autonomy, collaboration, and adaptability. Their benefits make both concepts highly relevant for organizations as they support agility, innovation, and attractiveness to existing and potential employees. Still, implementing these concepts within large, established companies remains challenging. Therefore, this case study investigates a Community of Practice that promotes New Work principles within a large software organization aiming to be agile and innovative. Our study explores the establishment and functioning of this community and its effectiveness in advancing New Work practices within the case company. The community has been growing and has achieved its first success in promoting New Work, but there remains potential for improvement. In particular, a clearer mandate from the upper management is needed. Franziska Tobisch, Florian Matthes |
XP | 2 |
| 2024 | Just Rewrite It Again: A Post-Processing Method for Enhanced Semantic Similarity and Privacy Preservation of Differentially Private Rewritten TextabstractThe study of Differential Privacy (DP) in Natural Language Processing often views the task of text privatization as a rewriting task, in which sensitive input texts are rewritten to hide explicit or implicit private information. In order to evaluate the privacy-preserving capabilities of a DP text rewriting mechanism, empirical privacy tests are frequently employed. In these tests, an adversary is modeled, who aims to infer sensitive information (e.g., gender) about the author behind a (privatized) text. Looking to improve the empirical protections provided by DP rewriting methods, we propose a simple post-processing method based on the goal of aligning rewritten texts with their original counterparts, where DP rewritten texts are rewritten again. Our results show that such an approach not only produces outputs that are more semantically reminiscent of the original inputs, but also texts which score on average better in empirical privacy evaluations. Therefore, our approach raises the bar for DP rewriting methods in their empirical privacy evaluations, providing an extra layer of protection against malicious adversaries. Stephen Meisenbacher, Florian Matthes |
ARES | 2 |
| 2024 | AGB-DE: A Corpus for the Automated Legal Assessment of Clauses in German Consumer ContractsabstractLegal tasks and datasets are often used as benchmarks for the capabilities of language models.However, openly available annotated datasets are rare.In this paper, we introduce AGB-DE, a corpus of 3,764 clauses from German consumer contracts that have been annotated and legally assessed by legal experts.Together with the data, we present a first baseline for the task of detecting potentially void clauses, comparing the performance of an SVM baseline with three fine-tuned open language models and the performance of GPT-3.5.Our results show the challenging nature of the task, with no approach exceeding an F1-score of 0.54.While the fine-tuned models often performed better with regard to precision, GPT-3.5 outperformed the other approaches with regard to recall.An analysis of the errors indicates that one of the main challenges could be the correct interpretation of complex clauses, rather than the decision boundaries of what is permissible and what is not. Daniel Braun 0003, Florian Matthes |
ACL (1) | 2 |
| 2024 | Who Wins Ethereum Block Building Auctions and Why?abstractThe MEV-Boost block auction contributes approximately 90% of all Ethereum blocks. Between October 2023 and March 2024, only three builders produced 80% of them, highlighting the concentration of power within the block builder market. To foster competition and preserve Ethereum's decentralized ethos and censorship-resistance properties, understanding the dominant players' competitive edges is essential. In this paper, we identify features that play a significant role in builders' ability to win blocks and earn profits by conducting a comprehensive empirical analysis of MEV-Boost auctions over a six-month period. We reveal that block market share positively correlates with order flow diversity, while profitability correlates with access to order flow from Exclusive Providers, such as integrated searchers and external providers with exclusivity deals. Additionally, we show a positive correlation between market share and profit margin among the top ten builders, with features such as exclusive signal, non-atomic arbitrages, and Telegram bot flow strongly correlating with both metrics. This highlights a "chicken-and-egg" problem where builders need differentiated order flow to profit, but only receive such flow if they have a significant market share. Overall, this work provides an in-depth analysis of the key features driving the builder market towards centralization and offers valuable insights for designing further iterations of Ethereum block auctions, preserving Ethereum's censorship resistance properties. Burak Öz, Danning Sui, Thomas Thiery, Florian Matthes |
AFT | 4 |
| 2024 | A Comparative Analysis of Word-Level Metric Differential Privacy: Benchmarking the Privacy-Utility Trade-offabstractThe application of Differential Privacy to Natural Language Processing techniques has emerged in relevance in recent years, with an increasing number of studies published in established NLP outlets. In particular, the adaptation of Differential Privacy for use in NLP tasks has first focused on the word-level, where calibrated noise is added to word embedding vectors to achieve “noisy” representations. To this end, several implementations have appeared in the literature, each presenting an alternative method of achieving word-level Differential Privacy. Although each of these includes its own evaluation, no comparative analysis has been performed to investigate the performance of such methods relative to each other. In this work, we conduct such an analysis, comparing seven different algorithms on two NLP tasks with varying hyperparameters, including the epsilon parameter, or privacy budget. In addition, we provide an in-depth analysis of the results with a focus on the privacy-utility trade-off, as well as open-source our implementation code for further reproduction. As a result of our analysis, we give insight into the benefits and challenges of word-level Differential Privacy, and accordingly, we suggest concrete steps forward for the research field. Stephen Meisenbacher, Nihildev Nandakumar, Alexandra Klymenko, Florian Matthes |
LREC/COLING | 4 |
| 2024 | HealthFC: Verifying Health Claims with Evidence-Based Medical Fact-CheckingabstractIn the digital age, seeking health advice on the Internet has become a common practice. At the same time, determining the trustworthiness of online medical content is increasingly challenging. Fact-checking has emerged as an approach to assess the veracity of factual claims using evidence from credible knowledge sources. To help advance automated Natural Language Processing (NLP) solutions for this task, in this paper we introduce a novel dataset HealthFC. It consists of 750 health-related claims in German and English, labeled for veracity by medical experts and backed with evidence from systematic reviews and clinical trials. We provide an analysis of the dataset, highlighting its characteristics and challenges. The dataset can be used for NLP tasks related to automated fact-checking, such as evidence retrieval, claim verification, or explanation generation. For testing purposes, we provide baseline systems based on different approaches, examine their performance, and discuss the findings. We show that the dataset is a challenging test bed with a high potential for future use. Juraj Vladika, Phillip Schneider, Florian Matthes |
LREC/COLING | 3 |
| 2024 | Comparing Knowledge Sources for Open-Domain Scientific Claim VerificationabstractThe increasing rate at which scientific knowledge is discovered and health claims shared online has highlighted the importance of developing efficient fact-checking systems for scientific claims.The usual setting for this task in the literature assumes that the documents containing the evidence for claims are already provided and annotated or contained in a limited corpus.This renders the systems unrealistic for real-world settings where knowledge sources with potentially millions of documents need to be queried to find relevant evidence.In this paper, we perform an array of experiments to test the performance of open-domain claim verification systems.We test the final verdict prediction of systems on four datasets of biomedical and health claims in different settings.While keeping the pipeline's evidence selection and verdict prediction parts constant, document retrieval is performed over three common knowledge sources (PubMed, Wikipedia, Google) and using two different information retrieval techniques.We show that PubMed works better with specialized biomedical claims, while Wikipedia is more suited for everyday health concerns.Likewise, BM25 excels in retrieval precision, while semantic search in recall of relevant evidence.We discuss the results, outline frequent retrieval patterns and challenges, and provide promising future directions. Juraj Vladika, Florian Matthes |
EACL (1) | 2 |
| 2024 | Thinking Outside of the Differential Privacy Box: A Case Study in Text Privatization with Language Model PromptingabstractThe field of privacy-preserving Natural Language Processing has risen in popularity, particularly at a time when concerns about privacy grow with the proliferation of Large Language Models.One solution consistently appearing in recent literature has been the integration of Differential Privacy (DP) into NLP techniques.In this paper, we take these approaches into critical view, discussing the restrictions that DP integration imposes, as well as bring to light the challenges that such restrictions entail.To accomplish this, we focus on DP-PROMPT, a recent method for text privatization leveraging language models to rewrite texts.In particular, we explore this rewriting task in multiple scenarios, both with DP and without DP.To drive the discussion on the merits of DP in NLP, we conduct empirical utility and privacy experiments.Our results demonstrate the need for more discussion on the usability of DP in NLP and its benefits over non-DP approaches. Stephen Meisenbacher, Florian Matthes |
EMNLP | 2 |
| 2024 | A Semi-Automatic Light-Weight Approach Towards Data Generation for a Domain-Specific FAQ Chatbot Using Human-in-the-Loop
Anum Afzal, Tao Xiang 0003, Florian Matthes |
ICAART (3) | 3 |
| 2024 | Evaluating Large Language Models in Semantic Parsing for Conversational Question Answering over Knowledge Graphs
Phillip Schneider, Manuel Klettner, Kristiina Jokinen, Elena Simperl, Florian Matthes |
ICAART (3) | 5 |
| 2024 | Diversifying Knowledge Enhancement of Biomedical Language Models Using Adapter Modules and Knowledge Graphs
Juraj Vladika, Alexander Fichtl, Florian Matthes |
ICAART (2) | 3 |
| 2024 | A Universal System for OpenID Connect Sign-ins with Verifiable Credentials and Cross-Device FlowabstractSelf-Sovereign Identity (SSI), as a new and promising identity management paradigm, needs mechanisms that can ease a gradual transition of existing services and developers towards it. Systems that bridge the gap between SSI and established identity and access management have been proposed but still lack adoption. We propose a comparatively simple system that enables SSI-based sign-ins for services that support the widespread OpenID Connect or OAuth 2.0 protocols. Its handling of claims is highly configurable through a single policy and designed for cross-device authentication flows involving a smartphone identity wallet. We evaluate our design by implementing and successfully integrating it with existing interfacing components. Felix Hoops, Florian Matthes |
ICBC | 2 |
| 2024 | Playing the MEV Game on a First-Come-First-Served BlockchainabstractFirst-Come-First-Served (FCFS) transaction ordering has been discussed as a fairness approach against harmful Maximal Extractable Value (MEV) strategies. However, such an ordering mechanism promotes latency optimizations, similar to High-Frequency Trading in Traditional Finance. This paper examines the dynamics of the MEV extraction game in an FCFS network, specifically Algorand. We introduce an arbitrage opportunity detection algorithm tailored to Algorand’s time constraints and assess its effectiveness. Our analysis reveals that while the states of exchange pools are updated approximately only every six blocks, pursuing MEV at the block state level is not viable, as arbitrage opportunities are typically closed within the block they appear. Additionally, we experiment on a private Algorand network to uncover latency optimization factors and show the importance of reducing latency in connections with relays well-connected to high-staked proposers. Burak Öz, Jonas Gebele, Parshant Singh, Filip Rezabek, Florian Matthes |
ICBC | 5 |
| 2024 | Bridging Information Gaps in Dialogues with Grounded Exchanges Using Knowledge GraphsabstractKnowledge models are fundamental to dialogue systems for enabling conversational interactions, which require handling domainspecific knowledge.Ensuring effective communication in information-providing conversations entails aligning user understanding with the knowledge available to the system.However, dialogue systems often face challenges arising from semantic inconsistencies in how information is expressed in natural language compared to how it is represented within the system's internal knowledge.To address this problem, we study the potential of large language models for conversational grounding, a mechanism to bridge information gaps by establishing shared knowledge between dialogue participants.Our approach involves annotating human conversations across five knowledge domains to create a new dialogue corpus called BridgeKG.Through a series of experiments on this dataset, we empirically evaluate the capabilities of large language models in classifying grounding acts and identifying grounded information items within a knowledge graph structure.Our findings offer insights into how these models use in-context learning for conversational grounding tasks and common prediction errors, which we illustrate with examples from challenging dialogues.We discuss how the models handle knowledge graphs as a semantic layer between unstructured dialogue utterances and structured information items. Phillip Schneider, Nektarios Machner, Kristiina Jokinen, Florian Matthes |
SIGDIAL | 4 |
| 2024 | Investigating Communities of Practice in Large-Scale Agile Software Development: An Interview StudyabstractAbstract Nowadays, responsiveness is essential to be competitive, particularly in software development. Traditional methods face limitations in meeting this demand for agility, which led to the rise of agile practices. Inspired by their success in small projects, organizations have begun to use agile methods in larger contexts. However, scaling agile practices introduces complexities and requires coordinating teams, managing dependencies, and collaboration. Communities of Practices (CoPs) are argued to address these issues and support organizations in adopting agile methods at scale. Still, empirical insights into the establishment of CoPs in scaled agile settings are limited. This study fills this gap by conducting expert interviews, exploring why organizations applying agile methods at scale adopt CoPs, and examining their characteristics. Our key findings include that, next to benefit from known advantages of CoPs, like knowledge sharing, organizations establish them to foster empowerment, strengthen alignment, and drive their agile transformation. Moreover, CoPs focus not only on agile but also on classical themes such as architecture. Communities are not necessarily established bottom-up but are often initiated by management, e.g., to empower employees. In general, CoPs are accepted by management and play an essential role in decision-making. Franziska Tobisch, Florian Matthes |
XP | 3 |
| 2024 | Investigating Effort Estimation in a Large-Scale Agile ERP Transformation ProgramabstractAbstract Adaptability is vital in today’s rapidly changing business environment, especially within IT. Agile methodologies have emerged to meet this demand and have thereby gained widespread adoption. While successful in smaller, co-located teams and low-criticality projects, applying agile methods in broader contexts poses challenges. Nevertheless, many organizations have started implementing agile methodologies in various areas, including large-scale Enterprise Resource Planning (ERP) projects. In contrast to traditional development, ERP projects involve deploying extensive integrated systems, are substantial in scale, and entail high risks and costs. Accurate predictions, like effort estimations, are crucial to meet customer satisfaction and deliver within plan and budget. However, estimating effort in an agile environment poses its own set of challenges. For instance, coordination efforts and dependencies among teams must be considered. While effort estimation is well-explored in classical software development and small-scale agile contexts, limited research exists in large-scale agile settings, particularly in projects rolling out and customizing standard ERP solutions. To address this gap, we conducted a case study on effort estimation in a large agile ERP transformation program, describing the estimation process, highlighting challenges, and proposing and evaluating mitigations. Franziska Tobisch, Karla Weigelt, Pascal Philipp, Florian Matthes |
XP | 4 |
| 2023 | Adoption of Information Security Practices in Large-Scale Agile Software Development: A Case Study in the Finance IndustryabstractAgile development methods have pervaded software engineering and are increasingly applied in large projects and organizations. At the same time, security threats and restrictive legislation regarding security and privacy are steadily rising. These two trends of agile software development at scale and increasingly important security requirements are often at odds with each other. Academic literature widely acknowledges the challenges therefrom and discusses approaches to integrate these two partly conflicting trends. However, several researchers point out a need for empirical studies and evaluations of these approaches in practice. To fill this research gap, we conducted a case study in the finance industry. We identified 27 agile security approaches in academic literature. Based on these theoretical findings, we carried out observations, document analysis, and unstructured interviews to identify which approaches the case company applies. We then conducted semi-structured interviews with 10 experts and a survey with 62 participants to evaluate 14 approaches. One of the key results is that role and knowledge approaches, such as dedicated security roles and communities, are especially important in scaled agile development environments. In addition, the most beneficial security activities are easy-to-integrate, such as a security tagging system, peer security code reviews, security stories, and threat poker. We also contribute evaluation criteria as well as drivers and obstacles for the adoption of agile security approaches that can be used for further research and practice. Sascha Nägele, Lorena Korn, Florian Matthes |
ARES | 3 |
| 2023 | Investigating Conversational Search Behavior for Domain Exploration
Phillip Schneider, Anum Afzal, Juraj Vladika, Daniel Braun 0003, Florian Matthes |
ECIR (2) | 5 |
| 2023 | Challenges in Domain-Specific Abstractive Summarization and How to Overcome ThemabstractLarge Language Models work quite well with general-purpose data and many tasks in Natural Language Processing. However, they show several limitations when used for a task such as domain-specific abstractive text summarization. This paper identifies three of those limitations as research problems in the context of abstractive text summarization: 1) Quadratic complexity of transformer-based models with respect to the input text length; 2) Model Hallucination, which is a model's ability to generate factually incorrect text; and 3) Domain Shift, which happens when the distribution of the model's training and test corpus is not the same. Along with a discussion of the open research questions, this paper also provides an assessment of existing state-of-the-art techniques relevant to domain-specific text summarization to address the research gaps. Anum Afzal, Juraj Vladika, Daniel Braun 0003, Florian Matthes |
ICAART (3) | 4 |
| 2023 | Acadela: A Domain-Specific Language for Modeling Clinical Pathways
Tri Huynh, Selin Erdem, Felix Martin Eckert, Florian Matthes |
ICSOFT | 4 |
| 2023 | A Method for Metric Management at a Large-Scale Agile Software Development Organization
Pascal Philipp, Franziska Tobisch, Leon Menzel, Florian Matthes |
IWSM-Mensura | 4 |
| 2023 | From Data to Dialogue: Leveraging the Structure of Knowledge Graphs for Conversational Exploratory Search
Phillip Schneider, Nils Rehtanz, Kristiina Jokinen, Florian Matthes |
PACLIC | 4 |
| 2023 | Lessons Learned: Surveying the Practicality of Differential Privacy in the IndustryabstractSince its introduction in 2006, differential privacy has emerged as a predominant statistical tool for quantifying data privacy in academic works. Yet despite the plethora of research and open-source utilities that have accompanied its rise, with limited exceptions, differential privacy has failed to achieve widespread adoption in the enterprise domain. Our study aims to shed light on the fundamental causes underlying this academic-industrial utilization gap through detailed interviews of 24 privacy practitioners across 9 major companies. We analyze the results of our survey to provide key findings and suggestions for companies striving to improve privacy protection in their data workflows and highlight the necessary and missing requirements of existing differential privacy tools, with the goal of guiding researchers working towards the broader adoption of differential privacy. Our findings indicate that analysts suffer from lengthy bureaucratic processes for requesting access to sensitive data, yet once granted, only scarcely-enforced privacy policies stand between rogue practitioners and misuse of private information. We thus argue that differential privacy can significantly improve the processes of requesting and conducting data exploration across silos, and conclude that with a few of the improvements suggested herein, the practical use of differential privacy across the enterprise is within striking distance. Gonzalo Munilla Garrido, Florian Matthes, Dawn Song |
Proc. Priv. Enhancing Technol. | 3 |
| 2022 | Understanding the Implementation of Technical Measures in the Process of Data Privacy Compliance: A Qualitative StudyabstractBackground: Modern privacy regulations, such as the General Data Protection Regulation (GDPR), address privacy in software systems in a technologically agnostic way by mentioning general ”technical measures” for data privacy compliance rather than dictating how these should be implemented. An understanding of the concept of technical measures and how exactly these can be handled in practice, however, is not trivial due to its interdisciplinary nature and the necessary technical-legal interactions. Alexandra Klymenko, Oleksandr Kosenkov, Stephen Meisenbacher, Parisa Elahidoost, Daniel Méndez 0001, Florian Matthes |
ESEM | 6 |
| 2022 | Barriers to the Practical Adoption of Federated Machine Learning in Cross-company Collaborations
Nadine Gaertner, Nemrude Verzano, Florian Matthes |
ICAART (3) | 4 |
| 2022 | How to Simplify Law Automatically? A Study on South Korean Legislation and Its Simplified Version
Stefanie Urchs, Akshaya Muralidharan, Florian Matthes |
ICAART (3) | 3 |
| 2022 | Investigating the Current State of Security in Large-Scale Agile DevelopmentabstractAbstract Agile methods have become the established way to successfully handle changing requirements and time-to-market pressure, even in large-scale environments. Simultaneously, security has become an increasingly important concern due to more frequent and impactful incidents, stricter regulations with growing fines, and reputational damages. Despite its importance, research on how to address security in large-scale agile development is scarce. Therefore, this paper provides an empirical investigation on tackling software product security in large-scale agile environments. Based on a literature review and preliminary interviews, we identified four essential categories that impact how to handle security: (i) the structure of the agile program, (ii) security governance, (iii) adaptions of security activities to agile processes, and (iv) tool-support and automation. We conducted semi-structured interviews with nine experts from nine companies in five industries based on these categories. We performed a content-structuring qualitative analysis to reveal recurring patterns of best practices and challenges in those categories and identify differences between organizations. Among the key findings is that the analyzed organizations introduce cross-team security-focused roles collaborating with agile teams and use automation where possible. Moreover, security governance is still driven top-down, which conflicts with team autonomy in agile settings. Sascha Nägele, Jan-Philipp Watzelt, Florian Matthes |
XP | 3 |
| 2022 | Preface to the EDOC 2016 Special Issue
Florian Matthes, Jan Mendling, Stefanie Rinderle-Ma |
Inf. Syst. | 1 |
| 2022 | Revealing the landscape of privacy-enhancing technologies in the context of data markets for the IoT: A systematic literature review
Gonzalo Munilla Garrido, Johannes Sedlmeir, Ömer Uludag, Ilias Soto Alaoui, André Luckow, Florian Matthes |
J. Netw. Comput. Appl. | 6 |
| 2022 | Revealing the state of the art of large-scale agile development research: A systematic mapping study
Ömer Uludag, Pascal Philipp, Abheeshta Putta, Maria Paasivaara, Casper Lassenius, Florian Matthes |
J. Syst. Softw. | 6 |
| 2021 | Exploring privacy-enhancing technologies in the automotive value chainabstractPrivacy-enhancing technologies (PETs) are becoming increasingly crucial for addressing customer needs, security, privacy (e. g., enhancing anonymity and confidentiality), and regulatory requirements. However, applying PETs in organizations requires a precise understanding of use cases, technologies, and limitations. This paper investigates several industrial use cases, their characteristics, and the potential applicability of PETs to these. We conduct expert interviews to identify and classify uses cases, a gray literature review of relevant open-source PET tools, and discuss how the use case characteristics can be addressed using PETs’ capabilities. While we focus mainly on automotive use cases, the results also apply to other use case domains. Gonzalo Munilla Garrido, Kaja Schmidt, Christopher Harth-Kitzerow, Johannes Klepsch, André Luckow, Florian Matthes |
IEEE BigData | 6 |
| 2021 | Automated Employee Objective Matching Using Pre-trained Word EmbeddingsabstractEmployee performance objectives are recurrent short-term goals established between an organization and its employees. These objectives are typically aligned with the strategic goals of the company. Human Resources professionals are responsible for supporting employees in achieving their objectives, which in turn mediates enterprise success. One form of such support is grouping employees who have similar objectives in teams so that they can cooperate with one another. This process becomes difficult in large organizations with thousands of employees.In this paper, we present a system to assist Human Resources professionals in finding employees having similar objectives. The system automates the process of matching employees based on their common objectives using pre-trained word embeddings and semantic similarity metrics. Even though word embeddings are used to solve a variety of business problems, their effectiveness in searching for similar employee objectives is not sufficiently studied. In accordance with the feedback received from the Human Resources department of a partner company, the system proved to be effective in reducing both the time and effort consumed in the matching process. Mohab Ghanem, Ahmed Elnaggar, Adam Mckinnon, Christian Debes, Olivier Boisard, Florian Matthes |
EDOC | 6 |
| 2021 | Sentence Boundary Detection in German Legal Documents
Ingo Glaser, Sebastian Moser, Florian Matthes |
ICAART (2) | 3 |
| 2021 | Anonymization of german legal court rulingsabstractIn the legal domain, many legal documents such as court decisions and contracts are regularly anonymized. This process requires text sequences with high sensitivity to be identified and neutralized to secure sensitive information from third parties. Usually, this process is performed manually by trained employees. Therefore, anonymization is generally considered an expensive and inefficient process. This work proposes a machine learning approach for the automatic identification of sensitive text elements in German legal court decisions and provides an implementation. For this task, different deep neural network architectures based on generally pre-trained contextual embeddings as well as trained word embeddings are evaluated. Because of the lack of non-anonymized data sets, an approach to create pseudonymized data sets is proposed as well. Ingo Glaser, Thomas Schamberger, Florian Matthes |
ICAIL | 3 |
| 2021 | Data Scarcity: Methods to Improve the Quality of Text Classification
Ingo Glaser, Shabnam Sadegharmaki, Basil Komboz, Florian Matthes |
ICPRAM | 4 |
| 2021 | Generation of Legal Norm Chains: Extracting the Most Relevant Norms from Court RulingsabstractVarious online databases exist to make judgments accessible in the digital age. Before a legal practitioner can utilize state-of-the-art information retrieval features to retrieve relevant court rulings, the textual document must be processed. More importantly, many verdicts lack crucial semantic information which can be utilized within the search process. One piece of information that is frequently missed, as the judge is not adding it during the publication process within the court, is the so-called norm chain. This list contains the most relevant norms for the underlying decision. Therefore this paper investigates the feasibility of automatically extracting the most relevant norms of a court ruling. A dataset constituting over 42k labeled court rulings was used in order to train different classifiers. While our models provide F1 performances of up to 0.77, they can undoubtedly be utilized within the editorial publication process to provide helpful suggestions. Ingo Glaser, Sebastian Moser, Florian Matthes |
JURIX | 3 |
| 2021 | Lbl2Vec: An Embedding-based Approach for Unsupervised Document Retrieval on Predefined TopicsabstractIn this paper, we consider the task of retrieving documents with predefined topics from an unlabeled document dataset using an unsupervised approach. The proposed unsupervised approach requires only a small number of keywords describing the respective topics and no labeled document. Existing approaches either heavily relied on a large amount of additionally encoded world knowledge or on term-document frequencies. Contrariwise, we introduce a method that learns jointly embedded document and word vectors solely from the unlabeled document dataset in order to find documents that are semantically similar to the topics described by the keywords. The proposed method requires almost no text preprocessing but is simultaneously effective at retrieving relevant documents with high probability. When successively retrieving documents on different predefined topics from publicly available and commonly used datasets, we achieved an average area under the receiver operating characteristic curve value of 0.95 on one dataset and 0.92 on another. Further, our method can be used for multiclass document classification, without the need to assign labels to the dataset in advance. Compared with an unsupervised classification baseline, we increased F1 scores from 76.6 to 82.7 and from 61.0 to 75.1 on the respective datasets. For easy replication of our approach, we make the developed Lbl2Vec code publicly available as a ready-to-use tool under the 3-Clause BSD license. Tim Schopf, Daniel Braun 0003, Florian Matthes |
WEBIST | 3 |
| 2021 | Evolution of the Agile Scaling FrameworksabstractAbstract Over the past decade, agile methods have become the favored choice for projects undertaken in rapidly changing environments. The success of agile methods in small, co-located projects has inspired companies to apply them in larger projects. Agile scaling frameworks, such as Large Scale Scrum and Scaled Agile Framework, have been invented by practitioners to scale agile to large projects and organizations. Given the importance of agile scaling frameworks, research on those frameworks is still limited. This paper presents our findings from an empirical survey answered by the methodologists of 15 agile scaling frameworks. We explored (i) framework evolution, (ii) main reasons behind their creation, (iii) benefits, and (iv) challenges of adopting these frameworks. The most common reasons behind creating the frameworks were improving the organization’s agility and collaboration between agile teams. The most commonly claimed benefits included enabling frequent deliveries and enhancing employee satisfaction, motivation, and engagement. The most mentioned challenges were using frameworks as cooking recipes instead of focusing on changing people’s culture and mindset. Ömer Uludag, Abheeshta Putta, Maria Paasivaara, Florian Matthes |
XP | 4 |
| 2020 | The Evolution of Architectural Decision Making as a Key Focus Area of Software Architecture Research: A Semi-Systematic Literature StudyabstractLiterature review studies are essential and form the foundation for any type of research. They serve as the point of departure for those seeking to understand a research topic, as well as, helps research communities to reflect on the ideas, fundamentals, and approaches that have emerged, been acknowledged, and formed the state-of-the-art. In this paper, we present a semi-systematic literature review of 218 papers published over the last four decades that have contributed to a better understanding of architectural design decisions (ADDs). These publications cover various related topics including tool support for managing ADDs, human aspects in architectural decision making (ADM), and group decision making. The results of this paper should be treated as a getting-started guide for researchers who are entering the investigation phase of research on ADM. In this paper, the readers will find a brief description of the contributions made by the established research community over the years. Based on those insights, we recommend our readers to explore the publications and the topics in depth. Manoj Bhat, Klym Shumaiev, Uwe Hohenstein, Andreas Biesdorf, Florian Matthes |
ICSA | 5 |
| 2020 | MucLex: A German Lexicon for Surface RealisationabstractLanguage resources for languages other than English are often scarce. Rule-based surface realisers need elaborate lexica in order to be able to generate correct language, especially in languages like German, which include many irregular word forms. In this paper, we present MucLex, a German lexicon for the Natural Language Generation task of surface realisation, based on the crowd-sourced online lexicon Wiktionary. MucLex contains more than 100,000 lemmata and more than 670,000 different word forms in a well-structured XML file and is available under the Creative Commons BY-SA 3.0 license. Kira Klimt, Daniel Braun 0003, Daniela Schneider, Florian Matthes |
LREC | 4 |
| 2020 | Automatic Detection of Terms and Conditions in German and English Online Shops
Daniel Braun 0003, Florian Matthes |
WEBIST | 2 |
| 2019 | Using Enterprise Architecture Models for Creating the Record of Processing Activities (Art. 30 GDPR)abstractThe record of processing activities (RPA) is a central document in demonstrating compliance with the General Data Protection Regulation (GDPR). Article 30 of the GDPR specifies the information that has to be made available to the supervisory authority upon request. Currently, data protection management experts conduct their own data collection and maintain isolated RPAs. We show how existing Enterprise Architecture models can be augmented with the necessary information to maintain and generate an RPA. We evaluate the completeness and usefulness of the approach together with data protection management experts. Dominik Huth, Ahmet Tanakol, Florian Matthes |
EDOC | 3 |
| 2019 | Challenges in Documenting Microservice-Based IT Landscape: A Survey from an Enterprise Architecture Management PerspectiveabstractThe microservice architecture style is currently receiving much attention in both industry and academia. Many companies are already using microservices with great success and the advantages of this architectural style are discussed in numerous papers. However, microservices also introduce a high level of complexity and new challenges regarding to Enterprise Architecture (EA) model maintenance. We conducted a survey among 58 IT practitioners in the German market to analyze the status quo in the adaption of microservices and what challenges organizations face while documenting microservice-based IT landscape from an EA perspective. The eleven identified challenges are synthesized to four categories (content, assignment, tooling, and business-related challenges) and mapped to the usage of microservices which constitute the foundation for future research efforts and pose new research questions not yet considered. Martin Kleehaus, Florian Matthes |
EDOC | 2 |
| 2019 | Investigating the Establishment of Architecture Principles for Supporting Large-Scale Agile TransformationsabstractThe widespread use of agile methods shows a fundamental shift in the way organizations try to cope with unpredictable competitive environments. In large-scale agile settings, multiple development activities need to be coordinated to achieve desirable enterprise-wide effects and agility. A powerful instrument to effectively guide and steer large-scale agile endeavors is the formulation and usage of architecture principles. Despite their raison d'être to guide large organizational transformations, extant studies on how principles can be used to support large-scale agile transformations are still lacking. Against this backdrop, we present a multiple-case study involving five German companies that aims to shed light on the establishment of architecture principles to support large-scale agile transformations. Based on our results from sixteen semi-structured interviews, we present current practices as well as challenges faced by organizations during the application of architecture principles. In addition, we show a set of principles used to support large-scale agile transformations. Ömer Uludag, Henderik A. Proper, Florian Matthes |
EDOC | 3 |
| 2019 | Towards Computer-aided Analysis of Readability and Comprehensibility of Patient Information in the Context of Clinical Research ProjectsabstractClinical trials aim to examine the safety and efficiency of novel procedures such as certain medication, specific treatments, or medical devices. They constitute a compulsory requirement for the marketing authorization of these trial objects. As such, they are a vital part of a modern and innovative health care system. However, clinical trials (CT) also involve risks for the participant's health. Ingo Glaser, Georg Bonczek, Jörg Landthaler, Florian Matthes |
ICAIL | 4 |
| 2019 | Investigating the adoption and application of large-scale scrum at a German automobile manufacturerabstractOver the last two decades, agile methods have been adopted by an increasing number of organizations to improve their software development processes. In contrast to traditional methods, agile methods place more emphasis on flexible processes than on detailed upfront plans and heavy documentation. Since agile methods have proved to be successful at the team level, large organizations are now aiming to scale agile methods to the enterprise level by adopting and applying so-called scaling agile frameworks such as Large-Scale Scrum (LeSS) or Scaled Agile Framework (SAFe). Although there is a growing body of literature on large-scale agile development, literature documenting actual experiences related to scaling agile frameworks is still scarce. This paper aims to fill this gap by providing a case study on the adoption and application of LeSS in four different products of a German automobile manufacturer. Based on seven interviews, we present how the organization adopted and applied LeSS, and discuss related challenges and success factors. The comparison of the products indicates that transparency, training courses and workshops, and change management are crucial for a successful adoption. Ömer Uludag, Martin Kleehaus, Niklas Dreymann, Christian Kabelin, Florian Matthes |
ICGSE | 5 |
| 2019 | SimpleNLG-DE: Adapting SimpleNLG 4 to GermanabstractSimpleNLG is a popular open source surface realiser for the English language.For German, however, the availability of open source and non-domain specific realisers is sparse, partly due to the complexity of the German language.In this paper, we present SimpleNLG-DE, an adaption of SimpleNLG to German.We discuss which parts of the German language have been implemented and how we evaluated our implementation using the TIGER Corpus and newly created data-sets. Daniel Braun 0003, Kira Klimt, Daniela Schneider, Florian Matthes |
INLG | 4 |
| 2019 | Using Social Network Analysis to Investigate the Collaboration Between Architects and Agile Teams: A Case Study of a Large-Scale Agile Development Program in a German Consumer Electronics CompanyabstractAbstract Over the past two decades, agile methods have transformed and brought unique changes to software development practice by strongly emphasizing team collaboration, customer involvement, and change tolerance. The success of agile methods for small, co-located teams has inspired organizations to increasingly use them on a larger scale to build complex software systems. The scaling of agile methods poses new challenges such as inter-team coordination, dependencies to other existing environments or distribution of work without a defined architecture. The latter is also the reason why large-scale agile development has been subject to criticism since it neglects detailed assistance on software architecting. Although there is a growing body of literature on large-scale agile development, literature documenting the collaboration between architects and agile teams in such development efforts is still scarce. As little research has been conducted on this issue, this paper aims to fill this gap by providing a case study of a German consumer electronics retailer’s large-scale agile development program. Based on social network analysis, this study describes the collaboration between architects and agile teams in terms of architecture sharing. Ömer Uludag, Martin Kleehaus, Soner Erçelik, Florian Matthes |
XP | 4 |
| 2019 | Modeling aspects of the language of life through transfer-learning protein sequencesabstractBACKGROUND: Predicting protein function and structure from sequence is one important challenge for computational biology. For 26 years, most state-of-the-art approaches combined machine learning and evolutionary information. However, for some applications retrieving related proteins is becoming too time-consuming. Additionally, evolutionary information is less powerful for small families, e.g. for proteins from the Dark Proteome. Both these problems are addressed by the new methodology introduced here. RESULTS: We introduced a novel way to represent protein sequences as continuous vectors (embeddings) by using the language model ELMo taken from natural language processing. By modeling protein sequences, ELMo effectively captured the biophysical properties of the language of life from unlabeled big data (UniRef50). We refer to these new embeddings as SeqVec (Sequence-to-Vector) and demonstrate their effectiveness by training simple neural networks for two different tasks. At the per-residue level, secondary structure (Q3 = 79% ± 1, Q8 = 68% ± 1) and regions with intrinsic disorder (MCC = 0.59 ± 0.03) were predicted significantly better than through one-hot encoding or through Word2vec-like approaches. At the per-protein level, subcellular localization was predicted in ten classes (Q10 = 68% ± 1) and membrane-bound were distinguished from water-soluble proteins (Q2 = 87% ± 1). Although SeqVec embeddings generated the best predictions from single sequences, no solution improved over the best existing method using evolutionary information. Nevertheless, our approach improved over some popular methods using evolutionary information and for some proteins even did beat the best. Thus, they prove to condense the underlying principles of protein sequences. Overall, the important novelty is speed: where the lightning-fast HHblits needed on average about two minutes to generate the evolutionary information for a target protein, SeqVec created embeddings on average in 0.03 s. As this speed-up is independent of the size of growing sequence databases, SeqVec provides a highly scalable approach for the analysis of big data in proteomics, i.e. microbiome or metaproteome analysis. CONCLUSION: Transfer-learning succeeded to extract information from unlabeled sequence databases relevant for various protein prediction tasks. SeqVec modeled the language of life, namely the principles underlying protein sequences better than any features suggested by textbooks and prediction methods. The exception is evolutionary information, however, that information is not available on the level of a single sequence. Michael Heinzinger, Ahmed Elnaggar, Yu Wang 0008, Christian Dallago, Dmitrii Nechaev, Florian Matthes, Burkhard Rost |
BMC Bioinform. | 6 |
| 2018 | Identifying and Structuring Challenges in Large-Scale Agile Development Based on a Structured Literature ReviewabstractOver the last two decades, agile methods have transformed and brought unique changes to software development practice by strongly emphasizing team collaboration, customer involvement, and change tolerance. The success of agile methods for small, co-located teams has inspired organizations to increasingly apply agile practices to large-scale efforts. Since these methods are originally designed for small teams, unprecedented challenges occur when introducing them at larger scale, such as inter-team coordination and communication, dependencies with other organizational units or general resistances to changes. Compared to the rich body of agile software development literature describing typical challenges, recurring challenges of stakeholders and initiatives in large-scale agile development has not yet been studied through secondary studies sufficiently. With this paper, we aim to fill this gap by presenting a structured literature review on challenges in large-scale agile development. We identified 79 challenges grouped into eleven categories. Ömer Uludag, Martin Kleehaus, Christoph Caprano, Florian Matthes |
EDOC | 4 |
| 2018 | An Expert Recommendation System for Design Decision Making: Who Should be Involved in Making a Design Decision?abstractIn large software engineering projects, designing software systems is a collaborative decision-making process where a group of architects and developers make design decisions on how to address design concerns by discussing alternative design solutions. For the decision-making process, involving appropriate individuals requires objectivity and awareness about their expertise. In this paper, we propose a novel expert recommendation system that identifies individuals who could be involved in tackling new design concerns in software engineering projects. The approach behind the proposed system addresses challenges such as identifying architectural skills, quantifying architectural expertise of architects and developers, and finally matching and recommending individuals with suitable expertise to discuss new design concerns. To validate our approach, a quantitative evaluation of the recommendation system was performed using design decisions from four software engineering projects. The evaluation not only indicates that individuals with architectural expertise can be identified for design decision making but also provides quantitative evidence for the existence of personal experience bias during the decision-making process. Manoj Bhat, Klym Shumaiev, Uwe Hohenstein, Andreas Biesdorf, Florian Matthes |
ICSA | 6 |
| 2018 | Classifying Semantic Types of Legal Sentences: Portability of Machine Learning ModelsabstractLegal contract analysis is an important research area. The classification of clauses or sentences enables valuable insights such as the extraction of rights and obligations. However, datasets consisting of contracts are quite rare, particularly regarding German language. Ingo Glaser, Elena Scepankova, Florian Matthes |
JURIX | 3 |
| 2018 | Towards Explainable Semantic Text MatchingabstractThe growing amount of textual data in the legal domain leads to a demand for better text analysis tools adapted to legal domain specific use cases. Semantic Text Matching (STM) is the general problem of linking text fragments of one or more document types. The STM problem is present in many legal document analysis tasks, such as argumentation mining. A common solution approach to the STM problem is to use text similarity measures to identify matching text fragments. In this paper, we recapitulate the STM problem and a use case in German tenancy law, where we match tenancy contract clauses and legal comment chapters. We propose an approach similar to local interpretable model-agnostic explanations (LIME) to better understand the behavior of text similarity measures like TFIDF and word embeddings. We call this approach eXplainable Semantic Text Matching (XSTM). Jörg Landthaler, Ingo Glaser, Florian Matthes |
JURIX | 3 |
| 2018 | A Model-driven Approach for Generating RESTful Web Services in Single-Page Applications
Adrian Hernandez-Mendez, Niklas Scholz, Florian Matthes |
MODELSWARD | 3 |
| 2018 | Supporting Large-Scale Agile Development with Domain-Driven Design
Ömer Uludag, Matheus Hauder, Martin Kleehaus, Christina Schimpfle, Florian Matthes |
XP | 5 |
| 2018 | Preface to the EDOC 2016 Special Issue
Florian Matthes, Jan Mendling, Stefanie Rinderle-Ma |
Inf. Syst. | 1 |
| 2017 | Process and Tool-Support to Collaboratively Formalize Statutory Texts by Executable Models
Bernhard Waltl, Thomas Reschenhofer, Florian Matthes |
DEXA (2) | 3 |
| 2017 | Automatic Extraction of Design Decisions from Issue Management Systems: A Machine Learning Based Approach
Manoj Bhat, Klym Shumaiev, Andreas Biesdorf, Uwe Hohenstein, Florian Matthes |
ECSA | 5 |
| 2017 | Investigating the Role of Architects in Scaling Agile FrameworksabstractThis study describes the roles of architects in scaling agile frameworks with the help of a structured literature review. We aim to provide a primary analysis of 20 identified scaling agile frameworks. Subsequently, we thoroughly describe three popular scaling agile frameworks: Scaled Agile Framework, Large Scale Scrum, and Disciplined Agile 2.0. After specifying the main concepts of scaling agile frameworks, we characterize roles of enterprise, software, solution, and information architects, as identified in four scaling agile frameworks. Finally, we provide a discussion of generalizable findings on the role of architects in scaling agile frameworks. Ömer Uludag, Martin Kleehaus, Florian Matthes |
EDOC | 4 |
| 2017 | SaToS: Assessing and Summarising Terms of Services from German WebshopsabstractEvery time we buy something online, we are confronted with Terms of Services.However, only a few people actually read these terms, before accepting them, often to their disadvantage.In this paper, we present the SaToS browser plugin which summarises and simplifies Terms of Services from German webshops. Daniel Braun 0003, Elena Scepankova, Patrick Holl, Florian Matthes |
INLG | 4 |
| 2017 | Classifying Legal Norms with Active Machine LearningabstractThis paper describes an extended machine learning approach to classify legal norms in German statutory texts. We implemented an active machine learning (AML) framework based on open-source software. Within the paper we discuss different query strategies to optimize the selection of instances during the learning phase to decrease the required training data. Bernhard Waltl, Johannes Muhr, Ingo Glaser, Georg Bonczek, Elena Scepankova, Florian Matthes |
JURIX | 6 |
| 2017 | Evaluating Natural Language Understanding Services for Conversational Question Answering SystemsabstractConversational interfaces recently gained a lot of attention.One of the reasons for the current hype is the fact that chatbots (one particularly popular form of conversational interfaces) nowadays can be created without any programming knowledge, thanks to different toolkits and socalled Natural Language Understanding (NLU) services.While these NLU services are already widely used in both, industry and science, so far, they have not been analysed systematically.In this paper, we present a method to evaluate the classification performance of NLU services.Moreover, we present two new corpora, one consisting of annotated questions and one consisting of annotated questions with the corresponding answers.Based on these corpora, we conduct an evaluation of some of the most popular NLU services.Thereby we want to enable both, researchers and companies to make more educated decisions about which service they should use. Daniel Braun 0003, Adrian Hernandez-Mendez, Florian Matthes, Manfred Langen |
SIGDIAL Conference | 3 |
| 2017 | Columbus: A Tool for Discovering User Interface Models in Component-basedWeb Applications
Adrian Hernandez-Mendez, Andreas Tielitz, Florian Matthes |
WEBIST | 3 |
| 2016 | Multi-Level Event and Anomaly Correlation Based on Enterprise Architecture Information
Jörg Landthaler, Martin Kleehaus, Florian Matthes |
EOMAS@CAiSE | 3 |
| 2016 | Extending Full Text Search for Legal Document Collections Using Word EmbeddingsabstractTraditional full text search allows fast search for exact matches. However, full text search is not optimal to deal with synonyms or semantically related terms and phrases. In this paper we explore a novel method that provides the ability to find not only exact matches, but also semantically similar parts for arbitrary length search queries. We achieve this without the application of ontologies, but base our approach on Word Embeddings. Recently, Word Embeddings have been applied successfully for many natural language processing tasks. We argue that our method is well suited for legal document collections and examine its applicability for two different use cases: We conduct a case study on a stand-alone law, in particular the EU Data Protection Directive 94/46/EC (EU-DPD) in order to extract obligations. Secondly, from a collection of publicly available templates for German rental contracts we retrieve similar provisions. Jörg Landthaler, Bernhard Waltl, Patrick Holl, Florian Matthes |
JURIX | 4 |
| 2016 | Differentiation and Empirical Analysis of Reference Types in Legal DocumentsabstractThis paper proposes an extensible model distinguishing between reference types within legal documents. It differentiates between four types of references, namely fully-explicit, semi-explicit, implicit, and tacit references. Bernhard Waltl, Jörg Landthaler, Florian Matthes |
JURIX | 3 |
| 2016 | Supporting end-users in defining complex queries on evolving and domain-specific data modelsabstractTo define queries on domain-specific data-structures, end-users have to be familiar with the underlying data model. This is challenging if the data model evolves continuously and through collaborative model management. While visual languages address this issue by providing strong guidance for non-technical end-users, they suffer from a limited expressiveness due to their focus on usability. Therefore, they are not applicable in cases where domain-experts have the need to dedline complex queries. This paper proposes an interactive and textual query editor which focuses on the formulation of complex queries and the familiarization of the end-user with the underlying data model during the query definition process. We demonstrate our approach by this query editor's application in the domain of Enterprise Architecture Management, where tech-savvy end-users need to define complex metrics based on an evolving enterprise architecture model. Thomas Reschenhofer, Florian Matthes |
VL/HCC | 2 |
| 2015 | A Data Science Environment for Legal TextsabstractSince the beginning of legal informatics, data analysis has been very attractive to the scientific community in different domains and areas. Modern algorithms and mining technologies are able to unveil structural, such as network-like features and semantic properties that might be contained explicitly or implicitly in texts. This becomes evident by screening the topics of recent scientific conferences and workshops (e.g., ICAIL 2015). However, only few efforts were spent on the development of a generic environment taking into account the characteristics of legislative systems, which would foster the re-usage of results, such as models, components or frameworks. This paper proposes a reference architecture, which can easily be extended to specific research questions and use cases. It is designed and implemented as a generic environment allowing state-of-the-art text analysis. We performed a case study on the German tenancy law. Thereby, we analyzed the evolution the law in the last 25 years. Bernhard Waltl, Marin Zec, Florian Matthes |
JURIX | 3 |
| 2014 | Collaborative Annotation of Recorded Teaching Video SessionsabstractDriven by technology advances the availability of digital video recordings of live training sessions increases at a fast pace. The goal of our research is to better understand the impact of these digital artifacts on the individual (and possibly collaborative) note-taking process of learners. In this paper, we develop a conceptual framework describing the augmentation of teaching sessions by computer-supported tools. We use the framework to describe related work and to outline our research design that involves the development of a minimum viable collaborative annotation tool and the study of the effects of variations in tool functionality (like visibility of annotations, kinds of annotations, or form of annotations) on the learning process. Florian Matthes, Klym Shumaiev |
CSEDU (1) | 1 |
| 2014 | Towards Measures of Complexity: Applying Structural and Linguistic Metrics to German LawsabstractThe increasing complexity of legal systems has many origins, which are worth a deeper analysis. This paper is an attempt to unveil the complexity in legal texts driven by structural, lexical and syntactical properties. Thereby we transferred established quantitative methods from structural network analysis and linguistics into the domain of legal text analysis. Based on 3 553 German laws, respectively regulations, we calculated several structural and lexical indicators for complexity and determined highly significant correlations (p ≤ 0.01). The papers' contribution is a set of metrics, enabling a structured and objective comparison of legal texts regarding their complexity. Bernhard Waltl, Florian Matthes |
JURIX | 2 |
| 2014 | Guest editorial to the Theme Section on enterprise modelling
Tony Clark 0001, Florian Matthes, Balbir S. Barn, Alan W. Brown |
Softw. Syst. Model. | 2 |
| 2012 | Towards Automated Enterprise Architecture Documentation: Data Quality Aspects of SAP PI
Sebastian Grunow, Florian Matthes, Sascha Roth |
ADBIS (2) | 2 |
| 2011 | Teaching Global Software Engineering and International Project Management - Experiences and Lessons Learned from Four Academic Projects
Florian Matthes, Christian Neubert, Christopher Schulz, Christian Lescher, José Contreras, Robert Laurini, Béatrice Rumpler, David Sol, Kai Warendorf |
CSEDU (2) | 1 |
| 2011 | Modeling the Supply and Demand of Architectural Information on Enterprise LevelabstractEnterprise architecture (EA) management aims at analyzing and improving the enterprise as a whole. A correct and consistent analysis is based on reliable EA data. However, current industrial practice shows that many persons need to collect, prepare, and disseminate EA relevant data while only a small group of persons actually benefits from this information. This state of affairs has a negative impact on the motivation of those who are in charge of gathering EA information. Additionally, the monetary value of this information is often implicit. To overcome this situation, this paper presents an approach to model the supply and demand situation for EA information. The resulting model helps to understand, explain, and ease EA-related information gathering. The applicability of the resulting model is demonstrated with the help of a real world case study from the German federal government. Sabine Buckl, Andreas Gehlert, Florian Matthes, Christopher Schulz, Christian M. Schweda |
EDOC | 3 |
| 2011 | Testing & quality assurance in data migration projectsabstractNew business models, constant technological progress, as well as ever-changing legal regulations require that companies replace their business applications from time to time. As a side effect, this demands for migrating the data from the existing source application to a target application. Since the success of the application replacement as a form of IT maintenance is contingent on the underlying data migration project, it is crucial to accomplish the migration in time and on budget. This however, calls for a stringent data migration process model combined with well-defined quality assurance measures. The paper presents, first, a field-tested process model for data migration projects. Secondly, it points out the typical risks we have frequently observed in the course of this type of project. Thirdly, the paper provides practice-based testing and quality assurance techniques to reduce or even eliminate these data migration risks. Florian Matthes, Christopher Schulz, Klaus Haller |
ICSM | 1 |
| 2011 | Hybrid Wikis: Empowering Users to Collaboratively Structure Information
Florian Matthes, Christian Neubert, Alexander Steinhoff |
ICSOFT (1) | 1 |
| 2010 | A Technique for Annotating EA Information Models with Goals
Sabine Buckl, Florian Matthes, Christian M. Schweda |
EOMAS | 2 |
| 2010 | A situated approach to enterprise architecture managementabstractToday's enterprises are confronted with a challenging environment that demands continuous transformations. Globalized markets, disruptive technological innovations, and new legal regulations call for enterprises, which flexibly adapt to these requirements. A commonly accepted means to guide such enterprise transformations is enterprise architecture (EA) management. Enterprises seeking to introduce and establish such a management function see themselves confronted via a plethora of tools, approaches, and frameworks that claim to provide “the definitive design prescriptions” for an EA management function. The applicability of the different prescriptions nevertheless heavily depends on the organizational context and the EA-related goals that the enterprise wants to pursue. This paper presents an extensible set of selection guidelines, that helps enterprises to choose the EA management approach best suited for their goals and context. These selection guidelines are linked to different EA management approaches and frameworks, which are related to organizational contexts, in which they can operate, and EA management goals, that the approaches can help to pursue. Utilizing the selection guidelines, an enterprise can specify the applicable context as well as its EA management goals, and is provided with a selection of suitable approaches. Finally an outlook critically reflects the findings of the paper and provides an outlook on future areas of research. Sabine Buckl, Christian M. Schweda, Florian Matthes |
SMC | 3 |
| 2009 | Using Enterprise Architecture Management Patterns to Complement TOGAFabstractThe design of an enterprise architecture (EA) management function for an enterprise is no easy task. Various frameworks exist as well as EA management tools, which promise to deliver guidance for performing EA management. Nevertheless, the approaches presented by them stay either on a level too abstract to provide realization support or are far too generic, neglecting enterprise-specific EA related concerns. In this article, we discuss the architecture framework of The Open Group (TOGAF) and detail on its promising but nevertheless highly generic architecture development method (ADM). This article shows how the generic development steps can be complemented by a pattern based approach to EA management providing guidance for addressing specific EA related concerns with step-by-step methodologies as well as with corresponding viewpoints and information models. Sabine Buckl, Alexander M. Ernst, Florian Matthes, René Ramacher, Christian M. Schweda |
EDOC | 3 |
| 2009 | Functional Analysis of Enterprise 2.0 Tools: A Services Catalog
Thomas Büchner, Florian Matthes, Christian Neubert |
IC3K | 2 |
| 2009 | A Viable System Perspective on Enterprise Architecture ManagementabstractA number of approaches towards enterprise architecture (EA) management is proposed in literature, differing in the underlying understanding of the EA as well as in the description of the function for performing EA management. These plurality of methods and models should be interpreted as an indicator of the low maturity of the research area. In contrast, some researchers see it as inevitable consequence of the diversity of the enterprises under consideration. Staying to this interpretation, we approach the topic of EA management from a cybernetic point of view. Thereby, we elicit constituents, which should be considered in every EA management function based on a viable system perspective on the topic. From this perspective, we further revisit selected EA management approaches and show to which extent they allude to the viable system nature of the EA. Sabine Buckl, Christian M. Schweda, Florian Matthes |
SMC | 3 |
| 2005 | Improving IT Management at the BMW Group by Integrating Existing IT Management ProcessesabstractThe management of IT landscapes consisting of thousands of business applications, different middleware systems, and supporting various business processes is a challenge for modern IT management. The BMW Group has addressed this challenge by establishing an integrated IT management process which covers strategy, architecture, planning and controlling. The BMW Group integrates the following preexisting IT processes in a continuous management process: "architecture and standardization" defines blueprints in terms of architectural patterns; "landscape management" maintains an overall model of the (present and future) IT landscape; "portfolio management" coordinates, evaluates and prioritizes action items with IT impact; "synchronization management" manages ongoing (IT) projects, their dependencies and the cross functions; "strategy and objectives" generates action items, defines guidelines for the other processes and adjusts strategies based on feedback from the other processes. This paper describes the problems that arise if such an integrated, continuous management process is lacking. Florian Matthes, André Wittenburg |
EDOC | 2 |
| 2000 | A Process-Oriented and Content-Based Perspective on Software Components
Holm Wegner, Patrick Hupe, Florian Matthes |
Inf. Syst. | 3 |
| 1999 | A Process-Oriented Approach to Software Component Definition
Florian Matthes, Holm Wegner, Patrick Hupe |
CAiSE | 1 |
| 1998 | SAP R/3: A Database Application System (Tutorial)abstractMany database applications in the real world are no longer built on top of a stand-alone database system. Rather, generic (standard) application systems are employed in which the database system is one integrated component. SAP is the market leader for integrated business administration systems, and its SAP R/3 product is a comprehensive software system which integrates modules for finance, material management, sales and distribution, etc. From an architectural point of view, SAP R/3 is a client/server application system with a relational database system as back-end. SAP supports a choice between a variety of commercial relational database products. Alfons Kemper, Donald Kossmann, Florian Matthes |
SIGMOD Conference | 3 |
| 1997 | On Migrating Threads
Bernd Mathiske, Florian Matthes, Joachim W. Schmidt |
J. Intell. Inf. Syst. | 2 |
| 1996 | Integrating Subtyping, Matching and Type Quantification: A Practical Perspective
Andreas Gawecki, Florian Matthes |
ECOOP | 2 |
| 1996 | Exploiting Persistent Intermediate Code Representations in Open Database Environments
Andreas Gawecki, Florian Matthes |
EDBT | 2 |
| 1994 | Persistent Threads
Florian Matthes, Joachim W. Schmidt |
VLDB | 1 |
| 1994 | The DBPL Project: Advances in Modular Database Programming
Joachim W. Schmidt, Florian Matthes |
Inf. Syst. | 2 |
| 1993 | Viewers: A Data-World Analogue of Procedure Calls
Kazimierz Subieta, Florian Matthes, Joachim W. Schmidt, Andreas Rudloff |
VLDB | 2 |