Jens Dörpinghaus

dblp:206/3321 · DBLP profile ↗
← Back
21ranked-venue papers
17as first author
13since 2021 · last 2026
0000-0003-0245-7752ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 13 first-author · 8 since 2021Software engineering, systems software and programming languages · 15 · 12 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 12 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Theory of computation · 2 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 The transformative potential of AI in software engineering: a case study on LeetCode and ChatGPT
abstract
Abstract The recent surge in the field of generative artificial intelligence (GenAI) has the potential to bring about transformative changes across a range of sectors, including software engineering and education. As GenAI tools, such as OpenAI’s ChatGPT, are increasingly utilised in software engineering, it becomes imperative to understand the impact of these technologies on the software product. This study employs a methodological approach, comprising web scraping and data mining from LeetCode, with the objective of comparing the software quality of Python programs produced by LeetCode users with that generated by GPT-4o. In order to gain insight into these matters, this study addresses the question whether GPT-4o produces software of superior quality to that produced by humans. The findings indicate that GPT-4o does not present a considerable impediment to code quality, understandability, or runtime when generating code on a limited scale. Indeed, the generated code even exhibits significantly better values across all the three code quality dimensions in comparison to the user-written code. However, no significantly superior values were observed for the generated code in terms of memory usage in comparison to the user code, which contravened the expectations. Furthermore, it will be demonstrated that GPT-4o encountered challenges in generalising to problems that were not included in the training data set. This contribution presents a first large-scale study comparing generated code with human-written code based on LeetCode platform based on multiple measures including code quality, code understandability, time behaviour and resource utilisation. All data is publicly available for further research.
Manuel Merkel, Jens Dörpinghaus
Empir. Softw. Eng.2
2025 Do LLMs dream of antique hermeneutics? Critical remarks on automated text interpretation
abstract
This study investigates the potential of large language models (LLMs) to apply hermeneutical methods rooted in philosophy, theology, sociology and literary studies in a meaningful manner.Utilising a comparative experimental design, four LLMs were prompted to interpret a variety of texts, encompassing religious, philosophical, poetic, and conversational material.The findings indicate considerable variability, an absence of reproducibility, and a substantial reliance on prompt design, model type, and language.This suggests that LLMs do not employ coherent hermeneutical strategies.Despite the automation of formal features, limitations in context sensitivity, interpretive intentionality, and epistemic grounding render LLMs ill-suited to authentic hermeneutics.The study concludes that current LLMs are incapable of replicating the depth of human interpretive practice, and calls for further interdisciplinary research to define evaluation standards for machine-assisted exegesis.
Jens Dörpinghaus, Michael Tiemann 0002
FedCSIS1
2024 Towards the analysis of errors in centrality measures in perpetuated networks
abstract
Centrality measures are essential tools for analyzing the structure and dynamics of graphs, such as knowledge or social networks.They reveal the significance and influence of individual nodes.However, their accuracy can be influenced by data quality, algorithms, and network properties.This study investigates errors in centrality measures within perpetuated networks.It focuses on network resilience and how these results may be used to develop efficient algorithms for centrality measures.It also investigates how perturbation strategies impact network resilience and predict connectivity in the perturbed network.By employing centrality measures (degree, betweenness, closeness, eigenvector), we identify critical nodes that significantly affect network connectivity and information flow.Additionally, statistical tests (Kolmogorov-Smirnov, Cramér-von Mises) assess network robustness and pinpoint critical transition points.This study, by outlining methods for error identification, quantification, and mitigation, offers valuable insights for enhancing network resilience across various domains, including infrastructure design and social network analysis.
Meetkumar Pravinbhai Mangroliya, Jens Dörpinghaus, Robert Rockenfeller
FedCSIS2
2024 IT Professionals in Germany. Labor Market Demands of Computer Science Education and their Perception on Social Media
abstract
The skills and qualifications of IT professionals are constantly changing and under discussion. In particular, we see the impact of emerging technologies and tools, such as AI, on the labor market and their reflection in the broader scientific community, media and everyday life [2, 4]. Several approaches have discussed how to assess the impact of computer science education from the perspective of education and labor market research [5, 7]. However, in order to uncover the complex dynamics surrounding computer science education, we need to take a closer look at the labor demand as well as the public perception and valuation of IT professionals. Therefore, the poster presents a new method that combines the analysis of online job advertisements (OJA) for computer science occupations at different qualification levels with social media data (Twitter/X and YouTube). As OJA and social media are usually described in unstructured natural language, text mining methods are key to extract information about skills and competencies. We used the Computer Science Ontology (CSO) for annotation [1] and outline further research questions in the wider context of political economy.
Stefan Udelhofen, Jens Dörpinghaus
ITiCSE (2)2
2024 Automated annotation of parallel bible corpora with cross-lingual semantic concordance
abstract
Abstract Here we present an improved approach for automated annotation of New Testament corpora with cross-lingual semantic concordance based on Strong’s numbers. Based on already annotated texts, they provide references to the original Greek words. Since scientific editions and translations of biblical texts are often not available for scientific purposes and are rarely freely available, there is a lack of up-to-date training data. In addition, since annotation, curation, and quality control of alignments between these texts are expensive, there is a lack of available biblical resources for scholars. We present two improved approaches to the problem, based on dictionaries and already annotated biblical texts. We provide a detailed evaluation of annotated and unannotated translations. We also discuss a proof of concept based on English and German New Testament translations. The results presented in this paper are novel and, to our knowledge, unique. They show promising performance, although further research is needed.
Jens Dörpinghaus
Nat. Lang. Eng.1
2023 Classifying Industrial Sectors from German Textual Data with a Domain Adapted Transformer
abstract
For economics and sociological research, lists of industries and their branches are widely used in research to categorize data and get an overview on different types of industries.However, many different taxonomies and ordering schema exist, due to different research focus but also due to different national scenarios and interests.In this paper, we will focus without loss of generality on regional data from Germany.Manual annotation of textual data is time-consuming and tedious, naturally giving rise to our initial research question, also highly inspired by questions from computational social sciences: How can we automatically categorize textual data, e.g.job advertisements or business profiles, by industrial sectors?We will present an approach towards classification using a pre-trained domain-adapted Transformer model.We find that domain-adapted models generalize better and outperform state of the art non domain-adapted Transformer models on Out-Of-Distribution data.Additionally, we open source two novel datasets mapping textual data to WZ2008 sections and divisions, enabling further research.
Richard Fechner, Jens Dörpinghaus, Anja Firll
FedCSIS2
2023 Lessons from Continuing Vocational Training Courses for Computer Science Education
abstract
The labor market heavily relies on both vocational and academic education and training, re-training and advanced vocational qualification to meet challenges, e.g. the advancing digitalization[2, 3]. Continuing education is a central prerequisite for securing skilled labor, for ensuring the employability of all employees and thus also for national competitiveness and innovation. From the perspective of education and labor market research, several approaches discuss how the impact of computer science education can be evaluated. Other research focuses on the needs of the labor market, by analysing job advertisements. In order to broaden the perspective on the entire range of CVET courses and to be able to gain new insights from this, our analysis is intended to provide an initial overview of the content of CVET courses in Germany. By that, we offer structured information on skills and competencies that are included in current CVET courses. In future research, this information can be compared to labor market needs, e.g. described in job advertisements, in order to identify education gaps.
Jens Dörpinghaus, Johanna Binnewitt, Kristine Hein
ITiCSE (2)1
2022 A novel link prediction approach on clinical knowledge graphs utilizing graph structures
abstract
This paper presents a novel approach towards link prediction in clinical knowledge graphs.They play a central role in linking data from different data sources and are widely used in big data integration, especially for connecting data from different domains.We present a knowledge graph initially built on data from a clinical trial on Spinocerebellar ataxia type 3 (SCA3), which is a rare autosomal dominant inherited disorder.The contributions of this paper are (1) to create a feasible data representation schema capable of handling clinical imaging data in a knowledge graph and to ( 2) convert the data efficiently into a knowledge graph.Due to the limited amount of patientnodes usually common methods for link prediction and graph embeddings are problematic and thus we will (3) present a novel approach for link prediction utilising graph structures and Conditional Random Fields.In addition, we present (4) an extensive evaluation underlining the importance of (a) data management and (b) further research on link prediction using graph structures.
Jens Dörpinghaus, Tobias Hübenthal, Jennifer Faber
FedCSIS1
2022 Analyzing longitudinal Data in Knowledge Graphs utilizing shrinking pseudo-triangles
abstract
This paper aims to analyze longitudinal data, serial data related to different time points, in knowledge graphs.Knowledge graphs play a central role for linking different data.While multiple layers for data from different sources are considered, there is only very limited research on longitudinal data in knowledge graphs.However, knowledge graphs are widely used in big data integration, especially for connecting data from different domains.Few studies have investigated the questions how multiple layers and time points within graphs impact methods and algorithms developed for single-purpose networks.This manuscript investigates the impact of a modeling of longitudinal data in multiple layers on retrieval algorithms.In particular, (a) we propose a first draft of a generic model for longitudinal data in multi-layer knowledge graphs, (b) we develop an experimental environment to evaluate a generic retrieval algorithm on random graphs inspired by computational social sciences.We present a knowledge graph generated on German job advertisements comprising data from different sources, both structured and unstructured, on data between 2011 and 2021.The data is linked using text mining and natural language processing methods.We further (c) present two different shrinking techniques for structured and unstructured layers in knowledge based on graph structures like triangles and pseudo-triangles.The presented approach (d) shows that on the one hand, the initial research questions, on the other hand the graph structures and topology have a great impact on the structures and efficiency for additional data stored.Although the experimental analysis of random graphs allows us to make some basic observations we will (e) make suggestions for additional research on particular graph structures that have a great impact on the analysis of knowledge graph structures.
Jens Dörpinghaus, Vera Weil, Johanna Binnewitt
FedCSIS1
2022 Towards the Analysis of Longitudinal Data in Knowledge Graphs on Job Ads
Jens Dörpinghaus, Vera Weil, Johanna Binnewitt
WCO1
2022 Context mining and graph queries on giant biomedical knowledge graphs
abstract
Abstract Contextual information is widely considered for NLP and knowledge discovery in life sciences since it highly influences the exact meaning of natural language. The scientific challenge is not only to extract such context data, but also to store this data for further query and discovery approaches. Classical approaches use RDF triple stores, which have serious limitations. Here, we propose a multiple step knowledge graph approach using labeled property graphs based on polyglot persistence systems to utilize context data for context mining, graph queries, knowledge discovery and extraction. We introduce the graph-theoretic foundation for a general context concept within semantic networks and show a proof of concept based on biomedical literature and text mining. Our test system contains a knowledge graph derived from the entirety of PubMed and SCAIView data and is enriched with text mining data and domain-specific language data using Biological Expression Language. Here, context is a more general concept than annotations. This dense graph has more than 71M nodes and 850M relationships. We discuss the impact of this novel approach with 27 real-world use cases represented by graph queries. Storing and querying a giant knowledge graph as a labeled property graph is still a technological challenge. Here, we demonstrate how our data model is able to support the understanding and interpretation of biomedical data. We present several real-world use cases that utilize our massive, generated knowledge graph derived from PubMed data and enriched with additional contextual data. Finally, we show a working example in context of biologically relevant information using SCAIView.
Jens Dörpinghaus, Andreas Stefan, Bruce Schultz, Marc Jacobs 0001
Knowl. Inf. Syst.1
2021 Automated creation of parallel Bible corpora with cross-lingual semantic concordance
abstract
Here we present a novel approach for automated creation of parallel New Testament corpora with cross-lingual semantic concordance based on Strong's numbers.As scientific editions and translations of Bible texts are often not free to use for scientific purposes and are rarely free to use, and due to the fact that the annotation, curation and quality control of alignments between these texts are quite expensive, there is a lack of available Biblical resources for scholars.We present two approaches to tackle the problem, a dictionary-based approach and a Conditional Random Field (CRF) model and a detailed evaluation on annotated and non-annotated translations.We discuss a proof-of-concept based on English and German New Testament translations.The results presented in this paper are novel and according to our knowledge unique.They present promising performance, although further research is necessary.
Jens Dörpinghaus, Carsten Düing
FedCSIS1
2021 An efficient approach towards the generation and analysis of interoperable clinical data in a knowledge graph
abstract
Knowledge graphs have been shown to play an important role in recent knowledge mining settings, for example in the fields of life sciences or bioinformatics.Contextual information is widely used for NLP and knowledge discovery tasks, since it highly influences the exact meaning of expressions and also queries on data.The contributions of this paper are (1) an efficient approach towards interoperable data, (2) a runtime analysis of 14 realworld use cases represented by graph queries and (3) a unique view on clinical data and its application, combining methods of algorithmic optimisation, graph theory and data science.
Jens Dörpinghaus, Sebastian Schaaf, Vera Weil, Tobias Hübenthal
FedCSIS1
2020 Knowledge Detection and Discovery using Semantic Graph Embeddings on Large Knowledge Graphs generated on Text Mining Results
abstract
Knowledge graphs play a central role in big data integration, especially for connecting data from different domains.Bringing unstructured texts, e.g. from scientific literature, into a structured, comparable format is one of the key assets.Here, we use knowledge graphs in the biomedical domain working together with text mining based document data for knowledge extraction and retrieval from text and natural language structures.For example cause and effect models, can potentially facilitate clinical decision making or help to drive research towards precision medicine.However, the power of knowledge graphs critically depends on context information.Here we provide a novel semantic approach towards a context enriched biomedical knowledge graph utilizing data integration with linked data applied to language technologies and text mining.This graph concept can be used for graph embedding applied in different approaches, e.g with focus on topic detection, document clustering and knowledge discovery.We discuss algorithmic approaches to tackle these challenges and show results for several applications like search query finding and knowledge discovery.The presented remarkable approaches lead to valuable results on large knowledge graphs.
Jens Dörpinghaus, Marc Jacobs 0001
FedCSIS1
2020 Optimization of Retrieval Algorithms on Large Scale Knowledge Graphs
abstract
Knowledge graphs have been shown to play an important role in recent knowledge mining and discovery, for example in the field of life sciences or bioinformatics. Although a lot of research has been done on the field of query optimization, query transformation and of course in storing and retrieving large scale knowledge graphs the field of algorithmic optimization is still a major challenge and a vital factor in using graph databases. Few researchers have addressed the problem of optimizing algorithms on large scale labeled property graphs. Here, we present two optimization approaches and compare them with a naive approach of directly querying the graph database. The aim of our work is to determine limiting factors of graph databases like Neo4j and we describe a novel solution to tackle these challenges. For this, we suggest a classification schema to differ between the complexity of a problem on a graph database. We evaluate our optimization approaches on a test system containing a knowledge graph derived biomedical publication data enriched with text mining data. This dense graph has more than 71M nodes and 850M relationships. The results are very encouraging and - depending on the problem - we were able to show a speedup of a factor between 44 and 3839.
Jens Dörpinghaus, Andreas Stefan
FedCSIS1
2020 Semantic Graph Queries on Linked Data in Knowledge Graphs
Jens Dörpinghaus, Andreas Stefan
WCO@FedCSIS1
2019 A Minimum Set-Cover Problem with several constraints
abstract
A lot of problems in natural language processing can be interpreted using structures from discrete mathematics.In this paper we will discuss the search query and topic finding problem using a generic context-based approach.This problem can be described as a Minimum Set Cover Problem with several constraints.The goal is to find a minimum covering of documents with the given context for a fixed weight function.The aim of this problem reformulation is a deeper understanding of both the hierarchical problem using union and cut as well as the nonhierarchical problem using the union.We thus choose a modeling using bipartite graphs and suggest a novel reformulation using an integer linear program as well as novel graph-theoretic approaches.
Jens Dörpinghaus, Carsten Düing, Vera Weil
FedCSIS1
2019 Knowledge Extraction and Applications utilizing Context Data in Knowledge Graphs
abstract
Context is widely considered for NLP and knowledge discovery since it highly influences the exact meaning of natural language. The scientific challenge is not only to extract such context data, but also to store this data for further NLP approaches. Here, we propose a multiple step knowledge graphbased approach to utilize context data for NLP and knowledge expression and extraction. We introduce the graph-theoretic foundation for a general context concept within semantic networks and show a proof-of-concept-based on biomedical literature and text mining. We discuss the impact of this novel approach on text analysis, various forms of text recognition and knowledge extraction and retrieval.
Jens Dörpinghaus, Andreas Stefan
FedCSIS1
2018 What was the Question? A Systematization of Information Retrieval and NLP Problems
abstract
In this paper we suggest a novel systematization of Information Retrieval and Natural Language Processing problems.Using this rather general description of problems we are able to discuss and proof the equivalence of some problems.We provide reformulations of well-known problems like Named Entity Recognition using our novel description and discuss further research and the expected outcome.We will discuss the relation of two problems, cluster labeling and search query finding.With these results we are able to provide a novel optimization approach to both problems.This novel systematization approach provides a yet unknown view generating new classes of problems in NLP.It brings application and algorithmic approaches together and offers a better description with concepts of theoretical computer science.
Jens Dörpinghaus, Johannes Darms, Marc Jacobs 0001
FedCSIS1
2018 A Graph-Theoretic Approach to the Train Marshalling Problem
abstract
Rearranging cars of an incoming train in a hump yard is a widely discussed topic.We focus on the train marshalling problem where the incoming cars of a train are distributed to a certain number of sorting tracks.When pulled out again to build the outgoing train, cars sharing the same destination should appear consecutively.The goal is to minimize the number of sorting tracks.We suggest a graph-theoretic approach for this N P-complete problem.The idea is to partition an associated directed graph into what we call pseudochains of minimum length.We describe a greedy-type heuristic to solve the partitioning problem which, on random instances, performs better than the known heuristics for the train marshalling problem.
Jens Dörpinghaus, Rainer Schrader
FedCSIS1
2017 Document Clustering using a Graph Covering with Pseudostable Sets
abstract
In text mining, document clustering describes the efforts to assign unstructured documents to clusters, which in turn usually refer to topics.Clustering is widely used in science for data retrieval and organisation.In this paper we present a new graph theoretical approach to document clustering and its application on a real-world data set.We will show that the wellknown graph partition to stable sets or cliques can be generalized to pseudostable sets or pseudocliques.This allows to make a soft clustering as well as a hard clustering.We will present an integer linear programming and a greedy approach for this NP-complete problem and discuss some results on random instances and some real world data for different similarity measures.
Jens Dörpinghaus, Sebastian Schaaf, Juliane Fluck, Marc Jacobs 0001
FedCSIS1