Matthias L. Hemmje

dblp:h/MatthiasHemmje · DBLP profile ↗
← Back
61ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0001-8293-2802ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 23 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 8 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 since 2021Artificial intelligence and machine learning · 6Software engineering, systems software and programming languages · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Trustable, Reproducible, and Intelligent Information Visualization Systems (TRI-IVIS)
abstract
TRI-IVIS focuses on Trustable, Reproducible and Intelligent Information Visualization Systems, aiming to advance methods and systems for managing and analyzing large-scale complex data, integrating Artificial Intelligence, Machine Learning, Big Data analytics and advanced visual interfaces. While there are many domains that may benefit from such characteristics, preferred domains for the workshop are, but not limited to, genomic applications and Meetings, Incentives, Conferences and Events (MICE). Many topics revolve around such domains, ranging from heterogeneous large-scale data, distributed storage, multi-stakeholder scenarios, and regulatory requirements.
Paolo Buono, Philippe Tamla, Thomas Krause 0004, Haithem Afli, Matthias L. Hemmje
AVI5
2025 From Reads to Reports: A Vision for a GFM-Powered Genomic Diagnostic Platform
abstract
Microbiome sequencing offers significant potential for advancing clinical diagnostics, but its adoption is hindered by challenges in data processing, standardization, and the translation of complex genomic data into actionable clinical insights. The GenDAI project addresses these challenges by developing a novel, integrated medical diagnostics platform that leverages Artificial Intelligence (AI), powered by Genomic Foundation Models (GFMs), to accelerate and improve the analysis of microbiome data. The platform's primary goal is to provide a fully automated, reproducible, and compliant end-to-end solution, from data ingestion to clinical reporting, to support personalized medicine, with an initial focus on Inflammatory Bowel Disease (IBD). This paper presents an overview of GenDAI's vision, outlining its user-centered methodological approach and the conceptual architecture of its core components. The architecture integrates four key pillars: (1) a fully automated and auditable diagnostics workflow, (2) a secure and compliant cloud platform for long-term data management based on Open Archival Information System (OAIS) and Findable, Accessible, Interoperable, Reusable (FAIR) principles, (3) an advanced AI engine for biomarker discovery using GFMs, and (4) interactive, usercentered reporting tools designed to enhance explainability and clinical trust. By providing a holistic and ethically-grounded framework, GenDAI aims to bridge the gap between advanced genomic research and practical clinical application.
Thomas Krause 0004, Philippe Tamla, Andrea Leoni, Flavia Monti, Francesca De Luzi, Jamie Fitz Gerald, Bruno G. Andrade, Haithem Afli, Massimo Mecella, Paolo Buono, Andrea Molinari, Matthias L. Hemmje
BIBM12
2025 Genomic Foundation Models for SNP Analysis
abstract
This paper presents SnipFlow, a workflow for knowledge-based analysis of single nucleotide polymorphisms in personalised medicine. SnipFlow combines structured knowledge derived from the genome sequence with unstructured literaturebased knowledge in a common data model and presents the results transparently via a knowledge management system. By integrating Genomic Foundation Models (GFMs), the workflow can be extended to predict functional effects, pathogenicity and context-dependent relationships even without existing database entries. The paper shows that GFMs do not replace classical approaches, but can complement them in a targeted manner, and outlines future work on the integration and evaluation of different models in the SNP analysis process.
Johanna Pethke, Thomas Krause 0004, Bruno G. Andrade, Michael Kramer, Matthias L. Hemmje
BIBM5
2025 Responsible Use of AI in Genomics and Ethical Implications
abstract
This paper focuses on the aspects of responsibility and ethics when using AI systems in the contexts of health and genomics. The present proposal also aims to address the important need to consider how the use of AI systems in the field of genomics affects human life and its implications. As AI systems increasingly participate in diagnostic and predictive decision-making, they introduce new challenges concerning moral agency, transparency, and human oversight. The discussion aims to examine how these technologies, while offering unprecedented analytical capabilities, simultaneously reshape the ethical landscape of medical practice and genomic research.
Giulia Ricci, Paolo Buono, Thomas Krause 0004, Philippe Tamla, Matthias L. Hemmje, Francesca De Luzi, Francesco Leotta, Andrea Marrella, Flavia Monti, Massimo Mecella
BIBM5
2025 LLM-Driven Cloud-Based Infrastructure Design
abstract
Artificial intelligence has become integral to the life sciences, enabling large-scale data analysis and accelerating discovery. However, deploying these analytical capabilities depends on cloud infrastructures that are often complex to design, validate, and operate. Current approaches rely heavily on manual, time-consuming configuration by experts. This paper explores how artificial intelligence, specifically Large Language Models, can support not only genomic data analysis but also the automated creation of the underlying computational infrastructure. We propose a methodology that translates naturallanguage user stories and requirements into cloud-architectural specifications and containerized deployments, bridging the gap between user intent and executable infrastructure. Preliminary results indicate that LLM-driven synthesis can reduce human effort in defining container relationships and deployment logic, advancing automation in cloud-native bioinformatics.
Andrea Sepielli, Marco Calamo, Filippo Bianchini, Francesca De Luzi, Matteo Marinacci, Flavia Monti, Jacopo Rossi, Massimo Mecella, Philippe Tamla, Thomas Krause 0004, Matthias L. Hemmje
BIBM11
2025 The GenDAI Cloud-Native Infrastructure and Data Stewardship for Clinical Metagenomic Diagnostics
abstract
This paper presents a cloud-native architecture for clinical metagenomic diagnostics developed as part of the Horizon Europe project GenDAI. The architecture integrates automated, reproducible, and auditable workflows with deterministic elasticity, compliant data stewardship, explainable Artificial Intelligence (AI), and verifiable reporting. Requirements derived from clinical practice inform a unified modeling, implementation, and evaluation strategy for a modular platform that combines workflow orchestration, data governance, AI-powered modeling, and cloud-native reporting. The system embeds provenance-bydesign, policy-as-code enforcement, and deterministic elasticity across all layers to enable reproducible, compliant, and trustworthy metagenomic diagnostics. This work provides a principled pathway for translating research-grade tools into regulator-ready diagnostic services while maintaining transparency, reproducibility, and long-term trust.
Philippe Tamla, Thomas Krause 0004, Matthias L. Hemmje, Flavia Monti, Francesca De Luzi, Massimo Mecella, Bruno Andrade, Paolo Buono, Andrea Molinari
BIBM3
2024 256 Metaverse Records Dataset
Patrick Steinert, Stefan Wagenpfeil, Ingo Frommholz, Matthias L. Hemmje
ACM Multimedia4
2024 In-Database Feature Extraction to Improve Early Detection of Problematic Online Gambling Behavior
abstract
This study involves a comprehensive analysis of an anonymized dataset provided by a Swiss online casino that adds to the identification of reliable early indicators for problematic online gambling. Targeting gambling addiction prevention, our objective was to model and evaluate behavioral characteristics that signal early stages of problem gambling. We scrutinized player behaviors against a list of gamblers previously excluded for problematic gambling, using this as our target variable. Our approach combined traditional gambling risk indicators, as outlined in the existing literature, with innovative exploratory feature engineering and feature selection. This involved computing moving aggregates over specific periods to capture nuanced gambling patterns. All features were evaluated by assessing mutual information with the target variable as well as the collinearity of each pairwise combination of features. Based on our data analysis, we found that the total losses in the previous seven days, total deposits in the previous 15 days, total duration played in the previous seven days, stakes (amount bet per game) over the previous seven days, and making a deposit 12 h after a loss (chasing) were the most informative and independent risk indicators. To assess the accuracy of these indicators for early detection of problematic gambling and accordingly for responsible gambling interventions, we combined them in a linear regression model and compared its performance with the casino's currently used model. We found that a binary decision model based on a linear combination of these indicators provided better recall, greater precision, and more timely decisions than the benchmark.
Gabriel Stechschulte, Malte Wintner, Matthias L. Hemmje, Jürg Schwarz, Suzanne Lischer, Michael Kaufmann 0004
IEEE Trans. Comput. Soc. Syst.3
2022 Towards a Work Task Simulation Supporting Training of Work Design Skills during Qualification-based Learning
Ramona Srbecky, Michael Winterhagen, Benjamin Wallenborn, Matthias Then, Wieland Fraas, Jan Dettmers, Matthias L. Hemmje
CSEDU (2)8
2022 CIE: A Cloud-Based Information Extraction System for Named Entity Recognition in AWS, Azure, and Medical Domain
Philippe Tamla, Benedict Hartmann, Calvin Kramer, Florian Freund, Matthias L. Hemmje
IC3K6
2021 GenDAI - AI-Assisted Laboratory Diagnostics for Genomic Applications
abstract
Genomic applications like gene expression analysis or metagenomics are increasingly useful tools for laboratory diagnostics as they provide biomarkers for physiological and pathophysiological states. Artificial intelligence can help to analyze, visualize and interpret results obtained by genomic instruments. However, most software solutions for the analysis of genomic data are not suitable for laboratory use as they face several legal and technical challenges. Therefore, we propose a conceptual architecture named “GenDAI” that combines artificial intelligence and genomic applications while taking into account the challenges of laboratory diagnostics. GenDAI is based on two previous conceptual models, one for gene expression analysis and one for metagenomics. An upcoming prestudy will explore more detailed regulatory and technical requirements and gather use cases from laboratory practice, which will allow practical validation of the proposed solution in the future.
Thomas Krause 0004, Elena Jolkver, Sebastian Bruchhaus, Michael Kramer, Matthias L. Hemmje
BIBM5
2020 Big Data Analysis, AI, and Visualization Workshop: Road Mapping Infrastructures for Artificial Intelligence Supporting Advanced Visual Big Data Analysis
abstract
The overall scope and goal of the workshop is to bring together researchers active in the areas of Artificial Intelligence (AI), Big Data Analysis, and Visualization to achieve a road map, which can support the acceleration in research and data science activities by means of transforming, enriching, and deploying AI models and algorithms as well as intelligent advanced visual user interfaces supporting creation, configuration, management, and usage of distributed Big Data Analysis. Big Data Analysis and AI mutually support each other: AI-powered algorithms empower data scientists to analyze Big Data and thereby exploit its full potential whereas Big Data enables AI experts to comfortably design, validate, and deploy AI models. One of the workshop's objectives is the examination of the importance and necessity of a third, a more straightforward relationship of Big Data and AI: AI supporting all user stereotypes and organizations involved in Big Data Analysis on their exploration journey from raw input data to insight and effectuation.
Thoralf Reis, Marco X. Bornschlegl, Matthias L. Hemmje
AVI3
2020 Emerging Named Entity Recognition in a Medical Knowledge Management Ecosystem
Christian Nawroth, Felix Engel 0002, Matthias L. Hemmje
KEOD3
2020 Supporting Named Entity Recognition and Document Classification in a Knowledge Management System for Applied Gaming
Philippe Tamla, Florian Freund, Matthias L. Hemmje
KEOD3
2020 Integrating Smart Production Planning Into Smart Production
abstract
Industry 4.0 R&D is currently driving the emergence of Smart Production Environments (SPE) where the manufacturing of new products takes place very flexibly and dynamically in several steps and locations, possibly distributed all over the world. Companies involved in such smart productions act in highly competitive global markets and always have to find new ways to, e.g., cut costs by changing to the best providers, or enabling, e.g., tax savings, to stay competitive. This can only be achieved by using the fastest, most effective, efficient, and flexible distributed collaborative production processes. In this context, a Collaborative Adaptive Production Process Planning (CAPPP) can be supported by semantic product data management approaches enabling production-knowledge representation, utilization, curation, and archival as well as knowledge sharing, access, and reuse in many flexible and efficient ways. To support such scenarios, semantic representation of production knowledge integrated into a machine-readable process formalization is a key enabling factor for sharing such explicit knowledge resources in cloud-based production knowledge repositories. Our approach of Knowledge-Based Production Planning (KPP) introduces such a method and already provides a corresponding prototypical Proof-of-Concept implementation. In this way, it can not only be used in production planning but also, e.g., in Collaborative Manufacturing Change Management (CMCM) as well as, e.g., in Collaborative Assembly-, Logistics- and Layout Planning (CALLP) both building on the results of CAPPP. Therefore, a collaborative planning and optimization from mass to lot-size one production in a machine readable and processable representation will be possible. On the other hand, KPP knowledge can be shared to support other types of innovation, collaboration, and cocreation within a cloud-based semantic production knowledge repository.
Tobias Vogel 0003, Benjamin Gernhardt, Matthias L. Hemmje
NOMS3
2019 Emerging Named Entity Recognition on Retrieval Features in an Affective Computing Corpus
abstract
Affective Computing (AC) is a relatively new, dynamic and interdisciplinary research field. Numerous contributions from fields like computer science, psychology, cognitive science, sociology, physiology and medical science have been made. Consequently, it is difficult to track all recently published trends for early insight utilisation in practise or as basis for innovative research. Even if this fact holds true for many other research fields, AC in this respect is stimulating, due to its dynamic and interdisciplinary characteristics. However, Emergent Entities Recognition is a new concept introduced for early detection and prediction of developing professional terminology. Initial software developments have been completed and briefly analysed in general databases (e.g. MEDLINE). Here, we are interested in its evaluation for AC. In this respect, we have created and used a new Benchmark for Emergent Entities recognition specifially for the field of AC and show evaluation results in comparison to state of the art trained named entity recognition models and to a generic corpus (MEDLINE).
Christian Nawroth, Felix Engel 0002, Paul Mc Kevitt, Matthias L. Hemmje
BIBM4
2019 Is there an Optimal Technology to Provide Personal Supportive Feedback in Prevention of Obesity?
abstract
Obesity is a global challenge that affects health and wellbeing worldwide. In this position paper, we review the digital technology used in prevention of obesity and present the proposed STOP project that integrates state-of-the-art wearable technology, chatbot, gamification data fusion, and machine learning with the aim to provide personalised supportive feedback for preventing obesity and maintaining healthy weight. Implication of sensitive data with General Data Protection Regulation (GDPR) is discussed. We conclude that machine learning plays an important role in data fusion, analytics, and providing optimal messaging tailored design to support healthy weight.
Simone Sandri, Matthias L. Hemmje, Huiru Zheng, Felix Engel 0002, Anne Moorhead, Haiying Wang 0001, Raymond R. Bond, Michael F. McTear, Andrea Molinari, Paolo Bouquet
BIBM2
2019 Using an Affective Computing Taxonomy Management System to Support Data Management in Personality Traits
abstract
Affective Computing is a rather new and multidisciplinary research field that seeks sophisticated automation in emotion detection for later analysis. However, the automated emotion detection and analysis require as well comprehensive data management support, e.g. to keep control of data produced, and to enable its efficient reuse through classification with established terminology. This paper contributes to data management aspects in Affective Computing and to automation support in emotion classification on the basis of a personal traits analysis. Hence, we describe the implementation of a taxonomy management system, derived from requirements of a case study that investigates the relationship between personality and emotions in Affective Computing. The study makes use of machine learning software developed by SenseCare, an EU-funded R&D project that applies Affective Computing to enhance and advance future healthcare processes and systems.
Ryan Donovan, Michael Healy, Paul Mc Kevitt, Paul Walsh, Felix Engel 0002, Michael Fuchs 0002, Matthias L. Hemmje
BIBM8
2019 A metagenomic content and knowledge management ecosystem platform
abstract
The reduced cost of DNA sequencing allows metagenomics to be applied on a larger scale. With metagenomic analysis, we have better insight into supplement usage, methane production, and feed conversion efficiency in livestock systems. Nevertheless, sequencing machines generate an enormous amount of complex data. Conventional methods used in the analysis of genomic data involve pre-processing and synchronous reconstruction by multiple systems, which is time consuming and prone to failure. Furthermore, the sequencing datasets and analysis results need to be organized and stored properly in order for scientists to search and access them. To tackle these challenges, a new workflow for metagenomic analysis with improved infrastructure is needed. The MetaPlat project supports experts in both academic and non-academic sectors dealing with challenges in the field of metagenomics by focusing on improved hardware and software platforms. High-performance, fault-tolerant, flexible, and scalable processors and analysis systems will help to increase the effectiveness and efficiency of current metagenomics studies. In this paper, we propose such as an infrastructure applying emerging technologies, such as Kafka, Docker, and Hadoop. Details of the infrastructure solution and some preliminary results are also discussed.
Yanxin Wu, Haithem Afli, Paul Mc Kevitt, Paul Walsh, Felix Engel 0002, Michael Fuchs 0002, Matthias L. Hemmje
BIBM8
2019 Using Topic Specific Features for Argument Stance Recognition
abstract
Argument detection and its representation through ontologies are important parts of today’s attempt in automated recognition and processing of useful information in the vast amount of constantly produced data. However, due to the highly complex nature of an argument and its characteristics, its automated recognition is hard to implement. Given this overall challenge, as part of the objectives of the RecomRatio project, we are interested in the traceable, automated stance detection of arguments, to enable the construction of explainable pro/con argument ontologies. In our research, we design and evaluate an explainable machine learning based classifier, trained on two publicly available data sets. The evaluation results proved that explainable argument stance recognition is possible with up to .96 F1 when working within the same set of topics and .6 F1 when working with entirely different topics. This informed our hypothesis, that there are two sets of features in argument stance recognition: General features and topic specific features.
Tobias Eljasik-Swoboda, Felix Engel 0002, Matthias L. Hemmje
DATA3
2019 Supporting Taxonomy Development and Evolution by Means of Crowdsourcing
abstract
Information overload continues to be a challenge. By dividing the material into many different small subsets, classification based on a taxonomy makes data exploration and retrieval faster and more accurate. Instead of having to know the exact keywords that describe the knowledge resource, users can browse and search for them by selecting the categories that the resource is most likely to belong. Nevertheless, developing taxonomies is not an easy task. It requires the authors to have a certain amount of knowledge in the domain. Furthermore, the workload will increase as any new taxonomy needs to be frequently updated to remain relevant and useful. To combat these problems, this paper proposes another approach to crowdsource taxonomy development and evolution. We describe in this paper the concept of this approach along with different types of evaluations targeting on the one hand to demonstrate the feasibility of the approach and the usability of the initial prototype as well as on the other hand the quality and effectiveness of the chosen method.
Matthias L. Hemmje
KEOD2
2018 SenseCare: Using Automatic Emotional Analysis to Provide Effective Tools for Supporting
Ryan Donovan, Michael Healy, Huiru Zheng, Felix Engel 0002, Michael Fuchs 0002, Paul Walsh, Matthias L. Hemmje, Paul Mc Kevitt
BIBM8
2018 An Interface to Heterogeneous Data Sources Based on the Mediator/Wrapper Architecture in the Hadoop Ecosystem
Klaus-Dieter Schmatz, Kevin Berwind, Felix Engel 0002, Matthias L. Hemmje
BIBM4
2018 ImmunoAdept - bringing blood microbiome profiling to the clinical practice
Paul Walsh, Bruno G. Andrade, Cintia C. Palu, Brendan Lawlor, Brian Kelly, Matthias L. Hemmje, Michael Kramer
BIBM7
2018 No Target Function Classifier - Fast Unsupervised Text Categorization using Semantic Spaces
Tobias Eljasik-Swoboda, Michael Kaufmann 0004, Matthias L. Hemmje
DATA3
2018 Towards Enabling Emerging Named Entity Recognition as a Clinical Information and Argumentation Support
Christian Nawroth, Felix Engel 0002, Tobias Eljasik-Swoboda, Matthias L. Hemmje
DATA4
2018 Interfaces and services for integrating games and game components into competence based courses within learning management systems
abstract
Development of an educational game that can be integrated into a course within a Learning Management System (LMS) is a challenging task both from the didactic and from the technical point of view. A thorough evaluation of students' playing performance on LMS-side requires informative user-specific data that have to be collected during gameplay and then transferred to the LMS. Standardized interaction technologies like Learning Tools Interoperability (LTI) specify mechanisms that can be applied for implementing functionality to exchange students' competence profiles and traces. The Knowledge-Management Eco-System Portal of the European funded RAGE (Realizing an Applied Gaming Ecosystem) Research and Innovation Action project provides capable components, assets and related utility packs for developing such functionality: besides a toolkit for creating competence based games that meet the reaquirements of Qualifications Based Learning (QBL), assets for storing individual player profiles and traces are available. Furthermore, tools for creating game-specific analytics are provided that can be embedded by LTI-compatible LMSs like Moodle. This paper is concerned with the integration of QBL-compliant competence based games into LMS-courses. A major topic is the description of an exemplar Moodle-plugin with its required functionality and interfaces to the above-mentioned RAGE assets. As the approach is not yet implemented, this paper remains on a conceptual level.
Matthias Then, Benjamin Wallenborn, Alexander Nussbaumer, Michael Fuchs 0002, Simon Fuchs, Matthias L. Hemmje
EDUCON6
2017 Approach to semi-automatic labeling of video sequences for affective computing-enabling the comprehensive assessment of emotion detection software from mimics
abstract
In recent years, most breakthroughs in fields such as image and video processing were based on machine learning technologies that allow computers to recognize objects in images with nearly human precision. In some application domains, computers even surpassed human level performance. These breakthroughs result from an exponential increase of computational resources and digitization of society (massive availability of video material) as well as significant progress in the field of artificial intelligence, namely deep learning. Despite these advances, however, automatically recognizing emotions of humans in video material is still a challenging, but important task in affective computing. Without doubt, the automated detection of social-emotional and psychiatrically relevant information from video material will make a valuable contribution to various domains. This is particularly true for the field of personalized medicine. Major obstacles are high efforts and costs of labeling single video frames manually to create training data for machine learning technologies. To tackle this, we outline and discuss a methodological approach to create labeled video data with a minimum of human involvement.
Thilo Böhm, Felix Engel 0002, Danilo Bzdok, Matthias L. Hemmje
BIBM5
2017 The role of reproducibility in affective computing
abstract
The use of Affective Computing in the medical domain is gaining momentum, but is challenged through requirements arising through the inherent processing of personal sensitive data, that will effect comprehensive analysis reproducibility. Reproducibility is a key element in good research practice and a key ingredient to comprehensively validate AC applications in a medical context. Various research has been undertaken to support reproducible analysis procedures through the establishment of a conceptual basis (definition and modeling) and by means of technology support. However, its realization is generally hardly achievable. Therefore, this workshop contribution will elaborate and document on reproducibility aspects related to Affective Computing in the medical domain, as we face it in the course of the EC co-funded SenseCare project. This contribution is meant as a starting point for further discussions and further reproducibility related research in AC.
Felix Engel 0002, Alphonsus Keary, Kevin Berwind, Marco X. Bornschlegl, Matthias L. Hemmje
BIBM5
2017 Modeling and Qualitative Evaluation of a Management Canvas for Big Data Applications
Michael Kaufmann 0004, Tobias Eljasik-Swoboda, Christian Nawroth, Kevin Berwind, Marco X. Bornschlegl, Matthias L. Hemmje
DATA6
2016 Road Mapping Infrastructures for Advanced Visual Interfaces Supporting Big Data Applications in Virtual Research Environments
abstract
Handling the complexity of relevant data requires new techniques about data access, visualization, perception, and interaction for innovative and successful strategies. In order to address human-computer interaction, cognitive eficiency, and interoperability problems, a generic information visualization, user empowerment, as well as service integration and mediation approach based on the existing state-of-the-art in the relevant areas of computer science has to be achieved.
Marco X. Bornschlegl, Andrea Manieri, Paul Walsh, Tiziana Catarci, Matthias L. Hemmje
AVI5
2016 Analysis of rumen microbial community in cattle through the integration of metagenomic and network-based approaches
abstract
A better understanding of the composition of rumen microbial communities and the association between host genetic and microbial activities has important applications and implication in bioscience. Being capable of revealing the full extent of microbial gene diversity, metagenomics-based approaches hold great promises in this endeavor. This study investigates the rumen microbial community in cattle through the integration of metagenomic and network-based approaches. Based on the relative abundance of 1570 microbial genes identified in a metagenomics analysis, the co-abundance network was constructed and functional modules of microbial genes were identified. One of the main contributions of this study is to develop a random matrix theory-based approach to automatically determine the correlation threshold used to construct the co-abundance network. It has been shown that the network exhibits a highly modular structure with each of the three main modules well separated. The involvement of KEGG pathways in each module was analysed. A close look at the abundance profiles highlights that Module B is strongly associated with methane emissions while Module C is highly enriched with microbial genes associated with feed conversion efficiency.
Haiying Wang 0001, Huiru Zheng, Fiona Browne, Rainer Roehe, Richard J. Dewhurst, Felix Engel 0002, Matthias L. Hemmje, Paul Walsh
BIBM7
2016 Toward Cloud-based Classification and Annotation Support
abstract
Manually annotating content-based categories to existing documents is a time-consuming task for human domain experts. In order to ease this effort, automated text categorization is used. This paper evaluates the state of the art in cloud-based text categorization and proposes an architecture for flexible cloud-based classification and annotation support, leveraging the advantages provided by cloud-based architectures.
Tobias Eljasik-Swoboda, Michael Kaufmann 0004, Matthias L. Hemmje
CLOSER (2)3
2016 EDISON Data Science Framework: A Foundation for Building Data Science Profession for Research and Industry
abstract
Data Science is an emerging field of science, which requires a multi-disciplinary approach and should be built with a strong link to emerging Big Data and data driven technologies, and consequently needs re-thinking and re-design of both traditional educational models and existing courses. The education and training of Data Scientists currently lacks a commonly accepted, harmonized instructional model that reflects by design the whole lifecycle of data handling in modern, data driven research and the digital economy. This paper presents the EDISON Data Science Framework (EDSF) that is intended to create a foundation for the Data Science profession definition. The EDSF includes the following core components: Data Science Competence Framework (CF-DS), Data Science Body of Knowledge (DS-BoK), Data Science Model Curriculum (MC-DS), and Data Science Professional profiles (DSP profiles). The MC-DS is built based on CF-DS and DS-BoK, where Learning Outcomes are defined based on CF-DS competences and Learning Units are mapped to Knowledge Units in DS-BoK. In its own turn, Learning Units are defined based on the ACM Classification of Computer Science (CCS2012) and reflect typical courses naming used by universities in their current programmes. The paper provides example how the proposed EDSF can be used for designing effective Data Science curricula and reports the experience of implementing EDSF by the Champion Universities that cooperate with the EDISON project.
Yuri Demchenko, Adam Belloum, Wouter Los, Tomasz Wiktorski, Andrea Manieri, Holger Brocks, Jana Becker, Dominic Heutelbeck, Matthias L. Hemmje, Steve Brewer
CloudCom9
2016 Combining Taxonomies using Word2vec
abstract
Taxonomies have gained a broad usage in a variety of fields due to their extensibility, as well as their use for classification and knowledge organization. Of particular interest is the digital document management domain in which their hierarchical structure can be effectively employed in order to organize documents into content-specific categories. Common or standard taxonomies (e.g., the ACM Computing Classification System) contain concepts that are too general for conceptualizing specific knowledge domains. In this paper we introduce a novel automated approach that combines sub-trees from general taxonomies with specialized seed taxonomies by using specific Natural Language Processing techniques. We provide an extensible and generalizable model for combining taxonomies in the practical context of two very large European research projects. Because the manual combination of taxonomies by domain experts is a highly time consuming task, our model measures the semantic relatedness between concept labels in CBOW or skip-gram Word2vec vector spaces. A preliminary quantitative evaluation of the resulting taxonomies is performed after applying a greedy algorithm with incremental thresholds used for matching and combining topic labels.
Tobias Eljasik-Swoboda, Matthias L. Hemmje, Mihai Dascalu, Stefan Trausan-Matu
DocEng2
2016 Recalot.com: Towards a Reusable, Modular, and RESTFul Social Recommender System
Matthäus Schmedding, Michael Fuchs 0002, Claus-Peter Klas, Felix Engel 0002, Holger Brocks, Dominic Heutelbeck, Matthias L. Hemmje
ICSR7
2016 Intuitive Knowledge Connectivity: Design and Prototyping of Cross-Platform Knowledge Networks
Michael Kaufmann 0004, Andreas Waldis, Patrick Siegfried, Gwendolin Wilke, Edy Portmann, Matthias L. Hemmje
KSEM6
2016 Towards a probabilistic model for supporting collaborative information access
Thilo Böhm, Claus-Peter Klas, Matthias L. Hemmje
Inf. Retr. J.3
2015 Data Science Professional Uncovered: How the EDISON Project will Contribute to a Widely Accepted Profile for Data Scientists
abstract
The digital revolution made available vast amounts of data both in industry and in the research landscape. The ability to manipulate and extract knowledge and value from this data represents a new profession called the Data Scientist: expected to be the most visible job in future years. The EDISON project has been established in order to support universities, research centers, industry and research infrastructure organisations to cope with the potential shortfall of Data Scientists, to define the framework of competences as well as the body of knowledge for this profession. In this paper the EDISON team describes how it intends to nurture the profession of Data Scientist to cope with the expected increase in demand. The strategy proposed is based on both the analysis of the demand side (industries, research centers and research infrastructure organisations) and the supply side (Universities and training centers) bridging between the providers and employers by cooperating on the establishment of a Competence Framework and a Body of Knowledge for the Data Scientist Professional. The project will exploit piloting initiatives in cooperation with pioneer universities and also involve external experts as evangelists.
Andrea Manieri, Steve Brewer, Ruben Riestra, Yuri Demchenko, Matthias L. Hemmje, Tomasz Wiktorski, Tiziana Ferrari, Jérémy Frey
CloudCom5
2015 Adaptive Information Retrieval Support for Multi-session Information Tasks
Daniel T. J. Backhausen, Claus-Peter Klas, Matthias L. Hemmje
TPDL3
2015 An Experimental Evaluation of Collaborative Search Result Division Strategies
Thilo Böhm, Claus-Peter Klas, Matthias L. Hemmje
TPDL3
2013 e-Infrastructures for Digital Libraries...the Future
Wim Jansen, Roberto Barbera, Michel Drescher, Antonella Fresa, Matthias L. Hemmje, Yannis E. Ioannidis, Norbert Meyer, Nick Poole, Peter Stanchev
TPDL5
2012 Supporting user of information retrieval systems by visualizing the information dialog context
abstract
In this paper we present first results for a new approach designing innovative user interfaces for information retrieval systems. The leading thought of this paper is based on the fact that a dialog between user and system during a search process establishes an information dialog context. We introduce a framework for information retrieval systems to handle the activities and sets elaborated within a search process in order to support the user. In this paper we present a prototype tool which provides and visualizes sets of information objects in order to bring users into the position to keep the developed sets within the search process at hand and actively manage them. Finally, a description of a user study and expert interviews and their evaluation results conducted on the basis of the prototypical tool is provided.
Paul Landwich, Claus-Peter Klas, Matthias L. Hemmje
AVI3
2012 Aesthetic visualisation of information: optimization of graph representations
abstract
In this work we will discuss the optimization of aesthetic criteria for graph representation. Most existing algorithms for drawing graphs generate a specific layout and are usually specialized in optimizing certain aesthetic criteria. Generated representations visualize data for specific applications. After all, it could be important to have an individual layout for the application. The communication objective should determine the representation. The individual representation should, however, also be optimized to be cognitively efficient. We suggest a clear distinction between the graph layout and its optimization. A useful field of application could be the visualization of search results for supporting information retrieval.
Dieter Meiller, Matthias L. Hemmje, Claus-Peter Klas
AVI2
2009 Architecture to Connect Tool-based Web Interfaces to Service-oriented Architectures
Claus-Peter Klas, Martin Mois, Matthias L. Hemmje
WEBIST3
2006 Distributed Leader Election in P2P Systems for Dynamic Sets
abstract
The collection of and search for location information is a core component in many pervasive and mobile computing applications. In distributed collaboration scenarios this location data is collected by different entities, e.g., users with GPS enabled mobile phones. Instead of using a centralized service for managing this distributed dynamic location data, we use a peer-to-peer data structure, the so-called distributed space partitioning tree (DSPT). A DSPT is a general use peer-to-peer data structure, similar to distributed hash tables (DHTs), that allows publishing, updating of, and searching for dynamic sets. In this paper we present an efficient distributed leader election algorithm that can be used in DSPTs to eliminate redundant network traffic.
Dominic Heutelbeck, Matthias L. Hemmje
MDM2
2006 RectNet - A Distributed Geometrical Data Structure
abstract
The collection of and search for location information is a core component in many pervasive and mobile computing applications. Instead of using a centralized service for managing distributed dynamic location data, we previously introduced the concept of a distributed data structure, the socalled distributed space partitioning tree (DSPT). A DSPT is a general use distributed data structure, similar to distributed hash tables (DHTs), that allows publishing, updating of, and searching for geometrical objects. The problem of range queries on a set of points in a distrubuted scenatio space has been well studied. DSPTs generalize this problem, by allowing the keys of the objects and queries to have a spatial extension with arbitrary boundaries. In this paper we describe RectNet, a first implementation of a DSPT. RectNet is based on an binary space partitioning and torus topology. We provide an overview of RectNets architecture, algorithms, and a brief evaluation.
Dominic Heutelbeck, Matthias L. Hemmje
MDM2
2006 A Peer-to-Peer Data Structure for Dynamic Location Data
abstract
The collection of and search for location information is a core component in many pervasive and mobile computing applications. In distributed collaboration scenarios this location data is collected by different entities, e.g., users with GPS enabled mobile phones. Instead of using a centralized service for managing this distributed dynamic location data, we present a new peer-to-peer data structure, the so-called distributed space partitioning tree (DSPT). A DSPT is a general use peer-to-peer data structure, similar to distributed hash tables (DHTs), that allows publishing, updating of, and searching for geometrical objects.
Dominic Heutelbeck, Matthias L. Hemmje
PerCom2
2003 Media and metadata management for capture and access systems in electronic lecturing environments
abstract
One of the main challenges facing Web-based multimedia content creators today is the development of cost-effective digital media content for courseware that is reusable and interoperable. Given this, we propose the design and implementation of a model to manage the multimedia and metadata within a system of related topics in a course that is delivered via electronic lectures. This model not only supports the re-purposing and interchange of a course's digital media content, but also preserves the semantic dependencies and associated media-mapping, of pre-requisite and co-requisite knowledge within a course. Coarse grain or topic-based segmentations of an electronic lecture are used as building blocks to automatically create new, media-rich, electronic lecture experiences of interest to a user and subject to the constraints imposed by the model.
Avare Stewart, Patrick Wolf, Matthias L. Hemmje
ICME3
2003 Supporting Model-Based Construction of Semantic-Enabled Web Applications
abstract
Semantic annotation of Web content is in the core of the current semantic Web activity. The operationalization of the semantic Web raises the challenge on how to systematically integrate semantic annotations into generated Web application pages. In this paper, we present VizCo, a tool for systematically integrating RDF based semantically enriched application domain models into the process of setting up dynamically generated Web application user interfaces and supporting their evolution. The form-based Web pages are annotated based on the underlying domain model. Combined with further mapping tools, VizCo is used in our Web application development framework for coupling the domain model views and, indirectly, the underlying application data with other Web application components, especially with elements of the user interface. In contrast to a direct coupling with the application data an additional semantic layer is introduced. We follow a pragmatic approach that utilizes domain information extracted from application schemata and data as a starting point. Both the domain model and its coupling to the data source can be manually refined and extended. The couplings are dynamically translated into bi-directional data bindings at runtime.
Michael Fuchs 0002, Claudia Niederée, Matthias L. Hemmje, Erich J. Neuhold
WISE3
2002 Analysing data trough visualizations in a web-based trade fair system
abstract
The enormous amount of data available on the web must be adequately exploited by company managers in order to improve their business. An important role is played by various techniques that are capable of extracting useful information. Our approach exploits visualization techniques. In this paper we present 2D and 3D visualizations that are used in a web-based system that supports the organization and management of trade fairs. We show how the main users of the system, namely fair organisers and companies that participate in the fair either as exhibitors or visitors, take advantage of the available visual tools in their business activities.
Paolo Buono, Maria Francesca Costabile, Gerald Jaeschke, Matthias L. Hemmje
SEKE4
2001 Distributed Information Search with Adaptive Meta-Search Engines
Lieming Huang, Ulrich Thiel, Matthias L. Hemmje, Erich J. Neuhold
CAiSE3
2001 Adaptively constructing the query interface for meta-search engines
abstract
With the exponential growth of information on the Internet, current information integration systems have become more and more unsuitable for this "Internet age" due to the great diversity among sources. This paper presents a constraint-based query user interface model, which can be applied to the construction of dynamically generated adaptive user interfaces for meta-search engines.
Lieming Huang, Ulrich Thiel, Matthias L. Hemmje, Erich J. Neuhold
IUI3
2001 Adaptive Web Meta-search Enhanced by Constraint-Based Query Constructing and Mapping
Lieming Huang, Ulrich Thiel, Matthias L. Hemmje, Erich J. Neuhold
WAIM3
2001 Using Link Types in Web Page Ranking and Filtering
abstract
Corresponding to the evolution of the Web from the poorly structured towards more structured, semantic rich network, search methods that apply to it are also evolving from using little structural and semantic information to using more such information that is available. This paper proposes a ranking and filtering mechanism that makes use of link types that is representable with new Web standards. We suggest that page ranking can be propagational through links and the propagational rates depend on the types of the links and users' specific set of interests. Page filtering can be decided based on link types combined with some other information relevant to links. For either a ranking or filtering task, a profile containing a set of ranking or filtering rules to be followed in the task can be specified to reflect users' specific interests. Technical issues in implementing the mechanism in a search system are also discussed.
Zhanzi Qiu, Matthias L. Hemmje, Erich J. Neuhold
WISE (1)2
2000 ConSearch: Using hypertext contexts as web search boundaries
Zhanzi Qiu, Matthias L. Hemmje, Erich J. Neuhold
CoopIS2
2000 ADMIRE: an adaptive data model for meta search engines
Lieming Huang, Matthias L. Hemmje, Erich J. Neuhold
Comput. Networks2
1999 Conference Notes - 1996: Foundations of Advanced Information Visualization for Visual Information (Retrieval) Systems
Mark E. Rorvig, Matthias L. Hemmje
J. Am. Soc. Inf. Sci.2
1996 Foundations of Advanced Information Visualization for Information Retrieval Szstems (Workshop Abstract)
abstract
No abstract available.
Mark E. Rorvig, Matthias L. Hemmje
SIGIR2
1994 LyberWorld - A Visualization User Interface Supporting Fulltext Retrieval
Matthias L. Hemmje, Clemens Kunkel, Alexander Willett
SIGIR1
1993 A Dynamic Gesture Language and Graphical Feedback for Interaction in a 3D User Interface
abstract
Abstract In user interfaces of modern systems, users get the impression of directly interacting with application objects. In 3D based user interfaces, novel input devices, like hand and force input devices, are being introduced. They aim at providing natural ways of interaction. The use of a hand input device allows the recognition of static poses and dynamic gestures performed by a user's hand. This paper describes the use of a hand input device for interacting with a 3D graphical application. A dynamic gesture language, which allows users to teach some hand gestures, is presented. Furthermore, a user interface integrating the recognition of these gestures and providing feedback for them, is introduced. Particular attention has been spent on implementing a tool for easy specification of dynamic gestures, and on strategies for providing graphical feedback to users' interactions. To demonstrate that the introduced 3D user interface features, and the way the system presents graphical feedback, are not restricted to a hand input device, a force input device has also been integrated into the user interface.
Monica Bordegoni, Matthias L. Hemmje
Comput. Graph. Forum2