VLDB 2026 Research / reviewers in the wild / expert
Alexander Löser
dblp:36/979
· DBLP profile ↗
35ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-4440-3261ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 12 since 2021Databases, data management, data science and information retrieval · 16 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Theory of computation · 2 · 2 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
Bogdan Kostic, Conor Fallon, Julian Risch, Alexander Löser |
LREC | 4 |
| 2026 | DeepICD-R1: Medical Reasoning through Hierarchical Rewards and Unsupervised Distillation
Tom Röhr, Thomas Steffek, Roman Teucher, Keno Bressem, Alexei Figueroa Rosero, Paul Grundmann, Peter Tröger, Felix A. Gers, Alexander Löser |
LREC | 9 |
| 2025 | Computational Models for Patient Stratification in Urologic Cancers - Creating Robust and Trustworthy Multimodal AI for Health CareabstractCurrent clinical approaches fail to fully utilize unstructured data in managing prostate cancer (PCa) and kidney cancer (KC), leading to inefficiencies in patient care and increased costs. Effective diagnostics and treatments depend on integrating multimodal data, yet progress is hampered by limited data accessibility and a lack of collaborative validation between clinicians and computer scientists. To address these challenges, the EU-funded COMFORT project aims to develop commercially viable, data-driven multimodal decision support systems. These systems will improve clinical prognostication, patient stratification, and personalized treatment while also assessing the trust that healthcare professionals and patients place in AI-driven tools. Antonis Billis, Alessa Hering, Cora Meyer, Lisa Adams, Renato Cuocolo, Alexander Löser, Jawed Nawabi, Keno Bressem |
CBMS | 7 |
| 2025 | "Where does it hurt?" - Dataset and Study on Physician Intent Trajectories in Doctor Patient DialoguesabstractIn a doctor-patient dialogue, the primary objective of physicians is to diagnose patients and propose a treatment plan. Medical doctors guide these conversations through targeted questioning to efficiently gather the information required to provide the best possible outcomes for patients. To the best of our knowledge, this is the first work that studies physician intent trajectories in doctor-patient dialogues. We use the ‘Ambient Clinical Intelligence Benchmark’ (Aci-bench) dataset for our study. We collaborate with medical professionals to develop a fine-grained taxonomy of physician intents based on the SOAP framework (Subjective, Objective, Assessment, and Plan). We then conduct a large-scale annotation effort to label over 5000 doctor-patient turns with the help of a large number of medical experts recruited using Prolific, a popular crowd-sourcing platform. This large labeled dataset is an important resource contribution that we use for benchmarking the state-of-the-art generative and encoder models for medical intent classification tasks. Our findings show that our models understand the general structure of medical dialogues with high accuracy, but often fail to identify transitions between SOAP categories. We also report for the first time common trajectories in medical dialogue structures that provide valuable insights for designing ‘differential diagnosis’ systems. Finally, we extensively study the impact of intent filtering for medical dialogue summarization and observe a significant boost in performance. We make the codes and data, including annotation guidelines, publicly available at https://github.com/DATEXIS/medical-intent-classification. Tom Röhr, Soumyadeep Roy, Fares Al Mohamad, Jens-Michalis Papaioannou, Wolfgang Nejdl, Felix A. Gers, Alexander Löser |
ECAI | 7 |
| 2024 | Data Drift in Clinical Outcome Prediction from Admission NotesabstractClinical NLP research faces a scarcity of publicly available datasets due to privacy concerns. MIMIC-III marked a significant milestone, enabling substantial progress, and now, with MIMIC-IV, the dataset has expanded significantly, offering a broader scope. In this paper, we focus on the task of predicting clinical outcomes from clinical text. This is crucial in modern healthcare, aiding in preventive care, differential diagnosis, and capacity planning. We introduce a novel clinical outcome prediction dataset derived from MIMIC-IV. Furthermore, we provide initial insights into the performance of models trained on MIMIC-III when applied to our new dataset, with specific attention to potential data drift. We investigate challenges tied to evolving documentation standards and changing codes in the International Classification of Diseases (ICD) taxonomy, such as the transition from ICD-9 to ICD-10. We also explore variations in clinical text across different hospital wards. Our study aims to probe the robustness and generalization of clinical outcome prediction models, contributing to the ongoing advancement of clinical NLP in healthcare. Paul Grundmann, Jens-Michalis Papaioannou, Tom Oberhauser, Thomas Steffek, Amy Siu, Wolfgang Nejdl, Alexander Löser |
LREC/COLING | 7 |
| 2024 | DDxGym: Online Transformer Policies in a Knowledge Graph Based Natural Language EnvironmentabstractDifferential diagnosis (DDx) is vital for physicians and challenging due to the existence of numerous diseases and their complex symptoms. Model training for this task is generally hindered by limited data access due to privacy concerns. To address this, we present DDxGym, a specialized OpenAI Gym environment for clinical differential diagnosis. DDxGym formulates DDx as a natural-language-based reinforcement learning (RL) problem, where agents emulate medical professionals, selecting examinations and treatments for patients with randomly sampled diseases. This RL environment utilizes data labeled from online resources, evaluated by medical professionals for accuracy. Transformers, while effective for encoding text in DDxGym, are unstable in online RL. For that reason we propose a novel training method using an auxiliary masked language modeling objective for policy optimization, resulting in model stabilization and significant performance improvement over strong baselines. Following this approach, our agent effectively navigates large action spaces and identifies universally applicable actions. All data, environment details, and implementation, including experiment reproduction code, are made publicly available. Benjamin Winter, Alexei Figueroa Rosero, Alexander Löser, Felix A. Gers, Nancy Katerina Figueroa Rosero, Ralf Krestel |
LREC/COLING | 3 |
| 2024 | Boosting Long-Tail Data Classification with Sparse Prototypical Networks
Alexei Figueroa Rosero, Jens-Michalis Papaioannou, Conor Fallon, Alexandra Bekiaridou, Keno Bressem, Stavros Zanos, Felix A. Gers, Wolfgang Nejdl, Alexander Löser |
ECML/PKDD (7) | 9 |
| 2024 | medBERT.de: A comprehensive German BERT model for the medical domainabstractThis paper presents medBERT.de , a pre-trained German BERT model specifically designed for the German medical domain. The model has been trained on a large corpus of 4.7 Million German medical documents and has been shown to achieve new state-of-the-art performance on eight different medical benchmarks covering a wide range of disciplines and medical document types. In addition to evaluating the overall performance of the model, this paper also conducts a more in-depth analysis of its capabilities. We investigate the impact of data deduplication on the model's performance, as well as the potential benefits of using more efficient tokenization methods. Our results indicate that domain-specific models such as medBERT.de are particularly useful for longer texts, and that deduplication of training data does not necessarily lead to improved performance. Furthermore, we found that efficient tokenization plays only a minor role in improving model performance, and attribute most of the improved performance to the large amount of training data. To encourage further research, the pre-trained model weights and new benchmarks based on radiological data are made publicly available for use by the scientific community. Keno Bressem, Jens-Michalis Papaioannou, Paul Grundmann, Florian Borchert, Lisa Adams, Leonhard Liu, Felix Busch, Jan P. Loyen, Stefan Markus Niehues, Moritz Augustin, Lennart Grosser, Marcus R. Makowski, Hugo J. W. L. Aerts, Alexander Löser |
Expert Syst. Appl. | 15 |
| 2023 | From symmetry to asymmetry: Generalizing TSP approximations by parametrization
Lukas Behrendt, Katrin Casel, Tobias Friedrich 0001, Gregor Lagodzinski, Alexander Löser, Marcus Wilhelm |
J. Comput. Syst. Sci. | 5 |
| 2022 | Attention Networks for Augmenting Clinical Text with Support Sets for Diagnosis PredictionabstractDiagnosis prediction on admission notes is a core clinical task. However, these notes may incompletely describe the patient. Also, clinical language models may suffer from idiosyncratic language or imbalanced vocabulary for describing diseases or symptoms. We tackle the task of diagnosis prediction, which consists of predicting future patient diagnoses from clinical texts at the time of admission. We improve the performance on this task by introducing an additional signal from support sets of diagnostic codes from prior admissions or as they emerge during differential diagnosis. To enhance the robustness of diagnosis prediction methods, we propose to augment clinical text with potentially complementary set data from diagnosis codes from previous patient visits or from codes that emerge from the current admission as they become available through diagnostics. We discuss novel attention network architectures and augmentation strategies to solve this problem. Our experiments reveal that support sets improve the performance drastically to predict less common diagnosis codes. Our approach clearly outperforms the previous state-of-the-art PubMedBERT baseline by up 3% points. Furthermore, we find that support sets drastically improve the performance for pregnancy- and gynecology-related diagnoses up to 32.9% points compared to the baseline. Paul Grundmann, Tom Oberhauser, Felix A. Gers, Alexander Löser |
COLING | 4 |
| 2022 | Cross-Lingual Knowledge Transfer for Clinical PhenotypingabstractClinical phenotyping enables the automatic extraction of clinical conditions from patient records, which can be beneficial to doctors and clinics worldwide. However, current state-of-the-art models are mostly applicable to clinical notes written in English. We therefore investigate cross-lingual knowledge transfer strategies to execute this task for clinics that do not use the English language and have a small amount of in-domain data available. Our results reveal two strategies that outperform the state-of-the-art: Translation-based methods in combination with domain-specific encoders and cross-lingual encoders plus adapters. We find that these strategies perform especially well for classifying rare phenotypes and we advise on which method to prefer in which situation. Our results show that using multilingual data overall improves clinical phenotyping models and can compensate for data sparseness. Jens-Michalis Papaioannou, Paul Grundmann, Betty van Aken, Athanasios Samaras, Ilias Kyparissidis, George Giannakoulas, Felix A. Gers, Alexander Löser |
LREC | 8 |
| 2022 | KIMERA: Injecting Domain Knowledge into Vacant Transformer HeadsabstractTraining transformer language models requires vast amounts of text and computational resources. This drastically limits the usage of these models in niche domains for which they are not optimized, or where domain-specific training data is scarce. We focus here on the clinical domain because of its limited access to training data in common tasks, while structured ontological data is often readily available. Recent observations in model compression of transformer models show optimization potential in improving the representation capacity of attention heads. We propose KIMERA (Knowledge Injection via Mask Enforced Retraining of Attention) for detecting, retraining and instilling attention heads with complementary structured domain knowledge. Our novel multi-task training scheme effectively identifies and targets individual attention heads that are least useful for a given downstream task and optimizes their representation with information from structured data. KIMERA generalizes well, thereby building the basis for an efficient fine-tuning. KIMERA achieves significant performance boosts on seven datasets in the medical domain in Information Retrieval and Clinical Outcome Prediction settings. We apply KIMERA to BERT-base to evaluate the extent of the domain transfer and also improve on the already strong results of BioBERT in the clinical domain. Benjamin Winter, Alexei Figueroa Rosero, Alexander Löser, Felix A. Gers, Amy Siu |
LREC | 3 |
| 2021 | Clinical Outcome Prediction from Admission Notes using Self-Supervised Knowledge IntegrationabstractBetty van Aken, Jens-Michalis Papaioannou, Manuel Mayrdorfer, Klemens Budde, Felix Gers, Alexander Loeser. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Betty van Aken, Jens-Michalis Papaioannou, Manuel Mayrdorfer, Klemens Budde, Felix A. Gers, Alexander Löser |
EACL | 6 |
| 2021 | Aspect-Based Passage Retrieval with Contextualized Discourse Vectors
Jens-Michalis Papaioannou, Manuel Mayrdorfer, Sebastian Arnold 0001, Felix A. Gers, Klemens Budde, Alexander Löser |
ECIR (2) | 6 |
| 2021 | From Symmetry to Asymmetry: Generalizing TSP Approximations by Parametrization
Lukas Behrendt, Katrin Casel, Tobias Friedrich 0001, Gregor Lagodzinski, Alexander Löser, Marcus Wilhelm |
FCT | 5 |
| 2020 | Discovering Biased News Articles Leveraging Multiple Human AnnotationsabstractUnbiased and fair reporting is an integral part of ethical journalism. Yet, political propaganda and one-sided views can be found in the news and can cause distrust in media. Both accidental and deliberate political bias affect the readers and shape their views. We contribute to a trustworthy media ecosystem by automatically identifying politically biased news articles. We introduce novel corpora annotated by two communities, i.e., domain experts and crowd workers, and we also consider automatic article labels inferred by the newspapers’ ideologies. Our goal is to compare domain experts to crowd workers and also to prove that media bias can be detected automatically. We classify news articles with a neural network and we also improve our performance in a self-supervised manner. Konstantina Lazaridou, Alexander Löser, Maria Mestre, Felix Naumann |
LREC | 2 |
| 2020 | Is Language Modeling Enough? Evaluating Effective Embedding CombinationsabstractUniversal embeddings, such as BERT or ELMo, are useful for a broad set of natural language processing tasks like text classification or sentiment analysis. Moreover, specialized embeddings also exist for tasks like topic modeling or named entity disambiguation. We study if we can complement these universal embeddings with specialized embeddings. We conduct an in-depth evaluation of nine well known natural language understanding tasks with SentEval. Also, we extend SentEval with two additional tasks to the medical domain. We present PubMedSection, a novel topic classification dataset focussed on the biomedical domain. Our comprehensive analysis covers 11 tasks and combinations of six embeddings. We report that combined embeddings outperform state of the art universal embeddings without any embedding fine-tuning. We observe that adding topic model based embeddings helps for most tasks and that differing pre-training tasks encode complementary features. Moreover, we present new state of the art results on the MPQA and SUBJ tasks in SentEval. Rudolf Schneider 0001, Tom Oberhauser, Paul Grundmann, Felix A. Gers, Alexander Löser, Steffen Staab |
LREC | 5 |
| 2020 | Learning Contextualized Document Representations for Healthcare Answer RetrievalabstractWe present Contextual Discourse Vectors (CDV), a distributed document representation for efficient answer retrieval from long healthcare documents. Our approach is based on structured query tuples of entities and aspects from free text and medical taxonomies. Our model leverages a dual encoder architecture with hierarchical LSTM layers and multi-task training to encode the position of clinical entities and aspects alongside the document discourse. We use our continuous representations to resolve queries with short latency using approximate nearest neighbor search on sentence level. We apply the CDV model for retrieving coherent answer passages from nine English public health resources from the Web, addressing both patients and medical professionals. Because there is no end-to-end training data available for all application scenarios, we train our model with self-supervised data from Wikipedia. We show that our generalized model significantly outperforms several state-of-the-art baselines for healthcare passage ranking and is able to adapt to heterogeneous domains without additional fine-tuning. Sebastian Arnold 0001, Betty van Aken, Paul Grundmann, Felix A. Gers, Alexander Löser |
WWW | 5 |
| 2019 | How Does BERT Answer Questions?: A Layer-Wise Analysis of Transformer RepresentationsabstractBidirectional Encoder Representations from Transformers (BERT) reach state-of-the-art results in a variety of Natural Language Processing tasks. However, understanding of their internal functioning is still insufficient and unsatisfactory. In order to better understand BERT and other Transformer-based models, we present a layer-wise analysis of BERT's hidden states. Unlike previous research, which mainly focuses on explaining Transformer models by their attention weights, we argue that hidden states contain equally valuable information. Specifically, our analysis focuses on models fine-tuned on the task of Question Answering (QA) as an example of a complex downstream task. We inspect how QA models transform token vectors in order to find the correct answer. To this end, we apply a set of general and QA-specific probing tasks that reveal the information stored in each representation layer. Our qualitative analysis of hidden state visualizations provides additional insights into BERT's reasoning process. Our results show that the transformations within BERT go through phases that are related to traditional pipeline tasks. The system can therefore implicitly incorporate task-specific information into its token representations. Furthermore, our analysis reveals that fine-tuning has little impact on the models' semantic abilities and that prediction errors can be recognized in the vector representations of even early layers. Betty van Aken, Benjamin Winter, Alexander Löser, Felix A. Gers |
CIKM | 3 |
| 2019 | SECTOR: A Neural Model for Coherent Topic Segmentation and ClassificationabstractWhen searching for information, a human reader first glances over a document, spots relevant sections, and then focuses on a few sentences for resolving her intention. However, the high variance of document structure complicates the identification of the salient topic of a given section at a glance. To tackle this challenge, we present SECTOR, a model to support machine reading systems by segmenting documents into coherent sections and assigning topic labels to each section. Our deep neural network architecture learns a latent topic embedding over the course of a document. This can be leveraged to classify local topics from plain text and segment a document at topic shifts. In addition, we contribute WikiSection, a publicly available data set with 242k labeled sections in English and German from two distinct domains: diseases and cities. From our extensive evaluation of 20 architectures, we report a highest score of 71.6% F1 for the segmentation and classification of 30 topics from the English city domain, scored by our SECTOR long short-term memory model with Bloom filter embeddings and bidirectional segmentation. This is a significant improvement of 29.5 points F1 over state-of-the-art CNN classifiers with baseline segmentation. Sebastian Arnold 0001, Rudolf Schneider 0001, Philippe Cudré-Mauroux, Felix A. Gers, Alexander Löser |
Trans. Assoc. Comput. Linguistics | 5 |
| 2015 | Resolving Common Analytical Tasks in Text DatabasesabstractWith the convergence of data warehousing, online analytical processing and the Semantic Web, analytical tasks are no longer only designed and executed by experts. Instead, various users expect to query keyword search engines with analytical intentions. One efficient approach to answer these tasks is to leverage the factual information stored in large-scale text databases. These systems enable analysts to access unstructured text sources from the Web with structured query languages. The challenge of mapping keyword queries to structured queries has been approached in various forms. However, these systems are not able to detect the underlying intent of a task. Thus, they cannot infer the user's expectations towards specificity and form of the results. Moreover, a large fraction of queries for retrieving analytical results is rare. As a result, services for intent-aware task recognition perform poorly or are not even triggered on these long-tail queries. We report from a study over 102,360 query and click patterns from a factual search engine. Our analysis reveals six common analytical tasks: explore, relate, resolve, list, compare and answer. To distinguish among these, we study the effects of syntactical structures in the query, methods for interactive entity detection and query segmentation techniques. We evaluate these features on language models and Naive Bayes classifiers. From our evaluation we report a combined F1 score of 90% for the prediction of task intent from keyword queries. Sebastian Arnold 0001, Alexander Löser, Torsten Kilias |
DOLAP | 2 |
| 2015 | INDREX: In-database relation extraction
Torsten Kilias, Alexander Löser, Periklis Andritsos |
Inf. Syst. | 2 |
| 2013 | INDREX: in-database distributional relation extractionabstractRelation extraction transforms the textual representation of a relationship into the relational model of a data warehouse. Early systems, such as SystemT by IBM or the open source system GATE solve this task with handcrafted rule sets that the system executes document-by-document. Thereby the user must execute a highly interactive and iterative process of reading a document, of expressing rules, of testing these rules on the next document and of refining rules. Until now, these systems do neither leverage the full potential of built-in declarative query languages nor the indexing and query optimization techniques of a modern RDBMS that would enable a user interactive rule refinement across documents and on the entire corpus. We propose the INDREX system that enables a user for the first time to describe corpus-wide extraction tasks in a declarative language and permits the user to run interactive rule refinement queries. For enabling this powerful functionality we extend a standard PostgreSQL with a set of white-box user-defined functions that enable corpus-wide transformations from sentences into relationships. We store the text corpus and rules in the same RDBMS that already holds domain specific structured data. As a result, (1) the user can leverage this data to further adapt rules to the target domain, (2) the user does not need an additional system for rule extraction and (3) the INDREX system can leverage the full power of built-in indexing and query optimization techniques of the underlaying RDBMS. In a preliminary study we report on the feasibility of this disruptive approach and show multiple queries in INDREX on the Reuters Corpus, Volume 1. Torsten Kilias, Alexander Löser, Periklis Andritsos |
DOLAP | 2 |
| 2013 | Effective Selectional Restrictions for Unsupervised Relation Extraction
Alan Akbik, Larysa Visengeriyeva, Johannes Kirschnick, Alexander Löser |
IJCNLP | 4 |
| 2012 | Unsupervised Discovery of Relations and Discriminative Extraction Patterns
Alan Akbik, Larysa Visengeriyeva, Priska Herger, Holmer Hemsen, Alexander Löser |
COLING | 5 |
| 2011 | FactCrawl: A Fact Retrieval Framework for Full-Text Indices
Christoph Boden, Alexander Löser, Christoph Nagel, Stephan Pieper |
WebDB | 2 |
| 2009 | Near-duplicate detection for web-forumsabstractCurrent forum search technologies lack the ability to identify threads with near-duplicate content and to group these threads in the search results. As a result, forum users are overloaded with duplicated search results and prefer to create new threads without trying to find existing ones. In this paper we therefore identify common reasons leading to near-duplicates and develop a new near-duplicate detection algorithm for forum threads. The algorithm is implemented using a large case study of a real-world forum serving more than one million users. We compare this work with current algorithms, similar to [4, 5], for detecting near-duplicates on machine generated web pages. Our preliminary results show, that we significantly outperform these algorithms and that we are able to group forum threads with a precision of 74%. Klemens Muthmann, Wojciech M. Barczynski, Falk Brauer, Alexander Löser |
IDEAS | 4 |
| 2009 | Beyond Search: Web-Scale Business Analytics
Alexander Löser |
WISE | 1 |
| 2007 | Navigating the intranet with high precisionabstractDespite the success of web search engines, search over large enterprise intranets still suffers from poor result quality. Earlier work [6] that compared intranets and the Internet from the view point of keyword search has pointed to several reasons why the search problem is quite different in these two domains. In this paper, we address the problem of providing high quality answers to navigational queries in the intranet (e.g., queries intended to find product or personal home pages, service pages, etc.). Our approach is based on offline identification of navigational pages, intelligent generation of term-variants to associate with each page, and the construction of separate indices exclusively devoted to answering navigational queries. Using a testbed of 5.5M pages from the IBM intranet, we present evaluation results that demonstrate that for navigational queries, our approach of using custom indices produces results of significantly higher precision than those produced by a general purpose search algorithm. Huaiyu Zhu 0001, Sriram Raghavan, Shivakumar Vaithyanathan, Alexander Löser |
WWW | 4 |
| 2007 | Semantic social overlay networksabstractPeer selection for query routing is a core task in peer-to-peer networks. Unstructured peer-to-peer systems (like Gnutella) ignore this problem, leading to an abundance of network traffic. Structured peer-to-peer systems (like Chord) enforce a particular, global way of distributing data among the peers in order to solve this problem, but then encounter problems of network volatility and conflicts with the autonomy of the peer data management. In this paper, we propose a new mechanism, INGA, which is based on the observation that query routing in social networks is made possible by locally available knowledge about the expertise of neighbors and a semantics-based peer selection function. We validate INGA by simulation experiments with different data sets. We compare INGA with competing peer selection mechanisms on resulting parameters like recall, message gain or number of messages produced. Alexander Löser, Steffen Staab, Christoph Tempich |
IEEE J. Sel. Areas Commun. | 1 |
| 2005 | Searching Dynamic Communities with Personal Indexes
Alexander Löser, Christoph Tempich, Bastian Quilitz, Wolf-Tilo Balke, Steffen Staab, Wolfgang Nejdl |
ISWC | 1 |
| 2004 | Taxonomy-Based Routing Overlays in P2P Networks
Alexander Löser, Kai Schubert, Frederik Zimmer |
IDEAS | 1 |
| 2004 | Super-peer-based routing strategies for RDF-based peer-to-peer networks
Wolfgang Nejdl, Martin Wolpers, Wolf Siberski, Christoph Schmitz 0001, Mario T. Schlosser, Ingo Brunkhorst, Alexander Löser |
J. Web Semant. | 7 |
| 2003 | Information Integration in Schema-Based Peer-To-Peer Networks
Alexander Löser, Wolf Siberski, Martin Wolpers, Wolfgang Nejdl |
CAiSE | 1 |
| 2003 | Super-peer-based routing and clustering strategies for RDF-based peer-to-peer networksabstractRDF-based P2P networks have a number of advantages compared with simpler P2P networks such as Napster, Gnutella or with approaches based on distributed indices such as CAN and CHORD. RDF-based P2P networks allow complex and extendable descriptions of resources instead of fixed and limited ones, and they provide complex query facilities against these metadata instead of simple keyword-based searches.In previous papers, we have described the Edutella infrastructure and different kinds of Edutella peers implementing such an RDF-based P2P network. In this paper we will discuss these RDF-based P2P networks as a specific example of a new type of P2P networks, schema-based P2P networks, and describe the use of super-peer based topologies for these networks. Super-peer based networks can provide better scalability than broadcast based networks, and do provide perfect support for inhomogeneous schema-based networks, which support different metadata schemas and ontologies (crucial for the Semantic Web). Furthermore, as we will show in this paper, they are able to support sophisticated routing and clustering strategies based on the metadata schemas, attributes and ontologies used. Especially helpful in this context is the RDF functionality to uniquely identify schemas, attributes and ontologies. The resulting routing indices can be built using dynamic frequency counting algorithms and support local mediation and transformation rules, and we will sketch some first ideas for implementing these advanced functionalities as well. Wolfgang Nejdl, Martin Wolpers, Wolf Siberski, Christoph Schmitz 0001, Mario T. Schlosser, Ingo Brunkhorst, Alexander Löser |
WWW | 7 |