VLDB 2026 Research / reviewers in the wild / expert
Wolfgang Nejdl
dblp:n/WolfgangNejdl
· DBLP profile ↗
146ranked-venue papers in the field
9as first author
21since 2021 · last 2026
0000-0003-3374-2193ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 79 (2 first)Database Systems & Data Management · 23 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 23 (5 first)Data Mining & Knowledge Discovery · 14Business Process & Enterprise Data · 4Other / Interdisciplinary · 2Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-Induced Bias Toward LLM-Generated Content in Dense Retrieval
William Xion, Wolfgang Nejdl |
ECIR (1) | 2 |
| 2026 | Cross-Modal Rationale Transfer for Explainable Humanitarian Classification on Social MediaabstractAdvances in social media data dissemination enable the provision of real-time information during a crisis. The information comes from different classes, such as infrastructure damages, persons missing or stranded in the affected zone, etc. Existing methods attempted to classify text and images into various humanitarian categories, but their decision-making process remains largely opaque, which affects their deployment in real-life applications. Recent work has sought to improve transparency by extracting textual rationales from tweets to explain predicted classes. However, such explainable classification methods have mostly focused on text, rather than crisis-related images. In this paper, we propose an interpretable-by-design multimodal classification framework. Our method first learns the joint representation of text and image using a visual language transformer model and extracts text rationales. Next, it extracts the image rationales via the mapping with text rationales. Our approach demonstrates how to learn rationales in one modality from another through cross-modal rationale transfer, which saves annotation effort. Finally, tweets are classified based on extracted rationales. Experiments are conducted over CrisisMMD benchmark dataset, and results show that our proposed method boosts the classification Macro-F1 by 2-35% while extracting accurate text tokens and image patches as rationales. Human evaluation also supports the claim that our proposed method is able to retrieve better image rationale patches (12%) that help to identify humanitarian classes. Our method adapts well to new, unseen datasets in zero-shot mode, achieving an accuracy of 80%. Koustav Rudra, Wolfgang Nejdl |
WWW | 3 |
| 2026 | TransLIME: Towards transfer explainability to explain black-box models on tabular datasetsabstractExplainable Artificial Intelligence methods have gained significant traction for their ability to elucidate the decision-making processes of black-box models, particularly in high-stakes fields such as healthcare and finance. Among these, Local Interpretable Model-agnostic Explanations (LIME) stands out as a widely adopted post-hoc, model-agnostic approach that interprets black-box predictions by constructing an interpretable surrogate model on perturbed instances to approximate the local behavior of the original model around a given instance. However, the effectiveness of LIME can depend on the quality of the training data used by the black-box model. When trained on limited or low-quality data, the black-box model may yield inaccurate predictions for perturbed samples, resulting in poorly defined local decision boundaries and consequently unreliable explanations. This limitation is especially problematic in data-scarce settings. To overcome this challenge, we propose TransLIME, a novel end-to-end explainable transfer learning framework that improves the local fidelity and stability of LIME on limited tabular datasets by transferring relevant explainability knowledge from a related auxiliary source domain with a shifted distribution. Also, in TransLIME, only representative source prototype explanations obtained through clustering are transferred to the target domain, thereby reducing cross-domain exposure of both data and explanatory information during transfer. Experimental evaluations on real-world datasets demonstrate the effectiveness of the proposed framework in improving explanation quality in target domains with limited data. Rehan Raza, Guanjin Wang, Hamid Laga, Kevin Kok Wai Wong, Wolfgang Nejdl |
Inf. Sci. | 5 |
| 2025 | A Systematic Evaluation of Single-Cell Foundation Models on Cell-Type Classification TaskabstractThis study presents a comprehensive benchmarking of three state-of-the-art single-cell foundation models scGPT, Geneformer, and scFoundation, on cell-type classification tasks. We evaluate the models on three datasets: myeloid, human pancreas, and multiple sclerosis, examining both standard fine-tuning and few-shot learning scenarios. Our work reveals that scFoundation consistently achieves the best performance while Geneformer performs poorly, yielding results sometimes even worse than those of the baseline models. Additionally, we demonstrate that a good foundation model can generalize well even when fine-tuned with out-of-distribution data, a capability that the baseline models lack. Our work highlights the potential of foundation models for addressing challenging biomedical questions, particularly in contexts where models are trained on one population but deployed on another. Nicolas Steiner, Omid Vosoughi, Johanna Schrader, Soumyadeep Roy, Wolfgang Nejdl, Ming Tang 0008 |
WSDM | 6 |
| 2025 | Fed-FUEL: fairness and utility enhancing agnostic federated learning frameworkabstractAbstract Federated learning (FL) is an emerging communication-efficient and collaborative learning paradigm of machine learning with privacy guarantees. As these advancements unfold, adapting FL for fairness-aware learning becomes crucial. In this context, we propose a pre-processing fairness and utility (balanced accuracy) enhancing agnostic federated learning framework (Fed-FUEL) that mitigates discrimination embedded in the non-independent identically distributed data. We contribute a novel adaptive data manipulation method that mitigates discrimination embedded in the data at client side during optimization, resulting in an optimized and fair centralized server. This pre-processing approach abstracts the model architecture from the equation, offering a significant advantage in a federated environment. This abstraction not only facilitates a broader application across diverse model architectures without necessitating modifications but also sidesteps the potential complexities and inefficiencies associated with model-specific in-processing methods. Extensive experiments with a range of publicly available datasets demonstrate that our method outperforms the competing baselines in terms of both discrimination mitigation and predictive performance. Our model effectively adapts to both statistical and causal fairness notions, as shown through our experiments. Maryam Badar, Raneen Younis, Sandipan Sikdar, Wolfgang Nejdl, Marco Fisichella |
Data Min. Knowl. Discov. | 4 |
| 2025 | Interpretable zero-shot stance detection with proactive content interventionabstractZero-Shot Stance Detection (ZSSD) identifies an author’s stance towards unseen targets. Existing works have mainly focused on contrastive, meta, adversarial learning, or data augmentation but face issues like data scarcity, generalizability , and lack of coherence between text and targets. Moreover, stance detection must be interpretable to ensure transparency. Recent works with large language models (LLMs) aim to enhance unseen target knowledge or generate explanations but often rely excessively on explicit reasoning or provide coarse explanations, overlooking implicit cues and complicating interpretation. To address these challenges, we propose a novel interpretable multi-stage ZSSD framework. Stage 1 decodes explanations (rationales) justifying the stance while Stage 2 provides the final stance label, thus providing inherent interpretability in predicting stances. Extensive experiments prove that our approach outperforms other baselines with an average improvement in F1 scores of 27.99% with LLMs and 23.60% without LLMs for SemEval and 14.62% with LLMs and 25.24% without LLMs for VAST datasets for the ZSSD task, benefiting from the proposed pipeline architecture and interpretable design. Furthermore, to mitigate the harmful effects of offensive content and promote a more respectful online environment, we integrate an intervention module that leverages the contextual insights derived from our ZSSD framework with the ethics-based text generation power of LLMs to develop interventions. Automatic and human evaluation of LLM-generated interventions based on various proposed criteria provide insights into how LLMs perceive similar information from different perspectives, which can help foster morally sound and respectful online discourse. Apoorva Upadhyaya, Wolfgang Nejdl, Marco Fisichella |
Inf. Process. Manag. | 2 |
| 2024 | Harnessing Empathy and Ethics for Relevance Detection and Information Categorization in Climate and COVID-19 TweetsabstractIn this work, we aim to understand the general public perception of societal issues related to the current climate crisis and the COVID-19 pandemic on Twitter (X). Social media discussions on such matters often lead to misleading information, resulting in delays in initiatives proposed by governments or policymakers. Hence, we focus on extracting relevant information from the conversations on climate change and COVID that could be useful for authorities to curb the spread of potentially biased information by proposing the classification tasks of relevance detection (RD) and information categorization (IC). We first curate the datasets for the RD and IC tasks for the climate domain and extend the COVID-19 benchmark attention-worthy Twitter dataset for the IC task through manual annotation. We initially conduct experiments with LLMs and observe that LLMs can extract the relevant information in zero and few-shot settings based on multi-perspective reasoning in the form of cognitive empathy and ethical standards, but still perform worse than fine-tuned small language models. Based on the initial findings, we conclude that LLMs may not be the best extractor of relevant information, but induce cognitive empathy and ethical reasonings that can intuitively guide supervised models. To achieve this idea, we develop a cognitive empathy and ethical reasoning-based multi-tasking pipelined network for RD and IC tasks. Our proposed approach provides valuable insights that could be useful in real-world scenarios for governments, policymakers, and other researchers to decode the overall public outlook on societal issues. Apoorva Upadhyaya, Wolfgang Nejdl, Marco Fisichella |
CIKM | 2 |
| 2024 | Adaptive Dispatching of Mobile Charging Stations using Multi-Agent Graph Convolutional Cooperative-Competitive Reinforcement LearningabstractBattery electric vehicles (BEV) offer an opportunity to decrease transportation and mobility emissions significantly. The availability of charging station networks and infrastructure is crucial for the proliferation of BEVs. While the expansion of the charging networks is still slow, optimal utilization of the existing infrastructure and dispatching of mobile charging stations can serve as a bypass while more charging stations are built. In this work, we propose a novel multi-agent reinforcement learning - AdapMCS - approach for optimizing the adaptive dispatching of mobile charging stations to maximize the number of served charging requests by a charging station operator while improving the customer experience. By combining graph neural networks with reinforcement learning our approach is able to adapt to dynamic spatio-temporal changes in the demand distribution, for example, during big events such as concerts or fairs. Furthermore, we conduct a thorough evaluation using a publicly available real-world dataset and simulation of dynamic demand distribution changes. The results show that our adaptive dispatching approach is able to deal with the demand shifts and achieve significant gains for both customers, in terms of reducing waiting and charging times, and operators, in terms of increasing their profit. Shimon Wonsak, Nils Henke, Mohammad Alrifai, Michael Nolting, Wolfgang Nejdl |
SIGSPATIAL/GIS | 5 |
| 2024 | Boosting Long-Tail Data Classification with Sparse Prototypical Networks
Alexei Figueroa Rosero, Jens-Michalis Papaioannou, Conor Fallon, Alexandra Bekiaridou, Keno Bressem, Stavros Zanos, Felix A. Gers, Wolfgang Nejdl, Alexander Löser |
ECML/PKDD (7) | 8 |
| 2024 | Beyond Accuracy: Investigating Error Types in GPT-4 Responses to USMLE QuestionsabstractGPT-4 demonstrates high accuracy in medical QA tasks, leading with an accuracy of 86.70%, followed by Med-PaLM 2 at 86.50%. However, around 14% of errors remain. Additionally, current works use GPT-4 to only predict the correct option without providing any explanation and thus do not provide any insight into the thinking process and reasoning used by GPT-4 or other LLMs. Therefore, we introduce a new domain-specific error taxonomy derived from collaboration with medical students. Our GPT-4 USMLE Error (G4UE) dataset comprises 4153 GPT-4 correct responses and 919 incorrect responses to the United States Medical Licensing Examination (USMLE) respectively. These responses are quite long (258 words on average), containing detailed explanations from GPT-4 justifying the selected option. We then launch a large-scale annotation study using the Potato annotation platform and recruit 44 medical experts through Prolific, a well-known crowdsourcing platform. We annotated 300 out of these 919 incorrect data points at a granular level for different classes and created a multi-label span to identify the reasons behind the error. In our annotated dataset, a substantial portion of GPT-4's incorrect responses is categorized as a "Reasonable response by GPT-4," by annotators. This sheds light on the challenge of discerning explanations that may lead to incorrect options, even among trained medical professionals. We also provide medical concepts and medical semantic predications extracted using the SemRep tool for every data point. We believe that it will aid in evaluating the ability of LLMs to answer complex medical questions. We make the resources available at https://github.com/roysoumya/usmle-gpt4-error-taxonomy. Soumyadeep Roy, Aparup Khatua, Fatemeh Ghoochani, Uwe Hadler, Wolfgang Nejdl, Niloy Ganguly |
SIGIR | 5 |
| 2024 | Adversarial Mask Explainer for Graph Neural NetworksabstractThe Graph Neural Networks (GNNs) model is a powerful tool for integrating node information with graph topology to learn representations and make predictions. However, the complex graph structure of GNNs has led to a lack of clear explainability in the decision-making process. Recently, there has been a growing interest in seeking instance-level explanations of the GNNs model, which aims to uncover the decision-making process of the GNNs model and provide insights into how it arrives at its final output. Previous works have focused on finding a set of weights (masks) for edges/nodes/node features to determine their importance. These works have adopted a regularization term and a hyperparameter K to control the explanation size during the training process and keep only the top-K weights as the explanation set. However, the true size of the explanation is typically unknown to users, making it difficult to provide reasonable values for the regularization term and K. In this work, we propose a novel framework AMExplainer which leverages the concept of adversarial networks to achieve a dual optimization objective in the target function. This approach ensures both accurate prediction of the mask and sparsity of the explanation set. In addition, we devise a novel scaling function to automatically sense and amplify the weights of the informative part of the graph, which filters out insignificant edges/nodes/node features for expediting the convergence of the solution during training. Our extensive experiments show that AMExplainer yields a more compelling explanation by generating a sparse set of masks while simultaneously maintaining fidelity. Wei Zhang 0309, Xiaofan Li 0004, Wolfgang Nejdl |
WWW | 3 |
| 2024 | Open benchmark for filtering techniques in entity resolution
Franziska Neuhof, Marco Fisichella, George Papadakis 0001, Konstantinos Nikoletos, Nikolaus Augsten, Wolfgang Nejdl, Manolis Koubarakis |
VLDB J. | 6 |
| 2023 | Benchmarking Filtering Techniques for Entity ResolutionabstractEntity Resolution is the task of identifying pairs of entity profiles that represent the same real-world object. To avoid checking a quadratic number of entity pairs, various filtering techniques have been proposed that fall into two main categories: (i) blocking workflows group together entity profiles with identical or similar signatures, and (ii) nearest-neighbor methods convert all entity profiles into vectors and identify the closest ones to every query entity. Unfortunately, the main techniques from these two categories have rarely been compared in the literature and, thus, their relative performance is unknown. We perform the first systematic experimental study that investigates the relative performance of the main representatives per category over numerous established datasets. Comparing techniques from different categories turns out to be a non-trivial task due to the various configuration parameters that are hard to fine-tune, but have a significant impact on performance. We consider a plethora of parameter configurations, optimizing each technique with respect to recall and precision targets. Both schema-agnostic and schema-based settings are evaluated. The experimental results provide novel insights into the effectiveness, the time efficiency and the scalability of the considered techniques. George Papadakis 0001, Marco Fisichella, Franziska Schoger, Georgios M. Mandilaras, Nikolaus Augsten, Wolfgang Nejdl |
ICDE | 6 |
| 2023 | A Multi-Task Model for Sentiment Aided Stance Detection of Climate Change TweetsabstractClimate change has become one of the biggest challenges of our time. Social media platforms such as Twitter play an important role in raising public awareness and spreading knowledge about the dangers of the current climate crisis. With the increasing number of campaigns and communication about climate change through social media, the information could create more awareness and reach the general public and policy makers. However, these Twitter communications lead to polarization of beliefs, opinion-dominated ideologies, and often a split into two communities of climate change deniers and believers. In this paper, we propose a framework that helps identify denier statements on Twitter and thus classifies the stance of the tweet into one of the two attitudes towards climate change (denier/believer). The sentimental aspects of Twitter data on climate change are deeply rooted in general public attitudes toward climate change. Therefore, our work focuses on learning two closely related tasks: Stance Detection and Sentiment Analysis of climate change tweets. We propose a multi-task framework that performs stance detection (primary task) and sentiment analysis (auxiliary task) simultaneously. The proposed model incorporates the feature-specific and shared-specific attention frameworks to fuse multiple features and learn the generalized features for both tasks. The experimental results show that the proposed framework increases the performance of the primary task, i.e., stance detection by benefiting from the auxiliary task, i.e., sentiment analysis compared to its uni-modal and single-task variants. Apoorva Upadhyaya, Marco Fisichella, Wolfgang Nejdl |
ICWSM | 3 |
| 2023 | A Multi-task Model for Emotion and Offensive Aided Stance Detection of Climate Change TweetsabstractIn this work, we address the United Nations Sustainable Development Goal 13: Climate Action by focusing on identifying public attitudes toward climate change on social media platforms such as Twitter. Climate change is threatening the health of the planet and humanity. Public engagement is critical to address climate change. However, climate change conversations on Twitter tend to polarize beliefs, leading to misinformation and fake news that influence public attitudes, often dividing them into climate change believers and deniers. Our paper proposes an approach to classify the attitude of climate change tweets (believe/deny/ambiguous) to identify denier statements on Twitter. Most existing approaches for detecting stances and classifying climate change tweets either overlook deniers’ tweets or do not have a suitable architecture. The relevant literature suggests that emotions and higher levels of toxicity are prevalent in climate change Twitter conversations, leading to a delay in appropriate climate action. Therefore, our work focuses on learning stance detection (main task) while exploiting the auxiliary tasks of recognizing emotions and offensive utterances. We propose a multimodal multitasking framework MEMOCLiC that captures the input data using different embedding techniques and attention frameworks, and then incorporates the learned emotional and offensive expressions to obtain an overall representation of the features relevant to the stance of the input tweet. Extensive experiments conducted on a novel curated climate change dataset and two benchmark stance detection datasets (SemEval-2016 and ClimateStance-2022) demonstrate the effectiveness of our approach. Apoorva Upadhyaya, Marco Fisichella, Wolfgang Nejdl |
WWW | 3 |
| 2023 | Towards sentiment and Temporal Aided Stance Detection of climate change tweets
Apoorva Upadhyaya, Marco Fisichella, Wolfgang Nejdl |
Inf. Process. Manag. | 3 |
| 2022 | Unraveling Social Perceptions & Behaviors towards Migrants on Twitter
Aparup Khatua, Wolfgang Nejdl |
ICWSM | 2 |
| 2021 | Efficient Scalable Temporal Web Graph StoreabstractTemporal web graphs have been attracting much attention recently due to their important applications in web search, data mining, and social network analysis. Accumulated over long periods, those graphs have grown gigantic in size and rich in temporal evolution, which poses tough challenges for data storage and management. Though a few temporal graph management systems were previously proposed, none of them can simultaneously satisfy both essential requirements when retrieving on temporal web graphs: very large data scalability and very low querying latency.In this work, we address the above gap in existing works by developing a highly efficient temporal graph management system which is dedicated to web graphs. To this end, we greatly extend the most efficient framework for managing large static web graphs to handle temporal information using the property matrix while preserving most of the outstanding features of the base framework. Ultimately, our proposed system can achieve a nearly instant response for vertex-centric temporal retrieval while still being scalable to huge datasets. Experiments on a real-world dataset with more than 43B nodes and 317B links show that using a small non-dedicated cluster, our system can reach a reduction of data storage space up to 88% of raw data size and reduce the retrieval time by 20%, compared to the baselines. We also demonstrate that our system also yields a significant reduction of computational costs for many graph ranking algorithms. Khoi Duy Vo, Sergej Zerr, Xiaofei Zhu, Wolfgang Nejdl |
IEEE BigData | 4 |
| 2021 | FARF: A Fair and Adaptive Random Forests Classifier
Wenbin Zhang 0002, Albert Bifet, Xiangliang Zhang 0001, Jeremy C. Weiss, Wolfgang Nejdl |
PAKDD (2) | 5 |
| 2021 | EduCOR: An Educational and Career-Oriented Recommendation OntologyabstractAbstract With the increased dependence on online learning platforms and educational resource repositories, a unified representation of digital learning resources becomes essential to support a dynamic and multi-source learning experience. We introduce the EduCOR ontology, an educational, career-oriented ontology that provides a foundation for representing online learning resources for personalised learning systems. The ontology is designed to enable learning material repositories to offer learning path recommendations, which correspond to the user’s learning goals and preferences, academic and psychological parameters, and labour-market skills. We present the multiple patterns that compose the EduCOR ontology, highlighting its cross-domain applicability and integrability with other ontologies. A demonstration of the proposed ontology on the real-life learning platform eDoer is discussed as a use case. We evaluate the EduCOR ontology using both gold standard and task-based approaches. The comparison of EduCOR to three gold schemata, and its application in two use-cases, shows its coverage and adaptability to multiple OER repositories, which allows generating user-centric and labour-market oriented recommendations. Resource: https://tibonto.github.io/educor/ . Eleni Ilkou, Hasan Abu-Rasheed, MohammadReza Tavakoli, Sherzod Hakimov, Gábor Kismihók, Sören Auer, Wolfgang Nejdl |
ISWC | 7 |
| 2021 | Hashing-Accelerated Graph Neural Networks for Link PredictionabstractNetworks are ubiquitous in the real world. Link prediction, as one of the key problems for network-structured data, aims to predict whether there exists a link between two nodes. The traditional approaches are based on the explicit similarity computation between the compact node representation by embedding each node into a low-dimensional space. In order to efficiently handle the intensive similarity computation in link prediction, the hashing technique has been successfully used to produce the node representation in the Hamming space. However, the hashing-based link prediction algorithms face accuracy loss from the randomized hashing techniques or inefficiency from the learning to hash techniques in the embedding process. Currently, the Graph Neural Network (GNN) framework has been widely applied to the graph-related tasks in an end-to-end manner, but it commonly requires substantial computational resources and memory costs due to massive parameter learning, which makes the GNN-based algorithms impractical without the help of a powerful workhorse. In this paper, we propose a simple and effective model called #GNN, which balances the trade-off between accuracy and efficiency. #GNN is able to efficiently acquire node representation in the Hamming space for link prediction by exploiting the randomized hashing technique to implement message passing and capture high-order proximity in the GNN framework. Furthermore, we characterize the discriminative power of #GNN in probability. The extensive experimental results demonstrate that the proposed #GNN algorithm achieves accuracy comparable to the learning-based algorithms and outperforms the randomized algorithm, while running significantly faster than the learning-based algorithms. Also, the proposed algorithm shows excellent scalability on a large-scale network with the limited resources. Wei Wu 0011, Bin Li 0015, Chuan Luo 0002, Wolfgang Nejdl |
WWW | 4 |
| 2020 | Matching Recruiters and Jobseekers on TwitterabstractAn efficient job recommendation framework needs to recommend an appropriate jobseeker to a recruiter and vice-versa. Prior studies have mostly considered datasets from commercial job portals such as LinkedIn or CareerBuilder. However, these datasets are proprietary and not publicly available. Moreover, these portals charge their clients for offering customized services. Hence, we explore whether publicly available Twitter data can be a viable alternative to commercial job portals. We have extracted 0.76 million job-related tweets. We have manually annotated tweet-pairs from recruiters and jobseekers in the domain of computer science jobs. Next, we have employed Siamese architecture and considered multiple artificial neural network models with different word embeddings. We have achieved around 97% accuracy for some of our models. Our study demonstrates the potential of the Twitter platform for job recommendations. Aparup Khatua, Wolfgang Nejdl |
ASONAM | 2 |
| 2019 | Node Representation Learning for Directed Graphs
Megha Khosla, Jurek Leonhardt, Wolfgang Nejdl, Avishek Anand |
ECML/PKDD (1) | 3 |
| 2019 | Efficient Summarizing of Evolving Events from Twitter StreamsabstractTwitter has been heavily used for users to report and share information about real-world events. However, understanding the multiple aspects of an event as it happens is a very challenging task due to the prevalent noise and redundant in tweets as well as the evolution of the event. In this paper, we present a graph-based method for summarizing evolutionary events from tweet streams. Unlike existing approaches that either require prior information, result in less readable summaries, or are not scalable, our proposed method can automatically extract sets of representative tweets as concise summaries for the events. Moreover, the method also allows the summaries to be updated efficiently using an incremental procedure, thus can scale up to large data streams. The experiments on five datasets reveal that our proposed method significantly outperforms several baselines. Tuan-Anh Hoang, Wolfgang Nejdl |
SDM | 3 |
| 2018 | W2E: A Worldwide-Event Benchmark Dataset for Topic Detection and TrackingabstractTopic detection and tracking in document streams is a critical task in many important applications, hence has been attracting research interest in recent decades. With the large size of data streams, there have been a number of works from different approaches that propose automatic methods for the task. However, there is only a few small benchmark datasets that are publicly available for evaluating the proposed methods. The lack of large datasets with fine-grained groundtruth implicitly restrains the development of more advanced methods. In this work, we address this issue by collecting and publishing W2E - a large dataset consisting of news articles from more than 50 prominent mass media channels worldwide. The articles cover a large set of popular events within a full year. W2E is more than 15 times larger than TREC's TDT2 dataset, which is widely used in prior work. We further conduct exploratory analysis to examine the dynamics and diversity of W2E and propose potential uses of the dataset in other research. Tuan-Anh Hoang, Khoi Duy Vo, Wolfgang Nejdl |
CIKM | 3 |
| 2018 | Multiple Models for Recommending Temporal Aspects of Entities
Tu Ngoc Nguyen, Nattiya Kanhabua, Wolfgang Nejdl |
ESWC | 3 |
| 2018 | LogCanvas: Visualizing Search History Using Knowledge GraphsabstractIn this demo paper, we introduce LogCanvas, a platform for user search history visualization.Different from the existing visualization tools, LogCanvas focuses on helping users re-construct the semantic relationship among their search activities. LogCanvas segments a user's search history into different sessions and generates a knowledge graph to represent the information exploration process in each session.A knowledge graph is composed of the most important concepts or entities discovered by each search query as well as their relationships. It thus captures the semantic relationship among the queries.LogCanvas offers a session timeline viewer and a snippets viewer to enable users to re-find their previous search results efficiently. LogCanvas also provides a collaborative perspective to support a group of users in sharing search results and experience. Luyan Xu, Zeon Trevor Fernando, Xuan Zhou 0001, Wolfgang Nejdl |
SIGIR | 4 |
| 2017 | RussianFlu-DE: A German Corpus for a Historical Epidemic with Temporal Annotation
Tran Van Canh, Katja Markert, Wolfgang Nejdl |
TPDL | 3 |
| 2016 | History by Diversity: Helping Historians search News ArchivesabstractLongitudinal corpora like newspaper archives are of immense value to historical research, and time as an important factor for historians strongly influences their search behaviour in these archives. While searching for articles published over time, a key preference is to retrieve documents which cover the important aspects from important points in time which is different from standard search behavior. To support this search strategy, we introduce the notion of a Historical Query Intent to explicitly model a historian's search task and define an aspect-time diversification problem over news archives. Wolfgang Nejdl, Avishek Anand |
CHIIR | 2 |
| 2016 | Finding News Citations for WikipediaabstractAn important editing policy in Wikipedia is to provide citations for added statements in Wikipedia pages, where statements can be arbitrary pieces of text, ranging from a sentence to a paragraph. In many cases citations are either outdated or missing altogether. Besnik Fetahu, Katja Markert, Wolfgang Nejdl, Avishek Anand |
CIKM | 3 |
| 2016 | ArchiveWeb: Collaboratively Extending and Exploring Web Archive Collections
Zeon Trevor Fernando, Ivana Marenzi, Wolfgang Nejdl, Rishita Kalyani |
TPDL | 3 |
| 2016 | How to Search the Internet Archive Without Indexing It
Nattiya Kanhabua, Philipp Kemkes, Wolfgang Nejdl, Tu Ngoc Nguyen, Felipe Reis, Nam Khanh Tran |
TPDL | 3 |
| 2016 | On the Applicability of Delicious for Temporal Search on Web ArchivesabstractWeb archives are large longitudinal collections that store webpages from the past, which might be missing on the current live Web. Consequently, temporal search over such collections is essential for finding prominent missing webpages and tasks like historical analysis. However, this has been challenging due to the lack of popularity information and proper ground truth to evaluate temporal retrieval models. In this paper we investigate the applicability of external longitudinal resources to identify important and popular websites in the past and analyze the social bookmarking service Delicious for this purpose. The timestamped bookmarks on Delicious provide explicit cues about popular time periods in the past along with relevant descriptors. These are valuable to identify important documents in the past for a given temporal query. Focusing purely on recall, we analyzed more than 12,000 queries and find that using Delicious yields average recall values from 46% up to 100%, when limiting ourselves to the best represented queries in the considered dataset. This constitutes an attractive and low-overhead approach for quick access into Web archives by not dealing with the actual contents. Helge Holzmann, Wolfgang Nejdl, Avishek Anand |
SIGIR | 2 |
| 2016 | Expedition: A Time-Aware Exploratory Search System Designed for ScholarsabstractArchives are an important source of study for various scholars. Digitization and the web have made archives more accessible and led to the development of several time-aware exploratory search systems. However these systems have been designed for more general users rather than scholars. Scholars have more complex information needs in comparison to general users. They also require support for corpus creation during their exploration process. In this paper we present Expedition - a time-aware exploratory search system that addresses the requirements and information needs of scholars. Expedition possesses a suite of ad-hoc and diversity based retrieval models to address complex information needs; a newspaper-style user interface to allow for larger textual previews and comparisons; entity filters to more naturally refine a result list and an interactive annotated timeline which can be used to better identify periods of importance. Wolfgang Nejdl, Avishek Anand |
SIGIR | 2 |
| 2015 | A Random Walk Model for Optimization of Search Impact in Web Frontier RankingabstractLarge-scale web search engines need to crawl the Web continuously to discover and download newly created web content. The speed at which the new content is discovered and the quality of the discovered content can have a big impact on the coverage and quality of the results provided by the search engine. In this paper, we propose a search-centric solution to the problem of prioritizing the pages in the frontier of a crawler for download. Our approach essentially orders the web pages in the frontier through a random walk model that takes into account the pages' potential impact on user-perceived search quality. In addition, we propose a link graph enrichment technique that extends this solution. Finally, we explore a machine learning approach that combines different frontier prioritization approaches. We conduct experiments using two very large, real-life web datasets to observe various search quality metrics. Comparisons with several baseline techniques indicate that the proposed approaches have the potential to improve the user-perceived quality of web search results considerably. Giang Tran, Ata Turk, Berkant Barla Cambazoglu, Wolfgang Nejdl |
SIGIR | 4 |
| 2014 | Optimizing Multi-Relational Factorization Models for Multiple Target RelationsabstractMulti-matrix factorization models provide a scalable and effective approach for multi-relational learning tasks such as link prediction, Linked Open Data (LOD) mining, recommender systems and social network analysis. Such models are learned by optimizing the sum of the losses on all relations in the data. Early models address the problem where there is only one target relation for which predictions should be made. More recent models address the multi-target variant of the problem and use the same set of parameters to make predictions for all target relations. In this paper, we argue that a model optimized for each target relation individually has better predictive performance than models optimized for a compromise on the performance on all target relations. We introduce specific parameters for each target but, instead of learning them independently from each other, we couple them through a set of shared auxiliary parameters, which has a regularizing effect on the target specific ones. Experiments on large Web datasets derived from DBpedia, Wikipedia and BlogCatalog show the performance improvement obtained by using target specific parameters and that our approach outperforms competitive state-of-the-art methods while being able to scale gracefully to big data. Lucas Drumond, Ernesto Diaz-Aviles, Lars Schmidt-Thieme, Wolfgang Nejdl |
CIKM | 4 |
| 2014 | Predicting Pair Similarities for Near-Duplicate Detection in High Dimensional Spaces
Marco Fisichella, Andrea Ceroni, Fan Deng 0004, Wolfgang Nejdl |
DEXA (2) | 4 |
| 2014 | A Scalable Approach for Efficiently Generating Structured Dataset Topic Profiles
Besnik Fetahu, Stefan Dietze, Bernardo Pereira Nunes, Marco A. Casanova, Davide Taibi 0002, Wolfgang Nejdl |
ESWC | 6 |
| 2014 | An adaptive teleportation random walk model for learning social tag relevanceabstractSocial tags are known to be a valuable source of information for image retrieval and organization. However, contrary to the conventional document retrieval, rich tag frequency information in social sharing systems, such as Flickr, is not available, thus we cannot directly use the tag frequency (analogous to the term frequency in a document) to represent the relevance of tags. Many heuristic approaches have been proposed to address this problem, among which the well-known neighbor voting based approaches are the most effective methods. The basic assumption of these methods is that a tag is considered as relevant to the visual content of a target image if this tag is also used to annotate the visual neighbor images of the target image by lots of different users. The main limitation of these approaches is that they treat the voting power of each neighbor image either equally or simply based on its visual similarity. In this paper, we cast the social tag relevance learning problem as an adaptive teleportation random walk process on the voting graph. In particular, we model the relationships among images by constructing a voting graph, and then propose an adaptive teleportation random walk, in which a confidence factor is introduced to control the teleportation probability, on the voting graph. Through this process, direct and indirect relationships among images can be explored to cooperatively estimate the tag relevance. To quantify the performance of our approach, we compare it with state-of-the-art methods on two publicly available datasets (NUS-WIDE and MIR Flickr). The results indicate that our method achieves substantial performance gains on these datasets. Xiaofei Zhu, Wolfgang Nejdl, Mihai Georgescu |
SIGIR | 2 |
| 2014 | Meta-Blocking: Taking Entity Resolutionto the Next LevelabstractEntity Resolution is an inherently quadratic task that typically scales to large data collections through blocking. In the context of highly heterogeneous information spaces, blocking methods rely on redundancy in order to ensure high effectiveness at the cost of lower efficiency (i.e., more comparisons). This effect is partially ameliorated by coarse-grained block processing techniques that discard entire blocks either a-priori or during the resolution process. In this paper, we introduce meta-blocking as a generic procedure that intervenes between the creation and the processing of blocks, transforming an initial set of blocks into a new one with substantially fewer comparisons and equally high effectiveness. In essence, meta-blocking aims at extracting the most similar pairs of entities by leveraging the information that is encapsulated in the block-to-entity relationships. To this end, it first builds an abstract graph representation of the original set of blocks, with the nodes corresponding to entity profiles and the edges connecting the co-occurring ones. During the creation of this structure all redundant comparisons are discarded, while the superfluous ones can be removed by pruning of the edges with the lowest weight. We analytically examine both procedures, proposing a multitude of edge weighting schemes, graph pruning algorithms as well as pruning criteria. Our approaches are schema-agnostic, thus accommodating any type of blocks. We evaluate their performance through a thorough experimental study over three large-scale, real-world data sets, with the outcomes verifying significant efficiency enhancements at a negligible cost in effectiveness. George Papadakis 0001, Georgia Koutrika, Themis Palpanas, Wolfgang Nejdl |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2014 | Analyzing and Mining Comments and Comment Ratings on the Social WebabstractAn analysis of the social video sharing platform YouTube and the news aggregator Yahoo! News reveals the presence of vast amounts of community feedback through comments for published videos and news stories, as well as through metaratings for these comments. This article presents an in-depth study of commenting and comment rating behavior on a sample of more than 10 million user comments on YouTube and Yahoo! News. In this study, comment ratings are considered first-class citizens. Their dependencies with textual content, thread structure of comments, and associated content (e.g., videos and their metadata) are analyzed to obtain a comprehensive understanding of the community commenting behavior. Furthermore, this article explores the applicability of machine learning and data mining to detect acceptance of comments by the community, comments likely to trigger discussions, controversial and polarizing content, and users exhibiting offensive commenting behavior. Results from this study have potential application in guiding the design of community-oriented online discussion platforms. Stefan Siersdorfer, Sergiu Chelaru, José San Pedro, Ismail Sengör Altingövde, Wolfgang Nejdl |
ACM Trans. Web | 5 |
| 2013 | Aligning freebase with the YAGO ontologyabstractLinked Open Data (LOD) has emerged as the de-facto standard for publishing data on the Web. The cross-domain large scale Freebase and YAGO datasets represent central hubs and reference points for the LOD cloud. Freebase is an open-world dataset, which contains about 22 million entities and more than 350 million facts in more than 100 domains. The scale of Freebase makes it difficult for the users to get an overview of the data and efficiently retrieve the desired information. Integration of Freebase with the YAGO ontology that contains more than 360,000 concepts enables us to provide more semantic information for Freebase and to facilitate novel applications, such as efficient query construction, over large scale data. In this paper we analyze the structure of YAGO in more depth and show how to match YAGO and Freebase categories. The new YAGO+F structure that results from our matching tightly connects both datasets and provides an important next step to systematically interconnect LOD subcollections. We make our YAGO+F structure available online in the hope that it can provide a good starting point for future applications, which can build upon a wide variety of Freebase data clearly arranged in the semantic categories of YAGO. Elena Demidova, Irina Oelze, Wolfgang Nejdl |
CIKM | 3 |
| 2013 | Extracting Event-Related Information from Article Updates in Wikipedia
Mihai Georgescu, Nattiya Kanhabua, Daniel Krause 0002, Wolfgang Nejdl, Stefan Siersdorfer |
ECIR | 4 |
| 2013 | Recommending High Utility Query via Session-Flow Graph
Xiaofei Zhu, Jiafeng Guo, Xueqi Cheng 0001, Yanyan Lan, Wolfgang Nejdl |
ECIR | 5 |
| 2013 | Combining a Co-occurrence-Based and a Semantic Measure for Entity Linking
Bernardo Pereira Nunes, Stefan Dietze, Marco A. Casanova, Ricardo Kawase, Besnik Fetahu, Wolfgang Nejdl |
ESWC | 6 |
| 2013 | Efficient query construction for large scale dataabstractIn recent years, a number of open databases have emerged on the Web, providing Web users with platforms to collaboratively create structured information. As these databases are intended to accommodate heterogeneous information and knowledge, they usually comprise a very large schema and billions of instances. Browsing and searching data on such a scale is not an easy task for a Web user. In this context, interactive query construction offers an intuitive interface for novice users to retrieve information from databases neither requiring any knowledge of structured query languages, nor any prior knowledge of the database schema. However, the existing mechanisms do not scale well on large scale datasets. This paper presents a set of techniques to boost the scalability of interactive query construction, from the perspective of both, user interaction cost and performance. We connect an abstract ontology layer to the database schema to shorten the process of user-computer interaction. We also introduce a search mechanism to enable efficient exploration of query interpretation spaces over large scale data. Extensive experiments show that our approach scales well on Freebase - an open database containing more than 7,000 relational tables in more than 100 domains. Elena Demidova, Xuan Zhou 0001, Wolfgang Nejdl |
SIGIR | 3 |
| 2013 | Introduction to the special section on twitter and microblogging servicesabstractNo abstract available. Irwin King, Wolfgang Nejdl |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2013 | A Blocking Framework for Entity Resolution in Highly Heterogeneous Information SpacesabstractIn the context of entity resolution (ER) in highly heterogeneous, noisy, user-generated entity collections, practically all block building methods employ redundancy to achieve high effectiveness. This practice, however, results in a high number of pairwise comparisons, with a negative impact on efficiency. Existing block processing strategies aim at discarding unnecessary comparisons at no cost in effectiveness. In this paper, we systemize blocking methods for clean-clean ER (an inherently quadratic task) over highly heterogeneous information spaces (HHIS) through a novel framework that consists of two orthogonal layers: the effectiveness layer encompasses methods for building overlapping blocks with small likelihood of missed matches; the efficiency layer comprises a rich variety of techniques that significantly restrict the required number of pairwise comparisons, having a controllable impact on the number of detected duplicates. We map to our framework all relevant existing methods for creating and processing blocks in the context of HHIS, and additionally propose two novel techniques: attribute clustering blocking and comparison scheduling. We evaluate the performance of each layer and method on two large-scale, real-world data sets and validate the excellent balance between efficiency and effectiveness that they achieve. George Papadakis 0001, Ekaterini Ioannou, Themis Palpanas, Claudia Niederée, Wolfgang Nejdl |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2013 | Analyzing, Detecting, and Exploiting Sentiment in Web QueriesabstractThe Web contains an increasing amount of biased and opinionated documents on politics, products, and polarizing events. In this article, we present an indepth analysis of Web search queries for controversial topics, focusing on query sentiment. To this end, we conduct extensive user assessments and discriminative term analyses, as well as a sentiment analysis using the SentiWordNet thesaurus, a lexical resource containing sentiment annotations. Furthermore, in order to detect the sentiment expressed in queries, we build different classifiers based on query texts, query result titles, and snippets. We demonstrate the virtue of query sentiment detection in two different use cases. First, we define a query recommendation scenario that employs sentiment detection of results to recommend additional queries for polarized queries issued by search engine users. The second application scenario is controversial topic discovery, where query sentiment classifiers are employed to discover previously unknown topics that trigger both highly positive and negative opinions among the users of a search engine. For both use cases, the results of our evaluations on real-world data are promising and show the viability and potential of query sentiment analysis in practical scenarios. Sergiu Chelaru, Ismail Sengör Altingövde, Stefan Siersdorfer, Wolfgang Nejdl |
ACM Trans. Web | 4 |
| 2012 | What is happening right now ... that interests me?: online topic discovery and recommendation in twitterabstractUsers engaged in the Social Web increasingly rely upon continuous streams of Twitter messages (tweets) for real-time access to information and fresh knowledge about current affairs. However, given the deluge of tweets, it is a challenge for individuals to find relevant and appropriately ranked information. We propose to address this knowledge management problem by going beyond the general perspective of information finding in Twitter, that asks: "What is happening right now?", towards an individual user perspective, and ask: "What is interesting to me right now?" In this paper, we consider collaborative filtering as an online ranking problem and present RMFO, a method that creates, in real-time, user-specific rankings for a set of tweets based on individual preferences that are inferred from the user's past system interactions. Experiments on the 476 million Twitter tweets dataset show that our online approach largely outperforms recommendations based on Twitter's global trend and Weighted Regularized Matrix Factorization (WRMF), a highly competitive state-of-the-art Collaborative Filtering technique, demonstrating the efficacy of our approach. Ernesto Diaz-Aviles, Lucas Drumond, Zeno Gantner, Lars Schmidt-Thieme, Wolfgang Nejdl |
CIKM | 5 |
| 2012 | Map to humans and reduce error: crowdsourcing for deduplication applied to digital librariesabstractDetecting duplicate entities, usually by examining metadata, has been the focus of much recent work. Several methods try to identify duplicate entities, while focusing either on accuracy or on efficiency and speed - with still no perfect solution. We propose a combined layered approach for duplicate detection with the main advantage of using Crowdsourcing as a training and feedback mechanism. By using Active Learning techniques on human provided examples, we fine tune our algorithm toward better duplicate detection accuracy. We keep the training cost low by gathering training data on demand for borderline cases or for inconclusive assessments. We apply our simple and powerful methods to an online publication search system: First, we perform a coarse duplicate detection relying on publication signatures in real time. Then, a second automatic step compares duplicate candidates and increases accuracy while adjusting based on both feedback from our online users and from Crowdsourcing platforms. Our approach shows an improvement of 14% over the untrained setting and is at only 4% difference to the human assessors in accuracy. Mihai Georgescu, Dang Duc Pham, Claudiu S. Firan, Wolfgang Nejdl, Julien Gaugaz |
CIKM | 4 |
| 2012 | Supporting temporal analytics for health-related events in microblogsabstractMicroblogging services, such as Twitter, are gaining interests as a means of sharing information in social networks. Numerous works have shown the potential of using Twitter posts (or tweets) in order to infer the existence and magnitude of real-world events. In the medical domain, there has been a surge in detecting public health related tweets for early warning so that a rapid response from health authorities can take place. In this paper, we present a temporal analytics tool for supporting a comparative, temporal analysis of disease outbreaks between Twitter and official sources, such as, World Health Organization (WHO) and ProMED-mail. We automatically extract and aggregate outbreak events from official outbreak reports, producing time series data. Our tool can support a correlation analysis and an understanding of the temporal developments of outbreak mentions in Twitter, based on comparisons with official sources. Nattiya Kanhabua, Sara Romano, Avare Stewart, Wolfgang Nejdl |
CIKM | 4 |
| 2012 | Epidemic Intelligence for the Crowd, by the Crowd
Ernesto Diaz-Aviles, Avare Stewart, Edward Velasco, Kerstin Denecke, Wolfgang Nejdl |
ICWSM | 5 |
| 2012 | Real-time top-n recommendation in social streamsabstractThe Social Web is successfully established, and steadily growing in terms of users, content and services. People generate and consume data in real-time within social networking services, such as Twitter, and increasingly rely upon continuous streams of messages for real-time access to fresh knowledge about current affairs. In this paper, we focus on analyzing social streams in real-time for personalized topic recommendation and discovery. We consider collaborative filtering as an online ranking problem and present Stream Ranking Matrix Factorization - RMFX -, which uses a pairwise approach to matrix factorization in order to optimize the personalized ranking of topics. Our novel approach follows a selective sampling strategy to perform online model updates based on active learning principles, that closely simulates the task of identifying relevant items from a pool of mostly uninteresting ones. RMFX is particularly suitable for large scale applications and experiments on the "476 million Twitter tweets" dataset show that our online approach largely outperforms recommendations based on Twitter's global trend, and it is also able to deliver highly competitive Top-N recommendations faster while using less space than Weighted Regularized Matrix Factorization (WRMF), a state-of-the-art matrix factorization technique for Collaborative Filtering, demonstrating the efficacy of our approach. Ernesto Diaz-Aviles, Lucas Drumond, Lars Schmidt-Thieme, Wolfgang Nejdl |
RecSys | 4 |
| 2012 | Swarming to rank for recommender systemsabstractRecommender systems make product suggestions that are tailored to the user's individual needs and represent powerful means to combat information overload. In this paper, we focus on the item prediction task of Recommender Systems and present SwarmRankCF, a method to automatically optimize the performance quality of recommender systems using a Swarm Intelligence perspective. Our approach, which is well-founded in a Particle Swarm Optimization framework, learns a ranking function by optimizing the combination of unique characteristics (i.e., features) of users, items and their interactions. In particular, we build feature vectors from a factorization of the user-item interaction matrix, and directly optimize Mean Average Precision metric in order to learn a linear ranking model for personalized recommendations. Our experimental evaluation, on a real world online radio dataset, indicates that our approach is able to find ranking functions that significantly improve the performance of the system for the Top-N recommendation task. Ernesto Diaz-Aviles, Mihai Georgescu, Wolfgang Nejdl |
RecSys | 3 |
| 2012 | Beyond 100 million entities: large-scale blocking-based resolution for heterogeneous dataabstractA prerequisite for leveraging the vast amount of data available on the Web is Entity Resolution, i.e., the process of identifying and linking data that describe the same real-world objects. To make this inherently quadratic process applicable to large data sets, blocking is typically employed: entities (records) are grouped into clusters - the blocks - of matching candidates and only entities of the same block are compared. However, novel blocking techniques are required for dealing with the noisy, heterogeneous, semi-structured, user-generateddata in the Web, as traditional blocking techniques are inapplicable due to their reliance on schema information. The introduction of redundancy, improves the robustness of blocking methods but comes at the price of additional computational cost. George Papadakis 0001, Ekaterini Ioannou, Claudia Niederée, Themis Palpanas, Wolfgang Nejdl |
WSDM | 5 |
| 2012 | Using site-level connections to estimate link confidenceabstractSearch engines are essential tools for web users today. They rely on a large number of features to compute the rank of search results for each given query. The estimated reputation of pages is among the effective features available for search engine designers, probably being adopted by most current commercial search engines. Page reputation is estimated by analyzing the linkage relationships between pages. This information is used by link analysis algorithms as a query‐independent feature, to be taken into account when computing the rank of the results. Unfortunately, several types of links found on the web may damage the estimated page reputation and thus cause a negative effect on the quality of search results. This work studies alternatives to reduce the negative impact of such noisy links. More specifically, the authors propose and evaluate new methods that deal with noisy links, considering scenarios where the reputation of pages is computed using the PageRank algorithm. They show, through experiments with real web content, that their methods achieve significant improvements when compared to previous solutions proposed in the literature. Jucimar Brito de Souza, André Luiz da Costa Carvalho, Marco Cristo, Edleno Silva de Moura, Pável Calado, Paul-Alexandru Chirita, Wolfgang Nejdl |
J. Assoc. Inf. Sci. Technol. | 7 |
| 2012 | A Probabilistic Scheme for Keyword-Based Incremental Query ConstructionabstractDatabases enable users to precisely express their informational needs using structured queries. However, database query construction is a laborious and error-prone process, which cannot be performed well by most end users. Keyword search alleviates the usability problem at the price of query expressiveness. As keyword search algorithms do not differentiate between the possible informational needs represented by a keyword query, users may not receive adequate results. This paper presents IQP- a novel approach to bridge the gap between usability of keyword search and expressiveness of database queries. IQPenables a user to start with an arbitrary keyword query and incrementally refine it into a structured query through an interactive interface. The enabling techniques of IQPinclude: 1) a probabilistic framework for incremental query construction; 2) a probabilistic model to assess the possible informational needs represented by a keyword query; 3) an algorithm to obtain the optimal query construction process. This paper presents the detailed design of IQP, and demonstrates its effectiveness and scalability through experiments over real-world data and a user study. Elena Demidova, Xuan Zhou 0001, Wolfgang Nejdl |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2012 | A hybrid approach for efficient Web service composition with end-to-end QoS constraintsabstractDynamic selection of Web services at runtime is important for building flexible and loosely-coupled service-oriented applications. An abstract description of the required services is provided at design-time, and matching service offers are located at runtime. With the growing number of Web services that provide the same functionality but differ in quality parameters (e.g., availability, response time), a decision needs to be made on which services should be selected such that the user's end-to-end QoS requirements are satisfied. Although very efficient, local selection strategy fails short in handling global QoS requirements. Solutions based on global optimization, on the other hand, can handle global constraints, but their poor performance renders them inappropriate for applications with dynamic and realtime requirements. In this article we address this problem and propose a hybrid solution that combines global optimization with local selection techniques to benefit from the advantages of both worlds. The proposed solution consists of two steps: first, we use mixed integer programming (MIP) to find the optimal decomposition of global QoS constraints into local constraints. Second, we use distributed local selection to find the best Web services that satisfy these local constraints. The results of experimental evaluation indicate that our approach significantly outperforms existing solutions in terms of computation time while achieving close-to-optimal results. Mohammad Alrifai, Thomas Risse 0001, Wolfgang Nejdl |
ACM Trans. Web | 3 |
| 2011 | Analyzing Political Trends in the Blogosphere
Gianluca Demartini, Stefan Siersdorfer, Sergiu Chelaru, Wolfgang Nejdl |
ICWSM | 4 |
| 2011 | Incremental diversification for very large sets: a streaming-based approachabstractResult diversification is an effective method to reduce the risk that none of the returned results satisfies a user's query intention. It has been shown to decrease query abandonment substantially. On the other hand, computing an optimally diverse set is NP-hard for the usual objectives. Existing greedy diversification algorithms require random access to the input set, rendering them impractical in the context of large result sets or continuous data. Enrico Minack, Wolf Siberski, Wolfgang Nejdl |
SIGIR | 3 |
| 2011 | LinkDB: a probabilistic linkage database systemabstractEntity linkage deals with the problem of identifying whether two pieces of information represent the same real world object. The traditional methodology computes the similarity among the entities, and then merges those with similarity above some specific threshold. We demonstrate LinkDB, an original entity storage and querying system that deals with the entity linkage problem in a novel way. LinkDB is a probabilistic linkage database that uses existing linkage techniques to generate linkages among entities, but instead of performing the merges based on these linkages, it stores them alongside the data and performs only the required merges at run-time, by effectively taking into consideration the query specifications. We explain the technical challenges behind this kind of query answering, and we show how this new mechanism is able to provide answers that traditional entity linkage mechanisms cannot. Ekaterini Ioannou, Wolfgang Nejdl, Claudia Niederée, Yannis Velegrakis |
SIGMOD Conference | 2 |
| 2010 | Query Ranking in Information Integration
Rodolfo Stecher, Stefania Costache 0001, Claudia Niederée, Wolfgang Nejdl |
CAiSE | 4 |
| 2010 | Bringing order to your photos: event-driven classification of flickr images based on social knowledgeabstractWith the rapidly increasing popularity of Social Media sites, a lot of user generated content has been injected in the Web, thus resulting in a large amount of both multimedia items (music - Last.fm, MySpace.com, pictures - Flickr, Picasa, videos - YouTube) and textual data (tags and other text-based documents). As a consequence, especially for multimedia content it has become more and more difficult to find exactly the objects that best match the users' information needs. The methods we propose in this paper try to alleviate this problem and we focus on the domain of pictures, in particular on a subset of Flickr data. Many of the photos posted by users on Flickr have been shot during events and our methods aim to allow browsing and organization of picture collections in a natural way, by events. The algorithms we introduce in this paper exploit the social information produced by users in form of tags, titles and photo descriptions, for classifying pictures into different event categories. The extensive automated experiments demonstrate that our approach is very effective and opens new possibilities for multimedia retrieval, in particular image search. Moreover, the direct comparison with previous event detection algorithms confirm once more the quality of our methods. Claudiu S. Firan, Mihai Georgescu, Wolfgang Nejdl, Raluca Paiu |
CIKM | 3 |
| 2010 | Unsupervised public health event detection for epidemic intelligenceabstractRecent pandemics such as Swine Flu have caused concern for public health officials. Given the ever increasing pace at which infectious diseases can spread globally, officials must be prepared to react sooner and with greater epidemic intelligence gathering capabilities. However, state-of-the-art systems for Epidemic Intelligence have not kept the pace with the growing need for more robust public health event detection. In this paper, we propose a game-changing approach where public health events are detected in an unsupervised manner. We address the problems associated with adapting an unsupervised learner to the medical domain and in doing so, propose an approach which combines aspects from different feature-based event detection methods. We evaluate our approach with a real world dataset with respect to the quality of article clusters. Our results show that we are able to achieve a precision of 66% and a recall of 81% when evaluated using manually annotated, real-world data. This shows promising results for the use of such techniques in this new problem setting. Marco Fisichella, Avare Stewart, Kerstin Denecke, Wolfgang Nejdl |
CIKM | 4 |
| 2010 | Evaluating Evidences for Keyword Query Disambiguation in Entity Centric Database Search
Elena Demidova, Xuan Zhou 0001, Irina Oelze, Wolfgang Nejdl |
DEXA (2) | 4 |
| 2010 | Efficient Incremental Near Duplicate Detection Based on Locality Sensitive Hashing
Marco Fisichella, Fan Deng 0004, Wolfgang Nejdl |
DEXA (1) | 3 |
| 2010 | Efficient Semantic-Aware Detection of Near Duplicate Resources
Ekaterini Ioannou, Odysseas Papapetrou, Dimitrios Skoutas 0001, Wolfgang Nejdl |
ESWC (2) | 4 |
| 2010 | IQP: Incremental query construction, a probabilistic approachabstractThis paper presents IQP - a novel approach to bridge the gap between usability of keyword search and expressiveness of database queries. IQP enables a user to start with an arbitrary keyword query and incrementally refine it into a structured query through an interactive interface. The enabling techniques of IQP include: (1) a conceptual framework for incremental query construction; (2) a probabilistic model to assess the possible informational needs represented by a keyword query; (3) an algorithm to perform an optimal query construction. Elena Demidova, Xuan Zhou 0001, Wolfgang Nejdl |
ICDE | 3 |
| 2010 | LDA for on-the-fly auto taggingabstractIn this paper, we propose a method for automatic tagging sparse and short textual resources. In the presence of a new resource, our method creates an ad hoc corpus of related resources, then applies Latent Dirichlet Allocation (LDA) to elicit latent topics for the resource and the associated corpus. This is done in order to automatically tag the resource based on the most likely tags derived from the latent topics identified. We evaluate our method, using an offline analysis on publicly available BibSonomy dataset and an online study, showing its effectiveness. Ernesto Diaz-Aviles, Mihai Georgescu, Avare Stewart, Wolfgang Nejdl |
RecSys | 4 |
| 2010 | DivQ: diversification for keyword search over structured databasesabstractKeyword queries over structured databases are notoriously ambiguous. No single interpretation of a keyword query can satisfy all users, and multiple interpretations may yield overlapping results. This paper proposes a scheme to balance the relevance and novelty of keyword search results over structured databases. Firstly, we present a probabilistic model which effectively ranks the possible interpretations of a keyword query over structured data. Then, we introduce a scheme to diversify the search results by re-ranking query interpretations, taking into account redundancy of query results. Finally, we propose α-nDCG-W and WS-recall, an adaptation of α-nDCG and S-recall metrics, taking into account graded relevance of subtopics. Our evaluation on two real-world datasets demonstrates that search results obtained using the proposed diversification algorithms better characterize possible answers available in the database than the results of the initial relevance ranking. Elena Demidova, Peter Fankhauser, Xuan Zhou 0001, Wolfgang Nejdl |
SIGIR | 4 |
| 2010 | Boilerplate detection using shallow text featuresabstractIn addition to the actual content Web pages consist of navi-gational elements, templates, and advertisements. This boil-erplate text typically is not related to the main content, may deteriorate search precision and thus needs to be detected properly. In this paper, we analyze a small set of shallow text features for classifying the individual text elements in a Web page. We compare the approach to complex, state-of-the-art techniques and show that competitive accuracy can be achieved, at almost no cost. Moreover, we derive a simple and plausible stochastic model for describing the boilerplate creation process. With the help of our model, we also quantify the impact of boilerplate removal to re-trieval performance and show significant improvements over the baseline. Finally, we extend the principled approach by straight-forward heuristics, achieving a remarkable accuracy. Christian Kohlschütter, Peter Fankhauser, Wolfgang Nejdl |
WSDM | 3 |
| 2010 | How useful are your comments?: analyzing and predicting youtube comments and comment ratingsabstractAn analysis of the social video sharing platform YouTube reveals a high amount of community feedback through comments for published videos as well as through meta ratings for these comments. In this paper, we present an in-depth study of commenting and comment rating behavior on a sample of more than 6 million comments on 67,000 YouTube videos for which we analyzed dependencies between comments, views, comment ratings and topic categories. In addition, we studied the influence of sentiment expressed in comments on the ratings for these comments using the SentiWordNet thesaurus, a lexical WordNet-based resource containing sentiment annotations. Finally, to predict community acceptance for comments not yet rated, we built different classifiers for the estimation of ratings for these comments. The results of our large-scale evaluations are promising and indicate that community feedback on already rated comments can help to filter new unrated comments or suggest particularly useful but still unrated comments. Stefan Siersdorfer, Sergiu Chelaru, Wolfgang Nejdl, José San Pedro |
WWW | 3 |
| 2010 | Cardinality estimation and dynamic length adaptation for Bloom filters
Odysseas Papapetrou, Wolf Siberski, Wolfgang Nejdl |
Distributed Parallel Databases | 3 |
| 2010 | Why finding entities in Wikipedia is difficult, sometimes
Gianluca Demartini, Claudiu S. Firan, Tereza Iofciu, Ralf Krestel, Wolfgang Nejdl |
Inf. Retr. | 5 |
| 2010 | On-the-Fly Entity-Aware Query Processing in the Presence of LinkageabstractEntity linkage is central to almost every data integration and data cleaning scenario. Traditional techniques use some computed similarity among data structure to perform merges and then answer queries on the merged data. We describe a novel framework for entity linkage with uncertainty. Instead of using the linkage information to merge structures a-priori, possible linkages are stored alongside the data with their belief value. A new probabilistic query answering technique is used to take the probabilistic linkage into consideration. The framework introduces a series of novelties: (i) it performs merges at run time based not only on existing linkages but also on the given query; (ii) it allows results that may contain structures not explicitly represented in the data, but generated as a result of a reasoning on the linkages; and (iii) enables an evaluation of the query conditions that spans across linked structures, offering a functionality not currently supported by any traditional probabilistic databases. We formally define the semantics, describe an efficient implementation and report on the findings of our experimental evaluation. Ekaterini Ioannou, Wolfgang Nejdl, Claudia Niederée, Yannis Velegrakis |
Proc. VLDB Endow. | 2 |
| 2010 | Bridging the gap between tagging and querying vocabularies: Analyses and applications for enhancing multimedia IR
Kerstin Bischoff, Claudiu S. Firan, Wolfgang Nejdl, Raluca Paiu |
J. Web Semant. | 3 |
| 2010 | Leveraging personal metadata for Desktop search: The Beagle++ system
Enrico Minack, Raluca Paiu, Stefania Costache 0001, Gianluca Demartini, Julien Gaugaz, Ekaterini Ioannou, Paul-Alexandru Chirita, Wolfgang Nejdl |
J. Web Semant. | 8 |
| 2009 | Social Knowledge-Driven Music Hit Prediction
Kerstin Bischoff, Claudiu S. Firan, Mihai Georgescu, Wolfgang Nejdl, Raluca Paiu |
ADMA | 4 |
| 2009 | Automatically Identifying Tag Types
Kerstin Bischoff, Claudiu S. Firan, Cristina Kadar, Wolfgang Nejdl, Raluca Paiu |
ADMA | 4 |
| 2009 | SUITS: Faceted User Interface for Constructing Structured Queries from Keywords
Elena Demidova, Xuan Zhou 0001, Gideon Zenz, Wolfgang Nejdl |
DASFAA | 4 |
| 2009 | Exploiting Flickr Tags and Groups for Finding Landmark Photos
Rabeeh Ayaz Abbasi, Sergey Chernov 0001, Wolfgang Nejdl, Raluca Paiu, Steffen Staab |
ECIR | 3 |
| 2009 | A Vector Space Model for Ranking Entities and Its Application to Expert Search
Gianluca Demartini, Julien Gaugaz, Wolfgang Nejdl |
ECIR | 3 |
| 2009 | Zerber+R: top-k retrieval from a confidential indexabstractPrivacy-preserving document exchange among collaboration groups in an enterprise as well as across enterprises requires techniques for sharing and search of access-controlled information through largely untrusted servers. In these settings search systems need to provide confidentiality guarantees for shared information while offering IR properties comparable to the ordinary search engines. Top-k is a standard IR technique which enables fast query execution on very large indexes and makes systems highly scalable. However, indexing access-controlled information for top-k retrieval is a challenging task due to the sensitivity of the term statistics used for ranking. Sergej Zerr, Daniel Olmedilla, Wolfgang Nejdl, Wolf Siberski |
EDBT | 3 |
| 2009 | How to Trace and Revise Identities
Julien Gaugaz, Jakub Zakrzewski 0002, Gianluca Demartini, Wolfgang Nejdl |
ESWC | 4 |
| 2009 | Benchmarking Fulltext Search Performance of RDF Stores
Enrico Minack, Wolf Siberski, Wolfgang Nejdl |
ESWC | 3 |
| 2009 | Latent dirichlet allocation for tag recommendationabstractTagging systems have become major infrastructures on the Web. They allow users to create tags that annotate and categorize content and share them with other users, very helpful in particular for searching multimedia content. However, as tagging is not constrained by a controlled vocabulary and annotation guidelines, tags tend to be noisy and sparse. Especially new resources annotated by only a few users have often rather idiosyncratic tags that do not reflect a common perspective useful for search. In this paper we introduce an approach based on Latent Dirichlet Allocation (LDA) for recommending tags of resources in order to improve search. Resources annotated by many users and thus equipped with a fairly stable and complete tag set are used to elicit latent topics to which new resources with only a few tags are mapped. Based on this, other tags belonging to a topic can be recommended for the new resource. Our evaluation shows that the approach achieves significantly better precision and recall than the use of association rules, suggested in previous work, and also recommends more specific tags. Moreover, extending resources with these recommended tags significantly improves search for new resources. Ralf Krestel, Peter Fankhauser, Wolfgang Nejdl |
RecSys | 3 |
| 2009 | Pharos: an audiovisual search platformabstract841 Alessandro Bozzon, Marco Brambilla 0001, Piero Fraternali, Francesco Nucci, Stefan Debald, Eric Moore, Wolfgang Nejdl, Michel Plu, Patrick Aichroth, Olli Pihlajamaa, Cyril Laurier, Serge Zagorac, Gerhard Backfried, Daniel Weinland, Vincenzo Croce |
SIGIR | 7 |
| 2009 | Improving music genre classification using collaborative tagging dataabstractAs a fundamental and critical component of music information retrieval (MIR) systems, music genre classification has attracted considerable research attention. Automatically classifying music by genre is, however, a challenging problem due to the fact that music is an evolving art. While most of the existing work categorizes music using features extracted from music audio signals, in this paper, we propose to exploit the semantic information embedded in tags supplied by users of social networking websites. Particularly, we consider the tag information by creating a graph of tracks so that tracks are neighbors if they are similar in terms of their associated tags. Two classification methods based on the track graph are developed. The first one employs a classification scheme which simultaneously considers the audio content and neighborhood of tracks. In contrast, the second one is a two-level classifier which initializes genre label for unknown tracks using their audio content, and then iteratively updates the genres considering the influence from their neighbors. A set of optimizing strategies are designed for the purpose of further enhancing the quality of the two-level classifier. Extensive experiments are conducted on real-world data collected from Last.fm. Promising experimental results demonstrate the benefit of using tags for accurate music genre classification. Ling Chen 0006, Phillip Wright, Wolfgang Nejdl |
WSDM | 3 |
| 2009 | COWES: Web user clustering based on evolutionary web sessions
Ling Chen 0006, Sourav S. Bhowmick, Wolfgang Nejdl |
Data Knowl. Eng. | 3 |
| 2009 | How valuable is medical social media data? Content analysis of the medical web
Kerstin Denecke, Wolfgang Nejdl |
Inf. Sci. | 2 |
| 2009 | NEAR-Miner: Mining Evolution Associations of Web Site Directories for Efficient Maintenance of Web ArchivesabstractWeb archives preserve the history of autonomous Web sites and are potential gold mines for all kinds of media and business analysts. The most common Web archiving technique uses crawlers to automate the process of collecting Web pages. However, (re)downloading entire collection of pages periodically from a large Web site is unfeasible. In this paper, we take a step towards addressing this problem. We devise a data mining-driven policy for selectively (re)downloading Web pages that are located in hierarchical directory structures which are believed to have changed significantly (e.g., a substantial percentage of pages are inserted to/removed from the directory). Consequently, there is no need to download and maintain pages that have not changed since the last crawl as they can be easily retrieved from the archive. In our approach, we propose an off-line data mining algorithm called near- Miner that analyzes the evolution history of Web directory structures of the original Web site stored in the archive and mines negatively correlated association rules (near) between ancestor-descendant Web directories. These rules indicate the evolution correlations between Web directories. Using the discovered rules, we propose an efficient Web archive maintenance algorithm called warm that optimally skips the subdirectories (during the next crawl) which are negatively correlated with it in undergoing significant changes. Our experimental results with real data show that our approach improves the efficiency of the archive maintenance process significantly while sacrificing slightly in keeping the "freshness" of the archives. Furthermore, our experiments demonstrate that it is not necessary to discover nears frequently as the mining rules can be utilized effectively for archive maintenance over multiple versions. Ling Chen 0006, Sourav S. Bhowmick, Wolfgang Nejdl |
Proc. VLDB Endow. | 3 |
| 2009 | From keywords to semantic queries - Incremental query construction on the semantic web
Gideon Zenz, Xuan Zhou 0001, Enrico Minack, Wolf Siberski, Wolfgang Nejdl |
J. Web Semant. | 5 |
| 2008 | Probabilistic Entity Linkage for Heterogeneous Information Spaces
Ekaterini Ioannou, Claudia Niederée, Wolfgang Nejdl |
CAiSE | 3 |
| 2008 | Can all tags be used for search?abstractCollaborative tagging has become an increasingly popular means for sharing and organizing Web resources, leading to a huge amount of user generated metadata. These tags represent quite a few different aspects of the resources they describe and it is not obvious whether and how these tags or subsets of them can be used for search. This paper is the first to present an in-depth study of tagging behavior for very different kinds of resources and systems - Web pages (Del.icio.us), music (Last.fm), and images (Flickr) - and compares the results with anchor text characteristics. We analyze and classify sample tags from these systems, to get an insight into what kinds of tags are used for different resources, and provide statistics on tag distributions in all three tagging environments. Since even relevant tags may not add new information to the search procedure, we also check overlap of tags with content, with metadata assigned by experts and from other sources. We discuss the potential of different kinds of tags for improving search, comparing them with user queries posted to search engines as well as through a user survey. The results are promising and provide more insight into both the use of different kinds of tags for improving search and possible extensions of tagging systems to support the creation of potentially search-relevant tags. Kerstin Bischoff, Claudiu S. Firan, Wolfgang Nejdl, Raluca Paiu |
CIKM | 3 |
| 2008 | A densitometric approach to web page segmentationabstractWeb Page segmentation is a crucial step for many applications in Information Retrieval, such as text classification, de-duplication and full-text search. In this paper we describe a new approach to segment HTML pages, building on methods from Quantitative Linguistics and strategies borrowed from the area of Computer Vision. We utilize the notion of text-density as a measure to identify the individual text segments of a web page, reducing the problem to solving a 1D-partitioning task. The distribution of segment-level text density seems to follow a negative hypergeometric distribution, described by Frumkina's Law. Our extensive evaluation confirms the validity and quality of our approach and its applicability to the Web. Christian Kohlschütter, Wolfgang Nejdl |
CIKM | 2 |
| 2008 | Wildcards for lightweight information integration in virtual desktopsabstractWe present a flexible information integration approach which addresses the dynamic integration needs in a personal desktop environment where only partial mappings are defined between the sources to be integrated. Our approach is based on query rewriting using substitution rules. In addition to exploiting defined mappings, we employ substitution strategies, which are inspired by the idea of using wildcards in querying and filtering tasks. Starting from a triple based query language as used for querying RDF data, unmapped ontological elements are substituted in a controlled way with variables, leading to a controlled form of query relaxation. In addition, the approach also provides evidences for refining the existing mapping based on the results of executing the relaxed queries. Different strategies for replacing non-matched ontology elements with variables are presented and evaluated over real-world data sets. Rodolfo Stecher, Claudia Niederée, Wolfgang Nejdl |
CIKM | 3 |
| 2008 | Ranking Categories for Web Search
Gianluca Demartini, Paul-Alexandru Chirita, Ingo Brunkhorst, Wolfgang Nejdl |
ECIR | 4 |
| 2008 | Zerber: r-confidential indexing for distributed documentsabstractTo carry out work assignments, small groups distributed within a larger enterprise often need to share documents among themselves while shielding those documents from others’ eyes. In this situation, users need an indexing facility that can quickly locate relevant documents that they are allowed to access, without (1) leaking information about the remaining documents, (2) imposing a large management burden as users, groups, and documents evolve, or (3) requiring users to agree on a central completely trusted authority. To address this problem, we propose the concept of r-confidentiality, which captures the degree of information leakage from an index about the terms contained in inaccessible documents. Then we propose the r-confidential Zerber indexing facility for sensitive documents, which uses secret splitting and term merging to provide tunable limits on information leakage, even under statistical attacks; requires only limited trust in a central indexing authority; and is extremely easy to use and administer. Experiments with real-world data show that Zerber offers excellent performance for index insertions and lookups while requiring only a modest amount of storage space and network bandwidth Sergej Zerr, Elena Demidova, Daniel Olmedilla, Wolfgang Nejdl, Marianne Winslett, Soumyadeb Mitra |
EDBT | 4 |
| 2008 | DECK: Detecting Events from Web Click-Through DataabstractIn the past few years, there has been increased research interest in detecting previously unidentified events from Web resources. Our focus in this paper is to detect events from the click-through data generated by Web search engines. Existing event detection algorithms, which mainly study the news archive data, cannot be employed directly because of the following two unique features of click-through data: 1) the information provided by click-through data is quite limited; 2) not every query issued to a Web search engine corresponds to an event in the real world. In this paper, we address this problem by proposing an effective algorithm which Detects Events from ClicK-through data DECK. We firstly transform click-through data to the 2D polar space by considering the semantic dimension and temporal dimension of queries. Robust subspace estimation is performed to detect subspaces such that each subspace consists of queries of similar semantics. Next, we prune uninteresting subspaces which do not contain queries corresponding to real events by simultaneously considering the respective distribution of queries along the semantic dimension and the temporal dimension in each subspace. Finally, events are detected from interesting subspaces using a nonparametric clustering technique. Compared with an existing approach, our experimental results based on real-life data have shown that the proposed approach is more accurate and effective in detecting real events from click-through data. Ling Chen 0006, Yiqun Hu, Wolfgang Nejdl |
ICDM | 3 |
| 2008 | Attack resistant collaborative filteringabstractThe widespread deployment of recommender systems has lead to user feedback of varying quality. While some users faithfully express their true opinion, many provide noisy ratings which can be detrimental to the quality of the generated recommendations. The presence of noise can violate modeling assumptions and may thus lead to instabilities in estimation and prediction. Even worse, malicious users can deliberately insert attack profiles in an attempt to bias the recommender system to their benefit. While previous research has attempted to study the robustness of various existing Collaborative Filtering (CF) approaches, this remains an unsolved problem. Approaches such as Neighbor Selection algorithms, Association Rules and Robust Matrix Factorization have produced unsatisfactory results. This work describes a new collaborative algorithm based on SVD which is accurate as well as highly stable to shilling. This algorithm exploits previously established SVD based shilling detection algorithms, and combines it with SVD based-CF. Experimental results show a much diminished effect of all kinds of shilling attacks. This work also offers significant improvement over previous Robust Collaborative Filtering frameworks. Bhaskar Mehta, Wolfgang Nejdl |
SIGIR | 2 |
| 2008 | Semantically Enhanced Entity Ranking
Gianluca Demartini, Claudiu S. Firan, Tereza Iofciu, Wolfgang Nejdl |
WISE | 4 |
| 2008 | Using subspace analysis for event detection from web click-through dataabstractAlthough most of existing research usually detects events by analyzing the content or structural information of Web documents, a recent direction is to study the usage data. In this paper, we focus on detecting events from Web click-through data generated by Web search engines. We propose a novel approach which effectively detects events from click-through data based on robust subspace analysis. We first transform click-through data to the 2D polar space. Next, an algorithm based on Generalized Principal Component Analysis (GPCA) is used to estimate subspaces of transformed data such that each subspace contains query sessions of similar topics. Then, we prune uninteresting subspaces which do not contain query sessions corresponding to real events by considering both the semantic certainty and the temporal certainty of query sessions in each subspace. Finally, various events are detected from interesting subspaces by utilizing a nonparametric clustering technique. Compared with existing approaches, our experimental results based on real-life click-through data have shown that the proposed approach is more accurate in detecting real events and more effective in determining the number of events. Ling Chen 0006, Yiqun Hu, Wolfgang Nejdl |
WWW | 3 |
| 2008 | Detecting image spam using visual features and near duplicate detectionabstractEmail spam is a much studied topic, but even though current email spam detecting software has been gaining a competitive edge against text based email spam, new advances in spam generation have posed a new challenge: image-based spam. Image based spam is email which includes embedded images containing the spam messages, but in binary format. In this paper, we study the characteristics of image spam to propose two solutions for detecting image-based spam, while drawing a comparison with the existing techniques. The first solution, which uses the visual features for classification, offers an accuracy of about 98%, i.e. an improvement of at least 6% compared to existing solutions. SVMs (Support Vector Machines) are used to train classifiers using judiciously decided color, texture and shape features. The second solution offers a novel approach for near duplication detection in images. It involves clustering of image GMMs (Gaussian Mixture Models) based on the Agglomerative Information Bottleneck (AIB) principle, using Jensen-Shannon divergence (JS) as the distance measure. Bhaskar Mehta, Saurabh Nangia, Manish Gupta 0006, Wolfgang Nejdl |
WWW | 4 |
| 2008 | Privacy preserving document indexing infrastructure for a distributed environmentabstractTo carry out work assignments, small groups distributed within a larger enterprise or collaborative community often need to share documents among themselves while shielding those documents from others' eyes. In this situation, users need an indexing facility that can quickly locate relevant documents that they are allowed to access, without (1) leaking information about the remaining documents, (2) imposing a large management burden as users, groups, and documents evolve, or (3) requiring users to agree on a central completely trusted authority. In order to achieve this aim user access levels and access control have to be reflected in the index structures and/or retrieval algorithms as well as in ranking the search results. My Ph.D. work focuses on building up an indexing infrastructure which supports confidential indexing, sharing and retrieval of unstructured information which is spread over a number of distributed access-controlled collections. In order to allow for effective and efficient indexing and retrieval in these settings, it considers aspects of confidentiality preservation within an outsourced inverted index, a DHT index structure in P2P networks as well as confidential top-k information retrieval. Sergej Zerr, Wolfgang Nejdl |
Proc. VLDB Endow. | 2 |
| 2008 | An environment for flexible advanced compensations of Web service transactionsabstractBusiness to business integration has recently been performed by employing Web service environments. Moreover, such environments are being provided by major players on the technology markets. Those environments are based on open specifications for transaction coordination. When a failure in such an environment occurs, a compensation can be initiated to recover from the failure. However, current environments have only limited capabilities for compensations, and are usually based on backward recovery. In this article, we introduce an environment to deal with advanced compensations based on forward recovery principles. We extend the existing Web service transaction coordination architecture and infrastructure in order to support flexible compensation operations. We use a contract-based approach, which allows the specification of permitted compensations at runtime. We introduce abstract service and adapter components, which allow us to separate the compensation logic from the coordination logic. In this way, we can easily plug in or plug out different compensation strategies based on a specification language defined on top of basic compensation activities and complex compensation types. Experiments with our approach and environment show that such an approach to compensation is feasible and beneficial. Additionally, we introduce a cost-benefit model to evaluate the proposed environment based on net value analysis. The evaluation shows in which circumstances the environment is economical. Peter Dolog, Wolfgang Nejdl |
ACM Trans. Web | 3 |
| 2007 | Personalizing PageRank-Based Ranking over Distributed Collections
Stefania Costache 0001, Wolfgang Nejdl, Raluca Paiu |
CAiSE | 2 |
| 2007 | Building a Desktop Search Test-Bed
Sergey Chernov 0001, Pavel Serdyukov, Paul-Alexandru Chirita, Gianluca Demartini, Wolfgang Nejdl |
ECIR | 5 |
| 2007 | Enhancing Expert Search Through Query Modeling
Pavel Serdyukov, Sergey Chernov 0001, Wolfgang Nejdl |
ECIR | 3 |
| 2007 | Integrating Databases, Search Engines and Web Applications: A Model-Driven Approach
Alessandro Bozzon, Tereza Iofciu, Wolfgang Nejdl, Sascha Tönnies |
ICWE | 3 |
| 2007 | Engineering Compensations in Web Service Environment
Peter Dolog, Wolfgang Nejdl |
ICWE | 3 |
| 2007 | Sub-Ontology Discovery for Adaptive Re-use
Rodolfo Stecher, Claudia Niederée, Wolfgang Nejdl, Paolo Bouquet |
iiWAS | 3 |
| 2007 | Robust collaborative filteringabstractThe widespread deployment of recommender systems has lead to user feedback of varying quality. While some users faithfully express their true opinion, many provide noisy ratings which can be detrimental to the quality of the generated recommendations. The presence of noise can violate modeling assumptions and may thus lead to instabilities in estimation and prediction. Even worse, malicious users can deliberately insert attack profiles in an attempt to bias the recommender system to their benefit. Robust statistics is an area within statistics where esti-mation methods have been developed that deteriorate more gracefully in the presence of unmodeled noise and slight departures from modeling assumptions. In this work, we study how such robust statistical methods, in particular M-estimators, can be used to generate stable recommendation even in the presence of noise and spam. To that extent, we present a Robust Matrix Factorization algorithm and study its stability. We conclude that M-estimators do not add significant stability to recommendation; however the pre-sented algorithm can outperform existing recommendation algorithms in its recommendation quality. Bhaskar Mehta, Thomas Hofmann 0001, Wolfgang Nejdl |
RecSys | 3 |
| 2007 | Lexical analysis for modeling web query reformulationabstractModeling Web query reformulation processes is still an unsolved problem. In this paper we argue that lexical analysis is highly beneficial for this purpose. We propose to use the variation in Query Clarity, as well as the Part-Of-Speech pattern transitions as indicators of user's search actions. Experiments with a log of 2.4 million queries showed our techniques to be more flexible than the current approaches, while also providing us with interesting insights into user's Web behavioral patterns. Alessandro Bozzon, Paul-Alexandru Chirita, Claudiu S. Firan, Wolfgang Nejdl |
SIGIR | 4 |
| 2007 | Personalized query expansion for the webabstractThe inherent ambiguity of short keyword queries demands for enhanced methods for Web retrieval. In this paper we propose to improve such Web queries by expanding them with terms collected from each user's Personal Information Repository, thus implicitly personalizing the search output. We introduce five broad techniques for generating the additional query keywords by analyzing user data at increasing granularity levels, ranging from term and compound level analysis up to global co-occurrence statistics, as well as to using external thesauri. Our extensive empirical analysis under four different scenarios shows some of these approaches to perform very well, especially on ambiguous queries, producing a very strong increase in the quality of the output rankings. Subsequently, we move this personalized search framework one step further and propose to make the expansion process adaptive to various features of each query. A separate set of experiments indicates the adaptive algorithms to bring an additional statistically significant improvement over the best static expansion approach. Paul-Alexandru Chirita, Claudiu S. Firan, Wolfgang Nejdl |
SIGIR | 3 |
| 2007 | Query relaxation using malleable schemasabstractIn contrast to classical databases and IR systems, real-world information systems have to deal increasingly with very vague and diverse structures for information management and storage that cannot be adequately handled yet. While current object-relational database systems require clear and unified data schemas, IR systems usually ignore the structured information completely. Malleable schemas, as recently introduced, provide a novel way to deal with vagueness, ambiguity and diversity by incorporating imprecise and overlapping definitions of data structures. In this paper, we propose a novel query relaxation scheme that enables users to find best matching information by exploiting malleable schemas to effectively query vaguely structured information. Our scheme utilizes duplicates in differently described data sets to discover the correlations within a malleable schema, and then uses these correlations to appropriately relax the users’ queries. In addition, it ranks results of the relaxed query according to their respective probability of satisfying the original query’s intent. We have implemented the scheme and conducted extensive experiments with real-world data to confirm its performance and practicality. Xuan Zhou 0001, Julien Gaugaz, Wolf-Tilo Balke, Wolfgang Nejdl |
SIGMOD Conference | 4 |
| 2007 | Mirror site maintenance based on evolution associations of web directoriesabstractMirroring Web sites is a well-known technique commonly used in the Web community. A mirror site should be updated frequently to ensure that it reflects the content of the original site. Existing mirroring tools apply page-level strategies to check each page of a site, which is inefficient and expensive. In this paper, we propose a novel site-level mirror maintenance strategy. Our approach studies the evolution of Web directorystructures and mines association rules between ancestor-descendant Web directories. Discovered rules indicate the evolution correlations between Web directories. Thus, when maintaining the mirror of a Web site (directory), we can optimally skipsubdirectories which are negatively correlated with it in undergoing significant changes. The preliminary experimental results show that our approach improves the efficiency of the mirror maintenance process significantly while sacrificing slightly in keeping the "freshness" of the mirrors. Ling Chen 0006, Sourav S. Bhowmick, Wolfgang Nejdl |
WWW | 3 |
| 2007 | P-TAG: large scale automatic generation of personalized annotation tags for the webabstractThe success of the Semantic Web depends on the availability of Web pages annotated with metadata. Free form metadata or tags, as used in social bookmarking and folksonomies, have become more and more popular and successful. Such tags are relevant keywords associated with or assigned to a piece of information (e.g., a Web page), describing the item and enabling keyword-based classification. In this paper we propose P-TAG, a method which automatically generates personalized tags for Web pages. Upon browsing a Web page, P-TAG produces keywords relevant both to its textual content, but also to the data residing on the surfer's Desktop, thus expressing a personalized viewpoint. Empirical evaluations with several algorithms pursuing this approach showed very promising results. We are therefore very confident that such a user oriented automatic tagging approach can provide large scale personalized metadata annotations as an important step towards realizing the Semantic Web. Paul-Alexandru Chirita, Stefania Costache 0001, Wolfgang Nejdl, Siegfried Handschuh |
WWW | 3 |
| 2007 | Utility analysis for topically biased PageRankabstractPageRank is known to be an efficient metric for computing general document importance in the Web. While commonly used as a one-size-fits-all measure, the ability to produce topically biased ranks has not yet been fully explored in detail. In particular, it was still unclear to what granularity of "topic" the computation of biased page ranks makes sense. In this paper we present the results of a thorough quantitative and qualitative analysis of biasing PageRank on Open Directory categories. We show that the MAP quality of Biased PageRank generally increases with the ODP level up to a certain point, thus sustaining the usage of more specialized categories to bias PageRank on, in order to improve topic specific search. Christian Kohlschütter, Paul-Alexandru Chirita, Wolfgang Nejdl |
WWW | 3 |
| 2006 | Summarizing local context to personalize global web searchabstractThe PC Desktop is a very rich repository of personal information, efficiently capturing user's interests. In this paper we propose a new approach towards an automatic personalization of web search in which the user specific information is extracted from such local desktops, thus allowing for an increased quality of user profiling, while sharing less private information with the search engine. More specifically, we investigate the opportunities to select personalized query expansion terms for web search using three different desktop oriented approaches: summarizing the entire desktop data, summarizing only the desktop documents relevant to each user query, and applying natural language processing techniques to extract dispersive lexical compounds from relevant desktop resources. Our experiments with the Google API showed at least the latter two techniques to produce a very strong improvement over current web search. Paul-Alexandru Chirita, Claudiu S. Firan, Wolfgang Nejdl |
CIKM | 3 |
| 2006 | Efficient Parallel Computation of PageRank
Christian Kohlschütter, Paul-Alexandru Chirita, Wolfgang Nejdl |
ECIR | 3 |
| 2006 | Semantic Web Policies - A Discussion of Requirements and Research Issues
Piero A. Bonatti, Claudiu Duma, Norbert E. Fuchs, Wolfgang Nejdl, Daniel Olmedilla, Joachim Peer, Nahid Shahmehri |
ESWC | 4 |
| 2006 | Beagle++: Semantically Enhanced Searching and Ranking on the Desktop
Paul-Alexandru Chirita, Stefania Costache 0001, Wolfgang Nejdl, Raluca Paiu |
ESWC | 3 |
| 2006 | Analyzing User Behavior to Rank Desktop Items
Paul-Alexandru Chirita, Wolfgang Nejdl |
SPIRE | 2 |
| 2006 | Site level noise removal for search enginesabstractThe currently booming search engine industry has determined many online organizations to attempt to artificially increase their ranking in order to attract more visitors to their web sites. In the same time, the growth of the web has also inherently generated several navigational hyperlink structures which have a negative impact on the importance measures employed by current search engines. In this paper we propose and evaluate algorithms for identifying all these noisy links over the web graph, may them be spam or simple relationships between real world entities represented by sites, replication of content, etc. Unlike prior work, we target a different type of noisy link structures, residing at the site level, instead of the page level. We thus investigate and annihilate site level mutual reinforcement relationships, abnormal support coming from one site towards another, as well as complex link alliances between web sites. Our experiments with the link database of the TodoBR search engine show a very strong increase in the quality of the output rankings after having applied our techniques. André Luiz da Costa Carvalho, Paul-Alexandru Chirita, Edleno Silva de Moura, Pável Calado, Wolfgang Nejdl |
WWW | 5 |
| 2005 | MailRank: using ranking for spam detectionabstractCan we use social networks to combat spam? This paper investigates the feasibility of MailRank, a new email ranking and classification scheme exploiting the social communication network created via email interactions. The underlying email network data is collected from the email contacts of all MailRank users and updated automatically based on their email activities to achieve an easy maintenance. MailRank is used to rate the sender address of arriving emails such that emails from trustworthy senders can be ranked and classified as spam or non-spam. The paper presents two variants: Basic MailRank computes a global reputation score for each email address, whereas in Personalized MailRank the score of each email address is different for each MailRank user. The evaluation shows that MailRank is highly resistant against spammer attacks, which obviously have to be considered right from the beginning in such an application scenario. MailRank also performs well even for rather sparse networks, i.e., where only a small set of peers actually take part in the ranking of email addresses. Paul-Alexandru Chirita, Jörg Diederich 0001, Wolfgang Nejdl |
CIKM | 3 |
| 2005 | Activity Based Metadata for Semantic Desktop Search
Paul-Alexandru Chirita, Rita Gavriloaie, Stefania Costache 0001, Wolfgang Nejdl, Raluca Paiu |
ESWC | 4 |
| 2005 | Ontology-Based Policy Specification and Management
Wolfgang Nejdl, Daniel Olmedilla, Marianne Winslett, Charles C. Zhang |
ESWC | 1 |
| 2005 | Progressive Distributed Top k Retrieval in Peer-to-Peer NetworksabstractQuery processing in traditional information management systems has moved from an exact match model to more flexible paradigms allowing cooperative retrieval by aggregating the database objects' degree of match for each different query predicate and returning the best matching objects only. In peer-to-peer systems such strategies are even more important, given the potentially large number of peers, which may contribute to the results. Yet current peer-to-peer research has barely started to investigate such approaches. In this paper we discuss the benefits of best match/top-k queries in the context of distributed peer-to-peer information infrastructures and show how to extend the limited query processing in current peer-to-peer networks by allowing the distributed processing of top-k queries, while maintaining a minimum of data traffic. Relying on a super-peer backbone organized in the HyperCuP topology we show how to use local indexes for optimizing the necessary query routing and how to process intermediate results in inner network nodes at the earliest possible point in time cutting down the necessary data traffic within the network. Our algorithm is based on dynamically collected query statistics only, no continuous index update processes are necessary, allowing it to scale easily to large numbers of peers, as well as dynamic additions/deletions of peers. We show our approach to always deliver correct result sets and to be optimal in terms of necessary object accesses and data traffic. Finally, we present simulation results for both static and dynamic network environments. Wolf-Tilo Balke, Wolfgang Nejdl, Wolf Siberski, Uwe Thaden |
ICDE | 2 |
| 2005 | The Personal Publication Reader
Fabian Abel, Robert Baumgartner, Adrian Brooks, Christian Enzi, Georg Gottlob, Nicola Henze, Marcus Herzog, Matthias Kriesell, Wolfgang Nejdl, Kai Tomaschewski |
ISWC | 9 |
| 2005 | Semantically Rich Recommendations in Social Networks for Sharing, Exchanging and Ranking Semantic Context
Stefania Costache 0001, Wolfgang Nejdl, Raluca Paiu |
ISWC | 2 |
| 2005 | Searching Dynamic Communities with Personal Indexes
Alexander Löser, Christoph Tempich, Bastian Quilitz, Wolf-Tilo Balke, Steffen Staab, Wolfgang Nejdl |
ISWC | 6 |
| 2005 | Using ODP metadata to personalize searchabstractThe Open Directory Project is clearly one of the largest collaborative efforts to manually annotate web pages. This effort involves over 65,000 editors and resulted in metadata specifying topic and importance for more than 4 million web pages. Still, given that this number is just about 0.05 percent of the Web pages indexed by Google, is this effort enough to make a difference? In this paper we discuss how these metadata can be exploited to achieve high quality personalized web search. First, we address this by introducing an additional criterion for web page ranking, namely the distance between a user profile defined using ODP topics and the sets of ODP topics covered by each URL returned in regular web search. We empirically show that this enhancement yields better results than current web search using Google. Then, in the second part of the paper, we investigate the boundaries of biasing PageRank on subtopics of the ODP in order to automatically extend these metadata to the whole web. Paul-Alexandru Chirita, Wolfgang Nejdl, Raluca Paiu, Christian Kohlschütter |
SIGIR | 2 |
| 2005 | Peer-Sensitive ObjectRank - Valuing Contextual Information in Social Networks
Andrei Damian, Wolfgang Nejdl, Raluca Paiu |
WISE | 2 |
| 2004 | Model-Driven Design of Web Applications with Client-Side Adaptation
Stefano Ceri, Peter Dolog, Maristella Matera, Wolfgang Nejdl |
ICWE | 4 |
| 2004 | How to Build Google2Google - An (Incomplete) Recipe
Wolfgang Nejdl |
ISWC | 1 |
| 2004 | Top-k Query Evaluation for Schema-Based Peer-to-Peer Networks
Wolfgang Nejdl, Wolf Siberski, Uwe Thaden, Wolf-Tilo Balke |
ISWC | 1 |
| 2004 | Finding Related Pages Using the Link Structure of the WWWabstractMost of the current algorithms for finding related pages are exclusively based on text corpora of the WWW or incorporate only authority or hub values of pages. In this paper, we present HubFinder, a new fast algorithm for finding related pages exploring the link structure of the Web graph. Its criterion for filtering output pages is "pluggable", depending on the user's interests, and may vary from global page ranks to text content, etc. We also introduce HubRank, a new ranking algorithm which gives a more complete view of page "importance" by biasing the authority measure of PageRank towards hub values of pages. Finally, we present an evaluation of these algorithms in order to prove their qualities experimentally. Paul-Alexandru Chirita, Daniel Olmedilla, Wolfgang Nejdl |
Web Intelligence | 3 |
| 2004 | Super-peer-based routing strategies for RDF-based peer-to-peer networks
Wolfgang Nejdl, Martin Wolpers, Wolf Siberski, Christoph Schmitz 0001, Mario T. Schlosser, Ingo Brunkhorst, Alexander Löser |
J. Web Semant. | 1 |
| 2003 | Information Integration in Schema-Based Peer-To-Peer Networks
Alexander Löser, Wolf Siberski, Martin Wolpers, Wolfgang Nejdl |
CAiSE | 4 |
| 2003 | Super-peer-based routing and clustering strategies for RDF-based peer-to-peer networksabstractRDF-based P2P networks have a number of advantages compared with simpler P2P networks such as Napster, Gnutella or with approaches based on distributed indices such as CAN and CHORD. RDF-based P2P networks allow complex and extendable descriptions of resources instead of fixed and limited ones, and they provide complex query facilities against these metadata instead of simple keyword-based searches.In previous papers, we have described the Edutella infrastructure and different kinds of Edutella peers implementing such an RDF-based P2P network. In this paper we will discuss these RDF-based P2P networks as a specific example of a new type of P2P networks, schema-based P2P networks, and describe the use of super-peer based topologies for these networks. Super-peer based networks can provide better scalability than broadcast based networks, and do provide perfect support for inhomogeneous schema-based networks, which support different metadata schemas and ontologies (crucial for the Semantic Web). Furthermore, as we will show in this paper, they are able to support sophisticated routing and clustering strategies based on the metadata schemas, attributes and ontologies used. Especially helpful in this context is the RDF functionality to uniquely identify schemas, attributes and ontologies. The resulting routing indices can be built using dynamic frequency counting algorithms and support local mediation and transformation rules, and we will sketch some first ideas for implementing these advanced functionalities as well. Wolfgang Nejdl, Martin Wolpers, Wolf Siberski, Christoph Schmitz 0001, Mario T. Schlosser, Ingo Brunkhorst, Alexander Löser |
WWW | 1 |
| 2002 | Integrating Schema-specific Native XML Repositories into a RDF-based E-Learning P2P Network
Changtao Qu, Wolfgang Nejdl, Holger Schinzel |
Dublin Core Conference | 2 |
| 2002 | Towards a Modification Exchange Language for Distributed RDF Repositories
Wolfgang Nejdl, Wolf Siberski, Bernd Simon, Julien Tane |
ISWC | 1 |
| 2002 | EDUTELLA: a P2P networking infrastructure based on RDFabstractMetadata for the World Wide Web is important, but metadata for Peer-to-Peer (P2P) networks is absolutely crucial. In this paper we discuss the open source project Edutella which builds upon metadata standards defined for the WWW and aims to provide an RDF-based metadata infrastructure for P2P applications, building on the recently announced JXTA Framework. We describe the goals and main services this infrastructure will provide and the architecture to connect Edutella Peers based on exchange of RDF metadata. As the query service is one of the core services of Edutella, upon which other services are built, we specify in detail the Edutella Common Data Model (ECDM) as basis for the Edutella query exchange language (RDF-QEL-i) and format implementing distributed queries over the Edutella network. Finally, we shortly discuss registration and mediation services, and introduce the prototype and application scenario for our current Edutella aware peers. Wolfgang Nejdl, Boris Wolf, Changtao Qu, Stefan Decker, Michael Sintek, Ambjörn Naeve, Mikael Nilsson, Matthias Palmér, Tore Risch |
WWW | 1 |
| 1993 | Evaluating Recursive Queries in Distributed DatabasesabstractThe execution of logic queries in a distributed database environment is studied. Conventional optimization strategies, such as the early evaluation of selection conditions and the clustering of processing to manipulate and exchange large sets of tuples, are redefined in view of the additional difficulties due to logic queries, in particular to recursive rules. In order to allow efficient processing of these logic queries, several program transformation techniques that attempt to minimize distribution costs based on the idea of semijoins and generalized semijoins in conventional databases are presented. Although local computation of semijoins is not possible for the general case, classes of programs are indicated for which these transformations succeed in producing set-oriented computation. Processes evaluating the recursive program in a distributed network are described, and an efficient method for testing the termination of the computation is developed. The approach is compared with sequential as well as dataflow-oriented evaluation.> Wolfgang Nejdl, Stefano Ceri, Gio Wiederhold |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1987 | Recursive Strategies for Answering Recursive Queries - The RQA/FQI Strategy
Wolfgang Nejdl |
VLDB | 1 |