VLDB 2026 Research / reviewers in the wild / expert
Ran Yu 0001
dblp:45/8919-1
· DBLP profile ↗
20ranked-venue papers in the field
4as first author
14since 2021 · last 2026
0000-0002-1619-3164ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 13 (3 first)Data Mining & Knowledge Discovery · 3Database Systems & Data Management · 2 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | INFUSE Workshop: The First Workshop on INFormation Access in Uncertainty ScEnarios
Alisa Rieger, Ran Yu 0001, Amir Ebrahimi Fard, Nicolas Mattis, Johanne R. Trippas |
ECIR (3) | 2 |
| 2025 | VISOR: VIsual Seizure Onset Detection PeRsonalized for Epilepsy Patients
Uttam Kumar 0002, Ran Yu 0001, Michael Wenzel, Elena Demidova |
PAKDD (2) | 2 |
| 2025 | IWILDS'25: The 5th International Workshop on Investigating Learning During Web SearchabstractWeb-based learning is evolving rapidly as traditional search engines are complemented by Large Language Models (LLMs) and other AI technologies. This evolution offers new opportunities, such as automated information synthesis and personalized learning experiences. However, this also presents new challenges, including the need for learners to be aware of potential biases and misinformation in AI-generated content, and to maintain focus and depth in their learning journeys. Anett Hoppe, Ran Yu 0001, Jiqun Liu, Nilavra Bhattacharya |
WSDM | 2 |
| 2024 | A Multimodal and Multitask Approach for Adaptive Geospatial Region Embeddings
Rajjat Dadwal, Ran Yu 0001, Elena Demidova |
PAKDD (5) | 2 |
| 2023 | SCANNER: A Spatio-temporal Correlation and Neighborhood-based Feature Enrichment for Traffic PredictionabstractAccurate traffic speed prediction is essential for road safety and effective traffic management. However, this task is challenging due to the complex spatial and temporal interactions. Current traffic speed prediction approaches typically focus on short-term patterns in a close spatial neighborhood and fail to fully exploit the potential of complex and more distant spatial and temporal interactions. In this paper, we propose SCANNER - a novel Spatio-temporal CorrelAtioN and Neighborhood-based feature EnRichment approach for traffic speed prediction. SCANNER explicitly captures the relationship between road segments at different times and brings additional contextual information into the prediction. Our evaluation on two real-world datasets demonstrates a clear and consistent advantage of the SCANNER approach for both short and long-term speed prediction over the state-of-the-art baselines. Steve Gounoue, Ran Yu 0001, Elena Demidova |
SIGSPATIAL/GIS | 2 |
| 2023 | Iterative Geographic Entity Alignment with Cross-AttentionabstractAbstract Aligning schemas and entities of community-created geographic data sources with ontologies and knowledge graphs is a promising research direction for making this data widely accessible and reusable for semantic applications. However, such alignment is challenging due to the substantial differences in entity representations and sparse interlinking across sources, as well as high heterogeneity of schema elements and sparse entity annotations in community-created geographic data. To address these challenges, we propose a novel cross-attention-based iterative alignment approach called IGEA in this paper. IGEA adopts cross-attention to align heterogeneous context representations across geographic data sources and knowledge graphs. Moreover, IGEA employs an iterative approach for schema and entity alignment to overcome annotation and interlinking sparsity. Experiments on real-world datasets from several countries demonstrate that our proposed approach increases entity alignment performance compared to baseline methods by up to 18% points in F1-score. IGEA increases the performance of the entity and tag-to-class alignment by 7 and 8% points in terms of F1-score, respectively, by employing the iterative method. Alishiba Dsouza, Ran Yu 0001, Moritz Windoffer, Elena Demidova |
ISWC | 2 |
| 2023 | Spatial Link Prediction with Spatial and Semantic EmbeddingsabstractAbstract Semantic geospatial applications, such as geographic question answering, have benefited from knowledge graphs incorporating information regarding geographic entities and their relations. However, one of the most critical limitations of geographic knowledge graphs is the lack of semantic relations between geographic entities. The most extensive knowledge graphs specifically tailored to geographic entities are extracted from unstructured sources, with these graphs often relying on datatype properties to describe the entities, resulting in a flat representation that lacks entity relationships. Therefore, predicting links between geographic entities is essential for advancing semantic geospatial applications. Existing neural link prediction methods for knowledge graphs typically rely on pre-existing entity relations, making them unsuitable for scenarios where such information is absent. In this paper, we tackle the challenge of predicting spatial links in sparsely interlinked knowledge graphs by introducing two novel approaches: supervised spatial link prediction (SSLP) and unsupervised inductive spatial link prediction (USLP). These approaches leverage the wealth of literal values in geographic knowledge graphs through spatial and semantic embeddings. To assess the effectiveness of our proposed methods, we conduct evaluations on the WorldKG geographic knowledge graph, which incorporates geospatial data extracted from OpenStreetMap. Our results demonstrate that the SSLP and USLP approaches substantially outperform state-of-the-art link prediction methods. Genivika Mann, Alishiba Dsouza, Ran Yu 0001, Elena Demidova |
ISWC | 3 |
| 2023 | Constructing and meta-evaluating state-aware evaluation metrics for interactive search systemsabstractAbstract Evaluation metrics such as precision, recall and normalized discounted cumulative gain have been widely applied in ad hoc retrieval experiments. They have facilitated the assessment of system performance in various topics over the past decade. However, the effectiveness of such metrics in capturing users’ in-situ search experience, especially in complex search tasks that trigger interactive search sessions, is limited. To address this challenge, it is necessary to adaptively adjust the evaluation strategies of search systems to better respond to users’ changing information needs and evaluation criteria. In this work, we adopt a taxonomy of search task states that a user goes through in different scenarios and moments of search sessions, and perform a meta-evaluation of existing metrics to better understand their effectiveness in measuring user satisfaction. We then built models for predicting task states behind queries based on in-session signals. Furthermore, we constructed and meta-evaluated new state-aware evaluation metrics. Our analysis and experimental evaluation are performed on two datasets collected from a field study and a laboratory study, respectively. Results demonstrate that the effectiveness of individual evaluation metrics varies across task states. Meanwhile, task states can be detected from in-session signals. Our new state-aware evaluation metrics could better reflect in-situ user satisfaction than an extensive list of the widely used measures we analyzed in this work in certain states. Findings of our research can inspire the design and meta-evaluation of user-centered adaptive evaluation metrics, and also shed light on the development of state-aware interactive search systems. Marco Markwald, Jiqun Liu, Ran Yu 0001 |
Inf. Retr. J. | 3 |
| 2022 | SaL-Lightning Dataset: Search and Eye Gaze Behavior, Resource Interactions and Knowledge Gain during Web SearchabstractThe emerging research field Search as Learning (SAL) investigates how the Web facilitates learning through modern information retrieval systems. SAL research requires significant amounts of data that capture both search behavior of users and their acquired knowledge in order to obtain conclusive insights or train supervised machine learning models. However, the creation of such datasets is costly and requires interdisciplinary efforts in order to design studies and capture a wide range of features. In this paper, we address this issue and introduce an extensive dataset based on a user study, in which 114 participants were asked to learn about the formation of lightning and thunder. Participants’ knowledge states were measured before and after Web search through multiple-choice questionnaires and essay-based free recall tasks. To enable future research in SAL-related tasks we recorded a plethora of features and person-related attributes. Besides the screen recordings, visited Web pages, and detailed browsing histories, a large number of behavioral features and resource features were monitored. We underline the usefulness of the dataset by describing three, already published, use cases. Christian Otto, Markus Rokicki, Georg Pardi, Wolfgang Gritz, Daniel Hienert, Ran Yu 0001, Johannes von Hoyer, Anett Hoppe, Stefan Dietze, Peter Holtz, Yvonne Kammerer, Ralph Ewerth |
CHIIR | 6 |
| 2022 | IWILDS'22 - Third International Workshop on Investigating Learning During Web SearchabstractSince its inception, the World Wide Web has become a major information source, consulted for a diversity of informational tasks. With an abundance of information available online, Web search engines have been a main entry point, supporting users in finding suitable Web content for ever more complex information needs. The IWILDS workshop series invites research on complex search activities related to human learning. It provides an interdisciplinary platform for the presentation and discussion of recent research on human learning on the Web, welcoming perspectives from computer & information science, education and psychology. Anett Hoppe, Ran Yu 0001, Jiqun Liu |
SIGIR | 2 |
| 2021 | WorldKG: A World-Scale Geographic Knowledge GraphabstractOpenStreetMap is a rich source of openly available geographic information. However, the representation of geographic entities, e.g., buildings, mountains, and cities, within OpenStreetMap is highly heterogeneous, diverse, and incomplete. As a result, this rich data source is hardly usable for real-world applications. This paper presents WorldKG - a new geographic knowledge graph aiming to provide a comprehensive semantic representation of geographic entities in OpenStreetMap. We describe the WorldKG knowledge graph, including its ontology that builds the semantic dataset backbone, the extraction procedure of the ontology and geographic entities from OpenStreetMap, and the methods to enhance entity annotation. We perform statistical and qualitative dataset assessment, demonstrating the large scale and high precision of the semantic geographic information in WorldKG. Alishiba Dsouza, Nicolas Tempelmeier, Ran Yu 0001, Simon Gottschalk 0001, Elena Demidova |
CIKM | 3 |
| 2021 | IWILDS'21: Second International Workshop on Learning During Web SearchabstractWeb search is one of the most ubiquitous online activities and often used as a starting point to learn, i. e., to acquire or extend one's knowledge about certain topics or procedures. When learning by searching the Web, individuals are confronted with an unprecedented amount of information in various forms and varying quality. Thus, successful learning on the Web requires high degrees of self-regulation and should be supported by the adequate design of search, recommendation, and training tools. This creates a highly interdisciplinary research area at the intersection of information retrieval, human-computer interaction, psychology, and educational sciences. Search as Learning (SAL) research examines the relationships between querying, navigation, media consumption behavior, and the learning outcomes during Web search, how they can be measured, predicted, and supported. Anett Hoppe, Ran Yu 0001, Irina R. Brich, Jiqun Liu |
CIKM | 2 |
| 2021 | State-Aware Meta-Evaluation of Evaluation Metrics in Interactive Information RetrievalabstractIn interactive IR (IIR), users often seek to achieve different goals (e.g. exploring a new topic, finding a specific known item) at different search iterations and thus may evaluate system performances differently. Without state-aware approach, it would be extremely difficult to simulate and achieve real-time adaptive search evaluation and recommendation. To address this gap, our work identifies users' task states from interactive search sessions and meta-evaluates a series of online and offline evaluation metrics under varying states based on a user study dataset consisting of 1548 unique query segments from 450 search sessions. Our results indicate that: 1) users' individual task states can be identified and predicted from search behaviors and implicit feedback; 2) the effectiveness of mainstream evaluation measures (measured based upon their respective correlations with user satisfaction) vary significantly across task states. This study demonstrates the implicit heterogeneity in user-oriented IR evaluation and connects studies on complex search tasks with evaluation techniques. It also informs future research on the design of state-specific, adaptive user models and evaluation metrics. Jiqun Liu, Ran Yu 0001 |
CIKM | 2 |
| 2021 | Topic-independent modeling of user knowledge in informational search sessionsabstractAbstract Web search is among the most frequent online activities. In this context, widespread informational queries entail user intentions to obtain knowledge with respect to a particular topic or domain. To serve learning needs better, recent research in the field of interactive information retrieval has advocated the importance of moving beyond relevance ranking of search results and considering a user’s knowledge state within learning oriented search sessions. Prior work has investigated the use of supervised models to predict a user’s knowledge gain and knowledge state from user interactions during a search session. However, the characteristics of the resources that a user interacts with have neither been sufficiently explored, nor exploited in this task. In this work, we introduce a novel set of resource-centric features and demonstrate their capacity to significantly improve supervised models for the task of predicting knowledge gain and knowledge state of users in Web search sessions. We make important contributions, given that reliable training data for such tasks is sparse and costly to obtain. We introduce various feature selection strategies geared towards selecting a limited subset of effective and generalizable features. Ran Yu 0001, Markus Rokicki, Ujwal Gadiraju, Stefan Dietze |
Inf. Retr. J. | 1 |
| 2020 | TweetsCOV19 - A Knowledge Base of Semantically Annotated Tweets about the COVID-19 PandemicabstractPublicly available social media archives facilitate research in the social sciences and provide corpora for training and testing a wide range of machine learning and natural language processing methods. With respect to the recent outbreak of the Coronavirus disease 2019 (COVID-19), online discourse on Twitter reflects public opinion and perception related to the pandemic itself as well as mitigating measures and their societal impact. Understanding such discourse, its evolution, and interdependencies with real-world events or (mis)information can foster valuable insights. On the other hand, such corpora are crucial facilitators for computational methods addressing tasks such as sentiment analysis, event detection, or entity recognition. However, obtaining, archiving, and semantically annotating large amounts of tweets is costly. In this paper, we describe TweetsCOV19, a publicly available knowledge base of currently more than 8 million tweets, spanning October 2019 - April 2020. Metadata about the tweets as well as extracted entities, hashtags, user mentions, sentiments, and URLs are exposed using established RDF/S vocabularies, providing an unprecedented knowledge base for a range of knowledge discovery tasks. Next to a description of the dataset and its extraction and annotation process, we present an initial analysis and use cases of the corpus. Dimitar Dimitrov 0002, Erdal Baran, Pavlos Fafalios, Ran Yu 0001, Xiaofei Zhu, Matthäus Zloch, Stefan Dietze |
CIKM | 4 |
| 2020 | IWILDS'20: The 1st International Workshop on Investigating Learning during Web SearchabstractWeb search is one of the most ubiquitous online activities and often used for learning purposes, i.e., to extend one's knowledge or skills about certain topics or procedures. The importance of learning as an outcome of Web search has been recognized in research at the intersection of information retrieval, human-computer interaction, psychology, and educational sciences. Search as Learning (SAL) research examines relationships between querying, navigation, and reading behavior during Web search and the resulting learning outcomes, and how they can be measured, predicted, and supported. IWILDS aims to provide a platform to the interdisciplinary SAL community, with the objective to bring together interested researchers, provide room for presentation and discussion of novel research insights, and to inspire future directions of SAL research. Anett Hoppe, Ran Yu 0001, Yvonne Kammerer, Ladislao Salmerón |
CIKM | 2 |
| 2018 | Analyzing Knowledge Gain of Users in Informational Search Sessions on the WebabstractWeb search is frequently used by people to acquire new knowledge and to satisfy learning-related objectives, but little is known about how a user»s knowledge evolves through the course of a search session. We present a study addressing the knowledge gain of users in informational search sessions. Using crowdsourcing, we recruited 500 distinct users and orchestrated real-world search sessions spanning 10 different topics and information needs. By using scientifically formulated knowledge tests we calibrated the knowledge of users before and after their search sessions, quantifying their knowledge gain. We investigated the impact of information needs on the search behavior and knowledge gain of users, revealing a significant effect of information need on user queries and navigational patterns, but no direct effect on the knowledge gain. Users on average exhibited a higher knowledge gain through search sessions pertaining to topics they were less familiar with. Ujwal Gadiraju, Ran Yu 0001, Stefan Dietze, Peter Holtz |
CHIIR | 2 |
| 2018 | Predicting User Knowledge Gain in Informational Search SessionsabstractWeb search is frequently used by people to acquire new knowledge and to satisfy learning-related objectives. In this context, informational search missions with an intention to obtain knowledge pertaining to a topic are prominent. The importance of learning as an outcome of web search has been recognized. Yet, there is a lack of understanding of the impact of web search on a user's knowledge state. Predicting the knowledge gain of users can be an important step forward if web search engines that are currently optimized for relevance can be molded to serve learning outcomes. In this paper, we introduce a supervised model to predict a user's knowledge state and knowledge gain from features captured during the search sessions. To measure and predict the knowledge gain of users in informational search sessions, we recruited 468 distinct users using crowdsourcing and orchestrated real-world search sessions spanning 11 different topics and information needs. By using scientifically formulated knowledge tests, we calibrated the knowledge of users before and after their search sessions, quantifying their knowledge gain. Our supervised models utilise and derive a comprehensive set of features from the current state of the art and compare performance of a range of feature sets and feature selection strategies. Through our results, we demonstrate the ability to predict and classify the knowledge state and gain using features obtained during search sessions, exhibiting superior performance to an existing baseline in the knowledge state prediction task. Ran Yu 0001, Ujwal Gadiraju, Peter Holtz, Markus Rokicki, Philipp Kemkes, Stefan Dietze |
SIGIR | 1 |
| 2017 | FuseM: Query-Centric Data Fusion on Structured Web MarkupabstractEmbedded markup based on Microdata, RDFa, and Microformats have become prevalent on the Web and constitute an unprecedented source of data. However, RDF statements extracted from markup are fundamentally different from traditional RDF graphs: entity descriptions are flat, facts are highly redundant, and despite very frequent co-references explicit links are missing. Therefore, carrying out typical entity-centric tasks such as retrieval and summarisation cannot be tackled sufficiently with state-of-the-art methods and require preliminary data fusion. Given the scale and dynamics of Web markup, the applicability of general data fusion approaches is limited. We present a novel query-centric data fusion approach which overcomes such issues through a combination of entity retrieval and fusion techniques geared towards the specific challenges associated with embedded markup. To ensure precise and diverse entity descriptions, we follow a supervised learning approach and train a classifier for data fusion of a pool of candidate facts relevant to a given query and obtained through a preliminary entity retrieval step. We perform a thorough evaluation on a subset of the Web Data Commons dataset and show significant improvement over existing baselines. In addition, an investigation into the coverage and complementarity of facts from the constructed entity descriptions compared to DBpedia, shows potential for aiding tasks such as knowledge base population. Ran Yu 0001, Ujwal Gadiraju, Besnik Fetahu, Stefan Dietze |
ICDE | 1 |
| 2015 | Adaptive Focused Crawling of Linked Data
Ran Yu 0001, Ujwal Gadiraju, Besnik Fetahu, Stefan Dietze |
WISE (1) | 1 |