Philipp Mayr 0001

dblp:59/561 · also Philipp Mayr-Schlegel · DBLP profile ↗
← Back
38ranked-venue papers in the field
7as first author
13since 2021 · last 2026
0000-0002-6656-1658ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 30 (4 first)Other / Interdisciplinary · 4 (2 first)Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1Big Data, Cloud & Distributed Data Systems · 1 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Cultural Analytics for Good: Building Inclusive Evaluation Frameworks for Historical IR
Suchana Datta, Dwaipayan Roy 0001, Derek Greene, Gerardine Meaney, Karen Wade, Philipp Mayr 0001
ECIR (3)6
2026 The Second International Workshop on Scholarly Information Access (SCOLIA 2026)
Ingo Frommholz, Christin Kreutz, Philipp Mayr 0001, Guillaume Cabanac
ECIR (3)3
2026 MIRA: An LLM-Assisted Benchmark for Multi-Category Integrated Retrieval
abstract
Users increasingly expect modern search systems to offer a unified interface that seamlessly retrieves information from diverse data sources and formats. However, current information retrieval (IR) evaluation benchmarks have not kept pace with this development, primarily due to the lack of test collections that represent the diversity of contemporary search domains. We address this critical gap with MIRA, a novel benchmark based on a large-scale social science search platform. MIRA is designed for category-aware ranking across heterogeneous categories – Publications, Research Data, Variables, and Instruments & Tools – within a single, unified evaluation framework. The proposed collection is distinctive in several ways: (1) it is built upon real user queries, providing a more realistic basis for evaluation; (2) it covers scholarly items from four distinct categories, enabling multi-faceted evaluation; and (3) it leverages a Large Language Model to generate topic descriptions and narratives, as well as for relevance assessment with respect to these topics, substantially reducing the labor and cost of test collection generation. We release this resource to benefit the community by providing a foundational testbed for the research on multi-faceted, category-aware, integrated, or cross-category information retrieval.
Mehmet Deniz Türkmen, Suchana Datta, Dwaipayan Roy 0001, Daniel Hienert, Philipp Mayr 0001, Derek Greene
SIGIR5
2025 Model Card Metadata Collection from Hugging Face to Foster Multidisciplinary AI Research: A Dataset
Muhammad Asif Suryani, Saurav Karmakar, Brigitte Mathiak, Philipp Mayr 0001
DATA4
2025 The First Workshop on Scholarly Information Access (SCOLIA)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne, Christin Kreutz
ECIR (5)2
2024 Bibliometric-Enhanced Information Retrieval: 14th International BIR Workshop (BIR 2024)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne
ECIR (5)2
2024 VADIS - A Variable Detection, Interlinking and Summarization System
Yavuz Selim Kartal, Muhammad Ahsan Shahid, Sotaro Takeshita, Tornike Tsereteli, Andrea Zielinski, Benjamin Zapilko, Philipp Mayr 0001
ECIR (5)7
2024 Patterns in the Growth and Thematic Evolution of Artificial Intelligence Research: A Study Using Bradford Distribution of Productivity and Path Analysis
abstract
Artificial intelligence (AI) has emerged as a transformative technology with applications across multiple domains. The corpus of work related to the field of AI has grown significantly in volume as well as in terms of the application of AI in wider domains. However, given the wide application of AI in diverse areas, the measurement and characterization of the span of AI research is often a challenging task. Bibliometrics is a well-established method in the scientific community to measure the patterns and impact of research. It however has also received significant criticism for its overemphasis on the macroscopic picture and the inability to provide a deep understanding of growth and thematic structure of knowledge-creation activities. Therefore, this study presents a framework comprising of two techniques, namely, Bradford’s distribution and path analysis to characterize the growth and thematic evolution of the discipline. While the Bradford distribution provides a macroscopic view of artificial intelligence research in terms of patterns of growth, the path analysis method presents a microscopic analysis of the thematic evolutionary trajectories, thereby completing the analytical framework. Detailed insights into the evolution of each subdomain are drawn, major techniques employed in various AI applications are identified, and some relevant implications are discussed to demonstrate the usefulness of the analyses.
Solanki Gupta, Anurag Kanaujia, Hiran H. Lathabai, Vivek Kumar Singh 0001, Philipp Mayr 0001
Int. J. Intell. Syst.5
2024 An editorial of "AI + informetrics": Robust models for large-scale analytics
Yi Zhang 0095, Philipp Mayr 0001, Arho Suominen, Ying Ding 0001
Inf. Process. Manag.3
2023 Bibliometric-Enhanced Information Retrieval: 13th International BIR Workshop (BIR 2023)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne
ECIR (3)2
2022 Bibliometric-enhanced Information Retrieval: 12th International BIR Workshop (BIR 2022)
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne
ECIR (2)2
2021 Bibliometric-Enhanced Information Retrieval: 11th International BIR Workshop
Ingo Frommholz, Philipp Mayr 0001, Guillaume Cabanac, Suzan Verberne
ECIR (2)2
2021 BiblioDAP'21: The 1st Workshop on Bibliographic Data Analysis and Processing
abstract
Automatic processing of bibliographic data becomes very important in digital libraries, data science and machine learning due to its importance in keeping pace with the significant increase of published papers every year from one side and to the inherent challenges from the other side. This processing has several aspects including but not limited to I) Automatic extraction of references from PDF documents, II) Building an accurate citation graph, III) Author name disambiguation, etc. Bibliographic data is heterogeneous by nature and occurs in both structured (e.g. citation graph) and unstructured (e.g. publications) formats. Therefore, it requires data science and machine learning techniques to be processed and analysed. Here we introduce BiblioDAP'21: The 1st Workshop on Bibliographic Data Analysis and Processing.
Zeyd Boukhers, Philipp Mayr 0001, Silvio Peroni
KDD2
2020 Bibliometric-Enhanced Information Retrieval 10th Anniversary Workshop Edition
Guillaume Cabanac, Ingo Frommholz, Philipp Mayr 0001
ECIR (2)3
2020 Characteristics of Dataset Retrieval Sessions: Experiences from a Real-Life Digital Library
Zeljko Carevic, Dwaipayan Roy 0001, Philipp Mayr 0001
TPDL3
2020 The OpenCitations Data Model
abstract
A variety of schemas and ontologies are currently used for the machine-readable description of bibliographic entities and citations. This diversity, and the reuse of the same ontology terms with different nuances, generates inconsistencies in data. Adoption of a single data model would facilitate data integration tasks regardless of the data supplier or context application. In this paper we present the OpenCitations Data Model (OCDM), a generic data model for describing bibliographic entities and citations, developed using Semantic Web technologies. We also evaluate the effective reusability of OCDM according to ontology evaluation practices, mention existing users of OCDM, and discuss the use and impact of OCDM in the wider open science community.
Marilena Daquino, Silvio Peroni, David M. Shotton, Giovanni Colavizza, Behnam Ghavimi, Anne Lauscher, Philipp Mayr 0001, Matteo Romanello, Philipp Zumstein
ISWC (2)7
2019 Bibliometric-Enhanced Information Retrieval: 8th International BIR Workshop
Guillaume Cabanac, Ingo Frommholz, Philipp Mayr 0001
ECIR (2)3
2019 An Evaluation of the Effect of Reference Strings and Segmentation on Citation Matching
Behnam Ghavimi, Wolfgang Otto 0002, Philipp Mayr 0001
TPDL3
2019 Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL 2019)
abstract
The deluge of scholarly publication poses a challenge for scholars find relevant research and policy makers to seek in-depth information and understand research impact. Information retrieval (IR), natural language processing (NLP) and bibliometrics could enhance scholarly search, retrieval and user experience, but their use in digital libraries is not widespread. To address this gap, we propose the 4th Joint Workshop on BIRNDL and the 5th CL-SciSumm Shared Task. We seek to foster collaboration among researchers in NLP, IR and Digital Libraries (DL), and to stimulate the development of new methods in NLP, IR, recommendation systems and scientometrics toward improved scholarly document understanding, analysis, and retrieval at scale.
Muthu Kumar Chandrasekaran, Philipp Mayr 0001, Michihiro Yasunaga, Dayne Freitag, Dragomir R. Radev, Min-Yen Kan
SIGIR2
2018 The Role of the Task Topic in Web Search of Different Task Types
abstract
When users are looking for information on the Web, they show different behavior for different task types, e.g., for fact finding vs. information gathering tasks. For example, related work in this area has investigated how this behavior can be measured and applied to distinguish between easy and difficult tasks. In this work, we look at the searcher's behavior in the domain of journalism for four different task types, and additionally, for two different topics in each task type. Search behavior is measured with a number of session variables and correlated to subjective measures such as task difficulty, task success and the usefulness of documents. We acknowledge prior results in this area that task difficulty is correlated to user effort and that easy and difficult tasks are distinguishable by session variables. However, in this work, we emphasize the role of the task topic - in and of itself - over parameters such as the search results and read content pages, dwell times, session variables and subjective measures such as task difficulty or task success. With this knowledge researchers should give more attention to the task topic as an important influence factor for user behavior.
Daniel Hienert, Matthew Mitsui, Philipp Mayr 0001, Chirag Shah 0001, Nicholas J. Belkin
CHIIR3
2018 Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL 2018)
abstract
The large scale of scholarly publications poses a challenge for scholars in information seeking and sensemaking. Information retrieval~(IR), bibliometric and natural language processing (NLP) techniques could enhance scholarly search, retrieval and user experience but are not yet widely used. To this purpose, we propose the third iteration of the Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL). The workshop is intended to stimulate IR, NLP researchers and Digital Library professionals to elaborate on new approaches in natural language processing, information retrieval, scientometrics, text mining and recommendation techniques that can advance the state-of-the-art in scholarly document understanding, analysis, and retrieval at scale. The BIRNDL workshop will incorporate multiple invited talks, paper sessions, a poster session and the 4th edition of the Computational Linguistics (CL) Scientific Summarization Shared Task.
Muthu Kumar Chandrasekaran, Kokil Jaidka, Philipp Mayr 0001
SIGIR3
2018 DATA: SEARCH'18 - Searching Data on the Web
abstract
This half day workshop explores challenges in data search, with a particular focus on data on the web. We want to stimulate an interdisciplinary discussion around how to improve the description, discovery, ranking and presentation of structured and semi-structured data, across data formats and domain applications. We welcome contributions describing algorithms and systems, as well as frameworks and studies in human data interaction. The workshop aims to bring together communities interested in making the web of data more discoverable, easier to search and more user friendly.
Paul Groth, Laura Koesten, Philipp Mayr 0001, Maarten de Rijke, Elena Simperl
SIGIR3
2017 A Complete Year of User Retrieval Sessions in a Social Sciences Academic Search Engine
Philipp Mayr 0001, Ameni Kacem
TPDL1
2017 Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL 2017)
abstract
The large scale of scholarly publications poses a challenge for scholars in information seeking and sensemaking. Bibliometrics, information retrieval (IR), text mining and NLP techniques could help in these search and look-up activities, but are not yet widely used. This workshop is intended to stimulate IR researchers and digital library professionals to elaborate on new approaches in natural language processing, information retrieval, scientometrics, text mining and recommendation techniques that can advance the state-of-the-art in scholarly document understanding, analysis, and retrieval at scale. The BIRNDL workshop at SIGIR 2017 will incorporate an invited talk, paper sessions and the third edition of the Computational Linguistics (CL) Scientific Summarization Shared Task.
Muthu Kumar Chandrasekaran, Kokil Jaidka, Philipp Mayr 0001
SIGIR3
2016 Bibliometric-Enhanced Information Retrieval: 3rd International BIR Workshop
Philipp Mayr 0001, Ingo Frommholz, Guillaume Cabanac
ECIR1
2016 Survey on High-level Search Activities Based on the Stratagem Level in Digital Libraries
Zeljko Carevic, Philipp Mayr 0001
TPDL2
2016 Evaluating Co-authorship Networks in Author Name Disambiguation for Common Names
Fakhri Momeni, Philipp Mayr 0001
TPDL2
2015 Bibliometric-Enhanced Information Retrieval: 2nd International BIR Workshop
Philipp Mayr 0001, Ingo Frommholz, Andrea Scharnhorst, Peter Mutschke
ECIR1
2014 Bibliometric-Enhanced Information Retrieval
Philipp Mayr 0001, Andrea Scharnhorst, Birger Larsen, Philipp Schaer, Peter Mutschke
ECIR1
2013 Bibliometric-enhanced retrieval models for big scholarly information systems
abstract
Bibliometric techniques are not yet widely used to enhance retrieval processes in digital libraries, although they offer value-added effects for users. In this paper we will explore how statistical modelling of scholarship, such as Bradfordizing or network analysis of coauthorship network, can improve retrieval services for specific communities, as well as for large, cross-domain large collections. This paper aims to raise awareness of the missing link between information retrieval (IR) and bibliometrics / scientometrics and to create a common ground for the incorporation of bibliometric-enhanced services into retrieval at the digital library interface.
Philipp Mayr 0001, Peter Mutschke
IEEE BigData1
2013 A framework for specific term recommendation systems
abstract
In this paper we present the IRSA framework that enables the automatic creation of search term suggestion or recommendation systems (TS). Such TS are used to operationalize interactive query expansion and help users in refining their information need in the query formulation phase. Our recent research has shown TS to be more effective when specific to a certain domain. The presented technical framework allows owners of Digital Libraries to create their own specific TS constructed via OAI-harvested metadata with very little effort.
Thomas Lüke, Philipp Schaer, Philipp Mayr 0001
SIGIR3
2012 Integrating Interactive Visualizations in the Search Process of Digital Libraries and IR Systems
Daniel Hienert, Frank Sawitzki, Philipp Schaer, Philipp Mayr 0001
ECIR4
2012 Improving Retrieval Results with Discipline-Specific Query Expansion
Thomas Lüke, Philipp Schaer, Philipp Mayr 0001
TPDL3
2012 Extending Term Suggestion with Author Names
Philipp Schaer, Philipp Mayr 0001, Thomas Lüke
TPDL2
2011 A Novel Combined Term Suggestion Service for Domain-Specific Digital Libraries
Daniel Hienert, Philipp Schaer, Johann Schaible, Philipp Mayr 0001
TPDL4
2010 Establishing a Multi-Thesauri-Scenario based on SKOS and Cross-Concordances
Philipp Mayr 0001, Benjamin Zapilko, York Sure-Vetter
Dublin Core Conference1
2008 Comparing Human and Automatic Thesaurus Mapping Approaches in the Agricultural Domain
Boris Lauser, Gudrun Johannsen, Caterina Caracciolo, Willem Robert van Hage, Johannes Keizer, Philipp Mayr 0001
Dublin Core Conference6
2008 Building a Terminology Network for Search: The KoMoHe Project
Philipp Mayr 0001, Vivien Petras
Dublin Core Conference1