Norbert Fuhr

dblp:f/NFuhr · DBLP profile ↗
← Back
73ranked-venue papers in the field
36as first author
9since 2021 · last 2026
0000-0002-0441-6949ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 67 (33 first)Database Systems & Data Management · 6 (3 first)
YearPublicationVenuePosition
2026 LISP - A Rich Interaction Dataset and Loggable Interactive Search Platform
Jana Isabelle Friese, Andreas Konstantin Kruff, Philipp Schaer, Norbert Fuhr, Nicola Ferro 0001
ECIR (4)4
2025 Towards Reproducibility of Interactive Retrieval Experiments: Framework and Case Study
Jana Isabelle Friese, Norbert Fuhr
ECIR (4)2
2024 Context-Driven Interactive Query Simulations Based on Generative Large Language Models
Björn Engelmann 0002, Timo Breuer 0002, Jana Isabelle Friese, Philipp Schaer, Norbert Fuhr
ECIR (2)5
2024 SIGIR 2024 Workshop on Simulations for Information Access (Sim4IA 2024)
abstract
Simulations in various forms have been used to evaluate information access systems, like search engines, recommender systems, or conversational agents. In the form of the Cranfield paradigm, a simulation setup is well-known in the IR community, but user simulations have recently gained interest. While user simulations help to reduce the complexity of evaluation experiments and help with reproducibility, they can also contribute to a better understanding of users. Building on recent developments in methods and toolkits, the Sim4IA workshop aims to bring together researchers and practitioners to form an interactive and engaging forum for discussions on the future perspectives of the field. An additional aim is to plan an upcoming TREC/CLEF campaign.
Philipp Schaer, Christin Kreutz, Krisztian Balog, Timo Breuer 0002, Norbert Fuhr
SIGIR5
2022 Detecting Significant Differences Between Information Retrieval Systems via Generalized Linear Models
abstract
Being able to compare Information Retrieval(IR) systems correctly is pivotal to improving their quality. Among the most popular tools for statistical significance testing, we list t-test and ANOVA that belong to the linear models family. Therefore, given the relevance of linear models for IR evaluation, a great effort has been devoted to studying how to improve them to better compare IR systems.
Guglielmo Faggioli, Nicola Ferro 0001, Norbert Fuhr
CIKM3
2022 Validating Simulations of User Query Variants
Timo Breuer 0002, Norbert Fuhr, Philipp Schaer
ECIR (1)2
2021 Starting Conversations with Search Engines - Interfaces that Elicit Natural Language Queries
abstract
Search systems on the Web rely on user input to generate relevant results. Since early information retrieval systems, users are trained to issue keyword searches and adapt to the language of the system. Recent research has shown that users often withhold detailed information about their initial information need, although they are able to express it in natural language. We therefore conduct a user study (N = 139) to investigate how four different design variants of search interfaces can encourage the user to reveal more information. Our results show that a chatbot-inspired search interface can increase the number of mentioned product attributes by 84% and promote natural language formulations by 139% in comparison to a standard search bar interface.
Andrea Papenmeier, Dagmar Kern, Daniel Hienert, Alfred Sliwa, Ahmet Aker, Norbert Fuhr
CHIIR6
2021 Dataset of Natural Language Queries for E-Commerce
abstract
Shopping online is more and more frequent in our everyday life. For e-commerce search systems, understanding natural language coming through voice assistants, chatbots or from conversational search is an essential ability to understand what the user really wants. However, evaluation datasets with natural and detailed information needs of product-seekers which could be used for research do not exist. Due to privacy issues and competitive consequences, only few datasets with real user search queries from logs are openly available. In this paper, we present a dataset of 3,540 natural language queries in two domains that describe what users want when searching for a laptop or a jacket of their choice. The dataset contains annotations of vague terms and key facts of 1,754 laptop queries. This dataset opens up a range of research opportunities in the fields of natural language processing and (interactive) information retrieval for product search.
Andrea Papenmeier, Dagmar Kern, Daniel Hienert, Alfred Sliwa, Ahmet Aker, Norbert Fuhr
CHIIR6
2021 A Transparent Logical Framework for Aspect-Oriented Product Ranking Based on User Reviews
Firas Sabbah, Norbert Fuhr
ECIR (1)2
2020 How to Measure the Reproducibility of System-oriented IR Experiments
abstract
Replicability and reproducibility of experimental results are primary concerns in all the areas of science and IR is not an exception. Besides the problem of moving the field towards more reproducible experimental practices and protocols, we also face a severe methodological issue: we do not have any means to assess when reproduced is reproduced. Moreover, we lack any reproducibility-oriented dataset, which would allow us to develop such methods.
Timo Breuer 0002, Nicola Ferro 0001, Norbert Fuhr, Maria Maistro, Tetsuya Sakai, Philipp Schaer, Ian Soboroff
SIGIR3
2020 Proof by Experimentation? Towards Better IR Research
abstract
The current fight against the COVID-19 pandemic illustrates the importance of proper scientific methods: Besides fake news lacking any factual evidence, reports on clinical trials with various drugs often yield contradicting results; here, only a closer look at the underlying empirical methodology can help in forming a clearer picture.
Norbert Fuhr
SIGIR1
2019 CENTRE@CLEF 2019
Nicola Ferro 0001, Norbert Fuhr, Maria Maistro, Tetsuya Sakai, Ian Soboroff
ECIR (2)2
2018 Personalised Session Difficulty Prediction in an Online Academic Search Engine
Vu Tran 0001, Norbert Fuhr
TPDL2
2017 Modeling Interactive Information Retrieval as a Stochastic Process
abstract
Stochastic models have a long history in IR, starting from probabilistic models for (non-interactive) IR, and later for evaluation measures like e.g. average precision or discounted cumulated gain. The probabilistic ranking principle defines the order in which a user should inspect the documents retrieved by the system, for maximizing the retrieval quality experienced by the user. In a similar way, for interactive IR, the interactive probabilistic ranking principle (IPRP) defines a framework for the optimum ordering of suggestions / possible actions to be considered by the user.
Norbert Fuhr
CHIIR1
2013 Integrating IR Technologies for Professional Search - (Full-Day Workshop)
Michail Salampasis, Norbert Fuhr, Allan Hanbury, Mihai Lupu, Birger Larsen, Henrik Strindberg
ECIR2
2012 Quantitative Analysis of Search Sessions Enhanced by Gaze Tracking with Dynamic Areas of Interest
Vu Tran 0001, Norbert Fuhr
TPDL2
2012 Salton award lecture: information retrieval as engineering science
abstract
No abstract available.
Norbert Fuhr
SIGIR1
2012 Using eye-tracking with dynamic areas of interest for analyzing interactive information retrieval
abstract
Based on a new framework for capturing dynamic areas of interest in eye-tracking, we model the user search process as a Markov-chain. The analysis indicates possible system improvements and yields parameter estimates for the Interactive Probability Ranking Principle (IPRP).
Vu Tran 0001, Norbert Fuhr
SIGIR2
2012 The optimum clustering framework: implementing the cluster hypothesis
Norbert Fuhr, Marc Lechtenfeld, Benno Stein 0001, Tim Gollub
Inf. Retr.1
2012 Decentralized Probabilistic Text Clustering
abstract
Text clustering is an established technique for improving quality in information retrieval, for both centralized and distributed environments. However, traditional text clustering algorithms fail to scale on highly distributed environments, such as peer-to-peer networks. Our algorithm for peer-to-peer clustering achieves high scalability by using a probabilistic approach for assigning documents to clusters. It enables a peer to compare each of its documents only with very few selected clusters, without significant loss of clustering quality. The algorithm offers probabilistic guarantees for the correctness of each document assignment to a cluster. Extensive experimental evaluation with up to 1 million peers and 1 million documents demonstrates the scalability and effectiveness of the algorithm.
Odysseas Papapetrou, Wolf Siberski, Norbert Fuhr
IEEE Trans. Knowl. Data Eng.3
2010 Evaluation of an Adaptive Search Suggestion System
Sascha Kriewel, Norbert Fuhr
ECIR2
2010 Text Clustering for Peer-to-Peer Networks with Probabilistic Guarantees
Odysseas Papapetrou, Wolf Siberski, Norbert Fuhr
ECIR3
2008 A probability ranking principle for interactive information retrieval
Norbert Fuhr
Inf. Retr.1
2007 A Decision-Theoretic Model for Decentralised Query Routing in Hierarchical Peer-to-Peer Networks
Henrik Nottelmann, Norbert Fuhr
ECIR2
2006 Generating Search Term Variants for Text Collections with Historic Spellings
Andrea Ernst-Gerlach, Norbert Fuhr
ECIR2
2006 Comparing Different Architectures for Query Routing in Peer-to-Peer Networks
Henrik Nottelmann, Norbert Fuhr
ECIR2
2006 Retrieval quality vs. effectiveness of specificity-oriented search in XML collections
Norbert Fuhr, Norbert Gövert
Inf. Retr.1
2006 Evaluating the effectiveness of content-oriented XML retrieval methods
Norbert Gövert, Norbert Fuhr, Mounia Lalmas-Roelleke, Gabriella Kazai
Inf. Retr.2
2006 Introduction to the special issue on XML retrieval
abstract
No abstract available.
Ricardo Baeza-Yates, Norbert Fuhr, Yoelle Maarek
ACM Trans. Inf. Syst.2
2005 Introduction to the Special Issue on INEX
Norbert Fuhr, Mounia Lalmas-Roelleke
Inf. Retr.1
2004 Applying the Divergence from Randomness Approach for Content-Only Search in XML Documents
Mohammad Abolhassani, Norbert Fuhr
ECIR2
2004 Combining CORI and the Decision-Theoretic Approach for Advanced Resource Selection
Henrik Nottelmann, Norbert Fuhr
ECIR2
2004 A report on the first year of the INitiative for the Evaluation of XML retrieval
abstract
Abstract The INitiative for the Evaluation of XML retrieval (INEX) aims at providing an infrastructure to evaluate the effectiveness of content‐oriented XML retrieval systems. To this end, in the first round of INEX in 2002, a test collection of real world XML documents along with a set of topics and respective relevance assessments have been created with the collaboration of 36 participating organizations. In this article, we provide an overview of the first round of the INEX initiative.
Gabriella Kazai, Mounia Lalmas-Roelleke, Norbert Fuhr, Norbert Gövert
J. Assoc. Inf. Sci. Technol.3
2004 XIRQL: An XML query language based on information retrieval concepts
abstract
XIRQL ("circle") is an XML query language that incorporates imprecision and vagueness for both structural and content-oriented query conditions. The corresponding uncertainty is handled by a consistent probabilistic model. The core features of XIRQL are (1) document ranking based on index term weighting, (2) specificity-oriented search for retrieving the most relevant parts of documents, (3) datatypes with vague predicates for dealing with specific types of content and (4) structural vagueness for vague interpretation of structural query conditions. A XIRQL database may contain several classes of documents, where all documents in a class conform to the same DTD; links between documents also are supported. XIRQL queries are translated into a path algebra, which can be processed by our HyREX retrieval engine.
Norbert Fuhr, Kai Großjohann
ACM Trans. Inf. Syst.1
2003 From Uncertain Inference to Probability of Relevance for Advanced IR Applications
Henrik Nottelmann, Norbert Fuhr
ECIR2
2003 Evaluating different methods of estimating retrieval quality for resource selection
abstract
In a federated digital library system, it is too expensive to query every accessible library. Resource selection is the task to decide to which libraries a query should be routed. Most existing resource selection algorithms compute a library ranking in a heuristic way. In contrast, the decision-theoretic framework (DTF) follows a different approach on a better theoretic foundation: It computes a selection which minimises the overall costs (e.g. retrieval quality, time, money) of the distributed retrieval. For estimating retrieval quality the recall-precision function is proposed. In this paper, we introduce two new methods: The first one computes the empirical distribution of the probabilities of relevance from a small library sample, and assumes it to be representative for the whole library. The second method assumes that the indexing weights follow a normal distribution, leading to a normal distribution for the document scores. Furthermore, we present the first evaluation of DTF by comparing this theoretical approach with the heuristical state-of-the-art system CORI; here we find that DTF outperforms CORI in most cases.
Henrik Nottelmann, Norbert Fuhr
SIGIR2
2003 From Retrieval Status Values to Probabilities of Relevance for Advanced IR Applications
Henrik Nottelmann, Norbert Fuhr
Inf. Retr.2
2002 Index compression vs. retrieval time of inverted files for XML documents
abstract
Query languages for retrieval of XML documents allow for conditions referring both to the content and the structure of documents. In this paper, we investigate two different approaches for reducing index space of inverted files for XML documents. First, we consider methods for compressing index entries. Second, we develop the new XS tree data structure which contains the structural description of a document in a rather compact form, such that these descriptions can be kept in main memory. Experimental results on two large XML document collections show that very high compression rates for indexes can be achieved, but any compression increases retrieval time. On the other hand, highly compressed indexes may be feasible for applications where storage is limited, such as in PDAs or E-book devices.
Norbert Fuhr, Norbert Gövert
CIKM1
2002 HyREX: hyper-media retrieval engine for XML
abstract
No abstract available.
Norbert Fuhr, Norbert Gövert, Kai Großjohann
SIGIR1
2001 Learning Probabilistic Datalog Rules for Information Classification and Transformation
abstract
Probabilistic Datalog is a combination of classical Datalog (i.e., function-free Horn clause predicate logic) with probability theory. Therefore, probabilistic weights may be attached to both facts and rules. But it is often impossible to assign exact rule weights or even to construct the rules themselves. Instead of specifying them manually, learning algorithms can be used to learn both rules and weights. In practice, these algorithms are very slow because they need a large example set and have to test a high number of rules. We apply a number of extensions to these algorithms in order to improve efficiency. Several applications demonstrate the power of learning probabilistic Datalog rules, showing that learning rules is suitable for low dimensional problems (e.g., schema mapping) but inappropriate for higher dimensions like e.g. in text classification.
Henrik Nottelmann, Norbert Fuhr
CIKM2
2001 XIRQL: A Query Language for Information Retrieval in XML Documents
abstract
Based on the document-centric view of XML, we present the query language XIRQL. Current proposals for XML query languages lack most IR-related features, which are weighting and ranking, relevance-oriented search, datatypes with vague predicates, and semantic relativism. XIRQL integrates these features by using ideas from logic-based probabilistic IR models, in combination with concepts from the database area. For processing XIRQL queries, a path algebra is presented, that also serves as a starting point for query optimization.
Norbert Fuhr, Kai Großjohann
SIGIR1
2000 Concepts for a Graphical User Interface for Hypermedia Retrieval
abstract
Which concepts are general and specific enough to support hypermedia retrieval? We present in this paper a graphical user interface that makes explicit the following three key concepts for hypermedia retrieval: content of documents, facts about documents and structure of documents. The underlying model is based on a probabilistic object-oriented logic (POOL) that enables content-based querying, fact-based querying, and the exploitation of the structural nature of hypermedia documents. In this paper, we focus on the graphical user interface which reflects the expressiveness of POOL for querying hypermedia documents. We report on the application of the system using a heterogeneous collection of hypermedia documents.
Mounia Lalmas-Roelleke, Thomas Roelleke, Frank Turra, Norbert Fuhr
FQAS4
2000 Probabilistic datalog: Implementing logical information retrieval for advanced applications
abstract
In the logical approach to information retrieval (IR), retrieval is considered as uncertain inference. Whereas classical IR models are based on propositional logic, we combine Datalog (function-free Horn clause predicate logic) with probability theory. Therefore, probabilistic weights may be attached to both facts and rules. The underlying semantics extends the well-founded semantics of modularly stratified Datalog to a possible worlds semantics. By using default independence assumptions with explicit specification of disjoint events, the inference process always yields point probabilities. We describe an evaluation method and present an implementation. This approach allows for easy formulation of specific retrieval models for arbitrary applications, and classical probabilistic IR models can be implemented by specifying the appropriate rules. In comparison to other approaches, the possibility of recursive rules allows for more powerful inferences, and predicate logic gives the expressiveness required for multimedia retrieval. Furthermore, probabilistic Datalog can be used as a query language for integrated information retrieval and database systems.
Norbert Fuhr
J. Am. Soc. Inf. Sci.1
1999 A Probabilistic Description-Oriented Approach for Categorizing Web Documents
abstract
The automatic categorisation of web documents is becoming crucial for organising the huge amount of information available in the Internet. We are facing a new challenge due to the fact that web documents have a rich structure and are highly heterogeneous. Two ways to respond to this challenge are (1) using a representation of the content of web documents that captures these two characteristics and (2) using more effective classifiers.
Norbert Gövert, Mounia Lalmas-Roelleke, Norbert Fuhr
CIKM3
1999 Towards Data Abstraction in Networked Information Retrieval Systems
Norbert Fuhr
Inf. Process. Manag.1
1999 A Decision-Theoretic Approach to Database Selection in Networked IR
abstract
In networked IR, a client submits a query to a broker, which is in contact with a large number of databases. In order to yield a maximum number of documents at minimum cost, the broker has to make estimates about the retrieval cost of each database, and then decide for each database whether or not to use it for the current query, and if, how many documents to retrieve from it. For this purpose, we develop a general decision-theoretic model and discuss different cost structures. Besides cost for retrieving relevant versus nonrelevant documents, we consider the following parameters for each database: expected retrieval quality, expected number of relevant documents in the database and cost factors for query processing and document delivery. For computing the overall optimum, a divide-and-conquer algorithm is given. If there are several brokers knowing different databases, a preselection of brokers can only be performed heuristically, but the computation of the optimum can be done similarily to the single-broker case. In addition, we derive a formula which estimates the number of relevant documents in a database based on dictionary information.
Norbert Fuhr
ACM Trans. Inf. Syst.1
1998 HySpirit - A Probabilistic Inference Engine for Hypermedia Retrieval in Large Databases
Norbert Fuhr, Thomas Roelleke
EDBT1
1998 Querying for Facts and Content in Hypermedia Documents
Thomas Roelleke, Norbert Fuhr
FQAS2
1998 DOLORES: A System for Logic-Based Retrieval of Multimedia Objects
abstract
We describe the design and implementation of a system for logic-based multimedia retrieval. As high-level logic for retrieval of hypermedia documents, we have developed a probabilistic object-oriented logic (POOL) which supports aggregated objects, different kinds of propositions (terms, classifications and attributes) and even rules as being contained in objects. Based on a probabilistic four-valued logic, POOL uses an implicit open world assumption, allows for closed world assumptions and is able to deal with inconsistent knowledge. POOL programs and queries are translated into probabilistic Datalog programs which can be interpreted by the HySpirit inference engine. For storing the multimedia data, we have developed a new basic IR engine which yields physical data abstraction. The overall architecture and the flexibility of each layer supports logic-based methods for multimedia information retrieval.
Norbert Fuhr, Norbert Gövert, Thomas Roelleke
SIGIR1
1997 A Probabilistic Relational Algebra for the Integration of Information Retrieval and Database Systems
abstract
We present a probabilistic relational algebra (PRA) which is a generalization of standard relational algebra. In PRA, tuples are assigned probabilistic weights giving the probability that a tuple belongs to a relation. Based on intensional semantics, the tuple weights of the result of a PRA expression always conform to the underlying probabilistic model. We also show for which expressions extensional semantics yields the same results. Furthermore, we discuss complexity issues and indicate possibilities for optimization. With regard to databases, the approach allows for representing imprecise attribute values, whereas for information retrieval, probabilistic document indexing and probabilistic search term weighting can be modeled. We introduce the concept of vague predicates which yield probabilistic weights instead of Boolean values, thus allowing for queries with vague selection conditions. With these features, PRA implements uncertainty and vagueness in combination with the relational model.
Norbert Fuhr, Thomas Roelleke
ACM Trans. Inf. Syst.1
1996 Object-Oriented and Database Concepts for the Design of Networked Information Retrieval Systems
abstract
By using data abstraction concepts from database and objectoriented systems and combining them with uncertain inference, we develop a new approach for the design of information retrieval (IR) systems, thus solving major problems with regard to networked IR. Different data types with vague predicates are required to allow for queries referring to arbitrary attributes of documents. Physical data independence helps in solving text search problems related to noun phrases, compound words and proper nouns. Logical data independence allows for different views on an IR database, thus implementing e.g. data security. Inheritance, especially on attributes, supports the creation of unified views on a set of IR databases. Uncertain inference allows for query processing even on incompatible database schemas. 1 Introduction Due to the rapid growth of the Internet, the development of powerful search tools is becoming a crucial issue for effective use of the sources that are available on the network....
Norbert Fuhr
CIKM1
1996 Network Information Retrieval (Workshop Abstract)
abstract
No abstract available.
Norbert Fuhr
SIGIR1
1996 Retrieval of Complex Objects Using a Four-Valued Logic
abstract
The aggregated structure of documents plays a key role in full-text, multimedia, and network Information Retrieval (IR). Considering aggregation provides new querying facilities and improves retrieval effectiveness. We present a knowledge representation for IR purposes which pays special attention to this aggregated structure of objects. In addition, further features of objects can be described. Thus, the structure of full-text documents, the heterogeneity and the spatial and temporal relationships of objects typical for multimedia IR, and meta information for network IR are representable within one integrated framework. The model we propose allows for querying on the content of documents (objects) as well as on other features. The query result may contain objects having different types. Instead of retrieving only whole documents, the retrieval process determines the least aggregated entities that imply the query. 1 Motivation and Background New IR applications like full-text, multime...
Thomas Roelleke, Norbert Fuhr
SIGIR2
1996 Retrieval Effectiveness of Proper Name Search Methods
Ulrich Pfeifer, Thomas Poersch, Norbert Fuhr
Inf. Process. Manag.3
1995 Probabilistic Datalog - A Logic For Powerful Retrieval Methods
abstract
ing with credit is permitted. To copy otherwise, to republish, to post on servers, or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from Publications Dept, ACM Inc., fax +1 (212) 869-0481, or 2 \\Delta ---Since Datalog P allows for recursive rules, it provides more powerful inference than any other (implemented) probabilistic IR model. ---Finally, since Datalog P is a generalization of (deterministic) Datalog, it can be used as a standard query language for both IR and database systems, and thus also for integration of these two types of systems on the logical level. 2. INFORMAL DESCRIPTION OF Datalog P Probabilistic Datalog is an extension of stratified Datalog (see e.g. [Ullman 88], [Ceri et al. 90]). On the syntactical level, the only difference is that with ground facts, also a probabilistic weight may be given, e.g. 0.7 indterm(d1,ir). 0.8 indterm(d1,db). Informally speaking, the probabilistic weight gives...
Norbert Fuhr
SIGIR1
1995 Efficient Processing of Vague Queries using a Data Stream Approach
abstract
Article Efficient processing of vague queries using a data stream approach Share on Authors: Ulrich Pfeifer University of Dortmund, Germany University of Dortmund, GermanyView Profile , Norbert Fuhr University of Dortmund, Germany University of Dortmund, GermanyView Profile Authors Info & Claims SIGIR '95: Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrievalJuly 1995 Pages 189–197https://doi.org/10.1145/215206.215359Online:01 July 1995Publication History 6citation320DownloadsMetricsTotal Citations6Total Downloads320Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Ulrich Pfeifer, Norbert Fuhr
SIGIR2
1994 Probabilistic Information Retrieval as a Combination of Abstraction, Inductive Learning, and Probabilistic Assumptions
abstract
We show that former approaches in probabilistic information retrieval are based on one or two of the three concepts abstraction, inductive learning , and probabilistic assumptions , and we propose a new approach which combines all three concepts. This approach is illustrated for the case of indexing with a controlled vocabulary. For this purpose, we describe a new probabilistic model first, which is then combined with logistic regression, thus yielding a generalization of the original model. Experimental results for the pure theoretical model as well as for heuristic variants are given. Furthermore, linear and logistic regression are compared.
Norbert Fuhr, Ulrich Pfeifer
ACM Trans. Inf. Syst.1
1993 A Probabilistic Relational Model for the Integration of IR and Databases
abstract
In this paper, a probabilistic relational model is presented which combines relational algebra with probabilistic retrieval. Based on certain independence assumptions, the operators of the relational algebra are redefined such that the probabilistic algebra is a generalization of the standard relational algebra. Furthermore, a special join operator implementing probabilistic retrieval is proposed. When applied to typical document databases, queries can not only ask for documents, but for any kind of object in the database. In addition, an implicit ranking of these objects is provided in case the query relates to probabilistic indexing or uses the probabilistic join operator. The proposed algebra is intended as a standard interface to combined database and IR systems, as a basis for implementing user-friendly interfaces.
Norbert Fuhr
SIGIR1
1992 Integration of Probabilistic Fact and Text Retrieval
abstract
In this paper, a model for combining text and fact retrieval is described. A query is a set of conditions, where a single condition is either a text or fact condition. Fact conditions can be interpreted as being vague, thus leading to nonbinary weights for fact conditions with respect to database objects. For text conditions, we use descriptions of the occurence of terms in documents instead of precomputed indexing weights, thus treating terms similar to attributes. Probabilistic indexing weights for conditions are computed by introducing the notion of correctness (or acceptability) of a condition w.r.t. an object. These indexing weights are used in retrieval for a probabilistic ranking of objects based on the retrieval for a probabilistic ranking of objects based on the retrieval-with-probabilistic-indexing (RPI) model, for which a new derivation is given here.
Norbert Fuhr
SIGIR1
1991 An Information Retrieval View of Environmental Information Systems
Norbert Fuhr
DEXA1
1991 Combining Model-Oriented and Description-Oriented Approaches for Probabilistic Indexing
abstract
We distinguish model-oriented and description-oriented approaches in probabilistic information retrieval.The former refer to certain representations of documents and
Norbert Fuhr, Ulrich Pfeifer
SIGIR1
1991 A Probabilistic Learning Approach for Document Indexing
abstract
We describe a method for probabilistic document indexing using relevance feedback data that has been collected from a set of queries.Our approach is based on three new concepts: (1) Abstraction from specific terms and documents, which overcomes the restriction of limited relevance information for parameter estimation.(2) Flexibility of the representation, which allows the integrationof new text analysis and knowledge-based methods in our approach as well as the consideration of document structures or different types of terms.(3) Probabilistic learning or classification methods for the estimation of the indexing weights making better use of the available relevance information, Our approach can be applied under restrictions that hold for real applications.We give experimental results for five test collections which show improvements over other methods.
Norbert Fuhr, Chris Buckley
ACM Trans. Inf. Syst.1
1990 Probabilistic Document Indexing from Relevance Feedback Data
abstract
Based on the binary independence indexing model, we apply three new concepts for probabilistic document indexing from relevance feedback data:
Norbert Fuhr, Chris Buckley
SIGIR1
1990 A Probabilistic Framework for Vague Queries and Imprecise Information in Databases
Norbert Fuhr
VLDB1
1989 Optimum Polynomial Retrieval Functions
abstract
We show that any approach to develop optimum retrieval functions is based on two kinds of assumptions: first, a certain form of representation for documents and requests, and second, additional simplifying assumptions that predefine the type of the retrieval function. Then we describe an approach for the development of optimum polynomial retrieval functions: request-document pairs (ƒl, dm) are mapped onto description vectors @@@@(ƒl, dm), and a polynomial function of the form @@@@T · @@@@(@@@@) is developed such that it yields estimates of the probability of relevance P(R|@@@@(ƒl, dm)) with minimum square errors. We give experimental results for the application of this approach to documents with weighted indexing as well as to documents with complex representations. In contrast to other probabilistic models, our approach yields estimates of the actual probabilities, it can handle very complex representations of documents and requests, and it can be easily applied to multi-valued relevance scales. On the other hand, this approach is not suited to log-linear probabilistic models, and it needs large samples of relevance feedback data for its application.
Norbert Fuhr
SIGIR1
1989 Models for retrieval with probabilistic indexing
Norbert Fuhr
Inf. Process. Manag.1
1989 Optimum probability estimation from empirical distributions
Norbert Fuhr, Hubert Hüther
Inf. Process. Manag.1
1989 Optimum Polynomial Retrieval Functions Based on the Probability Ranking Principle
abstract
We show that any approach to developing optimum retrieval functions is based on two kinds of assumptions: first, a certain form of representation for documents and requests, and second, additional simplifying assumptions that predefine the type of the retrieval function. Then we describe an approach for the development of optimum polynomial retrieval functions: request-document pairs ( f l , d m ) are mapped onto description vectors x ( f l , d m ), and a polynomial function e ( x ) is developed such that it yields estimates of the probability of relevance P( R | x ( f l , d m ) with minimum square errors. We give experimental results for the application of this approach to documents with weighted indexing as well as to documents with complex representations. In contrast to other probabilistic models, our approach yields estimates of the actual probabilities, it can handle very complex representations of documents and requests, and it can be easily applied to multivalued relevance scales. On the other hand, this approach is not suited to log-linear probabilistic models and it needs large samples of relevance feedback data for its application.
Norbert Fuhr
ACM Trans. Inf. Syst.1
1988 The Automatic Indexing System AIR/PHYS -- From Research to Application
abstract
Since October 1985, the automatic indexing system AIR/PHYS has been used in the input production of the physics data base of the Fachinformationsentrum Karlsruhe/West Germany. The texts to be indexed are abstracts written in English. The system of descriptors is prescribed. For the application of the AIR/PHYS system a large-scale dictionary containing more than 600 000 word-descriptor relations reap. phrase-descriptor relations has been developed. Most of these relations have been obtained by means of statistical and heuristical methods. In consequence, the relation system is rather imperfect. Therefore, the indexing system needs some fault- tolerating features. An appropriate indexing approach and the corresponding structure of the AIR/PHYS system are described. Finally, the conditions of the application as well as problems of further development are discussed.
Peter Biebricher, Norbert Fuhr, Gerhard Lustig, Michael Schwantner, Gerhard Knorz
SIGIR2
1988 Optimum Probability Estimation Based on Expectations
abstract
Probability estimation is important for the application of probabilistic models as well as for any evaluation in IR. We discuss the interdependencies between parameter estimation and other properties of probabilistic models. Then we define an optimum estimate which can be applied to various typical estimation problems in IR. A method for the computation of this estimate is described which uses expectations from empirical distributions. Some experiments show the applicability of our method, whereas comparable approaches are partially based on false assumptions or yield estimates with systematic errors.
Norbert Fuhr, Hubert Hüther
SIGIR1
1987 Probabilistic Search Term Weighting-Some Negative Results
abstract
The effect of probabilistic search term weighting on the improvement of retrieval quality has been demonstrated in various experiments described in the literature. In this paper, we investigate the feasibility of this method for boolean retrieval with terms from a prescribed indexing vocabulary. This is a quite different test setting in comparison to other experiments where linear retrieval with free text terms was used. The experimental results show that in our case no improvement over a simple coordination match function can be achieved. On the other hand, models based on probabilistic indexing outperform the ranking procedures using search term weights.
Norbert Fuhr, Peter Müller 0004
SIGIR1
1986 Two Models of Retrieval with Probabilistic Indexing
abstract
We describe two retrieval models for probabilistic indexing. The binary independence indexing (BII) model is a generalized version of the Maron & Kuhns indexing model. In this model, the indexing weight of a descriptor in a document is an estimate of the probability of relevance of this document with respect to queries using this descriptor. The retrieval-with-probabilistic-indexing (RPI) model is suited to different kinds of probabilistic indexing. Therefore we assume that each indexing model has its own concept of 'correctness' to which the probabilities relate. The concept of correctness is not necessarily identical with the concept of relevance, it is only required to depend on relevance. In addition to the probabilistic indexing weights, the RPI model provides the possibility of relevance weighting of search terms. Both retrieval models are compared in experiments, showing equally good results.
Norbert Fuhr
SIGIR1
1984 Retrieval Test Evaluation of a Rule Based Automatic Index (AIR/PHYS)
Norbert Fuhr, Gerhard Knorz
SIGIR1