EDBT 2026 Demo / reviewers in the wild / expert
Kalervo Järvelin
dblp:j/KalervoJarvelin
· DBLP profile ↗
89ranked-venue papers in the field
24as first author
5since 2021 · last 2025
0000-0001-7655-8930ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 81 (20 first)Database Systems & Data Management · 6 (3 first)Knowledge Engineering, Semantic Web & Information Systems · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating Multi-Dimensional Cumulated Utility in Information RetrievalabstractTraditional Information Retrieval (IR) effectiveness metrics assume that a relevant document satisfies the information need as a whole. Nevertheless, if the information need is faceted or contains subtopics, this notion of relevance cannot model documents relevant only to one or a few subtopics. Furthermore, faceted documents in a ranked list may focus on the same subtopics, and their content may overlap while neglecting other subtopics. Hence, a search result, where topranked documents deal with different subtopics should be preferred over a result where documents are thematically limited and provide overlapping information. The Multi-Dimensional Cumulated Utility (MDCU) metric, recently formulated theoretically by Järvelin and Sormunen, extends the evaluation of novelty and diversity by considering content overlapping among documents. While Järvelin and Sormunen described the theory of MDCU and illustrated its application on a toy example, they did not investigate its empirical use. In this paper, we show the practical feasibility and validity of the MDCU by applying it to publicly available TREC test collections. Furthermore, we analyse its relation with the well-established α-nDCG, and finally, we provide a Python implementation of the MDCU, fostering its adoption as an evaluation framework. Our results indicate a positive correlation between α-nDCG and MDCU, suggesting that both measures correctly identify similar trends when evaluating the IR systems. Finally, compared to α-nDCG, MDCU exhibits a stronger statistical power and identifies up to 9 times more statistically significantly different pairs of systems. Francesco Luigi De Faveri, Guglielmo Faggioli, Nicola Ferro 0001, Kalervo Järvelin |
SIGIR | 4 |
| 2024 | TraQuLA: Transparent Question Answering Over RDF Through Linguistic Analysis
Elizaveta Zimina, Kalervo Järvelin, Jaakko Peltonen, Aarne Ranta, Jyrki Nummenmaa |
ICWE | 2 |
| 2024 | A Blueprint of IR Evaluation Integrating Task and User CharacteristicsabstractTraditional search result evaluation metrics in information retrieval, such as MAP and NDCG, naively focus on topical relevance between a document and search topic and assume this relationship as mono-dimensional and often binary. They neglect document content overlap and assume gains piling up as the searcher examines the ranked list at greater length. We propose a novel search result evaluation framework based on multidimensional, graded relevance assessments, explicit modelling of document overlaps and attributes affecting document usability beyond relevance. Document relevance to a search task is seen to consist of several content themes and document usability attributes. Documents may also overlap regarding their content themes. Attributes such as document readability, trustworthiness, or language represent the entire document’s usability in the search task context, for a given searcher and her motivating task. The proposed framework evaluates the quality of a ranked search result, taking into account the contribution of each successive document, with estimated overlap across themes, and usability based on its attributes. Kalervo Järvelin, Eero Sormunen |
ACM Trans. Inf. Syst. | 1 |
| 2023 | The association of disciplinary background with the evolution of topics and methods in Library and Information Science research 1995-2015abstractAbstract The paper reports a longitudinal analysis of the topical and methodological development of Library and Information Science (LIS). Its focus is on the effects of researchers' disciplines on these developments. The study extends an earlier cross‐sectional study (Vakkari et al., Journal of the Association for Information Science and Technology, 2022a, 73, 1706–1722) by a coordinated dataset representing a content analysis of articles published in 31 scholarly LIS journals in 1995, 2005, and 2015. It is novel in its coverage of authors' disciplines, topical and methodological aspects in a coordinated dataset spanning two decades thus allowing trend analysis. The findings include a shrinking trend in the share of LIS from 67 to 36% while Computer Science, and Business and Economics increase their share from 9 and 6% to 21 and 16%, respectively. The earlier cross‐sectional study (Vakkari et al., Journal of the Association for Information Science and Technology, 2022a, 73, 1706–1722) for the year 2015 identified three topical clusters of LIS research, focusing on topical subfields, methodologies, and contributing disciplines. Correspondence analysis confirms their existence already in 1995 and traces their development through the decades. The contributing disciplines infuse their concepts, research questions, and approaches to LIS and may also subsume vital parts of LIS in their own structures of knowledge production. Pertti Vakkari, Kalervo Järvelin, Yu-Wei Chang 0001 |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2022 | Disciplinary contributions to research topics and methodology in Library and Information Science - Leading to fragmentation?abstractAbstract The study analyses contributions to Library and Information Science (LIS) by researchers representing various disciplines. How are such contributions associated with the choice of research topics and methodology? The study employs a quantitative content analysis of articles published in 31 scholarly LIS journals in 2015. Each article is seen as a contribution to LIS by the authors' disciplines, which are inferred from their affiliations. The unit of analysis is the article‐discipline pair. Of the contribution instances, the share of LIS is one third. Computer Science contributes one fifth and Business and Economics one sixth. The latter disciplines dominate the contributions in information retrieval, information seeking, and scientific communication indicating strong influences in LIS. Correspondence analysis reveals three clusters of research, one focusing on traditional LIS with contributions from LIS and Humanities and survey‐type research; another on information retrieval with contributions from Computer Science and experimental research; and the third on scientific communication with contributions from Natural Sciences and Medicine and citation analytic research. The strong differentiation of scholarly contributions in LIS hints to the fragmentation of LIS as a discipline. Pertti Vakkari, Yu-Wei Chang 0001, Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2020 | Probabilistic Dynamic Non-negative Group Factor Model for Multi-source Text MiningabstractNonnegative matrix factorization (NMF) is a popular approach to model data, however, most models are unable to flexibly take into account multiple matrices across sources and time or apply only to integer-valued data. We introduce a probabilistic, Gaussian Process-based, more inclusive NMF-based model which jointly analyzes nonnegative data such as text data word content from multiple sources in a temporal dynamic manner. The model collectively models observed matrix data, source-wise latent variables, and their dependencies and temporal evolution with a full-fledged hierarchical approach including flexible nonparametric temporal dynamics. Experiments on simulated data and real data show the model out-performs, comparable models. A case study on social media and news demonstrates the model discovers semantically meaningful topical factors and their evolution Chien Lu, Jaakko Peltonen, Jyrki Nummenmaa, Kalervo Järvelin |
CIKM | 4 |
| 2018 | QWERTY: The Effects of Typing on Web Search BehaviorabstractTyping is a common form of query input for search engines and other information retrieval systems; we therefore investigate the relationship between typing behavior and search interactions. The search process is interactive and typically requires entering one or more queries, and assessing both summaries from Search Engine Result Pages and the underlying documents, to ultimately satisfy some information need. Under the Search Economic Theory model of interactive information retrieval, differences in query costs will result in search behavior changes. We investigate how differences in query inputs themselves may relate to Search Economic Theory by conducting a lab-based experiment to observe how text entries influence subsequent search interactions. Our results indicate that for faster typing speeds, more queries are entered in a session, while both query lengths and assessment times are lower. Kevin Ong, Kalervo Järvelin, Mark Sanderson, Falk Scholer |
CHIIR | 2 |
| 2018 | Salton Award Keynote: Information Interaction in Contextabstractkeynote Share on Salton Award Keynote: Information Interaction in Context Author: Kalervo P. Jarvelin Univerity of Tampere, Tampere, Finland Univerity of Tampere, Tampere, FinlandView Profile Authors Info & Claims SIGIR '18: The 41st International ACM SIGIR Conference on Research & Development in Information RetrievalJune 2018 Pages 1–2https://doi.org/10.1145/3209978.3210230Published:27 June 2018Publication History 0citation451DownloadsMetricsTotal Citations0Total Downloads451Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Kalervo Järvelin |
SIGIR | 1 |
| 2017 | Using Information Scent to Understand Mobile and Desktop Web Search BehaviorabstractThis paper investigates if Information Foraging Theory can be used to understand differences in user behavior when searching on mobile and desktop web search systems. Two groups of thirty-six participants were recruited to carry out six identical web search tasks on desktop or on mobile. The search tasks were prepared with a different number and distribution of relevant documents on the first result page. Search behaviors on mobile and desktop were measurably different. Desktop participants viewed and clicked on more results but saved fewer as relevant, compared to mobile participants, when information scent level increased. Mobile participants achieved higher search accuracy than desktop participants for tasks with increasing numbers of relevant search results. Conversely, desktop participants were more accurate than mobile participants for tasks with an equal number of relevant results that were more distributed across the results page. Overall, both an increased number and better positioning of relevant search results improved the ability of participants to locate relevant results on both desktop and mobile. Participants spent more time and issued more queries on desktop, but abandoned less and saved more results for initial queries on mobile. Kevin Ong, Kalervo Järvelin, Mark Sanderson, Falk Scholer |
SIGIR | 2 |
| 2017 | Validating simulated interaction for retrieval evaluation
Teemu Pääkkönen, Jaana Kekäläinen, Heikki Keskustalo, Leif Azzopardi, David Maxwell 0001, Kalervo Järvelin |
Inf. Retr. J. | 6 |
| 2017 | Search task features in work tasks of varying types and complexityabstractInformation searching in practice seldom is an end in itself. In work, work task (WT) performance forms the context, which information searching should serve. Therefore, information retrieval (IR) systems development/evaluation should take the WT context into account. The present paper analyzes how WT features: task complexity and task types, affect information searching in authentic work: the types of information needs, search processes, and search media. We collected data on 22 information professionals in authentic work situations in three organization types: city administration, universities, and companies. The data comprise 286 WTs and 420 search tasks (STs). The data include transaction logs, video recordings, daily questionnaires, interviews. and observation. The data were analyzed quantitatively. Even if the participants used a range of search media, most STs were simple throughout the data, and up to 42% of WTs did not include searching. WT's effects on STs are not straightforward: different WT types react differently to WT complexity. Due to the simplicity of authentic searching, the WT/ST types in interactive IR experiments should be reconsidered. Miamaria Saastamoinen, Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2016 | Adaptive Distributional Extensions to DFR RankingabstractDivergence From Randomness (DFR) ranking models assume that informative terms are distributed in a corpus differently than non-informative terms. Different statistical models (e.g. Poisson, geometric) are used to model the distribution of non-informative terms, producing different DFR models. An informative term is then detected by measuring the divergence of its distribution from the distribution of non-informative terms. However, there is little empirical evidence that the distributions of non-informative terms used in DFR actually fit current datasets. Practically this risks providing a poor separation between informative and non-informative terms, thus compromising the discriminative power of the ranking model. We present a novel extension to DFR, which first detects the best-fitting distribution of non-informative terms in a collection, and then adapts the ranking computation to this best-fitting distribution. We call this model Adaptive Distributional Ranking (ADR) because it adapts the ranking to the statistics of the specific dataset being processed each time. Experiments on TREC data show ADR to outperform DFR models (and their extensions) and be comparable in performance to a query likelihood language model (LM). Casper Petersen, Jakob Grue Simonsen, Kalervo Järvelin, Christina Lioma |
CIKM | 3 |
| 2016 | The twist measure for IR evaluation: Taking user's effort into accountabstractWe present a novel measure for ranking evaluation, called Twist (τ). It is a measure for informational intents, which handles both binary and graded relevance. τ stems from the observation that searching is currently a that searching is currently taken for granted and it is natural for users to assume that search engines are available and work well. As a consequence, users may assume the utility they have in finding relevant documents, which is the focus of traditional measures, as granted. On the contrary, they may feel uneasy when the system returns nonrelevant documents because they are then forced to do additional work to get the desired information, and this causes avoidable effort. The latter is the focus of τ, which evaluates the effectiveness of a system from the point of view of the effort required to the users to retrieve the desired information. We provide a formal definition of τ, a demonstration of its properties, and introduce the notion of effort/gain plots, which complement traditional utility‐based measures. By means of an extensive experimental evaluation, τ is shown to grasp different aspects of system performances, to not require extensive and costly assessments, and to be a robust tool for detecting differences between systems. Nicola Ferro 0001, Gianmaria Silvello, Heikki Keskustalo, Ari Pirkola, Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2015 | Searching and Stopping: An Analysis of Stopping Rules and StrategiesabstractSearching naturally involves stopping points, both at a query level (how far down the ranked list should I go?) and at a session level (how many queries should I issue?). Understanding when searchers stop has been of much interest to the community because it is fundamental to how we evaluate search behaviour and performance. Research has shown that searchers find it difficult to formalise stopping criteria, and typically resort to their intuition of what is "good enough". While various heuristics and stopping criteria have been proposed, little work has investigated how well they perform, and whether searchers actually conform to any of these rules. In this paper, we undertake the first large scale study of stopping rules, investigating how they influence overall session performance, and which rules best match actual stopping behaviour. Our work is focused on stopping at the query level in the context of ad-hoc topic retrieval, where searchers undertake search tasks within a fixed time period. We show that stopping strategies based upon the disgust or frustration point rules - both of which capture a searcher's tolerance to non-relevance - typically result in (i) the best overall performance, and (ii) provide the closest approximation to actual searcher behaviour, although a fixed depth approach also performs remarkably well. Findings from this study have implications regarding how we build measures, and how we conduct simulations of search behaviours. David Maxwell 0001, Leif Azzopardi, Kalervo Järvelin, Heikki Keskustalo |
CIKM | 3 |
| 2015 | User Simulations for Interactive Search: Evaluating Personalized Query Suggestion
Suzan Verberne, Maya Sappelli, Kalervo Järvelin, Wessel Kraaij |
ECIR | 3 |
| 2015 | An Initial Investigation into Fixed and Adaptive Stopping StrategiesabstractMost models, measures and simulations often assume that a searcher will stop at a predetermined place in a ranked list of results. However, during the course of a search session, real-world searchers will vary and adapt their interactions with a ranked list. These interactions depend upon a variety of factors, including the content and quality of the results returned, and the searcher's information need. In this paper, we perform a preliminary simulated analysis into the influence of stopping strategies when query quality varies. Placed in the context of ad-hoc topic retrieval during a multi-query search session, we examine the influence of fixed and adaptive stopping strategies on overall performance. Surprisingly, we find that a fixed strategy can perform as well as the examined adaptive strategies, but the fixed depth needs to be adjusted depending on the querying strategy used. Further work is required to explore how well the stopping strategies reflect actual search behaviour, and to determine whether one stopping strategy is dominant. David Maxwell 0001, Leif Azzopardi, Kalervo Järvelin, Heikki Keskustalo |
SIGIR | 3 |
| 2015 | Task-Based Information Interaction Evaluation: The Viewpoint of Program TheoryabstractEvaluation is central in research and development of information retrieval (IR). In addition to designing and implementing new retrieval mechanisms, one must also show through rigorous evaluation that they are effective. A major focus in IR is IR mechanisms’ capability of ranking relevant documents optimally for the users, given a query. Searching for information in practice involves searchers, however, and is highly interactive. When human searchers have been incorporated in evaluation studies, the results have often suggested that better ranking does not necessarily lead to better search task, or work task, performance. Therefore, it is not clear which system or interface features should be developed to improve the effectiveness of human task performance. In the present article, we focus on the evaluation of task-based information interaction (TBII). We give special emphasis to learning tasks to discuss TBII in more concrete terms. Information interaction is here understood as behavioral and cognitive activities related to task planning, searching information items, selecting between them, working with them, and synthesizing and reporting. These five generic activities contribute to task performance and outcome and can be supported by information systems. In an attempt toward task-based evaluation, we introduce program theory as the evaluation framework. Such evaluation can investigate whether a program consisting of TBII activities and tools works and how it works and, further, provides a causal description of program (in)effectiveness. Our goal in the present article is to structure TBII on the basis of the five generic activities and consider the evaluation of each activity using the program theory framework. Finally, we combine these activity-based program theories in an overall evaluation framework for TBII. Such an evaluation is complex due to the large number of factors affecting information interaction. Instead of presenting tested program theories, we illustrate how the evaluation of TBII should be accomplished using the program theory framework in the evaluation of systems and behaviors, and their interactions, comprehensively in context. Kalervo Järvelin, Pertti Vakkari, Paavo Arvola, Feza Baskaya, Anni Järvelin, Jaana Kekäläinen, Heikki Keskustalo, Sanna Kumpulainen, Miamaria Saastamoinen, Reijo Savolainen, Eero Sormunen |
ACM Trans. Inf. Syst. | 1 |
| 2014 | Evolution of library and information science, 1965-2005: Content analysis of journal articlesabstractThis article first analyzes library and information science ( LIS ) research articles published in core LIS journals in 2005. It also examines the development of LIS from 1965 to 2005 in light of comparable data sets for 1965, 1985, and 2005. In both cases, the authors report (a) how the research articles are distributed by topic and (b) what approaches, research strategies, and methods were applied in the articles. In 2005, the largest research areas in LIS by this measure were information storage and retrieval, scientific communication, library and information‐service activities, and information seeking. The same research areas constituted the quantitative core of LIS in the previous years since 1965. Information retrieval has been the most popular area of research over the years. The proportion of research on library and information‐service activities decreased after 1985, but the popularity of information seeking and of scientific communication grew during the period studied. The viewpoint of research has shifted from library and information organizations to end users and development of systems for the latter. The proportion of empirical research strategies was high and rose over time, with the survey method being the single most important method. However, attention to evaluation and experiments increased considerably after 1985. Conceptual research strategies and system analysis, description, and design were quite popular, but declining. The most significant changes from 1965 to 2005 are the decreasing interest in library and information‐service activities and the growth of research into information seeking and scientific communication. Otto Tuomaala, Kalervo Järvelin, Pertti Vakkari |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2013 | Modeling behavioral factors ininteractive information retrievalabstractIn real-life, information retrieval consists of sessions of one or more query iterations. Each iteration has several subtasks like query formulation, result scanning, document link clicking, document reading and judgment, and stopping. Each of the subtasks has behavioral factors associated with them. These factors include search goals and cost constraints, query formulation strategies, scanning and stopping strategies, and relevance assessment behav-ior. Traditional IR evaluation focuses on retrieval and result presentation methods, and interaction within a single-query session. In the present study we aim at assessing the effects of the behavioral factors on retrieval effectiveness. Our research questions include how effective is human behavior employing search strategies compared to various baselines under various search goals and time constraints. We examine both ideal as well as fallible human behavior and wish to identify robust behaviors, if any. Methodologically, we use extensive simulation of human behavior in a test collection. Our findings include that (a) human behavior using multi-query sessions may exceed in effectiveness comparable single-query sessions, (b) the same empirically observed behavioral patterns are reasonably effective under various search goals and constraints, but (c) remain on average clearly below the best possible ones. Moreover, there is no behavioral pattern for sessions that would be even close to winning in most cases; the information need (or topic) in relation to the test collection is a determining factor. Feza Baskaya, Heikki Keskustalo, Kalervo Järvelin |
CIKM | 3 |
| 2012 | Time drives interaction: simulating sessions in diverse searching environmentsabstractReal life information retrieval takes place in sessions, where users search by iterating between various cognitive, perceptual and motor subtasks through an interactive interface. The sessions may follow diverse strategies, which, together with the interface characteristics, affect user effort (cost), experience and session effectiveness. In this paper we propose a pragmatic evaluation approach based on scenarios with explicit subtask costs. We study the limits of effectiveness of diverse interactive searching strategies in two searching environments (the scenarios) under overall cost constraints. This is based on a comprehensive simulation of 20 million sessions in each scenario. We analyze the effectiveness of the session strategies over time, and the properties of the most and the least effective sessions in each case. Furthermore, we will also contrast the proposed evaluation approach with the traditional one, rank based evaluation, and show how the latter may hide essential factors that affect users' performance and satisfaction - and gives even counter-intuitive results. Feza Baskaya, Heikki Keskustalo, Kalervo Järvelin |
SIGIR | 3 |
| 2012 | Barriers to task-based information access in molecular medicineabstractAbstract We analyze barriers to task‐based information access in molecular medicine, focusing on research tasks, which provide task performance sessions of varying complexity. Molecular medicine is a relevant domain because it offers thousands of digital resources as the information environment. Data were collected through shadowing of real work tasks. Thirty work task sessions were analyzed and barriers in these identified. The barriers were classified by their character (conceptual, syntactic, and technological) and by their context of appearance (work task, system integration, or system). Also, work task sessions were grouped into three complexity classes and the frequency of barriers of varying types across task complexity levels were analyzed. Our findings indicate that although most of the barriers are on system level, there is a quantum of barriers in integration and work task contexts. These barriers might be overcome through attention to the integrated use of multiple systems at least for the most frequent uses. This can be done by means of standardization and harmonization of the data and by taking the requirements of the work tasks into account in system design and development, because information access is seldom an end itself, but rather serves to reach the goals of work tasks. Sanna Kumpulainen, Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2011 | Simulating Simple and Fallible Relevance Feedback
Feza Baskaya, Heikki Keskustalo, Kalervo Järvelin |
ECIR | 3 |
| 2011 | IR Research: Systems, Interaction, Evaluation and Theories
Kalervo Järvelin |
ECIR | 1 |
| 2011 | GRAS: An effective and efficient stemming algorithm for information retrievalabstractA novel graph-based language-independent stemming algorithm suitable for information retrieval is proposed in this article. The main features of the algorithm are retrieval effectiveness, generality, and computational efficiency. We test our approach on seven languages (using collections from the TREC, CLEF, and FIRE evaluation platforms) of varying morphological complexity. Significant performance improvement over plain word-based retrieval, three other language-independent morphological normalizers, as well as rule-based stemmers is demonstrated. Jiaul H. Paik, Mandar Mitra, Swapan K. Parui, Kalervo Järvelin |
ACM Trans. Inf. Syst. | 4 |
| 2010 | Measuring impact of twelve information scientists using the DCI indexabstractAbstract The Discounted Cumulated Impact (DCI) index has recently been proposed for research evaluation. In the present work an earlier dataset by Cronin and Meho (2007) is reanalyzed, with the aim of exemplifying the salient features of the DCI index. We apply the index on, and compare our results to, the outcomes of the Cronin‐Meho (2007) study. Both authors and their top publications are used as units of analysis, which suggests that, by adjusting the parameters of evaluation according to the needs of research evaluation, the DCI index delivers data on an author's (or publication's) “lifetime” impact or current impact at the time of evaluation on an author's (or publication's) capability of inviting citations from highly cited later publications as an indication of impact, and on the relative impact across a set of authors (or publications) over their “lifetime” or currently. Per Ahlgren, Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2010 | Derived types in semantic association discovery
Janne Jämsen, Timo Niemi, Kalervo Järvelin |
J. Intell. Inf. Syst. | 3 |
| 2009 | Interactive relevance feedback with graded relevance and sentence extraction: simulated user experimentsabstractResearch on relevance feedback (RFB) in information retrieval (IR) has given mixed results. Success in RFB seems to depend on the searcher's willingness to provide feedback and ability to identify relevant documents or query keys. The paper is based on simulating many user scenarios regarding the amount and quality of RFB. In addition, we experiment with query-biased sentence extraction for query reformulation. The baselines are initial no-feedback queries and queries based on pseudo-relevance feedback. The core question is: under which conditions would RFB based on sentence extraction be successful? The answer depends on user's behavior, implementation of feedback query formulation, and the evaluation methods. A small amount of feedback from a short browsing window seems to improve the final ranking the most. Longer browsing allows more feedback and better queries but also consumes the available relevant documents. Kalervo Järvelin |
CIKM | 1 |
| 2009 | Including summaries in system evaluationabstractIn batch evaluation of retrieval systems, performance is calculated based on predetermined relevance judgements applied to a list of documents returned by the system for a query. This evaluation paradigm, however, ignores the current standard operation of search systems which require the user to view summaries of documents prior to reading the documents themselves. Andrew Turpin, Falk Scholer, Kalervo Järvelin, Mingfang Wu, J. Shane Culpepper |
SIGIR | 3 |
| 2008 | Discounted Cumulated Gain Based Evaluation of Multiple-Query IR Sessions
Kalervo Järvelin, Susan Price, Lois M. L. Delcambre, Marianne Lykke |
ECIR | 1 |
| 2008 | A Novel Implementation of the FITE-TRT Translation Method
Aki Loponen, Ari Pirkola, Kalervo Järvelin, Heikki Keskustalo |
ECIR | 3 |
| 2008 | Intuition-supporting visualization of user's performance based on explicit negative higher-order relevanceabstractModeling the beyond-topical aspects of relevance are currently gaining popularity in IR evaluation. For example, the discounted cumulated gain (DCG) measure implicitly models some aspects of higher-order relevance via diminishing the value of relevant documents seen later during retrieval (e.g., due to information cumulated, redundancy, and effort). In this paper, we focus on the concept of negative higher-order relevance (NHOR) made explicit via negative gain values in IR evaluation. We extend the computation of DCG to allow negative gain values, perform an experiment in a laboratory setting, and demonstrate the characteristics of NHOR in evaluation. The approach leads to intuitively reasonable performance curves emphasizing, from the user's point of view, the progression of retrieval towards success or failure. We discuss normalization issues when both positive and negative gain values are allowed and conclude by discussing the usage of NHOR to characterize test collections. Heikki Keskustalo, Kalervo Järvelin, Ari Pirkola, Jaana Kekäläinen |
SIGIR | 2 |
| 2008 | Evaluating the effectiveness of relevance feedback based on a user simulation model: effects of a user scenario on cumulated gain value
Heikki Keskustalo, Kalervo Järvelin, Ari Pirkola |
Inf. Retr. | 2 |
| 2008 | Focused web crawling in the acquisition of comparable corpora
Tuomas Talvensaari, Ari Pirkola, Kalervo Järvelin, Martti Juhola, Jorma Laurikkala |
Inf. Retr. | 3 |
| 2008 | The DCI index: Discounted cumulated impact-based research evaluationabstractAbstract Research evaluation is increasingly popular and important among research funding bodies and science policy makers. Various indicators have been proposed to evaluate the standing of individual scientists, institutions, journals, or countries. A simple and popular one among the indicators is the h‐index, the Hirsch index (Hirsch 2005), which is an indicator for lifetime achievement of a scholar. Several other indicators have been proposed to complement or balance the h‐index. However, these indicators have no conception of aging. The AR‐index (Jin et al. 2007) incorporates aging but divides the received citation counts by the raw age of the publication. Consequently, the decay of a publication is very steep and insensitive to disciplinary differences. In addition, we believe that a publication becomes outdated only when it is no longer cited, not because of its age. Finally, all indicators treat citations as equally material when one might reasonably think that a citation from a heavily cited publication should weigh more than a citation froma non‐cited or little‐cited publication.We propose a new indicator, the Discounted Cumulated Impact (DCI) index, which devalues old citations in a smooth way. It rewards an author for receiving new citations even if the publication is old. Further, it allows weighting of the citations by the citation weight of the citing publication. DCI can be used to calculate research performance on the basis of the h‐core of a scholar or any other publication data set. Finally, it supports comparing research performance to the average performance in the domain and across domains as well. Kalervo Järvelin, Olle Persson |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | Erratum re: "The DCI-index: Discounted cumulated impact-based research evaluation", JASIST 59(9), 1433-1440abstractAbstract The article by K. Järvelin & O. Persson published in JASIST 59(9), “The DCI‐Index: Discounted Cumulated Impact‐Based Research Evaluation,” (pp. 1433–1440) contains an unfortunate error in one of its formulas, Equation 3 . The present paper gives the correction and an example of impact analysis based on the corrected formula. Kalervo Järvelin, Olle Persson |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2008 | Experiments with transitive dictionary translation and pseudo-relevance feedback using graded relevance assessmentsabstractAbstract In this article, the authors present evaluation results for transitive dictionary‐based cross‐language information retrieval (CLIR) using graded relevance assessments in a best match retrieval environment. A text database containing newspaper articles and a related set of 35 search topics were used in the tests. Source language topics (in English, German, and Swedish) were automatically translated into the target language (Finnish) via an intermediate (or pivot) language. Effectiveness of the transitively translated queries was compared to that of the directly translated and monolingual Finnish queries. Pseudo‐relevance feedback (PRF) was also used to expand the original transitive target queries. Cross‐language information retrieval performance was evaluated on three relevance thresholds: stringent, regular, and liberal. The transitive translations performed well achieving, on the average, 85–93% of the direct translation performance, and 66–72% of monolingual performance. Moreover, PRF was successful in raising the performance of transitive translation routes in absolute terms as well as in relation to monolingual and direct translation performance applying PRF. Raija Lehtokangas, Heikki Keskustalo, Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2008 | A tool for data cube construction from structurally heterogeneous XML documentsabstractAbstract Data cubes for OLAP (On‐Line Analytical Processing) often need to be constructed from data located in several distributed and autonomous information sources. Such a data integration process is challenging due to semantic, syntactic, and structural heterogeneity among the data. While XML (extensible markup language) is the de facto standard for data exchange, the three types of heterogeneity remain. Moreover, popular path‐oriented XML query languages, such as XQuery, require the user to know in much detail the structure of the documents to be processed and are, thus, effectively impractical in many real‐world data integration tasks. Several Lowest Common Ancestor (LCA)‐based XML query evaluation strategies have recently been introduced to provide a more structure‐independent way to access XML documents. We shall, however, show that this approach leads in the context of certain—not uncommon—types of XML documents to undesirable results. This article introduces a novel high‐level data extraction primitive that utilizes the purpose‐built Smallest Possible Context (SPC) query evaluation strategy. We demonstrate, through a system prototype for OLAP data cube construction and a sample application in informetrics, that our approach has real advantages in data integration. Turkka Näppilä, Kalervo Järvelin, Timo Niemi |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2007 | s-grams: Defining generalized n-grams for information retrieval
Anni Järvelin, Antti Järvelin, Kalervo Järvelin |
Inf. Process. Manag. | 3 |
| 2007 | Restricted inflectional form generation in management of morphological keyword variation
Kimmo Kettunen 0001, Eija Airio, Kalervo Järvelin |
Inf. Retr. | 3 |
| 2007 | An analysis of two approaches in information retrieval: From frameworks to study designsabstractAbstract There is a well‐known gap between systems‐oriented information retrieval (IR) and user‐oriented IR, which cognitive IR seeks to bridge. It is therefore interesting to analyze approaches at the level of frameworks, models, and study designs. This article is an exercise in such an analysis, focusing on two significant approaches to IR: the lab IR approach and P. Ingwersen's (1996) cognitive IR approach. The article focuses on their research frameworks, models, hypotheses, laws and theories, study designs, and possible contributions. The two approaches are quite different, which becomes apparent in the use of independent, controlled, and dependent variables in the study designs of each approach. Thus, each approach is capable of contributing very differently to understanding and developing information access. The article also discusses integrating the approaches at the study‐design level. Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2007 | Corpus-based cross-language information retrieval in retrieval of highly relevant documentsabstractAbstract Information retrieval systems' ability to retrieve highly relevant documents has become more and more important in the age of extremely large collections, such as the World Wide Web (WWW). The authors' aim was to find out how corpus‐based cross‐language information retrieval (CLIR) manages in retrieving highly relevant documents. They created a Finnish–Swedish comparable corpus from two loosely related document collections and used it as a source of knowledge for query translation. Finnish test queries were translated into Swedish and run against a Swedish test collection. Graded relevance assessments were used in evaluating the results and three relevance criterion levels—liberal, regular, and stringent—were applied. The runs were also evaluated with generalized recall and precision, which weight the retrieved documents according to their relevance level. The performance of the Comparable Corpus Translation system (COCOT) was compared to that of a dictionary‐based query translation program; the two translation methods were also combined. The results indicate that corpus‐based CLIR performs particularly well with highly relevant documents. In average precision, COCOT even matched the monolingual baseline on the highest relevance level. The performance of the different query translation methods was further analyzed by finding out reasons for poor rankings of highly relevant documents. Tuomas Talvensaari, Martti Juhola, Jorma Laurikkala, Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2007 | Frequency-based identification of correct translation equivalents (FITE) obtained through transformation rulesabstractWe devised a novel statistical technique for the identification of the translation equivalents of source words obtained by transformation rule based translation (TRT). The effectiveness of the technique called frequency-based identification of translation equivalents ( FITE ) was tested using biological and medical cross-lingual spelling variants and out-of-vocabulary (OOV) words in Spanish-English and Finnish-English TRT. The results showed that, depending on the source language and frequency corpus, FITE-TRT (the identification of translation equivalents from TRT's translation set by means of the FITE technique) may achieve high translation recall. In the case of the Web as the frequency corpus, translation recall was 89.2%--91.0% for Spanish-English FITE-TRT. For both language pairs FITE-TRT achieved high translation precision: 95.0%--98.8%. The technique also reliably identified native source language words: source words that cannot be correctly translated by TRT. Dictionary-based CLIR augmented with FITE-TRT performed substantially better than basic dictionary-based CLIR where OOV keys were kept intact. FITE-TRT with Web document frequencies was the best technique among several fuzzy translation/matching approaches tested in cross-language retrieval experiments. We also discuss the application of FITE-TRT in the automatic construction of multilingual dictionaries. Ari Pirkola, Jarmo Toivonen, Heikki Keskustalo, Kalervo Järvelin |
ACM Trans. Inf. Syst. | 4 |
| 2007 | Creating and exploiting a comparable corpus in cross-language information retrievalabstractWe present a method for creating a comparable text corpus from two document collections in different languages. The collections can be very different in origin. In this study, we build a comparable corpus from articles by a Swedish news agency and a U.S. newspaper. The keys with best resolution power were extracted from the documents of one collection, the source collection, by using the relative average term frequency (RATF) value. The keys were translated into the language of the other collection, the target collection, with a dictionary-based query translation program. The translated queries were run against the target collection and an alignment pair was made if the retrieved documents matched given date and similarity score criteria. The resulting comparable collection was used as a similarity thesaurus to translate queries along with a dictionary-based translator. The combined approaches outperformed translation schemes where dictionary-based translation or corpus translation was used alone. Tuomas Talvensaari, Jorma Laurikkala, Kalervo Järvelin, Martti Juhola, Heikki Keskustalo |
ACM Trans. Inf. Syst. | 3 |
| 2006 | The Effects of Relevance Feedback Quality and Quantity in Interactive Relevance Feedback: A Simulation Based on User Modeling
Heikki Keskustalo, Kalervo Järvelin, Ari Pirkola |
ECIR | 2 |
| 2006 | Hierarchical clustering of a Finnish newspaper article collection with graded relevance assessments
Tuomo Korenius, Jorma Laurikkala, Martti Juhola, Kalervo Järvelin |
Inf. Retr. | 4 |
| 2006 | Experiments with dictionary-based CLIR using graded relevance assessments: Improving effectiveness by pseudo-relevance feedback
Raija Lehtokangas, Heikki Keskustalo, Kalervo Järvelin |
Inf. Retr. | 3 |
| 2006 | "Irrational" searchers and IR-rational researchersabstractAbstract In this article the authors look at the prescriptions advocated by Web search textbooks in the light of a selection of empirical data of real Web information search processes. They use the strategy of disjointed incrementalism, which is a theoretical foundation from decision making, to focus on how people face complex problems, and claim that such problem solving can be compared to the tasks searchers perform when interacting with the Web. The findings suggest that textbooks on Web searching should take into account that searchers only tend to take a certain number of sources into consideration, that the searchers adjust their goals and objectives during searching, and that searchers reconsider the usefulness of sources at different stages of their work tasks as well as their search tasks. Nils Pharo, Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2005 | Dictionary-Based CLIR Loses Highly Relevant Documents
Raija Lehtokangas, Heikki Keskustalo, Kalervo Järvelin |
ECIR | 3 |
| 2005 | Assessing learning outcomes in two information retrieval learning environments
Kai Halttunen, Kalervo Järvelin |
Inf. Process. Manag. | 2 |
| 2005 | Collaborative Information Retrieval in an information-intensive domain
Preben Hansen, Kalervo Järvelin |
Inf. Process. Manag. | 2 |
| 2005 | Translating cross-lingual spelling variants using transformation rules
Jarmo Toivonen, Ari Pirkola, Heikki Keskustalo, Kari Visala, Kalervo Järvelin |
Inf. Process. Manag. | 5 |
| 2004 | Stemming and lemmatization in the clustering of finnish text documentsabstractStemming and lemmatization were compared in the clustering of Finnish text documents. Since Finnish is a highly inflectional and agglutinative language, we hypothesized that lemmatization, involving splitting of the compound words, would be more appropriate normalization approach than the straightforward stemming. The relevance of the documents were evaluated with a four-point relevance assessment scale, which was collapsed into binary one by considering all the relevant and only the highly relevant documents relevant, respectively. Experiments with four hierarchical clustering methods supported the hypothesis. The stringent relevance scale showed that lemmatization allowed the single and complete linkage methods to recover especially the highly relevant documents better than stemming. In comparison with stemming, lemmatization together with the average linkage and Ward's methods produced higher precision. We conclude that lemmatization is a better word normalization method than stemming, when Finnish text documents are clustered for information retrieval. Tuomo Korenius, Jorma Laurikkala, Kalervo Järvelin, Martti Juhola |
CIKM | 3 |
| 2004 | Transitive dictionary translation challenges direct dictionary translation in CLIR
Raija Lehtokangas, Eija Airio, Kalervo Järvelin |
Inf. Process. Manag. | 3 |
| 2004 | Advanced query language for manipulating complex entities
Timo Niemi, Marko Junkkari, Kalervo Järvelin, Samu Viita |
Inf. Process. Manag. | 3 |
| 2004 | The SST method: a tool for analysing Web information search processes
Nils Pharo, Kalervo Järvelin |
Inf. Process. Manag. | 2 |
| 2004 | Dictionary-Based Cross-Language Information Retrieval: Learning Experiences from CLEF 2000-2002
Turid Hedlund, Eija Airio, Heikki Keskustalo, Raija Lehtokangas, Ari Pirkola, Kalervo Järvelin |
Inf. Retr. | 6 |
| 2003 | Fuzzy translation of cross-lingual spelling variantsabstractWe will present a novel two-step fuzzy translation technique for cross-lingual spelling variants. In the first stage, transformation rules are applied to source words to render them more similar to their target language equivalents. The rules are generated automatically using translation dictionaries as source data. In the second stage, the intermediate forms obtained in the first stage are translated into a target language using fuzzy matching. The effectiveness of the technique was evaluated empirically using five source languages and English as a target language. The target word list contained 189 000 English words with the correct equivalents for the source words among them. The source words were translated using the two-step fuzzy translation technique, and the results were compared with those of plain fuzzy matching based translation. The combined technique performed better, sometimes considerably better, than fuzzy matching alone. Ari Pirkola, Jarmo Toivonen, Heikki Keskustalo, Kari Visala, Kalervo Järvelin |
SIGIR | 5 |
| 2003 | Non-adjacent Digrams Improve Matching of Cross-Lingual Spelling Variants
Heikki Keskustalo, Ari Pirkola, Kari Visala, Erkka Leppänen, Kalervo Järvelin |
SPIRE | 5 |
| 2003 | Applying query structuring in cross-language retrieval
Ari Pirkola, Deniz Puolamäki, Kalervo Järvelin |
Inf. Process. Manag. | 3 |
| 2003 | Multidimensional Data Model and Query Language for InformetricsabstractAbstract Multidimensional data analysis or On‐line analytical processing (OLAP) offers a single subject‐oriented source for analyzing summary data based on various dimensions. We demonstrate that the OLAP approach gives a promising starting point for advanced analysis and comparison among summary data in informetrics applications. At the moment there is no single precise, commonly accepted logical/conceptual model for multidimensional analysis. This is because the requirements of applications vary considerably. We develop a conceptual/logical multidimensional model for supporting the complex and unpredictable needs of informetrics. Summary data are considered with respect of some dimensions. By changing dimensions the user may construct other views on the same summary data. We develop a multidimensional query language whose basic idea is to support the definition of views in a way, which is natural and intuitive for lay users in the informetrics area. We show that this view‐oriented query language has a great expressive power and its degree of declarativity is greater than in contemporary operation‐oriented or SQL (Structured Query Language)‐like OLAP query languages. Timo Niemi, Lasse Hirvonen, Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2002 | Using graded relevance assessments in IR evaluationabstractAbstract This article proposes evaluation methods based on the use of nondichotomous relevance judgements in IR experiments. It is argued that evaluation methods should credit IR methods for their ability to retrieve highly relevant documents. This is desirable from the user point of view in modern large IR environments. The proposed methods are (1) a novel application of P‐R curves and average precision computations based on separate recall bases for documents of different degrees of relevance, and (2) generalized recall and precision based directly on multiple grade relevance assessments (i.e., not dichotomizing the assessments). We demonstrate the use of the traditional and the novel evaluation measures in a case study on the effectiveness of query types, based on combinations of query structures and expansion, in retrieving documents of various degrees of relevance. The test was run with a best match retrieval system (InQuery 1 ) in a text database consisting of newspaper articles. To gain insight into the retrieval process, one should use both graded relevance assessments and effectiveness measures that enable one to observe the differences, if any, between retrieval methods in retrieving documents of different levels of relevance. In modern times of information overload, one should pay attention, in particular, to the capability of retrieval methods retrieving highly relevant documents. Jaana Kekäläinen, Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2002 | Cumulated gain-based evaluation of IR techniques
Kalervo Järvelin, Jaana Kekäläinen |
ACM Trans. Inf. Syst. | 1 |
| 2001 | Aspects of Swedish morphology and semantics from the perspective of mono- and cross-language information retrieval
Turid Hedlund, Ari Pirkola, Kalervo Järvelin |
Inf. Process. Manag. | 3 |
| 2001 | ExpansionTool: Concept-Based Query Expansion and Construction
Kalervo Järvelin, Jaana Kekäläinen, Timo Niemi |
Inf. Retr. | 1 |
| 2001 | Dictionary-Based Cross-Language Information Retrieval: Problems, Methods, and Research Findings
Ari Pirkola, Turid Hedlund, Heikki Keskustalo, Kalervo Järvelin |
Inf. Retr. | 4 |
| 2001 | Employing the resolution power of search keysabstractAbstract Search key resolution power is analyzed in the context of a request, i.e., among the set of search keys for the request. Methods of characterizing the resolution power of keys automatically are studied, and the effects search keys of varying resolution power have on retrieval effectiveness are analyzed. It is shown that it often is possible to identify the best key of a query while the discrimination between the remaining keys presents problems. It is also shown that query performance is improved by suitably using the best key in a structured query. The tests were run with InQuery 1 in a subcollection of the TREC collection, which contained some 515,000 documents. Ari Pirkola, Kalervo Järvelin |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2000 | IR evaluation methods for retrieving highly relevant documentsabstractThis paper proposes evaluation methods based on the use of non-dichotomous relevance judgements in IR experiments. It is argued that evaluation methods should credit IR methods for their ability to retrieve highly relevant documents. This is desirable from the user point of view in modern large IR environments. The proposed methods are (1) a novel application of P-R curves and average precision computations based on separate recall bases for documents of different degrees of relevance, and (2) two novel measures computing the cumulative gain the user obtains by examining the retrieval result up to a given ranked position. We then demonstrate the use of these evaluation methods in a case study on the effectiveness of query types, based on combinations of query structures and expansion, in retrieving documents of various degrees of relevance. The test was run with a best match retrieval system (In-Query1) in a text database consisting of newspaper articles. The results indicate that the tested strong query structures are most effective in retrieving highly relevant documents. The differences between the query types are practically essential and statistically significant. More generally, the novel evaluation methods and the case demonstrate that non-dichotomous relevance assessments are applicable in IR experiments, may reveal interesting phenomena, and allow harder testing of IR methods. Kalervo Järvelin, Jaana Kekäläinen |
SIGIR | 1 |
| 2000 | The visual query language CQL for transitive and relational computation
Kalervo Järvelin, Timo Niemi, Airi Salminen |
Data Knowl. Eng. | 1 |
| 2000 | The Co-Effects of Query Structure and Expansion on Retrieval Performance in Probabilistic Text Retrieval
Jaana Kekäläinen, Kalervo Järvelin |
Inf. Retr. | 2 |
| 1999 | Integration of complex objects and transitive relationships for information retrieval
Kalervo Järvelin, Timo Niemi |
Inf. Process. Manag. | 1 |
| 1999 | The Effects of Conjunction, Facet Structure, and Dictionary Combinations in Concept-Based Cross-Language Retrieval
Ari Pirkola, Heikki Keskustalo, Kalervo Järvelin |
Inf. Retr. | 3 |
| 1998 | The Impact of Query Structure and Query Expansion on Retrieval PerformanceabstractThe effects of query structures and query expansion (QE) on retrieval performance were tested with a best match retrieval system (INQUERY').Query structure means the use of operators to express the relations between search keys.Eight different structures were tested, representing weak structures (averages and weighted averages of the weights of the keys) and strong structures (e.g., queries with more elaborated search key relations).QE was based on concepts, which were first selected from a conceptual model, and then expanded by semantic relationships given in the model.The expansion levels were (a) no expansion, (b) a synonym expansion, (c) a narrower concept expansion, (d) an associative concept expansion, and (e) a cumulative expansion of all other expansions.With weak structures and Boolean structured queries, QE was not very effective.The best performance was achieved with one of the strong structures at the largest expansion level. Jaana Kekäläinen, Kalervo Järvelin |
SIGIR | 2 |
| 1996 | A Deductive Data Model for Query ExpansionabstractWe present a deductive data model for conceptbased query expansion.It ]s based on three abstraction levels: the conceptual, linguistic and occurrence levels.Concepts and relationships among them are represented at the conceptual level.The expression level represents natural language expressions for concepts.Each expression has one or more matchmg models at the occurrence level.The models specify the matching of the expression in database Indices built in varying ways.The data model supports a concept-based query expansion and formulation tool, the ExpansionTool, for heterogeneous IR system envirom ments, Expansion is controlled by adjustable matching reliability. Kalervo Järvelin, Jaana Kristensen, Timo Niemi, Eero Sormunen, Heikki Keskustalo |
SIGIR | 1 |
| 1996 | The Effect of Anaphor and Ellipsis Resolution on Proximity Searching in a Text Database
Ari Pirkola, Kalervo Järvelin |
Inf. Process. Manag. | 2 |
| 1995 | An NF2 Relational Interface For Document Retrieval, Restructuring and AggregationabstractComplexdocuments are used in many environments, e.g., information retrieval (IR).Such documents contain subdocuments, which may contain further subdocuments, etc. Powerful tools are needed to facilitate their retrieval, restructuring, and analysis.Existing IR systems are poor in complex document restructuring and data aggregation.However, in practice, IR system users would often want to obtain aggregation information on subdocuments of complex documents.In this paper we address this problem and provide a truly declarative and powerful interface for the users.Our interface is based on the non-fkst-normal-form (NF2) relational model.It allows intuitive and systematic modeling of complex documents. Kalervo Järvelin, Timo Niemi |
SIGIR | 1 |
| 1995 | Task Complexity Affects Information Seeking and Use
Katriina Byström, Kalervo Järvelin |
Inf. Process. Manag. | 2 |
| 1995 | A Straightforward NF² Relational Interface with Applications in Information Retrieval
Timo Niemi, Kalervo Järvelin |
Inf. Process. Manag. | 2 |
| 1993 | An Entity-Based Approach to Query Processing in Relational Databases. Part I: Entity Type Representation
Kalervo Järvelin, Timo Niemi |
Data Knowl. Eng. | 1 |
| 1993 | An Entity-Based Approach to Query Processing in Relational Databases. Part II: Entity Query Construction and Updating
Kalervo Järvelin, Timo Niemi |
Data Knowl. Eng. | 1 |
| 1993 | The Evolution of Library and Information Science 1965-1985: A Content Analysis of Journal Articles
Kalervo Järvelin, Pertti Vakkari |
Inf. Process. Manag. | 1 |
| 1993 | Deductive Information Retrieval Based on ClassificationsabstractModern fact databases contain abundant data classified through several classifications. Typically, users must consult these classifications in separate manuals or files, thus making their effective use difficult. Contemporary database systems do little to support deductive use of classifications. In this study we show how deductive data management techniques can be applied to the utilization of data value classifications. Computation of transitive class relationships is of primary importance here. We define a representation of classifications which supports transitive computation and present an operation-oriented deductive query language tailored for classification-based deductive information retrieval. The operations of this language are on the same abstraction level as relational algebra operations and can be integrated with these to form a powerful and flexible query language for deductive information retrieval. We define the integration of the operations and demonstrate the usefulness of the language in terms of several sample queries. © 1993 John Wiley & Sons, Inc. Kalervo Järvelin, Timo Niemi |
J. Am. Soc. Inf. Sci. | 1 |
| 1992 | Advanced Query Formulation in Deductive Databases
Timo Niemi, Kalervo Järvelin |
Inf. Process. Manag. | 2 |
| 1992 | Operation-oriented query language approach for recursive queries - Part 1. Functional definition
Timo Niemi, Kalervo Järvelin |
Inf. Syst. | 2 |
| 1992 | Operation-oriented query language approach for recursive queries - Part 2. Prototype implementation and its integration with relational databases
Timo Niemi, Kalervo Järvelin |
Inf. Syst. | 2 |
| 1991 | Advanced Retrieval From Heterogeneous Fact Databases: Integration of Data Retrieval, Conversion, Aggregation and Deductive TechniquesabstractArticle Data conversion, aggregation and deduction for advanced retrieval from the heterogeneous fact databases Share on Authors: Kalervo Järvelin Dept. of Information Studies, University of Tampere, P.O.Box 607 SF-33101 TAMPERE, Finland Dept. of Information Studies, University of Tampere, P.O.Box 607 SF-33101 TAMPERE, FinlandView Profile , Timo Niemi Dept. of Computer Science, University of Tampere, P.O.Box 607 SF-33101 TAMPERE, Finland Dept. of Computer Science, University of Tampere, P.O.Box 607 SF-33101 TAMPERE, FinlandView Profile Authors Info & Claims SIGIR '91: Proceedings of the 14th annual international ACM SIGIR conference on Research and development in information retrievalSeptember 1991 Pages 173–182https://doi.org/10.1145/122860.122877Online:01 September 1991Publication History 2citation356DownloadsMetricsTotal Citations2Total Downloads356Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Kalervo Järvelin, Timo Niemi |
SIGIR | 1 |
| 1989 | An approach to query cost modelling in numeric databases
Kalervo Järvelin |
JASIS | 1 |
| 1987 | Advanced tools for data conversion and database cost modelling
Kalervo Järvelin, Timo Niemi |
Inf. Manag. | 1 |
| 1986 | Cardinality estimation in numeric on-line databases
Kalervo Järvelin |
Inf. Process. Manag. | 1 |
| 1985 | Straightforward formalization of the relational model
Timo Niemi, Kalervo Järvelin |
Inf. Syst. | 2 |