Ari Pirkola

dblp:95/239 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 8 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 100%
Artificial intelligence
1 paper
Machine translation · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
cross-language information retrieval
0.132007
Frequency-based identification of correct translation equivalents (FITE) obtained through transformation rules · ACM Trans. Inf. Syst. 2007
Fuzzy translation of cross-lingual spelling variants · SIGIR 2003
The Effects of Query Structure and Dictionary Setups in Dictionary-Based Cross-Language Information Retrieval · SIGIR 1998
Information retrieval › evaluation › effectiveness metrics
discounted cumulative gain
0.112008
Intuition-supporting visualization of user's performance based on explicit negative higher-order relevance · SIGIR 2008
Information retrieval
evaluation
0.112008
Intuition-supporting visualization of user's performance based on explicit negative higher-order relevance · SIGIR 2008
Information retrieval › cross-language information retrieval
out-of-vocabulary term translation
0.112007
Frequency-based identification of correct translation equivalents (FITE) obtained through transformation rules · ACM Trans. Inf. Syst. 2007
Information retrieval › evaluation
test collection
0.012008
Intuition-supporting visualization of user's performance based on explicit negative higher-order relevance · SIGIR 2008
Information retrieval
multilingual dictionary construction
0.012007
Frequency-based identification of correct translation equivalents (FITE) obtained through transformation rules · ACM Trans. Inf. Syst. 2007
Information retrieval › cross-language information retrieval
query translation
0.011998
The Effects of Query Structure and Dictionary Setups in Dictionary-Based Cross-Language Information Retrieval · SIGIR 1998

Methods — techniques the papers use, named apart from their topics

transformation rule based translation · 0.1frequency-based identification · 0.1normalization · 0.1laboratory experiment · 0.1transformation rules · 0.0fuzzy matching · 0.0machine readable dictionary · 0.0domain-specific dictionary · 0.0
YearPublicationVenuePosition
2016 The twist measure for IR evaluation: Taking user's effort into account
abstract
We present a novel measure for ranking evaluation, called Twist (τ). It is a measure for informational intents, which handles both binary and graded relevance. τ stems from the observation that searching is currently a that searching is currently taken for granted and it is natural for users to assume that search engines are available and work well. As a consequence, users may assume the utility they have in finding relevant documents, which is the focus of traditional measures, as granted. On the contrary, they may feel uneasy when the system returns nonrelevant documents because they are then forced to do additional work to get the desired information, and this causes avoidable effort. The latter is the focus of τ, which evaluates the effectiveness of a system from the point of view of the effort required to the users to retrieve the desired information. We provide a formal definition of τ, a demonstration of its properties, and introduce the notion of effort/gain plots, which complement traditional utility‐based measures. By means of an extensive experimental evaluation, τ is shown to grasp different aspects of system performances, to not require extensive and costly assessments, and to be a robust tool for detecting differences between systems.
Nicola Ferro 0001, Gianmaria Silvello, Heikki Keskustalo, Ari Pirkola, Kalervo Järvelin
J. Assoc. Inf. Sci. Technol.4
2012 Topic-specific Web Searching based on a Real-text Dictionary
Ari Pirkola
WEBIST1
2009 Effects of Crawling Strategies on the Performance of Focused Web Crawling
Ari Pirkola, Tuomas Talvensaari
WEBIST1
2008 A Novel Implementation of the FITE-TRT Translation Method
Aki Loponen, Ari Pirkola, Kalervo Järvelin, Heikki Keskustalo
ECIR2
2008 Intuition-supporting visualization of user's performance based on explicit negative higher-order relevance
abstract
Modeling the beyond-topical aspects of relevance are currently gaining popularity in IR evaluation. For example, the discounted cumulated gain (DCG) measure implicitly models some aspects of higher-order relevance via diminishing the value of relevant documents seen later during retrieval (e.g., due to information cumulated, redundancy, and effort). In this paper, we focus on the concept of negative higher-order relevance (NHOR) made explicit via negative gain values in IR evaluation. We extend the computation of DCG to allow negative gain values, perform an experiment in a laboratory setting, and demonstrate the characteristics of NHOR in evaluation. The approach leads to intuitively reasonable performance curves emphasizing, from the user's point of view, the progression of retrieval towards success or failure. We discuss normalization issues when both positive and negative gain values are allowed and conclude by discussing the usage of NHOR to characterize test collections.
Heikki Keskustalo, Kalervo Järvelin, Ari Pirkola, Jaana Kekäläinen
SIGIR3
2008 Evaluating the effectiveness of relevance feedback based on a user simulation model: effects of a user scenario on cumulated gain value
Heikki Keskustalo, Kalervo Järvelin, Ari Pirkola
Inf. Retr.3
2008 Focused web crawling in the acquisition of comparable corpora
Tuomas Talvensaari, Ari Pirkola, Kalervo Järvelin, Martti Juhola, Jorma Laurikkala
Inf. Retr.2
2007 Frequency-based identification of correct translation equivalents (FITE) obtained through transformation rules
abstract
We devised a novel statistical technique for the identification of the translation equivalents of source words obtained by transformation rule based translation (TRT). The effectiveness of the technique called frequency-based identification of translation equivalents ( FITE ) was tested using biological and medical cross-lingual spelling variants and out-of-vocabulary (OOV) words in Spanish-English and Finnish-English TRT. The results showed that, depending on the source language and frequency corpus, FITE-TRT (the identification of translation equivalents from TRT's translation set by means of the FITE technique) may achieve high translation recall. In the case of the Web as the frequency corpus, translation recall was 89.2%--91.0% for Spanish-English FITE-TRT. For both language pairs FITE-TRT achieved high translation precision: 95.0%--98.8%. The technique also reliably identified native source language words: source words that cannot be correctly translated by TRT. Dictionary-based CLIR augmented with FITE-TRT performed substantially better than basic dictionary-based CLIR where OOV keys were kept intact. FITE-TRT with Web document frequencies was the best technique among several fuzzy translation/matching approaches tested in cross-language retrieval experiments. We also discuss the application of FITE-TRT in the automatic construction of multilingual dictionaries.
Ari Pirkola, Jarmo Toivonen, Heikki Keskustalo, Kalervo Järvelin
ACM Trans. Inf. Syst.1
2006 The Effects of Relevance Feedback Quality and Quantity in Interactive Relevance Feedback: A Simulation Based on User Modeling
Heikki Keskustalo, Kalervo Järvelin, Ari Pirkola
ECIR3
2005 Translating cross-lingual spelling variants using transformation rules
Jarmo Toivonen, Ari Pirkola, Heikki Keskustalo, Kari Visala, Kalervo Järvelin
Inf. Process. Manag.2
2004 Dictionary-Based Cross-Language Information Retrieval: Learning Experiences from CLEF 2000-2002
Turid Hedlund, Eija Airio, Heikki Keskustalo, Raija Lehtokangas, Ari Pirkola, Kalervo Järvelin
Inf. Retr.5
2003 Fuzzy translation of cross-lingual spelling variants
abstract
We will present a novel two-step fuzzy translation technique for cross-lingual spelling variants. In the first stage, transformation rules are applied to source words to render them more similar to their target language equivalents. The rules are generated automatically using translation dictionaries as source data. In the second stage, the intermediate forms obtained in the first stage are translated into a target language using fuzzy matching. The effectiveness of the technique was evaluated empirically using five source languages and English as a target language. The target word list contained 189 000 English words with the correct equivalents for the source words among them. The source words were translated using the two-step fuzzy translation technique, and the results were compared with those of plain fuzzy matching based translation. The combined technique performed better, sometimes considerably better, than fuzzy matching alone.
Ari Pirkola, Jarmo Toivonen, Heikki Keskustalo, Kari Visala, Kalervo Järvelin
SIGIR1
2003 Non-adjacent Digrams Improve Matching of Cross-Lingual Spelling Variants
Heikki Keskustalo, Ari Pirkola, Kari Visala, Erkka Leppänen, Kalervo Järvelin
SPIRE2
2003 Applying query structuring in cross-language retrieval
Ari Pirkola, Deniz Puolamäki, Kalervo Järvelin
Inf. Process. Manag.1
2001 Aspects of Swedish morphology and semantics from the perspective of mono- and cross-language information retrieval
Turid Hedlund, Ari Pirkola, Kalervo Järvelin
Inf. Process. Manag.2
2001 Dictionary-Based Cross-Language Information Retrieval: Problems, Methods, and Research Findings
Ari Pirkola, Turid Hedlund, Heikki Keskustalo, Kalervo Järvelin
Inf. Retr.1
2001 Employing the resolution power of search keys
abstract
Abstract Search key resolution power is analyzed in the context of a request, i.e., among the set of search keys for the request. Methods of characterizing the resolution power of keys automatically are studied, and the effects search keys of varying resolution power have on retrieval effectiveness are analyzed. It is shown that it often is possible to identify the best key of a query while the discrimination between the remaining keys presents problems. It is also shown that query performance is improved by suitably using the best key in a structured query. The tests were run with InQuery 1 in a subcollection of the TREC collection, which contained some 515,000 documents.
Ari Pirkola, Kalervo Järvelin
J. Assoc. Inf. Sci. Technol.1
1999 The Effects of Conjunction, Facet Structure, and Dictionary Combinations in Concept-Based Cross-Language Retrieval
Ari Pirkola, Heikki Keskustalo, Kalervo Järvelin
Inf. Retr.1
1998 The Effects of Query Structure and Dictionary Setups in Dictionary-Based Cross-Language Information Retrieval
abstract
this paper, the translation polysemy and the dictionary coverage problems were attacked by means of the combination of a general language MRD and a domain specific MR D i.e., a medical dictionary. The domain was restricted to medicine and health by choosing as test requests TREC's (see Harman, 1996) health related topics. The performance of translated Finnish queries against English documents was compared to the performance of original English queries against English documents. Because the domain was medicine and health, it was assumed that the medical dictionary disambiguates word senses, giving less incorrect senses than the general dictionary, and that it contains such search keys that are not found in the general dictionary
Ari Pirkola
SIGIR1
1996 The Effect of Anaphor and Ellipsis Resolution on Proximity Searching in a Text Database
Ari Pirkola, Kalervo Järvelin
Inf. Process. Manag.1