Jarmo Toivonen

dblp:21/1451 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Artificial intelligence
1 paper
Machine translation · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
cross-language information retrieval
0.122007
Frequency-based identification of correct translation equivalents (FITE) obtained through transformation rules · ACM Trans. Inf. Syst. 2007
Fuzzy translation of cross-lingual spelling variants · SIGIR 2003
Information retrieval › cross-language information retrieval
out-of-vocabulary term translation
0.112007
Frequency-based identification of correct translation equivalents (FITE) obtained through transformation rules · ACM Trans. Inf. Syst. 2007
Information retrieval
multilingual dictionary construction
0.012007
Frequency-based identification of correct translation equivalents (FITE) obtained through transformation rules · ACM Trans. Inf. Syst. 2007

Methods — techniques the papers use, named apart from their topics

transformation rule based translation · 0.1frequency-based identification · 0.1transformation rules · 0.0fuzzy matching · 0.0
YearPublicationVenuePosition
2007 Frequency-based identification of correct translation equivalents (FITE) obtained through transformation rules
abstract
We devised a novel statistical technique for the identification of the translation equivalents of source words obtained by transformation rule based translation (TRT). The effectiveness of the technique called frequency-based identification of translation equivalents ( FITE ) was tested using biological and medical cross-lingual spelling variants and out-of-vocabulary (OOV) words in Spanish-English and Finnish-English TRT. The results showed that, depending on the source language and frequency corpus, FITE-TRT (the identification of translation equivalents from TRT's translation set by means of the FITE technique) may achieve high translation recall. In the case of the Web as the frequency corpus, translation recall was 89.2%--91.0% for Spanish-English FITE-TRT. For both language pairs FITE-TRT achieved high translation precision: 95.0%--98.8%. The technique also reliably identified native source language words: source words that cannot be correctly translated by TRT. Dictionary-based CLIR augmented with FITE-TRT performed substantially better than basic dictionary-based CLIR where OOV keys were kept intact. FITE-TRT with Web document frequencies was the best technique among several fuzzy translation/matching approaches tested in cross-language retrieval experiments. We also discuss the application of FITE-TRT in the automatic construction of multilingual dictionaries.
Ari Pirkola, Jarmo Toivonen, Heikki Keskustalo, Kalervo Järvelin
ACM Trans. Inf. Syst.2
2005 Translating cross-lingual spelling variants using transformation rules
Jarmo Toivonen, Ari Pirkola, Heikki Keskustalo, Kari Visala, Kalervo Järvelin
Inf. Process. Manag.1
2003 Fuzzy translation of cross-lingual spelling variants
abstract
We will present a novel two-step fuzzy translation technique for cross-lingual spelling variants. In the first stage, transformation rules are applied to source words to render them more similar to their target language equivalents. The rules are generated automatically using translation dictionaries as source data. In the second stage, the intermediate forms obtained in the first stage are translated into a target language using fuzzy matching. The effectiveness of the technique was evaluated empirically using five source languages and English as a target language. The target word list contained 189 000 English words with the correct equivalents for the source words among them. The source words were translated using the two-step fuzzy translation technique, and the results were compared with those of plain fuzzy matching based translation. The combined technique performed better, sometimes considerably better, than fuzzy matching alone.
Ari Pirkola, Jarmo Toivonen, Heikki Keskustalo, Kari Visala, Kalervo Järvelin
SIGIR2