Bartholomäus Wloka

dblp:65/8936 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0002-7484-878XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Eliciting analogical reasoning from language models in retrieval-augmented translation under low-resource scenarios
abstract
Retrieval-Augmented Neural Machine Translation (RANMT), which augments translation models with relevant examples fetched by a similarity retriever, is proficient in well-resourced translations. However, its inherent advantages are not fully realized in low-resource contexts, where the sparsity of data can often result in less relevant or less useful information for translation. Recent literature indicates that nearest-neighbor examples from small training data have the unfortunate effect of impairing RANMT performance. Our examination of 16 low-resource tasks reveals a sharp deterioration in performance of a multilingual language model when conditioned on retrieved examples compared to direct translation. To address this problem, we explore a framework based on analogical reasoning, aiming to enhance the capacity of language models to infer translations from parallel examples in limited data settings. This framework mimics a cognitive process of human translation by structuring examples in analogy patterns. We propose a multi-objective learning strategy that augments vanilla training for conditional translation to learn latent knowledge from examples. We also investigate different retrieval methods for selecting translation examples based on lexical similarity, semantic relatedness, or a combination of both. The results show that our approach is effective in optimizing RANMT in low-resource settings, delivering notable improvements across all retrieval settings. In particular, augmented training akin to reasoning with analogies in two directions, contributes significantly to deriving benefits from examples, even when their relevance is limited. Moreover, our approach demonstrates superior performance in low-resource translation tasks compared to prompting large language models in few-shot contexts. It also proves to be competitive with models that have been extensively trained using substantial amounts of supervised data.
Bartholomäus Wloka, Yves Lepage
Neurocomputing2
2022 WAPITI - Web-based Assignment Preparation and Instruction Tool for Interpreters
Bartholomäus Wloka, Yves Lepage, Werner Winiwarter
iiWAS1
2021 DARE - a Comprehensive Methodology for Mastering Kanji
abstract
In this paper we introduce DARE, a methodology for learning Japanese kanji characters, the study of which is an integral part of acquiring the Japanese language. Our methodology covers aspects of Dissecting, Assembling, Retrieving, and Exploring (DARE) to acquire four complementary skills, i.e. to decompose a kanji into its components, put together a kanji from basic building blocks, visually identify a kanji, and understand its usage in context. We have developed a Web-based learning environment to implement this methodology using augmented browsing to train these skills in the context of Wikipedia pages. We have created a workflow which enables language students to choose a specific topic of interest. This is done by selecting keywords which start a crawl process to gather a corpus of Wikipedia pages for this domain. This collection of text is the basis for the dynamically created learning experience.
Bartholomäus Wloka, Werner Winiwarter
iiWAS1
2021 AAA4LLL - Acquisition, Annotation, Augmentation for Lively Language Learning
abstract
HCI and NLP traditionally focus on different evaluation methods. While HCI involves a small number of people directly and deeply, NLP traditionally relies on standardized benchmark evaluations that involve a larger number of people indirectly. We present five methodological proposals at the intersection of HCI and NLP and situate them in the context of ML-based NLP models. Our goal is to foster interdisciplinary collaboration and progress in both fields by emphasizing what the fields can learn from each other.
Bartholomäus Wloka, Werner Winiwarter
LDK1
2015 Towards automated creation of high quality domain-specific machine translation resources
abstract
In this paper we present a workflow for the automated creation of parallel domain-specific corpora, i.e. multilingual translated text collections of a certain domain, in which the text pieces are aligned at sentence level. The source for the text extraction are Wikipedia articles. This workflow will be adaptable to any language pair, though the first implementation is targeting English/Japanese. The workflow consists of intelligent text acquisition, text alignment including novel techniques, and large scale quality evaluation by human experts. This will enable us to create an adaptable, fine-tuned system as well as high quality corpora, which will be compiled during the implementation.
Bartholomäus Wloka
iiWAS1
2013 DASISH: An Initiative for a European Data Humanities Infrastructure
abstract
Collaborative research, preservation of data, common access to information and recently an increasingly growing amount of data have been an issue in disciplines in which the studied material is located in many places. This is especially the case in the humanities. Digital libraries are a step towards turning hard copies into digitally available material. The need to manage, share, and preserve this information gave birth to a new term, the Digital Humanities. In this paper we present the current state, goals and future outlook of the European Digital Humanities initiative.
Bartholomäus Wloka, Werner Winiwarter, Gerhard Budin
iiWAS1
2010 Enhancing language learning and translation with ubiquitous applications
abstract
In this paper we introduce a comprehensive framework for a ubiquitous translation and language learning environment utilizing the capabilities of modern cell phone technology. We present a partial first realization of this framework: an application for learning Japanese characters and Japanese-English translation; the latter is based on results from our previous work. For the implementation, we have used the open-source Maemo platform on the Nokia N900, a state-of-the-art mobile device. We present the architecture of our framework, our current state of implementation, and the findings we have gathered so far.
Bartholomäus Wloka, Werner Winiwarter
MoMM1