Wondimagegnhue Tufa

dblp:224/5988 · also Wondimagegnhue Tsegaye, Wondimagegnhue Tsegaye Tufa · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2024
0009-0008-6883-9941ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Grounding Toxicity in Real-World Events Across Languages
Wondimagegnhue Tufa, Ilia Markov, Piek Vossen
NLDB (1)1
2023 Cross-Domain Toxic Spans Detection
Stefan F. Schouten, Baran Barbarestani, Wondimagegnhue Tufa, Piek Vossen, Ilia Markov
NLDB3
2023 A WordNet View on Crosslingual Transformers
abstract
WordNet is a database that represents relations between words and concepts as an abstraction of the contexts in which words are used.Contextualized language models represent words in contexts but leave the underlying concepts implicit.In this paper, we investigate how different layers of a pre-trained language model shape the abstract lexical relationship toward the actual contextual concept.Can we define the amount of contextualized concept forming needed given the abstracted representation of a word?Specifically, we consider samples of words with different polysemy profiles shared across three languages, assuming that words with a different polysemy profile require a different degree of concept shaping by context.We conduct probing experiments to investigate the impact of prior polysemy profiles on the representation in different layers.We analyze how contextualized models can approximate meaning through context and examine crosslingual interference effects.
Wondimagegnhue Tufa, Lisa Beinborn, Piek Vossen
GWC1
2018 Parallel Corpora for bi-lingual English-Ethiopian Languages Statistical Machine Translation
abstract
In this paper, we describe an attempt towards the development of parallel corpora for English and Ethiopian Languages, such as Amharic, Tigrigna, Afan-Oromo, Wolaytta and Ge’ez. The corpora are used for conducting a bi-directional statistical machine translation experiments. The BLEU scores of the bi-directional Statistical Machine Translation (SMT) systems show a promising result. The morphological richness of the Ethiopian languages has a great impact on the performance of SMT specially when the targets are Ethiopian languages. Now we are working towards an optimal alignment for a bi-directional English-Ethiopian languages SMT.
Solomon Teferra Abate, Michael Melese Woldeyohannis, Martha Yifiru Tachbelie, Million Meshesha, Solomon Atinafu, Wondwossen Mulugeta, Yaregal Assabie, Hafte Abera, Binyam Ephrem Seyoum, Tewodros Abebe, Wondimagegnhue Tufa, Amanuel Lemma, Tsegaye Andargie, Seifedin Shifaw
COLING11