Marek Maziarz

dblp:22/8936 · DBLP profile ↗
← Back
11ranked-venue papers in the field
6as first author
5since 2021 · last 2023
0000-0003-0318-2869ORCID · corroborated

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 10 (6 first)Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2023 Data Augmentation Method for Boosting Multilingual Word Sense Disambiguation
abstract
Recent advances in Word Sense Disambiguation suggest neural language models can be successfully improved by incorporating knowledge base structure.Such class of models are called hybrid solutions.We propose a method of improving hybrid WSD models by harnessing data augmentation techniques and bilingual training.The data augmentation consist of structure augmentation using interlingual connections between wordnets and text data augmentation based on multilingual glosses and usage examples.We utilise language-agnostic neural model trained both with SemCor and Princeton WordNet gloss and example corpora, as well as with Polish WordNet glosses and usage examples.This augmentation technique proves to make well-known hybrid WSD architecture to be competitive, when compared to current State-of-the-Art models, even more complex.
Arkadiusz Janz, Marek Maziarz
GWC2
2023 Lexicalised and non-lexicalized multi-word expressions in WordNet: a cross-encoder approach
abstract
Focusing on recognition of multi-word expressions (MWEs), we address the problem of recording MWEs in WordNet.In fact, not all MWEs recorded in that lexical database could with no doubt be considered as lexicalised (e.g.elements of wordnet taxonomy, quantifier phrases, certain collocations).In this paper, we use a cross-encoder approach to improve our earlier method of distinguishing between lexicalised and non-lexicalised MWEs found in WordNet using custom-designed rulebased and statistical approaches.We achieve F1-measure for the class of lexicalised word combinations close to 80%, easily beating two baselines (random and a majority class one).Language model also proves to be better than a feature-based logistic regression model.
Marek Maziarz, Lukasz Grabowski, Tadeusz Piotrowski, Ewa Rudnicka, Maciej Piasecki
GWC1
2021 Discriminating Homonymy from Polysemy in Wordnets: English, Spanish and Polish Nouns
abstract
We propose a novel method of homonymypolysemy discrimination for three Indo-European Languages (English, Spanish and Polish).Support vector machines and LASSO logistic regression were successfully used in this task, outperforming baselines.The feature set utilised lemma properties, gloss similarities, graph distances and polysemy patterns.The proposed ML models performed equally well for English and the other two languages (constituting testing data sets).The algorithms not only ruled out most cases of homonymy but also were efficacious in distinguishing between closer and indirect semantic relatedness.
Arkadiusz Janz, Marek Maziarz
GWC2
2021 Testing agreement between lexicographers: A case of homonymy and polysemy
abstract
In this paper we compare Oxford Lexico and Merriam Webster dictionaries with Princeton WordNet with respect to the description of semantic (dis)similarity between polysemous and homonymous senses that could be inferred from them.WordNet lacks any explicit description of polysemy or homonymy, but as a network of linked senses it may be used to compute semantic distances between word senses.To compare WordNet with the dictionaries, we transformed sample entry microstructures of the latter into graphs and crosslinked them with the equivalent senses of the former.We found that dictionaries are in high agreement with each other, if one considers polysemy and homonymy altogether, and in moderate concordance, if one focuses merely on polysemy descriptions.Measuring the shortest path lengths on WordNet gave results comparable to those on the dictionaries in predicting semantic dissimilarity between polysemous senses, but was less felicitous while recognising homonymy.
Marek Maziarz, Francis Bond, Ewa Rudnicka
GWC1
2021 Mapping WordNet onto human brain connectome in emotion processing and semantic similarity recognition
abstract
In this article we extend a WordNet structure with relations linking synsets to Desikan’s brain regions. Based on lexicographer files and WordNet Domains the mapping goes from synset semantic categories to behavioural and cognitive functions and then directly to brain lobes. A human brain connectome (HBC) adjacency matrix was utilised to capture transition probabilities between brain regions. We evaluated the new structure in several tasks related to semantic similarity and emotion processing using brain-expanded Princeton WordNet (207k LUs) and Polish WordNet (285k LUs, 30k annotated with valence, arousal and 8 basic emotions). A novel HBC vector representation turned out to be significantly better than proposed baselines.
Jan Kocon, Marek Maziarz
Inf. Process. Manag.2
2019 Testing Zipf's meaning-frequency law with wordnets as sense inventories
abstract
According to George K. Zipf, more frequent words have more senses.We have tested this law using corpora and wordnets of English, Spanish, Portuguese, French, Polish, Japanese, Indonesian and Chinese.We have proved that the law works pretty well for all of these languages if we takeas Zipf did -mean values of meaning count and averaged ranks.On the other hand, the law disastrously fails in predicting the number of senses for a single lemma.We have also provided the evidence that slope coefficients of Zipfian log-log linear model may vary from language to language.
Francis Bond, Arkadiusz Janz, Marek Maziarz, Ewa Rudnicka
GWC3
2018 Towards Mapping Thesauri onto plWordNet
abstract
plWordNet, the wordnet of Polish, has become a very comprehensive description of the Polish lexical system.This paper presents a plan of its semi-automated integration with thesauri, terminological databases and ontologies, as a further necessary step in its development.This will improve linking of plWordNet into Linked Open Data, and facilitate applications in, e.g., WSD, keyword extraction or automated metadata generation.We present an overview of resources relevant to Polish and a plan for their linking to plWordNet.
Marek Maziarz, Maciej Piasecki
GWC1
2016 Adverbs in plWordNet: Theory and Implementation
abstract
Adverbs are seldom well represented in wordnets.Princeton WordNet, for example, derives from adjectives practically all its adverbs and whatever involvement they have.GermaNet stays away from this part of speech.Adverbs in plWordNet will be emphatically present in all their semantic and syntactic distinctness.We briefly discuss the linguistic background of the lexical system of Polish adverbs.We describe an automated generator of accurate candidate adverbs, and introduce the lexicographic procedures which will ensure high consistency of wordnet editors' decisions about adverbs.lation hypo(gor ączkowo 1 , nerwowo 1 ), then, is an instance of hyponymy in plWordNet.Listing 1: Hyponymy.Modifier of intentional verbs.Jeżeli ktoś/coś robi coś x, to robi to y. Jeżeli ktoś/coś robi coś y, to niekoniecznie robi to x.'If someone/something does something x, they do it y.' 'If someone/something does something y, they do not necessarily do it x.' Listing 2: Hyponymy.Modifier of unintentional verbs.Jeżeli coś dzieje się x, to dzieje się y.Jeżeli coś dzieje się y, to niekoniecznie dzieje się x.'If something happens x, it happens y.' 'If something happens y, it does not necessarily happen x.' Listing 3: Hyponymy.Adjective modifier.Jeżeli ktoś/coś jest x jakiś, to jest też y jakiś.Jeżeli ktoś/coś jest y jakiś, to niekoniecznie jest x jakiś.'If someone/something is x so, they are also y so.' 'If someone/something is y so, they are not necessarily x so.' Listing 4: Hyponymy.Predicative adverb.Jeżeli jest x, to jest też y.Jeżeli jest y, to niekoniecznie jest x.
Marek Maziarz, Stan Szpakowicz, Michal Kalinski
GWC1
2016 plWordNet 3.0 - Almost There
abstract
It took us nearly ten years to get from no wordnet for Polish to the largest wordnet ever built.We started small but quickly learned to dream big.Now we are about to release plWordNet 3.0-emo -complete with sentiment and emotions annotatedand a domestic version of Princeton Word-Net, larger than WordNet 3.1 by nearly ten thousand newly added words.The paper retraces the road we travelled and talks a little about the future.
Maciej Piasecki, Stan Szpakowicz, Marek Maziarz, Ewa Rudnicka
GWC3
2014 plWordNet as the Cornerstone of a Toolkit of Lexico-semantic Resources
abstract
A wordnet is many things to many people: a graph of inter-related lexicalised concepts, a taxonomy, a thesaurus, and so on.A wordnet makes good sense as the mainstay of any deep automated semantic analysis of text.We have begun the construction of a multi-component, multi-use toolkit of natural language processing tools with plWordNet, a very large Polish wordnet, at its centre.The components will include plWordNet and its mapping onto an ontology (the upper level and elements of the middle level), a lexicon of proper names and a semantic valency lexicon.Some of those elements will be aligned with plWordNet, and there will be a mapping onto Princeton WordNet.Several challenging applications will show the utility of the toolkit in practice.
Marek Maziarz, Maciej Piasecki, Ewa Rudnicka, Stan Szpakowicz
GWC1
2014 Registers in the System of Semantic Relations in plWordNet
abstract
Lexicalised concepts are represented in wordnets by word-sense pairs.The strength of markedness is one of the factors which influence word use.Stylistically unmarked words are largely contextneutral.Technical terms, obsolete words, "officialese", slangs, obscenities and so on are all marked, often strongly, and that limits their use considerably.We discuss the position of register and markedness in wordnets with respect to semantic relations, and we list typical values of register.We illustrate the discussion with the system of registers in plWordNet, the largest Polish wordnet.We present a decision tree for the assignment of marking labels, and examine the consistency of the editing decisions based on that tree.
Marek Maziarz, Maciej Piasecki, Ewa Rudnicka, Stan Szpakowicz
GWC1