Maciej Piasecki

dblp:55/5861 · DBLP profile ↗
← Back
46ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0003-1503-0993ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 9 first-author · 12 since 2021Databases, data management, data science and information retrieval · 24 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Towards Complex Question Answering in Polish Language
Konrad Wojtasik, Aleksandra Domagala, Marcin Oleksy, Maciej Piasecki
ICCCI (1)4
2024 BEIR-PL: Zero Shot Information Retrieval Benchmark for the Polish Language
abstract
The BEIR dataset is a large, heterogeneous benchmark for Information Retrieval (IR), garnering considerable attention within the research community. However, BEIR and analogous datasets are predominantly restricted to English language. Our objective is to establish extensive large-scale resources for IR in the Polish language, thereby advancing the research in this NLP area. In this work, inspired by mMARCO and Mr. TyDi datasets, we translated all accessible open IR datasets into Polish, and we introduced the BEIR-PL benchmark – a new benchmark which comprises 13 datasets, facilitating further development, training and evaluation of modern Polish language models for IR tasks. We executed an evaluation and comparison of numerous IR models on the newly introduced BEIR-PL benchmark. Furthermore, we publish pre-trained open IR models for Polish language, marking a pioneering development in this field. The BEIR-PL is included in MTEB Benchmark and also available with trained models at URL https://huggingface.co/clarin-knext.
Konrad Wojtasik, Kacper Wolowiec, Vadim Shishkin, Arkadiusz Janz, Maciej Piasecki
LREC/COLING5
2023 Word Sense Disambiguation Based on Iterative Activation Spreading with Contextual Embeddings for Sense Matching
abstract
Many knowledge-based solutions were proposed to solve Word Sense Disambiguation (WSD) problem with limited annotated resources.Such WSD algorithms are able to cover very large sense repositories, but still being outperformed by supervised ones on benchmark data.In this paper, we start with analysis identifying key properties and issues in application of spreading activation algorithms in knowledge-based WSD, e.g.influence of the network local structures, interaction with context information and sense frequency.Taking our observations as a point of departure, we introduce a novel solution with new contextto-sense matching using BERT embeddings, iterative parallel spreading activation function and selective sense alignment using contextual BERT embeddings.The proposed solution obtains performance beyond the state-of-the-art for the contemporary knowledge-based WSD approaches for both English and Polish data.
Arkadiusz Janz, Maciej Piasecki
GWC2
2023 Lexicalised and non-lexicalized multi-word expressions in WordNet: a cross-encoder approach
abstract
Focusing on recognition of multi-word expressions (MWEs), we address the problem of recording MWEs in WordNet.In fact, not all MWEs recorded in that lexical database could with no doubt be considered as lexicalised (e.g.elements of wordnet taxonomy, quantifier phrases, certain collocations).In this paper, we use a cross-encoder approach to improve our earlier method of distinguishing between lexicalised and non-lexicalised MWEs found in WordNet using custom-designed rulebased and statistical approaches.We achieve F1-measure for the class of lexicalised word combinations close to 80%, easily beating two baselines (random and a majority class one).Language model also proves to be better than a feature-based logistic regression model.
Marek Maziarz, Lukasz Grabowski, Tadeusz Piotrowski, Ewa Rudnicka, Maciej Piasecki
GWC5
2023 Wordnet-oriented recognition of derivational relations
abstract
Derivational relations are an important element in defining meanings, as they help to explore word-formation schemes and predict senses of derivates (derived words).In this work, we analyse different methods of representing derivational forms obtained from WordNetfrom quantitative vectors to contextual learned embedding methods -and compare ways of classifying the derivational relations occurring between them.Our research focuses on the explainability of the obtained representations and results.The data source for our research is plWordNet, which is the wordnet of the Polish language and includes a rich set of derivation examples.
Wiktor Walentynowicz, Maciej Piasecki
GWC2
2023 Wordnet for Definition Augmentation with Encoder-Decoder Architecture
abstract
Data augmentation is a difficult task in Natural Language Processing.Simple methods that can be relatively easily applied in other domains like insertion, deletion or substitution, mostly result in changing the sentence meaning significantly and obtaining an incorrect example.Wordnets are potentially a perfect source of rich and high quality data that when integrated with the powerful capacity of generative models can help to solve this complex task.In this work, we use plWordNet, which is a wordnet of the Polish language, to explore the capability of encoder-decoder architectures in data augmentation of sense glosses.We discuss the limitations of generative methods and perform qualitative review of generated data samples.
Konrad Wojtasik, Arkadiusz Janz, Maciej Piasecki
GWC3
2022 Non-Contextual vs Contextual Word Embeddings in Multiword Expressions Detection
Maciej Piasecki, Kamil Kanclerz
ICCCI1
2022 Context-free Transformer-based Generative Lemmatiser for Polish
Wiktor Walentynowicz, Maciej Piasecki, Artur Kot
ICCCI2
2022 This is the way: designing and compiling LEPISZCZE, a comprehensive NLP benchmark for Polish
abstract
The availability of compute and data to train larger and larger language models increases the demand for robust methods of benchmarking the true progress of LM training. Recent years witnessed significant progress in standardized benchmarking for English. Benchmarks such as GLUE, SuperGLUE, or KILT have become a de facto standard tools to compare large language models. Following the trend to replicate GLUE for other languages, the KLEJ benchmark\ (klej is the word for glue in Polish) has been released for Polish. In this paper, we evaluate the progress in benchmarking for low-resourced languages. We note that only a handful of languages have such comprehensive benchmarks. We also note the gap in the number of tasks being evaluated by benchmarks for resource-rich English/Chinese and the rest of the world.In this paper, we introduce LEPISZCZE (lepiszcze is the Polish word for glew, the Middle English predecessor of glue), a new, comprehensive benchmark for Polish NLP with a large variety of tasks and high-quality operationalization of the benchmark.We design LEPISZCZE with flexibility in mind. Including new models, datasets, and tasks is as simple as possible while still offering data versioning and model tracking. In the first run of the benchmark, we test 13 experiments (task and dataset pairs) based on the five most recent LMs for Polish. We use five datasets from the Polish benchmark and add eight novel datasets. As the paper's main contribution, apart from LEPISZCZE, we provide insights and experiences learned while creating the benchmark for Polish as the blueprint to design similar benchmarks for other low-resourced languages.
Lukasz Augustyniak, Kamil Tagowski, Albert Sawczyn, Denis Janiak, Roman Bartusiak, Adrian Szymczak, Arkadiusz Janz, Piotr Szymanski, Marcin Watroba, Mikolaj Morzy, Tomasz Kajdanowicz, Maciej Piasecki
NeurIPS12
2021 Literary Genre Recognition among Polish Blog Posts
abstract
Robust methods have been proposed for content and topic-based text classification, as well authorship attribution in stylometry. However, the problem of a fine-grained literary genre (style) recognition is much less studied. We present several approaches to the recognition of eight literary genres manually annotated in a large corpus of Polish blogs. Different text representations were combined with neural network classifiers, including deep, recursive neural networks. Very good results were achieved for the representation of blog posts with the help of pre-trained fastText word embeddings and the Bi-GRU recursive deep neural network as a classifier. As the observed good performance of this classifier could be a result of topical bias across genres, experiments on a selected sub-corpus with a reduced dominance of the most frequent topic were also conducted with no significant change observed.
Edyta Rogula, Maciej Piasecki, Tomasz Naskret
KES2
2021 Neural Language Models vs Wordnet-based Semantically Enriched Representation in CST Relation Recognition
abstract
Neural language models, including transformer-based models, that are pretrained on very large corpora became a common way to represent text in various tasks, including recognition of textual semantic relations, e.g.Cross-document Structure Theory.Pre-trained models are usually fine tuned to downstream tasks and the obtained vectors are used as an input for deep neural classifiers.No linguistic knowledge obtained from resources and tools is utilised.In this paper we compare such universal approaches with a combination of rich graph-based linguistically motivated sentence representation and a typical neural network classifier applied to a task of recognition of CST relation in Polish.The representation describes selected levels of the sentence structure including description of lexical meanings on the basis of the wordnet (plWordNet) synsets and connected SUMO concepts.The obtained results show that in the case of difficult relations and medium size training corpus semantically enriched text representation leads to significantly better results.
Arkadiusz Janz, Maciej Piasecki, Piotr Watorski
GWC2
2021 A (Non)-Perfect Match: Mapping plWordNet onto PrincetonWordNet
abstract
The paper reports on the methodology and final results of a large-scale synset mapping between plWordNet and Princeton WordNet.Dedicated manual and semi-automatic mapping procedures as well as interlingual relation types for nouns, verbs, adjectives and adverbs are described.The statistics of all types of interlingual relations are also provided.
Ewa Rudnicka, Wojciech Witkowski, Maciej Piasecki
GWC3
2020 Automated Bilingual Linking of Wordnet Senses
Maciej Piasecki, Roman Dyszlewski, Ewa Rudnicka
ICCCI1
2020 Brand-Product Relation Extraction Using Heterogeneous Vector Space Representations
abstract
Relation Extraction is a fundamental NLP task. In this paper we investigate the impact of underlying text representation on the performance of neural classification models in the task of Brand-Product relation extraction. We also present the methodology of preparing annotated textual corpora for this task and we provide valuable insight into the properties of Brand-Product relations existing in textual corpora. The problem is approached from a practical angle of applications Relation Extraction in facilitating commercial Internet monitoring.
Arkadiusz Janz, Lukasz Kopoci'nski, Maciej Piasecki, Agnieszka Pluwak
LREC3
2019 A Comparison of Sense-level Sentiment Scores
abstract
In this paper, we compare a variety of sense-tagged sentiment resources, including SentiWordNet, ML-Senticon, plWord-Net emo and the NTU Multilingual Corpus.The goal is to investigate the quality of the resources and see how well the sentiment polarity annotation maps across languages.
Francis Bond, Arkadiusz Janz, Maciej Piasecki
GWC3
2019 plWordNet 4.1 - a Linguistically Motivated, Corpus-based Bilingual Resource
abstract
The paper presents the latest release of the Polish WordNet, namely plWord-Net 4.1.The most significant developments since 3.0 version include new relations for nouns and verbs, mapping semantic role-relations from the valency lexicon Walenty onto the plWord-Net structure and sense-level interlingual mapping.Several statistics are presented in order to illustrate the development and contemporary state of the wordnet.
Agnieszka Dziob, Maciej Piasecki, Ewa Rudnicka
GWC2
2019 WordNet2Vec: Corpora agnostic word vectorization method
Roman Bartusiak, Lukasz Augustyniak, Tomasz Kajdanowicz, Przemyslaw Kazienko, Maciej Piasecki
Neurocomputing5
2018 Classifier-based Polarity Propagation in a WordNet
Jan Kocon, Arkadiusz Janz, Maciej Piasecki
LREC3
2018 Recognition of Hyponymy and Meronymy Relations in Word Embeddings for Polish
abstract
Word embeddings were used for the extraction of hyponymy relation in several approaches, but also it was recently shown that they should not work, in fact.In our work we verified both claims using a very large wordnet of Polish as a gold standard for lexico-semantic relations and word embeddings extracted from a very large corpus of Polish.We showed that a hyponymy extraction method based on linear regression classifiers trained on clusters of vectors can be successfully applied on large scale.We presented also a possible explanation for contradictory findings in the literature.Moreover, in order to show the feasibility of the method we extended it to the recognition of meronymy.
Gabriela Czachor, Maciej Piasecki, Arkadiusz Janz
GWC2
2018 Implementation of the Verb Model in plWordNet 4.0
abstract
The paper presents an expansion of the verb model for plWordNet -the wordnet of Polish.A modified system of constitutive features (register, aspect and verb classes), synset and lexical relations is presented.A special attention is given to the proposed new relations and changes in the verb classification.We discuss also the results of its verification by application to the description of a relatively large sample of Polish verbs.The model introduces a new class of relations, namely non-constitutive synset relations that are shared among lexical units, but describe, not define synsets.The proposed model is compared to the entailment relations in other wordnets, and the description of verbs based on valency frames.
Agnieszka Dziob, Maciej Piasecki
GWC2
2018 Context-sensitive Sentiment Propagation in WordNet
abstract
In this paper we present a comprehensive overview of recent methods of the sentiment propagation in a wordnet.Next, we propose a fully automated method called Classifier-based Polarity Propagation, which utilises a very rich set of features, where most of them are based on wordnet relation types, multi-level bag-ofsynsets and bag-of-polarities.We have evaluated our solution using manually annotated part of plWordNet 3.1 emo, which contains more than 83k manual sentiment annotations, covering more than 41k synsets.We have demonstrated that in comparison to existing rule-based methods using a specific narrow set of semantic relations our method has achieved statistically significant and better results starting with the same seed synsets.
Jan Kocon, Arkadiusz Janz, Maciej Piasecki
GWC3
2018 Towards Mapping Thesauri onto plWordNet
abstract
plWordNet, the wordnet of Polish, has become a very comprehensive description of the Polish lexical system.This paper presents a plan of its semi-automated integration with thesauri, terminological databases and ontologies, as a further necessary step in its development.This will improve linking of plWordNet into Linked Open Data, and facilitate applications in, e.g., WSD, keyword extraction or automated metadata generation.We present an overview of resources relevant to Polish and a plan for their linking to plWordNet.
Marek Maziarz, Maciej Piasecki
GWC2
2018 WordnetLoom - a Multilingual Wordnet Editing System Focused on Graph-based Presentation
abstract
The paper presents a new re-built and expanded, version 2.0 of WordnetLoom -an open wordnet editor.It facilitates work on a multilingual system of wordnets, is based on efficient software architecture of thin client, and offers more flexibility in enriching wordnet representation.This new version is built on the experience collected during the use of the previous one for more than 10 years of plWordNet development.We discuss its extensions motivated by the collected experience.A special focus is given to the development of a variant for the needs of MultiWordnet of Portuguese, which is based on a very different wordnet development model.
Tomasz Naskret, Agnieszka Dziob, Maciej Piasecki, Chakaveh Saedi, António Branco
GWC3
2018 Wordnet-based Evaluation of Large Distributional Models for Polish
abstract
The paper presents construction of large scale test datasets for word embeddings on the basis of a very large wordnet.They were next applied for evaluation of word embedding models and used to assess and compare the usefulness of different word embeddings extracted from a very large corpus of Polish.We analysed also and compared several publicly available models described in literature.In addition, several large word embeddings models built on the basis of a very large Polish corpus are presented.
Maciej Piasecki, Gabriela Czachor, Arkadiusz Janz, Dominik Kaszewski, Pawel Kedzia
GWC1
2018 Lexical Perspective on Wordnet to Wordnet Mapping
abstract
The paper presents a feature-based model of equivalence targeted at (manual) sense linking between Princeton WordNet and plWordNet.The model incorporates insights from lexicographic and translation theories on bilingual equivalence and draws on the results of earlier synsetlevel mapping of nouns between Princeton WordNet and plWordNet.It takes into account all basic aspects of language such as form, meaning and function and supplements them with (parallel) corpus frequency and translatability.Three types of equivalence are distinguished, namely strong, regular and weak depending on the conformity with the proposed features.The presented solutions are languageneutral and they can be easily applied to language pairs other than Polish and English.Sense-level mapping is a more finegrained mapping than the existing synset mappings and is thus of great potential to human and machine translation.
Ewa Rudnicka, Francis Bond, Lukasz Grabowski, Maciej Piasecki, Tadeusz Piotrowski
GWC4
2018 Towards Emotive Annotation in plWordNet 4.0
abstract
The paper presents an approach to building a very large emotive lexicon for Polish based on plWordNet.An expanded annotation model is discussed, in which lexical units (word senses) are annotated with basic emotions, fundamental human values and sentiment polarisation.The annotation process is performed manually in the 2+1 scheme by pairs of linguists and psychologies.Guidelines referring to the usage in corpora, substitution tests as well linguistic properties of lexical units (e.g.derivational associations) are discussed.Application of the model in a substantial extension of the emotive annotation of plWordNet is presented.The achieved high inter-annotator agreement shows that with relatively small workload a promising emotive resource can be created.
Monika Zasko-Zielinska, Maciej Piasecki
GWC2
2016 plWordNet 3.0 - a Comprehensive Lexical-Semantic Resource
abstract
We have released plWordNet 3.0, a very large wordnet for Polish. In addition to what is expected in wordnets – richly interrelated synsets – it contains sentiment and emotion annotations, a large set of multi-word expressions, and a mapping onto WordNet 3.1. Part of the release is enWordNet 1.0, a substantially enlarged copy of WordNet 3.1, with material added to allow for a more complete mapping. The paper discusses the design principles of plWordNet, its content, its statistical portrait, a comparison with similar resources, and a partial list of applications.
Marek Maziarz, Maciej Piasecki, Ewa Rudnicka, Stan Szpakowicz, Pawel Kedzia
COLING2
2016 plWordNet in Word Sense Disambiguation task
abstract
The paper explores the application of plWordNet, a very large wordnet of Polish, in weakly supervised Word Sense Disambiguation (WSD).Because plWord-Net provides only partial descriptions by glosses and usage examples, and does not include sense-disambiguated glosses, PageRank-based WSD methods perform slightly worse than for English.However, we show that the use of weights for the relation types and the order in which lexical units have been added for sense re-ranking can significantly improve WSD precision.The evaluation was done on two Polish corpora (KPWr and Składnica) including manual WSD.We discuss the fundamental difference in the construction of both corpora and very different test results.
Maciej Piasecki, Pawel Kedzia, Marlena Orlinska
GWC1
2016 plWordNet 3.0 - Almost There
abstract
It took us nearly ten years to get from no wordnet for Polish to the largest wordnet ever built.We started small but quickly learned to dream big.Now we are about to release plWordNet 3.0-emo -complete with sentiment and emotions annotatedand a domestic version of Princeton Word-Net, larger than WordNet 3.1 by nearly ten thousand newly added words.The paper retraces the road we travelled and talks a little about the future.
Maciej Piasecki, Stan Szpakowicz, Marek Maziarz, Ewa Rudnicka
GWC1
2014 Ruled-based, Interlingual Motivated Mapping of plWordNet onto SUMO Ontology
Pawel Kedzia, Maciej Piasecki
LREC2
2014 plWordNet as the Cornerstone of a Toolkit of Lexico-semantic Resources
abstract
A wordnet is many things to many people: a graph of inter-related lexicalised concepts, a taxonomy, a thesaurus, and so on.A wordnet makes good sense as the mainstay of any deep automated semantic analysis of text.We have begun the construction of a multi-component, multi-use toolkit of natural language processing tools with plWordNet, a very large Polish wordnet, at its centre.The components will include plWordNet and its mapping onto an ontology (the upper level and elements of the middle level), a lexicon of proper names and a semantic valency lexicon.Some of those elements will be aligned with plWordNet, and there will be a mapping onto Princeton WordNet.Several challenging applications will show the utility of the toolkit in practice.
Marek Maziarz, Maciej Piasecki, Ewa Rudnicka, Stan Szpakowicz
GWC2
2014 Registers in the System of Semantic Relations in plWordNet
abstract
Lexicalised concepts are represented in wordnets by word-sense pairs.The strength of markedness is one of the factors which influence word use.Stylistically unmarked words are largely contextneutral.Technical terms, obsolete words, "officialese", slangs, obscenities and so on are all marked, often strongly, and that limits their use considerably.We discuss the position of register and markedness in wordnets with respect to semantic relations, and we list typical values of register.We illustrate the discussion with the system of registers in plWordNet, the largest Polish wordnet.We present a decision tree for the assignment of marking labels, and examine the consistency of the editing decisions based on that tree.
Marek Maziarz, Maciej Piasecki, Ewa Rudnicka, Stan Szpakowicz
GWC2
2013 Relational Propagation of Word Sentiment in WordNet
Andrzej Misiaszek, Tomasz Kajdanowicz, Przemyslaw Kazienko, Maciej Piasecki
CDVE4
2012 Tools for plWordNet Development. Presentation and Perspectives
Bartosz Broda, Marek Maziarz, Maciej Piasecki
LREC3
2012 Constraint Based Description of Polish Multiword Expressions
Roman Kurc, Maciej Piasecki, Bartosz Broda
LREC2
2012 Recognition of Polish Derivational Relations Based on Supervised Learning Scheme
Maciej Piasecki, Radoslaw Ramocki, Marek Maziarz
LREC1
2011 Linguistically Informed Mining Lexical Semantic Relations from Wikipedia Structure
Maciej Piasecki, Agnieszka Indyka-Piasecka, Roman Kurc
ACIIDS (1)1
2011 Heterogeneous Knowledge Sources in Graph-Based Expansion of the Polish Wordnet
Maciej Piasecki, Roman Kurc, Bartosz Broda
ACIIDS (1)1
2010 Building a Node of the Accessible Language Technology Infrastructure
Bartosz Broda, Michal Marcinczuk, Maciej Piasecki
LREC3
2010 Resource and Service Centres as the Backbone for a Sustainable Service Infrastructure
Peter Wittenburg, Núria Bel, Lars Borin, Gerhard Budin, Nicoletta Calzolari, Eva Hajicová, Kimmo Koskenniemi, Lothar Lemnitzer, Bente Maegaard, Maciej Piasecki, Jean-Marie Pierrel, Stelios Piperidis, Inguna Skadina, Dan Tufis, Remco van Veenendaal, Tamás Váradi, Martin Wynne
LREC10
2008 Corpus-based Semantic Relatedness for the Construction of Polish WordNet
Bartosz Broda, Magdalena Derwojedowa, Maciej Piasecki, Stan Szpakowicz
LREC3
2007 Correction of Medical Handwriting OCR Based on Semantic Similarity
Bartosz Broda, Maciej Piasecki
IDEAL2
2006 Application of syntactic properties to three-level recognition of polish hand-written medical texts
abstract
In the paper, three-level hand-writing recognition using language syntactic properties on the upper level is presented. Isolated characters are recognized on the lowest level. The character classification from the lowest level is used in words recognition. Words are recognized using a combined classifier based on possibly incomplete unigram lexicon. Word classifier builds a rank of the most likely words. Ranks created for subsequent words are input to the syntactic classifier, which recognizes the whole sentences. Here the local syntactic constraints are used to build a syntactically consistent sentence. The method has been applied to recognition of hand-written medical texts describing fixed aspects of patient treatment. Due to narrow area of topics explained in the texts and peculiarity of style characteristic for physicians writing texts, the syntax of expected sentences is relatively simple, what makes the problem of checking the syntactic consistency simpler.
Grzegorz Godlewski, Maciej Piasecki, Jerzy Sas
ACM Symposium on Document Engineering2
2005 Distributed Service - Oriented Architecture for Information Extraction System "Semanta"
abstract
Our objective is to provide a flexible, scalable, distributed architecture that assures a high performance for information extraction (IE) systems working in Internet. The architecture is based on both the general paradigm of the service-oriented architecture, client-server approach and strong separation of concerns between storage and processing components. An experimental IE system, named Semanta, utilising the proposed architecture is also presented. In the following document, we describe five main Semanta services, which are Web user interface (WebUI), Web crawler service (WCS), parsing service (PS), IE service and manager
Lukasz Jastrzebski, Maciej Piasecki, Grzegorz Strzelecki
ISDA2
2005 Logo -The Modular Conversational Agent Understanding Polish
abstract
Our paper presents a modular architecture for a natural dialogue system. The architecture is applied in the construction of Logo - a dialogue system for Polish, based on the Discourse Representation theory, implementing some aspects of the pragmatic analysis and enabling communication with an agent acting in the virtual reality of a blocks world. Logo deals also with the selected issues of coreference resolution. The proposed architecture is intended to be flexible and open for utilisation of diverged language resources.
Maciej Piasecki, Ireneusz Matysiak, Anna Rusak
ISDA1
2000 Modelling Multimedia Presentation in UML
Ludwik Kuzniarz, Maciej Piasecki
EJC2