VLDB 2026 Research / reviewers in the wild / expert
Maciej Ogrodniczuk
dblp:116/4961
· DBLP profile ↗
28ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0002-3467-9424ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 10 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LocalGovPL: A Corpus of Speaker-Attributed Polish Local Government Transcripts
Dariusz Czerski, Maciej Ogrodniczuk |
LREC | 2 |
| 2026 | Fables-DTR: A Corpus of Fables Annotated for Discourse and Temporal Relations
Purificação Silvano, António Leal, Maciej Ogrodniczuk, Aleksandra Tomaszewska, Joana Gomes, Luís Filipe Cunha, Evelin Amorim, Martyna Lewandowska, Anna Sliwicka, Alípio Mário Jorge |
LREC | 3 |
| 2024 | Polish Discourse Corpus (PDC): Corpus Design, ISO-Compliant Annotation, Data Highlights, and Parser DevelopmentabstractThis paper presents the Polish Discourse Corpus, a pioneering resource of this kind for Polish and the first corpus in Poland to employ the ISO standard for discourse relation annotation. The Polish Discourse Corpus adopts ISO 24617-8, a segment of the Language Resource Management – Semantic Annotation Framework (SemAF), which outlines a set of core discourse relations adaptable for diverse languages and genres. The paper overviews the corpus architecture, annotation procedures, the challenges that the annotators have encountered, as well as key statistical data concerning discourse relations and connectives in the corpus. It further discusses the initial phases of the discourse parser tailored for the ISO 24617-8 framework. Evaluations on the efficacy and potential refinement areas of the corpus annotation and parsing strategies are also presented. The final part of the paper touches upon anticipated research plans to improve discourse analysis techniques in the project and to conduct discourse studies involving multiple languages. Maciej Ogrodniczuk, Aleksandra Tomaszewska, Daniel Ziembicki, Sebastian Zurowski, Ryszard Tuora, Aleksandra Zwierzchowska |
LREC/COLING | 1 |
| 2024 | Universal Anaphora: The First Three YearsabstractThe aim of the Universal Anaphora initiative is to push forward the state of the art in anaphora and anaphora resolution by expanding the aspects of anaphoric interpretation which are or can be reliably annotated in anaphoric corpora, producing unified standards to annotate and encode these annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation. Although several papers on aspects of the initiative have appeared, no overall description of the initiative’s goals, proposals and achievements has been published yet except as an online draft. This paper aims to fill this gap, as well as to discuss its progress so far. Massimo Poesio, Maciej Ogrodniczuk, Vincent Ng 0001, Sameer Pradhan, Juntao Yu, Nafise Sadat Moosavi, Silviu Paun, Amir Zeldes, Anna Nedoluzhko, Michal Novák 0001, Martin Popel, Zdenek Zabokrtský, Daniel Zeman |
LREC/COLING | 2 |
| 2024 | Silver Retriever: Advancing Neural Passage Retrieval for Polish Question AnsweringabstractModern open-domain question answering systems often rely on accurate and efficient retrieval components to find passages containing the facts necessary to answer the question. Recently, neural retrievers have gained popularity over lexical alternatives due to their superior performance. However, most of the work concerns popular languages such as English or Chinese. For others, such as Polish, few models are available. In this work, we present Silver Retriever, a neural retriever for Polish trained on a diverse collection of manually or weakly labeled datasets. Silver Retriever achieves much better results than other Polish models and is competitive with larger multilingual models. Together with the model, we open-source five new passage retrieval datasets. Piotr Rybak, Maciej Ogrodniczuk |
LREC/COLING | 2 |
| 2024 | PolQA: Polish Question Answering DatasetabstractRecently proposed systems for open-domain question answering (OpenQA) require large amounts of training data to achieve state-of-the-art performance. However, data annotation is known to be time-consuming and therefore expensive to acquire. As a result, the appropriate datasets are available only for a handful of languages (mainly English and Chinese). In this work, we introduce and publicly release PolQA, the first Polish dataset for OpenQA. It consists of 7,000 questions, 87,525 manually labeled evidence passages, and a corpus of over 7,097,322 candidate passages. Each question is classified according to its formulation, type, as well as entity type of the answer. This resource allows us to evaluate the impact of different annotation choices on the performance of the QA system and propose an efficient annotation strategy that increases the passage retrieval accuracy@10 by 10.55 p.p. while reducing the annotation cost by 82%. Piotr Rybak, Piotr Przybyla, Maciej Ogrodniczuk |
LREC/COLING | 3 |
| 2023 | PolEval 2022/23 Challenge Tasks and ResultsabstractThis paper summarizes the 2022/2023 edition of PolEval -an evaluation campaign for natural language processing tools for Polish.We describe the tasks organized in this edition, which are: Punctuation prediction from conversational language, Abbreviation disambiguation and Passage Retrieval.We also discuss the datasets prepared for each of the tasks, evaluation metrics chosen to rank the submissions and also sum up the approaches chosen by the participants to tackle the tasks. Lukasz Kobylinski, Maciej Ogrodniczuk, Piotr Rybak, Piotr Przybyla, Piotr Pezik, Agnieszka Mikolajczyk, Wojciech Janowski, Michal Marcinczuk, Aleksander Smywinski-Pohl |
FedCSIS | 2 |
| 2023 | Adopting ISO 24617-8 for Discourse Relations Annotation in Polish:Challenges and Future Directions
Sebastian Zurowski, Daniel Ziembicki, Aleksandra Tomaszewska, Maciej Ogrodniczuk, Agata Drozd |
LDK | 4 |
| 2022 | Curated Multilingual Language Resources for CEF AT (CURLICAT): overall viewabstractThe work in progress on the CEF Action CURLICAT is presented. The general aim of the Action is to compile curated datasets in seven languages of the consortium in domains of relevance to European Digital Service Infrastructures (DSIs) in order to enhance the eTranslation services. Tamás Váradi, Marko Tadic, Svetla Koeva, Maciej Ogrodniczuk, Dan Tufis, Radovan Garabík, Simon Krek, Andraz Repar |
EAMT | 4 |
| 2022 | Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated SourcesabstractThis article presents the current outcomes of the CURLICAT CEF Telecom project, which aims to collect and deeply annotate a set of large corpora from selected domains. The CURLICAT corpus includes 7 monolingual corpora (Bulgarian, Croatian, Hungarian, Polish, Romanian, Slovak and Slovenian) containing selected samples from respective national corpora. These corpora are automatically tokenized, lemmatized and morphologically analysed and the named entities annotated. The annotations are uniformly provided for each language specific corpus while the common metadata schema is harmonised across the languages. Additionally, the corpora are annotated for IATE terms in all languages. The file format is CoNLL-U Plus format, containing the ten columns specific to the CoNLL-U format and three extra columns specific to our corpora as defined by Varádi et al. (2020). The CURLICAT corpora represent a rich and valuable source not just for training NMT models, but also for further studies and developments in machine learning, cross-lingual terminological data extraction and classification. Tamás Váradi, Bence Nyéki, Svetla Koeva, Marko Tadic, Vanja Stefanec, Maciej Ogrodniczuk, Bartlomiej Niton, Piotr Pezik, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan, Dan Tufis, Radovan Garabík, Simon Krek, Andraz Repar |
LREC | 6 |
| 2020 | The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual EuropeabstractMultilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cross-lingual and cross-cultural communication are still omnipresent. Language Technologies (LTs) are a powerful means to break down these barriers. While the last decade has seen various initiatives that created a multitude of approaches and technologies tailored to Europe’s specific needs, there is still an immense level of fragmentation. At the same time, AI has become an increasingly important concept in the European Information and Communication Technology area. For a few years now, AI – including many opportunities, synergies but also misconceptions – has been overshadowing every other topic. We present an overview of the European LT landscape, describing funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. We present a brief overview of the main LT-related activities on the EU level in the last ten years and develop strategic guidance with regard to four key dimensions. Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajic 0001, Khalid Choukri, Andrejs Vasiljevs, Gerhard Backfried, Christoph Prinz, José Manuél Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriute, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavrilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Marko Tadic, Dan Tufis, Tamás Váradi, Kadri Vider, Andy Way, François Yvon |
LREC | 37 |
| 2020 | The MARCELL Legislative CorpusabstractThis article presents the current outcomes of the MARCELL CEF Telecom project aiming to collect and deeply annotate a large comparable corpus of legal documents. The MARCELL corpus includes 7 monolingual sub-corpora (Bulgarian, Croatian, Hungarian, Polish, Romanian, Slovak and Slovenian) containing the total body of respective national legislative documents. These sub-corpora are automatically sentence split, tokenized, lemmatized and morphologically and syntactically annotated. The monolingual sub-corpora are complemented by a thematically related parallel corpus (Croatian-English). The metadata and the annotations are uniformly provided for each language specific sub-corpus. Besides the standard morphosyntactic analysis plus named entity and dependency annotation, the corpus is enriched with the IATE and EUROVOC labels. The file format is CoNLL-U Plus Format, containing the ten columns specific to the CoNLL-U format and four extra columns specific to our corpora. The MARCELL corpora represents a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification. Tamás Váradi, Svetla Koeva, Martin Yamalov, Marko Tadic, Bálint Sass, Bartlomiej Niton, Maciej Ogrodniczuk, Piotr Pezik, Verginica Barbu Mititelu, Radu Ion, Elena Irimia, Maria Mitrofan, Vasile Florian Pais, Dan Tufis, Radovan Garabík, Simon Krek, Andraz Repar, Matjaz Rihtar, Janez Brank |
LREC | 7 |
| 2018 | Deep Neural Networks for Coreference Resolution for Polish
Bartlomiej Niton, Pawel Morawiecki, Maciej Ogrodniczuk |
LREC | 3 |
| 2018 | Multislownik: Linking plWordNet-based Lexical Data for Lexicography and Educational PurposesabstractMultisłownik is an automated integrator of Polish lexical data retrieved from multiple available online sources intended to be used in various scenarios requiring access to such data, most prominently dictionary creation, linguistic studies and education.In contrast to many available internet dictionaries Multisłownik is WordNet-centric, capturing the core definitions from Słowosieć, the Polish Word-Net, and linking external resources to particular synsets.The paper provides details of construction of the resource, discussed the difficulties related to linking different logical structures of underlying data and investigates two sample scenarios for using the resulting platform. Maciej Ogrodniczuk, Joanna Bili'nska, Zbigniew Bronk, Witold Kieras |
GWC | 1 |
| 2017 | Multi-pass Sieve Coreference Resolution System for Polish
Bartlomiej Niton, Maciej Ogrodniczuk |
LDK | 2 |
| 2014 | Measuring Readability of Polish Texts: Baseline Experiments
Bartosz Broda, Bartlomiej Niton, Wlodzimierz Gruszczynski, Maciej Ogrodniczuk |
LREC | 4 |
| 2014 | Digital Library 2.0: Source of Knowledge and Research Collaboration Platform
Wlodzimierz Gruszczynski, Maciej Ogrodniczuk |
LREC | 2 |
| 2014 | The Polish Summaries Corpus
Maciej Ogrodniczuk, Mateusz Kopec |
LREC | 1 |
| 2014 | Polish Coreference Corpus in Numbers
Maciej Ogrodniczuk, Mateusz Kopec, Agata Savary |
LREC | 1 |
| 2014 | The Strategic Impact of META-NET on the Regional, National and International Level
Georg Rehm, Hans Uszkoreit, Sophia Ananiadou, Núria Bel, Audroné Bieleviciené, Lars Borin, António Branco, Gerhard Budin, Nicoletta Calzolari, Walter Daelemans, Radovan Garabík, Marko Grobelnik, Carmen García-Mateo, Josef van Genabith, Jan Hajic 0001, Inma Hernáez Rioja, John Judge, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Joseph Mariani, John McNaught, Maite Melero, Monica Monachini, Asunción Moreno, Jan Odijk, Maciej Ogrodniczuk, Piotr Pezik, Stelios Piperidis, Adam Przepiórkowski, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Koenraad De Smedt, Marko Tadic, Paul Thompson 0002, Dan Tufis, Tamás Váradi, Andrejs Vasiljevs, Kadri Vider, Jolanta Zabarskaite |
LREC | 29 |
| 2013 | Coreference Annotation Schema for an Inflectional Language
Maciej Ogrodniczuk, Magdalena Zawislawska, Katarzyna Glowinska, Agata Savary |
CICLing (1) | 1 |
| 2013 | A Multi-purpose Online Toolset for NLP Applications
Maciej Ogrodniczuk, Michal Lenart |
NLDB | 1 |
| 2012 | Creating a Coreference Resolution System for Polish
Mateusz Kopec, Maciej Ogrodniczuk |
LREC | 2 |
| 2012 | The Polish Sejm Corpus
Maciej Ogrodniczuk |
LREC | 1 |
| 2012 | Web Service integration platform for Polish linguistic resources
Maciej Ogrodniczuk, Michal Lenart |
LREC | 1 |
| 2012 | Towards a comprehensive open repository of Polish language resources
Maciej Ogrodniczuk, Piotr Pezik, Adam Przepiórkowski |
LREC | 1 |
| 2012 | PoliMorf: a (not so) new open morphological dictionary for Polish
Marcin Wolinski, Marcin Milkowski, Maciej Ogrodniczuk, Adam Przepiórkowski |
LREC | 3 |
| 2012 | Polish Language Processing Chains for Multilingual Information Systems
Maciej Ogrodniczuk, Adam Przepiórkowski |
NLDB | 1 |