Marek Kubis

dblp:15/8157 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-2016-2598ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 What Students Want from AI Assistants: Insights from a Short Self-Directed Machine Learning Course
Jacek Marciniak, Barbara Kolodziejczak, Marcin Szczepanski, Marek Kubis, Dorota Marciniak, Michal Gulczynski, Adam Szpilkowski, Adam Wieczarek
CSEDU (2)4
2024 Using Bibliodata LODification to Create Metadata-Enriched Literary Corpora in Line with FAIR Principles
abstract
This paper discusses the design principles and procedures for creating a balanced corpus for research in computational literary studies, building on the experience of computational linguistics but adapting it to the specificities of the digital humanities. It showcases the development of the Metadata-enriched Polish Novel Corpus from the 19th and 20th centuries (19/20MetaPNC), consisting of 1,000 novels from 1854–1939, as an illustrative case and proposes a comprehensive workflow for the creation and reuse of literary corpora. What sets 19/20MetaPNC apart is its approach to balance, which considers the spatial dimension, the inclusion of non-canonical texts previously overlooked by other corpora, and the use of a complex, multi-stage metadata enrichment and verification process. Emphasis is placed on research-oriented metadata design, efficient data collection and data sharing according to the FAIR principles as well as 5- and 7-star data standards to increase the visibility and reusability of the corpus. A knowledge graph-based solution for the creation of exchangeable and machine-readable metadata describing corpora has been developed. For this purpose, metadata from bibliographic catalogs and other sources were transformed into Linked Data following the bibliodata LODification approach.
Agnieszka Karlinska, Cezary Rosinski, Marek Kubis, Patryk Hubar, Jan Wieczorek
LREC/COLING3
2024 Spoken Language Corpora Augmentation with Domain-Specific Voice-Cloned Speech
Mateusz Czyznikiewicz, Lukasz Bondaruk, Jakub Kubiak, Adam Wiacek, Lukasz Degórski, Marek Kubis, Pawel Skórzewski
FedCSIS6
2023 Back Transcription as a Method for Evaluating Robustness of Natural Language Understanding Models to Speech Recognition Errors
abstract
In a spoken dialogue system, an NLU model is preceded by a speech recognition system that can deteriorate the performance of natural language understanding.This paper proposes a method for investigating the impact of speech recognition errors on the performance of natural language understanding models.The proposed method combines the back transcription procedure with a fine-grained technique for categorizing the errors that affect the performance of NLU models.The method relies on the usage of synthesized speech for NLU evaluation.We show that the use of synthesized speech in place of audio recording does not change the outcomes of the presented technique in a significant way.
Marek Kubis, Pawel Skórzewski, Marcin Sowanski, Tomasz Zietkiewicz
EMNLP1
2023 Center for Artificial Intelligence Challenge on Conversational AI Correctness
abstract
This paper describes a challenge on Conversational AI correctness with the goal to develop Natural Language Understanding models that are robust against speech recognition errors.The data for the competition consist of natural language utterances along with semantic frames that represent the commands targeted at a virtual assistant.The specification of the task is given along with the data preparation procedure and the evaluation rules.The baseline models for the task are discussed and the results of the competition are reported.
Marek Kubis, Pawel Skórzewski, Marcin Sowanski, Tomasz Zietkiewicz
FedCSIS1
2022 Part-of-Speech Models Compression Methods for on-Device Grapheme-to-Phoneme Conversion
abstract
The paper investigates methods of compressing part-of-speech models that are developed for an on-device grapheme-to-phoneme conversion module. The performance of part-of-speech models is analyzed under different compression regimes. The evaluation is done with respect to French, German and Italian datasets that consist of TTS input prompts. The study shows that a proper selection of a compression method reduces the model size significantly without deteriorating the grapheme-to-phoneme conversion performance.
Marek Kubis, Maxime Méloux, Pawel Skórzewski, Marcin Lewandowski, Gunu Jho, Hyoungmin Park
ICASSP1
2020 Development of Kazakh Named Entity Recognition Models
Darkhan Akhmed-Zaki, Madina Mansurova, Vladimir B. Barakhnin, Marek Kubis, Darya Chikibayeva, Marzhan Kyrgyzbayeva
ICCCI4
2020 Noetic end-to-end response selection with supervised neural network based classifiers and unsupervised similarity models
abstract
This paper describes a solution for the Noetic End-to-End Response Selection challenge – one of the tasks of the 7th Dialog System Technology Challenge. The goal of the task is to select the most appropriate continuation of a dialogue from a given set of responses. We approach this problem by building an ensemble of supervised neural network based classifiers and unsupervised similarity models. The dialogue continuation is selected according to a score that aggregates the rankings of candidate responses determined by the models in the ensemble.
Pawel Skórzewski, Weronika Sieinska, Marek Kubis
Comput. Speech Lang.3
2012 A Query Language for WordNet-Like Lexical Databases
Marek Kubis
ACIIDS (3)1
2010 PolNet - Polish WordNet: Data and Tools
Zygmunt Vetulani, Marek Kubis, Tomasz Obrêbski
LREC2