Milan Straka

dblp:18/1522 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-3295-5576ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 MorfFlex: Handling Rich Morphology
abstract
We present MorfFlex, a morphological dictionary architecture suitable for languages with extensive regularity in both inflection and derivation. As the primary example of MorfFlex in use we introduce MorfFlex CZ, a morphological dictionary of Czech. It is distributed as a simple, unstructured list of triplets, however, its manually maintained, unpublished source files and conversion scripts encode a sophisticated system of inflectional and derivational patterns. These patterns dramatically reduce the otherwise enormous size of the dictionary, which currently contains over 100 million wordforms and more than 1 million lemmas. The MorfFlex CZ dictionary serves as an essential resource for ensuring the consistency of manual morphological annotation in the Prague Dependency Treebanks and underpins state-of-the-art automatic tools such as MorphoDiTa. In this paper, we focus on: (i) presenting an effective method for managing the rich morphological system within the dictionary, and (ii) demonstrating the utility of such a language resource for maintaining annotation consistency in corpora and supporting the development of advanced NLP applications.
Jaroslava Hlavácová, Marie Mikulová, Barbora Stepánková, Milan Straka, Jan Hajic 0001
LREC4
2026 Prague Dependency Treebank - Consolidated 2.0: Enriching a Complex Annotation Scheme
abstract
The Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several types of inter-sentential phenomena, especially coreference and discourse relations. We present its second consolidated version (PDT-C 2.0), which concludes almost 30-years long project of sustained development of the resource to a uniformly and coherently annotated, genre-diversified, almost 4 million token language resource of Czech language, with accompanying fully compatible lexicons. In addition to continuous linguistic research, the richly linguistically annotated corpus is also widely used in international comparisons of the development of traditional and novel NLP tools as well as in conversions into other formalisms. The corpus and the trained parsers are available under the CC BY-NC-SA licence.
Marie Mikulová, Jirí Mírovský, Milan Straka, Pavlína Synková, Jan Stepánek, Barbora Stepánková, Jan Hajic 0001
LREC3
2026 Meet UD_Czech-PDTC: A Large and Genre-Rich Treebank in Universal Dependencies
abstract
Czech has been part of Universal Dependencies since its first release in 2015. It has also been one of the best represented languages, with the Prague Dependency Treebank being order of magnitude larger than most other UD treebanks. More recently, three other datasets from the Prague family were added and the annotations thoroughly revisited, forming the "Prague Dependency Treebank-Consolidated" (PDT-C). In comparison to the original PDT, PDT-C is more than twice as large, but it is also much more diverse in terms of genres and domains. In this paper, we describe the conversion of the new resource to Universal Dependencies. While the two annotation schemes are relatively similar at the first sight, there are numerous small differences in topology of the dependency structures and in granularity of the POS and relation type inventories. We demonstrate a selection of such differences on examples, discuss the diverging motivations, as well as ways to overcome the differences during conversion. We argue that while PDT is less "universal" and more tightly bound to one language, its multi-layer annotation is rich and provides all information needed for basic UD trees, and much more.
Marie Mikulová, Barbora Stepánková, Daniel Zeman, Jan Stepánek, Milan Straka
LREC5
2024 Practical End-to-End Optical Music Recognition for Pianoform Music
Jirí Mayer, Milan Straka, Jan Hajic jr., Pavel Pecina
ICDAR (6)2
2024 beeFormer: Bridging the Gap Between Semantic and Interaction Similarity in Recommender Systems
abstract
Recommender systems often use text-side information to improve their predictions, especially in cold-start or zero-shot recommendation scenarios, where traditional collaborative filtering approaches cannot be used. Many approaches to text-mining side information for recommender systems have been proposed over recent years, with sentence Transformers being the most prominent one. However, these models are trained to predict semantic similarity without utilizing interaction data with hidden patterns specific to recommender systems. In this paper, we propose beeFormer, a framework for training sentence Transformer models with interaction data. We demonstrate that our models trained with beeFormer can transfer knowledge between datasets while outperforming not only semantic similarity sentence Transformers but also traditional collaborative filtering methods. We also show that training on multiple datasets from different domains accumulates knowledge in a single model, unlocking the possibility of training universal, domain-agnostic sentence Transformer models to mine text representations for recommender systems. We release the source code, trained models, and additional details allowing replication of our experiments at https://github.com/recombee/beeformer.
Vojtech Vancura, Pavel Kordík, Milan Straka
RecSys3
2024 CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking
abstract
We present CWRCzech, Click Web Ranking dataset for Czech, a 100M query-document Czech click dataset for relevance ranking with user behavior data collected from search engine logs of Seznam.cz. To the best of our knowledge, CWRCzech is the largest click dataset with raw text published so far. It provides document positions in the search results as well as information about user behavior: 27.6M clicked documents and 10.8M dwell times. In addition, we also publish a manually annotated Czech test for the relevance task, containing nearly 50k query-document pairs, each annotated by at least 2 annotators. Finally, we analyze how the user behavior data improve relevance ranking and show that models trained on data automatically harnessed at sufficient scale can surpass the performance of models trained on human annotated data. CWRCzech is published under an academic non-commercial license and is available to the research community at https://github.com/seznam/CWRCzech.
Josef Vonásek, Milan Straka, Rostislav Krc, Lenka Lasonová, Ekaterina Egorova, Jana Straková, Jakub Náplava
SIGIR2
2022 Quality and Efficiency of Manual Annotation: Pre-annotation Bias
abstract
This paper presents an analysis of annotation using an automatic pre-annotation for a mid-level annotation complexity task - dependency syntax annotation. It compares the annotation efforts made by annotators using a pre-annotated version (with a high-accuracy parser) and those made by fully manual annotation. The aim of the experiment is to judge the final annotation quality when pre-annotation is used. In addition, it evaluates the effect of automatic linguistically-based (rule-formulated) checks and another annotation on the same data available to the annotators, and their influence on annotation quality and efficiency. The experiment confirmed that the pre-annotation is an efficient tool for faster manual syntactic annotation which increases the consistency of the resulting annotation without reducing its quality.
Marie Mikulová, Milan Straka, Jan Stepánek, Barbora Stepánková, Jan Hajic 0001
LREC2
2022 Czech Grammar Error Correction with a Large and Diverse Corpus
abstract
Abstract We introduce a large and diverse Czech corpus annotated for grammatical error correction (GEC) with the aim to contribute to the still scarce data resources in this domain for languages other than English. The Grammar Error Correction Corpus for Czech (GECCC) offers a variety of four domains, covering error distributions ranging from high error density essays written by non-native speakers, to website texts, where errors are expected to be much less common. We compare several Czech GEC systems, including several Transformer-based ones, setting a strong baseline to future research. Finally, we meta-evaluate common GEC metrics against human judgments on our data. We make the new Czech GEC corpus publicly available under the CC BY-SA 4.0 license at http://hdl.handle.net/11234/1-4639.
Jakub Náplava, Milan Straka, Jana Straková, Alexandr Rosen
Trans. Assoc. Comput. Linguistics2
2020 Prague Dependency Treebank - Consolidated 1.0
abstract
We present a richly annotated and genre-diversified language resource, the Prague Dependency Treebank-Consolidated 1.0 (PDT-C 1.0), the purpose of which is - as it always been the case for the family of the Prague Dependency Treebanks - to serve both as a training data for various types of NLP tasks as well as for linguistically-oriented research. PDT-C 1.0 contains four different datasets of Czech, uniformly annotated using the standard PDT scheme (albeit not everything is annotated manually, as we describe in detail here). The texts come from different sources: daily newspaper articles, Czech translation of the Wall Street Journal, transcribed dialogs and a small amount of user-generated, short, often non-standard language segments typed into a web translator. Altogether, the treebank contains around 180,000 sentences with their morphological, surface and deep syntactic annotation. The diversity of the texts and annotations should serve well the NLP applications as well as it is an invaluable resource for linguistic research, including comparative studies regarding texts of different genres. The corpus is publicly and freely available.
Jan Hajic 0001, Eduard Bejcek, Jaroslava Hlavácová, Marie Mikulová, Milan Straka, Jan Stepánek, Barbora Stepánková
LREC5
2019 Neural Architectures for Nested NER through Linearization
abstract
We propose two neural network architectures for nested named entity recognition (NER), a setting in which named entities may overlap and also be labeled with more than one label.We encode the nested labels using a linearized scheme.In our first proposed approach, the nested labels are modeled as multilabels corresponding to the Cartesian product of the nested labels in a standard LSTM-CRF architecture.In the second one, the nested NER is viewed as a sequence-to-sequence problem, in which the input sequence consists of the tokens and output sequence of the labels, using hard attention on the word whose label is being predicted.The proposed methods outperform the nested NER state of the art on four corpora: ACE-2004, ACE-2005, GENIA and Czech CNEC.We also enrich our architectures with the recently published contextual embeddings: ELMo, BERT and Flair, reaching further improvements for the four nested entity corpora.In addition, we report flat NER stateof-the-art results for CoNLL-2002 Dutch and Spanish and for CoNLL-2003 English.
Jana Straková, Milan Straka, Jan Hajic 0001
ACL (1)2
2019 75 Languages, 1 Model: Parsing Universal Dependencies Universally
abstract
Dan Kondratyuk, Milan Straka. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Daniel Kondratyuk, Milan Straka
EMNLP/IJCNLP (1)2
2018 LemmaTag: Jointly Tagging and Lemmatizing for Morphologically Rich Languages with BRNNs
abstract
We present LemmaTag, a featureless neural network architecture that jointly generates part-of-speech tags and lemmas for sentences by using bidirectional RNNs with characterlevel and word-level embeddings.We demonstrate that both tasks benefit from sharing the encoding part of the network, predicting tag subcategories, and using the tagger output as an input to the lemmatizer.We evaluate our model across several languages with complex morphology, which surpasses state-of-the-art accuracy in both part-of-speech tagging and lemmatization in Czech, German, and Arabic.
Daniel Kondratyuk, Tomas Gavenciak, Milan Straka, Jan Hajic 0001
EMNLP3
2018 Using Adversarial Examples in Natural Language Processing
Petr Belohlávek, Ondrej Plátek, Zdenek Zabokrtský, Milan Straka
LREC4
2018 Diacritics Restoration Using Neural Networks
Jakub Náplava, Milan Straka, Pavel Stranák, Jan Hajic 0001
LREC2
2018 SumeCzech: Large Czech News-Based Summarization Dataset
Milan Straka, Nikita Mediankin, Tom Kocmi, Zdenek Zabokrtský, Vojtech Hudecek, Jan Hajic 0001
LREC1
2016 UDPipe: Trainable Pipeline for Processing CoNLL-U Files Performing Tokenization, Morphological Analysis, POS Tagging and Parsing
Milan Straka, Jan Hajic 0001, Jana Straková
LREC1
2016 Merging Data Resources for Inflectional and Derivational Morphology in Czech
Zdenek Zabokrtský, Magda Sevcíková, Milan Straka, Jonás Vidra, Adéla Limburská
LREC3
2013 Stop-probability estimates computed on a large corpus improve Unsupervised Dependency Parsing
David Marecek, Milan Straka
ACL (1)2
2010 The performance of the Haskell containers package
abstract
In this paper, we perform a thorough performance analysis of the containers package, the de facto standard Haskell containers library, comparing it to the most of existing alternatives on HackageDB. We then significantly improve its performance, making it comparable to the best implementations available. Additionally, we describe a new persistent data structure based on hashing, which offers the best performance out of available data structures containing Strings and ByteStrings.
Milan Straka
Haskell1
2007 Linear-Time Ranking of Permutations
Martin Mares, Milan Straka
ESA2