VLDB 2026 Research / reviewers in the wild / expert
Jirí Mírovský
dblp:87/8158
· DBLP profile ↗
33ranked-venue papers
12as first author
10since 2021 · last 2026
0000-0003-2741-1347ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 12 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Using Valency Inheritance in Building a Valency Lexicon
Václava Kettnerová, Veronika Kolárová, Jirí Mírovský, Michal Olbrich |
LREC | 3 |
| 2026 | Prague Dependency Treebank - Consolidated 2.0: Enriching a Complex Annotation SchemeabstractThe Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several types of inter-sentential phenomena, especially coreference and discourse relations. We present its second consolidated version (PDT-C 2.0), which concludes almost 30-years long project of sustained development of the resource to a uniformly and coherently annotated, genre-diversified, almost 4 million token language resource of Czech language, with accompanying fully compatible lexicons. In addition to continuous linguistic research, the richly linguistically annotated corpus is also widely used in international comparisons of the development of traditional and novel NLP tools as well as in conversions into other formalisms. The corpus and the trained parsers are available under the CC BY-NC-SA licence. Marie Mikulová, Jirí Mírovský, Milan Straka, Pavlína Synková, Jan Stepánek, Barbora Stepánková, Jan Hajic 0001 |
LREC | 2 |
| 2026 | SouDeC: Source Detection and Classification in Czech
Jirí Mírovský, Barbora Hladká |
LREC | 1 |
| 2026 | Presenting the Prague Discourse Treebank 4.0
Jirí Mírovský, Pavlína Synková |
LREC | 1 |
| 2026 | DReUD: Discourse Relations in Universal Dependencies
Jirí Mírovský, Pavlína Synková |
LREC | 1 |
| 2025 | PONK: Tool for Client-Oriented Legal Writing in CzechabstractAdministrative and legal communication is often difficult for laypersons to understand, creating barriers to justice and undermining trust in public institutions. The PONK tool helps authors identify and revise unclear text using a multimodular approach that combines linguistic rules and lexical surprisal based on large language models. This integration ensures precise and adaptable detection of problematic passages. Human evaluation confirmed the tool’s effectiveness in improving legal text comprehensibility. PONK’s user-friendly interface highlights unclear segments at the word level, offering a practical solution for clearer and more accessible legal communication. Tereza Novotná, Ján Cerný, Ivan Kraus, Ivana Kvapilíková, Jirí Mírovský, Arnold Stanovský, Barbora Hladká |
JURIX | 5 |
| 2024 | Cost-Effective Discourse Annotation in the Prague Czech-English Dependency TreebankabstractWe present a cost-effective method for obtaining a high-quality annotation of explicit discourse relations in the Czech part of the Prague Czech–English Dependency Treebank, a corpus of almost 50 thousand sentences coming from the Czech translation of the Wall Street Journal part of the Penn Treebank. We use three different sources of information and combine them to obtain the discourse annotation: (i) annotation projection from the Penn Discourse Treebank 3.0, (ii) manual tectogrammatical (deep syntax) representation of sentences of the corpus, and (iii) the Lexicon of Czech Discourse Connectives CzeDLex. After solving as many discrepancies as possible automatically, the final discourse annotation is achieved by manual inspection of the remaining problematic cases. The discourse annotation of the corpus will be available both in the Prague format (on top of tectogrammatical trees) with the Prague taxonomy of discourse types, and in the Penn format (on plain texts) with the Penn Discourse Treebank 3.0 sense taxonomy. Jirí Mírovský, Pavlína Synková, Lucie Poláková, Marie Paclíková |
LREC/COLING | 1 |
| 2024 | Developing a Rhetorical Structure Theory Treebank for CzechabstractWe introduce the first version of the Czech RST Discourse Treebank, a collection of Czech journalistic texts manually annotated using the Rhetorical Structure Theory (RST), a global coherence model proposed by Mann and Thompson (1988). Each document in the corpus is represented as a single tree-like structure, where discourse units are interconnected through hierarchical rhetorical relations and their relative importance for the main purpose of a text is modeled by the nuclearity principle. The treebank is freely available in the LINDAT/CLARIAH-CZ repository under the Creative Commons license; for some documents, it includes two gold annotations representing divergent yet relevant interpretations. The paper outlines the annotation process, provides corpus statistics and evaluation, and discusses the issue of consistency associated with the global level of textual interpretation. In general, good agreement on the structure and labeling could be achieved on the lowest, local tree level and on the identification of the most central (nuclear) elementary discourse units. Disagreements mostly concerned segmentation and, in the structure, differences in the stepwise process of linking the largest text blocks. The project contributes to the advancement of RST research and its application to real-world text analysis challenges. Lucie Poláková, Jirí Mírovský, Sárka Zikánová, Eva Hajicová |
LREC/COLING | 2 |
| 2024 | Announcing the Prague Discourse Treebank 3.0abstractWe present the Prague Discourse Treebank 3.0 – a new version of the annotation of discourse relations marked by primary and secondary discourse connectives in the data of the Prague Dependency Treebank. Compared to the previous version (PDiT 2.0), the version 3.0 comes with three types of major updates: (i) it brings a largely revised annotation of discourse relations: pragmatic relations have been thoroughly reworked, many inconsistencies across all discourse types have been fixed and previously unclear cases marked in annotators’ comments have been resolved, (ii) it achieves consistency with a Lexicon of Czech Discourse Connectives (CzeDLex), and (iii) it provides the data not only in its native format (Prague Markup Language, discourse relations annotated at the top of the dependency trees), but also in the Penn Discourse Treebank 3.0 format (plain text plus a stand-off discourse annotation) and sense taxonomy. PDiT 3.0 contains 21,662 discourse relations (plus 445 list relations) in 49 thousand sentences. Pavlína Synková, Jirí Mírovský, Lucie Poláková, Magdalena Rysova |
LREC/COLING | 2 |
| 2022 | Annotating Attribution in Czech News Server ArticlesabstractThis paper focuses on detection of sources in the Czech articles published on a news server of Czech public radio. In particular, we search for attribution in sentences and we recognize attributed sources and their sentence context (signals). We organized a crowdsourcing annotation task that resulted in a data set of 2,167 stories with manually recognized signals and sources. In addition, the sources were classified into the classes of named and unnamed sources. Barbora Hladká, Jirí Mírovský, Matyás Kopp, Václav Moravec |
LREC | 2 |
| 2020 | CzeDLex 0.6 and its Representation in the PML-TQabstractCzeDLex is an electronic lexicon of Czech discourse connectives with its data coming from a large treebank annotated with discourse relations. Its new version CzeDLex 0.6 (as compared with the previous version 0.5, which was published in 2017) is significantly larger with respect to manually processed entries. Also, its structure has been modified to allow for primary connectives to appear with multiple entries for a single discourse sense. The lexicon comes in several formats, being both human and machine readable, and is available for searching in PML Tree Query, a user-friendly and powerful search tool for all kinds of linguistically annotated treebanks. The main purpose of this paper/demo is to present the new version of the lexicon and to demonstrate possibilities of mining various types of information from the lexicon using PML Tree Query; we present several examples of search queries over the lexicon data along with their results. The new version of the lexicon, CzeDLex 0.6, is available on-line and was officially released in December 2019 under the Creative Commons License. Jirí Mírovský, Lucie Poláková, Pavlína Synková |
LREC | 1 |
| 2020 | GeCzLex: Lexicon of Czech and German Anaphoric ConnectivesabstractWe introduce the first version of GeCzLex, an online electronic resource for translation equivalents of Czech and German discourse connectives. The lexicon is one of the outcomes of the research on anaphoricity and long-distance relations in discourse, it contains at present anaphoric connectives (ACs) for Czech and German connectives, and further their possible translations documented in bilingual parallel corpora (not necessarily anaphoric). As a basis, we use two existing monolingual lexicons of connectives: the Lexicon of Czech Discourse Connectives (CzeDLex) and the Lexicon of Discourse Markers (DiMLex) for German, interlink their relevant entries via semantic annotation of the connectives (according to the PDTB 3 sense taxonomy) and statistical information of translation possibilities from the Czech and German parallel data of the InterCorp project. The lexicon is, as far as we know, the first bilingual inventory of connectives with linkage on the level of individual entries, and a first attempt to systematically describe devices engaged in long-distance, non-local discourse coherence. The lexicon is freely available under the Creative Commons License. Lucie Poláková, Katerina Rysová, Magdalena Rysova, Jirí Mírovský |
LREC | 4 |
| 2019 | Connectives with Both Arguments External: A Survey on Czech
Lucie Poláková, Jirí Mírovský |
CICLing (1) | 2 |
| 2018 | Discourse Coherence Through the Lens of an Annotated Text Corpus: A Case Study
Eva Hajicová, Jirí Mírovský |
LREC | 2 |
| 2017 | Extracting a Lexicon of Discourse Connectives in Czech from an Annotated Corpus
Pavlína Synková, Magdalena Rysova, Lucie Poláková, Jirí Mírovský |
PACLIC | 4 |
| 2016 | Analysis of Word Order in Multiple Treebanks
Vladislav Kubon, Markéta Lopatková, Jirí Mírovský |
CICLing (1) | 3 |
| 2016 | Searching in the Penn Discourse Treebank Using the PML-Tree Query
Jirí Mírovský, Lucie Poláková, Jan Stepánek |
LREC | 1 |
| 2016 | Coreference in Prague Czech-English Dependency Treebank
Anna Nedoluzhko, Michal Novák 0001, Silvie Cinková, Marie Mikulová, Jirí Mírovský |
LREC | 5 |
| 2016 | Designing CzeDLex - A Lexicon of Czech Discourse Connectives
Jirí Mírovský, Pavlína Jínová, Magdalena Rysova, Lucie Poláková |
PACLIC | 1 |
| 2015 | Deletions and Node Reconstructions in a Dependency-Based Multilevel Annotation Scheme
Jan Hajic 0001, Eva Hajicová, Marie Mikulová, Jirí Mírovský, Jarmila Panevová, Daniel Zeman |
CICLing (1) | 4 |
| 2014 | Genres in the Prague Discourse Treebank
Lucie Poláková, Pavlína Jínová, Jirí Mírovský |
LREC | 3 |
| 2014 | Valency and Word Order in Czech ― A Corpus Probe
Katerina Rysová, Jirí Mírovský |
LREC | 2 |
| 2013 | (Pre-)Annotation of Topic-Focus Articulation in Prague Czech-English Dependency Treebank
Jirí Mírovský, Katerina Rysová, Magdalena Rysova, Eva Hajicová |
IJCNLP | 1 |
| 2013 | Introducing the Prague Discourse Treebank 1.0
Lucie Poláková, Jirí Mírovský, Anna Nedoluzhko, Pavlína Jínová, Sárka Zikánová, Eva Hajicová |
IJCNLP | 2 |
| 2013 | A Case Study of a Free Word Order
Vladislav Kubon, Markéta Lopatková, Jirí Mírovský |
PACLIC | 3 |
| 2012 | Interplay of Coreference and Discourse Relations: Discourse Connectives with a Referential Component
Lucie Poláková, Pavlína Jínová, Jirí Mírovský |
LREC | 3 |
| 2010 | Annotation Tool for Extended Textual Coreference and Bridging Anaphora
Jirí Mírovský, Petr Pajas, Anna Nedoluzhko |
LREC | 1 |
| 2010 | Typical Cases of Annotators' Disagreement in Discourse Annotations in Prague Dependency Treebank
Sárka Zikánová, Lucie Mladová, Jirí Mírovský, Pavlína Jínová |
LREC | 3 |
| 2008 | PDT 2.0 Requirements on a Query Language
Jirí Mírovský |
ACL | 1 |
| 2008 | Netgraph - Making Searching in Treebanks Easy
Jirí Mírovský |
IJCNLP | 1 |
| 2008 | Does Netgraph Fit Prague Dependency Treebank?
Jirí Mírovský |
LREC | 1 |
| 2005 | Automatic transcription of Czech, Russian, and Slovak spontaneous speech in the MALACH projectabstractThis paper describes the 3.5-years effort put into building LVCSR systems for recognition of spontaneous speech of Czech, Russian, and Slovak witnesses of the Holocaust in the MALACH project. For processing of colloquial, highly emotional and heavily accented speech of elderly people containing many non-speech events we have developed techniques that very effectively handle both non-speech events and colloquial and accented variants of uttered words. Manual transcripts as one of the main sources for language modeling were automatically „normalized ” using standardized lexicon, which brought about 2 to 3 % reduction of the word error rate (WER). The subsequent interpolation of such LMs with models built from an additional collection (consisting of topically selected sentences from general text corpora) resulted into an additional improvement of performance of up to 3 %. 1. Josef Psutka, Pavel Ircing, Josef V. Psutka, Jan Hajic 0001, William J. Byrne, Jirí Mírovský |
INTERSPEECH | 6 |
| 2003 | Large vocabulary ASR for spontaneous czech in the MALACH projectabstractThis paper describes LVCSR research into the automatic transcription of spontaneous Czech speech in the MALACH (Multilingual Access to Large Spoken Archives) project. This project attempts to provide improved access to the large multilingual spoken archives collected by the Survivors of the Shoah Visual History Foundation (VHF) (www.vhf.org) by advancing the state of the art in automated speech recognition. We describe a baseline ASR system and discuss the problems in language modeling that arise from the nature of Czech as a highly inflectional language that also exhibits diglossia between its written and spontaneous forms. The difficulties of this task are compounded by heavily accented, emotional and disfluent speech along with frequent switching between languages. To overcome the limited amount of relevant language model data we use statistical techniques for selecting an appropriate training corpus from a large unstructured text collection resulting in significant reductions in word error rate. 1. Josef Psutka, Pavel Ircing, Josef V. Psutka, Vlasta Radová, William J. Byrne, Jan Hajic 0001, Jirí Mírovský, Samuel Gustman |
INTERSPEECH | 7 |