VLDB 2026 Research / reviewers in the wild / expert
Jan Hajic 0001
dblp:40/1286 · also Jan Hajic Sr.
· DBLP profile ↗
64ranked-venue papers
11as first author
12since 2021 · last 2026
0000-0002-3503-7730ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 61 · 10 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MorfFlex: Handling Rich MorphologyabstractWe present MorfFlex, a morphological dictionary architecture suitable for languages with extensive regularity in both inflection and derivation. As the primary example of MorfFlex in use we introduce MorfFlex CZ, a morphological dictionary of Czech. It is distributed as a simple, unstructured list of triplets, however, its manually maintained, unpublished source files and conversion scripts encode a sophisticated system of inflectional and derivational patterns. These patterns dramatically reduce the otherwise enormous size of the dictionary, which currently contains over 100 million wordforms and more than 1 million lemmas. The MorfFlex CZ dictionary serves as an essential resource for ensuring the consistency of manual morphological annotation in the Prague Dependency Treebanks and underpins state-of-the-art automatic tools such as MorphoDiTa. In this paper, we focus on: (i) presenting an effective method for managing the rich morphological system within the dictionary, and (ii) demonstrating the utility of such a language resource for maintaining annotation consistency in corpora and supporting the development of advanced NLP applications. Jaroslava Hlavácová, Marie Mikulová, Barbora Stepánková, Milan Straka, Jan Hajic 0001 |
LREC | 5 |
| 2026 | Prague Dependency Treebank - Consolidated 2.0: Enriching a Complex Annotation SchemeabstractThe Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several types of inter-sentential phenomena, especially coreference and discourse relations. We present its second consolidated version (PDT-C 2.0), which concludes almost 30-years long project of sustained development of the resource to a uniformly and coherently annotated, genre-diversified, almost 4 million token language resource of Czech language, with accompanying fully compatible lexicons. In addition to continuous linguistic research, the richly linguistically annotated corpus is also widely used in international comparisons of the development of traditional and novel NLP tools as well as in conversions into other formalisms. The corpus and the trained parsers are available under the CC BY-NC-SA licence. Marie Mikulová, Jirí Mírovský, Milan Straka, Pavlína Synková, Jan Stepánek, Barbora Stepánková, Jan Hajic 0001 |
LREC | 7 |
| 2026 | HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained ModelsabstractWe present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for some 20 languages, and through end-to-end evaluation of various language model architectures trained on this data. For multilingual LLM evaluation, we provide a comprehensive collection of benchmarks for nine European languages, with special emphasis on natively created tasks, mechanisms to mitigate prompt sensitivity, and refined normalization and aggregation of scores. Additionally, we train and evaluate a family of 57 monolingual encoder–decoder models, as well as about 30 “smallish” monolingual GPT-like reference models. Besides the monolingual data and models, we also present a very large collection of parallel texts automatically mined from this data, together with a novel parallel corpus synthesized via machine translation. Stephan Oepen, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Maja Buljan, Laurie Burchell, Lucas Georges Gabriel Charpentier, Pinzhen Chen, Mariia Fedorova, Ona de Gibert Bonet, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Andrey Kutuzov, Veronika Laippala, Bhavitvya Malik, Vladislav Mikhailov, Amanda Myntti, Dayyán O'Brien, Lucie Poláková, Gema Ramírez-Sánchez, Janine Siewert, Pavel Stepachev, Jörg Tiedemann, Teemu Vahtola, Dusan Varis, Fedor Vitiugin, Jaume Zaragoza |
LREC | 12 |
| 2026 | Automatic Suggestions Help Extending Eventive Ontology: A Case Study on SynSemClass
Jana Straková, Eva Fucíková, Zdenka Uresová, Jan Hajic 0001 |
LREC | 4 |
| 2025 | An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)abstractLaurie Burchell, Ona De Gibert Bonet, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajič, Jindřich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Petter Mæhlum, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Nikita Moghe, Amanda Myntti, Dayyán O’Brien, Stephan Oepen, Proyag Pal, Jousia Piha, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dušan Variš, Tereza Vojtěchová, Jaume Zaragoza-Bernabeu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Laurie Burchell, Ona de Gibert Bonet, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Petter Mæhlum, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Nikita Moghe, Amanda Myntti, Dayyán O'Brien, Stephan Oepen, Proyag Pal, Jousia Piha, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dusan Varis, Tereza Vojtechová, Jaume Zaragoza-Bernabeu |
ACL (1) | 10 |
| 2025 | HPLT's Second Data ReleaseabstractWe describe the progress of the High Performance Language Technologies (HPLT) project, a 3-year EU-funded project that started in September 2022. We focus on the up-to-date results on the release of free text datasets derived from web crawls, one of the central objectives of the project. The second release used a revised processing pipeline, and an enlarged set of input crawls. From 4.5 petabytes of web crawls we extracted 7.6T tokens of monolingual text in 193 languages, plus 380 million parallel sentences in 51 language pairs. We also release MultiHPLT, a cross-combination of the parallel data, which produces 1,275 pairs, as well as releasing the containing documents for all parallel sentences in order to enable research in document-level MT. We report changes in the pipeline, analysis and evaluation results for the second parallel data release based on machine translation systems. All datasets are released under a permissive CC0 licence. Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Laurie Burchell, Pinzhen Chen, Mariia Fedorova, Ona de Gibert Bonet, Liane Guillou, Barry Haddow, Jan Hajic 0001, Jindrich Helcl, Erik Henriksson, Andrey Kutuzov, Veronika Laippala, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Amanda Myntti, Dayyán O'Brien, Stephan Oepen, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dusan Varis, Jaume Zaragoza-Bernabeu |
MTSummit (2) | 10 |
| 2024 | Building a Broad Infrastructure for Uniform Meaning RepresentationsabstractThis paper reports the first release of the UMR (Uniform Meaning Representation) data set. UMR is a graph-based meaning representation formalism consisting of a sentence-level graph and a document-level graph. The sentence-level graph represents predicate-argument structures, named entities, word senses, aspectuality of events, as well as person and number information for entities. The document-level graph represents coreferential, temporal, and modal relations that go beyond sentence boundaries. UMR is designed to capture the commonalities and variations across languages and this is done through the use of a common set of abstract concepts, relations, and attributes as well as concrete concepts derived from words from invidual languages. This UMR release includes annotations for six languages (Arapaho, Chinese, English, Kukama, Navajo, Sanapana) that vary greatly in terms of their linguistic properties and resource availability. We also describe on-going efforts to enlarge this data set and extend it to other genres and modalities. We also briefly describe the available infrastructure (UMR annotation guidelines and tools) that others can use to create similar data sets. Julia Bonn, Matthew J. Buchholz, Jayeol Chun, Andrew Cowell, William Croft 0001, Lukas Denk, Sijia Ge, Jan Hajic 0001, Kenneth Lai, James H. Martin, Skatje Myers, Alexis Palmer, Martha Palmer, Claire Benet Post, James Pustejovsky, Kristine Stenzel, Haibo Sun, Zdenka Uresová, Rosa Vallejos, Jens E. L. Van Gysel, Meagan Vigus, Nianwen Xue, Jin Zhao 0009 |
LREC/COLING | 8 |
| 2024 | Textual Coverage of Eventive Entries in Lexical Semantic ResourcesabstractThis short paper focuses on the coverage of eventive entries (verbs, predicates, etc.) of some well-known lexical semantic resources when applied to random running texts taken from the internet. While coverage gaps are often reported for manually created lexicons (which is the case of most semantically-oriented lexical ones), it was our aim to quantify these gaps, cross-lingually, on a new purely textual resource set produced by the HPLT Project from crawled internet data. Several English, German, Spanish and Czech lexical semantic resources (which, for the most part, focus on verbs and predicates) have been selected for this experiment. We also describe the challenges related to the fact that these resources are (to a varying extent) semantically oriented, meaning that the texts have to be preprocessed to obtain lemmas (base forms) and some types of MWEs before the coverage can be reasonably evaluated, and thus the results are necessarily only approximate. The coverage of these resources, with some exclusions as described in the paper, range from 41.00% to 97.33%, confirming the need to expand at least some - even well-known - resources to cover the prevailing source of today’s textual resources with regard to lexical units describing events or states (or possibly other eventive mentions). Eva Fucíková, Cristina Fernández Alcaina, Jan Hajic 0001, Zdenka Uresová |
LREC/COLING | 3 |
| 2023 | What's the Meaning of Superhuman Performance in Today's NLU?abstractSimone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajič, Daniel Hershcovich, Eduard Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, Roberto Navigli. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajic 0001, Daniel Hershcovich, Eduard H. Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, Roberto Navigli |
ACL (1) | 4 |
| 2022 | Overview of the ELE ProjectabstractThis paper provides an overview of the ongoing European Language Equality(ELE) project, an 18-month action funded by the European Commission which involves 52 partners. The primary goal of ELE is to prepare the European Language Equality Programme, in the form of a strategic research, innovation and implementation agenda and a roadmap for achieving full digital language equality (DLE) in Europe by 2030. Itziar Aldabe, Jane Dunne, Aritz Farwell, Owen Gallagher, Federico Gaspari, Maria Giagkou, Jan Hajic 0001, Jens Peter Kückens, Teresa Lynn, Georg Rehm, German Rigau, Katrin Marheinecke, Stelios Piperidis, Natália Resende, Tereza Vojtechová, Andy Way |
EAMT | 7 |
| 2022 | Quality and Efficiency of Manual Annotation: Pre-annotation BiasabstractThis paper presents an analysis of annotation using an automatic pre-annotation for a mid-level annotation complexity task - dependency syntax annotation. It compares the annotation efforts made by annotators using a pre-annotated version (with a high-accuracy parser) and those made by fully manual annotation. The aim of the experiment is to judge the final annotation quality when pre-annotation is used. In addition, it evaluates the effect of automatic linguistically-based (rule-formulated) checks and another annotation on the same data available to the annotators, and their influence on annotation quality and efficiency. The experiment confirmed that the pre-annotation is an efficient tool for faster manual syntactic annotation which increases the consistency of the resulting annotation without reducing its quality. Marie Mikulová, Milan Straka, Jan Stepánek, Barbora Stepánková, Jan Hajic 0001 |
LREC | 5 |
| 2022 | Making a Semantic Event-type Ontology MultilingualabstractWe present an extension of the SynSemClass Event-type Ontology, originally conceived as a bilingual Czech-English resource. We added German entries to the classes representing the concepts of the ontology. Having a different starting point than the original work (unannotated parallel corpus without links to a valency lexicon and, of course, different existing lexical resources), it was a challenge to adapt the annotation guidelines, the data model and the tools used for the original version. We describe the process and results of working in such a setup. We also show the next steps to adapt the annotation process, data structures and formats and tools necessary to make the addition of a new language in the future more smooth and efficient, and possibly to allow for various teams to work on SynSemClass extensions to many languages concurrently. We also present the latest release which contains the results of adding German, freely available for download as well as for online access. Zdenka Uresová, Karolina Zaczynska, Peter Bourgonje, Eva Fucíková, Georg Rehm, Jan Hajic 0001 |
LREC | 6 |
| 2020 | Prague Dependency Treebank - Consolidated 1.0abstractWe present a richly annotated and genre-diversified language resource, the Prague Dependency Treebank-Consolidated 1.0 (PDT-C 1.0), the purpose of which is - as it always been the case for the family of the Prague Dependency Treebanks - to serve both as a training data for various types of NLP tasks as well as for linguistically-oriented research. PDT-C 1.0 contains four different datasets of Czech, uniformly annotated using the standard PDT scheme (albeit not everything is annotated manually, as we describe in detail here). The texts come from different sources: daily newspaper articles, Czech translation of the Wall Street Journal, transcribed dialogs and a small amount of user-generated, short, often non-standard language segments typed into a web translator. Altogether, the treebank contains around 180,000 sentences with their morphological, surface and deep syntactic annotation. The diversity of the texts and annotations should serve well the NLP applications as well as it is an invaluable resource for linguistic research, including comparative studies regarding texts of different genres. The corpus is publicly and freely available. Jan Hajic 0001, Eduard Bejcek, Jaroslava Hlavácová, Marie Mikulová, Milan Straka, Jan Stepánek, Barbora Stepánková |
LREC | 1 |
| 2020 | Universal Dependencies v2: An Evergrowing Multilingual Treebank CollectionabstractUniversal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, universal part-of-speech tags, and standardized morphological features; and a syntactic layer focusing on syntactic relations between predicates, arguments and modifiers. In this paper, we describe version 2 of the universal guidelines (UD v2), discuss the major changes from UD v1 to UD v2, and give an overview of the currently available treebanks for 90 languages. Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajic 0001, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster 0001, Francis M. Tyers, Daniel Zeman |
LREC | 4 |
| 2020 | European Language Grid: An OverviewabstractWith 24 official EU and many additional languages, multilingualism in Europe and an inclusive Digital Single Market can only be enabled through Language Technologies (LTs). European LT business is dominated by hundreds of SMEs and a few large players. Many are world-class, with technologies that outperform the global players. However, European LT business is also fragmented – by nation states, languages, verticals and sectors, significantly holding back its impact. The European Language Grid (ELG) project addresses this fragmentation by establishing the ELG as the primary platform for LT in Europe. The ELG is a scalable cloud platform, providing, in an easy-to-integrate way, access to hundreds of commercial and non-commercial LTs for all European languages, including running tools and services as well as data sets and resources. Once fully operational, it will enable the commercial and non-commercial European LT community to deposit and upload their technologies and data sets into the ELG, to deploy them through the grid, and to connect with other resources. The ELG will boost the Multilingual Digital Single Market towards a thriving European LT community, creating new jobs and opportunities. Furthermore, the ELG project organises two open calls for up to 20 pilot projects. It also sets up 32 national competence centres and the European LT Council for outreach and coordination purposes. Georg Rehm, Maria Berger, Ela Elsholz, Stefanie Hegele, Florian Kintzel, Katrin Marheinecke, Stelios Piperidis, Miltos Deligiannis, Dimitrios Galanis, Katerina Gkirtzou, Penny Labropoulou, Kalina Bontcheva, Jan Hajic 0001, Jana Hamrlová, Lukás Kacena, Khalid Choukri, Victoria Arranz, Andrejs Vasiljevs, Orians Anvari, Andis Lagzdins, Julija Melnika, Gerhard Backfried, Erinç Dikici, Miroslav Jánosík, Katja Prinz, Christoph Prinz, Severin Stampler, Dorothea Thomas-Aniola, José Manuél Gómez-Pérez, Andrés García-Silva, Cristian Berrio, Ulrich Germann, Steve Renals, Ondrej Klejch |
LREC | 15 |
| 2020 | The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual EuropeabstractMultilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cross-lingual and cross-cultural communication are still omnipresent. Language Technologies (LTs) are a powerful means to break down these barriers. While the last decade has seen various initiatives that created a multitude of approaches and technologies tailored to Europe’s specific needs, there is still an immense level of fragmentation. At the same time, AI has become an increasingly important concept in the European Information and Communication Technology area. For a few years now, AI – including many opportunities, synergies but also misconceptions – has been overshadowing every other topic. We present an overview of the European LT landscape, describing funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. We present a brief overview of the main LT-related activities on the EU level in the last ten years and develop strategic guidance with regard to four key dimensions. Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajic 0001, Khalid Choukri, Andrejs Vasiljevs, Gerhard Backfried, Christoph Prinz, José Manuél Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriute, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavrilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Marko Tadic, Dan Tufis, Tamás Váradi, Kadri Vider, Andy Way, François Yvon |
LREC | 6 |
| 2019 | Neural Architectures for Nested NER through LinearizationabstractWe propose two neural network architectures for nested named entity recognition (NER), a setting in which named entities may overlap and also be labeled with more than one label.We encode the nested labels using a linearized scheme.In our first proposed approach, the nested labels are modeled as multilabels corresponding to the Cartesian product of the nested labels in a standard LSTM-CRF architecture.In the second one, the nested NER is viewed as a sequence-to-sequence problem, in which the input sequence consists of the tokens and output sequence of the labels, using hard attention on the word whose label is being predicted.The proposed methods outperform the nested NER state of the art on four corpora: ACE-2004, ACE-2005, GENIA and Czech CNEC.We also enrich our architectures with the recently published contextual embeddings: ELMo, BERT and Flair, reaching further improvements for the four nested entity corpora.In addition, we report flat NER stateof-the-art results for CoNLL-2002 Dutch and Spanish and for CoNLL-2003 English. Jana Straková, Milan Straka, Jan Hajic 0001 |
ACL (1) | 3 |
| 2018 | Synonymy in Bilingual Context: The CzEngClass LexiconabstractThis paper describes CzEngClass, a bilingual lexical resource being built to investigate verbal synonymy in bilingual context and to relate semantic roles common to one synonym class to verb arguments (verb valency). In addition, the resource is linked to existing resources with the same of a similar aim: English and Czech WordNet, FrameNet, PropBank, VerbNet (SemLink), and valency lexicons for Czech and English (PDT-Vallex, Vallex, and EngVallex). There are several goals of this work and resource: (a) to provide gold standard data for automatic experiments in the future (such as automatic discovery of synonym classes, word sense disambiguation, assignment of classes to occurrences of verbs in text, coreferential linking of verb and event arguments in text, etc.), (b) to build a core (bilingual) lexicon linked to existing resources, for comparative studies and possibly for training automatic tools, and (c) to enrich the annotation of a parallel treebank, the Prague Czech English Dependency Treebank, which so far contained valency annotation but has not linked synonymous senses of verbs together. The method used for extracting the synonym classes is a semi-automatic process with a substantial amount of manual work during filtering, role assignment to classes and individual Class members’ arguments, and linking to the external lexical resources. We present the first version with 200 classes (about 1800 verbs) and evaluate interannotator agreement using several metrics. Zdenka Uresová, Eva Fucíková, Eva Hajicová, Jan Hajic 0001 |
COLING | 4 |
| 2018 | LemmaTag: Jointly Tagging and Lemmatizing for Morphologically Rich Languages with BRNNsabstractWe present LemmaTag, a featureless neural network architecture that jointly generates part-of-speech tags and lemmas for sentences by using bidirectional RNNs with characterlevel and word-level embeddings.We demonstrate that both tasks benefit from sharing the encoding part of the network, predicting tag subcategories, and using the tagger output as an input to the lemmatizer.We evaluate our model across several languages with complex morphology, which surpasses state-of-the-art accuracy in both part-of-speech tagging and lemmatization in Czech, German, and Arabic. Daniel Kondratyuk, Tomas Gavenciak, Milan Straka, Jan Hajic 0001 |
EMNLP | 4 |
| 2018 | Bridging the LAPPS Grid and CLARIN
Erhard W. Hinrichs, Nancy Ide, James Pustejovsky, Jan Hajic 0001, Marie Hinrichs, Mohammad Fazleh Elahi, Keith Suderman, Marc Verhagen, Kyeongmin Rim, Pavel Stranák, Jozef Misutka |
LREC | 4 |
| 2018 | Diacritics Restoration Using Neural Networks
Jakub Náplava, Milan Straka, Pavel Stranák, Jan Hajic 0001 |
LREC | 4 |
| 2018 | SumeCzech: Large Czech News-Based Summarization Dataset
Milan Straka, Nikita Mediankin, Tom Kocmi, Zdenek Zabokrtský, Vojtech Hudecek, Jan Hajic 0001 |
LREC | 6 |
| 2018 | Creating a Verb Synonym Lexicon Based on a Parallel Corpus
Zdenka Uresová, Eva Fucíková, Eva Hajicová, Jan Hajic 0001 |
LREC | 4 |
| 2018 | Tools for Building an Interlinked Synonym Lexicon Network
Zdenka Uresová, Eva Fucíková, Eva Hajicová, Jan Hajic 0001 |
LREC | 4 |
| 2016 | Universal Dependencies v1: A Multilingual Treebank Collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajic 0001, Christopher D. Manning, Ryan T. McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, Daniel Zeman |
LREC | 5 |
| 2016 | Towards Comparability of Linguistic Graph Banks for Semantic Parsing
Stephan Oepen, Marco Kuhlmann, Yusuke Miyao, Daniel Zeman, Silvie Cinková, Dan Flickinger, Jan Hajic 0001, Angelina Ivanova, Zdenka Uresová |
LREC | 7 |
| 2016 | QTLeap WSD/NED Corpora: Semantic Annotation of Parallel Corpora in Six Languages
Arantxa Otegi, Nora Aranberri, António Branco, Jan Hajic 0001, Martin Popel, Kiril Ivanov Simov, Eneko Agirre, Petya Osenova, Rita Valadas Pereira, João Silva 0004, Steven Neale |
LREC | 4 |
| 2016 | Fostering the Next Generation of European Language Technology: Recent Developments ― Emerging Initiatives ― Challenges and Opportunities
Georg Rehm, Jan Hajic 0001, Josef van Genabith, Andrejs Vasiljevs |
LREC | 2 |
| 2016 | UDPipe: Trainable Pipeline for Processing CoNLL-U Files Performing Tokenization, Morphological Analysis, POS Tagging and Parsing
Milan Straka, Jan Hajic 0001, Jana Straková |
LREC | 2 |
| 2015 | Deletions and Node Reconstructions in a Dependency-Based Multilevel Annotation Scheme
Jan Hajic 0001, Eva Hajicová, Marie Mikulová, Jirí Mírovský, Jarmila Panevová, Daniel Zeman |
CICLing (1) | 1 |
| 2014 | The Strategic Impact of META-NET on the Regional, National and International Level
Georg Rehm, Hans Uszkoreit, Sophia Ananiadou, Núria Bel, Audroné Bieleviciené, Lars Borin, António Branco, Gerhard Budin, Nicoletta Calzolari, Walter Daelemans, Radovan Garabík, Marko Grobelnik, Carmen García-Mateo, Josef van Genabith, Jan Hajic 0001, Inma Hernáez Rioja, John Judge, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Joseph Mariani, John McNaught, Maite Melero, Monica Monachini, Asunción Moreno, Jan Odijk, Maciej Ogrodniczuk, Piotr Pezik, Stelios Piperidis, Adam Przepiórkowski, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Koenraad De Smedt, Marko Tadic, Paul Thompson 0002, Dan Tufis, Tamás Váradi, Andrejs Vasiljevs, Kadri Vider, Jolanta Zabarskaite |
LREC | 15 |
| 2014 | CLARA: A New Generation of Researchers in Common Language Resources and Their Applications
Koenraad De Smedt, Erhard W. Hinrichs, Detmar Meurers, Inguna Skadina, Bolette S. Pedersen, Costanza Navarretta, Núria Bel, Krister Lindén, Markéta Lopatková, Jan Hajic 0001, Gisle Andersen, Przemyslaw Lenkiewicz |
LREC | 10 |
| 2014 | Multilingual Test Sets for Machine Translation of Search Queries for Cross-Lingual Information Retrieval in the Medical Domain
Zdenka Uresová, Jan Hajic 0001, Pavel Pecina, Ondrej Dusek |
LREC | 2 |
| 2014 | Not an Interlingua, But Close: Comparison of English AMRs to Chinese and Czech
Nianwen Xue, Ondrej Bojar, Jan Hajic 0001, Martha Palmer, Zdenka Uresová, Xiuhong Zhang |
LREC | 3 |
| 2014 | Adaptation of machine translation for multilingual information retrieval in the medical domain
Pavel Pecina, Ondrej Dusek, Lorraine Goeuriot, Jan Hajic 0001, Jaroslava Hlavácová, Gareth J. F. Jones, Liadh Kelly, Johannes Leveling, David Marecek, Michal Novák 0001, Martin Popel, Rudolf Rosa, Ales Tamchyna, Zdenka Uresová |
Artif. Intell. Medicine | 4 |
| 2013 | Joint Morphological and Syntactic Analysis for Richly Inflected LanguagesabstractJoint morphological and syntactic analysis has been proposed as a way of improving parsing accuracy for richly inflected languages. Starting from a transition-based model for joint part-of-speech tagging and dependency parsing, we explore different ways of integrating morphological features into the model. We also investigate the use of rule-based morphological analyzers to provide hard or soft lexical constraints and the use of word clusters to tackle the sparsity of lexical features. Evaluation on five morphologically rich languages (Czech, Finnish, German, Hungarian, and Russian) shows consistent improvements in both morphological and syntactic accuracy for joint prediction over a pipeline model, with further improvements thanks to lexical constraints and word clusters. The final results improve the state of the art in dependency parsing for all languages. Bernd Bohnet, Joakim Nivre, Igor Boguslavsky, Richárd Farkas, Filip Ginter, Jan Hajic 0001 |
Trans. Assoc. Comput. Linguistics | 6 |
| 2012 | Announcing Prague Czech-English Dependency Treebank 2.0
Jan Hajic 0001, Eva Hajicová, Jarmila Panevová, Petr Sgall, Ondrej Bojar, Silvie Cinková, Eva Fucíková, Marie Mikulová, Petr Pajas, Jan Popelka, Jirí Semecký, Jana Sindlerová, Jan Stepánek, Josef Toman, Zdenka Uresová, Zdenek Zabokrtský |
LREC | 1 |
| 2012 | HamleDT: To Parse or Not to Parse?
Daniel Zeman, David Marecek, Martin Popel, Loganathan Ramasamy, Jan Stepánek, Zdenek Zabokrtský, Jan Hajic 0001 |
LREC | 7 |
| 2009 | Semi-Supervised Training for the Averaged Perceptron POS Tagger
Drahomíra "johanka" Spoustová, Jan Hajic 0001, Jan Raab, Miroslav Spousta |
EACL | 2 |
| 2008 | Validating the Quality of Full Morphological Annotation
Drahomíra "johanka" Spoustová, Pavel Pecina, Jan Hajic 0001, Miroslav Spousta |
LREC | 3 |
| 2008 | PDTSL: An annotated resource for speech reconstructionabstractWe present a description of a new resource (Prague Dependency Treebank of Spoken Language) being created for English and Czech to be used for the task of speech understanding, broad natural language analysis for dialog systems and other speech-related tasks, including speech editing. The resources we have created so far contain audio and a standard transcription of spontaneous speech, but as a novel layer, we add an edited (ldquoreconstructedrdquo) version of the spoken utterances. These edits go beyond the scope of current speech reconstruction efforts in that we allow, on top of the usual deletions of speech artifacts, fillers, etc. also for word modifications, insertions and word order changes. We have used both monologue and dialogue recordings in English and Czech to verify the feasibility of such transcription. We have also assessed the quality of the resulting annotation since the relative freedom of the editing raises an issue of what a ldquocorrectrdquo annotation is. Jan Hajic 0001, Silvie Cinková, Marie Mikulová, Petr Pajas, Jan Ptácek, Josef Toman, Zdenka Uresová |
SLT | 1 |
| 2006 | Leveraging Reusability: Cost-Effective Lexical Acquisition for Large-Scale Ontology TranslationabstractThesauri and ontologies provide important value in facilitating access to digital archives by representing underlying principles of organization. Translation of such resources into multiple languages is an important component for providing multilingual access. However, the specificity of vocabulary terms in most ontologies precludes fully-automated machine translation using general-domain lexical resources. In this paper, we present an efficient process for leveraging human translations when constructing domain-specific lexical resources. We evaluate the effectiveness of this process by producing a probabilistic phrase dictionary and translating a thesaurus of 56,000 concepts used to catalogue a large archive of oral histories. Our experiments demonstrate a cost-effective technique for accurate machine translation of large ontologies. G. Craig Murray, Bonnie J. Dorr, Jimmy Lin, Jan Hajic 0001, Pavel Pecina |
ACL | 4 |
| 2006 | Leveraging Recurrent Phrase Structure in Large-scale Ontology Translation
G. Craig Murray, Bonnie J. Dorr, Jimmy Lin, Jan Hajic 0001, Pavel Pecina |
EAMT | 4 |
| 2006 | Perspectives of Turning Prague Dependency Treebank into a Knowledge Base
Václav Novák, Jan Hajic 0001 |
LREC | 2 |
| 2005 | Prague Czech-English dependency treebank: resource for structure-based MT
Martin Cmejrek, Jan Curín, Jan Hajic 0001, Jirí Havelka |
EAMT | 3 |
| 2005 | Automatic transcription of Czech, Russian, and Slovak spontaneous speech in the MALACH projectabstractThis paper describes the 3.5-years effort put into building LVCSR systems for recognition of spontaneous speech of Czech, Russian, and Slovak witnesses of the Holocaust in the MALACH project. For processing of colloquial, highly emotional and heavily accented speech of elderly people containing many non-speech events we have developed techniques that very effectively handle both non-speech events and colloquial and accented variants of uttered words. Manual transcripts as one of the main sources for language modeling were automatically „normalized ” using standardized lexicon, which brought about 2 to 3 % reduction of the word error rate (WER). The subsequent interpolation of such LMs with models built from an additional collection (consisting of topically selected sentences from general text corpora) resulted into an additional improvement of performance of up to 3 %. 1. Josef Psutka, Pavel Ircing, Josef V. Psutka, Jan Hajic 0001, William J. Byrne, Jirí Mírovský |
INTERSPEECH | 4 |
| 2005 | Cross-language text classificationabstractNo abstract available. J. Scott Olsson, Douglas W. Oard, Jan Hajic 0001 |
SIGIR | 3 |
| 2004 | The development of ASR for Slavic languages in the MALACH projectabstractThe development of acoustic training material for Slavic languages within the MALACH project is described. Initial experience with the variety of speakers and the difficulties encountered in transcribing Czech, Slovak, and Russian language oral history are described along with automatic speech recognition results intended to investigate the effectiveness of different transcription conventions that address language specific phenomena within the task domain. Josef Psutka, Jan Hajic 0001, William J. Byrne |
ICASSP (3) | 2 |
| 2004 | Prague Czech-English Dependency Treebank. Syntactically Annotated Resources for Machine Translation
Martin Cmejrek, Jan Curín, Jirí Havelka, Jan Hajic 0001, Vladislav Kubon |
LREC | 4 |
| 2004 | Issues in Annotation of the Czech Spontaneous Speech Corpus in the MALACH project
Josef Psutka, Pavel Ircing, Jan Hajic 0001, Vlasta Radová, Josef V. Psutka, William J. Byrne, Samuel Gustman |
LREC | 3 |
| 2004 | Automatic recognition of spontaneous speech for access to multilingual oral history archivesabstractMuch is known about the design of automated systems to search broadcast news, but it has only recently become possible to apply similar techniques to large collections of spontaneous speech. This paper presents initial results from experiments with speech recognition, topic segmentation, topic categorization, and named entity detection using a large collection of recorded oral histories. The work leverages a massive manual annotation effort on 10 000 h of spontaneous speech to evaluate the degree to which automatic speech recognition (ASR)-based segmentation and categorization techniques can be adapted to approximate decisions made by human annotators. ASR word error rates near 40% were achieved for both English and Czech for heavily accented, emotional and elderly spontaneous speech based on 65-84 h of transcribed speech. Topical segmentation based on shifts in the recognized English vocabulary resulted in 80% agreement with manually annotated boundary positions at a 0.35 false alarm rate. Categorization was considerably more challenging, with a nearest-neighbor technique yielding F=0.3. This is less than half the value obtained by the same technique on a standard newswire categorization benchmark, but replication on human-transcribed interviews showed that ASR errors explain little of that difference. The paper concludes with a description of how these capabilities could be used together to search large collections of recorded oral histories. William J. Byrne, David S. Doermann, Martin Franz, Samuel Gustman, Jan Hajic 0001, Douglas W. Oard, Michael Picheny, Josef Psutka, Bhuvana Ramabhadran, Dagobert Soergel, Todd Ward, Wei-Jing Zhu |
IEEE Trans. Speech Audio Process. | 5 |
| 2003 | Combination of a hidden tag model and a traditional n-gram model: a case study in czech speech recognition
Pavel Krbec, Petr Podveský, Jan Hajic 0001 |
INTERSPEECH | 3 |
| 2003 | Large vocabulary ASR for spontaneous czech in the MALACH projectabstractThis paper describes LVCSR research into the automatic transcription of spontaneous Czech speech in the MALACH (Multilingual Access to Large Spoken Archives) project. This project attempts to provide improved access to the large multilingual spoken archives collected by the Survivors of the Shoah Visual History Foundation (VHF) (www.vhf.org) by advancing the state of the art in automated speech recognition. We describe a baseline ASR system and discuss the problems in language modeling that arise from the nature of Czech as a highly inflectional language that also exhibits diglossia between its written and spontaneous forms. The difficulties of this task are compounded by heavily accented, emotional and disfluent speech along with frequent switching between languages. To overcome the limited amount of relevant language model data we use statistical techniques for selecting an appropriate training corpus from a large unstructured text collection resulting in significant reductions in word error rate. 1. Josef Psutka, Pavel Ircing, Josef V. Psutka, Vlasta Radová, William J. Byrne, Jan Hajic 0001, Jirí Mírovský, Samuel Gustman |
INTERSPEECH | 6 |
| 2003 | A simple multilingual machine translation systemabstractThe multilingual machine translation system described in the first part of this paper demonstrates that the translation memory (TM) can be used in a creative way for making the translation process more automatic (in a way which in fact does not depend on the languages used). The MT system is based upon exploitation of syntactic similarities between more or less related natural languages. It currently covers the translation from Czech to Slovak, Polish and Lithuanian. The second part of the paper also shows that one of the most popular TM based commercial systems, TRADOS, can be used not only for the translation itself, but also for a relatively fast and natural method of evaluation of the translation quality of MT systems. Jan Hajic 0001, Petr Homola, Vladislav Kubon |
MTSummit | 1 |
| 2001 | Serial Combination of Rules and Statistics: A Case Study in Czech TaggingabstractA hybrid system is described which combines the strength of manual rule-writing and statistical learning, obtaining results superior to both methods if applied separately. The combination of a rule-based system and a statistical one is not parallel but serial: the rule-based system performing partial disambiguation with recall close to 100% is applied first, and a trigram HMM tagger runs on its results. An experiment in Czech tagging has been performed with encouraging results. Jan Hajic 0001, Pavel Krbec, Pavel Kveton, Karel Oliva, Vladimír Petkevic |
ACL | 1 |
| 2001 | On large vocabulary continuous speech recognition of highly inflectional language - czechabstractThe thesis concerns the development of a large vocabulary continuous speech recognition (LVCSR) system for highly inflectional languages, with special emphasis on the language modeling. An idea and usage of the automatic speech recognition is introduced and the basic principles of the statistical approach to the speech recognition and the decomposition of the system into basic components are explained. An overview of the existing statistical language modeling techniques is given and methods of inferring reliable probability estimates from sparse data and measures of the language model quality are described. There are offered a theoretical background to the finite-state machinery and the application of the finite-state machine framework to LVCSR. The goals of the thesis were to build a LVCSR system for the Czech language using standard techniques that were used for English and to analyze the system performance and propose and implement techniques that would improve the recognition accuracy. The development of the baseline system is described. The Czech language properties, especially from the automatic speech recognition point of view, were analyzed. The outcomes of this theoretical analysis are exploited and language models that take into account the specific features of the Czech language are presented. There is given a description of the class-based language models that strengthen the language model robustness and therefore reduce the perplexity and consequently improve the recognition accuracy. And finally a model that uses subword parts (morphemes) as the basic language modeling units is introduced. Such model offers a better coverage of an unknown text in comparison with standard word-based models given the same vocabulary size. Pavel Ircing, Pavel Krbec, Jan Hajic 0001, Josef Psutka, Sanjeev Khudanpur, Frederick Jelinek, William J. Byrne |
INTERSPEECH | 3 |
| 1999 | A Statistical Parser for CzechabstractThis paper considers statistical parsing of Czech, which differs radically from English in at least two respects: (1) it is a highly inflected language, and (2) it has relatively free word order. These differences are likely to pose new problems for techniques that have been developed on English. We describe our experience in building on the parsing model of (Collins 97). Our final results- 80% dependency accuracy - represent good progress towards the 91% accuracy of the parser on English (Wall Street Journal) text. Michael Collins 0001, Jan Hajic 0001, Lance A. Ramshaw, Christoph Tillmann |
ACL | 2 |
| 1998 | Czech language processing, POS tagging
Jan Hajic 0001, Barbora Hladká |
LREC | 1 |
| 1995 | Machine Translation in the Czech Republic: history, methods, systems
Jan Hajic 0001 |
MTSummit | 1 |
| 1992 | Derivation Of Underlying Valency Frames From A Learner's DictionaryabstractThe authors collect lexical data for a module of English syntactic analysis in the context of a bilingual research project. The computer usable version of OALD (Hornby, 1974) is used as the primary source. The main focus is on the structure and derivation of valency frames for verbal entries in the target lexicon. Illustration of the complex relation between OALD’s verb subcategorization codes and the target complementation paradigms is provided, and an approach to the derivation procedure design suggested. Alexandr Rosen, Eva Hajicová, Jan Hajic 0001 |
COLING | 3 |
| 1990 | Spelling-checking for Highly Inflective Languages
Jan Hajic 0001, Janus Drózd |
COLING | 1 |
| 1988 | Formal morphology
Jan Hajic 0001 |
COLING | 1 |
| 1987 | RUSLAN - An MT System Between Closely Related Languages
Jan Hajic 0001 |
EACL | 1 |
| 1982 | Inferencing And Search For An Answer In TIBAQ
Petr Jirku, Jan Hajic 0001 |
COLING | 2 |