Ondrej Bojar

dblp:73/3097 · DBLP profile ↗
← Back
59ranked-venue papers
8as first author
28since 2021 · last 2026
0000-0002-0606-0050ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 8 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 ELOQUENT Lab at CLEF 2026: Evaluation of Generative Language Model Quality
Jussi Karlgren, Maria Barrett, Ondrej Bojar, Marie Isabel Engels, Diandra Fabre, Lorraine Goeuriot, Josiane Mothe, Philippe Mulhem, Mario Piacentini, Luis Francisco Vargas Madriz, Didier Schwab, Pavel Sindelár, George Stampoulidis, Katherina Thomas, Markarit Vartampetian
ECIR (4)3
2026 CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia
Josef Jon, Ondrej Bojar
LREC2
2025 MockConf: A Student Interpretation Dataset: Analysis, Word- and Span-level Alignment and Baselines
abstract
International audience
Dávid Javorský, Ondrej Bojar, François Yvon
ACL (1)2
2025 ELOQUENT CLEF Shared Tasks for Evaluation of Generative Language Model Quality, 2025 Edition
Jussi Karlgren, Ekaterina Artemova, Ondrej Bojar, Vladislav Mikhailov, Magnus Sahlgren, Erik Velldal, Lilja Øvrelid
ECIR (5)3
2025 Intrinsic vs. Extrinsic Evaluation of Czech Sentence Embeddings: Semantic Relevance Doesn't Help with MT Evaluation
abstract
In this paper, we compare Czech-specific and multilingual sentence embedding models through intrinsic and extrinsic evaluation paradigms. For intrinsic evaluation, we employ Costra, a complex sentence transformation dataset, and several Semantic Textual Similarity (STS) benchmarks to assess the ability of the embeddings to capture linguistic phenomena such as semantic similarity, temporal aspects, and stylistic variations. In the extrinsic evaluation, we fine-tune each embedding model using COMET-based metrics for machine translation evaluation. Our experiments reveal an interesting disconnect: models that excel in intrinsic semantic similarity tests do not consistently yield superior performance on downstream translation evaluation tasks. Conversely, models with seemingly over-smoothed embedding spaces can, through fine-tuning, achieve excellent results. These findings highlight the complex relationship between semantic property probes and downstream task, emphasizing the need for more research into “operationalizable semantics” in sentence embeddings, or more in-depth downstream tasks datasets (here translation evaluation).
Petra Barancíková, Ondrej Bojar
MTSummit (1)2
2025 How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
abstract
Abstract Simultaneous speech-to-text translation (SimulST) translates source-language speech into target-language text concurrently with the speaker’s speech, ensuring low latency for better user comprehension. Despite its intended application to unbounded speech, most research has focused on human pre-segmented speech, simplifying the task and overlooking significant challenges. This narrow focus, coupled with widespread terminological inconsistencies, is limiting the applicability of research outcomes to real-world applications, ultimately hindering progress in the field. Our extensive literature review of 110 papers not only reveals these critical issues in current research but also serves as the foundation for our key contributions. We: 1) define the steps and core components of a SimulST system, proposing a standardized terminology and taxonomy; 2) conduct a thorough analysis of community trends; and 3) offer concrete recommendations and future directions to bridge the gaps in existing literature, from evaluation frameworks to system architectures, for advancing the field towards more realistic and effective SimulST solutions.
Sara Papi, Peter Polak, Dominik Machácek, Ondrej Bojar
Trans. Assoc. Comput. Linguistics4
2024 Khan Academy Corpus: A Multilingual Corpus of Khan Academy Lectures
abstract
We present the Khan Academy Corpus totalling 10122 hours in 87394 recordings across 29 languages, where 43% of recordings (4252 hours) are equipped with human-written subtitles. The subtitle texts cover a total of 137 languages. The dataset was collected from open access Khan Academy lectures, benefiting from their manual transcripts and manual translations of the transcripts. The dataset can serve in creation or evaluation of multilingual speech recognition or translation systems, featuring a diverse set of subject domains.
Dominika Durisková, Daniela Jurásová, Matús Zilinec, Eduard Subert, Ondrej Bojar
LREC/COLING5
2024 GAATME: A Genetic Algorithm for Adversarial Translation Metrics Evaluation
abstract
Building on a recent method for decoding translation candidates from a Machine Translation (MT) model via a genetic algorithm, we modify it to generate adversarial translations to test and challenge MT evaluation metrics. The produced translations score very well in an arbitrary MT evaluation metric selected beforehand, despite containing serious, deliberately introduced errors. The method can be used to create adversarial test sets to analyze the biases and shortcomings of the metrics. We publish various such test sets for the Czech to English language pair, as well as the code to convert any parallel data into a similar adversarial test set.
Josef Jon, Ondrej Bojar
LREC/COLING2
2024 Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
abstract
Human evaluation is a critical component in machine translation system development and has received much attention in text translation research. However, little prior work exists on the topic of human evaluation for speech translation, which adds additional challenges such as noisy data and segmentation mismatches. We take the first steps to fill this gap by conducting a comprehensive human evaluation of the results of several shared tasks from the last International Workshop on Spoken Language Translation (IWSLT 2023). We propose an effective evaluation strategy based on automatic resegmentation and direct assessment with segment context. Our analysis revealed that: 1) the proposed evaluation strategy is robust and scores well-correlated with other types of human judgements; 2) automatic metrics are usually, but not always, well-correlated with direct assessment scores; and 3) COMET as a slightly stronger automatic metric than chrF, despite the segmentation noise introduced by the resegmentation step systems. We release the collected human-annotated data in order to encourage further investigation.
Matthias Sperber, Ondrej Bojar, Barry Haddow, Dávid Javorský, Xutai Ma, Matteo Negri, Jan Niehues, Peter Polak, Elizabeth Salesky, Katsuhito Sudoh, Marco Turchi
LREC/COLING2
2024 Long-Form End-To-End Speech Translation VIA Latent Alignment Segmentation
abstract
Contemporary datasets provide an oracle segmentation into sentences based on human-annotated transcripts and translations. However, the segmentation into sentences is not available in the real world. Current speech segmentation approaches offer poor segmentation quality in the low-latency regime. This paper proposes a novel segmentation approach for a low-latency end-to-end speech translation. We leverage an existing speech translation encoder-decoder architecture with speech translation CTC (ST CTC) and show that the latent alignments produced by ST CTC can guide the segmentation. To the best of our knowledge, our method is the first that allows an actual long-form end-to-end simultaneous speech translation, as one neural model translates and segments simultaneously. On a diverse set of language pairs and in- and out-of-domain data, we show that the proposed approach outperforms current state-of-the-art segmentation methods at no additional computational cost.
Peter Polak, Ondrej Bojar
SLT2
2023 Breeding Machine Translations: Evolutionary approach to survive and thrive in the world of automated evaluation
abstract
We propose a genetic algorithm (GA) based method for modifying n-best lists produced by a machine translation (MT) system.Our method offers an innovative approach to improving MT quality and identifying weaknesses in evaluation metrics.Using common GA operations (mutation and crossover) on a list of hypotheses in combination with a fitness function (an arbitrary MT metric), we obtain novel and diverse outputs with high metric scores.With a combination of multiple MT metrics as the fitness function, the proposed method leads to an increase in translation quality as measured by other held-out automatic metrics.With a single metric (including popular ones such as COMET) as the fitness function, we find blind spots and flaws in the metric.This allows for an automated search for adversarial examples in an arbitrary metric, without prior assumptions on the form of such example.As a demonstration of the method, we create datasets of adversarial examples and use them to show that reference-free COMET is substantially less robust than the reference-based version.
Josef Jon, Ondrej Bojar
ACL (1)2
2023 Incremental Blockwise Beam Search for Simultaneous Speech Translation with Controllable Quality-Latency Tradeoff
abstract
Blockwise self-attentional encoder models have recently emerged as one promising end-to-end approach to simultaneous speech translation. These models employ a blockwise beam search with hypothesis reliability scoring to determine when to wait for more input speech before translating further. However, this method maintains multiple hypotheses until the entire speech input is consumed -- this scheme cannot directly show a single \textit{incremental} translation to users. Further, this method lacks mechanisms for \textit{controlling} the quality vs. latency tradeoff. We propose a modified incremental blockwise beam search incorporating local agreement or hold-$n$ policies for quality-latency control. We apply our framework to models trained for online or offline translation and demonstrate that both types can be effectively used in online mode. Experimental results on MuST-C show 0.6-3.6 BLEU improvement without changing latency or 0.8-1.4 s latency improvement without changing quality.
Peter Polak, Brian Yan, Shinji Watanabe 0001, Alex Waibel, Ondrej Bojar
INTERSPEECH5
2023 Character-level NMT and language similarity
abstract
We explore the effectiveness of character-level neural machine translation using Transformer architecture for various levels of language similarity and size of the training dataset. We evaluate the models using automatic MT metrics and show that translation between similar languages benefits from character-level input segmentation, while for less related languages, character-level vanilla Transformer-base often lags behind subword-level segmentation. We confirm previous findings that it is possible to close the gap by finetuning the already trained subword-level models to character-level.
Josef Jon, Ondrej Bojar
MTSummit (1)2
2023 Negative Lexical Constraints in Neural Machine Translation
abstract
This paper explores negative lexical constraining in English to Czech neural machine translation. Negative lexical constraining is used to prohibit certain words or expressions in the translation produced by the NMT model. We compared various methods based on modifying either the decoding process or the training data. The comparison was performed on two tasks: paraphrasing and feedback-based translation refinement. We also studied how the methods “evade” the constraints, meaning that the disallowed expression is still present in the output, but in a changed form, most interestingly the case where a different surface form (for example different inflection) is produced. We propose a way to mitigate the issue through training with stemmed negative constraints, so that the ability of the model to induce different forms of a word might be used to prohibit the usage of all possible forms of the constraint. This helps to some extent, but the problem still persists in many cases.
Josef Jon, Dusan Varis, Michal Novák 0001, João Paulo Aires, Ondrej Bojar
MTSummit (1)5
2023 Boosting Unsupervised Machine Translation with Pseudo-Parallel Data
abstract
Even with the latest developments in deep learning and large-scale language modeling, the task of machine translation (MT) of low-resource languages remains a challenge. Neural MT systems can be trained in an unsupervised way without any translation resources but the quality lags behind, especially in truly low-resource conditions. We propose a training strategy that relies on pseudo-parallel sentence pairs mined from monolingual corpora in addition to synthetic sentence pairs back-translated from monolingual corpora. We experiment with different training schedules and reach an improvement of up to 14.5 BLEU points (English to Ukrainian) over a baseline trained on back-translated data only.
Ivana Kvapilíková, Ondrej Bojar
MTSummit (1)2
2023 The Role of Compounds in Human vs. Machine Translation Quality
abstract
We focus on the production of German compounds in English-to-German manual and automatic translation. On the example of WMT21 news translation test set, we observe that even the best MT systems produce much fewer compounds compared to three independent manual translations. Despite this striking difference, we observe that this insufficiency is not apparent in manual evaluation methods that target the overall translation quality (DA and MQM). Simple automatic methods like BLEU somewhat surprisingly provide a better indication of this quality aspect. Our manual analysis of system outputs, including our freshly trained Transformer models, confirms that current deep neural systems operating at the level of subword units are capable of constructing novel words, including novel compounds. This effect however cannot be measured using static dictionaries of compounds such as GermaNet. German compounds thus pose an interesting challenge for future development of MT systems.
Kristyna Neumannova, Ondrej Bojar
MTSummit (1)2
2023 Bad MT Systems are Good for Quality Estimation
abstract
Quality estimation (QE) is the task of predicting quality of outputs produced by machine translation (MT) systems. Currently, the highest-performing QE systems are supervised and require training on data with golden quality scores. In this paper, we investigate the impact of the quality of the underlying MT outputs on the performance of QE systems. We find that QE models trained on datasets with lower-quality translations often outperform those trained on higher-quality data. We also demonstrate that good performance can be achieved by using a mix of data from different MT systems.
Iryna Tryhubyshyn, Ales Tamchyna, Ondrej Bojar
MTSummit (1)3
2022 Hausa Visual Genome: A Dataset for Multi-Modal English to Hausa Machine Translation
abstract
Multi-modal Machine Translation (MMT) enables the use of visual information to enhance the quality of translations, especially where the full context is not available to enable the unambiguous translation in standard machine translation. Despite the increasing popularity of such technique, it lacks sufficient and qualitative datasets to maximize the full extent of its potential. Hausa, a Chadic language, is a member of the Afro-Asiatic language family. It is estimated that about 100 to 150 million people speak the language, with more than 80 million indigenous speakers. This is more than any of the other Chadic languages. Despite the large number of speakers, the Hausa language is considered as a low resource language in natural language processing (NLP). This is due to the absence of enough resources to implement most of the tasks in NLP. While some datasets exist, they are either scarce, machine-generated or in the religious domain. Therefore, there is the need to create training and evaluation data for implementing machine learning tasks and bridging the research gap in the language. This work presents the Hausa Visual Genome (HaVG), a dataset that contains the description of an image or a section within the image in Hausa and its equivalent in English. The dataset was prepared by automatically translating the English description of the images in the Hindi Visual Genome (HVG). The synthetic Hausa data was then carefully postedited, taking into cognizance the respective images. The data is made of 32,923 images and their descriptions that are divided into training, development, test, and challenge test set. The Hausa Visual Genome is the first dataset of its kind and can be used for Hausa-English machine translation, multi-modal research, image description, among various other natural language processing and generation tasks.
Idris Abdulmumin, Satya Ranjan Dash, Musa Abdullahi Dawud, Shantipriya Parida, Shamsuddeen Hassan Muhammad, Ibrahim Said Ahmad, Subhadarshi Panda, Ondrej Bojar, Bashir Shehu Galadanci, Bello Shehu Bello
LREC8
2022 Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers
abstract
With the development of multimodal systems and natural language generation techniques, the resurgence of multimodal datasets has attracted significant research interests, which aims to provide new information to enrich the representation of textual data. However, there remains a lack of a comprehensive survey for this task. To this end, we take the first step and present a thorough review of this research field. This paper provides an overview of a publicly available dataset with different modalities according to the applications. Furthermore, we discuss the new frontier and give our thoughts. We hope this survey of multimodal datasets can provide the community with quick access and a general picture of the multimodal dataset for specific Natural Language Processing (NLP) applications and motivates future researches. In this context, we release the collection of all multimodal datasets easily accessible here: https://github.com/drmuskangarg/Multimodal-datasets
Muskan Garg, Seema Wazarkar, Muskaan Singh, Ondrej Bojar
LREC4
2022 ELITR Minuting Corpus: A Novel Dataset for Automatic Minuting from Multi-Party Meetings in English and Czech
abstract
Taking minutes is an essential component of every meeting, although the goals, style, and procedure of this activity (“minuting” for short) can vary. Minuting is a rather unstructured writing activity and is affected by who is taking the minutes and for whom the intended minutes are. With the rise of online meetings, automatic minuting would be an important benefit for the meeting participants as well as for those who might have missed the meeting. However, automatically generating meeting minutes is a challenging problem due to a variety of factors including the quality of automatic speech recorders (ASRs), availability of public meeting data, subjective knowledge of the minuter, etc. In this work, we present the first of its kind dataset on Automatic Minuting. We develop a dataset of English and Czech technical project meetings which consists of transcripts generated from ASRs, manually corrected, and minuted by several annotators. Our dataset, AutoMin, consists of 113 (English) and 53 (Czech) meetings, covering more than 160 hours of meeting content. Upon acceptance, we will publicly release (aaa.bbb.ccc) the dataset as a set of meeting transcripts and minutes, excluding the recordings for privacy reasons. A unique feature of our dataset is that most meetings are equipped with more than one minute, each created independently. Our corpus thus allows studying differences in what people find important while taking the minutes. We also provide baseline experiments for the community to explore this novel problem further. To the best of our knowledge AutoMin is probably the first resource on minuting in English and also in a language other than English (Czech).
Anna Nedoluzhko, Muskaan Singh, Marie Hledíková, Tirthankar Ghosal, Ondrej Bojar
LREC5
2022 ALIGNMEET: A Comprehensive Tool for Meeting Annotation, Alignment, and Evaluation
abstract
Summarization is a challenging problem, and even more challenging is to manually create, correct, and evaluate the summaries. The severity of the problem grows when the inputs are multi-party dialogues in a meeting setup. To facilitate the research in this area, we present ALIGNMEET, a comprehensive tool for meeting annotation, alignment, and evaluation. The tool aims to provide an efficient and clear interface for fast annotation while mitigating the risk of introducing errors. Moreover, we add an evaluation mode that enables a comprehensive quality evaluation of meeting minutes. To the best of our knowledge, there is no such tool available. We release the tool as open source. It is also directly installable from PyPI.
Peter Polak, Muskaan Singh, Anna Nedoluzhko, Ondrej Bojar
LREC4
2022 Automatic Minuting: A Pipeline Method for Generating Minutes from Multi-Party Meeting Proceedings
Kartik Shinde, Tirthankar Ghosal, Muskaan Singh, Ondrej Bojar
PACLIC4
2021 End-to-End Lexically Constrained Machine Translation for Morphologically Rich Languages
abstract
Lexically constrained machine translation allows the user to manipulate the output sentence by enforcing the presence or absence of certain words and phrases. Although current approaches can enforce terms to appear in the translation, they often struggle to make the constraint word form agree with the rest of the generated output. Our manual analysis shows that 46% of the errors in the output of a baseline constrained model for English to Czech translation are related to agreement. We investigate mechanisms to allow neural machine translation to infer the correct word inflection given lemmatized constraints. In particular, we focus on methods based on training the model with constraints provided as part of the input sequence. Our experiments on the English-Czech language pair show that this approach improves the translation of constrained terms in both automatic and manual evaluation by reducing errors in agreement. Our approach thus eliminates inflection errors, without introducing new errors or decreasing the overall quality of the translation.
Josef Jon, João Paulo Aires, Dusan Varis, Ondrej Bojar
ACL/IJCNLP (1)4
2021 Sequence Length is a Domain: Length-based Overfitting in Transformer Models
abstract
Transformer-based sequence-to-sequence architectures, while achieving state-of-the-art results on a large number of NLP tasks, can still suffer from overfitting during training.In practice, this is usually countered either by applying regularization methods (e.g.dropout, L2regularization) or by providing huge amounts of training data.Additionally, Transformer and other architectures are known to struggle when generating very long sequences.For example, in machine translation, the neuralbased systems perform worse on very long sequences when compared to the preceding phrase-based translation approaches (Koehn and Knowles, 2017).We present results which suggest that the issue might also be in the mismatch between the length distributions of the training and validation data combined with the aforementioned tendency of the neural networks to overfit to the training data.We demonstrate on a simple string editing task and a machine translation task that the Transformer model performance drops significantly when facing sequences of length diverging from the length distribution in the training data.Additionally, we show that the observed drop in performance is due to the hypothesis length corresponding to the lengths seen by the model during training rather than the length of the input sequence.
Dusan Varis, Ondrej Bojar
EMNLP (1)2
2021 Neural Machine Translation Quality and Post-Editing Performance
abstract
We test the natural expectation that using MT in professional translation saves human processing time.The last such study was carried out by Sanchez-Torron and Koehn (2016) with phrase-based MT, artificially reducing the translation quality.In contrast, we focus on neural MT (NMT) of high quality, which has become the state-of-the-art approach since then and also got adopted by most translation companies.Through an experimental study involving over 30 professional translators for English→Czech translation, we examine the relationship between NMT performance and post-editing time and quality.Across all models, we found that better MT systems indeed lead to fewer changes in the sentences in this industry setting.The relation between system quality and post-editing time is however not straightforward and, contrary to the results on phrase-based MT, BLEU is definitely not a stable predictor of the time or final output quality.
Vilém Zouhar, Martin Popel, Ondrej Bojar, Ales Tamchyna
EMNLP (1)3
2021 Lost in Interpreting: Speech Translation from Source or Interpreter?
abstract
Interpreters facilitate multi-lingual meetings but the affordable set of languages is often smaller than what is needed. Automatic simultaneous speech translation can extend the set of provided languages. We investigate if such an automatic system should rather follow the original speaker, or an interpreter to achieve better translation quality at the cost of increased delay. To answer the question, we release Europarl Simultaneous Interpreting Corpus (ESIC), 10 hours of recordings and transcripts of European Parliament speeches in English, with simultaneous interpreting into Czech and German. We evaluate quality and latency of speaker-based and interpreter-based spoken translation systems from English to Czech. We study the differences in implicit simplification and summarization of the human interpreter compared to a machine translation system trained to shorten the output to some extent. Finally, we perform human evaluation to measure information loss of each of these approaches.
Dominik Machácek, Matús Zilinec, Ondrej Bojar
Interspeech3
2021 Backtranslation Feedback Improves User Confidence in MT, Not Quality
abstract
Vilém Zouhar, Michal Novák, Matúš Žilinec, Ondřej Bojar, Mateo Obregón, Robin L. Hill, Frédéric Blain, Marina Fomicheva, Lucia Specia, Lisa Yankovskaya. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Vilém Zouhar, Michal Novák 0001, Matús Zilinec, Ondrej Bojar, Mateo Obregón, Robin L. Hill, Frédéric Blain, Marina Fomicheva, Lucia Specia, Lisa Yankovskaya
NAACL-HLT4
2021 An Empirical Performance Analysis of State-of-the-Art Summarization Models for Automatic Minuting
Muskaan Singh, Tirthankar Ghosal, Ondrej Bojar
PACLIC3
2020 ELITR: European Live Translator
abstract
ELITR (European Live Translator) project aims to create a speech translation system for simultaneous subtitling of conferences and online meetings targetting up to 43 languages. The technology is tested by the Supreme Audit Office of the Czech Republic and by alfaview®, a German online conferencing system. Other project goals are to advance document-level and multilingual machine translation, automatic speech recognition, and automatic minuting.
Ondrej Bojar, Dominik Machácek, Sangeet Sagar, Otakar Smrz, Jonás Kratochvíl, Ebrahim Ansari, Dario Franceschini, Chiara Canton, Ivan Simonini, Thai Son Nguyen, Sebastian Stüker, Alex Waibel, Barry Haddow, Rico Sennrich, Philip Williams
EAMT1
2020 Efficiently Reusing Old Models Across Languages via Transfer Learning
abstract
Recent progress in neural machine translation (NMT) is directed towards larger neural networks trained on an increasing amount of hardware resources. As a result, NMT models are costly to train, both financially, due to the electricity and hardware cost, and environmentally, due to the carbon footprint. It is especially true in transfer learning for its additional cost of training the “parent” model before transferring knowledge and training the desired “child” model. In this paper, we propose a simple method of re-using an already trained model for different language pairs where there is no need for modifications in model architecture. Our approach does not need a separate parent model for each investigated language pair, as it is typical in NMT transfer learning. To show the applicability of our method, we recycle a Transformer model trained by different researchers and use it to seed models for different language pairs. We achieve better translation quality and shorter convergence times than when training from random initialization.
Tom Kocmi, Ondrej Bojar
EAMT2
2020 COSTRA 1.0: A Dataset of Complex Sentence Transformations
abstract
We present COSTRA 1.0, a dataset of complex sentence transformations. The dataset is intended for the study of sentence-level embeddings beyond simple word alternations or standard paraphrasing. This first version of the dataset is limited to sentences in Czech but the construction method is universal and we plan to use it also for other languages. The dataset consist of 4,262 unique sentences with average length of 10 words, illustrating 15 types of modifications such as simplification, generalization, or formal and informal language variation. The hope is that with this dataset, we should be able to test semantic properties of sentence embeddings and perhaps even to find some topologically interesting “skeleton” in the sentence embedding space. A preliminary analysis using LASER, multi-purpose multi-lingual sentence embeddings suggests that the LASER space does not exhibit the desired properties.
Petra Barancíková, Ondrej Bojar
LREC2
2020 Two Huge Title and Keyword Generation Corpora of Research Articles
abstract
Recent developments in sequence-to-sequence learning with neural networks have considerably improved the quality of automatically generated text summaries and document keywords, stipulating the need for even bigger training corpora. Metadata of research articles are usually easy to find online and can be used to perform research on various tasks. In this paper, we introduce two huge datasets for text summarization (OAGSX) and keyword generation (OAGKX) research, containing 34 million and 23 million records, respectively. The data were retrieved from the Open Academic Graph which is a network of research profiles and publications. We carefully processed each record and also tried several extractive and abstractive methods of both tasks to create performance baselines for other researchers. We further illustrate the performance of those methods previewing their outputs. In the near future, we would like to apply topic modeling on the two sets to derive subsets of research articles from more specific disciplines.
Erion Çano, Ondrej Bojar
LREC2
2020 Large Corpus of Czech Parliament Plenary Hearings
abstract
We present a large corpus of Czech parliament plenary sessions. The corpus consists of approximately 1200 hours of speech data and corresponding text transcriptions. The whole corpus has been segmented to short audio segments making it suitable for both training and evaluation of automatic speech recognition (ASR) systems. The source language of the corpus is Czech, which makes it a valuable resource for future research as only a few public datasets are available in the Czech language. We complement the data release with experiments of two baseline ASR systems trained on the presented data: the more traditional approach implemented in the Kaldi ASRtoolkit which combines hidden Markov models and deep neural networks (NN) and a modern ASR architecture implemented in Jaspertoolkit which uses deep NNs in an end-to-end fashion.
Jonás Kratochvíl, Peter Polak, Ondrej Bojar
LREC3
2020 Outbound Translation User Interface Ptakopet: A Pilot Study
Vilém Zouhar, Ondrej Bojar
LREC2
2019 Sentiment Analysis of Czech Texts: An Algorithmic Survey
abstract
In the area of online communication, commerce and transactions, analyzing sentiment polarity of texts written in various natural languages has become crucial. While there have been a lot of contributions in resources and studies for the English language, "smaller" languages like Czech have not received much attention. In this survey, we explore the effectiveness of many existing machine learning algorithms for sentiment analysis of Czech Facebook posts and product reviews. We report the sets of optimal parameter values for each algorithm and the scores in both datasets. We finally observe that support vector machines are the best classifier and efforts to increase performance even more with bagging, boosting or voting ensemble schemes fail to do so.
Erion Çano, Ondrej Bojar
ICAART (2)2
2019 Efficiency Metrics for Data-Driven Models: A Text Summarization Case Study
abstract
Using data-driven models for solving text summarization or similar tasks has become very common in the last years.Yet most of the studies report basic accuracy scores only, and nothing is known about the ability of the proposed models to improve when trained on more data.In this paper, we define and propose three data efficiency metrics: data score efficiency, data time deficiency and overall data efficiency.We also propose a simple scheme that uses those metrics and apply it for a more comprehensive evaluation of popular methods on text summarization and title generation tasks.For the latter task, we process and release a huge collection of 35 million abstract-title pairs from scientific articles.Our results reveal that among the tested models, the Transformer is the most efficient on both tasks.
Erion Çano, Ondrej Bojar
INLG2
2019 Representation of sentence meaning (A JNLE Special Issue)
abstract
Abstract This paper serves as a short overview of the JNLE special issue on representation of the meaning of the sentence, bringing together traditional symbolic and modern continuous approaches. We indicate notable aspects of sentence meaning and their compatibility with the two streams of research and then summarize the papers selected for this special issue.
Ondrej Bojar, Raffaella Bernardi, Bonnie L. Webber
Nat. Lang. Eng.1
2018 Are BLEU and Meaning Representation in Opposition?
abstract
One of possible ways of obtaining continuous-space sentence representations is by training neural machine translation (NMT) systems.The recent attention mechanism however removes the single point in the neural network from which the source sentence representation can be extracted.We propose several variations of the attentive NMT architecture bringing this meeting point back.Empirical evaluation suggests that the better the translation quality, the worse the learned sentence representations serve in a wide range of classification and similarity tasks.
Ondrej Cífka, Ondrej Bojar
ACL (1)2
2018 Translating Short Segments with NMT: A Case Study in English-to-Hindi
abstract
This paper presents a case study in translating short image captions of the Visual Genome dataset from English into Hindi using out-of-domain data sets of varying size. We experiment with three NMT models: the shallow and deep sequence-tosequence and the Transformer model as implemented in Marian toolkit. Phrase-based Moses serves as the baseline. The results indicate that the Transformer model outperforms others in the large data setting in a number of automatic metrics and manual evaluation, and it also produces the fewest truncated sentences. Transformer training is however very sensitive to the hyperparameters, so it requires more experimenting. The deep sequence-to-sequence model produced more flawless outputs in the small data setting and it was generally more stable, at the cost of more training iterations.
Shantipriya Parida, Ondrej Bojar
EAMT2
2017 LanideNN: Multilingual Language Identification on Text Stream
abstract
In language identification, a common first step in natural language processing, we want to automatically determine the language of some input text.Monolingual language identification assumes that the given document is written in one language.In multilingual language identification, the document is usually in two or three languages and we just want their names.We aim one step further and propose a method for textual language identification where languages can change arbitrarily and the goal is to identify the spans of each of the languages.Our method is based on Bidirectional Recurrent Neural Networks and it performs well in monolingual and multilingual language identification tasks on six datasets covering 131 languages.The method keeps the accuracy also for short documents and across domains, so it is ideal for off-the-shelf use without preparation of training data.
Tom Kocmi, Ondrej Bojar
EACL (1)2
2017 Paying Attention to Multi-Word Expressions in Neural Machine Translation
Matiss Rikters, Ondrej Bojar
MTSummit (1)2
2016 Target-Side Context for Discriminative Models in Statistical Machine Translation
abstract
Discriminative translation models utilizing source context have been shown to help statistical machine translation performance.We propose a novel extension of this work using target context information.Surprisingly, we show that this model can be efficiently integrated directly in the decoding process.Our approach scales to large training data sizes and results in consistent improvements in translation quality on four language pairs.We also provide an analysis comparing the strengths of the baseline source-context model with our extended source-context and targetcontext model and we show that our extension allows us to better capture morphological coherence.Our work is freely available as part of Moses.
Ales Tamchyna, Alexander Fraser 0001, Ondrej Bojar, Marcin Junczys-Dowmunt
ACL (1)3
2016 Pivoting Methods and Data for Czech-Vietnamese Translation via English
Duc Tam Hoang, Ondrej Bojar
EAMT2
2016 HUME: Human UCCA-Based Evaluation of Machine Translation
abstract
Human evaluation of machine translation normally uses sentence-level measures such as relative ranking or adequacy scales. However, these provide no insight into possible errors, and do not scale well with sentence length. We argue for a semantics-based evaluation, which captures what meaning components are retained in the MT output, thus providing a more fine-grained analysis of translation quality, and enabling the construction and tuning of semantics-based MT. We present a novel human semantic evaluation measure, Human UCCA-based MT Evaluation (HUME), building on the UCCA semantic representation scheme. HUME covers a wider range of semantic phenomena than previous methods and does not rely on semantic annotation of the potentially garbled MT output. We experiment with four language pairs, demonstrating HUME’s broad applicability, and report good inter-annotator agreement rates and correlation with human adequacy scores.
Alexandra Birch, Omri Abend, Ondrej Bojar, Barry Haddow
EMNLP3
2014 HindEnCorp - Hindi-English and Hindi-only Corpus for Machine Translation
Ondrej Bojar, Vojtech Diatka, Pavel Rychlý, Pavel Stranák, Vit Suchomel, Ales Tamchyna, Daniel Zeman
LREC1
2014 Two-Step Machine Translation with Lattices
Bushra Jawaid, Ondrej Bojar
LREC2
2014 A Tagged Corpus and a Tagger for Urdu
Bushra Jawaid, Amir Kamran, Ondrej Bojar
LREC3
2014 Not an Interlingua, But Close: Comparison of English AMRs to Chinese and Czech
Nianwen Xue, Ondrej Bojar, Jan Hajic 0001, Martha Palmer, Zdenka Uresová, Xiuhong Zhang
LREC2
2013 No Free Lunch in Factored Phrase-Based Machine Translation
Ales Tamchyna, Ondrej Bojar
CICLing (2)2
2012 Automatic MT Error Analysis: Hjerson Helping Addicter
Jan Berka, Ondrej Bojar, Mark Fishel, Maja Popovic, Daniel Zeman
LREC2
2012 The Joy of Parallelism with CzEng 1.0
Ondrej Bojar, Zdenek Zabokrtský, Ondrej Dusek, Petra Galuscáková, Martin Majlis, David Marecek, Jirka Marsík, Michal Novák 0001, Martin Popel, Ales Tamchyna
LREC1
2012 Terra: a Collection of Translation Error-Annotated Corpora
Mark Fishel, Ondrej Bojar, Maja Popovic
LREC2
2012 Announcing Prague Czech-English Dependency Treebank 2.0
Jan Hajic 0001, Eva Hajicová, Jarmila Panevová, Petr Sgall, Ondrej Bojar, Silvie Cinková, Eva Fucíková, Marie Mikulová, Petr Pajas, Jan Popelka, Jirí Semecký, Jana Sindlerová, Jan Stepánek, Josef Toman, Zdenka Uresová, Zdenek Zabokrtský
LREC5
2010 Evaluating Utility of Data Sources in a Large Parallel Czech-English Corpus CzEng 0.9
Ondrej Bojar, Adam Liska, Zdenek Zabokrtský
LREC1
2010 Data Issues in English-to-Hindi Machine Translation
Ondrej Bojar, Pavel Stranák, Daniel Zeman
LREC1
2010 Building a Bilingual ValLex Using Treebank Token Alignment: First Observations
Jana Sindlerová, Ondrej Bojar
LREC2
2008 CzEng 0.7: Parallel Corpus with Community-Supplied Translations
Ondrej Bojar, Miroslav Janícek, Zdenek Zabokrtský, Pavel Ceska, Peter Bena
LREC1
2007 Moses: Open Source Toolkit for Statistical Machine Translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondrej Bojar, Alexandra Constantin, Evan Herbst
ACL12
2006 Czech-English Word Alignment
Ondrej Bojar, Magdalena Prokopová
LREC1