Jana Straková

dblp:75/8155 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-0075-2408ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 94% Web and social media mining · 6%
Artificial intelligence
1 paper
Information extraction and text analysis · 70% Deep learning architectures and training · 23% Representation and self-supervised learning · 7%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › query log analysis
clickthrough data
0.812024
CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking · SIGIR 2024
Information retrieval
evaluation
0.812024
CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking · SIGIR 2024
Information retrieval › ranking › search ranking
relevance ranking
0.812024
CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking · SIGIR 2024
Information retrieval › evaluation
test collection
0.812024
CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking · SIGIR 2024
Information retrieval
web search
0.812024
CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking · SIGIR 2024
Natural language and speech › Information extraction and text analysis
named entity recognition
0.412019
Neural Architectures for Nested NER through Linearization · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › named entity recognition
nested named entity recognition
0.412019
Neural Architectures for Nested NER through Linearization · ACL (1) 2019
Natural language and speech › Information extraction and text analysis
sequence labeling
0.412019
Neural Architectures for Nested NER through Linearization · ACL (1) 2019
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning
0.412019
Neural Architectures for Nested NER through Linearization · ACL (1) 2019
Web and social media mining
user behavior analysis
0.212024
CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking · SIGIR 2024
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation
0.112019
Neural Architectures for Nested NER through Linearization · ACL (1) 2019

Methods — techniques the papers use, named apart from their topics

linearization · 0.4hard attention · 0.4contextual embeddings · 0.4LSTM-CRF · 0.4
YearPublicationVenuePosition
2026 Automatic Suggestions Help Extending Eventive Ontology: A Case Study on SynSemClass
Jana Straková, Eva Fucíková, Zdenka Uresová, Jan Hajic 0001
LREC1
2024 OOVs in the Spotlight: How to Inflect Them?
abstract
We focus on morphological inflection in out-of-vocabulary (OOV) conditions, an under-researched subtask in which state-of-the-art systems usually are less effective. We developed three systems: a retrograde model and two sequence-to-sequence (seq2seq) models based on LSTM and Transformer. For testing in OOV conditions, we automatically extracted a large dataset of nouns in the morphologically rich Czech language, with lemma-disjoint data splits, and we further manually annotated a real-world OOV dataset of neologisms. In the standard OOV conditions, Transformer achieves the best results, with increasing performance in ensemble with LSTM, the retrograde model and SIGMORPHON baselines. On the real-world OOV dataset of neologisms, the retrograde model outperforms all neural models. Finally, our seq2seq models achieve state-of-the-art results in 9 out of 16 languages from SIGMORPHON 2022 shared task data in the OOV evaluation (feature overlap) in the large data condition. We release the Czech OOV Inflection Dataset for rigorous evaluation in OOV conditions. Further, we release the inflection system with the seq2seq models as a ready-to-use Python library.
Tomás Sourada, Jana Straková, Rudolf Rosa
LREC/COLING2
2024 CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking
abstract
We present CWRCzech, Click Web Ranking dataset for Czech, a 100M query-document Czech click dataset for relevance ranking with user behavior data collected from search engine logs of Seznam.cz. To the best of our knowledge, CWRCzech is the largest click dataset with raw text published so far. It provides document positions in the search results as well as information about user behavior: 27.6M clicked documents and 10.8M dwell times. In addition, we also publish a manually annotated Czech test for the relevance task, containing nearly 50k query-document pairs, each annotated by at least 2 annotators. Finally, we analyze how the user behavior data improve relevance ranking and show that models trained on data automatically harnessed at sufficient scale can surpass the performance of models trained on human annotated data. CWRCzech is published under an academic non-commercial license and is available to the research community at https://github.com/seznam/CWRCzech.
Josef Vonásek, Milan Straka, Rostislav Krc, Lenka Lasonová, Ekaterina Egorova, Jana Straková, Jakub Náplava
SIGIR6
2022 Czech Grammar Error Correction with a Large and Diverse Corpus
abstract
Abstract We introduce a large and diverse Czech corpus annotated for grammatical error correction (GEC) with the aim to contribute to the still scarce data resources in this domain for languages other than English. The Grammar Error Correction Corpus for Czech (GECCC) offers a variety of four domains, covering error distributions ranging from high error density essays written by non-native speakers, to website texts, where errors are expected to be much less common. We compare several Czech GEC systems, including several Transformer-based ones, setting a strong baseline to future research. Finally, we meta-evaluate common GEC metrics against human judgments on our data. We make the new Czech GEC corpus publicly available under the CC BY-SA 4.0 license at http://hdl.handle.net/11234/1-4639.
Jakub Náplava, Milan Straka, Jana Straková, Alexandr Rosen
Trans. Assoc. Comput. Linguistics3
2019 Neural Architectures for Nested NER through Linearization
abstract
We propose two neural network architectures for nested named entity recognition (NER), a setting in which named entities may overlap and also be labeled with more than one label.We encode the nested labels using a linearized scheme.In our first proposed approach, the nested labels are modeled as multilabels corresponding to the Cartesian product of the nested labels in a standard LSTM-CRF architecture.In the second one, the nested NER is viewed as a sequence-to-sequence problem, in which the input sequence consists of the tokens and output sequence of the labels, using hard attention on the word whose label is being predicted.The proposed methods outperform the nested NER state of the art on four corpora: ACE-2004, ACE-2005, GENIA and Czech CNEC.We also enrich our architectures with the recently published contextual embeddings: ELMo, BERT and Flair, reaching further improvements for the four nested entity corpora.In addition, we report flat NER stateof-the-art results for CoNLL-2002 Dutch and Spanish and for CoNLL-2003 English.
Jana Straková, Milan Straka, Jan Hajic 0001
ACL (1)1
2016 UDPipe: Trainable Pipeline for Processing CoNLL-U Files Performing Tokenization, Morphological Analysis, POS Tagging and Parsing
Milan Straka, Jan Hajic 0001, Jana Straková
LREC3
2010 Czech Information Retrieval with Syntax-based Language Models
Jana Straková, Pavel Pecina
LREC1