VLDB 2026 Research / reviewers in the wild / expert
Jan Kocon
dblp:117/2896
· DBLP profile ↗
14ranked-venue papers
8as first author
7since 2021 · last 2023
0000-0002-7665-6896ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 5 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | PALS: Personalized Active Learning for Subjective Tasks in NLPabstractKamil Kanclerz, Konrad Karanowski, Julita Bielaniewicz, Marcin Gruza, Piotr Miłkowski, Jan Kocon, Przemyslaw Kazienko. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Kamil Kanclerz, Konrad Karanowski, Julita Bielaniewicz, Marcin Gruza, Piotr Milkowski, Jan Kocon, Przemyslaw Kazienko |
EMNLP | 6 |
| 2022 | Multi-model Analysis of Language-Agnostic Sentiment Classification on MultiEmo Data
Piotr Milkowski, Marcin Gruza, Przemyslaw Kazienko, Joanna Szolomicka, Stanislaw Wozniak, Jan Kocon |
ICCCI | 6 |
| 2021 | Controversy and Conformity: from Generalized to Personalized Aggressiveness DetectionabstractKamil Kanclerz, Alicja Figas, Marcin Gruza, Tomasz Kajdanowicz, Jan Kocon, Daria Puchalska, Przemyslaw Kazienko. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kamil Kanclerz, Alicja Figas, Marcin Gruza, Tomasz Kajdanowicz, Jan Kocon, Daria Puchalska, Przemyslaw Kazienko |
ACL/IJCNLP (1) | 5 |
| 2021 | Learning Personal Human Biases and Representations for Subjective Tasks in Natural Language ProcessingabstractMany tasks in natural language processing like offensive, toxic, or emotional text classification are subjective by nature. Humans tend to perceive textual content in their own individual way. Existing methods commonly rely on the agreed output values, the same for all consumers. Here, we propose personalized solutions to subjective tasks. Our four new deep learning models take into account not only the content but also the specificity of a given human. The models represent different approaches to learning the representation and processing data about text readers. The experiments were carried out on four datasets: Wikipedia discussion texts labelled with attack, aggression, and toxicity, as well as opinions annotated with ten numerical emotional categories. Emotional data was considered as multivariate regression (multitask), whereas Wikipedia data as independent classifications. All our models based on human biases and their representations significantly improve the prediction quality in subjective tasks evaluated from the individual’s perspective. Jan Kocon, Marcin Gruza, Julita Bielaniewicz, Damian Grimling, Kamil Kanclerz, Piotr Milkowski, Przemyslaw Kazienko |
ICDM | 1 |
| 2021 | Multi-task Sequence Classification for Disjoint Tasks in Low-resource LanguagesabstractMulti-task learning (MTL) has been successfully utilized in numerous NLP tasks, including sequence labeling. In this work, we utilize three transformer-based models (XLM-R, HerBERT, mBERT) to improve recognition quality using MTL for selected low-resource language (Polish) and three disjoint sequence labeling tasks with different levels of inter-annotator agreement. Our best MTL model outperforms single-task models both within the tasks domain and overall performance. Jarema Radom, Jan Kocon |
KES | 2 |
| 2021 | Offensive, aggressive, and hate speech analysis: From data-centric to human-centered approachabstractAnalysis of subjective texts like offensive content or hate speech is a great challenge, especially regarding annotation process. Most of current annotation procedures are aimed at achieving a high level of agreement in order to generate a high quality reference source. However, the annotation guidelines for subjective content may restrict the annotators’ freedom of decision making . Motivated by a moderate annotation agreement in offensive content datasets, we hypothesize that personalized approaches to offensive content identification should be in place. Thus, we propose two novel perspectives of perception: group-based and individual. Using demographics of annotators as well as embeddings of their previous decisions (annotated texts), we are able to train multimodal models (including transformer-based) adjusted to personal or community profiles. Based on the agreement of individuals and groups, we experimentally showed that annotator group agreeability strongly correlates with offensive content recognition quality. The proposed personalized approaches enabled us to create models adaptable to personal user beliefs rather than to agreed offensiveness understanding. Overall, our individualized approaches to offensive content classification outperform classic data-centric methods that generalize offensiveness perception and it refers to all six tested models. Additionally, we developed requirements for annotation procedures, personalization and content processing to make the solutions human-centered. Jan Kocon, Alicja Figas, Marcin Gruza, Daria Puchalska, Tomasz Kajdanowicz, Przemyslaw Kazienko |
Inf. Process. Manag. | 1 |
| 2021 | Mapping WordNet onto human brain connectome in emotion processing and semantic similarity recognitionabstractIn this article we extend a WordNet structure with relations linking synsets to Desikan’s brain regions. Based on lexicographer files and WordNet Domains the mapping goes from synset semantic categories to behavioural and cognitive functions and then directly to brain lobes. A human brain connectome (HBC) adjacency matrix was utilised to capture transition probabilities between brain regions. We evaluated the new structure in several tasks related to semantic similarity and emotion processing using brain-expanded Princeton WordNet (207k LUs) and Polish WordNet (285k LUs, 30k annotated with valence, arousal and 8 basic emotions). A novel HBC vector representation turned out to be significantly better than proposed baselines. Jan Kocon, Marek Maziarz |
Inf. Process. Manag. | 1 |
| 2020 | Cross-lingual deep neural transfer learning in sentiment analysisabstractIn this article, we present a novel technique for the use of language-agnostic sentence representations to adapt the model trained on texts in Polish (as a low-resource language) to recognize polarity in texts in other (high-resource) languages. The first model focuses on the creation of a language-agnostic representation of each sentence. The second one aims to predict the sentiment of the text based on these sentence representations. Besides models evaluation on PolEmo 1.0 Sentiment Corpus, we also conduct a proof of concept for using a deep neural network model trained only on language-agnostic embeddings of texts in Polish to predict the sentiment of the texts in MultiEmo-Test 1.0 Sentiment Corpus, containing PolEmo 1.0 test datasets translated into eight different languages: Dutch, English, French, German, Italian, Portuguese, Russian and Spanish. Both corpora are publicly available under a Creative Commons copyright license. Kamil Kanclerz, Piotr Milkowski, Jan Kocon |
KES | 3 |
| 2019 | Multi-Level Sentiment Analysis of PolEmo 2.0: Extended Corpus of Multi-Domain Consumer ReviewsabstractIn this article we present an extended version of PolEmo -a corpus of consumer reviews from 4 domains: medicine, hotels, products and school.Current version (PolEmo 2.0) contains 8,216 reviews having 57,466 sentences.Each text and sentence was manually annotated with sentiment in 2+1 scheme, which gives a total of 197,046 annotations.We obtained a high value of Positive Specific Agreement, which is 0.91 for texts and 0.88 for sentences.PolEmo 2.0 is publicly available under a Creative Commons copyright license.We explored recent deep learning approaches for the recognition of sentiment, such as Bidirectional Long Short-Term Memory (BiL-STM) and Bidirectional Encoder Representations from Transformers (BERT). Jan Kocon, Piotr Milkowski, Monika Zasko-Zielinska |
CoNLL | 1 |
| 2019 | Propagation of emotions, arousal and polarity in WordNet using Heterogeneous Structured Synset EmbeddingsabstractIn this paper we present a novel method for emotive propagation in a wordnet based on a large emotive seed.We introduce a sense-level emotive lexicon annotated with polarity, arousal and emotions.The data were annotated as a part of a large study involving over 20,000 participants.A total of 30,000 lexical units in Polish WordNet were described with metadata, each unit received about 50 annotations concerning polarity, arousal and 8 basic emotions, marked on a multilevel scale.We present a preliminary approach to propagating emotive metadata to unlabeled lexical units based on the distribution of manual annotations using logistic regression and description of mixed synset embeddings based on our Heterogeneous Structured Synset Embeddings. Jan Kocon, Arkadiusz Janz |
GWC | 1 |
| 2018 | Classifier-based Polarity Propagation in a WordNet
Jan Kocon, Arkadiusz Janz, Maciej Piasecki |
LREC | 1 |
| 2018 | Context-sensitive Sentiment Propagation in WordNetabstractIn this paper we present a comprehensive overview of recent methods of the sentiment propagation in a wordnet.Next, we propose a fully automated method called Classifier-based Polarity Propagation, which utilises a very rich set of features, where most of them are based on wordnet relation types, multi-level bag-ofsynsets and bag-of-polarities.We have evaluated our solution using manually annotated part of plWordNet 3.1 emo, which contains more than 83k manual sentiment annotations, covering more than 41k synsets.We have demonstrated that in comparison to existing rule-based methods using a specific narrow set of semantic relations our method has achieved statistically significant and better results starting with the same seed synsets. Jan Kocon, Arkadiusz Janz, Maciej Piasecki |
GWC | 1 |
| 2017 | Supervised approach to recognise Polish temporal expressions and rule-based interpretation of timexesabstractAbstract A key challenge of the Information Extraction in Natural Language Processing is the ability to recognise and classify temporal expressions (timexes). It is a crucial source of information about when something happens, how often something occurs or how long something lasts. Timexes extracted automatically from text, play a major role in many Information Extraction systems, such as question answering or event recognition. We prepared a broad specification of Polish timexes – PLIMEX. It is based on the state-of-the-art annotation guidelines for English, mainly TIMEX2 and TIMEX3 (a part of TimeML – Markup Language for Temporal and Event Expressions). We have expanded our specification for a description of the local meaning of timexes, based on LTIMEX annotation guidelines for English. Temporal description supports further event identification and extends event description model, focussing on anchoring events in time, events ordering and reasoning about the persistence of events. We prepared the specification, which is designed to address these issues, and we annotated all documents in Polish Corpus of Wroclaw University of Technology (KPWr) using our annotation guidelines. We also adapted our Liner2 machine learning system to recognise Polish timexes and we propose two-phase method to select a subset of features for Conditional Random Fields sequence labelling method. This article presents the whole process of corpus annotation, evaluation of inter-annotator agreement, extending Liner2 system with new features and evaluation of the recognition models before and after feature selection with the analysis of statistical significance of differences. Liner2 with presented models is available as open source software under the GNU General Public License. Jan Kocon, Michal Marcinczuk |
Nat. Lang. Eng. | 1 |
| 2012 | Inforex - a web-based tool for text corpus management and semantic annotation
Michal Marcinczuk, Jan Kocon, Bartosz Broda |
LREC | 2 |