Grigori Sidorov

dblp:68/4195 · DBLP profile ↗
← Back
42ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0003-3901-3522ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Enhancing Multi-Label Emotion Analysis and Corresponding Intensities for Ethiopian Languages
Tadesse Destaw Belay, Dawit Ketema Gete, Abinew Ali Ayele, Olga Kolesnikova, Iqra Ameer, Grigori Sidorov, Seid Muhie Yimam
LREC6
2026 Detecting Opioid Misuse on Social Media via Named Entity Recognition ( NER ) With Deep Learning
abstract
ABSTRACT The opioid overdose epidemic constitutes a critical public health crisis, necessitating advanced surveillance tools to enable timely intervention. Social media platforms provide a real‐time source of information on drug‐related behaviours. However, extracting structured knowledge from their informal, slang‐heavy and fragmented text presents significant technical challenges. While Named Entity Recognition (NER) enables the automated identification of drug‐related entities. Prior work has focused mainly on general biomedical domains, with limited exploration of domain‐specific scenarios. To address this gap, this study makes three key contributions. First, we present the first manually annotated, domain‐specific NER dataset for opioid overdose, comprising eight unique entity types, with a particular focus on routes of administration, sourced from Reddit posts spanning January 2021 to December 2023. Second, we provide a detailed description of the annotation process and guidelines, and systematically discuss challenges encountered during annotation, including slang, fragmented expressions and ambiguous language commonly found in social media posts. Third, we conduct an extensive series of experiments, including classical machine learning, deep learning with pretrained embeddings combined with CRF decoding, and transformer‐based models, evaluated at both token and entity levels. The proposed model achieved strong and balanced performance, with F1‐scores of 0.891 at both token and entity levels, highlighting the effectiveness of domain‐specific modelling for opioid overdose surveillance on social media.
Muhammad Ahmad 0004, Rita Orji, Fida Ullah, Ildar Z. Batyrshin, Grigori Sidorov
Expert Syst. J. Knowl. Eng.5
2026 FC-KAN: Function combinations in Kolmogorov-Arnold networks
Hoang Thang Ta, Duy-Quy Thai, Abu Bakar Siddiqur Rahman, Grigori Sidorov, Alexander F. Gelbukh
Inf. Sci.4
2025 CULEMO: Cultural Lenses on Emotion - Benchmarking LLMs for Cross-Cultural Emotion Understanding
abstract
NLP research has increasingly focused on subjective tasks such as emotion analysis. However, existing emotion benchmarks suffer fromtwo major shortcomings: (1) they largely rely on keyword-based emotion recognition, overlooking crucial cultural dimensions required fordeeper emotion understanding, and (2) many are created by translating English-annotated data into other languages, leading to potentially unreliable evaluation. To address these issues, we introduce Cultural Lenses on Emotion (CuLEmo), the first benchmark designedto evaluate culture-aware emotion prediction across six languages: Amharic, Arabic, English, German, Hindi, and Spanish. CuLEmocomprises 400 crafted questions per language, each requiring nuanced cultural reasoning and understanding. We use this benchmark to evaluate several state-of-the-art LLMs on culture-aware emotion prediction and sentiment analysis tasks. Our findings reveal that (1) emotion conceptualizations vary significantly across languages and cultures, (2) LLMs performance likewise varies by language and cultural context, and (3) prompting in English with explicit country context often outperforms in-language prompts for culture-aware emotion and sentiment understanding. The dataset and evaluation code is available.
Tadesse Destaw Belay, Ahmed Haj Ahmed, Alvin Grissom II, Iqra Ameer, Grigori Sidorov, Olga Kolesnikova, Seid Muhie Yimam
ACL (1)5
2025 Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding
abstract
Large Language Models (LLMs) show promising learning and reasoning abilities. Compared to other NLP tasks, multilingual and multi-label emotion evaluation tasks are under-explored in LLMs. In this paper, we present EthioEmo, a multi-label emotion classification dataset for four Ethiopian languages, namely, Amharic (amh), Afan Oromo (orm), Somali (som), and Tigrinya (tir). We perform extensive experiments with an additional English multi-label emotion dataset from SemEval 2018 Task 1. Our evaluation includes encoder-only, encoder-decoder, and decoder-only language models. We compare zero and few-shot approaches of LLMs to fine-tuning smaller language models. The results show that accurate multi-label emotion classification is still insufficient even for high-resource languages such as English, and there is a large gap between the performance of high-resource and low-resource languages. The results also show varying performance levels depending on the language and model type. EthioEmo is available publicly to further improve the understanding of emotions in language models and how people convey emotions through various languages.
Tadesse Destaw Belay, Israel Abebe Azime, Abinew Ali Ayele, Grigori Sidorov, Dietrich Klakow, Philipp Slusallek, Olga Kolesnikova, Seid Muhie Yimam
COLING4
2025 Interpretation of Myers-Briggs Type Indicator personality profiles based on ambivert continuum scale
Sabur Butt, Grigori Sidorov, Alexander F. Gelbukh
Expert Syst. Appl.2
2025 PoliGuilt: Two level guilt detection from social media texts
Abdul Gafar Manuel Meque, Fazlourrahman Balouchzahi, Alexander F. Gelbukh, Grigori Sidorov
Expert Syst. Appl.4
2025 UrduHope: Analysis of hope and hopelessness in Urdu texts
Fazlourrahman Balouchzahi, Sabur Butt, Maaz Amjad, Grigori Sidorov, Alexander F. Gelbukh
Knowl. Based Syst.4
2024 CLEF 2024 JOKER Lab: Automatic Humour Analysis
Liana Ermakova, Anne-Gwenn Bosser, Tristan Miller, Tremaine Thomas-Young, Victor Manuel Palma-Preciado, Grigori Sidorov, Adam Jatowt
ECIR (6)6
2024 Composer classification using melodic combinatorial n-grams
Daniel Alejandro Pérez Alvarez, Alexander F. Gelbukh, Grigori Sidorov
Expert Syst. Appl.3
2023 Science for Fun: The CLEF 2023 JOKER Track on Automatic Wordplay Analysis
Liana Ermakova, Tristan Miller, Anne-Gwenn Bosser, Victor Manuel Palma-Preciado, Grigori Sidorov, Adam Jatowt
ECIR (3)5
2023 Multi-label emotion classification in texts using transfer learning
Iqra Ameer, Necva Bölücü, Muhammad Hammad Fahim Siddiqui, Burcu Can, Grigori Sidorov, Alexander F. Gelbukh
Expert Syst. Appl.5
2023 ReDDIT: Regret detection and domain identification from text
Fazlourrahman Balouchzahi, Sabur Butt, Grigori Sidorov, Alexander F. Gelbukh
Expert Syst. Appl.3
2023 PolyHope: Two-level hope speech detection from tweets
Fazlourrahman Balouchzahi, Grigori Sidorov, Alexander F. Gelbukh
Expert Syst. Appl.2
2023 Sarcasm detection framework using context, emotion and sentiment features
Oxana Vitman, Yevhen Kostiuk, Grigori Sidorov, Alexander F. Gelbukh
Expert Syst. Appl.3
2020 Data Augmentation using Machine Translation for Fake News Detection in the Urdu Language
abstract
The task of fake news detection is to distinguish legitimate news articles that describe real facts from those which convey deceiving and fictitious information. As the fake news phenomenon is omnipresent across all languages, it is crucial to be able to efficiently solve this problem for languages other than English. A common approach to this task is supervised classification using features of various complexity. Yet supervised machine learning requires substantial amount of annotated data. For English and a small number of other languages, annotated data availability is much higher, whereas for the vast majority of languages, it is almost scarce. We investigate whether machine translation at its present state could be successfully used as an automated technique for annotated corpora creation and augmentation for fake news detection focusing on the English-Urdu language pair. We train a fake news classifier for Urdu on (1) the manually annotated dataset originally in Urdu and (2) the machine-translated version of an existing annotated fake news dataset originally in English. We show that at the present state of machine translation quality for the English-Urdu language pair, the fully automated data augmentation through machine translation did not provide improvement for fake news detection in Urdu.
Maaz Amjad, Grigori Sidorov, Alisa Zhila
LREC2
2019 Unsupervised sentence representations as word information series: Revisiting TF-IDF
Ignacio Arroyo-Fernández, Carlos-Francisco Méndez-Cruz, Gerardo Sierra, Juan-Manuel Torres-Moreno, Grigori Sidorov
Comput. Speech Lang.5
2018 Searching for Cerebrovascular Disease Optimal Treatment Recommendations Applying Partially Observable Markov Decision Processes
abstract
Partially observable Markov decision processes (POMDPs) are mathematical models for the planning of action sequences under conditions of uncertainty. Uncertainty in POMDPs is manifested in two ways: uncertainty in the perception of model states and uncertainty in the effects of actions on states. The diagnosis and treatment of cerebral vascular diseases (CVD) present this double condition of uncertainty, so we think that POMDP is the most suitable method to model them. In this paper, we propose a model of CVD that is based on observations obtained from neuroimaging studies such as computed tomography, magnetic resonance and ultrasound. The model is designed as a POMDP because the health status of the patient is not directly observable, and only can be deduced, with some probability, from the observations in the cerebral images. The components of the model (states, observations, actions, etc.) were defined based on specialized literature. A diagnosis of the patient’s health status is made and the most appropriate action for the recovery of health is recommended after introducing the observations when operating the model. Consultation of the probable state of health of the patient and alternative actions is also allowed.
Hermilo Victorio Meza, Manuel Mejía-Lavalle, Alicia Martínez Rebollar, Andrés Blanco-Ortega, Obdulia Pichardo-Lagunas, Grigori Sidorov
Int. J. Pattern Recognit. Artif. Intell.6
2018 Plagiarism Detection with Genetic-Based Parameter Tuning
abstract
A crucial step in plagiarism detection is text alignment. This task consists in finding similar text fragments between two given documents. We introduce an optimization methodology based on genetic algorithms to improve the performance of a plagiarism detection model by optimizing its input parameters. The implementation of the genetic algorithm is based on nonbinary representation of individuals, elitism selection, uniform crossover, and high mutation rate. The obtained parameter settings allow the plagiarism detection model to achieve better results than the state-of-the-art approaches.
Miguel A. Sánchez-Pérez, Alexander F. Gelbukh, Grigori Sidorov, Helena Gómez-Adorno
Int. J. Pattern Recognit. Artif. Intell.3
2017 Improving Cross-Topic Authorship Attribution: The Role of Pre-Processing
Ilia Markov, Efstathios Stamatatos, Grigori Sidorov
CICLing (2)3
2017 Application of the distributed document representation in the authorship attribution task for small corpora
Juan Pablo Posadas-Durán, Helena Gómez-Adorno, Grigori Sidorov, Ildar Z. Batyrshin, David Pinto 0001, Liliana Chanona-Hernández
Soft Comput.3
2014 Syntactic N-grams as machine learning features for natural language processing
Grigori Sidorov, Francisco Velasquez, Efstathios Stamatatos, Alexander F. Gelbukh, Liliana Chanona-Hernández
Expert Syst. Appl.1
2013 Syntactic Dependency-Based N-grams: More Evidence of Usefulness in Classification
Grigori Sidorov, Francisco Velasquez, Efstathios Stamatatos, Alexander F. Gelbukh, Liliana Chanona-Hernández
CICLing (1)1
2010 English-Spanish Large Statistical Dictionary of Inflectional Forms
Grigori Sidorov, Alberto Barrón-Cedeño, Paolo Rosso
LREC1
2010 Analysis of Definitions of Verbs in an Explanatory Dictionary for Automatic Extraction of Actants Based on Detection of Patterns
Noé Alejandro Castro-Sánchez, Grigori Sidorov
NLDB2
2010 Automatic Term Extraction Using Log-Likelihood Based Comparison with General Reference Corpus
Alexander F. Gelbukh, Grigori Sidorov, Eduardo Lavin-Villa, Liliana Chanona-Hernández
NLDB2
2009 Incorporating Linguistic Information to Statistical Word-Level Alignment
Eduardo Cendejas, Grettel Barceló, Alexander F. Gelbukh, Grigori Sidorov
CIARP4
2009 Formal Grammar for Hispanic Named Entities Analysis
Grettel Barceló, Eduardo Cendejas, Grigori Sidorov, Igor A. Bolshakov
CICLing3
2009 Hybrid Algorithm for Word-Level Alignment of Parallel Texts
Eduardo Cendejas, Grettel Barceló, Alexander F. Gelbukh, Grigori Sidorov
NLDB4
2009 Search Interface to a Mayan Glyph Database Based on Visual Characteristics
Grigori Sidorov, Obdulia Pichardo-Lagunas, Liliana Chanona-Hernández
NLDB1
2008 Division of Spanish Words into Morphemes with a Genetic Algorithm
Alexander F. Gelbukh, Grigori Sidorov, Diego Lara-Reyes, Liliana Chanona-Hernández
NLDB2
2007 Lexical-Based Alignment for Reconstruction of Structure in Parallel Texts
Alexander F. Gelbukh, Grigori Sidorov, Liliana Chanona-Hernández
NLDB2
2006 Alignment of Paragraphs in Bilingual Texts Using Bilingual Dictionaries and Dynamic Programming
Alexander F. Gelbukh, Grigori Sidorov
CIARP2
2006 Generation of Natural Language Explanations of Rules in an Expert System
María de los Ángeles Alonso-Lavernia, Argelio De-la-Cruz-Rivera, Grigori Sidorov
CICLing3
2005 On Some Optimization Heuristics for Lesk-Like WSD Algorithms
Alexander F. Gelbukh, Grigori Sidorov, Sang-Yong Han
NLDB2
2004 Automatic Syntactic Analysis for Detection of Word Combinations
Alexander F. Gelbukh, Grigori Sidorov, Sang-Yong Han, Erika Hernández-Rubio
CICLing2
2003 Approach to Construction of Automatic Morphological Analysis Systems for Inflective Languages with Little Effort
Alexander F. Gelbukh, Grigori Sidorov
CICLing2
2003 Tool for Computer-Aided Spanish Word Sense Disambiguation
Yoel Ledo Mezquita, Grigori Sidorov, Alexander F. Gelbukh
CICLing2
2002 Automatic Selection of Defining Vocabulary in an Explanatory Dictionary
Alexander F. Gelbukh, Grigori Sidorov
CICLing2
2002 Compilation of a Spanish Representative Corpus
Alexander F. Gelbukh, Grigori Sidorov, Liliana Chanona-Hernández
CICLing2
2001 Zipf and Heaps Laws' Coefficients Depend on Language
Alexander F. Gelbukh, Grigori Sidorov
CICLing2
2001 Automatic detection of semantically primitive words using their reachability in an explanatory dictionary
abstract
We suggest the method that permits building a set of candidates to be considered semantic primitives from the standard explanatory dictionary. Our method is based on the frequencies of the words that are reachable in a semantic network constructed from the dictionary. The method implements word sense disambiguation techniques, network construction, and reachability analysis. In part of word sense disambiguation we use an improved Lesk's algorithm. In the part of analysis of reachability we show that the words to which our algorithm assigns high weight, are plausible candidates to be semantic primitives. It is also shown that better candidates to semantic primitives should be included in short vicious cycles, which is detected by our algorithm. We applied the method to a rather large Spanish explanatory dictionary.
Grigori Sidorov, Alexander F. Gelbukh
SMC1