VLDB 2026 Research / reviewers in the wild / expert
Grigori Sidorov
dblp:68/4195
· DBLP profile ↗
42ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0003-3901-3522ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Multi-Label Emotion Analysis and Corresponding Intensities for Ethiopian Languages
Tadesse Destaw Belay, Dawit Ketema Gete, Abinew Ali Ayele, Olga Kolesnikova, Iqra Ameer, Grigori Sidorov, Seid Muhie Yimam |
LREC | 6 |
| 2026 | Detecting Opioid Misuse on Social Media via Named Entity Recognition ( NER ) With Deep LearningabstractABSTRACT The opioid overdose epidemic constitutes a critical public health crisis, necessitating advanced surveillance tools to enable timely intervention. Social media platforms provide a real‐time source of information on drug‐related behaviours. However, extracting structured knowledge from their informal, slang‐heavy and fragmented text presents significant technical challenges. While Named Entity Recognition (NER) enables the automated identification of drug‐related entities. Prior work has focused mainly on general biomedical domains, with limited exploration of domain‐specific scenarios. To address this gap, this study makes three key contributions. First, we present the first manually annotated, domain‐specific NER dataset for opioid overdose, comprising eight unique entity types, with a particular focus on routes of administration, sourced from Reddit posts spanning January 2021 to December 2023. Second, we provide a detailed description of the annotation process and guidelines, and systematically discuss challenges encountered during annotation, including slang, fragmented expressions and ambiguous language commonly found in social media posts. Third, we conduct an extensive series of experiments, including classical machine learning, deep learning with pretrained embeddings combined with CRF decoding, and transformer‐based models, evaluated at both token and entity levels. The proposed model achieved strong and balanced performance, with F1‐scores of 0.891 at both token and entity levels, highlighting the effectiveness of domain‐specific modelling for opioid overdose surveillance on social media. Muhammad Ahmad 0004, Rita Orji, Fida Ullah, Ildar Z. Batyrshin, Grigori Sidorov |
Expert Syst. J. Knowl. Eng. | 5 |
| 2026 | FC-KAN: Function combinations in Kolmogorov-Arnold networks
Hoang Thang Ta, Duy-Quy Thai, Abu Bakar Siddiqur Rahman, Grigori Sidorov, Alexander F. Gelbukh |
Inf. Sci. | 4 |
| 2025 | CULEMO: Cultural Lenses on Emotion - Benchmarking LLMs for Cross-Cultural Emotion UnderstandingabstractNLP research has increasingly focused on subjective tasks such as emotion analysis. However, existing emotion benchmarks suffer fromtwo major shortcomings: (1) they largely rely on keyword-based emotion recognition, overlooking crucial cultural dimensions required fordeeper emotion understanding, and (2) many are created by translating English-annotated data into other languages, leading to potentially unreliable evaluation. To address these issues, we introduce Cultural Lenses on Emotion (CuLEmo), the first benchmark designedto evaluate culture-aware emotion prediction across six languages: Amharic, Arabic, English, German, Hindi, and Spanish. CuLEmocomprises 400 crafted questions per language, each requiring nuanced cultural reasoning and understanding. We use this benchmark to evaluate several state-of-the-art LLMs on culture-aware emotion prediction and sentiment analysis tasks. Our findings reveal that (1) emotion conceptualizations vary significantly across languages and cultures, (2) LLMs performance likewise varies by language and cultural context, and (3) prompting in English with explicit country context often outperforms in-language prompts for culture-aware emotion and sentiment understanding. The dataset and evaluation code is available. Tadesse Destaw Belay, Ahmed Haj Ahmed, Alvin Grissom II, Iqra Ameer, Grigori Sidorov, Olga Kolesnikova, Seid Muhie Yimam |
ACL (1) | 5 |
| 2025 | Evaluating the Capabilities of Large Language Models for Multi-label Emotion UnderstandingabstractLarge Language Models (LLMs) show promising learning and reasoning abilities. Compared to other NLP tasks, multilingual and multi-label emotion evaluation tasks are under-explored in LLMs. In this paper, we present EthioEmo, a multi-label emotion classification dataset for four Ethiopian languages, namely, Amharic (amh), Afan Oromo (orm), Somali (som), and Tigrinya (tir). We perform extensive experiments with an additional English multi-label emotion dataset from SemEval 2018 Task 1. Our evaluation includes encoder-only, encoder-decoder, and decoder-only language models. We compare zero and few-shot approaches of LLMs to fine-tuning smaller language models. The results show that accurate multi-label emotion classification is still insufficient even for high-resource languages such as English, and there is a large gap between the performance of high-resource and low-resource languages. The results also show varying performance levels depending on the language and model type. EthioEmo is available publicly to further improve the understanding of emotions in language models and how people convey emotions through various languages. Tadesse Destaw Belay, Israel Abebe Azime, Abinew Ali Ayele, Grigori Sidorov, Dietrich Klakow, Philipp Slusallek, Olga Kolesnikova, Seid Muhie Yimam |
COLING | 4 |
| 2025 | Interpretation of Myers-Briggs Type Indicator personality profiles based on ambivert continuum scale
Sabur Butt, Grigori Sidorov, Alexander F. Gelbukh |
Expert Syst. Appl. | 2 |
| 2025 | PoliGuilt: Two level guilt detection from social media texts
Abdul Gafar Manuel Meque, Fazlourrahman Balouchzahi, Alexander F. Gelbukh, Grigori Sidorov |
Expert Syst. Appl. | 4 |
| 2025 | UrduHope: Analysis of hope and hopelessness in Urdu texts
Fazlourrahman Balouchzahi, Sabur Butt, Maaz Amjad, Grigori Sidorov, Alexander F. Gelbukh |
Knowl. Based Syst. | 4 |
| 2024 | CLEF 2024 JOKER Lab: Automatic Humour Analysis
Liana Ermakova, Anne-Gwenn Bosser, Tristan Miller, Tremaine Thomas-Young, Victor Manuel Palma-Preciado, Grigori Sidorov, Adam Jatowt |
ECIR (6) | 6 |
| 2024 | Composer classification using melodic combinatorial n-grams
Daniel Alejandro Pérez Alvarez, Alexander F. Gelbukh, Grigori Sidorov |
Expert Syst. Appl. | 3 |
| 2023 | Science for Fun: The CLEF 2023 JOKER Track on Automatic Wordplay Analysis
Liana Ermakova, Tristan Miller, Anne-Gwenn Bosser, Victor Manuel Palma-Preciado, Grigori Sidorov, Adam Jatowt |
ECIR (3) | 5 |
| 2023 | Multi-label emotion classification in texts using transfer learning
Iqra Ameer, Necva Bölücü, Muhammad Hammad Fahim Siddiqui, Burcu Can, Grigori Sidorov, Alexander F. Gelbukh |
Expert Syst. Appl. | 5 |
| 2023 | ReDDIT: Regret detection and domain identification from text
Fazlourrahman Balouchzahi, Sabur Butt, Grigori Sidorov, Alexander F. Gelbukh |
Expert Syst. Appl. | 3 |
| 2023 | PolyHope: Two-level hope speech detection from tweets
Fazlourrahman Balouchzahi, Grigori Sidorov, Alexander F. Gelbukh |
Expert Syst. Appl. | 2 |
| 2023 | Sarcasm detection framework using context, emotion and sentiment features
Oxana Vitman, Yevhen Kostiuk, Grigori Sidorov, Alexander F. Gelbukh |
Expert Syst. Appl. | 3 |
| 2020 | Data Augmentation using Machine Translation for Fake News Detection in the Urdu LanguageabstractThe task of fake news detection is to distinguish legitimate news articles that describe real facts from those which convey deceiving and fictitious information. As the fake news phenomenon is omnipresent across all languages, it is crucial to be able to efficiently solve this problem for languages other than English. A common approach to this task is supervised classification using features of various complexity. Yet supervised machine learning requires substantial amount of annotated data. For English and a small number of other languages, annotated data availability is much higher, whereas for the vast majority of languages, it is almost scarce. We investigate whether machine translation at its present state could be successfully used as an automated technique for annotated corpora creation and augmentation for fake news detection focusing on the English-Urdu language pair. We train a fake news classifier for Urdu on (1) the manually annotated dataset originally in Urdu and (2) the machine-translated version of an existing annotated fake news dataset originally in English. We show that at the present state of machine translation quality for the English-Urdu language pair, the fully automated data augmentation through machine translation did not provide improvement for fake news detection in Urdu. Maaz Amjad, Grigori Sidorov, Alisa Zhila |
LREC | 2 |
| 2019 | Unsupervised sentence representations as word information series: Revisiting TF-IDF
Ignacio Arroyo-Fernández, Carlos-Francisco Méndez-Cruz, Gerardo Sierra, Juan-Manuel Torres-Moreno, Grigori Sidorov |
Comput. Speech Lang. | 5 |
| 2018 | Searching for Cerebrovascular Disease Optimal Treatment Recommendations Applying Partially Observable Markov Decision ProcessesabstractPartially observable Markov decision processes (POMDPs) are mathematical models for the planning of action sequences under conditions of uncertainty. Uncertainty in POMDPs is manifested in two ways: uncertainty in the perception of model states and uncertainty in the effects of actions on states. The diagnosis and treatment of cerebral vascular diseases (CVD) present this double condition of uncertainty, so we think that POMDP is the most suitable method to model them. In this paper, we propose a model of CVD that is based on observations obtained from neuroimaging studies such as computed tomography, magnetic resonance and ultrasound. The model is designed as a POMDP because the health status of the patient is not directly observable, and only can be deduced, with some probability, from the observations in the cerebral images. The components of the model (states, observations, actions, etc.) were defined based on specialized literature. A diagnosis of the patient’s health status is made and the most appropriate action for the recovery of health is recommended after introducing the observations when operating the model. Consultation of the probable state of health of the patient and alternative actions is also allowed. Hermilo Victorio Meza, Manuel Mejía-Lavalle, Alicia Martínez Rebollar, Andrés Blanco-Ortega, Obdulia Pichardo-Lagunas, Grigori Sidorov |
Int. J. Pattern Recognit. Artif. Intell. | 6 |
| 2018 | Plagiarism Detection with Genetic-Based Parameter TuningabstractA crucial step in plagiarism detection is text alignment. This task consists in finding similar text fragments between two given documents. We introduce an optimization methodology based on genetic algorithms to improve the performance of a plagiarism detection model by optimizing its input parameters. The implementation of the genetic algorithm is based on nonbinary representation of individuals, elitism selection, uniform crossover, and high mutation rate. The obtained parameter settings allow the plagiarism detection model to achieve better results than the state-of-the-art approaches. Miguel A. Sánchez-Pérez, Alexander F. Gelbukh, Grigori Sidorov, Helena Gómez-Adorno |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2017 | Improving Cross-Topic Authorship Attribution: The Role of Pre-Processing
Ilia Markov, Efstathios Stamatatos, Grigori Sidorov |
CICLing (2) | 3 |
| 2017 | Application of the distributed document representation in the authorship attribution task for small corpora
Juan Pablo Posadas-Durán, Helena Gómez-Adorno, Grigori Sidorov, Ildar Z. Batyrshin, David Pinto 0001, Liliana Chanona-Hernández |
Soft Comput. | 3 |
| 2014 | Syntactic N-grams as machine learning features for natural language processing
Grigori Sidorov, Francisco Velasquez, Efstathios Stamatatos, Alexander F. Gelbukh, Liliana Chanona-Hernández |
Expert Syst. Appl. | 1 |
| 2013 | Syntactic Dependency-Based N-grams: More Evidence of Usefulness in Classification
Grigori Sidorov, Francisco Velasquez, Efstathios Stamatatos, Alexander F. Gelbukh, Liliana Chanona-Hernández |
CICLing (1) | 1 |
| 2010 | English-Spanish Large Statistical Dictionary of Inflectional Forms
Grigori Sidorov, Alberto Barrón-Cedeño, Paolo Rosso |
LREC | 1 |
| 2010 | Analysis of Definitions of Verbs in an Explanatory Dictionary for Automatic Extraction of Actants Based on Detection of Patterns
Noé Alejandro Castro-Sánchez, Grigori Sidorov |
NLDB | 2 |
| 2010 | Automatic Term Extraction Using Log-Likelihood Based Comparison with General Reference Corpus
Alexander F. Gelbukh, Grigori Sidorov, Eduardo Lavin-Villa, Liliana Chanona-Hernández |
NLDB | 2 |
| 2009 | Incorporating Linguistic Information to Statistical Word-Level Alignment
Eduardo Cendejas, Grettel Barceló, Alexander F. Gelbukh, Grigori Sidorov |
CIARP | 4 |
| 2009 | Formal Grammar for Hispanic Named Entities Analysis
Grettel Barceló, Eduardo Cendejas, Grigori Sidorov, Igor A. Bolshakov |
CICLing | 3 |
| 2009 | Hybrid Algorithm for Word-Level Alignment of Parallel Texts
Eduardo Cendejas, Grettel Barceló, Alexander F. Gelbukh, Grigori Sidorov |
NLDB | 4 |
| 2009 | Search Interface to a Mayan Glyph Database Based on Visual Characteristics
Grigori Sidorov, Obdulia Pichardo-Lagunas, Liliana Chanona-Hernández |
NLDB | 1 |
| 2008 | Division of Spanish Words into Morphemes with a Genetic Algorithm
Alexander F. Gelbukh, Grigori Sidorov, Diego Lara-Reyes, Liliana Chanona-Hernández |
NLDB | 2 |
| 2007 | Lexical-Based Alignment for Reconstruction of Structure in Parallel Texts
Alexander F. Gelbukh, Grigori Sidorov, Liliana Chanona-Hernández |
NLDB | 2 |
| 2006 | Alignment of Paragraphs in Bilingual Texts Using Bilingual Dictionaries and Dynamic Programming
Alexander F. Gelbukh, Grigori Sidorov |
CIARP | 2 |
| 2006 | Generation of Natural Language Explanations of Rules in an Expert System
María de los Ángeles Alonso-Lavernia, Argelio De-la-Cruz-Rivera, Grigori Sidorov |
CICLing | 3 |
| 2005 | On Some Optimization Heuristics for Lesk-Like WSD Algorithms
Alexander F. Gelbukh, Grigori Sidorov, Sang-Yong Han |
NLDB | 2 |
| 2004 | Automatic Syntactic Analysis for Detection of Word Combinations
Alexander F. Gelbukh, Grigori Sidorov, Sang-Yong Han, Erika Hernández-Rubio |
CICLing | 2 |
| 2003 | Approach to Construction of Automatic Morphological Analysis Systems for Inflective Languages with Little Effort
Alexander F. Gelbukh, Grigori Sidorov |
CICLing | 2 |
| 2003 | Tool for Computer-Aided Spanish Word Sense Disambiguation
Yoel Ledo Mezquita, Grigori Sidorov, Alexander F. Gelbukh |
CICLing | 2 |
| 2002 | Automatic Selection of Defining Vocabulary in an Explanatory Dictionary
Alexander F. Gelbukh, Grigori Sidorov |
CICLing | 2 |
| 2002 | Compilation of a Spanish Representative Corpus
Alexander F. Gelbukh, Grigori Sidorov, Liliana Chanona-Hernández |
CICLing | 2 |
| 2001 | Zipf and Heaps Laws' Coefficients Depend on Language
Alexander F. Gelbukh, Grigori Sidorov |
CICLing | 2 |
| 2001 | Automatic detection of semantically primitive words using their reachability in an explanatory dictionaryabstractWe suggest the method that permits building a set of candidates to be considered semantic primitives from the standard explanatory dictionary. Our method is based on the frequencies of the words that are reachable in a semantic network constructed from the dictionary. The method implements word sense disambiguation techniques, network construction, and reachability analysis. In part of word sense disambiguation we use an improved Lesk's algorithm. In the part of analysis of reachability we show that the words to which our algorithm assigns high weight, are plausible candidates to be semantic primitives. It is also shown that better candidates to semantic primitives should be included in short vicious cycles, which is detected by our algorithm. We applied the method to a rather large Spanish explanatory dictionary. Grigori Sidorov, Alexander F. Gelbukh |
SMC | 1 |