Alexander F. Gelbukh

dblp:g/AlexanderFGelbukh · DBLP profile ↗
← Back
106ranked-venue papers
15as first author
17since 2021 · last 2026
0000-0001-7845-9039ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 91 · 15 first-author · 12 since 2021Databases, data management, data science and information retrieval · 25 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 8Human-computer interaction and ubiquitous computing · 7 · 2 since 2021
YearPublicationVenuePosition
2026 FC-KAN: Function combinations in Kolmogorov-Arnold networks
Hoang Thang Ta, Duy-Quy Thai, Abu Bakar Siddiqur Rahman, Grigori Sidorov, Alexander F. Gelbukh
Inf. Sci.5
2026 Beyond Majority Voting: Agreement-Based Clustering to Model Annotator Perspectives in Subjective NLP Tasks
abstract
Abstract Disagreement in annotation is a common phenomenon in the development of NLP datasets and serves as a valuable source of insight. While majority voting remains the dominant strategy for aggregating labels, recent work has explored modeling individual annotators to preserve their perspectives. However, modeling each annotator is resource-intensive and remains underexplored across various NLP tasks. We propose an agreement-based clustering technique to model the disagreement between the annotators. We conduct comprehensive experiments in 40 datasets in 18 typologically diverse languages, covering three subjective NLP tasks: sentiment analysis, emotion classification, and hate speech detection. We evaluate four aggregation approaches: majority vote, ensemble, multi-label, and multitask. The results demonstrate that agreement-based clustering can leverage the full spectrum of annotator perspectives and significantly enhance classification performance in subjective NLP tasks compared to majority voting and individual annotator modeling. Regarding the aggregation approach, the multi-label and multitask approaches are better for modeling clustered annotators than an ensemble and model majority vote. The dataset is publicly available in GitHub: https://github.com/Tadesse-Destaw/Beyond-Majority-Voting.
Tadesse Destaw Belay, Ibrahim Said Ahmad, Idris Abdulmumin, Abinew Ali Ayele, Alexander F. Gelbukh, Eusebio Ricárdez-Vázquez, Olga Kolesnikova, Shamsuddeen Hassan Muhammad, Seid Muhie Yimam
Trans. Assoc. Comput. Linguistics5
2025 Interpretation of Myers-Briggs Type Indicator personality profiles based on ambivert continuum scale
Sabur Butt, Grigori Sidorov, Alexander F. Gelbukh
Expert Syst. Appl.3
2025 PoliGuilt: Two level guilt detection from social media texts
Abdul Gafar Manuel Meque, Fazlourrahman Balouchzahi, Alexander F. Gelbukh, Grigori Sidorov
Expert Syst. Appl.3
2025 UrduHope: Analysis of hope and hopelessness in Urdu texts
Fazlourrahman Balouchzahi, Sabur Butt, Maaz Amjad, Grigori Sidorov, Alexander F. Gelbukh
Knowl. Based Syst.5
2024 Towards Interpretable Emotion Classification: Evaluating LIME, SHAP, and Generative AI for Decision Explanations
abstract
This paper explores the classification of multi-label emotions utilizing fine-tuned RoBERTa base and zero-shot GPT4 models, with experiments conducted on the SemEval 2018 E-c dataset encompassing 11 emotions, where more than one label is allowed for a text. Employing SHAP and LIME for RoBERTa explanations and generative AI for GPT4, we assess the sufficiency of explanations using the BERT score metric. We show the explanations generated by LIME and SHAP visually using different plots. The BERT score indicates that generative AI produces better explanations than the statistical models, providing deeper insights into emotion selection, with a BERT score of 59.66% compared to SHAP-RoBERTa's 54.17% and LIME-RoBERTa's 53.22%. This shows the potential of generative AI in revealing the reasoning behind decisions within complex emotional contexts. Though the performance is superior, we also discuss the limitations of these models that hinder wide-scale adoption.
Muhammad Hammad Fahim Siddiqui, Diana Inkpen, Alexander F. Gelbukh
IV3
2024 NLP Progress in Indigenous Latin American Languages
abstract
Atnafu Tonja, Fazlourrahman Balouchzahi, Sabur Butt, Olga Kolesnikova, Hector Ceballos, Alexander Gelbukh, Thamar Solorio. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Atnafu Lambebo Tonja, Fazlourrahman Balouchzahi, Sabur Butt, Olga Kolesnikova, Hector G. Ceballos, Alexander F. Gelbukh, Thamar Solorio
NAACL-HLT6
2024 Composer classification using melodic combinatorial n-grams
Daniel Alejandro Pérez Alvarez, Alexander F. Gelbukh, Grigori Sidorov
Expert Syst. Appl.2
2023 Multi-label emotion classification in texts using transfer learning
Iqra Ameer, Necva Bölücü, Muhammad Hammad Fahim Siddiqui, Burcu Can, Grigori Sidorov, Alexander F. Gelbukh
Expert Syst. Appl.6
2023 ReDDIT: Regret detection and domain identification from text
Fazlourrahman Balouchzahi, Sabur Butt, Grigori Sidorov, Alexander F. Gelbukh
Expert Syst. Appl.4
2023 PolyHope: Two-level hope speech detection from tweets
Fazlourrahman Balouchzahi, Grigori Sidorov, Alexander F. Gelbukh
Expert Syst. Appl.3
2023 Sarcasm detection framework using context, emotion and sentiment features
Oxana Vitman, Yevhen Kostiuk, Grigori Sidorov, Alexander F. Gelbukh
Expert Syst. Appl.4
2023 Measuring semantic gap between user-generated content and product descriptions through compression comparison in e-commerce
Carlos A. Rodriguez-Diaz, Sergio Jiménez 0001, Daniel Bejarano, Julio A. Bernal-Chávez, Alexander F. Gelbukh
Inf. Sci.5
2022 ProficiencyRank: Automatically ranking expertise in online collaborative social networks
Sergio Jiménez 0001, Fabio N. Silva, George Dueñas, Alexander F. Gelbukh
Inf. Sci.4
2022 Improving aspect-level sentiment analysis with aspect extraction
Navonil Majumder, Rishabh Bhardwaj, Soujanya Poria, Alexander F. Gelbukh, Amir Hussain 0001
Neural Comput. Appl.4
2021 Bi-Bimodal Modality Fusion for Correlation-Controlled Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis aims to extract and integrate semantic information collected from multiple modalities to recognize the expressed emotions and sentiment in multimodal data. This research area’s major concern lies in developing an extraordinary fusion scheme that can extract and integrate key information from various modalities. However, previous work is restricted by the lack of leveraging dynamics of independence and correlation between modalities to reach top performance. To mitigate this, we propose the Bi-Bimodal Fusion Network (BBFN), a novel end-to-end network that performs fusion (relevance increment) and separation (difference increment) on pairwise modality representations. The two parts are trained simultaneously such that the combat between them is simulated. The model takes two bimodal pairs as input due to the known information imbalance among modalities. In addition, we leverage a gated control mechanism in the Transformer architecture to further improve the final output. Experimental results on three datasets (CMU-MOSI, CMU-MOSEI, and UR-FUNNY) verifies that our model significantly outperforms the SOTA. The implementation of this work is available at https://github.com/declare-lab/multimodal-deep-learning and https://github.com/declare-lab/BBFN.
Wei Han 0002, Hui Chen 0023, Alexander F. Gelbukh, Amir Zadeh 0001, Louis-Philippe Morency, Soujanya Poria
ICMI3
2021 Leveraging label hierarchy using transfer and multi-task learning: A case study on patent classification
Segun Taofeek Aroyehun, Jason Angel, Navonil Majumder, Alexander F. Gelbukh, Amir Hussain 0001
Neurocomputing4
2020 MIME: MIMicking Emotions for Empathetic Response Generation
abstract
Navonil Majumder, Pengfei Hong, Shanshan Peng, Jiankun Lu, Deepanway Ghosal, Alexander Gelbukh, Rada Mihalcea, Soujanya Poria. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Navonil Majumder, Pengfei Hong, Shanshan Peng, Jiankun Lu, Deepanway Ghosal, Alexander F. Gelbukh, Rada Mihalcea, Soujanya Poria
EMNLP (1)6
2019 DialogueRNN: An Attentive RNN for Emotion Detection in Conversations
abstract
Emotion detection in conversations is a necessary step for a number of applications, including opinion mining over chat history, social media threads, debates, argumentation mining, understanding consumer feedback in live conversations, and so on. Currently systems do not treat the parties in the conversation individually by adapting to the speaker of each utterance. In this paper, we describe a new method based on recurrent neural networks that keeps track of the individual party states throughout the conversation and uses this information for emotion classification. Our model outperforms the state-of-the-art by a significant margin on two different datasets.
Navonil Majumder, Soujanya Poria, Devamanyu Hazarika, Rada Mihalcea, Alexander F. Gelbukh, Erik Cambria
AAAI5
2019 Comparison of Text Classification Methods Using Deep Learning Neural Networks
Maaz Amjad, Alexander F. Gelbukh, Ilia Voronkov, Anna Saenko
CICLing (2)2
2019 DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation
abstract
Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, Alexander Gelbukh. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, Alexander F. Gelbukh
EMNLP/IJCNLP (1)5
2018 Addressing the Issue of Unavailability of Parallel Corpus Incorporating Monolingual Corpus on PBSMT System for English-Manipuri Translation
Amika Achom, Partha Pakray, Alexander F. Gelbukh
CICLing (1)3
2018 An Abstractive Text Summarization Using Recurrent Neural Network
Dipanwita Debnath, Partha Pakray, Ranjita Das, Alexander F. Gelbukh
CICLing (2)4
2018 IARM: Inter-Aspect Relation Modeling with Memory Networks in Aspect-Based Sentiment Analysis
abstract
Sentiment analysis has immense implications in modern businesses through user-feedback mining.Large product-based enterprises like Samsung and Apple make crucial business decisions based on the large quantity of user reviews and suggestions available in different e-commerce websites and social media platforms like Amazon and Facebook.Sentiment analysis caters to these needs by summarizing user sentiment behind a particular object.In this paper, we present a novel approach of incorporating the neighboring aspects related information into the sentiment classification of the target aspect using memory networks.Our method outperforms the state of the art by 1.6% on average in two distinct domains.
Navonil Majumder, Soujanya Poria, Alexander F. Gelbukh, Md. Shad Akhtar, Erik Cambria, Asif Ekbal
EMNLP3
2018 Plagiarism Detection with Genetic-Based Parameter Tuning
abstract
A crucial step in plagiarism detection is text alignment. This task consists in finding similar text fragments between two given documents. We introduce an optimization methodology based on genetic algorithms to improve the performance of a plagiarism detection model by optimizing its input parameters. The implementation of the genetic algorithm is based on nonbinary representation of individuals, elitism selection, uniform crossover, and high mutation rate. The obtained parameter settings allow the plagiarism detection model to achieve better results than the state-of-the-art approaches.
Miguel A. Sánchez-Pérez, Alexander F. Gelbukh, Grigori Sidorov, Helena Gómez-Adorno
Int. J. Pattern Recognit. Artif. Intell.2
2018 Multimodal sentiment analysis using hierarchical fusion with context modeling
Navonil Majumder, Devamanyu Hazarika, Alexander F. Gelbukh, Erik Cambria, Soujanya Poria
Knowl. Based Syst.3
2017 Designing an Ontology for Physical Exercise Actions
Sandeep Kumar Dash, Partha Pakray, Robert Porzel, Jan D. Smeddinck, Rainer Malaka, Alexander F. Gelbukh
CICLing (1)6
2017 Adaptation of Sentiment Analysis Techniques to Persian Language
Kia Dashtipour, Amir Hussain 0001, Alexander F. Gelbukh
CICLing (2)3
2017 Cross-domain deception detection using support vector networks
Ángel Hernández-Castañeda, Hiram Calvo, Alexander F. Gelbukh, Jorge García Flores
Soft Comput.3
2016 Adam Kilgarriff's Legacy to Computational Linguistics and Beyond
Roger Evans, Alexander F. Gelbukh, Gregory Grefenstette, Patrick Hanks, Milos Jakubícek, Diana McCarthy, Martha Palmer, Ted Pedersen, Michael Rundell, Pavel Rychlý, Serge Sharoff, David Tugwell
CICLing (1)2
2016 Multiword Expressions (MWE) for Mizo Language: Literature Survey
Goutam Majumder, Partha Pakray, Zoramdinthara Khiangte, Alexander F. Gelbukh
CICLing (1)4
2016 Mathematical properties of soft cardinality: Enhancing Jaccard, Dice and cosine similarity measures with element-wise distance
Sergio Jiménez 0001, Fabio A. González 0001, Alexander F. Gelbukh
Inf. Sci.3
2016 Aspect extraction for opinion mining with a deep convolutional neural network
Soujanya Poria, Erik Cambria, Alexander F. Gelbukh
Knowl. Based Syst.3
2015 Modelling Public Sentiment in Twitter: Using Linguistic Patterns to Enhance Supervised Learning
Prerna Chikersal, Soujanya Poria, Erik Cambria, Alexander F. Gelbukh, Chng Eng Siong
CICLing (2)4
2015 Mining Parallel Resources for Machine Translation from Comparable Corpora
Santanu Pal, Partha Pakray, Alexander F. Gelbukh, Josef van Genabith
CICLing (1)3
2015 Deep Convolutional Neural Network Textual Features and Multiple Kernel Learning for Utterance-level Multimodal Sentiment Analysis
abstract
We present a novel way of extracting features from short texts, based on the activation values of an inner layer of a deep convolutional neural network.We use the extracted features in multimodal sentiment analysis of short video clips representing one sentence each.We use the combined feature vectors of textual, visual, and audio modalities to train a classifier based on multiple kernel learning, which is known to be good at heterogeneous data.We obtain 14% performance improvement over the state of the art and present a parallelizable decision-level data fusion method, which is much faster, though slightly less accurate.
Soujanya Poria, Erik Cambria, Alexander F. Gelbukh
EMNLP3
2015 Is the Most Frequent Sense of a Word Better Connected in a Semantic Network?
Hiram Calvo, Alexander F. Gelbukh
ICIC (3)2
2014 An IR-Based Strategy for Supporting Chinese-Portuguese Translation Services in Off-line Mode
Jordi Centelles, Marta R. Costa-jussà, Rafael E. Banchs, Alexander F. Gelbukh
CICLing (2)4
2014 Graph Ranking on Maximal Frequent Sequences for Single Extractive Text Summarization
Yulia Ledeneva, René Arnulfo García-Hernández, Alexander F. Gelbukh
CICLing (2)3
2014 Dependency-Based Semantic Parsing for Concept-Level Text Analysis
Soujanya Poria, Basant Agarwal, Alexander F. Gelbukh, Amir Hussain 0001, Newton Howard
CICLing (1)3
2014 Statistical Relational Learning to Recognise Textual Entailment
Miguel Ángel Ríos-Gaona, Lucia Specia, Alexander F. Gelbukh, Ruslan Mitkov
CICLing (1)3
2014 Syntactic N-grams as machine learning features for natural language processing
Grigori Sidorov, Francisco Velasquez, Efstathios Stamatatos, Alexander F. Gelbukh, Liliana Chanona-Hernández
Expert Syst. Appl.4
2014 EmoSenticSpace: A novel framework for affective common-sense reasoning
Soujanya Poria, Alexander F. Gelbukh, Erik Cambria, Amir Hussain 0001, Guang-Bin Huang
Knowl. Based Syst.2
2013 Link Analysis for Representing and Retrieving Legal Information
Alfredo Monroy, Hiram Calvo, Alexander F. Gelbukh, Georgina García Pacheco
CICLing (2)3
2013 Syntactic Dependency-Based N-grams: More Evidence of Usefulness in Classification
Grigori Sidorov, Francisco Velasquez, Efstathios Stamatatos, Alexander F. Gelbukh, Liliana Chanona-Hernández
CICLing (1)4
2012 Age-Related Temporal Phrases in Spanish and Italian
Sofía N. Galicia-Haro, Alexander F. Gelbukh
CICLing (1)2
2011 Dependency Syntax Analysis Using Grammar Induction and a Lexical Categories Precedence System
Hiram Calvo, Omar Juárez-Gambino, Alexander F. Gelbukh, Kentaro Inui
CICLing (1)3
2011 Answer Validation Using Textual Entailment
Partha Pakray, Alexander F. Gelbukh, Sivaji Bandyopadhyay
CICLing (2)2
2010 A Syntactic Textual Entailment System Based on Dependency Parser
Partha Pakray, Alexander F. Gelbukh, Sivaji Bandyopadhyay
CICLing2
2010 Automatic Term Extraction Using Log-Likelihood Based Comparison with General Reference Corpus
Alexander F. Gelbukh, Grigori Sidorov, Eduardo Lavin-Villa, Liliana Chanona-Hernández
NLDB1
2010 Text Comparison Using Soft Cardinality
Sergio Jiménez 0001, Fabio A. González 0001, Alexander F. Gelbukh
SPIRE3
2010 Unsupervised WSD by Finding the Predominant Sense Using Context as a Dynamic Thesaurus
Javier Tejada-Cárcamo, Hiram Calvo, Alexander F. Gelbukh, Kazuo Hara
J. Comput. Sci. Technol.3
2009 Incorporating Linguistic Information to Statistical Word-Level Alignment
Eduardo Cendejas, Grettel Barceló, Alexander F. Gelbukh, Grigori Sidorov
CIARP3
2009 Generalized Mongue-Elkan Method for Approximate Text String Comparison
Sergio Jiménez 0001, Claudia Jeanneth Becerra, Alexander F. Gelbukh, Fabio A. González 0001
CICLing3
2009 NLP for Shallow Question Answering of Legal Documents Using Graphs
Alfredo Monroy, Hiram Calvo, Alexander F. Gelbukh
CICLing3
2009 Adaptive evolution: an efficient heuristic for global optimization
abstract
This paper presents a novel evolutionary approach to solve numerical optimization problems, called Adaptive Evolution (AEv). AEv is a new micro-population-like technique because it uses small populations (less than 10 individuals). The two main mechanisms of AEv are elitism and adaptive behavior. It has an adaptive parameter to adjust the balance between global exploration, local exploitation and elitism. Its two crossover operators allow a newly-generated offspring to be parent of other offspring in the same generation. AEv requires the fine-tuning of two parameters (several state-of-the-art approaches use at least three). AEv is tested on a set of 10 benchmark functions with 30 decision variables and it is compared with respect to some state-of-the-art algorithms to show its competitive performance.
Francisco Viveros Jiménez, Efrén Mezura-Montes, Alexander F. Gelbukh
GECCO3
2009 Hybrid Algorithm for Word-Level Alignment of Parallel Texts
Eduardo Cendejas, Grettel Barceló, Alexander F. Gelbukh, Grigori Sidorov
NLDB3
2008 Various Criteria of Collocation Cohesion in Internet: Comparison of Resolving Power
Igor A. Bolshakov, Elena I. Bolshakova, Alexey P. Kotlyarov, Alexander F. Gelbukh
CICLing4
2008 Terms Derived from Frequent Sequences for Extractive Text Summarization
Yulia Ledeneva, Alexander F. Gelbukh, René Arnulfo García-Hernández
CICLing2
2008 Division of Spanish Words into Morphemes with a Genetic Algorithm
Alexander F. Gelbukh, Grigori Sidorov, Diego Lara-Reyes, Liliana Chanona-Hernández
NLDB1
2007 Distribution-Based Semantic Similarity of Nouns
Igor A. Bolshakov, Alexander F. Gelbukh
CIARP2
2007 Case-Sensitivity of Classifiers for WSD: Complex Systems Disambiguate Tough Words Better
Harri M. T. Saarikoski, Steve Legrand, Alexander F. Gelbukh
CICLing3
2007 Improving the Customization of Natural Language Interface to Databases Using an Ontology
M. Jose A. Zarate, Rodolfo A. Pazos Rangel, Alexander F. Gelbukh, Joaquín Pérez Ortega
ICCSA (1)3
2007 Two Methods of Evaluation of Semantic Similarity of Nouns Based on Their Modifier Sets
Igor A. Bolshakov, Alexander F. Gelbukh
NLDB2
2007 Lexical-Based Alignment for Reconstruction of Structure in Parallel Texts
Alexander F. Gelbukh, Grigori Sidorov, Liliana Chanona-Hernández
NLDB1
2006 Alignment of Paragraphs in Bilingual Texts Using Bilingual Dictionaries and Dynamic Programming
Alexander F. Gelbukh, Grigori Sidorov
CIARP1
2006 DILUCT: An Open-Source Spanish Dependency Parser Based on Rules, Heuristics, and Selectional Preferences
Hiram Calvo, Alexander F. Gelbukh
NLDB2
2006 Studying Evolution of a Branch of Knowledge by Constructing and Analyzing Its Ontology
Pavel Makagonov, Alejandro Ruiz Figueroa, Alexander F. Gelbukh
NLDB3
2005 Distributional Thesaurus Versus WordNet: A Comparison of Backoff Techniques for Unsupervised PP Attachment
Hiram Calvo, Alexander F. Gelbukh, Adam Kilgarriff
CICLing2
2005 Unsupervised Learning of P NP P Word Combinations
Sofía N. Galicia-Haro, Alexander F. Gelbukh
CICLing2
2005 Experiment on Combining Sources of Evidence for Passage Retrieval
Alexander F. Gelbukh, Namo Kang, Sang-Yong Han
CICLing1
2005 Natural Language Processing
abstract
Summary form only given. Natural language processing (NLP) is a major area of artificial intelligence research, which in its turn serves as a field of application and interaction of a number of other traditional AI areas. Until recently, the focus in AI applications in NLP was on knowledge representation, logical reasoning, and constraint satisfaction - first applied to semantics and later to the grammar. In the last decade, a dramatic shift in the NLP research has led to the prevalence of very large scale applications of statistical methods, such as machine learning and data mining. Naturally, this also opened the way to the learning and optimization methods that constitute the core of modern AI, most notably genetic algorithms and neural networks. In this paper we give an overview of the current trends in NLP and discuss the possible applications of traditional AI techniques and their combination in this fascinating area.
Alexander F. Gelbukh
HIS1
2005 An Approach to Clustering Abstracts
Mikhail Alexandrov, Alexander F. Gelbukh, Paolo Rosso
NLDB2
2005 On Some Optimization Heuristics for Lesk-Like WSD Algorithms
Alexander F. Gelbukh, Grigori Sidorov, Sang-Yong Han
NLDB1
2004 Unsupervised Learning of Ontology-Linked Selectional Preferences
Hiram Calvo, Alexander F. Gelbukh
CIARP2
2004 Detecting Inflection Patterns in Natural Language by Minimization of Morphological Model
Alexander F. Gelbukh, Mikhail Alexandrov, Sang-Yong Han
CIARP1
2004 Advanced Relevance Feedback Query Expansion Strategy for Information Retrieval in MEDLINE
Kwangcheol Shin, Sang-Yong Han, Alexander F. Gelbukh, Jaehwa Park
CIARP3
2004 Extracting Semantic Categories of Nouns for Syntactic Disambiguation from Human-Oriented Explanatory Dictionaries
Hiram Calvo, Alexander F. Gelbukh
CICLing2
2004 Automatic Syntactic Analysis for Detection of Word Combinations
Alexander F. Gelbukh, Grigori Sidorov, Sang-Yong Han, Erika Hernández-Rubio
CICLing1
2004 Synonymous Paraphrasing Using WordNet and Internet
Igor A. Bolshakov, Alexander F. Gelbukh
NLDB2
2004 Acquiring Selectional Preferences from Untagged Text for Prepositional Phrase Attachment Disambiguation
Hiram Calvo, Alexander F. Gelbukh
NLDB2
2004 Identification of Composite Named Entities in a Spanish Textual Database
Sofía N. Galicia-Haro, Alexander F. Gelbukh, Igor A. Bolshakov
NLDB2
2003 Improving Prepositional Phrase Attachment Disambiguation Using the Web as Corpus
Hiram Calvo, Alexander F. Gelbukh
CIARP2
2003 Approach to Construction of Automatic Morphological Analysis Systems for Inflective Languages with Little Effort
Alexander F. Gelbukh, Grigori Sidorov
CICLing1
2003 A Portable Natural Language Interface for Diverse Databases Using Ontologies
J. Antonio Zárate M., Rodolfo A. Pazos Rangel, Alexander F. Gelbukh, J. Isabel Padrón C.
CICLing3
2003 Tool for Computer-Aided Spanish Word Sense Disambiguation
Yoel Ledo Mezquita, Grigori Sidorov, Alexander F. Gelbukh
CICLing3
2003 On Detection of Malapropisms by Multistage Collocation Testing
Igor A. Bolshakov, Alexander F. Gelbukh
NLDB2
2003 Selection of Representative Documents for Clusters in a Document Collection
Alexander F. Gelbukh, Mikhail Alexandrov, Ales Bourek, Pavel Makagonov
NLDB1
2002 Quantitative Comparison of Homonymy in Spanish EuroWordNet and Traditional Dictionaries
Igor A. Bolshakov, Sofía N. Galicia-Haro, Alexander F. Gelbukh
CICLing3
2002 Automatic Selection of Defining Vocabulary in an Explanatory Dictionary
Alexander F. Gelbukh, Grigori Sidorov
CICLing1
2002 Compilation of a Spanish Representative Corpus
Alexander F. Gelbukh, Grigori Sidorov, Liliana Chanona-Hernández
CICLing1
2002 On Semantic Classification of Modifiers
Igor A. Bolshakov, Alexander F. Gelbukh
NLDB2
2001 Chi-Square Classifier for Document Categorization
Mikhail Alexandrov, Alexander F. Gelbukh, George Lozovoi
CICLing2
2001 Three Mechanisms of Parser Driving for Structure Disambiguation
Sofía N. Galicia-Haro, Alexander F. Gelbukh, Igor A. Bolshakov
CICLing2
2001 Zipf and Heaps Laws' Coefficients Depend on Language
Alexander F. Gelbukh, Grigori Sidorov
CICLing1
2001 Finding Correlative Associations among News Topics
Manuel Montes-y-Gómez, Aurelio López-López, Alexander F. Gelbukh
CICLing3
2001 A Statistical Approach to the Discovery of Ephemeral Associations among News Topics
Manuel Montes-y-Gómez, Alexander F. Gelbukh, Aurelio López-López
DEXA2
2001 Flexible Comparison of Conceptual GraphsWork done under partial support of CONACyT, CGEPI-IPN, and SNI, Mexico
Manuel Montes-y-Gómez, Alexander F. Gelbukh, Aurelio López-López, Ricardo Baeza-Yates
DEXA2
2001 Yet another application of inference in computational linguistics
abstract
Texts in natural languages consist of words that are syntactically linked and semantically combinable-like political party, pay attention, or brick wall. Such semantically plausible combinations of two content words, which we hereafter refer to as collocations, are important knowledge in many areas of computational linguistics. We consider a lexical resource that provides such knowledge-a collocation database (CBD). Since such databases cannot be complete under any reasonable compilation procedure, we consider heuristic-based inference mechanisms that predict new plausible collocations based on the ones present in the CDB, with the help of a WordNet-like thesaurus. If an available collocation combines the entries A and B, and B is 'similar' to C, then A and C supposedly constitute a collocation of the same category. Also, we touch upon semantically induced morphological categories suiting for such inferences. Several heuristics for filtering out wrong hypotheses are also given and the experience in inferences obtained with CrossLexica CDB is briefly discussed.
Igor A. Bolshakov, Alexander F. Gelbukh
SMC2
2001 Acquiring syntactic information for a government pattern dictionary from large text corpora
abstract
There are some research lines in automatic subcategorization frame acquisition and the importance of their work could not be doubted. However, almost all automatic work has been done in the constituent approach. Conversely, manual work is the traditional way for syntactic information acquisition in the dependency approach, which considers the correspondence between semantic valences and theirs syntactic realizations. The last approximation has some advantages for description of languages with relaxed word order constraints and a vast prepositional use. Our work is intended to compile automatically a government patterns dictionary in what syntactic information is referred to and to give a tool to facilitate linking of valences and meaning.
Sofía N. Galicia-Haro, Alexander F. Gelbukh, Igor A. Bolshakov
SMC2
2001 Combining dependency and constituent-based resources for structure disambiguation
abstract
Unrestricted text analysis requires an accurate syntactic analysis but structural ambiguity is one of the most difficult problems to resolve. Researchers have tried different approaches to obtain the correct syntactic structure from analyzed sentences but no successful results have been obtained. Two different approaches have traditionally applied to syntactic analysis: constituent grammars and dependency grammars. We propose a model for syntactic analysis and disambiguation combining lexical dependencies and semantic proximity. Lexical dependencies are applied by means of a government pattern dictionary following the dependency approach. The semantic proximity is introduced by means of semantic closeness among constituents. Examples are given to illustrate the method's contributions.
Sofía N. Galicia-Haro, Alexander F. Gelbukh, Igor A. Bolshakov
SMC2
2001 Text mining with conceptual graphs
abstract
A method for conceptual clustering of a collection of texts represented with conceptual graphs is presented. It uses an incremental strategy to construct the cluster hierarchy and incorporates some characteristics attractive for text mining purposes. For instance, it considers the structural information of the graphs, uses domain knowledge to detect the clusters with generalized descriptions, and uses a user-defined similarity measure between the graphs.
Manuel Montes-y-Gómez, Alexander F. Gelbukh, Aurelio López-López, Ricardo Baeza-Yates
SMC2
2001 Automatic detection of semantically primitive words using their reachability in an explanatory dictionary
abstract
We suggest the method that permits building a set of candidates to be considered semantic primitives from the standard explanatory dictionary. Our method is based on the frequencies of the words that are reachable in a semantic network constructed from the dictionary. The method implements word sense disambiguation techniques, network construction, and reachability analysis. In part of word sense disambiguation we use an improved Lesk's algorithm. In the part of analysis of reachability we show that the words to which our algorithm assigns high weight, are plausible candidates to be semantic primitives. It is also shown that better candidates to semantic primitives should be included in short vicious cycles, which is detected by our algorithm. We applied the method to a rather large Spanish explanatory dictionary.
Grigori Sidorov, Alexander F. Gelbukh
SMC2
2000 Lazy Query Enrichment: A Method for Indexing Large Specialized Document Bases with Morphology and Concept Hierarchy
Alexander F. Gelbukh
DEXA1
2000 Information Retrieval with Conceptual Graph Matching
Manuel Montes-y-Gómez, Aurelio López-López, Alexander F. Gelbukh
DEXA3
2000 A Very Large Database of Collocations and Semantic Links
Igor A. Bolshakov, Alexander F. Gelbukh
NLDB2