Carlos J. Pérez 0001

dblp:154/1294 · also C. J. Pérez 0001, Carlos J. Pérez Sánchez 0001, Carlos Javier Pérez Sánchez · DBLP profile ↗
← Back
25ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0001-6385-9080ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 A swarm-based multi-objective approach for sentiment-oriented generic summarization: application to tweets
abstract
Abstract Nowadays, automatic text summarization task is a matter that has acquired special relevance in numerous contexts. Particularly, sentiment analysis and opinion mining need summarization methods to quickly analyze public opinion about any event. In this way, the aim of the sentiment-oriented summarization approach is to produce a summary reflecting the sentiment of the authors’ opinions, covering the main content, and reducing the redundancy. In this work, a Sentiment-Oriented Dominance-based Bee Algorithm (SODBA) has been designed, developed, and applied for solving this problem. Experimentation has been conducted with datasets provided by Document Understanding Conferences. The evaluation of the results has been carried out by using the Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metrics and the Pearson correlation coefficient. The reported results have outperformed those obtained in the scientific literature in terms of ROUGE metrics. Moreover, SODBA has been applied to the tweets concerning the COVID-19 pandemic to obtain the summaries of the days with the most positive and the most negative sentiment.
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
Soft Comput.3
2025 Enhancing noise robustness of automatic Parkinson's disease detection in diadochokinesis tests using multicondition training
abstract
Despite significant advances on automatic detection of Parkinson’s disease (PD) based on speech, several open challenges still need to be addressed before a validated computer-aided diagnosis system can be used practically. One of these challenges lies in considering the potential corruption of speech caused by environmental noises, which may be nonstationary and exhibit varied characteristics. Speech features automatically extracted from diadochokinetic (DDK) tests have shown utility in assessing articulatory aspects of speech impairment in PD. The authors propose an automatic PD detection system based on a multicondition training (MCT) framework. The approach considers various types of realistic acoustic noise in addition to DDK recordings and uses machine learning for feature selection and classification. For each experiment, the noise addition process did not artificially increase the dataset size, as each subject’s recordings were either affected by a single noise type or had no injected noise. To compare with this MCT-based approach, an alternative method is examined where training involves speech samples affected by uniform noise conditions. This method, referred to as single-condition training (SCT), involves training with features either from the original waveforms or from waveforms altered by noise addition, ensuring uniformity by using the same type of realistic noise across all the training samples. The benefit of the MCT approach is demonstrated by showing the results obtained in classification tests to discriminate patients affected by PD from healthy individuals. The experiments performed were based on an in-house voice recording database composed of 30 individuals diagnosed with PD and 30 healthy controls. The speech samples were recorded using a smartphone as a data collection device so that the samples were not affected by speech compression algorithms. Both approaches (SCT and MCT) were tested against each specific type of noise under consideration. The mean accuracy rates showed improvements of 1.68%, 5.18%, and 4.39% for/pa/,/ta/, and/ka/ syllables, respectively, when using MCT compared with SCT. To the best of the authors’ knowledge, this is the first strategy published in the literature to deal with the potential corruption of speech by environmental noise in automatic PD detection aid systems based on DDK tests. • New automatic Parkinson’s disease detection system based on voice recordings. • New articulatory database based on diadochokinesis tests recorded on smartphones. • Multicondition training to improve noise robustness. • Performance comparison of multicondition training vs single-condition training. • Comparison of detection performance of/p/,/t/ and/k/ plosive consonants.
Mario Madruga, Yolanda Campos-Roca, Carlos J. Pérez 0001
Expert Syst. Appl.3
2025 A keyword extraction model study in the movie domain with synopsis and reviews
abstract
Abstract The use of keywords is increasingly being applied across diverse domains, including the movie industry, whose main platforms are adopting advanced natural language processing techniques. Algorithms for automatic extraction of keywords can provide relevant information in this domain. The most novel approaches covering several categories (statistics, graphs, word embedding, and hybrid) have been considered in a model study framework. They have been implemented, applied, and evaluated with standard datasets. In addition, a movie dataset with gold standard keywords, based on textual metadata from synopses and reviews, has been specifically developed for this scope. Keyword extraction models have been evaluated in terms of F-score and computation time. Furthermore, content analysis, both quantitative and qualitative, of the extracted keywords in the movie context has been performed. Results show a great variability in model performance and computation time among the different models. Qualitative results, in addition to F-score and computation time, demonstrate that keyword extraction works better with synopses than with reviews. The quantitative content analysis revealed that EmbedRank effectively reduces redundancy and limits the use of proper nouns, leading to high-quality keywords.
Carlos González-Santos, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001, Iñaki Martínez-Sarriegui, Joaquín M. López-Muñoz
Knowl. Inf. Syst.3
2025 Generating automatic summaries with an indicator and decomposition-based hybrid evolutionary approach
abstract
• IDHEA (Indicator and Decomposition-based Hybrid Evolutionary Algorithm) is proposed. • IDHEA is designed, implemented, and applied to solve automatic text summarization. • IDHEA’s performance is assessed by multi-objective and text summarization metrics. • IDHEA is compared with multi-objective algorithms and proposals from other authors. • IDHEA outperforms these other optimization algorithms from the scientific literature. The field of multi-objective optimization is experiencing a relevant growth due to its successful applications in numerous real-life problems. Two prominent trends are indicator-based and decomposition-based search strategies. However, the performance of these search strategies depends on the specific problem to solve. In the context of automatic summarization, hybridization of these techniques is an interesting and challenging proposal, which aims to improve the performance. For this reason, an Indicator and Decomposition-based Hybrid Evolutionary Algorithm (IDHEA) has been designed, implemented, and tested for addressing the automatic summarization problem. The proposed hybrid multi-objective approach has integrated the fundamentals of both indicator-based and decomposition-based techniques to solve this particular problem. Experimentation has been conducted using Document Understanding Conferences (DUC) dataset. The performance has been assessed by means of multi-objective metrics such as hypervolume, inverted generational distance, distance to ideal point, and set coverage, while summary quality has been evaluated with Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metrics. The proposed approach has outperformed standard algorithms in multi-objective evaluation, in addition to improving the existing results in the scientific literature in terms of summary quality.
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
Knowl. Based Syst.3
2024 Automatic assignment of microgenres to movies using a word embedding-based approach
abstract
Abstract Streaming services are increasingly leveraging Artificial Intelligence (AI) technologies for improved content cataloging, user experiences in content discovery, and personalization. A significant challenge in this domain is the automated assignment of microgenres to movies. This study introduces and evaluates approaches based on clustering, topic modeling, and word embedding to address this task. The evaluation employs a preprocessed dataset containing movie-related data—title tags, synopses, genres, and reviews—alongside a predefined microgenre list. Comparisons of three activation functions (binary step, ramp, and sigmoid) gauge their effectiveness in augmenting microgenre tags. Results demonstrate the superiority of the word embedding approach over clustering and topic modeling in terms of mean accuracy. Even more, the word embedding approach stands as the sole fully automated solution. Analysis indicates that incorporating review-based tags introduces noise and undermines accuracy. Besides, the word embedding approach yields optimal outcomes using the sigmoid function, effectively doubling assigned tags while maintaining matching quality. This sheds light on the potential of word embedding methods within the movie domain.
Carlos González-Santos, Miguel A. Vega-Rodríguez, Joaquín M. López-Muñoz, Iñaki Martínez-Sarriegui, Carlos J. Pérez 0001
Multim. Tools Appl.5
2023 A new multi-objective evolutionary algorithm for citation-based summarization: Comprehensive analysis of the generated summaries
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
Eng. Appl. Artif. Intell.3
2023 A multi-objective artificial bee colony approach for profit-aware recommender systems
José A. Concha-Carrasco, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
Inf. Sci.3
2023 Automatic assignment of moral foundations to movies by word embedding
abstract
Morality is a topic that people are increasingly concerned about. Morality is observed and measured during public acts or when developing and consuming products, such as movies. The Moral Foundations Theory (MFT) was developed to rigorously perform these measurements with the support of the Moral Foundations Dictionary (MFD). In this paper, a Word Embedding-based Moral Foundation Assignment (WEMFA) approach has been designed, implemented, and applied to the movie domain for multiple assignment of moral foundations. WEMFA may use any dictionary, and it has been applied to a movie collection generated from movie synopses. A comparison between WEMFA and MoralStrength, the only approach found in the scientific literature, has been carried out. The proposed approach provided a percentage improvement of 41.7% with respect to the best version of MoralStrength, which uses an extension of the original MFD almost 10 times larger in number of terms. In addition, an extension of the original MFD (MFD24) has been built by adding 14 new moral foundations to the 10 original ones, enriching the moral context. WEMFA provided a mean accuracy of 78% with MFD24 despite the increment of the number of moral foundations. Besides, new extended dictionaries or even totally different ones can be used with WEMFA, since it does not need any training.
Carlos González-Santos, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001, Joaquín M. López-Muñoz, Iñaki Martínez-Sarriegui
Knowl. Based Syst.3
2023 Automatic Update Summarization by a Multiobjective Number-One-Selection Genetic Approach
abstract
Currently, the explosive growth of the information available on the Internet makes automatic text summarization systems increasingly important. A particularly relevant challenge is the update summarization task. Update summarization differs from traditional summarization in its dynamic nature. While traditional summarization is static, that is, the document collections about a specific topic remain unchanged, update summarization addresses dynamic document collections based on a specific topic. Therefore, update summarization consists of summarizing the new document collection under the assumption that the user has already read a previous summarization and only the new information is interesting. The multiobjective number-one-selection genetic algorithm (MONOGA) has been designed and implemented to address this problem. The proposed algorithm produces a summary that is relevant to the user's given query, and it also contains updates information. Experiments were conducted on Text Analysis Conference (TAC) datasets, and Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metrics were considered to assess the model performance. The results obtained by the proposed approach outperform those from the existing approaches in the scientific literature, obtaining average percentage improvements between 12.74% and 55.03% in the ROUGE scores.
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
IEEE Trans. Cybern.3
2022 A multi-objective memetic algorithm for query-oriented text summarization: Medicine texts as a case study
abstract
Automatic text summarization is a topic of great interest in many fields of knowledge. Particularly, query-oriented extractive multi-document text summarization methods have increased their importance recently, since they can automatically generate a summary according to a query given by the user. One way to address this problem is by multi-objective optimization approaches. In this paper, a memetic algorithm, specifically a Multi-Objective Shuffled Frog-Leaping Algorithm (MOSFLA) has been developed, implemented, and applied to solve the query-oriented extractive multi-document text summarization problem. Experiments have been conducted with datasets from Text Analysis Conference (TAC), and the obtained results have been evaluated with Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metrics. The results have shown that the proposed approach has achieved important improvements with respect to the works of scientific literature. Specifically, 25.41%, 7.13%, and 30.22% of percentage improvements in ROUGE-1, ROUGE-2, and ROUGE-SU4 scores have been respectively reached. In addition, MOSFLA has been applied to medicine texts from the Topically Diverse Query Focus Summarization (TD-QFS) dataset as a case study.
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
Expert Syst. Appl.3
2021 Replication-based regularization approaches to diagnose Reinke's edema by using voice recordings
abstract
Reinke's edema is one of the most prevalent laryngeal pathologies. Its detection can be addressed by using computer-aided diagnosis systems based on features extracted from speech recordings. When extracting acoustic features from different voice recordings of a particular subject at a concrete moment, imperfections in technology and the very biological variability result in values that are close, but they are not identical. This suggests that the within-subject variability must be properly addressed in the statistical methodology. Regularization-based regression approaches can be used to reduce the classification errors by favoring the best predictors and penalizing the worst ones. Three replication-based regularization approaches for variable selection and classification have been specifically designed and implemented to take into account the underlying within-subject variability. In order to illustrate the applicability of these approaches, an experiment has been specifically conducted to discriminate Reinke's edema patients (30 subjects) from healthy people (30 subjects) in a hospital environment. The features have been extracted from four phonations of the sustained vowel /a/ recorded for each subject, leading to a database that has fed the proposed machine learning approaches. The proposed replication-based approaches have been proved to be reliable in terms of selected features and predictive ability, leading to a stable accuracy rate of 0.89 under a cross-validation framework. Also, a comparison with traditional independence-based regularization methods reports a great variability of the latter in terms of selected features and accuracy metrics. Therefore, the proposed approaches contribute to fill a gap in the scientific literature on statistical approaches considering within-subject variability and can be used to build a robust expert system.
Lizbeth Naranjo, Carlos J. Pérez 0001, Yolanda Campos-Roca, Mario Madruga
Artif. Intell. Medicine2
2021 The impact of term-weighting schemes and similarity measures on extractive multi-document text summarization
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
Expert Syst. Appl.3
2021 Addressing topic modeling with a multi-objective optimization approach based on swarm intelligence
Carlos González-Santos, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
Knowl. Based Syst.3
2020 Robustness Assessment of Automatic Reinke's Edema Diagnosis Systems
abstract
In the past few years there has been a great interest in computer aided diagnosis research. In the field of voice quality assessment, signal processing gives us tools to analyze and extract numeric characteristics describing the analyzed signal. These features might be used to tell an impaired voice from a healthy one for many different voice conditions, being Reinke's edema one of the most severe ones. Most studies have been carried out under strict laboratory conditions, making use of professional sound equipment and facilities, trying to minimize the influence of external conditions. However, real world situations are exposed to adverse acoustic environments. The goal of this paper is to build automatic detection systems for Reinke's edema based on a novel in-house dataset and, alternatively, on the Massachusetts Eye and Ear Infirmary Voice Disorders Database, and assess noise robustness in both cases.
Mario Madruga, Yolanda Campos-Roca, Carlos J. Pérez 0001
ICASSP3
2020 Experimental analysis of multiple criteria for extractive multi-document text summarization
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
Expert Syst. Appl.3
2019 Parallelizing a multi-objective optimization approach for extractive multi-document text summarization
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
J. Parallel Distributed Comput.3
2019 Comparison of automatic methods for reducing the Pareto front to a single solution applied to multi-document text summarization
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
Knowl. Based Syst.3
2018 Graphs and Key Players in an Educational Social Network
Fernando Calle-Alonso, Vicente Botón-Fernández, Dimas de la Fuente, Carlos J. Pérez 0001, Miguel A. Vega-Rodríguez, Daniel de la Mata Lara
CSEDU (2)4
2018 Word Clouds as a Learning Analytic Tool for the Cooperative e-Learning Platform NeuroK
Fernando Calle-Alonso, Vicente Botón-Fernández, Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001, Daniel de la Mata Lara
CSEDU (2)5
2018 Automatic selection of a single solution from the Pareto front to identify key players in social networks
Dimas de la Fuente, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
Knowl. Based Syst.3
2018 Extractive multi-document text summarization using a multi-objective artificial bee colony optimization approach
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
Knowl. Based Syst.3
2017 NeuroK: A Collaborative e-Learning Platform based on Pedagogical Principles from Neuroscience
Fernando Calle-Alonso, Agustín Cuenca-Guevara, Daniel de la Mata Lara, Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001
CSEDU (1)6
2016 Addressing voice recording replications for Parkinson's disease detection
Lizbeth Naranjo, Carlos J. Pérez 0001, Yolanda Campos-Roca, Jacinto Martín
Expert Syst. Appl.2
2012 Bayesian Supervised Image Classification based on a Pairwise Comparison Method
Fernando Calle-Alonso, José Pablo Arias-Nicolás, Carlos J. Pérez 0001, Jacinto Martín
ICPRAM (1)3
2009 Bayesian robustness for decision making problems: Applications in medical contexts
Jacinto Martín, Carlos J. Pérez 0001, Peter Müller 0003
Int. J. Approx. Reason.2