VLDB 2026 Research / reviewers in the wild / expert
Carlos J. Pérez 0001
dblp:154/1294 · also C. J. Pérez 0001, Carlos J. Pérez Sánchez 0001, Carlos Javier Pérez Sánchez
· DBLP profile ↗
25ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0001-6385-9080ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A swarm-based multi-objective approach for sentiment-oriented generic summarization: application to tweetsabstractAbstract Nowadays, automatic text summarization task is a matter that has acquired special relevance in numerous contexts. Particularly, sentiment analysis and opinion mining need summarization methods to quickly analyze public opinion about any event. In this way, the aim of the sentiment-oriented summarization approach is to produce a summary reflecting the sentiment of the authors’ opinions, covering the main content, and reducing the redundancy. In this work, a Sentiment-Oriented Dominance-based Bee Algorithm (SODBA) has been designed, developed, and applied for solving this problem. Experimentation has been conducted with datasets provided by Document Understanding Conferences. The evaluation of the results has been carried out by using the Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metrics and the Pearson correlation coefficient. The reported results have outperformed those obtained in the scientific literature in terms of ROUGE metrics. Moreover, SODBA has been applied to the tweets concerning the COVID-19 pandemic to obtain the summaries of the days with the most positive and the most negative sentiment. Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
Soft Comput. | 3 |
| 2025 | Enhancing noise robustness of automatic Parkinson's disease detection in diadochokinesis tests using multicondition trainingabstractDespite significant advances on automatic detection of Parkinson’s disease (PD) based on speech, several open challenges still need to be addressed before a validated computer-aided diagnosis system can be used practically. One of these challenges lies in considering the potential corruption of speech caused by environmental noises, which may be nonstationary and exhibit varied characteristics. Speech features automatically extracted from diadochokinetic (DDK) tests have shown utility in assessing articulatory aspects of speech impairment in PD. The authors propose an automatic PD detection system based on a multicondition training (MCT) framework. The approach considers various types of realistic acoustic noise in addition to DDK recordings and uses machine learning for feature selection and classification. For each experiment, the noise addition process did not artificially increase the dataset size, as each subject’s recordings were either affected by a single noise type or had no injected noise. To compare with this MCT-based approach, an alternative method is examined where training involves speech samples affected by uniform noise conditions. This method, referred to as single-condition training (SCT), involves training with features either from the original waveforms or from waveforms altered by noise addition, ensuring uniformity by using the same type of realistic noise across all the training samples. The benefit of the MCT approach is demonstrated by showing the results obtained in classification tests to discriminate patients affected by PD from healthy individuals. The experiments performed were based on an in-house voice recording database composed of 30 individuals diagnosed with PD and 30 healthy controls. The speech samples were recorded using a smartphone as a data collection device so that the samples were not affected by speech compression algorithms. Both approaches (SCT and MCT) were tested against each specific type of noise under consideration. The mean accuracy rates showed improvements of 1.68%, 5.18%, and 4.39% for/pa/,/ta/, and/ka/ syllables, respectively, when using MCT compared with SCT. To the best of the authors’ knowledge, this is the first strategy published in the literature to deal with the potential corruption of speech by environmental noise in automatic PD detection aid systems based on DDK tests. • New automatic Parkinson’s disease detection system based on voice recordings. • New articulatory database based on diadochokinesis tests recorded on smartphones. • Multicondition training to improve noise robustness. • Performance comparison of multicondition training vs single-condition training. • Comparison of detection performance of/p/,/t/ and/k/ plosive consonants. Mario Madruga, Yolanda Campos-Roca, Carlos J. Pérez 0001 |
Expert Syst. Appl. | 3 |
| 2025 | A keyword extraction model study in the movie domain with synopsis and reviewsabstractAbstract The use of keywords is increasingly being applied across diverse domains, including the movie industry, whose main platforms are adopting advanced natural language processing techniques. Algorithms for automatic extraction of keywords can provide relevant information in this domain. The most novel approaches covering several categories (statistics, graphs, word embedding, and hybrid) have been considered in a model study framework. They have been implemented, applied, and evaluated with standard datasets. In addition, a movie dataset with gold standard keywords, based on textual metadata from synopses and reviews, has been specifically developed for this scope. Keyword extraction models have been evaluated in terms of F-score and computation time. Furthermore, content analysis, both quantitative and qualitative, of the extracted keywords in the movie context has been performed. Results show a great variability in model performance and computation time among the different models. Qualitative results, in addition to F-score and computation time, demonstrate that keyword extraction works better with synopses than with reviews. The quantitative content analysis revealed that EmbedRank effectively reduces redundancy and limits the use of proper nouns, leading to high-quality keywords. Carlos González-Santos, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001, Iñaki Martínez-Sarriegui, Joaquín M. López-Muñoz |
Knowl. Inf. Syst. | 3 |
| 2025 | Generating automatic summaries with an indicator and decomposition-based hybrid evolutionary approachabstract• IDHEA (Indicator and Decomposition-based Hybrid Evolutionary Algorithm) is proposed. • IDHEA is designed, implemented, and applied to solve automatic text summarization. • IDHEA’s performance is assessed by multi-objective and text summarization metrics. • IDHEA is compared with multi-objective algorithms and proposals from other authors. • IDHEA outperforms these other optimization algorithms from the scientific literature. The field of multi-objective optimization is experiencing a relevant growth due to its successful applications in numerous real-life problems. Two prominent trends are indicator-based and decomposition-based search strategies. However, the performance of these search strategies depends on the specific problem to solve. In the context of automatic summarization, hybridization of these techniques is an interesting and challenging proposal, which aims to improve the performance. For this reason, an Indicator and Decomposition-based Hybrid Evolutionary Algorithm (IDHEA) has been designed, implemented, and tested for addressing the automatic summarization problem. The proposed hybrid multi-objective approach has integrated the fundamentals of both indicator-based and decomposition-based techniques to solve this particular problem. Experimentation has been conducted using Document Understanding Conferences (DUC) dataset. The performance has been assessed by means of multi-objective metrics such as hypervolume, inverted generational distance, distance to ideal point, and set coverage, while summary quality has been evaluated with Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metrics. The proposed approach has outperformed standard algorithms in multi-objective evaluation, in addition to improving the existing results in the scientific literature in terms of summary quality. Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
Knowl. Based Syst. | 3 |
| 2024 | Automatic assignment of microgenres to movies using a word embedding-based approachabstractAbstract Streaming services are increasingly leveraging Artificial Intelligence (AI) technologies for improved content cataloging, user experiences in content discovery, and personalization. A significant challenge in this domain is the automated assignment of microgenres to movies. This study introduces and evaluates approaches based on clustering, topic modeling, and word embedding to address this task. The evaluation employs a preprocessed dataset containing movie-related data—title tags, synopses, genres, and reviews—alongside a predefined microgenre list. Comparisons of three activation functions (binary step, ramp, and sigmoid) gauge their effectiveness in augmenting microgenre tags. Results demonstrate the superiority of the word embedding approach over clustering and topic modeling in terms of mean accuracy. Even more, the word embedding approach stands as the sole fully automated solution. Analysis indicates that incorporating review-based tags introduces noise and undermines accuracy. Besides, the word embedding approach yields optimal outcomes using the sigmoid function, effectively doubling assigned tags while maintaining matching quality. This sheds light on the potential of word embedding methods within the movie domain. Carlos González-Santos, Miguel A. Vega-Rodríguez, Joaquín M. López-Muñoz, Iñaki Martínez-Sarriegui, Carlos J. Pérez 0001 |
Multim. Tools Appl. | 5 |
| 2023 | A new multi-objective evolutionary algorithm for citation-based summarization: Comprehensive analysis of the generated summaries
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | A multi-objective artificial bee colony approach for profit-aware recommender systems
José A. Concha-Carrasco, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
Inf. Sci. | 3 |
| 2023 | Automatic assignment of moral foundations to movies by word embeddingabstractMorality is a topic that people are increasingly concerned about. Morality is observed and measured during public acts or when developing and consuming products, such as movies. The Moral Foundations Theory (MFT) was developed to rigorously perform these measurements with the support of the Moral Foundations Dictionary (MFD). In this paper, a Word Embedding-based Moral Foundation Assignment (WEMFA) approach has been designed, implemented, and applied to the movie domain for multiple assignment of moral foundations. WEMFA may use any dictionary, and it has been applied to a movie collection generated from movie synopses. A comparison between WEMFA and MoralStrength, the only approach found in the scientific literature, has been carried out. The proposed approach provided a percentage improvement of 41.7% with respect to the best version of MoralStrength, which uses an extension of the original MFD almost 10 times larger in number of terms. In addition, an extension of the original MFD (MFD24) has been built by adding 14 new moral foundations to the 10 original ones, enriching the moral context. WEMFA provided a mean accuracy of 78% with MFD24 despite the increment of the number of moral foundations. Besides, new extended dictionaries or even totally different ones can be used with WEMFA, since it does not need any training. Carlos González-Santos, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001, Joaquín M. López-Muñoz, Iñaki Martínez-Sarriegui |
Knowl. Based Syst. | 3 |
| 2023 | Automatic Update Summarization by a Multiobjective Number-One-Selection Genetic ApproachabstractCurrently, the explosive growth of the information available on the Internet makes automatic text summarization systems increasingly important. A particularly relevant challenge is the update summarization task. Update summarization differs from traditional summarization in its dynamic nature. While traditional summarization is static, that is, the document collections about a specific topic remain unchanged, update summarization addresses dynamic document collections based on a specific topic. Therefore, update summarization consists of summarizing the new document collection under the assumption that the user has already read a previous summarization and only the new information is interesting. The multiobjective number-one-selection genetic algorithm (MONOGA) has been designed and implemented to address this problem. The proposed algorithm produces a summary that is relevant to the user's given query, and it also contains updates information. Experiments were conducted on Text Analysis Conference (TAC) datasets, and Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metrics were considered to assess the model performance. The results obtained by the proposed approach outperform those from the existing approaches in the scientific literature, obtaining average percentage improvements between 12.74% and 55.03% in the ROUGE scores. Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | A multi-objective memetic algorithm for query-oriented text summarization: Medicine texts as a case studyabstractAutomatic text summarization is a topic of great interest in many fields of knowledge. Particularly, query-oriented extractive multi-document text summarization methods have increased their importance recently, since they can automatically generate a summary according to a query given by the user. One way to address this problem is by multi-objective optimization approaches. In this paper, a memetic algorithm, specifically a Multi-Objective Shuffled Frog-Leaping Algorithm (MOSFLA) has been developed, implemented, and applied to solve the query-oriented extractive multi-document text summarization problem. Experiments have been conducted with datasets from Text Analysis Conference (TAC), and the obtained results have been evaluated with Recall-Oriented Understudy for Gisting Evaluation (ROUGE) metrics. The results have shown that the proposed approach has achieved important improvements with respect to the works of scientific literature. Specifically, 25.41%, 7.13%, and 30.22% of percentage improvements in ROUGE-1, ROUGE-2, and ROUGE-SU4 scores have been respectively reached. In addition, MOSFLA has been applied to medicine texts from the Topically Diverse Query Focus Summarization (TD-QFS) dataset as a case study. Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
Expert Syst. Appl. | 3 |
| 2021 | Replication-based regularization approaches to diagnose Reinke's edema by using voice recordingsabstractReinke's edema is one of the most prevalent laryngeal pathologies. Its detection can be addressed by using computer-aided diagnosis systems based on features extracted from speech recordings. When extracting acoustic features from different voice recordings of a particular subject at a concrete moment, imperfections in technology and the very biological variability result in values that are close, but they are not identical. This suggests that the within-subject variability must be properly addressed in the statistical methodology. Regularization-based regression approaches can be used to reduce the classification errors by favoring the best predictors and penalizing the worst ones. Three replication-based regularization approaches for variable selection and classification have been specifically designed and implemented to take into account the underlying within-subject variability. In order to illustrate the applicability of these approaches, an experiment has been specifically conducted to discriminate Reinke's edema patients (30 subjects) from healthy people (30 subjects) in a hospital environment. The features have been extracted from four phonations of the sustained vowel /a/ recorded for each subject, leading to a database that has fed the proposed machine learning approaches. The proposed replication-based approaches have been proved to be reliable in terms of selected features and predictive ability, leading to a stable accuracy rate of 0.89 under a cross-validation framework. Also, a comparison with traditional independence-based regularization methods reports a great variability of the latter in terms of selected features and accuracy metrics. Therefore, the proposed approaches contribute to fill a gap in the scientific literature on statistical approaches considering within-subject variability and can be used to build a robust expert system. Lizbeth Naranjo, Carlos J. Pérez 0001, Yolanda Campos-Roca, Mario Madruga |
Artif. Intell. Medicine | 2 |
| 2021 | The impact of term-weighting schemes and similarity measures on extractive multi-document text summarization
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
Expert Syst. Appl. | 3 |
| 2021 | Addressing topic modeling with a multi-objective optimization approach based on swarm intelligence
Carlos González-Santos, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
Knowl. Based Syst. | 3 |
| 2020 | Robustness Assessment of Automatic Reinke's Edema Diagnosis SystemsabstractIn the past few years there has been a great interest in computer aided diagnosis research. In the field of voice quality assessment, signal processing gives us tools to analyze and extract numeric characteristics describing the analyzed signal. These features might be used to tell an impaired voice from a healthy one for many different voice conditions, being Reinke's edema one of the most severe ones. Most studies have been carried out under strict laboratory conditions, making use of professional sound equipment and facilities, trying to minimize the influence of external conditions. However, real world situations are exposed to adverse acoustic environments. The goal of this paper is to build automatic detection systems for Reinke's edema based on a novel in-house dataset and, alternatively, on the Massachusetts Eye and Ear Infirmary Voice Disorders Database, and assess noise robustness in both cases. Mario Madruga, Yolanda Campos-Roca, Carlos J. Pérez 0001 |
ICASSP | 3 |
| 2020 | Experimental analysis of multiple criteria for extractive multi-document text summarization
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
Expert Syst. Appl. | 3 |
| 2019 | Parallelizing a multi-objective optimization approach for extractive multi-document text summarization
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
J. Parallel Distributed Comput. | 3 |
| 2019 | Comparison of automatic methods for reducing the Pareto front to a single solution applied to multi-document text summarization
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
Knowl. Based Syst. | 3 |
| 2018 | Graphs and Key Players in an Educational Social Network
Fernando Calle-Alonso, Vicente Botón-Fernández, Dimas de la Fuente, Carlos J. Pérez 0001, Miguel A. Vega-Rodríguez, Daniel de la Mata Lara |
CSEDU (2) | 4 |
| 2018 | Word Clouds as a Learning Analytic Tool for the Cooperative e-Learning Platform NeuroK
Fernando Calle-Alonso, Vicente Botón-Fernández, Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001, Daniel de la Mata Lara |
CSEDU (2) | 5 |
| 2018 | Automatic selection of a single solution from the Pareto front to identify key players in social networks
Dimas de la Fuente, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
Knowl. Based Syst. | 3 |
| 2018 | Extractive multi-document text summarization using a multi-objective artificial bee colony optimization approach
Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
Knowl. Based Syst. | 3 |
| 2017 | NeuroK: A Collaborative e-Learning Platform based on Pedagogical Principles from Neuroscience
Fernando Calle-Alonso, Agustín Cuenca-Guevara, Daniel de la Mata Lara, Jesús M. Sánchez-Gómez, Miguel A. Vega-Rodríguez, Carlos J. Pérez 0001 |
CSEDU (1) | 6 |
| 2016 | Addressing voice recording replications for Parkinson's disease detection
Lizbeth Naranjo, Carlos J. Pérez 0001, Yolanda Campos-Roca, Jacinto Martín |
Expert Syst. Appl. | 2 |
| 2012 | Bayesian Supervised Image Classification based on a Pairwise Comparison Method
Fernando Calle-Alonso, José Pablo Arias-Nicolás, Carlos J. Pérez 0001, Jacinto Martín |
ICPRAM (1) | 3 |
| 2009 | Bayesian robustness for decision making problems: Applications in medical contexts
Jacinto Martín, Carlos J. Pérez 0001, Peter Müller 0003 |
Int. J. Approx. Reason. | 2 |