VLDB 2026 Research / reviewers in the wild / expert
Ernestina Menasalvas Ruiz
dblp:58/1631 · also Ernestina Menasalva Ruiz, Ernestina Menasalvas
· DBLP profile ↗
71ranked-venue papers
5as first author
17since 2021 · last 2025
0000-0002-5615-6798ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 24 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 23 · 2 first-author · 11 since 2021Theory of computation · 2 · 2 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ELADAIS: An Integrated Platform for High-Impact Clinical Data Extraction, Standardization and Advanced Analytics Using OMOP-CDMabstractClinical data generated in healthcare systems is increasingly recognized as a key resource for biomedical research, healthcare optimization, and population health monitoring. However, its full potential remains underexploited due to fragmentation, heterogeneity, and lack of interoperability between data sources. The ELADAIS project addresses this challenge by designing, developing, and deploying a scalable, modular, and interoperable technological platform for the extraction, transformation, storage, and advanced analysis of high-impact clinical data. Grounded in the OMOP Common Data Model (OMOP-CDM), ELADAIS integrates a microservice-based architecture, analytical environments, workflow orchestration, and federated data capabilities. The platform will be deployed at two major hospitals in Madrid, Spain, with the expectation of standardizing access to over 1 million patient records and more than 19 million clinical events. This paper presents the architectural principles and expectations of ELADAIS, highlighting its potential to accelerate reproducible and collaborative clinical research. Alejandro Rodríguez González, Víctor Robles, Juan José Cubillas Mercado, Juan Manuel Martínez Pérez, Jose Luis González Mendez, Ernestina Menasalvas Ruiz |
CBMS | 6 |
| 2025 | Comparison of ConvNeXt and Vision-Language Models for Breast Density Assessment in Screening MammographyabstractMammographic breast density classification is essential for cancer risk assessment but remains challenging due to subjective interpretation and inter-observer variability. This study compares multimodal and CNN-based methods for automated classification using the BI-RADS system, evaluating BioMedCLIP and ConvNeXt across three learning scenarios: zero-shot classification, linear probing with textual descriptions, and fine-tuning with numerical labels. Results show that zero-shot classification achieved modest performance, while the fine-tuned ConvNeXt model outperformed the BioMedCLIP linear probe. Although linear probing demonstrated potential with pretrained embeddings, it was less effective than full fine-tuning. These findings suggest that despite the promise of multimodal learning, CNN-based models with end-to-end fine-tuning provide stronger performance for specialized medical imaging. The study underscores the need for more detailed textual representations and domain-specific adaptations in future radiology applications. Yusdivia Molina-Román, David Gómez-Ortiz, Ernestina Menasalvas Ruiz, José G. Tamez-Peña, Alejandro Santos-Díaz |
CBMS | 3 |
| 2025 | GPT for medical entity recognition in SpanishabstractAbstract In recent years, there has been a remarkable surge in the development of Natural Language Processing (NLP) models, particularly in the realm of Named Entity Recognition (NER). Models such as BERT have demonstrated exceptional performance, leveraging annotated corpora for accurate entity identification. However, the question arises: Can newer Large Language Models (LLMs) like GPT be utilized without the need for extensive annotation, thereby enabling direct entity extraction? In this study, we explore this issue, comparing the efficacy of fine-tuning techniques with prompting methods to elucidate the potential of GPT in the identification of medical entities within Spanish electronic health records (EHR). This study utilized a dataset of Spanish EHRs related to breast cancer and implemented both a traditional NER method using BERT, and a contemporary approach that combines few shot learning and integration of external knowledge, driven by LLMs using GPT, to structure the data. The analysis involved a comprehensive pipeline that included these methods. Key performance metrics, such as precision, recall, and F-score, were used to evaluate the effectiveness of each method. This comparative approach aimed to highlight the strengths and limitations of each method in the context of structuring Spanish EHRs efficiently and accurately.The comparative analysis undertaken in this article demonstrates that both the traditional BERT-based NER method and the few-shot LLM-driven approach, augmented with external knowledge, provide comparable levels of precision in metrics such as precision, recall, and F score when applied to Spanish EHR. Contrary to expectations, the LLM-driven approach, which necessitates minimal data annotation, performs on par with BERT’s capability to discern complex medical terminologies and contextual nuances within the EHRs. The results of this study highlight a notable advance in the field of NER for Spanish EHRs, with the few shot approach driven by LLM, enhanced by external knowledge, slightly edging out the traditional BERT-based method in overall effectiveness. GPT’s superiority in F-score and its minimal reliance on extensive data annotation underscore its potential in medical data processing. Alvaro Garcia-Barragán, Alberto González Calatayud, Oswaldo Solarte Pabón, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles |
Multim. Tools Appl. | 5 |
| 2024 | Named Entity Recognition in Mammography Radiology Reports using a Multilingual Transfer Learning ApproachabstractThis study explores a multilingual transfer learning strategy for Named Entity Recognition (NER) in mammography radiology reports, aiming to improve breast cancer diagnosis. By utilizing a dataset from TecSalud, which includes mammograms and Electronic Health Records (EHRs) over ten years, this study seeks to address the linguistic barriers in medical documentation through advanced Natural Language Processing (NLP) models. Our approach involves meticulously labeling twenty-four distinct entities within the predominantly Spanish dataset, covering a range of diagnostic features and interpretive findings, highlighting the challenge of linguistic diversity in medical records and the potential of NLP to bridge this gap.The results demonstrate that fine-tuning on the last layer offers a balanced approach between simplicity and accuracy, avoiding overfitting and achieving state-of-art results. Esteban Ricardo Salazar Cabrera, Alejandro Santos-Díaz, Ernestina Menasalvas Ruiz, José G. Tamez-Peña, Víctor Robles |
CBMS | 3 |
| 2024 | Step-forward structuring disease phenotypic entities with LLMs for disease understandingabstractIn the rapidly evolving field of biomedical text mining, the extraction of phenotypic entities from unstructured texts remains a pivotal challenge. This paper introduces a novel method that leverage Large Language Models (LLMs) to extract phenotypical entities from freely available texts such as Wikipedia. Our approach goes beyond traditional Named Entity Recognition (NER) techniques by utilizing both local and cloud-based LLMs. We present a comprehensive comparison with state-of-the-art tools. Our study confirms the significant advantages of LLMs in identifying relevant phenotypic entities, thus enhancing the ability of researchers and clinicians to understand and respond to disease dynamics more effectively. Therefore, this work underscores the potential of next-generation LLMs to redefine the standards for the extraction of phenotypic entities in biomedical research. Alvaro Garcia-Barragán, Alberto González Calatayud, Lucía Prieto Santamaría, Víctor Robles, Ernestina Menasalvas Ruiz |
CBMS | 5 |
| 2023 | Structuring Breast Cancer Spanish Electronic Health Records Using Deep LearningabstractUsing Natural Language Processing (NLP) in the clinical domain has increased the possibility of automatically extracting information from oncology clinical narratives. Specifically, deep learning methods have been used to extract information in the cancer domain. However, most of the above proposals have concentrated only on extracting named entities from clinical narratives, but those proposals do not include a methodology for structuring the information after an information extraction step. In this paper, we propose an automatic pipeline based on deep learning for structuring breast cancer information from clinical narratives written in Spanish. The pipeline inputs a set of clinical documents written in narrative form and automatically generates a structured JSON file that contains the information for each patient. This pipeline integrates both clinical entity extraction and negation and uncertainty detection. Obtained results have shown that deep learning methods are feasible for structuring information in the breast cancer domain. Alvaro Garcia-Barragán, Oswaldo Solarte Pabón, Georgiy Nedostup, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles |
CBMS | 5 |
| 2023 | Clustering-based Pattern Discovery in Lung Cancer TreatmentsabstractLung cancer is the leading cause of cancer death. More than 238,340 new cases of lung cancer patients are expected in 2023, with an estimation of more than 127,070 deaths. Choosing the correct treatment is an important element to enhance the probability of survival and to improve patient's quality of life. Cancer treatments might provoke secondary effects. These toxicities cause different health problems that impact the patient's quality of life. Hence, reducing treatments toxicities while maintaining or improving their effectiveness is an important goal that aims to be pursued from the clinical perspective. On the other hand, clinical guidelines include general knowledge about cancer treatment recommendations to assist clinicians. Although they provide treatment recommendations based on cancer disease aspects and individual patient features, a statistical analysis taking into account treatment outcomes is not provided here. Therefore, the comparison between clinical guidelines with treatment patterns found in clinical data, would allow to validate the patterns found, as well as discovering alternative treatment patterns. In this work, we have analyzed a dataset containing lung cancer patients information including patients' data, prescribed treatments and their outcomes. Using a Chi-square test and K-Modes clustering algorithm in combination with Pattern Discovery metrics we identify patterns, within the clusters, based on cancer stage and treatment outcomes. Obtained results are analyzed based on statistical and clinical relevance and compared with lung cancer clinical guidelines. The comparison reveals that all patterns found coincide with clinical guidelines recommendations, assessing the validity of the proposed method for pattern discovery in a clinical dataset. Daniel Gómez-Bravo, Aaron García, Guillermo Vigueras, Belén Ríos-Sánchez, Alejandra Pérez-García, Vanessa Ospina, Maria Torrente, Ernestina Menasalvas Ruiz, Mariano Provencio, Alejandro Rodríguez González |
CBMS | 8 |
| 2023 | Transformers for extracting breast cancer information from Spanish clinical narrativesabstractThe wide adoption of electronic health records (EHRs) offers immense potential as a source of support for clinical research. However, previous studies focused on extracting only a limited set of medical concepts to support information extraction in the cancer domain for the Spanish language. Building on the success of deep learning for processing natural language texts, this paper proposes a transformer-based approach to extract named entities from breast cancer clinical notes written in Spanish and compares several language models. To facilitate this approach, a schema for annotating clinical notes with breast cancer concepts is presented, and a corpus for breast cancer is developed. Results indicate that both BERT-based and RoBERTa-based language models demonstrate competitive performance in clinical Named Entity Recognition (NER). Specifically, BETO and multilingual BERT achieve F-scores of 93.71% and 94.63%, respectively. Additionally, RoBERTa Biomedical attains an F-score of 95.01%, while RoBERTa BNE achieves an F-score of 94.54%. The findings suggest that transformers can feasibly extract information in the clinical domain in the Spanish language, with the use of models trained on biomedical texts contributing to enhanced results. The proposed approach takes advantage of transfer learning techniques by fine-tuning language models to automatically represent text features and avoiding the time-consuming feature engineering process. Oswaldo Solarte Pabón, Orlando Montenegro, Alvaro Garcia-Barragán, Maria Torrente, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles |
Artif. Intell. Medicine | 6 |
| 2022 | Subgroup Discovery Analysis of Treatment Patterns in Lung Cancer PatientsabstractLung cancer is the leading cause of cancer death. More than 236,740 new cases of lung cancer patients are expected in 2022, with an estimation of more than 130,180 deaths. Improving the survival rates or the patient's quality of life is partially covered by a common element: treatments. Cancer treatments are well known for the toxic outcomes and secondary effects on the patients. These toxicities cause different health problems that impact the patient's quality of life. Reducing toxicities without a decline on the positive survival effect is an important goal that aims to be pursued from the clinical perspective. On the other hand, clinical guidelines include general knowl-edge about cancer treatment recommendations to assist clinicians. Although they provide treatment recommendations based on cancer disease aspects and individual patient features, a statistical analysis taking into account treatment outcomes is not provided here. Therefore, the comparison between clinical guidelines with treatment patterns found in clinical data, would allow to validate the patterns found, as well as discovering alternative treatment patterns. In this work, we have analyzed a dataset containing lung cancer patients information including patients' data, prescribed treatments and outcomes obtained. Using a Subgroup Discovery method we identify patterns based on cancer stage while relying on treatment outcomes. Results are compared with clinical guide-lines and analyzed based on statistical and medical relevance using Subgroup Discovery metrics. Daniel Gómez-Bravo, Aaron García, Guillermo Vigueras, Belén Ríos-Sánchez, Belén Otero-Carrasco, Roberto Hernández López, Maria Torrente, Ernestina Menasalvas Ruiz, Mariano Provencio, Alejandro Rodríguez González |
CBMS | 8 |
| 2022 | REDIRECTION: Generating drug repurposing hypotheses using link prediction with DISNET dataabstractIn recent years and due to COVID-19 pandemic, drug repurposing or repositioning has been placed in the spotlight. Giving new therapeutic uses to already existing drugs, this discipline allows to streamline the drug discovery process, reducing the costs and risks inherent to de novo development. Computational approaches have gained momentum, and emerging techniques from the machine learning domain have proved themselves as highly exploitable means for repurposing prediction. Against this backdrop, one can find that biomedical data can be represented in terms of graphs, which allow depicting in a very expressive manner the underlying structure of the information. Combining these graph data structures with deep learning models enhances the prediction of new links, such as potential disease-drug connections. In this paper, we present a new model named REDIRECTION, which aims to predict new disease-drug links in the context of drug repurposing. It has been trained with a part of the DISNET biomedical graph, formed by diseases, symptoms, drugs, and their relationships. The reserved testing graph for the evaluation has yielded to an AUROC of 0.93 and an AUPRC of 0.90. We have performed a secondary validation of REDIRECTION using RepoDB data as the testing set, which has led to an AUROC of 0.87 and a AUPRC of 0.83. In the light of these results, we believe that REDIRECTION can be a meaningful and promising tool to generate drug repurposing hypotheses. Adrián Ayuso Muñoz, Esther Ugarte Carro, Lucía Prieto Santamaría, Belén Otero-Carrasco, Ernestina Menasalvas Ruiz, Yuliana Pérez-Gallardo, Alejandro Rodríguez González |
CBMS | 5 |
| 2022 | Drug repositioning with gender perspective focused on Adverse Drug ReactionsabstractDrug repositioning is a novel, useful, and crucial technique to find new uses for existing drugs. In this field of study, when the clinical trials necessary to obtain successful drug repositioning have been carried out, the female gender has not been given much consideration. Thus far, the participation of women in clinical trials has been very limited. There were several argued reasons to exclude them from trials, like the likelihood of pregnancy or sudden hormonal changes. This has meant that for a long time the adverse effects of a drug on women were unknown. Scientifically, it was known that due to the biological processes of pharmacokinetics and pharmacodynamics, the response to drugs was not the same in both genders, but despite this evidence, there is still no difference in the dosage or form of using a drug between men and women. In this study, we made a preliminary analysis where the main goal is to investigate gender differences within the drug repositioning field through the adverse effects produced by such treatments. A special section on specific cases of drug repositioning in rare diseases will also be considered to carry out the same verification previously mentioned in the text. Belén Otero-Carrasco, Aurora Pérez-Pérez, Ernestina Menasalvas Ruiz, Juan Pedro Valente, Lucía Prieto Santamaría, Alejandro Rodríguez González |
CBMS | 3 |
| 2022 | Deep learning to extract Breast Cancer diagnosis conceptsabstractThe wide adoption of electronic health records (EHRs) provides a potential source to support clinical research. The Bidirectional Encoder Representations from Transformers (BERT) has shown promising results in extracting information in the biomedical domain, including the cancer field. However, one of the challenges in the cancer domain is annotating resources to support information extraction. In this paper, we will show how models trained in a lung cancer corpus can be used to extract cancer concepts even in other cancer types. In particular, we will show the performance of BERT models on breast cancer data that was not used to train the models. Results are very promising as they show the possibility of applying deep learning-based models to predict cancer concepts in a different dataset to the one they were trained on, representing a considerable save of time and resources. Oswaldo Solarte Pabón, Maria Torrente, Alvaro Garcia-Barragán, Mariano Provencio, Ernestina Menasalvas Ruiz, Víctor Robles |
CBMS | 5 |
| 2022 | A Trusted Platform Module-based, Pre-emptive and Dynamic Asset Discovery ToolabstractThis paper presents an original Intelligent and Secure Asset Discovery Tool (ISADT) that uses artificial intelligence and TPM-based technologies to: (i) detect the network assets, and (ii) detect suspicious pattern in the use of the network. The architecture has specifically been designed to discover the assets of medium and large size companies and institutions, such as hospitals, universities, or government buildings. Given the distributed design of the architecture, it can cope with the problem of the isolation of different Virtual Local Area Networks (VLANs). This is done by collecting information from all the VLANs and storing it in a central node, which can be accessed by the network administrator, who may consult and visualize the status in any moment, or even by other authorized applications. The collected data is kept in a secure warehouse by the use of a Trusted Platform Module. Moreover, collected data is processed by the use of artificial intelligence in two ways: (i) the traffic of each network is analysed so that suspicious patterns can be detected, and (ii) identified ports and status are analysed to detect anomalous combinations of open ports in a device. Antonio Jesús Díaz-Honrubia, Alberto Blázquez-Herranz, Lucía Prieto Santamaría, Ernestina Menasalvas Ruiz, Alejandro Rodríguez González, Gustavo Gonzalez Granadillo, Emmanouil A. Panaousis, Christos Xenakis |
J. Inf. Secur. Appl. | 4 |
| 2021 | A Meta-Path-Based Prediction Method for Disease ComorbiditiesabstractThe simultaneous presence of diseases worsens the prognosis of patients and makes their treatment difficult. Identifying the co-occurrence of diseases is key to improving the situation of patients and designing effective therapeutic strategies. On the one hand, the increasing availability of clinical information opens new ways to unveil hidden relationships between diseases. On the other hand, heterogeneous information networks have been used in recent years to discover novel knowledge from disease data, including symptoms, genes or drugs. The use of meta-paths allows the complex semantics of the relationships between the different types of nodes to be included in heterogeneous networks. In this study, we propose a system to predict disease comorbidities through the use of meta-paths in a heterogeneous network of diseases and symptoms, built from textual sources of public access. The results obtained improve those of similar studies based on biological data, and the predictions calculated for diabetes and Crohn's disease are supported by medical literature. Both the used data and the obtained prediction model are publicly accessible. Eduardo P. García del Valle, Lucía Prieto Santamaría, Gerardo Lagunes García, Massimiliano Zanin, Ernestina Menasalvas Ruiz, Alejandro Rodríguez González |
CBMS | 5 |
| 2021 | Extracting Cancer Treatments from Clinical Text written in Spanish: A Deep Learning ApproachabstractExtracting accurate information about cancer patients' treatments is crucial to support clinical research, treatment planning, and to improve clinical care outcomes. However, treatment information resides in unstructured clinical text, making the task of data structuring especially challenging. Although several approaches have been proposed to extract treatments from clinical text, most of these proposals have focused on the English language. In this paper, we propose a deep learning-based approach to extract cancer treatments from clinical text written in Spanish. This approach uses a Bidirectional Long Short Memory (BiLSTM) neural net with a CRF layer to perform Named Entity Recognition. An annotated corpus from clinical text written about lung cancer patients is used to train the BiLSTM-based model. Performed tests have shown a performance of 90% in the F1-score, suggesting the feasibility of our approach to extract cancer treatments from clinical narratives. Oswaldo Solarte Pabón, Alberto Blázquez-Herranz, Maria Torrente, Alejandro Rodríguez González, Mariano Provencio, Ernestina Menasalvas Ruiz |
DSAA | 6 |
| 2021 | Towards Treatment Patterns Validation in Lung Cancer PatientsabstractLung cancer is the leading cause of cancer death. From the estimation of cases that will be in 2021, more than 230,000 new cases are expected to be of lung cancer patients, with an estimation of more than 131,000 deaths. Improving the survival rates or the patient's quality of life is partially covered by a common element: treatments. Collective knowledge about cancer treatment recommendations is typically included in clinical guidelines, intended to optimize patient care and assist clinicians in lung cancer treatment. These guidelines define a set of treatment paths, where recommendations depend on cancer disease aspects and individual features for a concrete patient. Although oncologists are expected to follow clinical guidelines, the inter and intrapatients' variability of response to the possible treatment combinations, makes it necessary to personalize different treatment-patterns on certain cases. Additionally, clinical guidelines are not frequently updated with new findings or lack a consistent methodology when they are frequently updated. For that reason, the analysis of patterns on both patients treated following the standard of care, or outside it, would allow to validate clinical guidelines and identify potential new treatment recommendations. In this work, we have analysed whether actual treatments prescribed to lung cancer patients follow clinical guidelines or not. Using a machine learning method that provides as output association rules (Apriori), we identify patterns based on cancer stage. These preliminary results show that treatments patterns found mostly match with clinical guidelines recommendations, validating the information included in the consulted guidelines. Arturo Redondo, Belén Ríos-Sánchez, Guillermo Vigueras, Belén Otero-Carrasco, Roberto Hernández López, Maria Torrente, Ernestina Menasalvas Ruiz, Mariano Provencio, Alejandro Rodríguez González |
DSAA | 7 |
| 2021 | Normal tissue content impact on the GBM molecular classificationabstractMolecular classification of glioblastoma has enabled a deeper understanding of the disease. The four-subtype model (including Proneural, Classical, Mesenchymal and Neural) has been replaced by a model that discards the Neural subtype, found to be associated with samples with a high content of normal tissue. These samples can be misclassified preventing biological and clinical insights into the different tumor subtypes from coming to light. In this work, we present a model that tackles both the molecular classification of samples and discrimination of those with a high content of normal cells. We performed a transcriptomic in silico analysis on glioblastoma (GBM) samples (n = 810) and tested different criteria to optimize the number of genes needed for molecular classification. We used gene expression of normal brain samples (n = 555) to design an additional gene signature to detect samples with a high normal tissue content. Microdissection samples of different structures within GBM (n = 122) have been used to validate the final model. Finally, the model was tested in a cohort of 43 patients and confirmed by histology. Based on the expression of 20 genes, our model is able to discriminate samples with a high content of normal tissue and to classify the remaining ones. We have shown that taking into consideration normal cells can prevent errors in the classification and the subsequent misinterpretation of the results. Moreover, considering only samples with a low content of normal cells, we found an association between the complexity of the samples and survival for the three molecular subtypes. Rodrigo Madurga, Noemí García-Romero, Beatriz Jiménez, Ana Collazo, Francisco Pérez-Rodríguez, Aurelio Hernández-Laín, Carlos Fernández-Carballal, Ricardo Prat-Acín, Massimiliano Zanin, Ernestina Menasalvas Ruiz, Ángel Ayuso-Sacido |
Briefings Bioinform. | 10 |
| 2020 | Creating a Metamodel Based on Machine Learning to Identify the Sentiment of Vaccine and Disease-Related Messages in Twitter: the MAVIS StudyabstractMAVIS was a project that aimed to study the interactions in social networks (Twitter and Instagram) between users regarding the sentiment expressed in their messages when they talked about specific vaccines or diseases. The study was performed during the period 2015-2018 and was initially technically done by using a set of commercial tools to identify the polarity of the messages. With the aim of improving the results provided by such tools, we performed a deep analysis of the results from such tools and provide a machine learning method as a metamodel over the results of the commercial tools. In this paper we explain both the technical process performed together with the main results that were obtained. Alejandro Rodríguez González, Juan Manuel Tuñas, Diego Fernandez Peces-Barba, Ernestina Menasalvas Ruiz, Almudena Jaramillo, Manuel Cotarelo, Antonio Conejo, Amalia Arce, Angel Gil |
CBMS | 4 |
| 2020 | Lung Cancer Diagnosis Extraction from Clinical Notes Written in SpanishabstractThe wide adoption of electronic health records (EHRs) offers a potential source to support research. Lung cancer is one of the most common cancer in the world. Although several tools have been developed to automatically extract concepts from oncology clinical notes, still there is a gap between concept extraction and concept understanding. The high number of clinical notes for the same patient, use of negation and proper date annotations lays in the root of the problem. In this paper, we propose an approach to accurate Lung cancer diagnosis extraction from clinical notes written in Spanish. The approach deals with a disambiguation process required to extract the correct date and diagnosis of a patient having hundreds of clinical notes and consequently hundreds of annotations. Results obtained on an annotated database of 1000 patients show an F-score of 90%. Oswaldo Solarte Pabón, Maria Torrente, Alejandro Rodríguez González, Mariano Provencio, Ernestina Menasalvas Ruiz, Juan Manuel Tuñas |
CBMS | 5 |
| 2020 | Analysis of New Nosological Models from Disease Similarities using ClusteringabstractWhile classical disease nosology is based on phenotypical characteristics, the increasing availability of biological and molecular data is providing new understanding of diseases and their underlying relationships, that could lead to a more comprehensive paradigm for modern medicine. In the present work, similarities between diseases are used to study the generation of new possible disease nosologic models that include both phenotypical and biological information. To this aim, disease similarity is measured in terms of disease feature vectors, that stood for genes, proteins, metabolic pathways and PPIs in the case of biological similarity, and for symptoms in the case of phenotypical similarity. An improvement in similarity computation is proposed, considering weighted instead of Booleans feature vectors. Unsupervised learning methods were applied to these data, specifically, density-based DBSCAN clustering algorithm. As evaluation metric silhouette coefficient was chosen, even though the number of clusters and the number of outliers were also considered. As a results validation, a comparison with randomly distributed data was performed. Results suggest that weighted biological similarities based on proteins, and computed according to cosine index, may provide a good starting point to rearrange disease taxonomy and nosology. Lucía Prieto Santamaría, Eduardo P. García del Valle, Gerardo Lagunes García, Massimiliano Zanin, Alejandro Rodríguez González, Ernestina Menasalvas Ruiz, Yuliana Pérez-Gallardo, Gandhi Hernández-Chan |
CBMS | 6 |
| 2020 | A Data-Driven Approach for Analyzing Healthcare Services Extracted from Clinical RecordsabstractCancer remains one of the major public health challenges worldwide. After cardiovascular diseases, cancer is one of the first causes of death and morbidity in Europe, with more than 4 million new cases and 1.9 million deaths per year. The suboptimal management of cancer patients during treatment and subsequent follows up are major obstacles in achieving better outcomes of the patients and especially regarding cost and quality of life In this paper, we present an initial data-driven approach to analyze the resources and services that are used more frequently by lung-cancer patients with the aim of identifying where the care process can be improved by paying a special attention on services before diagnosis to being able to identify possible lung-cancer patients before they are diagnosed and by reducing the length of stay in the hospital. Our approach has been built by analyzing the clinical notes of those oncological patients to extract this information and their relationships with other variables of the patient. Although the approach shown in this manuscript is very preliminary, it shows that quite interesting outcomes can be derived from further analysis. Manuel Scurti, Ernestina Menasalvas Ruiz, Maria-Esther Vidal, Maria Torrente, Dimitrios Vogiatzis, Georgios Paliouras, Mariano Provencio, Alejandro Rodríguez González |
CBMS | 2 |
| 2020 | Reconstructing the patient's natural history from electronic health records
Marjan Najafabadipour, Massimiliano Zanin, Alejandro Rodríguez González, Maria Torrente, Beatriz Nuñez, Juan Luis Cruz-Bermúdez, Mariano Provencio, Ernestina Menasalvas Ruiz |
Artif. Intell. Medicine | 8 |
| 2020 | How Wikipedia disease information evolve over time? An analysis of disease-based articles changes
Gerardo Lagunes García, Alejandro Rodríguez González, Lucía Prieto Santamaría, Eduardo P. García del Valle, Massimiliano Zanin, Ernestina Menasalvas Ruiz |
Inf. Process. Manag. | 6 |
| 2019 | Wikipedia Disease Articles: An Analysis of their Content and EvolutionabstractNowadays there is a huge amount of medical information that can be retrieved from different sources, both structured and unstructured. Internet has plenty of textual sources with medical knowledge (books, scientific papers, specialized web pages, etc.), but not all of them are publicly available. Wikipedia is a free, open and worldwide accessible source of knowledge. It contains more than 150,000 articles of medical content in the form of texts (non-structured information) that can be mined. The aim of this work is to study whether the evolution of information contained in Wikipedia medical articles can be used in a research context. The study has been focused on extracting the elements, from Wikipedia disease articles, that can be used to guide a diagnosis process, support the creation of diagnostic systems, or analyze the similarities between diseases, among others. Gerardo Lagunes García, Lucía Prieto Santamaría, Eduardo P. García del Valle, Massimiliano Zanin, Ernestina Menasalvas Ruiz, Alejandro Rodríguez González |
CBMS | 5 |
| 2019 | iASiS: Towards Heterogeneous Big Data Analysis for Personalized MedicineabstractThe vision of IASIS project is to turn the wave of big biomedical data heading our way into actionable knowledge for decision makers. This is achieved by integrating data from disparate sources, including genomics, electronic health records and bibliography, and applying advanced analytics methods to discover useful patterns. The goal is to turn large amounts of available data into actionable information to authorities for planning public health activities and policies. The integration and analysis of these heterogeneous sources of information will enable the best decisions to be made, allowing for diagnosis and treatment to be personalised to each individual. The project offers a common representation schema for the heterogeneous data sources. The iASiS infrastructure is able to convert clinical notes into usable data, combine them with genomic data, related bibliography, image data and more, and create a global knowledge base. This facilitates the use of intelligent methods in order to discover useful patterns across different resources. Using semantic integration of data gives the opportunity to generate information that is rich, auditable and reliable. This information can be used to provide better care, reduce errors and create more confidence in sharing data, thus providing more insights and opportunities. Data resources for two different disease categories are explored within the iASiS use cases, dementia and lung cancer. Anastasia Krithara, Fotis Aisopos, Vassiliki Rentoumi, Anastasios Nentidis, Konstantinos Bougiatiotis, Maria-Esther Vidal, Ernestina Menasalvas Ruiz, Alejandro Rodríguez González, Eleftherios Samaras, Peter Garrard, Maria Torrente, Mariano Provencio, Nikos Dimakopoulos, Rui Mauricio, Jordi Rambla De Argila, Gian Gaetano Tartaglia, Georgios Paliouras |
CBMS | 7 |
| 2019 | Recognition of Time Expressions in Spanish Electronic Health RecordsabstractThe widespread adoption of Electronic Health Records (EHRs) is generating an ever-increasing amount of unstructured clinical texts. Processing time expressions from these domain-specific-texts is crucial for the discovery of patterns that can help in the detection of medical events and building the patient's natural history. In medical domain, the recognition of time information from texts is challenging due to their lack of structure; usage of various formats, styles and abbreviations; their domain specific nature; writing quality; and the presence of ambiguous expressions. Furthermore, despite of Spanish occupying the second position in the world ranking of number of native speakers, to the best of our knowledge, no Natural Language Processing (NLP) tools have been introduced for the recognition of time expressions from clinical texts, written in this particular language. Therefore, in this paper, we propose a Temporal Tagger for identifying and normalizing time expressions appeared in Spanish clinical texts. We further compare our Temporal Tagger with the Spanish version of SUTime. By using a large dataset comprising EHRs of people suffering from lung cancer, we show that our developed Temporal Tagger, with an F1 score of 0.93, outperforms SUTime, with an F1 score of 0.797. Marjan Najafabadipour, Massimiliano Zanin, Alejandro Rodríguez González, Consuelo Gonzalo-Martín, Beatriz Nuñez, Virginia Calvo, Juan Luis Cruz-Bermúdez, Mariano Provencio, Ernestina Menasalvas Ruiz |
CBMS | 9 |
| 2019 | Radiomics Textural Features Extracted from Subcortical Structures of Grey Matter Probability for Alzheimers Disease DetectionabstractAlzheimer's disease (AD) is characterized by a progressive deterioration of cognitive and behavioral functions as a result of the atrophy of specific regions of the brain. It is estimated that by 2050 there will be 131.5 million people affected. Thus, there is an urgent need to find biological markers for its early detection and monitoring. In this work, it is present an analysis of textural radiomics features extracted from a gray matter probability volume, in a set of individual subcortical regions, from a number of different atlases, to identify subject with AD in a MRI. Also, significant subcortical regions for AD detection have been identified using a ReliefF relevance test. Experimental results using the ADNI1 database have proven the potential of some of the tested radiomic features as possible biomarkers for AD/CN differentiation. César Antonio Ortiz, Nuria Gutiérrez Sánchez, Consuelo Gonzalo-Martín, Roberto Garrido García, Alejandro Rodríguez González, Ernestina Menasalvas Ruiz |
CBMS | 6 |
| 2019 | Completing Missing MeSH Code Mappings in UMLS Through Alternative Expert-Curated SourcesabstractThe increasing availability of biological, clinical and literary sources enables the study of diseases from a more comprehensive approach. However, the interoperability of these sources, particularly of the codes used to identify diseases, poses a major challenge. Because of its role as a hub of multiple medical vocabularies, the Unified Medical Language System (UMLS) has become one of the most widely used resources for mapping diseases from different classification systems. The coverage of these mappings, nevertheless, is still insufficient, so researchers must resort to other methods to fill these gaps. In this article we analyze the limitations in UMLS mappings for the MeSH, ICD-10 and SNOMED CT vocabularies and propose the exploitation of alternative expert-curated sources to complete it. As a result, we demonstrate that this approach allows resolving more than 50% of the missing mappings for these vocabularies. All the findings are shared for validation and reuse. Eduardo P. García del Valle, Gerardo Lagunes García, Ernestina Menasalvas Ruiz, Lucía Prieto Santamaría, Massimiliano Zanin, Alejandro Rodríguez González |
CBMS | 3 |
| 2019 | A Proposal for Recognizing Skills in Data Science Using Open BadgesabstractThe growing demand for data scientists and their importance in industrial trends has resulted in a number of recent initiatives to promote skills in data science. One of these initiatives is a badge based scheme for the recognition of skills in data science developed as part of the BDVe project with the collaboration of the BDVA. This initiative is based upon a committee of experts, from both academia and industry, which will define and maintain a list of skills required to play key roles in the data science ecosystem. Ernestina Menasalvas Ruiz, Ana María Moreno 0001, Nik Swoboda |
ITiCSE | 1 |
| 2019 | Disease networks and their contribution to disease understanding: A review of their evolution, techniques and data sourcesabstractOver a decade ago, a new discipline called network medicine emerged as an approach to understand human diseases from a network theory point-of-view. Disease networks proved to be an intuitive and powerful way to reveal hidden connections among apparently unconnected biomedical entities such as diseases, physiological processes, signaling pathways, and genes. One of the fields that has benefited most from this improvement is the identification of new opportunities for the use of old drugs, known as drug repurposing. The importance of drug repurposing lies in the high costs and the prolonged time from target selection to regulatory approval of traditional drug development. In this document we analyze the evolution of disease network concept during the last decade and apply a data science pipeline approach to evaluate their functional units. As a result of this analysis, we obtain a list of the most commonly used functional units and the challenges that remain to be solved. This information can be very valuable for the generation of new prediction models based on disease networks. Eduardo P. García del Valle, Gerardo Lagunes García, Lucía Prieto Santamaría, Massimiliano Zanin, Ernestina Menasalvas Ruiz, Alejandro Rodríguez González |
J. Biomed. Informatics | 5 |
| 2018 | Evaluating Wikipedia as a Source of Information for Disease UnderstandingabstractThe increasing availability of biological data is improving our understanding of diseases and providing new insight into their underlying relationships. Thanks to the improvements on both text mining techniques and computational capacity, the combination of biological data with semantic information obtained from medical publications has proven to be a very promising path. However, the limitations in the access to these data and their lack of structure pose challenges to this approach. In this document we propose the use of Wikipedia - the free online encyclopedia - as a source of accessible textual information for disease understanding research. To check its validity, we compare its performance in the determination of relationships between diseases with that of PubMed, one of the most consulted data sources of medical texts. The obtained results suggest that the information extracted from Wikipedia is as relevant as that obtained from PubMed abstracts (i.e. the free access portion of its articles), although further research is proposed to verify its reliability for medical studies. Eduardo P. García del Valle, Gerardo Lagunes García, Lucía Prieto Santamaría, Massimiliano Zanin, Ernestina Menasalvas Ruiz, Alejandro Rodríguez González |
CBMS | 5 |
| 2018 | À-trous wavelet transform-based hybrid image fusion for face recognition using region classifiersabstractAbstract This paper presents a new hybrid fusion framework based on thermal and visible face images. Fusion of information is done here in two phases, first at the pixel level and then at the decision level. For the pixel level fusion process, à‐trous wavelet transform is applied on both the thermal and visible face images. In decision level fusion, 34 region classifiers, each concentrating on a specified region of the face image, are tested individually for their ability to identify a person from the face image. The region classifiers, which contribute significantly in recognizing the face image, are considered for decision level fusion using majority voting. All experiments have been conducted on the UGC‐JU face database and IRIS benchmark face database. The maximum recognition rate is about 97.22% for both the databases whereas decisions of 17 region classifiers among 34 are considered. Experimental results and comparative study show that the proposed fusion method provides a framework for recognition of face images in uncontrolled environments such as variations in illumination conditions, pose, and facial expressions. Ayan Seal, Debotosh Bhattacharjee, Mita Nasipuri, Consuelo Gonzalo-Martín, Ernestina Menasalvas Ruiz |
Expert Syst. J. Knowl. Eng. | 5 |
| 2017 | OncoCall: Analyzing the Outcomes of the Oncology Telephone Patient AssistanceabstractHospital Puerta del Hierro in Madrid, Spain, implemented in November 2011 a new service that aim to aid the patients of the oncology service with their doubts during their treatments through the use of a centralized call center. This service was created with the goal of provide a more personalized patient attention as well as to try to reduce the number of re-entries in the hospital in the emergencies. The aim of this paper is to present the main result of the analysis of the data produced by their call service in order to verify if the objectives were fulfilled as well as to gather what improvements can be done. Ernestina Menasalvas Ruiz, Consuelo Gonzalo-Martín, Juan Manuel Tuñas, Alejandro Rodríguez González, Mariano Provencio, Cristina Gonzalez de Pedro, Marta Mendez, Olga Zaretskaia, Juan Luis Cruz, Jesús Rey, Consuelo Parejo |
CBMS | 1 |
| 2017 | Fusion of Visible and Thermal Images Using a Directed Search Method for Face RecognitionabstractA new image fusion algorithm based on the visible and thermal images for face recognition is presented in this paper. The new fusion algorithm derives the benefit from both the modalities images. The proposed fusion process is the weighted sum of thermal and visible face information with two weighting factors [Formula: see text] and [Formula: see text], respectively. The weighting factors are calculated using a directed search algorithm automatically. The proposed fusion framework is evaluated through extensive experiments using UGC-JU face database. Experiments are of three fold. Firstly, individual modalities images are used separately for human face recognition. Secondly, fused face images using the proposed method are used for recognition purpose. The highest level of accuracy achieved by using the proposed method is about 98.42%. Lastly, the three existing fusion methods are applied on the same face database for comparison with the results of the proposed method. All the results demonstrate significant performance improvements in recognition over individual modalities and some of the existing fusion approaches, suggesting that fusion is a viable approach that deserves further study and consideration. Ayan Seal, Debotosh Bhattacharjee, Mita Nasipuri, Consuelo Gonzalo-Martín, Ernestina Menasalvas Ruiz |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2016 | Predicting recurring concepts on data-streams by means of a meta-model and a fuzzy similarity function
Miguel Ángel Abad, João Bártolo Gomes, Ernestina Menasalvas Ruiz |
Expert Syst. Appl. | 3 |
| 2016 | Local optimal scale in a hierarchical segmentation method for satellite images - An OBIA approach for the agricultural landscape
Consuelo Gonzalo-Martín, Mario Lillo-Saavedra, Ernestina Menasalvas Ruiz, David Fonseca 0002, Angel García-Pedrero, Roberto Costumero |
J. Intell. Inf. Syst. | 3 |
| 2015 | Recurrent drifts: Aapplying fuzzy logic to concept similarity functionabstractRecurrent drift, as a specific type of concept drift, is characterised by the appearance of previously seen concepts.Therefore, in those cases the learning process could be saved or at least minimized by applying an already trained classification model.In this paper we propose Fuzzy-Rec, a framework that is able to deal with recurrent concept drifts by means of a repository of classification models and a similarity function.Fuzzy logic is used in the framework to implement the similarity function needed to compare different classification models.This is a crucial aspect when dealing with drift recurrence, as long as some measure must be implemented to determine which model better fits a previously seen context.As it can be seen in the experimentation results of this paper, this fuzzy similarity function provides excellent results both in synthetic and real datasets.As a conclusion, we can state that the introduction of fuzzy logic comparisons between models could lead to a better efficient reuse of previously seen concepts, saving computational resources by applying not just equal models, but also similar ones. Miguel Ángel Abad, Ernestina Menasalvas Ruiz |
FedCSIS | 2 |
| 2015 | Medical Mining: KDD 2015 TutorialabstractIn year 2015, we experience a proliferation of scientific publications, conferences and funding programs on KDD for medicine and healthcare. However, medical scholars and practitioners work differently from KDD researchers: their research is mostly hypothesis-driven, not data-driven. KDD researchers need to understand how medical researchers and practitioners work, what questions they have and what methods they use, and how mining methods can fit into their research frame and their everyday business. Purpose of this tutorial is to contribute to this learning process. We address medicine and healthcare; there the expertise of KDD scholars is needed and familiarity with medical research basics is a prerequisite. We aim to provide basics for (1) mining in epidemiology and (2) mining in the hospital. We also address, to a lesser extent, the subject of (3) preparing and annotating Electronic Health Records for mining. Myra Spiliopoulou, Pedro Pereira Rodrigues, Ernestina Menasalvas Ruiz |
KDD | 3 |
| 2014 | Recurring concept detection for spam filtering
Miguel Ángel Abad, João Bártolo Gomes, Ernestina Menasalvas Ruiz |
FUSION | 3 |
| 2014 | A methodology to compare Dimensionality Reduction algorithms in terms of loss of quality
Antonio Gracia Berná, Santiago González, Víctor Robles, Ernestina Menasalvas Ruiz |
Inf. Sci. | 4 |
| 2014 | Mining Recurring Concepts in a Dynamic Feature SpaceabstractMost data stream classification techniques assume that the underlying feature space is static. However, in real-world applications the set of features and their relevance to the target concept may change over time. In addition, when the underlying concepts reappear, reusing previously learnt models can enhance the learning process in terms of accuracy and processing time at the expense of manageable memory consumption. In this paper, we propose mining recurring concepts in a dynamic feature space (MReC-DFS), a data stream classification system to address the challenges of learning recurring concepts in a dynamic feature space while simultaneously reducing the memory cost associated with storing past models. MReC-DFS is able to detect and adapt to concept changes using the performance of the learning process and contextual information. To handle recurring concepts, stored models are combined in a dynamically weighted ensemble. Incremental feature selection is performed to reduce the combined feature space. This contribution allows MReC-DFS to store only the features most relevant to the learnt concepts, which in turn increases the memory efficiency of the technique. In addition, an incremental feature selection method is proposed that dynamically determines the threshold between relevant and irrelevant features. Experimental results demonstrating the high accuracy of MReC-DFS compared with state-of-the-art techniques on a variety of real datasets are presented. The results also show the superior memory efficiency of MReC-DFS. João Bártolo Gomes, Mohamed Medhat Gaber, Pedro A. C. Sousa, Ernestina Menasalvas Ruiz |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2013 | Multi-agent location system in wireless networks
Luis Mengual, Oscar Marbán, Santiago Eibe, Ernestina Menasalvas Ruiz |
Expert Syst. Appl. | 4 |
| 2013 | Self-configuring data mining for ubiquitous computing
Aysegul Cayci, Ernestina Menasalvas Ruiz, Yücel Saygin, Santiago Eibe |
Inf. Sci. | 2 |
| 2012 | Mobile Activity Recognition Using Ubiquitous Data Stream Mining
João Bártolo Gomes, Shonali Krishnaswamy, Mohamed Medhat Gaber, Pedro A. C. Sousa, Ernestina Menasalvas Ruiz |
DaWaK | 5 |
| 2012 | Framework for the Establishment of Resource-Aware Data Mining Techniques on Critical Infrastructures
Miguel Ángel Abad, Ernestina Menasalvas Ruiz |
IPMU (2) | 2 |
| 2012 | MARS: A Personalised Mobile Activity Recognition SystemabstractMobile activity recognition focuses on inferring the current activities of a mobile user by leveraging the sensory data that is available on today's smart phones. The state of the art in mobile activity recognition uses traditional classification techniques. Thus, the learning process typically involves: i) collection of labelled sensory data that is transferred and collated in a centralised repository, ii) model building where the classification model is trained and tested using the collected data, iii) a model deployment stage where the learnt model is deployed on-board a mobile device for identifying activities based on new sensory data. In this paper, we demonstrate the Mobile Activity Recognition System (MARS) where for the first time the model is built and continuously updated on-board the mobile device itself using data stream mining. The advantages of the on-board approach are that it allows model personalisation and increased privacy as the data is not sent to any external site. Furthermore, when the user or its activity profile changes MARS enables quick model adaptation. One of the stand out features of MARS is that training/updating the model takes less than 30 seconds per activity. MARS has been implemented on the Android platform to demonstrate that it can achieve accurate mobile activity recognition. Moreover, we can show in practice that MARS quickly adapts to user profile changes while at the same time being scalable and efficient in terms of consumption of the device resources. João Bártolo Gomes, Shonali Krishnaswamy, Mohamed Medhat Gaber, Pedro A. C. Sousa, Ernestina Menasalvas Ruiz |
MDM | 5 |
| 2012 | A management Ad Hoc networks model for rescue and emergency scenarios
Rommel Torres Tandazo, Luis Mengual, Oscar Marbán, Santiago Eibe, Ernestina Menasalvas Ruiz, Byron Maza |
Expert Syst. Appl. | 5 |
| 2012 | SOMAR: A SOcial Mobile Activity Recommender
Andrea Zanda, Santiago Eibe, Ernestina Menasalvas Ruiz |
Expert Syst. Appl. | 3 |
| 2012 | Tracking recurrent concepts using contextabstractThe problem of recurring concepts in data stream classification is a special case of concept drift where concepts may reappear. Although several existing methods are able to learn in the presence of concept drift, few consider contextual information João Bártolo Gomes, Pedro A. C. Sousa, Ernestina Menasalvas Ruiz |
Intell. Data Anal. | 3 |
| 2011 | Context-Aware Collaborative Data Stream Mining in Ubiquitous Devices
João Bártolo Gomes, Mohamed Medhat Gaber, Pedro A. C. Sousa, Ernestina Menasalvas Ruiz |
IDA | 4 |
| 2011 | A social network activity recommender system for ubiquitous devicesabstractThe increasing demand to access social networks by mobile devices together with increasing computation power of these mobile devices motivate the need of local recommendation services for social network users. Social networking is generating an incredible amount of information that is sometimes difficult for users to process, especially from mobile phones. Several links, activities, and recommendations are proposed by networked friends every hour, which together are nearly impossible to manage. There is a need to filter and make accessible such information to users, which is the motivation behind developing a mobile recommender that exploits social network information. Thus, in this paper, we propose the design and the implementation of a SOcial Mobile Activity Recommender (SOMAR) that can integrate Facebook social network mobile data and sensor data to propose activities to the user. The recommendations are completely calculated in situ in the mobile device with an embedded data mining component that is the basis to compute a social graph of the user relationships that will be later used for the recommendation process. The paper also presents some experiments that analyze the performance of the proposed method. Andrea Zanda, Ernestina Menasalvas Ruiz, Santiago Eibe |
ISDA | 2 |
| 2010 | Situation-Aware Data Stream Mining Service for Ubiquitous ApplicationsabstractAdvances in data mining, particularly in anytime anywhere data stream mining, make on-board data analysis possible in mobile devices with resource constraints. In this work, we propose a data stream mining service to support knowledge discovery in ubiquitous applications while addressing resource constraints on mobile devices. As the basis for our service we describe a general mechanism, which autonomously adapts the execution of the data stream mining process to each situation, using context and resource awareness. We describe the main components to achieve adaptability and propose a decision mechanism based on machine learning to support the configuration selection task, as we consider this to be a key element to achieve autonomy and adaptation of the mining service. We then present an instantiation of the proposed approach for the particular case of classification using the VFDT algorithm and analyze which factors influence it. Experimental results show how the adaptable data stream mining service improves resource consumption while increasing the quality of the anytime mining model. João Bártolo Gomes, Ernestina Menasalvas Ruiz, Pedro A. C. Sousa |
Mobile Data Management | 2 |
| 2010 | Adapting Batch Learning Algorithms Execution in Ubiquitous DevicesabstractIn order to provide context aware, adaptive, and anticipatory services, data mining services are required to provide them with intelligence. The data mining could be either executed in a central server or locally. In either case, adaptability to the changing environment is required. In the stream mining scenario, some solutions have been proposed to provide mechanisms to adapt the execution to available resources and context. Here, we propose a cost model mechanism to adapt the algorithm execution according to available resources and context information for the case of static data. The mechanism based on analyzing efficacy and efficiency (EE-Model) of the algorithm, is a two step process in which first the efficiency and efficacy of the algorithm are calculated for predefined algorithm configurations and dataset input. In a second step, taking into account the available resources and context, the best configuration of the algorithm is chosen. The paper describes the mechanism and presents an EE-Model instantiation for C4.5 algorithm. Further, we demonstrate the convenience of the proposed approach with a simulation of synthetic data. Andrea Zanda, Santiago Eibe, Ernestina Menasalvas Ruiz |
Mobile Data Management | 3 |
| 2010 | Density-based semi-supervised clustering
Myra Spiliopoulou, Ernestina Menasalvas Ruiz |
Data Min. Knowl. Discov. | 3 |
| 2009 | C-DenStream: Using Domain Knowledge on a Data Stream
Ernestina Menasalvas Ruiz, Myra Spiliopoulou |
Discovery Science | 2 |
| 2009 | Emerging User Intentions: Matching User Queries with Topic Evolution in News Text StreamsabstractTrend detection analysis from unstructured data poses a huge challenge to current advanced, web-enabled knowledge-based systems (KBS). Consolidated studies in topic and trend detection from text streams have concentrated so far mainly on identifying and visualizing dynamically evolving text patterns. From the knowledge modeling perspective identifying and defining new, relevant features that are able to synchronize the emergent user intentions to the dynamicity of the system's structure is a need. Additionally the advanced KBS have to remain highly sensitive to the content change, marked by evolution of trends in topics extracted from text streams. In this paper, we are describing a three-layered approach called the "user-system-content method" that is helping us to identify the most relevant knowledge mapping features derived from the USER, SYSTEM and CONTENT perspectives into an overall "context model", that will enable the advanced KBS to automatically streamline the query enrichment process in a much more user-centered, dynamical and flexible way. After a general introduction to our three-layered approach, we will describe into detail the necessary process steps for the implementation of our method and will present a case study for its integration on a real multimedia web-content portal using news streams as major source of unstructured information. Maria Valencia, Codrina Lauth, Ernestina Menasalvas Ruiz |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 3 |
| 2009 | Toward data mining engineering: A software engineering approach
Oscar Marbán, Javier Segovia, Ernestina Menasalvas Ruiz, Covadonga Fernández-Baizán |
Inf. Syst. | 3 |
| 2008 | A cost model to estimate the effort of data mining projects (DMCoMo)
Oscar Marbán, Ernestina Menasalvas Ruiz, María C. Fernández-Baizán |
Inf. Syst. | 2 |
| 2007 | An Engineering Approach to Data Mining Projects
Oscar Marbán, Gonzalo Mariscal, Ernestina Menasalvas Ruiz, Javier Segovia |
IDEAL | 3 |
| 2005 | Using Genetic Algorithms to Improve Accuracy of Economical Indexes Prediction
Óscar Cubo, Víctor Robles, Javier Segovia, Ernestina Menasalvas Ruiz |
IDA | 4 |
| 2004 | A Report of Activities at the WIC-Spain Research CentreabstractThe WIC(Web Intelligence Consortium)-Spain Research Centre is an interdisciplinary research initiative which brings together people with the aim of covering all the aspects related to the discovery of hidden knowledge in big volumes of data as long as exploring the practical impacts of advanced Information Technology (IT) and Artificial Intelligence (AI) on this volume of information The group consist of seven researchers and five Ph.D. students led by Dr. Javier Segovia and coordinated by Dr. Ernestina Menasalvas. We are funded from a variety of sources including the Spanish Ministry of Science and Technology as well as national and international industry. Pilar Herrero, María S. Pérez 0001, Ernestina Menasalvas Ruiz, Javier Segovia |
Web Intelligence | 3 |
| 2004 | Bayesian network multi-classifiers for protein secondary structure prediction
Víctor Robles, Pedro Larrañaga, José M. Peña 0002, Ernestina Menasalvas Ruiz, María S. Pérez 0001, Vanessa Herves, Anita Wasilewska |
Artif. Intell. Medicine | 4 |
| 2004 | Subsessions: A granular approach to click path analysisabstractThe fiercely competitive web-based electronic commerce (e-commerce) environment has made necessary the application of intelligent methods to gather and analyze information collected from consumer web sessions. Knowledge about user behavior and session goals can be discovered from the information gathered about user activities, as tracked by web clicks. Most current approaches to customer behavior analysis study the user session by examining each web page access. However, the abstraction of subsessions provides a more granular view of user activity. Here, we propose a method of increasing the granularity of the user session analysis by isolating useful subsessions within sessions. Each subsession represents a high-level user activity such as performing a purchase or searching for a particular type of information. Given a set of previously identified subsessions, we can determine at which point the user begins a preidentified subsession by tracking user clicks. With this information we can (1) optimize the user experience by precaching pages or (2) provide an adaptive user experience by presenting pages according to our estimation of the user's ultimate goal. To identify subsessions, we present an algorithm to compute frequent click paths from which subsessions then can be isolated. The algorithm functions by scanning all user sessions and extracting all frequent subpaths by using a distance function to determining subpath similarity. Each frequent subpath represents a subsession. An analysis of the pages represented by the subsession provides additional information about semantically related activities commonly performed by users. © 2004 Wiley Periodicals, Inc. Ernestina Menasalvas Ruiz, Socorro Millán, José M. Peña 0002, Michael Hadjimichael, Oscar Marbán |
Int. J. Intell. Syst. | 1 |
| 2003 | Interval Estimation Naïve Bayes
Víctor Robles, Pedro Larrañaga, José M. Peña 0002, Ernestina Menasalvas Ruiz, María S. Pérez 0001 |
IDA | 4 |
| 2003 | Improvement of Naïve Bayes Collaborative Filtering Using Interval EstimationabstractRecommender systems emerged to help users choose among the large amount of options that ecommerce sites offer. Collaborative filtering is one of the most successful recommender techniques. Here we propose an approach to collaborative filtering based on the simple Bayesian classifier. We propose a method of increasing the efficiency of naive Bayes by applying a new semi naive Bayes approach based on interval estimation. To evaluate our algorithm we use a database of Microsoft anonymous Web data from the UCl repository. Our empirical results show that our proposed Interval based naive Bayes approach outperforms typical naive Bayes. Víctor Robles, Pedro Larrañaga, Ernestina Menasalvas Ruiz, María S. Pérez 0001, Vanessa Herves |
Web Intelligence | 3 |
| 2003 | Expected Value of User Sessions: Limitations to the Non-Semantic ApproachabstractThe amazing evolution of e-commerce and the fierce competitive environment it has produced have encouraged commercial firms to apply intelligent methods to take advantage of competitors by gathering and analyzing information collected from consumer Web sessions. Knowledge about user objectives and session goals can be discovered from the information collected regarding user activities, as tracked by Web clicks. Most current approaches to customer behaviour analysis study the user session by examining only Web page accesses. To find out about navigators behaviour is crucial to Web sites sponsors attempting to evaluate the performance of their sites. Nevertheless, knowing the current navigation patterns is not always enough. Very often it is also necessary to measure sessions value according to business goals perspectives. We present two different measures to include business goals inside click stream analysis. Each of the alternatives is discussed and evaluated in terms of how company's objectives and expectations are taken into account as well as how this approach could be achieved. Ernestina Menasalvas Ruiz, B. Pardo, Socorro Millán, Esther Hochsztain, José M. Peña 0002 |
Web Intelligence | 1 |
| 2003 | Calculating economic indexes per household and censal section from official Spanish databases
Sonia Frutos, Ernestina Menasalvas Ruiz, César Montes, Javier Segovia |
Intell. Data Anal. | 2 |
| 2002 | Data mining - a semantic modelabstractData mining techniques applied to decision support in real-life problems require a multi-step process. Inputs and outputs of these steps require some standard format to be followed in order to achieve a useful platform for the execution of data mining algorithms. There is a need to develop a uniform model where every operation can be expressed in a standard way, allowing algorithms to cooperate and to reuse results. We present, first, a common structure for the representation of inter-step results, and second, a model of the operator, i.e. the entity that handles and transforms this common structure according to a basic data mining algorithm. Covadonga Fernández 0001, Juan F. Martínez, Anita Wasilewska, Michael Hadjimichael, Ernestina Menasalvas Ruiz |
FUZZ-IEEE | 5 |
| 2002 | Subsessions: a granular approach to click path analysisabstractElectronic, web-based commerce enables and demands the application of intelligent methods to analyze information collected from consumer web sessions. We propose a method of increasing the granularity of the user session analysis by isolating useful subsessions within web page access sessions, where each subsession represents a frequently traversed path indicating high-level user activity. The subsession approximates user state information as well as anticipated user activity, and as a result is useful for personalization and pre-caching. Ernestina Menasalvas Ruiz, Socorro Millán, José M. Peña 0002, Michael Hadjimichael, Oscar Marbán |
FUZZ-IEEE | 1 |
| 2002 | Automatic implementation system of security protocols based on formal description techniquesabstractWe present an automatic implementation system of security protocols based in formal description techniques. A sufficiently complete and concise formal specification that has allowed us to define the state machine that corresponds to a security protocol has been designed to achieve our goals. This formal specification makes it possible to incorporate in a flexible way the security mechanisms and functions (random numbers generation, timestamps, symmetric-key encryption, public-key cryptography, etc). Our solution implies the incorporation of an additional security layer LEI (Logical Element of Implementation) in the TCP/IP architecture. This additional layer be able both to interpret and to implement any security protocol from its formal specification. Our system provides an applications programming interface (API) for the development of distributed applications in the Internet like the e-commerce, bank transfers, network management or distribution information services that makes transparent to them the problem of security in the communications. Luis Mengual, Nicolás Barcia, Ernesto Jiménez, Ernestina Menasalvas Ruiz, Julio Setién, Javier Yágüez |
ISCC | 4 |
| 1999 | Rough Dependencies as a Particular Case of Correlation: Application to the Calculation of Approximative Reducts
María C. Fernández-Baizán, Ernestina Menasalvas Ruiz, José M. Peña 0002, Socorro Millán, Eloina Mesa |
PKDD | 2 |