EDBT 2026 Demo / reviewers in the wild / expert
Miguel García-Remesal
dblp:51/2035
· DBLP profile ↗
14ranked-venue papers
5as first author
3since 2021 · last 2026
0000-0002-5948-8691ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 4 first-authorArtificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Learning paradigms · 61% Probabilistic and Bayesian machine learning · 30% Face, body and person analysis · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning paradigms
imbalanced learning |
1.0 | 1 | 2026 | A Selective Under-Sampling (SUS) Method for Imbalanced Regression (Abstract Reprint) · AAAI 2026 |
Machine learning › Learning paradigms › imbalanced learning
imbalanced regression |
1.0 | 1 | 2026 | A Selective Under-Sampling (SUS) Method for Imbalanced Regression (Abstract Reprint) · AAAI 2026 |
Machine learning › Probabilistic and Bayesian machine learning
sampling |
1.0 | 1 | 2026 | A Selective Under-Sampling (SUS) Method for Imbalanced Regression (Abstract Reprint) · AAAI 2026 |
Computer vision › Face, body and person analysis › facial age estimation
age estimation |
0.3 | 1 | 2026 | A Selective Under-Sampling (SUS) Method for Imbalanced Regression (Abstract Reprint) · AAAI 2026 |
Bioinformatics and computational biology
biomedical text mining |
0.1 | 1 | 2010 | PubDNA Finder: a web database linking full-text articles to sequences of nucleic acids · Bioinform. 2010 |
Bioinformatics and computational biology
biological database |
0.0 | 1 | 2010 | PubDNA Finder: a web database linking full-text articles to sequences of nucleic acids · Bioinform. 2010 |
Methods — techniques the papers use, named apart from their topics
selective under-sampling · 1.0random under-sampling · 1.0SMOGN · 1.0text mining · 0.1sequence extraction · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Selective Under-Sampling (SUS) Method for Imbalanced Regression (Abstract Reprint)abstractMany mainstream machine learning approaches, such as neural networks, are not well suited to work with imbalanced data. Yet, this problem is frequently present in many real-world data sets. Collection methods are imperfect, and often not able to capture enough data in a specific range of the target variable. Furthermore, in certain tasks data is inherently imbalanced with many more normal events than edge cases. This problem is well studied within the classification context. However, only several methods have been proposed to deal with regression tasks. In addition, the proposed methods often not yield good performance with high-dimensional data, while imbalanced high-dimensional regression has scarcely been explored. In this paper we present a selective under-sampling (SUS) algorithm for dealing with imbalanced regression and its iterative version SUSiter. We assessed this method on 15 regression data sets from different imbalanced domains, 5 synthetic high-dimensional imbalanced data sets and 2 more complex imbalanced age estimation image data sets. Our results suggest that SUS and SUSiter typically outperform other state-of-the-art techniques like SMOGN, or random under-sampling, when used with neural networks as learners. Jovana Aleksic, Miguel García-Remesal |
AAAI | 2 |
| 2025 | A Selective Under-Sampling (SUS) Method for Imbalanced RegressionabstractMany mainstream machine learning approaches, such as neural networks, are not well suited to work with imbalanced data. Yet, this problem is frequently present in many real-world data sets. Collection methods are imperfect, and often not able to capture enough data in a specific range of the target variable. Furthermore, in certain tasks data is inherently imbalanced with many more normal events than edge cases. This problem is well studied within the classification context. However, only several methods have been proposed to deal with regression tasks. In addition, the proposed methods often do not yield good performance with high-dimensional data, while imbalanced high-dimensional regression has scarcely been explored. In this paper we present a selective under-sampling (SUS) algorithm for dealing with imbalanced regression and its iterative version SUSiter. We assessed this method on 15 regression data sets from different imbalanced domains, 5 synthetic high-dimensional imbalanced data sets and 2 more complex imbalanced age estimation image data sets. Our results suggest that SUS and SUSiter typically outperform other state-of-the-art techniques like SMOGN, or random under-sampling, when used with neural networks as learners. Jovana Aleksic, Miguel García-Remesal |
J. Artif. Intell. Res. | 2 |
| 2024 | End-to-end entity extraction from OCRed texts using summarization models
Pedro A. Villa-García, Raúl Alonso-Calvo, Miguel García-Remesal |
Neural Comput. Appl. | 3 |
| 2016 | A method and software framework for enriching private biomedical sources with data from public online repositories
Alberto Anguita, Miguel García-Remesal, Norbert Graf 0001, Victor Maojo |
J. Biomed. Informatics | 2 |
| 2012 | On a meaningful integration of web services in data-intensive biomedical environments: The DICODE approachabstractThis paper reports on an innovative approach that aims to reduce information management costs in dataintensive and cognitively-complex biomedical environments. Recognizing the importance of prominent high-performance computing paradigms and large data processing technologies as well as collaboration support systems to remedy data-intensive issues, it adopts a hybrid approach by building on the synergy of these technologies. The proposed approach provides innovative Web-based workbenches that integrate and orchestrate a set of interoperable services that reduce the data-intensiveness and complexity overload at critical decision points to a manageable level, thus permitting stakeholders to be more productive and concentrate on creative activities. Maurizio de la Calle, Miguel García-Remesal, Manolis Tzagarakis, Spyros Christodoulou, Georgia Tsiliki, Nikos I. Karacapilidis |
CBMS | 2 |
| 2010 | PubDNA Finder: a web database linking full-text articles to sequences of nucleic acidsabstractSUMMARY: PubDNA Finder is an online repository that we have created to link PubMed Central manuscripts to the sequences of nucleic acids appearing in them. It extends the search capabilities provided by PubMed Central by enabling researchers to perform advanced searches involving sequences of nucleic acids. This includes, among other features (i) searching for papers mentioning one or more specific sequences of nucleic acids and (ii) retrieving the genetic sequences appearing in different articles. These additional query capabilities are provided by a searchable index that we created by using the full text of the 176 672 papers available at PubMed Central at the time of writing and the sequences of nucleic acids appearing in them. To automatically extract the genetic sequences occurring in each paper, we used an original method we have developed. The database is updated monthly by automatically connecting to the PubMed Central FTP site to retrieve and index new manuscripts. Users can query the database via the web interface provided. AVAILABILITY: PubDNA Finder can be freely accessed at http://servet.dia.fi.upm.es:8080/pubdnafinder Miguel García-Remesal, Alejandro Cuevas, David Pérez-Rey, Luis Martín, Alberto Anguita, Diana de la Iglesia, Guillermo de la Calle, José Crespo, Victor Maojo |
Bioinform. | 1 |
| 2010 | A method for automatically extracting infectious disease-related primers and probes from the literatureabstractBACKGROUND: Primer and probe sequences are the main components of nucleic acid-based detection systems. Biologists use primers and probes for different tasks, some related to the diagnosis and prescription of infectious diseases. The biological literature is the main information source for empirically validated primer and probe sequences. Therefore, it is becoming increasingly important for researchers to navigate this important information. In this paper, we present a four-phase method for extracting and annotating primer/probe sequences from the literature. These phases are: (1) convert each document into a tree of paper sections, (2) detect the candidate sequences using a set of finite state machine-based recognizers, (3) refine problem sequences using a rule-based expert system, and (4) annotate the extracted sequences with their related organism/gene information. RESULTS: We tested our approach using a test set composed of 297 manuscripts. The extracted sequences and their organism/gene annotations were manually evaluated by a panel of molecular biologists. The results of the evaluation show that our approach is suitable for automatically extracting DNA sequences, achieving precision/recall rates of 97.98% and 95.77%, respectively. In addition, 76.66% of the detected sequences were correctly annotated with their organism name. The system also provided correct gene-related information for 46.18% of the sequences assigned a correct organism name. CONCLUSIONS: We believe that the proposed method can facilitate routine tasks for biomedical researchers using molecular methods to diagnose and prescribe different infectious diseases. In addition, the proposed method can be expanded to detect and extract other biological sequences from the literature. The extracted information can also be used to readily update available primer/probe databases or to create new databases from scratch. Miguel García-Remesal, Alejandro Cuevas, Victoria López-Alonso, Guillermo López-Campos, Guillermo de la Calle, Diana de la Iglesia, David Pérez-Rey, José Crespo, Fernando Martín-Sánchez, Victor Maojo |
BMC Bioinform. | 1 |
| 2009 | BIRI: a new approach for automatically discovering and indexing available public bioinformatics resources from the literatureabstractBACKGROUND: The rapid evolution of Internet technologies and the collaborative approaches that dominate the field have stimulated the development of numerous bioinformatics resources. To address this new framework, several initiatives have tried to organize these services and resources. In this paper, we present the BioInformatics Resource Inventory (BIRI), a new approach for automatically discovering and indexing available public bioinformatics resources using information extracted from the scientific literature. The index generated can be automatically updated by adding additional manuscripts describing new resources. We have developed web services and applications to test and validate our approach. It has not been designed to replace current indexes but to extend their capabilities with richer functionalities. RESULTS: We developed a web service to provide a set of high-level query primitives to access the index. The web service can be used by third-party web services or web-based applications. To test the web service, we created a pilot web application to access a preliminary knowledge base of resources. We tested our tool using an initial set of 400 abstracts. Almost 90% of the resources described in the abstracts were correctly classified. More than 500 descriptions of functionalities were extracted. CONCLUSION: These experiments suggest the feasibility of our approach for automatically discovering and indexing current and future bioinformatics resources. Given the domain-independent characteristics of this tool, it is currently being applied by the authors in other areas, such as medical nanoinformatics. BIRI is available at http://edelman.dia.fi.upm.es/biri/. Guillermo de la Calle, Miguel García-Remesal, Stefano Chiesa, Diana de la Iglesia, Victor Maojo |
BMC Bioinform. | 2 |
| 2008 | Building an Index of Nanomedical Resources: An Automatic Approach Based on Text Mining
Stefano Chiesa, Miguel García-Remesal, Guillermo de la Calle, Diana de la Iglesia, Vaida Bankauskaite, Victor Maojo |
KES (2) | 2 |
| 2008 | Using Hierarchical Task Network Planning Techniques to Create Custom Web Search Services over Multiple Biomedical Databases
Miguel García-Remesal |
KES (2) | 1 |
| 2007 | Logical Schema Acquisition from Text-Based Sources for Structured and Non-Structured Biomedical Sources Integration
Miguel García-Remesal, Victor Maojo, José Crespo, Holger Billhardt |
AMIA | 1 |
| 2007 | An agent- and ontology-based system for integrating public gene, protein, and disease databases
Raúl Alonso-Calvo, Victor Maojo, Holger Billhardt, Fernando Martín-Sánchez, Miguel García-Remesal, David Pérez-Rey |
J. Biomed. Informatics | 5 |
| 2004 | ARMEDA II: Supporting Genomic Medicine through the Integration of Medical and Genetic DatabasesabstractIn this paper we present ARMEDA II, a project designed to integrate distributed heterogeneous medical and genetic databases in support of genomic medicine. In this project, we have followed a "virtual repository" or VR approach. Although VRs are entities that do not contain any data, but metadata, they give users the perception of being working with local repositories that integrate data from different and remote sources. Our approach is based on two basic operators employed to connect new databases to the system: mapping and unification. The mapping process produces what is called the "virtual conceptual schema" of the newly created VR while the unification process provides tools to create an integrated virtual schema for at least two pre-existing VRs. We tested the current implementation of ARMEDA Il using two tumor databases, one containing information from a hospital and the other containing genetic data associated to the tumor samples. The performance of the system was also evaluated using a pre-created set of 30 queries. For all queries the test yielded promising results since the system successfully retrieved the correct information. The ARMEDA II project is the current version of an ongoing project developed in the framework of an European Commission funded project. Miguel García-Remesal, Victor Maojo, Holger Billhardt, José Crespo, Raúl Alonso-Calvo, David Pérez-Rey, Fernando Martin, A. Sousa |
BIBE | 1 |
| 2004 | Biomedical Ontologies in Post-Genomic Information SystemsabstractAfter the completion of the Human Genome Project, a new, post genomic era, is beginning to analyze and interpret the huge amount of genomic information. Information methods and techniques from areas such as database integration, information retrieval, knowledge discovery in databases (KDD) and decision support systems (DSS) are needed. These systems should take into account idiosyncratic differences between these two interacting fields, medicine and biology. Their correspondent medical informatics (MI) and bioinformatics (BI) should also interact and there is a need for a point to support the communication. Biomedical ontologies can be used to enhance biomedical information systems, providing a knowledge sharing framework. However, ontology tools are still in its infancy and there is a need of standards, services, automatic management tools, etc... to be able to properly apply this technology environment. Nevertheless, ontologies are just the technical framework the most important issue is the content and the use policy. David Pérez-Rey, Victor Maojo, Miguel García-Remesal, Raúl Alonso-Calvo |
BIBE | 3 |