VLDB 2026 Research / reviewers in the wild / expert
Alexander G. Ioannidis
dblp:328/5925
· DBLP profile ↗
12ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-4735-7803ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Chromosome Parallelization for Precision Medicine Genomic WorkflowsabstractLarge-scale genomic workflows used in precision medicine can process datasets spanning tens to hundreds of gigabytes per sample, leading to high memory spikes, intensive disk I/O, and task failures due to out-of-memory errors. Simple static resource allocation methods struggle to handle the variability in per-chromosome RAM demands, resulting in poor resource utilization and long runtimes. In this work, we propose multiple mechanisms for adaptive, RAM-efficient parallelization of chromosome-level bioinformatics workflows. First, we develop a symbolic regression model that estimates per-chromosome memory consumption for a given task and introduces an interpolating bias to conservatively minimize over-allocation. Second, we present a dynamic scheduler that adaptively predicts RAM usage with a polynomial regression model, treating task packing as a Knapsack problem to optimally batch jobs based on predicted memory requirements. Additionally, we present a static scheduler that optimizes chromosome processing order to minimize peak memory while preserving throughput. Our proposed methods, evaluated on simulations and real-world genomic pipelines, provide new mechanisms to reduce memory overruns and balance load across threads. We thereby achieve faster end-to-end execution, showcasing the potential to optimize large-scale genomic workflows. Daniel Mas Montserrat, Ray Verma, Míriam Barrabés, Francisco M. de la Vega, Carlos D. Bustamante, Alexander G. Ioannidis |
AAAI | 6 |
| 2025 | Feature Shift Localization NetworkabstractFeature shifts between data sources are present in many applications involving healthcare, biomedical, socioeconomic, financial, survey, and multi-sensor data, among others, where unharmonized heterogeneous data sources, noisy data measurements, or inconsistent processing and standardization pipelines can lead to erroneous features. Localizing shifted features is important to address the underlying cause of the shift and correct or filter the data to avoid degrading downstream analysis. While many techniques can detect distribution shifts, localizing the features originating them is still challenging, with current solutions being either inaccurate or not scalable to large and high-dimensional datasets. In this work, we introduce the Feature Shift Localization Network (FSL-Net), a neural network that can localize feature shifts in large and high-dimensional datasets in a fast and accurate manner. The network, trained with a large number of datasets, learns to extract the statistical properties of the datasets and can localize feature shifts from previously unseen datasets and shifts without the need for re-training. The code and ready-to-use trained model are available at \url{https://github.com/AI-sandbox/FSL-Net}. Míriam Barrabés, Daniel Mas Montserrat, Kapal Dev, Alexander G. Ioannidis |
ICML | 4 |
| 2025 | Compressive Meta-LearningabstractThe rapid expansion in the size of new datasets has created a need for fast and efficient parameter-learning techniques.Compressive learning is a framework that enables efficient processing by using random, nonlinear features to project large-scale databases onto compact, information-preserving representations whose dimensionality is independent of the number of samples and can be easily stored, transferred, and processed.These database-level summaries are then used to decode parameters of interest from the underlying data distribution without requiring access to the original samples, offering an efficient and privacy-friendly learning framework.However, both the encoding and decoding techniques are typically randomized and data-independent, failing to exploit the underlying structure of the data.In this work, we propose a framework that meta-learns both the encoding and decoding stages of compressive learning methods by using neural networks that provide faster and more accurate systems than the current state-of-the-art approaches.To demonstrate the potential of the presented Compressive Meta-Learning framework, we explore multiple applications-including neural network-based compressive PCA, compressive ridge regression, compressive k-means, and autoencoders. Daniel Mas Montserrat, David Bonet, Maria Perera, Xavier Giró-i-Nieto, Alexander G. Ioannidis |
KDD (2) | 5 |
| 2024 | HyperFast: Instant Classification for Tabular DataabstractTraining deep learning models and performing hyperparameter tuning can be computationally demanding and time-consuming. Meanwhile, traditional machine learning methods like gradient-boosting algorithms remain the preferred choice for most tabular data applications, while neural network alternatives require extensive hyperparameter tuning or work only in toy datasets under limited settings. In this paper, we introduce HyperFast, a meta-trained hypernetwork designed for instant classification of tabular data in a single forward pass. HyperFast generates a task-specific neural network tailored to an unseen dataset that can be directly used for classification inference, removing the need for training a model. We report extensive experiments with OpenML and genomic data, comparing HyperFast to competing tabular data neural networks, traditional ML methods, AutoML systems, and boosting machines. HyperFast shows highly competitive results, while being significantly faster. Additionally, our approach demonstrates robust adaptability across a variety of classification tasks with little to no fine-tuning, positioning HyperFast as a strong solution for numerous applications and rapid model deployment. HyperFast introduces a promising paradigm for fast classification, with the potential to substantially decrease the computational burden of deep learning. Our code, which offers a scikit-learn-like interface, along with the trained HyperFast model, can be found at https://github.com/AI-sandbox/HyperFast. David Bonet, Daniel Mas Montserrat, Xavier Giró-i-Nieto, Alexander G. Ioannidis |
AAAI | 4 |
| 2023 | Genomic Databases Homogenization with Machine LearningabstractLarge-scale and increasingly diverse datasets power modern genomic studies, yet robust data integration and homogenization across varying sources remains a challenge. The multiplicity of file formats and the computational requirements imposed by large genomic datasets make it difficult to deal with multiple data sources. Furthermore, there is a lack of open-source customizable tools to merge genomic databases while providing quality control functionalities. To fill this gap, we present MergeGenome, a machine learning-based method designed to integrate DNA sequences from multiple variant call format (VCF) files while maintaining data quality. By leveraging pre-existing VCF manipulation and imputation software, MergeGenome provides a robust pipeline of comprehensive steps to standardize nomenclature, remove ambiguities, correct strand alignment, eliminate mismatches, impute missing positions, and filter and correct erroneous variants with machine learning, among other functionalities. We demonstrate MergeGenome’s ability to obtain a high-quality combined dataset by merging two databases containing dog DNA and effectively detecting and correcting imputation errors. Finally, we show that using the homogenized dataset provides a boost in phenotype prediction performance. Míriam Barrabés, David Bonet, Víctor Novelle Moriano, Xavier Giró-i-Nieto, Daniel Mas Montserrat, Alexander G. Ioannidis |
BIBM | 6 |
| 2023 | Assessing Tree-Based Phenotype Prediction on the UK BiobankabstractPrecision medicine relies on the ability to identify associations between genomic data and its phenotypic expression in order to provide personalized predictions. Phenotype prediction using statistical models trained on large-scale genomic and phenotypic data is a critical research area at the intersection of machine learning and genomics. Current genotype-to-phenotype models, such as polygenic risk scores, only account for linear relationships, and the use of nonlinear methods is still partially unexplored. In this work, we evaluate the prediction accuracy and scalability of nine nonlinear decision tree-based algorithms, including ensembling and boosting mechanisms, and compare them to linear prediction models. We assess the prediction performance for 24 anthropometric and disease-related phenotypes present in the UK Biobank. By using random feature selection, we explore how accuracy and computational time vary for each method as a function of the number of genetic variants selected. Our results show that tree-based methods, especially gradientboosted trees, can offer superior predictions with computational times comparable to those of linear methods. Thus, models able to capture nonlinear relationships between genotypes and phenotypes merit consideration for integration in upcoming computational systems for personalized medicine. Alex Meléndez, Cayetana López, David Bonet, Gerard Sant, Daniel Mas Montserrat, Jordi Abante, Manuel Rivas Pérez, Ferran Marqués, Alexander G. Ioannidis |
BIBM | 9 |
| 2023 | Adversarial Attacks on Genotype SequencesabstractAdversarial attacks can drastically change the output of a method by small alterations to its input. While this can be a useful framework to analyze worst-case robustness, it can also be used by malicious agents to damage machine learning-based applications. The proliferation of platforms that allow users to share their DNA sequences and phenotype information to enable association studies has led to an increase in large genomic databases. Such open platforms are, however, vulnerable to malicious users uploading corrupted genetic sequence files that could damage downstream studies. Such studies commonly include steps involving the analysis of the genomic sequences’ structure using dimensionality reduction techniques and ancestry inference methods. In this paper we show how white-box gradient-based adversarial attacks can be used to corrupt the output of genomic analyses, and we explore different machine learning techniques to detect such manipulations. Daniel Mas Montserrat, Alexander G. Ioannidis |
ICASSP | 2 |
| 2023 | Adversarial Learning for Feature Shift Detection and CorrectionabstractData shift is a phenomenon present in many real-world applications, and while there are multiple methods attempting to detect shifts, the task of localizing and correcting the features originating such shifts has not been studied in depth. Feature shifts can occur in many datasets, including in multi-sensor data, where some sensors are malfunctioning, or in tabular and structured data, including biomedical, financial, and survey data, where faulty standardization and data processing pipelines can lead to erroneous features. In this work, we explore using the principles of adversarial learning, where the information from several discriminators trained to distinguish between two distributions is used to both detect the corrupted features and fix them in order to remove the distribution shift between datasets. We show that mainstream supervised classifiers, such as random forest or gradient boosting trees, combined with simple iterative heuristics, can localize and correct feature shifts, outperforming current statistical and neural network-based techniques. The code is available at https://github.com/AI-sandbox/DataFix. Míriam Barrabés, Daniel Mas Montserrat, Margarita Geleta, Xavier Giró-i-Nieto, Alexander G. Ioannidis |
NeurIPS | 5 |
| 2022 | SALAI-Net: species-agnostic local ancestry inference networkabstractMOTIVATION: Local ancestry inference (LAI) is the high resolution prediction of ancestry labels along a DNA sequence. LAI is important in the study of human history and migrations, and it is beginning to play a role in precision medicine applications including ancestry-adjusted genome-wide association studies (GWASs) and polygenic risk scores (PRSs). Existing LAI models do not generalize well between species, chromosomes or even ancestry groups, requiring re-training for each different setting. Furthermore, such methods can lack interpretability, which is an important element in each of these applications. RESULTS: We present SALAI-Net, a portable statistical LAI method that can be applied on any set of species and ancestries (species-agnostic), requiring only haplotype data and no other biological parameters. Inspired by identity by descent methods, SALAI-Net estimates population labels for each segment of DNA by performing a reference matching approach, which leads to an interpretable and fast technique. We benchmark our models on whole-genome data of humans and we test these models' ability to generalize to dog breeds when trained on human data. SALAI-Net outperforms previous methods in terms of balanced accuracy, while generalizing between different settings, species and datasets. Moreover, it is up to two orders of magnitude faster and uses considerably less RAM memory than competing methods. AVAILABILITY AND IMPLEMENTATION: We provide an open source implementation and links to publicly available data at github.com/AI-sandbox/SALAI-Net. Data is publicly available as follows: https://www.internationalgenome.org (1000 Genomes), https://www.simonsfoundation.org/simons-genome-diversity-project (Simons Genome Diversity Project), https://www.sanger.ac.uk/resources/downloads/human/hapmap3.html (HapMap), ftp://ngs.sanger.ac.uk/production/hgdp/hgdp_wgs.20190516 (Human Genome Diversity Project) and https://www.ncbi.nlm.nih.gov/bioproject/PRJNA448733 (Canid genomes). SUPPLEMENTARY INFORMATION: Supplementary data are available from Bioinformatics online. Benet Oriol Sabat, Daniel Mas Montserrat, Xavier Giró-i-Nieto, Alexander G. Ioannidis |
Bioinform. | 4 |
| 2022 | Archetypal Analysis for population geneticsabstractThe estimation of genetic clusters using genomic data has application from genome-wide association studies (GWAS) to demographic history to polygenic risk scores (PRS) and is expected to play an important role in the analyses of increasingly diverse, large-scale cohorts. However, existing methods are computationally-intensive, prohibitively so in the case of nationwide biobanks. Here we explore Archetypal Analysis as an efficient, unsupervised approach for identifying genetic clusters and for associating individuals with them. Such unsupervised approaches help avoid conflating socially constructed ethnic labels with genetic clusters by eliminating the need for exogenous training labels. We show that Archetypal Analysis yields similar cluster structure to existing unsupervised methods such as ADMIXTURE and provides interpretative advantages. More importantly, we show that since Archetypal Analysis can be used with lower-dimensional representations of genetic data, significant reductions in computational time and memory requirements are possible. When Archetypal Analysis is run in such a fashion, it takes several orders of magnitude less compute time than the current standard, ADMIXTURE. Finally, we demonstrate uses ranging across datasets from humans to canids. Julia Gimbernat-Mayol, Albert Dominguez Mantes, Carlos D. Bustamante, Daniel Mas Montserrat, Alexander G. Ioannidis |
PLoS Comput. Biol. | 5 |
| 2021 | Discovering prescription patterns in pediatric acute-onset neuropsychiatric syndrome patientsabstractOBJECTIVE: Pediatric acute-onset neuropsychiatric syndrome (PANS) is a complex neuropsychiatric syndrome characterized by an abrupt onset of obsessive-compulsive symptoms and/or severe eating restrictions, along with at least two concomitant debilitating cognitive, behavioral, or neurological symptoms. A wide range of pharmacological interventions along with behavioral and environmental modifications, and psychotherapies have been adopted to treat symptoms and underlying etiologies. Our goal was to develop a data-driven approach to identify treatment patterns in this cohort. MATERIALS AND METHODS: In this cohort study, we extracted medical prescription histories from electronic health records. We developed a modified dynamic programming approach to perform global alignment of those medication histories. Our approach is unique since it considers time gaps in prescription patterns as part of the similarity strategy. RESULTS: This study included 43 consecutive new-onset pre-pubertal patients who had at least 3 clinic visits. Our algorithm identified six clusters with distinct medication usage history which may represent clinician's practice of treating PANS of different severities and etiologies i.e., two most severe groups requiring high dose intravenous steroids; two arthritic or inflammatory groups requiring prolonged nonsteroidal anti-inflammatory drug (NSAID); and two mild relapsing/remitting group treated with a short course of NSAID. The psychometric scores as outcomes in each cluster generally improved within the first two years. DISCUSSION AND CONCLUSION: Our algorithm shows potential to improve our knowledge of treatment patterns in the PANS cohort, while helping clinicians understand how patients respond to a combination of drugs. Arturo L. Pineda, Armin Pourshafeie, Alexander G. Ioannidis, Collin McCloskey Leibold, Avis L. Chan, Carlos D. Bustamante, Jennifer Frankovich, Genevieve L. Wojcik |
J. Biomed. Informatics | 3 |
| 2020 | Lai-Net: Local-Ancestry Inference with Neural NetworksabstractLocal-ancestry inference (LAI), also referred to as ancestry deconvolution, provides high-resolution ancestry estimation along the human genome. In both research and industry, LAI is emerging as a critical step in DNA sequence analysis with applications extending from polygenic risk scores (used to predict traits in embryos and disease risk in adults) to genome-wide association studies, and from pharmacogenomics to inference of human population history. While many LAI methods have been developed, advances in computing hardware (GPUs) combined with machine learning techniques, such as neural networks, are enabling the development of new methods that are fast, robust and easily shared and stored. In this paper we develop the first neural network based LAI method, named LAI-Net, providing competitive accuracy with state-of-the-art methods and robustness to missing or noisy data, while having a small number of layers. Daniel Mas Montserrat, Carlos D. Bustamante, Alexander G. Ioannidis |
ICASSP | 3 |