VLDB 2026 Research / reviewers in the wild / expert
Patrik Waldmann
dblp:45/5734
· DBLP profile ↗
7ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-2390-6609ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Proximal regularization of deep residual neural networks applied to high-dimensional genomic dataabstractHigh-dimensional genomic datasets contain complex patterns shaped by substantial biological noise, which pose major challenges for predictive modeling in genetics and breeding. Residual neural networks (ResNets) provide a powerful framework for capturing nonlinear genomic effects, but often overfit in settings where marker numbers greatly exceed sample sizes. As a solution, a range of regularization methods have been proposed. One promising approach relies on the proximal mapping technique, which is computationally efficient since it can be directly incorporated into the optimization algorithm. However, the performance of ResNets with various convex or non-convex proximal regularizers remains under-explored on high-dimensional data. In this study, we propose an extended stochastic adaptive proximal gradient ResNet method that can handle both convex and non-convex regularizers that range from $L_{0}$ to $L_{\infty }$ and give more analysis of the convergence guarantee for the convex and non-convex regularizers. Moreover, we evaluate the prediction performance in a supervised regression setting on four real high-dimensional genomic datasets from mice, pig, wheat, and loblolly pine. For comparison, we also implement and evaluate traditional sparse linear proximal methods with the same regularizers, as well as LightGBM. Experimental results demonstrate that an 18-layer ResNet with $L_{\frac{1}{2}}$ regularization outperforms other configurations on both mice and pig datasets. For the wheat and loblolly pine data, the 15-layer ResNet $L_{\frac{1}{2}}$ configuration achieves the lowest test mean squared errors and the highest distance correlation (dCor). These findings highlight the effectiveness of the regularized adaptive proximal gradient ResNet method and its potential for prediction tasks on high-dimensional genomic data. Yuhua Fan, Ilkka Launonen, Mikko J. Sillanpää, Patrik Waldmann |
Briefings Bioinform. | 4 |
| 2026 | A Proximal Multi-Objective Optimization Method for Incorporation of Polygenic Breeding Values in Genomic PredictionabstractTraditional quantitative genetics rely on known pedigrees for estimating polygenic breeding values and variance components. With the advent of high-throughput sequencing, estimation of empirical relationships between individuals have become feasible. The single-step genomic BLUP procedure integrates polygenic and genomic information but requires intricate pedigree connections, leading to computationally demanding matrix operations. Here, we propose DYpREG, a flexible proximal operator multi-objective regularization method. DYpREG is a splitting algorithm that accommodates two distinct loss functions and a regularization function that can be chosen as LASSO, ridge regression (RR) or elastic net (EN). Notably, it incorporates pre-calculated breeding values without the need for extensive relationship matrix computations. Evaluated on two traits with different heritabilities from mice and pig, DYpREG with RR regularizer showed the best predictive performance compared to DYpREG with LASSO and EN regularizers as well as to all methods lacking polygenic information. The mean squared error (MSE) reduction was notable for both mice traits (1.5% and 24.77%) and both pig traits (2.8% and 2.1%). Comparative analyses with Bayesian reproducing kernel Hilbert space (RKHS) and spike-and-slab (BayesC) regression also revealed consistent improvements. Hence, DYpREG effectively integrates polygenic information, enhancing prediction accuracy in genomic evaluations. Patrik Waldmann, Yuhua Fan |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2025 | Multi-task genomic prediction using gated residual variable selection neural networksabstractBACKGROUND: The recent development of high-throughput sequencing techniques provide massive data that can be used in genome-wide prediction (GWP). Although GWP is effective on its own, the incorporation of traditional polygenic pedigree information into GWP has been shown to further improve prediction accuracy. However, most of the methods developed in this field require that individuals with genomic information can be connected to the polygenic pedigree within a standard linear mixed model framework that involves calculation of computationally demanding matrix inverses of the combined pedigrees. The extension of this integrated approach to more flexible machine learning methods has been slow. METHODS: This study aims to enhance genomic prediction by implementing gated residual variable selection neural networks (GRVSNN) for multi-task genomic prediction. By integrating low-rank information from pedigree-based relationship matrices with genomic markers, we seek to improve predictive accuracy and interpretability compared to conventional regression and deep learning (DL) models. The prediction properties of the GRVSNN model are evaluated on several real-world datasets, including loblolly pine, mouse and pig. RESULTS: The experimental results demonstrate that the GRVSNN model outperforms traditional tabular genomic prediction models, including Bayesian regression methods and LassoNet. Using genomic and pedigree information, GRVSNN achieves a lower mean squared error (MSE), and higher Pearson (r) and distance (dCor) correlation between predicted and true phenotypic values in the test data. Moreover, GRVSNN selects fewer genetic markers and pedigree loadings which improves interpretability. CONCLUSION: The suggested GRVSNN framework provides a novel and computationally effective approach to improve genomic prediction accuracy by integrating information from traditional pedigrees with genomic data. The model's ability to conduct multi-task predictions underscores its potential to enhance selection processes in agricultural species and predict multiple diseases in precision medicine. Yuhua Fan, Patrik Waldmann |
BMC Bioinform. | 2 |
| 2024 | Tabular deep learning: a comparative study applied to multi-task genome-wide predictionabstractPURPOSE: More accurate prediction of phenotype traits can increase the success of genomic selection in both plant and animal breeding studies and provide more reliable disease risk prediction in humans. Traditional approaches typically use regression models based on linear assumptions between the genetic markers and the traits of interest. Non-linear models have been considered as an alternative tool for modeling genomic interactions (i.e. non-additive effects) and other subtle non-linear patterns between markers and phenotype. Deep learning has become a state-of-the-art non-linear prediction method for sound, image and language data. However, genomic data is better represented in a tabular format. The existing literature on deep learning for tabular data proposes a wide range of novel architectures and reports successful results on various datasets. Tabular deep learning applications in genome-wide prediction (GWP) are still rare. In this work, we perform an overview of the main families of recent deep learning architectures for tabular data and apply them to multi-trait regression and multi-class classification for GWP on real gene datasets. METHODS: The study involves an extensive overview of recent deep learning architectures for tabular data learning: NODE, TabNet, TabR, TabTransformer, FT-Transformer, AutoInt, GANDALF, SAINT and LassoNet. These architectures are applied to multi-trait GWP. Comprehensive benchmarks of various tabular deep learning methods are conducted to identify best practices and determine their effectiveness compared to traditional methods. RESULTS: Extensive experimental results on several genomic datasets (three for multi-trait regression and two for multi-class classification) highlight LassoNet as a standout performer, surpassing both other tabular deep learning models and the highly efficient tree based LightGBM method in terms of both best prediction accuracy and computing efficiency. CONCLUSION: Through series of evaluations on real-world genomic datasets, the study identifies LassoNet as a standout performer, surpassing decision tree methods like LightGBM and other tabular deep learning architectures in terms of both predictive accuracy and computing efficiency. Moreover, the inherent variable selection property of LassoNet provides a systematic way to find important genetic markers that contribute to phenotype expression. Yuhua Fan, Patrik Waldmann |
BMC Bioinform. | 2 |
| 2021 | A proximal LAVA method for genome-wide association and prediction of traits with mixed inheritance patternsabstractBACKGROUND: The genetic basis of phenotypic traits is highly variable and usually divided into mono-, oligo- and polygenic inheritance classes. Relatively few traits are known to be monogenic or oligogeneic. The majority of traits are considered to have a polygenic background. To what extent there are mixtures between these classes is unknown. The rapid advancement of genomic techniques makes it possible to directly map large amounts of genomic markers (GWAS) and predict unknown phenotypes (GWP). Most of the multi-marker methods for GWAS and GWP falls into one of two regularization frameworks. The first framework is based on [Formula: see text]-norm regularization (e.g. the LASSO) and is suitable for mono- and oligogenic traits, whereas the second framework regularize with the [Formula: see text]-norm (e.g. ridge regression; RR) and thereby is favourable for polygenic traits. A general framework for mixed inheritance is lacking. RESULTS: We have developed a proximal operator algorithm based on the recent LAVA regularization method that jointly performs [Formula: see text]- and [Formula: see text]-norm regularization. The algorithm is built on the alternating direction method of multipliers and proximal translation mapping (LAVA ADMM). When evaluated on the simulated QTLMAS2010 data, it is shown that the LAVA ADMM together with Bayesian optimization of the regularization parameters provides an efficient approach with lower test prediction mean-squared-error (65.89) than the LASSO (66.11), Ridge regression (83.41) and Elastic net (66.11). For the real pig data the test MSE of the LAVA ADMM is 0.850 compared to the LASSO, RR and EN with 0.875, 0.853 and 0.853, respectively. CONCLUSIONS: This study presents the LAVA ADMM that is capable of joint modelling of monogenic major genetic effects and polygenic minor genetic effects which can be used for both genome-wide assoiciation and prediction purposes. The statistical evaluations based on both simulated and real pig data set shows that the LAVA ADMM has better prediction properies than the LASSO, RR and EN. Julia code for the LAVA ADMM is available at: https://github.com/patwa67/LAVAADMM . Patrik Waldmann |
BMC Bioinform. | 1 |
| 2019 | AUTALASSO: an automatic adaptive LASSO for genome-wide predictionabstractBACKGROUND: Genome-wide prediction has become the method of choice in animal and plant breeding. Prediction of breeding values and phenotypes are routinely performed using large genomic data sets with number of markers on the order of several thousands to millions. The number of evaluated individuals is usually smaller which results in problems where model sparsity is of major concern. The LASSO technique has proven to be very well-suited for sparse problems often providing excellent prediction accuracy. Several computationally efficient LASSO algorithms have been developed, but optimization of hyper-parameters can be demanding. RESULTS: We have developed a novel automatic adaptive LASSO (AUTALASSO) based on the alternating direction method of multipliers (ADMM) optimization algorithm. The two major hyper-parameters of ADMM are the learning rate and the regularization factor. The learning rate is automatically tuned with line search and the regularization factor optimized using Golden section search. Results show that AUTALASSO provides superior prediction accuracy when evaluated on simulated and real bull data compared to the adaptive LASSO, LASSO and ridge regression implemented in the popular glmnet software. CONCLUSIONS: The AUTALASSO provides a very flexible and computationally efficient approach to GWP, especially when it is important to obtain high prediction accuracy and genetic gain. The AUTALASSO also has the capability to perform GWAS of both additive and dominance effects with smaller prediction error than the ordinary LASSO. Patrik Waldmann, Maja Ferencakovic, Gábor Mészáros, Negar Khayatzadeh, Ino Curik, Johann Sölkner |
BMC Bioinform. | 1 |
| 2004 | BAPS 2: enhanced possibilities for the analysis of genetic population structureabstractUNLABELLED: Bayesian statistical methods based on simulation techniques have recently been shown to provide powerful tools for the analysis of genetic population structure. We have previously developed a Markov chain Monte Carlo (MCMC) algorithm for characterizing genetically divergent groups based on molecular markers and geographical sampling design of the dataset. However, for large-scale datasets such algorithms may get stuck to local maxima in the parameter space. Therefore, we have modified our earlier algorithm to support multiple parallel MCMC chains, with enhanced features that enable considerably faster and more reliable estimation compared to the earlier version of the algorithm. We consider also a hierarchical tree representation, from which a Bayesian model-averaged structure estimate can be extracted. The algorithm is implemented in a computer program that features a user-friendly interface and built-in graphics. The enhanced features are illustrated by analyses of simulated data and an extensive human molecular dataset. AVAILABILITY: Freely available at http://www.rni.helsinki.fi/~jic/bapspage.html. Jukka Corander, Patrik Waldmann, Pekka Marttinen, Mikko J. Sillanpää |
Bioinform. | 2 |