Pedro J. Ballester

dblp:85/4533 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-4078-743XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 6 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Conformal prediction of molecule-induced cancer cell growth inhibition challenged by strong distribution shifts
Saiveth Hernández-Hernández, Qianrong Guo, Pedro J. Ballester
Pattern Recognit.3
2025 Deconstructing biomarker generalisation failure in cancer anti-pd1 immunotherapy: toward adaptive patient stratification
abstract
Abstract Introduction Predictive biomarkers are central to precision oncology [1], yet transcriptomic signatures for predicting response to Anti-PD1 immunotherapy rarely generalise beyond their discovery cohort. This persistent failure, often attributed to dataset shift, remains poorly understand and hampers clinical translation. We sought to move beyond simple validation and systematically examine the causes of generalisation failure in melanoma and renal cell carcinoma, hypothesising that the brittleness of fitted models is a key barrier. Methods We developed 20 gene-panel biomarkers using a consensus [2] pipeline (BioAdapt) across four melanoma and renal cell carcinoma cohorts (N=360). To rigorously test generalisability, each cohort served as both training and independent test data, producing 75 cross-cohort validation experiments. Our framework comprised three stages: (1) diagnosis of dataset dissimilarity using the Maximum Mean Discrepancy test, (2) technical correction via label-free normalisation, and (3) biological stratification of patients by embedding similarity to training cohorts using UMAP. Results Direct transfer of models across cohorts yielded near-random performance (mean Matthews Correlation Coefficient ≈ 0). Statistical testing confirmed severe dataset shift (mean Maximum Mean Discrepancy = 0.584). Technical correction modestly improved performance but remained weak. In contrast, biological stratification identified patient subgroups where predictive accuracy nearly doubled compared to corrected models alone, with best-case MCC = 0.41. Discussion & Conclusion Our results reveal that biomarker failure is driven less by batch effects than by inherent model brittleness: fitted weights are highly local to the discovery cohort. While technical correction is necessary, stratifying patients by biological similarity enables meaningful recovery of predictive signal. This study establishes a systematic framework for quantifying biomarker limitations and highlights adaptive stratification as a promising route toward subgroup-specific translation in cancer immunotherapy. References Piyawajanusorn C., Ghislat G., Ballester P.J. “Predicting atezolizumab response in metastatic urothelial carcinoma patients using machine learning on integrated tumour gene expression and clinical data.” NPJ Precis Oncol. 2025;9(1):170. Abeel T., Helleputte T., Van de Peer Y., Dupont P., Saeys Y. “Robust biomarker identification for cancer diagnosis with ensemble feature selection methods.” Bioinformatics. 2010;26(3):392–398.
Elizabeth Amelia, Chayanit Piyawajanusorn, Pedro J. Ballester
Briefings Bioinform.3
2025 Deconstructing biomarker generalisation failure in anti-pd1 cancer immunotherapy response: a multi-cohort framework
abstract
Abstract Introduction Predictive biomarkers are central to precision oncology [1], yet transcriptomic signatures for predicting response to Anti-PD1 immunotherapy rarely generalise beyond their discovery cohort. This persistent failure, often attributed to dataset shift, remains poorly understand and hampers clinical translation. We sought to move beyond simple validation and systematically examine the causes of generalisation failure in melanoma and renal cell carcinoma, hypothesising that the brittleness of fitted models is a key barrier. Methods We assembled five independent transcriptomic cohorts of melanoma and renal cell carcinoma patients treated with Anti-PD1 therapy. Four cohorts (N = 360) were designated as discovery sets from which we generated 20 distinct gene-panel biomarkers using our BioAdapt pipeline. This consensus framework [2] repeatedly partitions data with different random seeds, applies feature selection, and fits logistic regression models with optimised thresholds, ensuring variability is captured rather than relying on a single split. For validation, each model trained on one cohort was systematically applied to the remaining independent datasets, producing 75 fully cross-cohort experiments where no model was ever tested on its own training data. To understand generalisation failure, we designed a three-stage framework: (1) Diagnosis of dataset dissimilarity using the Maximum Mean Discrepancy test (a statistical test of distributional shift), (2) Technical correction with label-free quantile normalisation, and (3) Biological stratification with UMAP to identify subgroups most similar to the training cohort. Results Direct model transfer failed, yielding near-random performance (mean MCC 0.016). For clinicians, this means that a biomarker trained on one hospital’s patients is very unlikely to work in another without adaptation. Diagnostic testing confirmed pervasive dataset shift (mean Maximum Mean Discrepancy 0.584). Correction produced only modest gains (mean corrected MCC 0.026). However, stratification revealed patient subgroups where performance nearly doubled (best-case MCC 0.41), demonstrating that predictive utility is often local to biologically defined subsets. In practice, this suggests biomarkers may still be valuable if applied to the right subset of patients, rather than universally. Discussion & Conclusion Our analysis shows that generalisation failure stems less from technical batch effects than from the intrinsic brittleness of fitted models. While correction is necessary, stratifying patients by biological similarity can recover predictive signal in defined contexts. These findings quantify the limits of current biomarker models and highlight the need for inherently robust, transferable approaches such as transfer learning or domain adaptation to achieve clinical translation. Importantly, this work provides a realistic roadmap for the field by showing where predictive biomarkers are most likely to succeed and in which patient populations they can be deployed with confidence. By reframing validation from a simple pass or fail exercise to a deeper diagnostic process, our framework offers a translationally relevant pathway towards biomarker deployment in clinical trials, ultimately enabling the stratification of patients who are most likely to respond to Anti-PD1 therapy and accelerating precision immuno-oncology. For clinicians, this means moving towards a future where biomarker-guided patient selection for Anti-PD1 treatment is both feasible and reliable, reducing the risk of exposing non-responders to ineffective therapies. References 1. Piyawajanusorn C., Ghislat G., Ballester P.J. ‘Predicting atezolizumab response in metastatic urothelial carcinoma patients using machine learning on integrated tumour gene expression and clinical data.’ NPJ Precis Oncol. 2025;9(1):170. 2. Abeel T., Helleputte T., Van de Peer Y., Dupont P., Saeys Y. ‘Robust biomarker identification for cancer diagnosis with ensemble feature selection methods.’ Bioinformatics. 2010;26(3):392–398.
Elizabeth Amelia, Chayanit Piyawajanusorn, Pedro J. Ballester
Briefings Bioinform.3
2025 Target-specific machine-learning scoring functions enhance virtual screening for Pf DHODH inhibitors
abstract
Abstract Aims To develop and evaluate target-specific machine learning (ML) scoring functions (SFs) for Plasmodium falciparum dihydroorotate dehydrogenase (PfDHODH), a validated antimalarial drug target. Based on evidence that target-specific ML SFs outperform generic ones [1], we aimed to create PfDHODH-tailored models to improve virtual screening and prioritise potent inhibitors. Methods Following established protocols for target-specific ML SF development [2], bioactivity data were gathered from PubChem and ChEMBL, augmented with property-matched decoys, split into training and test sets and docked into the ubiquinone pocket of PfDHODH using Smina. Features included protein-ligand extended connectivity (PLEC) fingerprints, Morgan fingerprints, Smina terms, and physicochemical descriptors. Random forest (RF), extreme gradient boosting, support vector machines, and neural network models were trained in regression and classification modes with optimised hyperparameters, then compared with generic SFs and the PfDHODH-specific EHIGN [3]. Results and conclusions Regression models outperformed classifiers, with RF using PLEC + Morgan fingerprints achieving the best performance (EF1% = 52.5, NEF1% = 0.617). Among the ten top-performing PfDHODH-specific ML SFs (best variants of algorithms with optimal feature sets), 8 outperformed Smina and 9 outperformed EHIGN. The best SF doubled virtual screening efficiency over generic methods, underscoring the power of ML-driven antimalarial discovery. References 1. Caba K., Tran-Nguyen VK., Rahman T., Ballester P.J. ‘Comprehensive machine learning boosts structure-based virtual screening for PARP1 inhibitors.’ J Cheminform. 2024;16. 2. Tran-Nguyen V.K., Junaid M., Simeon S., Ballester P.J. ‘A practical guide to machine-learning scoring for structure-based virtual screening.’ Nat Protoc. 2023;18:3460–511. 3. Yang Z., Zhong W., Lv Q., Dong T., Chen G., Chen C.YC. ‘Interaction-Based Inductive Bias in Graph Neural Networks: Enhancing Protein-Ligand Binding Affinity Predictions From 3D Structures.’ IEEE Trans Pattern Anal Mach Intell. 2024;46:8191–208.
Klaudia Caba, Pedro J. Ballester
Briefings Bioinform.2
2025 Comprehensive machine learning boosts structure-based virtual screening for PARP1 inhibitors
Klaudia Caba, Viet-Khoa Tran-Nguyen, Taufiqur Rahman, Pedro J. Ballester
Briefings Bioinform.4
2024 Scaffold Splits Overestimate Virtual Screening Performance
Qianrong Guo, Saiveth Hernández-Hernández, Pedro J. Ballester
ICANN (10)3
2021 The impact of compound library size on the performance of scoring functions for structure-based virtual screening
abstract
Larger training datasets have been shown to improve the accuracy of machine learning (ML)-based scoring functions (SFs) for structure-based virtual screening (SBVS). In addition, massive test sets for SBVS, known as ultra-large compound libraries, have been demonstrated to enable the fast discovery of selective drug leads with low-nanomolar potency. This proof-of-concept was carried out on two targets using a single docking tool along with its SF. It is thus unclear whether this high level of performance would generalise to other targets, docking tools and SFs. We found that screening a larger compound library results in more potent actives being identified in all six additional targets using a different docking tool along with its classical SF. Furthermore, we established that a way to improve the potency of the retrieved molecules further is to rank them with more accurate ML-based SFs (we found this to be true in four of the six targets; the difference was not significant in the remaining two targets). A 3-fold increase in average hit rate across targets was also achieved by the ML-based SFs. Lastly, we observed that classical and ML-based SFs often find different actives, which supports using both types of SFs on those targets.
Louison Fresnais, Pedro J. Ballester
Briefings Bioinform.2
2021 A gentle introduction to understanding preclinical data for cancer pharmaco-omic modeling
abstract
A central goal of precision oncology is to administer an optimal drug treatment to each cancer patient. A common preclinical approach to tackle this problem has been to characterize the tumors of patients at the molecular and drug response levels, and employ the resulting datasets for predictive in silico modeling (mostly using machine learning). Understanding how and why the different variants of these datasets are generated is an important component of this process. This review focuses on providing such introduction aimed at scientists with little previous exposure to this research area.
Chayanit Piyawajanusorn, Linh C. Nguyen, Ghita Ghislat, Pedro J. Ballester
Briefings Bioinform.4
2019 Classical scoring functions for docking are unable to exploit large volumes of structural and interaction data
abstract
MOTIVATION: Studies have shown that the accuracy of random forest (RF)-based scoring functions (SFs), such as RF-Score-v3, increases with more training samples, whereas that of classical SFs, such as X-Score, does not. Nevertheless, the impact of the similarity between training and test samples on this matter has not been studied in a systematic manner. It is therefore unclear how these SFs would perform when only trained on protein-ligand complexes that are highly dissimilar or highly similar to the test set. It is also unclear whether SFs based on machine learning algorithms other than RF can also improve accuracy with increasing training set size and to what extent they learn from dissimilar or similar training complexes. RESULTS: We present a systematic study to investigate how the accuracy of classical and machine-learning SFs varies with protein-ligand complex similarities between training and test sets. We considered three types of similarity metrics, based on the comparison of either protein structures, protein sequences or ligand structures. Regardless of the similarity metric, we found that incorporating a larger proportion of similar complexes to the training set did not make classical SFs more accurate. In contrast, RF-Score-v3 was able to outperform X-Score even when trained on just 32% of the most dissimilar complexes, showing that its superior performance owes considerably to learning from dissimilar training complexes to those in the test set. In addition, we generated the first SF employing Extreme Gradient Boosting (XGBoost), XGB-Score, and observed that it also improves with training set size while outperforming the rest of SFs. Given the continuous growth of training datasets, the development of machine-learning SFs has become very appealing. AVAILABILITY AND IMPLEMENTATION: https://github.com/HongjianLi/MLSF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiangjun Peng, Pavel Sidorov, Yee Leung, Kwong-Sak Leung, Man Hon Wong 0001, Pedro J. Ballester
Bioinform.8
2018 A Stochastic Spiking Neural Network for Virtual Screening
abstract
Virtual screening (VS) has become a key computational tool in early drug design and screening performance is of high relevance due to the large volume of data that must be processed to identify molecules with the sought activity-related pattern. At the same time, the hardware implementations of spiking neural networks (SNNs) arise as an emerging computing technique that can be applied to parallelize processes that normally present a high cost in terms of computing time and power. Consequently, SNN represents an attractive alternative to perform time-consuming processing tasks, such as VS. In this brief, we present a smart stochastic spiking neural architecture that implements the ultrafast shape recognition (USR) algorithm achieving two order of magnitude of speed improvement with respect to USR software implementations. The neural system is implemented in hardware using field-programmable gate arrays allowing a highly parallelized USR implementation. The results show that, due to the high parallelization of the system, millions of compounds can be checked in reasonable times. From these results, we can state that the proposed architecture arises as a feasible methodology to efficiently enhance time-consuming data-mining processes such as 3-D molecular similarity search.
Antoni Morro, Vicent Canals, Antoni Oliver 0002, Miquel L. Alomar, Fabio Galán-Prado, Pedro J. Ballester, José Luis Rosselló
IEEE Trans. Neural Networks Learn. Syst.6
2016 Correcting the impact of docking pose generation error on binding affinity prediction
abstract
BACKGROUND: Pose generation error is usually quantified as the difference between the geometry of the pose generated by the docking software and that of the same molecule co-crystallised with the considered protein. Surprisingly, the impact of this error on binding affinity prediction is yet to be systematically analysed across diverse protein-ligand complexes. RESULTS: Against commonly-held views, we have found that pose generation error has generally a small impact on the accuracy of binding affinity prediction. This is also true for large pose generation errors and it is not only observed with machine-learning scoring functions, but also with classical scoring functions such as AutoDock Vina. Furthermore, we propose a procedure to correct a substantial part of this error which consists of calibrating the scoring functions with re-docked, rather than co-crystallised, poses. In this way, the relationship between Vina-generated protein-ligand poses and their binding affinities is directly learned. As a result, test set performance after this error-correcting procedure is much closer to that of predicting the binding affinity in the absence of pose generation error (i.e. on crystal structures). We evaluated several strategies, obtaining better results for those using a single docked pose per ligand than those using multiple docked poses per ligand. CONCLUSIONS: Binding affinity prediction is often carried out on the docked pose of a known binder rather than its co-crystallised pose. Our results suggest than pose generation error is in general far less damaging for binding affinity prediction than it is currently believed. Another contribution of our study is the proposal of a procedure that largely corrects for this error. The resulting machine-learning scoring function is freely available at http://istar.cse.cuhk.edu.hk/rf-score-4.tgz and http://ballester.marseille.inserm.fr/rf-score-4.tgz .
Kwong-Sak Leung, Man Hon Wong 0001, Pedro J. Ballester
BMC Bioinform.4
2014 Substituting random forest for multiple linear regression improves binding affinity prediction of scoring functions: Cyscore as a case study
abstract
BACKGROUND: State-of-the-art protein-ligand docking methods are generally limited by the traditionally low accuracy of their scoring functions, which are used to predict binding affinity and thus vital for discriminating between active and inactive compounds. Despite intensive research over the years, classical scoring functions have reached a plateau in their predictive performance. These assume a predetermined additive functional form for some sophisticated numerical features, and use standard multivariate linear regression (MLR) on experimental data to derive the coefficients. RESULTS: In this study we show that such a simple functional form is detrimental for the prediction performance of a scoring function, and replacing linear regression by machine learning techniques like random forest (RF) can improve prediction performance. We investigate the conditions of applying RF under various contexts and find that given sufficient training samples RF manages to comprehensively capture the non-linearity between structural features and measured binding affinities. Incorporating more structural features and training with more samples can both boost RF performance. In addition, we analyze the importance of structural features to binding affinity prediction using the RF variable importance tool. Lastly, we use Cyscore, a top performing empirical scoring function, as a baseline for comparison study. CONCLUSIONS: Machine-learning scoring functions are fundamentally different from classical scoring functions because the former circumvents the fixed functional form relating structural features with binding affinities. RF, but not MLR, can effectively exploit more structural features and more training samples, leading to higher prediction performance. The future availability of more X-ray crystal structures will further widen the performance gap between RF-based and MLR-based scoring functions. This further stresses the importance of substituting RF for MLR in scoring function development.
Kwong-Sak Leung, Man Hon Wong 0001, Pedro J. Ballester
BMC Bioinform.4
2010 A machine learning approach to predicting protein-ligand binding affinity with applications to molecular docking
abstract
MOTIVATION: Accurately predicting the binding affinities of large sets of diverse protein-ligand complexes is an extremely challenging task. The scoring functions that attempt such computational prediction are essential for analysing the outputs of molecular docking, which in turn is an important technique for drug discovery, chemical biology and structural biology. Each scoring function assumes a predetermined theory-inspired functional form for the relationship between the variables that characterize the complex, which also include parameters fitted to experimental or simulation data and its predicted binding affinity. The inherent problem of this rigid approach is that it leads to poor predictivity for those complexes that do not conform to the modelling assumptions. Moreover, resampling strategies, such as cross-validation or bootstrapping, are still not systematically used to guard against the overfitting of calibration data in parameter estimation for scoring functions. RESULTS: We propose a novel scoring function (RF-Score) that circumvents the need for problematic modelling assumptions via non-parametric machine learning. In particular, Random Forest was used to implicitly capture binding effects that are hard to model explicitly. RF-Score is compared with the state of the art on the demanding PDBbind benchmark. Results show that RF-Score is a very competitive scoring function. Importantly, RF-Score's performance was shown to improve dramatically with training set size and hence the future availability of more high-quality structural and interaction data is expected to lead to improved versions of RF-Score. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pedro J. Ballester, John B. O. Mitchell
Bioinform.1
2007 Model calibration of a real petroleum reservoir using a parallel real-coded genetic algorithm
abstract
An application of a Real-coded Genetic Algorithm (GA) to the model calibration of a real petroleum reservoir is presented. In order to shorten the computation time, the possible solutions generated by the GA are evaluated in parallel on a group of computers. This required the GA to be adapted to a multi-processor structure, so that the scalability of the computation is maximised. The best solutions of each run enter the ensemble of calibrated models, which is finally analysed using a clustering algorithm. The aim is to identify the optimal regions contained in the ensemble and thus to reveal the distinct types of reservoir models consistent with the historic production data, as a way to assess the uncertainty in the Reservoir Characterisation due to the limited reliability of optimisation algorithms. The developed methodology is applied to the characterisation of a real petroleum reservoir. Results show a large improvement with respect to previous studies on that reservoir in terms of the quality and diversity of the obtained calibrated models. Our main conclusion is that, even with regularisation, many distinct calibrated models are possible, which highlights the importance of applying optimisation methods capable of identifying all such solutions.
Pedro J. Ballester, Jonathan N. Carter
IEEE Congress on Evolutionary Computation1
2006 A Multiparent Version of the Parent-Centric Normal Crossover for Multimodal Optimization
abstract
A new multiparent parent-centric crossover (mPNX), which is a development of a previous two-parent operator called the Parent-centric Normal crossover (PNX), is presented. Both crossovers are studied in combination with the SPC population model. The resulting Genetic Algorithms (GAs) are tested on a benchmark of particularly hard, nonseparable, shifted, multimodal analytical optimisation problems. This benchmark contains notoriously difficult test problems, which have been already attempted by a number of high quality optimisation methods with limited success. It is shown that the multiparent GA (SPC-mPNX) generally results in significant improvements with respect to the two-parent GA (SPC-PNX), while solving for the first time one of the used test problems.
Pedro J. Ballester, W. Graham Richards
IEEE Congress on Evolutionary Computation1
2005 Real-parameter optimization performance study on the CEC-2005 benchmark with SPC-PNX
abstract
This paper presents a performance study of a real-parameter genetic algorithm (SPC-PNX) on a new benchmark of real-parameter optimisation problems. This benchmark provides a systematic way to compare different optimisation methods on exactly the same test problems. These problems were designed to be hard as they incorporate features that have been shown to pose great difficulty to many optimisation methods.
Pedro J. Ballester, John Stephenson, Jonathan N. Carter, Kerry Gallagher
Congress on Evolutionary Computation1
2004 An Effective Real-Parameter Genetic Algorithm with Parent Centric Normal Crossover for Multimodal Optimisation
Pedro J. Ballester, Jonathan N. Carter
GECCO (1)1
2004 Tackling an Inverse Problem from the Petroleum Industry with a Genetic Algorithm for Sampling
Pedro J. Ballester, Jonathan N. Carter
GECCO (2)1
2003 Real-Parameter Genetic Algorithms for Finding Multiple Optimal Solutions in Multi-modal Optimization
Pedro J. Ballester, Jonathan N. Carter
GECCO1