VLDB 2026 Research / reviewers in the wild / expert
Maria De Iorio
dblp:07/362
· DBLP profile ↗
7ranked-venue papers
0as first author
1since 2021 · last 2025
0000-0003-3109-0478ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
metabolomics |
1.0 | 2 | 2025 | MetAssimulo 2.0: a web app for simulating realistic 1D and 2D metabolomic 1H NMR spectra · Bioinform. 2025 BATMAN - an R package for the automated quantification of metabolites from nuclear magnetic resonance spectra using a Bayesian model · Bioinform. 2012 |
Bioinformatics and computational biology › metabolomics
metabolite quantification |
0.1 | 1 | 2012 | BATMAN - an R package for the automated quantification of metabolites from nuclear magnetic resonance spectra using a Bayesian model · Bioinform. 2012 |
Bioinformatics and computational biology › structural biology
NMR spectroscopy |
0.1 | 1 | 2012 | BATMAN - an R package for the automated quantification of metabolites from nuclear magnetic resonance spectra using a Bayesian model · Bioinform. 2012 |
Bioinformatics and computational biology › statistical genetics
genetic association study |
0.1 | 1 | 2008 | Bayesian survival analysis in genetic association studies · Bioinform. 2008 |
Bioinformatics and computational biology
statistical genetics |
0.1 | 1 | 2008 | Bayesian survival analysis in genetic association studies · Bioinform. 2008 |
Bioinformatics and computational biology
survival analysis |
0.1 | 1 | 2008 | Bayesian survival analysis in genetic association studies · Bioinform. 2008 |
Methods — techniques the papers use, named apart from their topics
spectral simulation · 0.9correlation modeling · 0.9markov chain monte carlo · 0.1bayesian model · 0.1cox regression · 0.1coalescent model · 0.1bayesian inference · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MetAssimulo 2.0: a web app for simulating realistic 1D and 2D metabolomic 1H NMR spectraabstractMOTIVATION: Metabolomics extensively utilizes nuclear magnetic resonance (NMR) spectroscopy due to its excellent reproducibility and high throughput. Both 1D and 2D NMR spectra provide crucial information for metabolite annotation and quantification, yet present complex overlapping patterns which may require sophisticated machine learning algorithms to decipher. Unfortunately, the limited availability of labeled spectra can hamper application of machine learning, especially deep learning algorithms which require large amounts of labeled data. In this context, simulation of spectral data becomes a tractable solution for algorithm development. RESULTS: Here, we introduce MetAssimulo 2.0, a comprehensive upgrade of the MetAssimulo 1.b metabolomic 1H NMR simulation tool, reimplemented as a Python-based web application. Where MetAssimulo 1.0 only simulated 1D 1H spectra of human urine, MetAssimulo 2.0 expands functionality to urine, blood, and cerebral spinal fluid, enhancing the realism of blood spectra by incorporating a broad protein background. This enhancement enables a closer approximation to real blood spectra, achieving a Pearson correlation of approximately 0.82. Moreover, this tool now includes simulation capabilities for 2D J-resolved (J-Res) and Correlation Spectroscopy spectra, significantly broadening its utility in complex mixture analysis. MetAssimulo 2.0 simulates both single, and groups, of spectra with both discrete (case-control, e.g. heart transplant versus healthy) and continuous (e.g. body mass index) outcomes and includes inter-metabolite correlations. It thus supports a range of experimental designs and demonstrating associations between metabolite profiles and biomedical responses.By enhancing NMR spectral simulations, MetAssimulo 2.0 is well positioned to support and enhance research at the intersection of deep learning and metabolomics. AVAILABILITY AND IMPLEMENTATION: The code and the detailed instruction/tutorial for MetAssimulo 2.0 is available at https://github.com/yanyan5420/MetAssimulo_2.git. The relevant NMR spectra for metabolites are deposited in MetaboLights with accession number MTBLS12081. Beatriz Jiménez, Michael T. Judge, Toby J. Athersuch, Maria De Iorio, Timothy M. D. Ebbels |
Bioinform. | 5 |
| 2012 | BATMAN - an R package for the automated quantification of metabolites from nuclear magnetic resonance spectra using a Bayesian modelabstractMOTIVATION: Nuclear Magnetic Resonance (NMR) spectra are widely used in metabolomics to obtain metabolite profiles in complex biological mixtures. Common methods used to assign and estimate concentrations of metabolites involve either an expert manual peak fitting or extra pre-processing steps, such as peak alignment and binning. Peak fitting is very time consuming and is subject to human error. Conversely, alignment and binning can introduce artefacts and limit immediate biological interpretation of models. RESULTS: We present the Bayesian automated metabolite analyser for NMR spectra (BATMAN), an R package that deconvolutes peaks from one-dimensional NMR spectra, automatically assigns them to specific metabolites from a target list and obtains concentration estimates. The Bayesian model incorporates information on characteristic peak patterns of metabolites and is able to account for shifts in the position of peaks commonly seen in NMR spectra of biological samples. It applies a Markov chain Monte Carlo algorithm to sample from a joint posterior distribution of the model parameters and obtains concentration estimates with reduced error compared with conventional numerical integration and comparable to manual deconvolution by experienced spectroscopists. AVAILABILITY AND IMPLEMENTATION: http://www1.imperial.ac.uk/medicine/people/t.ebbels/ CONTACT: [email protected]. William J. Astle, Maria De Iorio, Timothy M. D. Ebbels |
Bioinform. | 3 |
| 2011 | Significance testing in ridge regression for genetic dataabstractBACKGROUND: Technological developments have increased the feasibility of large scale genetic association studies. Densely typed genetic markers are obtained using SNP arrays, next-generation sequencing technologies and imputation. However, SNPs typed using these methods can be highly correlated due to linkage disequilibrium among them, and standard multiple regression techniques fail with these data sets due to their high dimensionality and correlation structure. There has been increasing interest in using penalised regression in the analysis of high dimensional data. Ridge regression is one such penalised regression technique which does not perform variable selection, instead estimating a regression coefficient for each predictor variable. It is therefore desirable to obtain an estimate of the significance of each ridge regression coefficient. RESULTS: We develop and evaluate a test of significance for ridge regression coefficients. Using simulation studies, we demonstrate that the performance of the test is comparable to that of a permutation test, with the advantage of a much-reduced computational cost. We introduce the p-value trace, a plot of the negative logarithm of the p-values of ridge regression coefficients with increasing shrinkage parameter, which enables the visualisation of the change in p-value of the regression coefficients with increasing penalisation. We apply the proposed method to a lung cancer case-control data set from EPIC, the European Prospective Investigation into Cancer and Nutrition. CONCLUSIONS: The proposed test is a useful alternative to a permutation test for the estimation of the significance of ridge regression coefficients, at a much-reduced computational cost. The p-value trace is an informative graphical tool for evaluating the results of a test of significance of ridge regression coefficients as the shrinkage parameter increases, and the proposed test makes its production computationally feasible. Erika Cule, Paolo Vineis, Maria De Iorio |
BMC Bioinform. | 3 |
| 2010 | MetAssimulo: Simulation of Realistic NMR Metabolic ProfilesabstractBACKGROUND: Probing the complex fusion of genetic and environmental interactions, metabolic profiling (or metabolomics/metabonomics), the study of small molecules involved in metabolic reactions, is a rapidly expanding 'omics' field. A major technique for capturing metabolite data is 1H-NMR spectroscopy and this yields highly complex profiles that require sophisticated statistical analysis methods. However, experimental data is difficult to control and expensive to obtain. Thus data simulation is a productive route to aid algorithm development. RESULTS: MetAssimulo is a MATLAB-based package that has been developed to simulate 1H-NMR spectra of complex mixtures such as metabolic profiles. Drawing data from a metabolite standard spectral database in conjunction with concentration information input by the user or constructed automatically from the Human Metabolome Database, MetAssimulo is able to create realistic metabolic profiles containing large numbers of metabolites with a range of user-defined properties. Current features include the simulation of two groups ('case' and 'control') specified by means and standard deviations of concentrations for each metabolite. The software enables addition of spectral noise with a realistic autocorrelation structure at user controllable levels. A crucial feature of the algorithm is its ability to simulate both intra- and inter-metabolite correlations, the analysis of which is fundamental to many techniques in the field. Further, MetAssimulo is able to simulate shifts in NMR peak positions that result from matrix effects such as pH differences which are often observed in metabolic NMR spectra and pose serious challenges for statistical algorithms. CONCLUSIONS: No other software is currently able to simulate NMR metabolic profiles with such complexity and flexibility. This paper describes the algorithm behind MetAssimulo and demonstrates how it can be used to simulate realistic NMR metabolic profiles with which to develop and test new data analysis techniques. MetAssimulo is freely available for academic use at http://cisbic.bioinformatics.ic.ac.uk/metassimulo/. Harriet J. Muncey, Rebecca Jones, Maria De Iorio, Timothy M. D. Ebbels |
BMC Bioinform. | 3 |
| 2008 | Bayesian survival analysis in genetic association studiesabstractMOTIVATION: Large-scale genetic association studies are carried out with the hope of discovering single nucleotide polymorphisms involved in the etiology of complex diseases. There are several existing methods in the literature for performing this kind of analysis for case-control studies, but less work has been done for prospective cohort studies. We present a Bayesian method for linking markers to censored survival outcome by clustering haplotypes using gene trees. Coalescent-based approaches are promising for LD mapping, as the coalescent offers a good approximation to the evolutionary history of mutations. RESULTS: We compare the performance of the proposed method in simulation studies to the univariate Cox regression and to dimension reduction methods, and we observe that it performs similarly in localizing the causal site, while offering a clear advantage in terms of false positive associations. Moreover, it offers computational advantages. Applying our method to a real prospective study, we observe potential association between candidate ABC transporter genes and epilepsy treatment outcomes. AVAILABILITY: R codes are available upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ioanna Tachmazidou, Toby Andrew, Claudio J. Verzilli, Michael R. Johnson, Maria De Iorio |
Bioinform. | 5 |
| 2008 | Fregene: Simulation of realistic sequence-level data in populations and ascertained samplesabstractBACKGROUND: FREGENE simulates sequence-level data over large genomic regions in large populations. Because, unlike coalescent simulators, it works forwards through time, it allows complex scenarios of selection, demography, and recombination to be modelled simultaneously. Detailed tracking of sites under selection is implemented in FREGENE and provides the opportunity to test theoretical predictions and gain new insights into mechanisms of selection. We describe here main functionalities of both FREGENE and SAMPLE, a companion program that can replicate association study datasets. RESULTS: We report detailed analyses of six large simulated datasets that we have made publicly available. Three demographic scenarios are modelled: one panmictic, one substructured with migration, and one complex scenario that mimics the principle features of genetic variation in major worldwide human populations. For each scenario there is one neutral simulation, and one with a complex pattern of selection. CONCLUSION: FREGENE and the simulated datasets will be valuable for assessing the validity of models for selection, demography and population genetic parameters, as well as the efficacy of association studies. Its principle advantages are modelling flexibility and computational efficiency. It is open source and object-oriented. As such, it can be customised and the range of models extended. Marc Chadeau-Hyam, Clive J. Hoggart, Paul F. O'Reilly, John C. Whittaker, Maria De Iorio, David J. Balding |
BMC Bioinform. | 5 |
| 2008 | An Evolutionary Algorithm to Find Associations in Dense Genetic MapsabstractDiscovering the genetic basis of common human diseases will be assisted by large-scale association studies with a large number of individuals and genetic markers, such as single-nucleotide polymorphisms (SNPs). The potential size of the data and the resulting model space require the development of efficient methodology to unravel associations between epidemiological outcomes and SNPs in dense genetic maps. We apply an evolutionary algorithm (EA) to construct models consisting of logic trees. These trees are Boolean expressions involving nodes that contain strings of SNPs in high linkage disequilibrium (LD), that is, SNPs that are highly correlated with each other. At each generation of the algorithm, a population of logic tree models is modified using selection, crossover, and mutation moves. Logic trees are selected for the next generation using a fitness function based on the marginal likelihood in a Bayesian regression framework. Mutation and crossover moves use LD measures to propose changes to the trees, and facilitate the movement through the model space. We demonstrate our method on data from a candidate gene study of quantitative genetic variation. T. G. Clark, Maria De Iorio, R. C. Griffths |
IEEE Trans. Evol. Comput. | 2 |