VLDB 2026 Research / reviewers in the wild / expert
Roberto Romero
dblp:82/2547
· DBLP profile ↗
12ranked-venue papers
0as first author
0since 2021 · last 2015
0000-0002-4448-5121ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9Artificial intelligence and machine learning · 3Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › systems bioinformatics
pathway analysis |
0.3 | 2 | 2015 | Inter-species pathway perturbation prediction via data-driven detection of functional homology · Bioinform. 2015 A novel signaling pathway impact analysis · Bioinform. 2009 |
Bioinformatics and computational biology
gene expression analysis |
0.3 | 2 | 2013 | Strengths and limitations of microarray-based phenotype prediction: lessons learned from the IMPROVER Diagnostic Signature Challenge · Bioinform. 2013 A novel signaling pathway impact analysis · Bioinform. 2009 |
Bioinformatics and computational biology
comparative genomics |
0.2 | 1 | 2015 | Inter-species pathway perturbation prediction via data-driven detection of functional homology · Bioinform. 2015 |
Bioinformatics and computational biology › gene expression analysis
gene expression prediction |
0.2 | 1 | 2015 | Predicting protein phosphorylation from gene expression: top methods from the IMPROVER Species Translation Challenge · Bioinform. 2015 |
Bioinformatics and computational biology › statistical genetics
phenotype prediction |
0.2 | 1 | 2013 | Strengths and limitations of microarray-based phenotype prediction: lessons learned from the IMPROVER Diagnostic Signature Challenge · Bioinform. 2013 |
Bioinformatics and computational biology › statistical genetics
gene-environment interaction |
0.1 | 1 | 2011 | Varying coefficient model for gene-environment interaction: a non-linear look · Bioinform. 2011 |
Bioinformatics and computational biology › biological database
microarray data management |
0.1 | 1 | 2008 | KUTE-BASE: storing, downloading and exporting MIAME-compliant microarray experiments in minutes rather than hours · Bioinform. 2008 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.0 | 1 | 2013 | Strengths and limitations of microarray-based phenotype prediction: lessons learned from the IMPROVER Diagnostic Signature Challenge · Bioinform. 2013 |
Bioinformatics and computational biology › gene expression analysis
differential expression analysis |
0.0 | 1 | 2009 | A novel signaling pathway impact analysis · Bioinform. 2009 |
Bioinformatics and computational biology › biological database
microarray database |
0.0 | 1 | 2008 | KUTE-BASE: storing, downloading and exporting MIAME-compliant microarray experiments in minutes rather than hours · Bioinform. 2008 |
Methods — techniques the papers use, named apart from their topics
transcription data analysis · 0.2predictive modeling · 0.2pathway enrichment · 0.2machine learning · 0.2feature selection · 0.2data preprocessing · 0.2classifier comparison · 0.2wild bootstrap · 0.1regression spline · 0.1bootstrap · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Predicting protein phosphorylation from gene expression: top methods from the IMPROVER Species Translation ChallengeabstractMOTIVATION: Using gene expression to infer changes in protein phosphorylation levels induced in cells by various stimuli is an outstanding problem. The intra-species protein phosphorylation challenge organized by the IMPROVER consortium provided the framework to identify the best approaches to address this issue. RESULTS: Rat lung epithelial cells were treated with 52 stimuli, and gene expression and phosphorylation levels were measured. Competing teams used gene expression data from 26 stimuli to develop protein phosphorylation prediction models and were ranked based on prediction performance for the remaining 26 stimuli. Three teams were tied in first place in this challenge achieving a balanced accuracy of about 70%, indicating that gene expression is only moderately predictive of protein phosphorylation. In spite of the similar performance, the approaches used by these three teams, described in detail in this article, were different, with the average number of predictor genes per phosphoprotein used by the teams ranging from 3 to 124. However, a significant overlap of gene signatures between teams was observed for the majority of the proteins considered, while Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways were enriched in the union of the predictor genes of the three teams for multiple proteins. AVAILABILITY AND IMPLEMENTATION: Gene expression and protein phosphorylation data are available from ArrayExpress (E-MTAB-2091). Software implementation of the approach of Teams 49 and 75 are available at http://bioinformaticsprb.med.wayne.edu and http://people.cs.clemson.edu/∼luofeng/sbv.rar, respectively. CONTACT: [email protected] or [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Adel Dayarian, Roberto Romero, Michael Biehl, Erhan Bilal, Sahand Hormoz, Pablo Meyer 0001, Raquel Norel, Kahn Rhrissorrakrai, Gyan Bhanot, Feng Luo 0001, Adi L. Tarca |
Bioinform. | 2 |
| 2015 | Inter-species pathway perturbation prediction via data-driven detection of functional homologyabstractMOTIVATION: Experiments in animal models are often conducted to infer how humans will respond to stimuli by assuming that the same biological pathways will be affected in both organisms. The limitations of this assumption were tested in the IMPROVER Species Translation Challenge, where 52 stimuli were applied to both human and rat cells and perturbed pathways were identified. In the Inter-species Pathway Perturbation Prediction sub-challenge, multiple teams proposed methods to use rat transcription data from 26 stimuli to predict human gene set and pathway activity under the same perturbations. Submissions were evaluated using three performance metrics on data from the remaining 26 stimuli. RESULTS: We present two approaches, ranked second in this challenge, that do not rely on sequence-based orthology between rat and human genes to translate pathway perturbation state but instead identify transcriptional response orthologs across a set of training conditions. The translation from rat to human accomplished by these so-called direct methods is not dependent on the particular analysis method used to identify perturbed gene sets. In contrast, machine learning-based methods require performing a pathway analysis initially and then mapping the pathway activity between organisms. Unlike most machine learning approaches, direct methods can be used to predict the activation of a human pathway for a new (test) stimuli, even when that pathway was never activated by a training stimuli. AVAILABILITY: Gene expression data are available from ArrayExpress (accession E-MTAB-2091), while software implementations are available from http://bioinformaticsprb.med.wayne.edu?p=50 and http://goo.gl/hJny3h. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Christoph Hafemeister, Roberto Romero, Erhan Bilal, Pablo Meyer 0001, Raquel Norel, Kahn Rhrissorrakrai, Richard Bonneau, Adi L. Tarca |
Bioinform. | 2 |
| 2013 | Strengths and limitations of microarray-based phenotype prediction: lessons learned from the IMPROVER Diagnostic Signature ChallengeabstractMOTIVATION: After more than a decade since microarrays were used to predict phenotype of biological samples, real-life applications for disease screening and identification of patients who would best benefit from treatment are still emerging. The interest of the scientific community in identifying best approaches to develop such prediction models was reaffirmed in a competition style international collaboration called IMPROVER Diagnostic Signature Challenge whose results we describe herein. RESULTS: Fifty-four teams used public data to develop prediction models in four disease areas including multiple sclerosis, lung cancer, psoriasis and chronic obstructive pulmonary disease, and made predictions on blinded new data that we generated. Teams were scored using three metrics that captured various aspects of the quality of predictions, and best performers were awarded. This article presents the challenge results and introduces to the community the approaches of the best overall three performers, as well as an R package that implements the approach of the best overall team. The analyses of model performance data submitted in the challenge as well as additional simulations that we have performed revealed that (i) the quality of predictions depends more on the disease endpoint than on the particular approaches used in the challenge; (ii) the most important modeling factor (e.g. data preprocessing, feature selection and classifier type) is problem dependent; and (iii) for optimal results datasets and methods have to be carefully matched. Biomedical factors such as the disease severity and confidence in diagnostic were found to be associated with the misclassification rates across the different teams. AVAILABILITY: The lung cancer dataset is available from Gene Expression Omnibus (accession, GSE43580). The maPredictDSC R package implementing the approach of the best overall team is available at www.bioconductor.org or http://bioinformaticsprb.med.wayne.edu/. Adi L. Tarca, Mario Lauria, Erhan Bilal, Stéphanie Boué, Kushal Kumar Dey, Julia Hoeng, Heinz Koeppl, Florian Martin 0002, Pablo Meyer 0001, Preetam Nandy, Raquel Norel, Manuel C. Peitsch, John Jeremy Rice, Roberto Romero, Gustavo Stolovitzky, Marja Talikka, Christoph Zechner |
Bioinform. | 15 |
| 2013 | Z-Bag: a Classification Ensemble System with posterior Probabilistic outputsabstractEnsemble systems improve the generalization of single classifiers by aggregating the prediction of a set of base classifiers. Assessing classification reliability (posterior probability) is crucial in a number of applications, such as biomedical and diagnosis applications, where the cost of a misclassified input vector can be unacceptable high. Available methods are limited to either calibrate the posterior probability on an aggregated decision value or obtain a posterior probability for each base classifier and aggregate the result. We propose a method that takes advantage of the distribution of the decision values from the base classifiers to summarize a statistic which is subsequently used to generate the posterior probability. Three approaches are considered to fit the probabilistic output to the statistic: the standard Gaussian CDF, isotonic regression, and linear logistic. Even though this study focuses on a bagged support vector machine ensemble (Z‐bag), our approach is not limited by the aggregation method selected, the choice of base classifiers, nor the statistic used. Performance is assessed on one artificial and 12 real‐world data sets from the UCI Machine Learning Repository. Our approach achieves comparable or better generalization on accuracy and posterior estimation to existing ensemble calibration methods although lowering computational cost. Zhonghui Xu, Calin Voichita, Sorin Draghici, Roberto Romero |
Comput. Intell. | 4 |
| 2012 | A method for analysis and correction of cross-talk effects in pathway analysisabstractMany analysis techniques are currently available to identify the signaling pathways significantly impacted in a given condition. All these approaches calculate a p-value that aims to quantify the significance of the involvement of a given pathway in the condition under study. These p-values were thought to be related to the likelihood of their respective pathways being involved in the given condition, and to be independent. Here we show that this is not true, and that many pathways are not independent and that can considerably affect each other's p-values through a phenomenon we refer to as “cross-talk.” Thus, the significance of a given pathway in a given experiment has to be interpreted in the context of the other pathways that appear to be significant. Using real data, we show that in same cases pathways with significant classical p-values are not biologically meaningful, and that some biologically meaningful pathways with insignificant p-values become significant when the cross-talk effects of other pathways are removed. We show that this phenomenon is related to the amount of common genes between different pathways, affecting the most widely used methods for pathway analysis, and we propose an analysis technique that is able to correct the over-enrichment significance of a pathway when the cross-talk effects of other pathways are removed. Michele Donato, Sorin Draghici, Alin Tomoiaga, Peter Westfall, Zhonghui Xu, Roberto Romero |
IJCNN | 6 |
| 2012 | Down-weighting overlapping genes improves gene set analysisabstractBACKGROUND: The identification of gene sets that are significantly impacted in a given condition based on microarray data is a crucial step in current life science research. Most gene set analysis methods treat genes equally, regardless how specific they are to a given gene set. RESULTS: In this work we propose a new gene set analysis method that computes a gene set score as the mean of absolute values of weighted moderated gene t-scores. The gene weights are designed to emphasize the genes appearing in few gene sets, versus genes that appear in many gene sets. We demonstrate the usefulness of the method when analyzing gene sets that correspond to the KEGG pathways, and hence we called our method Pathway Analysis with Down-weighting of Overlapping Genes (PADOG). Unlike most gene set analysis methods which are validated through the analysis of 2-3 data sets followed by a human interpretation of the results, the validation employed here uses 24 different data sets and a completely objective assessment scheme that makes minimal assumptions and eliminates the need for possibly biased human assessments of the analysis results. CONCLUSIONS: PADOG significantly improves gene set ranking and boosts sensitivity of analysis using information already available in the gene expression profiles and the collection of gene sets to be analyzed. The advantages of PADOG over other existing approaches are shown to be stable to changes in the database of gene sets to be analyzed. PADOG was implemented as an R package available at: http://bioinformaticsprb.med.wayne.edu/PADOG/or http://www.bioconductor.org. Adi L. Tarca, Sorin Draghici, Gaurav Bhatti, Roberto Romero |
BMC Bioinform. | 4 |
| 2011 | Varying coefficient model for gene-environment interaction: a non-linear lookabstractMOTIVATION: The genetic basis of complex traits often involves the function of multiple genetic factors, their interactions and the interaction between the genetic and environmental factors. Gene-environment (G×E) interaction is considered pivotal in determining trait variations and susceptibility of many genetic disorders such as neurodegenerative diseases or mental disorders. Regression-based methods assuming a linear relationship between a disease response and the genetic and environmental factors as well as their interaction is the commonly used approach in detecting G×E interaction. The linearity assumption, however, could be easily violated due to non-linear genetic penetrance which induces non-linear G×E interaction. RESULTS: In this work, we propose to relax the linear G×E assumption and allow for non-linear G×E interaction under a varying coefficient model framework. We propose to estimate the varying coefficients with regression spline technique. The model allows one to assess the non-linear penetrance of a genetic variant under different environmental stimuli, therefore help us to gain novel insights into the etiology of a complex disease. Several statistical tests are proposed for a complete dissection of G×E interaction. A wild bootstrap method is adopted to assess the statistical significance. Both simulation and real data analysis demonstrate the power and utility of the proposed method. Our method provides a powerful and testable framework for assessing non-linear G×E interaction. Shujie Ma, Roberto Romero, Yuehua Cui |
Bioinform. | 3 |
| 2009 | A novel signaling pathway impact analysisabstractMOTIVATION: Gene expression class comparison studies may identify hundreds or thousands of genes as differentially expressed (DE) between sample groups. Gaining biological insight from the result of such experiments can be approached, for instance, by identifying the signaling pathways impacted by the observed changes. Most of the existing pathway analysis methods focus on either the number of DE genes observed in a given pathway (enrichment analysis methods), or on the correlation between the pathway genes and the class of the samples (functional class scoring methods). Both approaches treat the pathways as simple sets of genes, disregarding the complex gene interactions that these pathways are built to describe. RESULTS: We describe a novel signaling pathway impact analysis (SPIA) that combines the evidence obtained from the classical enrichment analysis with a novel type of evidence, which measures the actual perturbation on a given pathway under a given condition. A bootstrap procedure is used to assess the significance of the observed total pathway perturbation. Using simulations we show that the evidence derived from perturbations is independent of the pathway enrichment evidence. This allows us to calculate a global pathway significance P-value, which combines the enrichment and perturbation P-values. We illustrate the capabilities of the novel method on four real datasets. The results obtained on these data show that SPIA has better specificity and more sensitivity than several widely used pathway analysis methods. AVAILABILITY: SPIA was implemented as an R package available at http://vortex.cs.wayne.edu/ontoexpress/ Adi L. Tarca, Sorin Draghici, Purvesh Khatri, Sonia S. Hassan, Pooja Mittal, Jung-Sun Kim, Chong Jai Kim, Juan Pedro Kusanovic, Roberto Romero |
Bioinform. | 9 |
| 2008 | KUTE-BASE: storing, downloading and exporting MIAME-compliant microarray experiments in minutes rather than hoursabstractMOTIVATION: The BioArray Software Environment (BASE) is a very popular MIAME-compliant, web-based microarray data repository. However in BASE, like in most other microarray data repositories, the experiment annotation and raw data uploading can be very timeconsuming, especially for large microarray experiments. RESULTS: We developed KUTE (Karmanos Universal daTabase for microarray Experiments), as a plug-in for BASE 2.0 that addresses these issues. KUTE provides an automatic experiment annotation feature and a completely redesigned data work-flow that dramatically reduce the human-computer interaction time. For instance, in BASE 2.0 a typical Affymetrix experiment involving 100 arrays required 4 h 30 min of user interaction time forexperiment annotation, and 45 min for data upload/download. In contrast, for the same experiment, KUTE required only 28 min of user interaction time for experiment annotation, and 3.3 min for data upload/download. AVAILABILITY: http://vortex.cs.wayne.edu/kute/index.html. Sorin Draghici, Adi L. Tarca, Stephen Ethier, Roberto Romero |
Bioinform. | 5 |
| 2007 | A System Biology Approach for the Steady-State Analysis of Gene Signaling Networks
Purvesh Khatri, Sorin Draghici, Adi L. Tarca, Sonia S. Hassan, Roberto Romero |
CIARP | 5 |
| 2007 | Machine Learning and Its Applications to BiologyabstractThe term machine learning refers to a set of topics dealing with the creation and evaluation of algorithms that facilitate pattern recognition, classification, and prediction, based on models derived from existing data. Two facets of mechanization should be acknowledged when considering machine learning in broad terms. Firstly, it is intended that the classification and prediction tasks can be accomplished by a suitably programmed computing machine. That is, the product of machine learning is a classifier that can be feasibly used on available hardware. Secondly, it is intended that the creation of the classifier should itself be highly mechanized, and should not involve too much human input. This second facet is inevitably vague, but the basic objective is that the use of automatic algorithm construction methods can minimize the possibility that human biases could affect the selection and performance of the algorithm. Both the creation of the algorithm and its operation to classify objects or predict events are to be based on concrete, observable data.
The history of relations between biology and the field of machine learning is long and complex. An early technique [1] for machine learning called the perceptron constituted an attempt to model actual neuronal behavior, and the field of artificial neural network (ANN) design emerged from this attempt. Early work on the analysis of translation initiation sequences [2] employed the perceptron to define criteria for start sites in Escherichia coli. Further artificial neural network architectures such as the adaptive resonance theory (ART) [3] and neocognitron [4] were inspired from the organization of the visual nervous system. In the intervening years, the flexibility of machine learning techniques has grown along with mathematical frameworks for measuring their reliability, and it is natural to hope that machine learning methods will improve the efficiency of discovery and understanding in the mounting volume and complexity of biological data.
This tutorial is structured in four main components. Firstly, a brief section reviews definitions and mathematical prerequisites. Secondly, the field of supervised learning is described. Thirdly, methods of unsupervised learning are reviewed. Finally, a section reviews methods and examples as implemented in the open source data analysis and visualization language R (http://www.r-project.org). Adi L. Tarca, Vincent Carey, Xue-wen Chen 0001, Roberto Romero, Sorin Draghici |
PLoS Comput. Biol. | 4 |
| 2006 | Preparing Proprietary Systems for Continuous Education e-Learning for Inter-Operability: Exporting to SCORMabstractThis document describes the procedure followed to prepare an existing framework for e-learning, to make the contents, produced during many years, exportable to a no proprietary standard (SCORM). Victor Manso, Roberto Romero, Carlos Palau |
ICALT | 2 |