VLDB 2026 Research / reviewers in the wild / expert
Arnoldo Frigessi
dblp:32/6861
· DBLP profile ↗
17ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0001-7103-7589ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 4 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TVineSynth: A Truncated C-Vine Copula Generator of Synthetic Tabular Data to Balance Privacy and UtilityabstractWe propose TVineSynth, a vine copula based synthetic tabular data generator, which is designed to balance privacy and utility, using the vine tree structure and its truncation to do the trade-off. Contrary to synthetic data generators that achieve DP by globally adding noise, TVineSynth performs a controlled approximation of the estimated data generating distribution, so that it does not suffer from poor utility of the resulting synthetic data for downstream prediction tasks. TVineSynth introduces a targeted bias into the vine copula model that, combined with the specific tree structure of the vine, causes the model to zero out privacy-leaking dependencies while relying on those that are beneficial for utility. Privacy is here measured with membership (MIA) and attribute inference attacks (AIA). Further, we theoretically justify how the construction of TVineSynth ensures AIA privacy under a natural privacy measure for continuous sensitive attributes. When compared to competitor models, with and without DP, on simulated and on real-world data, TVineSynth achieves a superior privacy-utility balance. Elisabeth Griesbauer, Claudia Czado, Arnoldo Frigessi, Ingrid Hobæk Haff |
AISTATS | 3 |
| 2024 | Modeling geographic vaccination strategies for COVID-19 in NorwayabstractVaccination was a key intervention in controlling the COVID-19 pandemic globally. In early 2021, Norway faced significant regional variations in COVID-19 incidence and prevalence, with large differences in population density, necessitating efficient vaccine allocation to reduce infections and severe outcomes. This study explored alternative vaccination strategies to minimize health outcomes (infections, hospitalizations, ICU admissions, deaths) by varying regions prioritized, extra doses prioritized, and implementation start time. Using two models (individual-based and meta-population), we simulated COVID-19 transmission during the primary vaccination period in Norway, covering the first 7 months of 2021. We investigated alternative strategies to allocate more vaccine doses to regions with a higher force of infection. We also examined the robustness of our results and highlighted potential structural differences between the two models. Our findings suggest that early vaccine prioritization could reduce COVID-19 related health outcomes by 8% to 20% compared to a baseline strategy without geographic prioritization. For minimizing infections, hospitalizations, or ICU admissions, the best strategy was to initially allocate all available vaccine doses to fewer high-risk municipalities, comprising approximately one-fourth of the population. For minimizing deaths, a moderate level of geographic prioritization, with approximately one-third of the population receiving doubled doses, gave the best outcomes by balancing the trade-off between vaccinating younger people in high-risk areas and older people in low-risk areas. The actual strategy implemented in Norway was a two-step moderate level aimed at maintaining the balance and ensuring ethical considerations and public trust. However, it did not offer significant advantages over the baseline strategy without geographic prioritization. Earlier implementation of geographic prioritization could have more effectively addressed the main wave of infections, substantially reducing the national burden of the pandemic. Louis Yat Hin Chan, Gunnar Rø, Jørgen Eriksson Midtbø, Francesco Di Ruscio, Sara Sofie Viksmoen Watle, Lene Kristine Juvet, Jasper Littmann, Preben Aavitsland, Karin Maria Nygård, Are Stuwitz Berg, Geir Bukholm, Anja Bråthen Kristoffersen, Kenth Engø-Monsen, Solveig Engebretsen, David Swanson, Alfonso Diz-Lois Palomares, Jonas Christoffer Lindstrøm, Arnoldo Frigessi, Birgitte Freiesleben de Blasio |
PLoS Comput. Biol. | 18 |
| 2024 | Using birth-death processes to infer tumor subpopulation structure from live-cell imaging drug screening dataabstractTumor heterogeneity is a complex and widely recognized trait that poses significant challenges in developing effective cancer therapies. In particular, many tumors harbor a variety of subpopulations with distinct therapeutic response characteristics. Characterizing this heterogeneity by determining the subpopulation structure within a tumor enables more precise and successful treatment strategies. In our prior work, we developed PhenoPop, a computational framework for unravelling the drug-response subpopulation structure within a tumor from bulk high-throughput drug screening data. However, the deterministic nature of the underlying models driving PhenoPop restricts the model fit and the information it can extract from the data. As an advancement, we propose a stochastic model based on the linear birth-death process to address this limitation. Our model can formulate a dynamic variance along the horizon of the experiment so that the model uses more information from the data to provide a more robust estimation. In addition, the newly proposed model can be readily adapted to situations where the experimental data exhibits a positive time correlation. We test our model on simulated data (in silico) and experimental data (in vitro), which supports our argument about its advantages. Einar Bjarki Gunnarsson, Even Moa Myklebust, Alvaro Köhn-Luque, Dagim Shiferaw Tadele, Jorrit Martijn Enserink, Arnoldo Frigessi, Jasmine Foo, Kevin Leder |
PLoS Comput. Biol. | 7 |
| 2023 | A real-time regional model for COVID-19: Probabilistic situational awareness and forecastingabstractThe COVID-19 pandemic is challenging nations with devastating health and economic consequences. The spread of the disease has revealed major geographical heterogeneity because of regionally varying individual behaviour and mobility patterns, unequal meteorological conditions, diverse viral variants, and locally implemented non-pharmaceutical interventions and vaccination roll-out. To support national and regional authorities in surveilling and controlling the pandemic in real-time as it unfolds, we here develop a new regional mathematical and statistical model. The model, which has been in use in Norway during the first two years of the pandemic, is informed by real-time mobility estimates from mobile phone data and laboratory-confirmed case and hospitalisation incidence. To estimate regional and time-varying transmissibility, case detection probabilities, and missed imported cases, we developed a novel sequential Approximate Bayesian Computation method allowing inference in useful time, despite the high parametric dimension. We test our approach on Norway and find that three-week-ahead predictions are precise and well-calibrated, enabling policy-relevant situational awareness at a local scale. By comparing the reproduction numbers before and after lockdowns, we identify spatially heterogeneous patterns in their effect on the transmissibility, with a stronger effect in the most populated regions compared to the national reduction estimated to be 85% (95% CI 78%-89%). Our approach is the first regional changepoint stochastic metapopulation model capable of real time spatially refined surveillance and forecasting during emergencies. Solveig Engebretsen, Alfonso Diz-Lois Palomares, Gunnar Rø, Anja Bråthen Kristoffersen, Jonas Christoffer Lindstrøm, Kenth Engø-Monsen, Meghana Kamineni, Louis Yat Hin Chan, Ørjan Dale, Jørgen Eriksson Midtbø, Kristian Lindalen Stenerud, Francesco Di Ruscio, Arnoldo Frigessi, Birgitte Freiesleben de Blasio |
PLoS Comput. Biol. | 14 |
| 2022 | Discovery of host-directed modulators of virus infection by probing the SARS-CoV-2-host protein-protein interaction networkabstractThe ongoing coronavirus disease 2019 (COVID-19) pandemic has highlighted the need to better understand virus-host interactions. We developed a network-based method that expands the severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2)-host protein interaction network and identifies host targets that modulate viral infection. To disrupt the SARS-CoV-2 interactome, we systematically probed for potent compounds that selectively target the identified host proteins with high expression in cells relevant to COVID-19. We experimentally tested seven chemical inhibitors of the identified host proteins for modulation of SARS-CoV-2 infection in human cells that express ACE2 and TMPRSS2. Inhibition of the epigenetic regulators bromodomain-containing protein 4 (BRD4) and histone deacetylase 2 (HDAC2), along with ubiquitin-specific peptidase (USP10), enhanced SARS-CoV-2 infection. Such proviral effect was observed upon treatment with compounds JQ1, vorinostat, romidepsin and spautin-1, when measured by cytopathic effect and validated by viral RNA assays, suggesting that the host proteins HDAC2, BRD4 and USP10 have antiviral functions. We observed marked differences in antiviral effects across cell lines, which may have consequences for identification of selective modulators of viral infection or potential antiviral therapeutics. While network-based approaches enable systematic identification of host targets and selective compounds that may modulate the SARS-CoV-2 interactome, further developments are warranted to increase their accuracy and cell-context specificity. Vandana Ravindran, Jessica Wagoner, Paschalis Athanasiadis, Andreas B. Den Hartigh, Julia M. Sidorova, Aleksandr Ianevski, Susan L. Fink, Arnoldo Frigessi, Judith White, Stephen J. Polyak, Tero Aittokallio |
Briefings Bioinform. | 8 |
| 2022 | Dynamic slate recommendation with gated recurrent units and Thompson samplingabstractAbstract We consider the problem of recommending relevant content to users of an internet platform in the form of lists of items, called slates. We introduce a variational Bayesian Recurrent Neural Net recommender system that acts on time series of interactions between the internet platform and the user, and which scales to real world industrial situations. The recommender system is tested both online on real users, and on an offline dataset collected from a Norwegian web-based marketplace, FINN.no, that is made public for research. This is one of the first publicly available datasets which includes all the slates that are presented to users as well as which items (if any) in the slates were clicked on. Such a data set allows us to move beyond the common assumption that implicitly assumes that users are considering all possible items at each interaction. Instead we build our likelihood using the items that are actually in the slate, and evaluate the strengths and weaknesses of both approaches theoretically and in experiments. We also introduce a hierarchical prior for the item parameters based on group memberships. Both item parameters and user preferences are learned probabilistically. Furthermore, we combine our model with bandit strategies to ensure learning, and introduce ‘in-slate Thompson sampling’ which makes use of the slates to maximise explorative opportunities. We show experimentally that explorative recommender strategies perform on par or above their greedy counterparts. Even without making use of exploration to learn more effectively, click rates increase simply because of improved diversity in the recommended slates. Simen Eide, David S. Leslie, Arnoldo Frigessi |
Data Min. Knowl. Discov. | 3 |
| 2021 | FINN.no Slates Dataset: A new Sequential Dataset Logging Interactions, all Viewed Items and Click Responses/No-Click for Recommender Systems ResearchabstractWe present a novel recommender systems dataset that records the sequential interactions between users and an online marketplace. The users are sequentially presented with both recommendations and search results in the form of ranked lists of items, called slates, from the marketplace. The dataset includes the presented slates at each round, whether the user clicked on any of these items and which item the user clicked on. Although the usage of exposure data in recommender systems is growing, to our knowledge there is no open large-scale recommender systems dataset that includes the slates of items presented to the users at each interaction. As a result, most articles on recommender systems do not utilize this exposure information. Instead, the proposed models only depend on the user's click responses, and assume that the user is exposed to all the items in the item universe at each step, often called uniform candidate sampling. This is an incomplete assumption, as it takes into account items the user might not have been exposed to. This way items might be incorrectly considered as not of interest to the user. Taking into account the actually shown slates allows the models to use a more natural likelihood, based on the click probability given the exposure set of items, as is prevalent in the bandit and reinforcement learning literature. \cite{Eide2021DynamicSampling} shows that likelihoods based on uniform candidate sampling (and similar assumptions) are implicitly assuming that the platform only shows the most relevant items to the user. This causes the recommender system to implicitly reinforce feedback loops and to be biased towards previously exposed items to the user. Simen Eide, David S. Leslie, Arnoldo Frigessi, Joakim Rishaug, Helge Jenssen, Sofie Verrewaere |
RecSys | 3 |
| 2019 | A Bayesian two-way latent structure model for genomic data integration reveals few pan-genomic cluster subtypes in a breast cancer cohortabstractMOTIVATION: Unsupervised clustering is important in disease subtyping, among having other genomic applications. As genomic data has become more multifaceted, how to cluster across data sources for more precise subtyping is an ever more important area of research. Many of the methods proposed so far, including iCluster and Cluster of Cluster Assignments (COCAs), make an unreasonable assumption of a common clustering across all data sources, and those that do not are fewer and tend to be computationally intensive. RESULTS: We propose a Bayesian parametric model for integrative, unsupervised clustering across data sources. In our two-way latent structure model, samples are clustered in relation to each specific data source, distinguishing it from methods like COCAs and iCluster, but cluster labels have across-dataset meaning, allowing cluster information to be shared between data sources. A common scaling across data sources is not required, and inference is obtained by a Gibbs Sampler, which we improve with a warm start strategy and modified density functions to robustify and speed convergence. Posterior interpretation allows for inference on common clusterings occurring among subsets of data sources. An interesting statistical formulation of the model results in sampling from closed-form posteriors despite incorporation of a complex latent structure. We fit the model with Gaussian and more general densities, which influences the degree of across-dataset cluster label sharing. Uniquely among integrative clustering models, our formulation makes no nestedness assumptions of samples across data sources so that a sample missing data from one genomic source can be clustered according to its existing data sources. We apply our model to a Norwegian breast cancer cohort of ductal carcinoma in situ and invasive tumors, comprised of somatic copy-number alteration, methylation and expression datasets. We find enrichment in the Her2 subtype and ductal carcinoma among those observations exhibiting greater cluster correspondence across expression and CNA data. In general, there are few pan-genomic clusterings, suggesting that models assuming a common clustering across genomic data sources might yield misleading results. AVAILABILITY AND IMPLEMENTATION: The model is implemented in an R package called twl ('two-way latent'), available on CRAN. Data for analysis are available within the R package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. David M. Swanson, Tonje Lien, Helga Bergholtz, Therese Sørlie, Arnoldo Frigessi |
Bioinform. | 5 |
| 2019 | Diverse personalized recommendations with uncertainty from implicit preference data with the Bayesian Mallows modelabstractClicking data, which exists in abundance and contains objective user preference information, is widely used to produce personalized recommendations in web-based applications. Current popular recommendation algorithms, typically based on matrix factorizations, often focus on achieving high accuracy. While achieving good clickthrough rates, diversity of the recommended items is often overlooked. Moreover, most algorithms do not produce interpretable uncertainty quantifications of the recommendations. In this work, we propose the Bayesian Mallows for Clicking Data (BMCD) method, which simultaneously considers accuracy and diversity. BMCD augments clicking data into compatible full ranking vectors by enforcing all the clicked items clicked by a user to be top-ranked regardless of their rarity. User preferences are learned using a Mallows ranking model. Bayesian inference leads to interpretable uncertainties of each individual recommendation, and we also propose a method to make personalized recommendations based on such uncertainties. With a simulation study and a real life data example, we demonstrate that compared to state-of-the-art matrix factorization, BMCD makes personalized recommendations with similar accuracy, while achieving much higher level of diversity, and producing interpretable and actionable uncertainty estimation. Andrew Henry Reiner, Arnoldo Frigessi, Ida Scheel |
Knowl. Based Syst. | 3 |
| 2019 | A theoretical single-parameter model for urbanisation to study infectious disease spread and interventionsabstractThe world is continuously urbanising, resulting in clusters of densely populated urban areas and more sparsely populated rural areas. We propose a method for generating spatial fields with controllable levels of clustering of the population. We build a synthetic country, and use this method to generate versions of the country with different clustering levels. Combined with a metapopulation model for infectious disease spread, this allows us to in silico explore how urbanisation affects infectious disease spread. In a baseline scenario with no interventions, the underlying population clustering seems to have little effect on the final size and timing of the epidemic. Under within-country restrictions on non-commuting travel, the final size decreases with increased population clustering. The effect of travel restrictions on reducing the final size is larger with higher clustering. The reduction is larger in the more rural areas. Within-country travel restrictions delay the epidemic, and the delay is largest for lower clustering levels. We implemented three different vaccination strategies-uniform vaccination (in space), preferentially vaccinating urban locations and preferentially vaccinating rural locations. The urban and uniform vaccination strategies were most effective in reducing the final size, while the rural vaccination strategy was clearly inferior. Visual inspection of some European countries shows that many countries already have high population clustering. In the future, they will likely become even more clustered. Hence, according to our model, within-country travel restrictions are likely to be less and less effective in delaying epidemics, while they will be more effective in decreasing final sizes. In addition, to minimise final sizes, it is important not to neglect urban locations when distributing vaccines. To our knowledge, this is the first study to systematically investigate the effect of urbanisation on infectious disease spread and in particular, to examine effectiveness of prevention measures as a function of urbanisation. Solveig Engebretsen, Kenth Engø-Monsen, Arnoldo Frigessi, Birgitte Freiesleben de Blasio |
PLoS Comput. Biol. | 3 |
| 2017 | Probabilistic preference learning with the Mallows rank model
Valeria Vitelli, Øystein Sørensen, Marta Crispino, Arnoldo Frigessi, Elja Arjas |
J. Mach. Learn. Res. | 4 |
| 2016 | Tilting the lasso by knowledge-based post-processingabstractBACKGROUND: It is useful to incorporate biological knowledge on the role of genetic determinants in predicting an outcome. It is, however, not always feasible to fully elicit this information when the number of determinants is large. We present an approach to overcome this difficulty. First, using half of the available data, a shortlist of potentially interesting determinants are generated. Second, binary indications of biological importance are elicited for this much smaller number of determinants. Third, an analysis is carried out on this shortlist using the second half of the data. RESULTS: We show through simulations that, compared with adaptive lasso, this approach leads to models containing more biologically relevant variables, while the prediction mean squared error (PMSE) is comparable or even reduced. We also apply our approach to bone mineral density data, and again final models contain more biologically relevant variables and have reduced PMSEs. CONCLUSION: Our method leads to comparable or improved predictive performance, and models with greater face validity and interpretability with feasible incorporation of biological knowledge into predictive models. Kukatharmini Tharmaratnam, Matthew Sperrin, Thomas Jaki, Sjur Reppe, Arnoldo Frigessi |
BMC Bioinform. | 5 |
| 2011 | Identifying elemental genomic track types and representing them uniformlyabstractBACKGROUND: With the recent advances and availability of various high-throughput sequencing technologies, data on many molecular aspects, such as gene regulation, chromatin dynamics, and the three-dimensional organization of DNA, are rapidly being generated in an increasing number of laboratories. The variation in biological context, and the increasingly dispersed mode of data generation, imply a need for precise, interoperable and flexible representations of genomic features through formats that are easy to parse. A host of alternative formats are currently available and in use, complicating analysis and tool development. The issue of whether and how the multitude of formats reflects varying underlying characteristics of data has to our knowledge not previously been systematically treated. RESULTS: We here identify intrinsic distinctions between genomic features, and argue that the distinctions imply that a certain variation in the representation of features as genomic tracks is warranted. Four core informational properties of tracks are discussed: gaps, lengths, values and interconnections. From this we delineate fifteen generic track types. Based on the track type distinctions, we characterize major existing representational formats and find that the track types are not adequately supported by any single format. We also find, in contrast to the XML formats, that none of the existing tabular formats are conveniently extendable to support all track types. We thus propose two unified formats for track data, an improved XML format, BioXSD 1.1, and a new tabular format, GTrack 1.0. CONCLUSIONS: The defined track types are shown to capture relevant distinctions between genomic annotation tracks, resulting in varying representational needs and analysis possibilities. The proposed formats, GTrack 1.0 and BioXSD 1.1, cater to the identified track distinctions and emphasize preciseness, flexibility and parsing convenience. Sveinung Gundersen, Matús Kalas, Osman Abul, Arnoldo Frigessi, Eivind Hovig, Geir Kjetil Sandve |
BMC Bioinform. | 4 |
| 2011 | Linear and non-linear dependencies between copy number aberrations and mRNA expression reveal distinct molecular pathways in Breast CancerabstractBACKGROUND: Elucidating the exact relationship between gene copy number and expression would enable identification of regulatory mechanisms of abnormal gene expression and biological pathways of regulation. Most current approaches either depend on linear correlation or on nonparametric tests of association that are insensitive to the exact shape of the relationship. Based on knowledge of enzyme kinetics and gene regulation, we would expect the functional shape of the relationship to be gene dependent and to be related to the gene regulatory mechanisms involved. Here, we propose a statistical approach to investigate and distinguish between linear and nonlinear dependences between DNA copy number alteration and mRNA expression. RESULTS: We applied the proposed method to DNA copy numbers derived from Illumina 109 K SNP-CGH arrays (using the log R values) and expression data from Agilent 44 K mRNA arrays, focusing on commonly aberrated genomic loci in a collection of 102 breast tumors. Regression analysis was used to identify the type of relationship (linear or nonlinear), and subsequent pathway analysis revealed that genes displaying a linear relationship were overall associated with substantially different biological processes than genes displaying a nonlinear relationship. In the group of genes with a linear relationship, we found significant association to canonical pathways, including purine and pyrimidine metabolism (for both deletions and amplifications) as well as estrogen metabolism (linear amplification) and BRCA-related response to damage (linear deletion). In the group of genes displaying a nonlinear relationship, the top canonical pathways were specific pathways like PTEN and PI13K/AKT (nonlinear amplification) and Wnt(B) and IL-2 signalling (nonlinear deletion). Both amplifications and deletions pointed to the same affected pathways and identified cancer as the top significant disease and cell cycle, cell signaling and cellular development as significant networks. CONCLUSIONS: This paper presents a novel approach to assessing the validity of the dependence of expression data on copy number data, and this approach may help in identifying the drivers of carcinogenesis. Hiroko K. Solvang, Ole Christian Lingjærde, Arnoldo Frigessi, Anne-Lise Børresen-Dale, Vessela N. Kristensen |
BMC Bioinform. | 3 |
| 2011 | Estimated Comparative Integration Hotspots Identify Different Behaviors of Retroviral Gene Transfer VectorsabstractIntegration of retroviral vectors in the human genome follows non random patterns that favor insertional deregulation of gene expression and may cause risks of insertional mutagenesis when used in clinical gene therapy. Understanding how viral vectors integrate into the human genome is a key issue in predicting these risks. We provide a new statistical method to compare retroviral integration patterns. We identified the positions where vectors derived from the Human Immunodeficiency Virus (HIV) and the Moloney Murine Leukemia Virus (MLV) show different integration behaviors in human hematopoietic progenitor cells. Non-parametric density estimation was used to identify candidate comparative hotspots, which were then tested and ranked. We found 100 significative comparative hotspots, distributed throughout the chromosomes. HIV hotspots were wider and contained more genes than MLV ones. A Gene Ontology analysis of HIV targets showed enrichment of genes involved in antigen processing and presentation, reflecting the high HIV integration frequency observed at the MHC locus on chromosome 6. Four histone modifications/variants had a different mean density in comparative hotspots (H2AZ, H3K4me1, H3K4me3, H3K9me1), while gene expression within the comparative hotspots did not differ from background. These findings suggest the existence of epigenetic or nuclear three-dimensional topology contexts guiding retroviral integration to specific chromosome areas. Alessandro Ambrosi, Ingrid Kristine Glad, Danilo Pellin, Claudia Cattoglio, Fulvio Mavilio, Clelia Di Serio, Arnoldo Frigessi |
PLoS Comput. Biol. | 7 |
| 2007 | Predicting survival from microarray data - a comparative studyabstractMOTIVATION: Survival prediction from gene expression data and other high-dimensional genomic data has been subject to much research during the last years. These kinds of data are associated with the methodological problem of having many more gene expression values than individuals. In addition, the responses are censored survival times. Most of the proposed methods handle this by using Cox's proportional hazards model and obtain parameter estimates by some dimension reduction or parameter shrinkage estimation technique. Using three well-known microarray gene expression data sets, we compare the prediction performance of seven such methods: univariate selection, forward stepwise selection, principal components regression (PCR), supervised principal components regression, partial least squares regression (PLS), ridge regression and the lasso. RESULTS: Statistical learning from subsets should be repeated several times in order to get a fair comparison between methods. Methods using coefficient shrinkage or linear combinations of the gene expression values have much better performance than the simple variable selection methods. For our data sets, ridge regression has the overall best performance. AVAILABILITY: Matlab and R code for the prediction methods are available at http://www.med.uio.no/imb/stat/bmms/software/microsurv/. Hege M. Bøvelstad, Ståle Nygård, H. L. Størvold, Magne Aldrin, Ørnulf Borgan, Arnoldo Frigessi, Ole Christian Lingjærde |
Bioinform. | 6 |
| 2005 | The influence of missing value imputation on detection of differentially expressed genes from microarray dataabstractMOTIVATION: Missing values are problematic for the analysis of microarray data. Imputation methods have been compared in terms of the similarity between imputed and true values in simulation experiments and not of their influence on the final analysis. The focus has been on missing at random, while entries are missing also not at random. RESULTS: We investigate the influence of imputation on the detection of differentially expressed genes from cDNA microarray data. We apply ANOVA for microarrays and SAM and look to the differentially expressed genes that are lost because of imputation. We show that this new measure provides useful information that the traditional root mean squared error cannot capture. We also show that the type of missingness matters: imputing 5% missing not at random has the same effect as imputing 10-30% missing at random. We propose a new method for imputation (LinImp), fitting a simple linear model for each channel separately, and compare it with the widely used KNNimpute method. For 10% missing at random, KNNimpute leads to twice as many lost differentially expressed genes as LinImp. AVAILABILITY: The R package for LinImp is available at http://folk.uio.no/idasch/imp. Ida Scheel, Magne Aldrin, Ingrid Kristine Glad, Ragnhild Sørum, Heidi Lyng, Arnoldo Frigessi |
Bioinform. | 6 |