EDBT 2026 Demo / reviewers in the wild / expert
Nima Aghaeepour
dblp:51/4668
· DBLP profile ↗
11ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-6117-8764ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 56% Medical and health informatics · 44% | |
| Artificial intelligence
1 paper |
Language models and text generation · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › single-cell analysis
cytometry data analysis |
0.9 | 4 | 2018 | GateFinder: projection-based gating strategy optimization for flow and mass cytometry · Bioinform. 2018 Deep profiling of multitube flow cytometry data · Bioinform. 2015 Enhanced flowType/RchyOptimyx: a Bioconductor pipeline for discovery in high-dimensional cytometry data · Bioinform. 2014 |
Natural language and speech › Language models and text generation
instruction following |
0.8 | 1 | 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records · AAAI 2024 |
Medical and health informatics
clinical assessment |
0.8 | 1 | 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records · AAAI 2024 |
Medical and health informatics › clinical text processing
clinical text generation |
0.8 | 1 | 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records · AAAI 2024 |
Bioinformatics and computational biology
immunophenotyping |
0.4 | 2 | 2015 | Deep profiling of multitube flow cytometry data · Bioinform. 2015 Enhanced flowType/RchyOptimyx: a Bioconductor pipeline for discovery in high-dimensional cytometry data · Bioinform. 2014 |
Bioinformatics and computational biology
multi-omics data integration |
0.4 | 1 | 2019 | Multiomics modeling of the immunome, transcriptome, microbiome, proteome and metabolome adaptations during human pregnancy · Bioinform. 2019 |
Methods — techniques the papers use, named apart from their topics
natural language generation metrics · 1.5large language model · 1.5stacked generalization · 0.4elastic net · 0.4projection-based dimensionality reduction · 0.3polygon gating · 0.3nearest-neighbour imputation · 0.2clustering · 0.2classification · 0.2feature selection · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mitigation of outcome conflation in predicting patient outcomes using electronic health recordsabstractOBJECTIVES: Artificial intelligence (AI) models utilizing electronic health record data for disease prediction can enhance risk stratification but may lack specificity, which is crucial for reducing the economic and psychological burdens associated with false positives. This study aims to evaluate the impact of confounders on the specificity of single-outcome prediction models and assess the effectiveness of a multi-class architecture in mitigating outcome conflation. MATERIALS AND METHODS: We evaluated a state-of-the-art model predicting pancreatic cancer from disease code sequences in an independent cohort of 2.3 million patients and compared this single-outcome model with a multi-class model designed to predict multiple cancer types simultaneously. Additionally, we conducted a clinical simulation experiment to investigate the impact of confounders on the specificity of single-outcome prediction models. RESULTS: While we were able to independently validate the pancreatic cancer prediction model, we found that its prediction scores were also correlated with ovarian cancer, suggesting conflation of outcomes due to underlying confounders. Building on this observation, we demonstrate that the specificity of single-outcome prediction models is impaired by confounders using a clinical simulation experiment. Introducing a multi-class architecture improves specificity in predicting cancer types compared to the single-outcome model while preserving performance, mitigating the conflation of outcomes in both the real-world and simulated contexts. DISCUSSION: Our results highlight the risk of outcome conflation in single-outcome AI prediction models and demonstrate the effectiveness of a multi-class approach in mitigating this issue. CONCLUSION: The number of predicted outcomes needs to be carefully considered when employing AI disease risk prediction models. S. Momsen Reincke, Camilo Espinosa, Philip Chung 0005, Tomin James, Eloïse Berson, Nima Aghaeepour |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical RecordsabstractThe ability of large language models (LLMs) to follow natural language instructions with human-level fluency suggests many opportunities in healthcare to reduce administrative burden and improve quality of care. However, evaluating LLMs on realistic text generation tasks for healthcare remains challenging. Existing question answering datasets for electronic health record (EHR) data fail to capture the complexity of information needs and documentation burdens experienced by clinicians. To address these challenges, we introduce MedAlign, a benchmark dataset of 983 natural language instructions for EHR data. MedAlign is curated by 15 clinicians (7 specialities), includes clinician-written reference responses for 303 instructions, and provides 276 longitudinal EHRs for grounding instruction-response pairs. We used MedAlign to evaluate 6 general domain LLMs, having clinicians rank the accuracy and quality of each LLM response. We found high error rates, ranging from 35% (GPT-4) to 68% (MPT-7B-Instruct), and 8.3% drop in accuracy moving from 32k to 2k context lengths for GPT-4. Finally, we report correlations between clinician rankings and automated natural language generation metrics as a way to rank LLMs without human review. MedAlign is provided under a research data use agreement to enable LLM evaluations on tasks aligned with clinician needs and preferences. Scott L. Fleming, Alejandro Lozano, William J. Haberkorn, Jenelle A. Jindal, Eduardo Pontes Reis, Rahul Thapa, Louis Blankemeier, Julian Z. Genkins, Ethan Steinberg, Ashwin Nayak 0002, Birju Patel, Chia-Chun Chiang, Alison Callahan, Zepeng Huo, Sergios Gatidis, Scott J. Adams, Oluseyi Fayanju, Shreya J. Shah, Thomas Savage, Ethan Goh, Akshay Chaudhari, Nima Aghaeepour, Christopher D. Sharp, Michael A. Pfeffer, Percy Liang, Jonathan H. Chen, Keith E. Morse, Emma Brunskill, Jason Alan Fries, Nigam H. Shah |
AAAI | 22 |
| 2024 | Generating pregnant patient biological profiles by deconvoluting clinical records with electronic health record foundation modelsabstractTranslational biology posits a strong bi-directional link between clinical phenotypes and a patient's biological profile. By leveraging this bi-directional link, we can efficiently deconvolute pre-existing clinical information into biological profiles. However, traditional computational tools are limited in their ability to resolve this link because of the relatively small sizes of paired clinical-biological datasets for training and the high dimensionality/sparsity of tabular clinical data. Here, we use state-of-the-art foundation models (FMs) for electronic health record (EHR) data to generate proteomics profiles of pregnant patients, thereby deconvoluting pre-existing clinical information into biological profiles without the cost and effort of running large-scale traditional omics studies. We show that FM-derived representations of a patient's EHR data coupled with a fully connected neural network prediction head can generate 206 blood protein expression levels. Interestingly, these proteins were enriched for developmental pathways, while proteins not able to be generated from EHR data were enriched for metabolic pathways. Finally, we show a proteomic signature of gestational diabetes that includes proteins with established and novel links to gestational diabetes. These results showcase the power of FM-derived EHR representations in efficiently generating biological states of pregnant patients. This capability can revolutionize disease understanding and therapeutic development, offering a cost-effective, time-efficient, and less invasive alternative to traditional methods of generating proteomics. David Seong, Samson Mataraso, Camilo Espinosa, Eloïse Berson, S. Momsen Reincke, Chloe Kashiwagi, Yeasul Kim, Chi-Hung Shu, Philip Chung 0005, Marc Ghanem, Feng Xie 0004, Ronald J. Wong, Martin S. Angst, Brice Gaudilliere, Gary M. Shaw, David K. Stevenson, Nima Aghaeepour |
Briefings Bioinform. | 18 |
| 2019 | Multiomics modeling of the immunome, transcriptome, microbiome, proteome and metabolome adaptations during human pregnancyabstractMotivation: Multiple biological clocks govern a healthy pregnancy. These biological mechanisms produce immunologic, metabolomic, proteomic, genomic and microbiomic adaptations during the course of pregnancy. Modeling the chronology of these adaptations during full-term pregnancy provides the frameworks for future studies examining deviations implicated in pregnancy-related pathologies including preterm birth and preeclampsia. Results: We performed a multiomics analysis of 51 samples from 17 pregnant women, delivering at term. The datasets included measurements from the immunome, transcriptome, microbiome, proteome and metabolome of samples obtained simultaneously from the same patients. Multivariate predictive modeling using the Elastic Net (EN) algorithm was used to measure the ability of each dataset to predict gestational age. Using stacked generalization, these datasets were combined into a single model. This model not only significantly increased predictive power by combining all datasets, but also revealed novel interactions between different biological modalities. Future work includes expansion of the cohort to preterm-enriched populations and in vivo analysis of immune-modulating interventions based on the mechanisms identified. Availability and implementation: Datasets and scripts for reproduction of results are available through: https://nalab.stanford.edu/multiomics-pregnancy/. Supplementary information: Supplementary data are available at Bioinformatics online. Mohammad Sajjad Ghaemi, Daniel B. DiGiulio, Kévin Contrepois, Benjamin J. Callahan, Thuy T. M. Ngo, Brittany Lee-McMullen, Benoit Lehallier, Anna Robaczewska, David Mcilwain, Yael Rosenberg-Hasson, Ronald J. Wong, Cecele Quaintance, Anthony Culos, Natalie Stanley, Athena Tanada, Amy Tsai, Dyani Gaudilliere, Edward Ganio, Xiaoyuan Han, Kazuo Ando, Leslie McNeil, Martha Tingle, Paul H. Wise, Ivana Maric, Marina Sirota, Tony Wyss-Coray, Virginia D. Winn, Maurice L. Druzin, Ronald Gibbs, Gary L. Darmstadt, David B. Lewis, Vahid Partovi Nia, Bruno Agard, Robert Tibshirani, Garry P. Nolan, Michael Snyder 0001, David A. Relman, Stephen R. Quake, Gary M. Shaw, David K. Stevenson, Martin S. Angst, Brice Gaudilliere, Nima Aghaeepour |
Bioinform. | 43 |
| 2018 | GateFinder: projection-based gating strategy optimization for flow and mass cytometryabstractMotivation: High-parameter single-cell technologies can reveal novel cell populations of interest, but studying or validating these populations using lower-parameter methods remains challenging. Results: Here, we present GateFinder, an algorithm that enriches high-dimensional cell types with simple, stepwise polygon gates requiring only two markers at a time. A series of case studies of complex cell types illustrates how simplified enrichment strategies can enable more efficient assays, reveal novel biomarkers and clarify underlying biology. Availability and implementation: The GateFinder algorithm is implemented as a free and open-source package for BioConductor: https://nalab.stanford.edu/gatefinder. Supplementary information: Supplementary data are available at Bioinformatics online. Nima Aghaeepour, Erin F. Simonds, David J. H. F. Knapp, Robert V. Bruggner, Karen Sachs, Anthony Culos, Pier Federico Gherardini, Nikolay Samusik, Gabriela K. Fragiadakis, Sean Bendall, Brice Gaudilliere, Martin S. Angst, Connie J. Eaves, William A. Weiss, Wendy J. Fantl, Garry P. Nolan |
Bioinform. | 1 |
| 2015 | Deep profiling of multitube flow cytometry dataabstractAbstract Motivation: Deep profiling the phenotypic landscape of tissues using high-throughput flow cytometry (FCM) can provide important new insights into the interplay of cells in both healthy and diseased tissue. But often, especially in clinical settings, the cytometer cannot measure all the desired markers in a single aliquot. In these cases, tissue is separated into independently analysed samples, leaving a need to electronically recombine these to increase dimensionality. Nearest-neighbour (NN) based imputation fulfils this need but can produce artificial subpopulations. Clustering-based NNs can reduce these, but requires prior domain knowledge to be able to parameterize the clustering, so is unsuited to discovery settings. Results: We present flowBin, a parameterization-free method for combining multitube FCM data into a higher-dimensional form suitable for deep profiling and discovery. FlowBin allocates cells to bins defined by the common markers across tubes in a multitube experiment, then computes aggregate expression for each bin within each tube, to create a matrix of expression of all markers assayed in each tube. We show, using simulated multitube data, that flowType analysis of flowBin output reproduces the results of that same analysis on the original data for cell types of >10% abundance. We used flowBin in conjunction with classifiers to distinguish normal from cancerous cells. We used flowBin together with flowType and RchyOptimyx to profile the immunophenotypic landscape of NPM1-mutated acute myeloid leukemia, and present a series of novel cell types associated with that mutation. Availability and implementation: FlowBin is available in Bioconductor under the Artistic 2.0 free open source license. All data used are available in FlowRepository under accessions: FR-FCM-ZZYA, FR-FCM-ZZZK and FR-FCM-ZZES. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Kieran O'Neill, Nima Aghaeepour, Jeremy Parker, Donna Hogge, Aly Karsan, Bakul Dalal, Ryan Remy Brinkman |
Bioinform. | 2 |
| 2014 | Enhanced flowType/RchyOptimyx: a Bioconductor pipeline for discovery in high-dimensional cytometry dataabstractAbstract Summary: We present a significantly improved version of the flowType and RchyOptimyx BioConductor-based pipeline that is both 14 times faster and can accommodate multiple levels of biomarker expression for up to 96 markers. With these improvements, the pipeline is positioned to be an integral part of data analysis for high-throughput experiments on high-dimensional single-cell assay platforms, including flow cytometry, mass cytometry and single-cell RT-qPCR. Availability: FlowType and RchyOptimyx are distributed under the Artistic 2.0 license through Bioconductor. Contact: [email protected] Kieran O'Neill, Adrin Jalali, Nima Aghaeepour, Holger H. Hoos, Ryan Remy Brinkman |
Bioinform. | 3 |
| 2013 | Ensemble-based prediction of RNA secondary structuresabstractBACKGROUND: Accurate structure prediction methods play an important role for the understanding of RNA function. Energy-based, pseudoknot-free secondary structure prediction is one of the most widely used and versatile approaches, and improved methods for this task have received much attention over the past five years. Despite the impressive progress that as been achieved in this area, existing evaluations of the prediction accuracy achieved by various algorithms do not provide a comprehensive, statistically sound assessment. Furthermore, while there is increasing evidence that no prediction algorithm consistently outperforms all others, no work has been done to exploit the complementary strengths of multiple approaches. RESULTS: In this work, we present two contributions to the area of RNA secondary structure prediction. Firstly, we use state-of-the-art, resampling-based statistical methods together with a previously published and increasingly widely used dataset of high-quality RNA structures to conduct a comprehensive evaluation of existing RNA secondary structure prediction procedures. The results from this evaluation clarify the performance relationship between ten well-known existing energy-based pseudoknot-free RNA secondary structure prediction methods and clearly demonstrate the progress that has been achieved in recent years. Secondly, we introduce AveRNA, a generic and powerful method for combining a set of existing secondary structure prediction procedures into an ensemble-based method that achieves significantly higher prediction accuracies than obtained from any of its component procedures. CONCLUSIONS: Our new, ensemble-based method, AveRNA, improves the state of the art for energy-based, pseudoknot-free RNA secondary structure prediction by exploiting the complementary strengths of multiple existing prediction procedures, as demonstrated using a state-of-the-art statistical resampling approach. In addition, AveRNA allows an intuitive and effective control of the trade-off between false negative and false positive base pair predictions. Finally, AveRNA can make use of arbitrary sets of secondary structure prediction procedures and can therefore be used to leverage improvements in prediction accuracy offered by algorithms and energy models developed in the future. Our data, MATLAB software and a web-based version of AveRNA are publicly available at http://www.cs.ubc.ca/labs/beta/Software/AveRNA. Nima Aghaeepour, Holger H. Hoos |
BMC Bioinform. | 1 |
| 2013 | Flow Cytometry BioinformaticsabstractFlow cytometry bioinformatics is the application of bioinformatics to flow cytometry data, which involves storing, retrieving, organizing, and analyzing flow cytometry data using extensive computational resources and tools. Flow cytometry bioinformatics requires extensive use of and contributes to the development of techniques from computational statistics and machine learning. Flow cytometry and related methods allow the quantification of multiple independent biomarkers on large numbers of single cells. The rapid growth in the multidimensionality and throughput of flow cytometry data, particularly in the 2000s, has led to the creation of a variety of computational analysis methods, data standards, and public databases for the sharing of results. Computational methods exist to assist in the preprocessing of flow cytometry data, identifying cell populations within it, matching those cell populations across samples, and performing diagnosis and discovery using the results of previous steps. For preprocessing, this includes compensating for spectral overlap, transforming data onto scales conducive to visualization and analysis, assessing data for quality, and normalizing data across samples and experiments. For population identification, tools are available to aid traditional manual identification of populations in two-dimensional scatter plots (gating), to use dimensionality reduction to aid gating, and to find populations automatically in higher dimensional space in a variety of ways. It is also possible to characterize data in more comprehensive ways, such as the density-guided binary space partitioning technique known as probability binning, or by combinatorial gating. Finally, diagnosis using flow cytometry data can be aided by supervised learning techniques, and discovery of new cell types of biological importance by high-throughput statistical methods, as part of pipelines incorporating all of the aforementioned methods. Open standards, data, and software are also key parts of flow cytometry bioinformatics. Data standards include the widely adopted Flow Cytometry Standard (FCS) defining how data from cytometers should be stored, but also several new standards under development by the International Society for Advancement of Cytometry (ISAC) to aid in storing more detailed information about experimental design and analytical steps. Open data is slowly growing with the opening of the CytoBank database in 2010 and FlowRepository in 2012, both of which allow users to freely distribute their data, and the latter of which has been recommended as the preferred repository for MIFlowCyt-compliant data by ISAC. Open software is most widely available in the form of a suite of Bioconductor packages, but is also available for web execution on the GenePattern platform. Kieran O'Neill, Nima Aghaeepour, Josef Spidlen, Ryan Remy Brinkman |
PLoS Comput. Biol. | 2 |
| 2012 | Early immunologic correlates of HIV protection can be identified from computational analysis of complex multivariate T-cell flow cytometry assaysabstractMOTIVATION: Polychromatic flow cytometry (PFC), has enormous power as a tool to dissect complex immune responses (such as those observed in HIV disease) at a single cell level. However, analysis tools are severely lacking. Although high-throughput systems allow rapid data collection from large cohorts, manual data analysis can take months. Moreover, identification of cell populations can be subjective and analysts rarely examine the entirety of the multidimensional dataset (focusing instead on a limited number of subsets, the biology of which has usually already been well-described). Thus, the value of PFC as a discovery tool is largely wasted. RESULTS: To address this problem, we developed a computational approach that automatically reveals all possible cell subsets. From tens of thousands of subsets, those that correlate strongly with clinical outcome are selected and grouped. Within each group, markers that have minimal relevance to the biological outcome are removed, thereby distilling the complex dataset into the simplest, most clinically relevant subsets. This allows complex information from PFC studies to be translated into clinical or resource-poor settings, where multiparametric analysis is less feasible. We demonstrate the utility of this approach in a large (n=466), retrospective, 14-parameter PFC study of early HIV infection, where we identify three T-cell subsets that strongly predict progression to AIDS (only one of which was identified by an initial manual analysis). AVAILABILITY: The 'flowType: Phenotyping Multivariate PFC Assays' package is available through Bioconductor. Additional documentation and examples are available at: www.terryfoxlab.ca/flowsite/flowType/ SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected]. Nima Aghaeepour, Pratip K. Chattopadhyay, Anuradha Ganesan, Kieran O'Neill, Habil Zare, Adrin Jalali, Holger H. Hoos, Mario Roederer, Ryan Remy Brinkman |
Bioinform. | 1 |
| 2005 | Dynamic Positioning Based on Voronoi Cells (DPVC)
HesamAddin Dashti, Nima Aghaeepour, Sahar Asadi, Meysam Bastani, Zahra Delafkar, Fatemeh Miri Disfani, Serveh Ghaderi, Shahin Kamali, Sepideh Pashami, Alireza Siahpirani |
RoboCup | 2 |