Anthony Gitter

dblp:94/10045 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-5324-9833ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 68% Medical and health informatics · 23% Computing education · 8%
Artificial intelligence
2 papers
Language models and text generation · 54% Learning paradigms · 23% Transfer learning and domain adaptation · 23%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
in-context learning
0.912025
Assay2Mol: Large Language Model-based Drug Design Using BioAssay Context · EMNLP 2025
Medical and health informatics › precision medicine
disease subtype discovery
0.912025
MPAC: a computational framework for inferring pathway activities from multi-omic data · Bioinform. 2025
Bioinformatics and computational biology › drug discovery
drug design
0.912025
Assay2Mol: Large Language Model-based Drug Design Using BioAssay Context · EMNLP 2025
Bioinformatics and computational biology › molecular informatics › cheminformatics
molecule generation
0.912025
Assay2Mol: Large Language Model-based Drug Design Using BioAssay Context · EMNLP 2025
Bioinformatics and computational biology
multi-omics data integration
0.912025
MPAC: a computational framework for inferring pathway activities from multi-omic data · Bioinform. 2025
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
pathway activity inference
0.912025
MPAC: a computational framework for inferring pathway activities from multi-omic data · Bioinform. 2025
Medical and health informatics › precision medicine
patient subgroup identification
0.912025
MPAC: a computational framework for inferring pathway activities from multi-omic data · Bioinform. 2025
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
signaling pathway inference
0.412020
Inferring signaling pathways with probabilistic programming · Bioinform. 2020
Bioinformatics and computational biology
systems biology
0.412020
Inferring signaling pathways with probabilistic programming · Bioinform. 2020
Machine learning › Learning paradigms
multi-task learning
0.412019
Loss-Balanced Task Weighting to Reduce Negative Transfer in Multi-Task Learning · AAAI 2019
Machine learning › Transfer learning and domain adaptation
negative transfer
0.412019
Loss-Balanced Task Weighting to Reduce Negative Transfer in Multi-Task Learning · AAAI 2019
Bioinformatics and computational biology
machine learning for biology
0.212022
An approachable, flexible and practical machine learning workshop for biologists · Bioinform. 2022
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference
0.212013
Identifying proteins controlling key disease signaling pathways · Bioinform. 2013
Bioinformatics and computational biology › biological network › network biology
signaling network modeling
0.212013
Identifying proteins controlling key disease signaling pathways · Bioinform. 2013
Bioinformatics and computational biology › proteomics
phosphoproteomics
0.112020
Inferring signaling pathways with probabilistic programming · Bioinform. 2020
Bioinformatics and computational biology
time course data analysis
0.112020
Inferring signaling pathways with probabilistic programming · Bioinform. 2020
Computational science and engineering
computational chemistry
0.112019
Loss-Balanced Task Weighting to Reduce Negative Transfer in Multi-Task Learning · AAAI 2019

Methods — techniques the papers use, named apart from their topics

retrieval · 1.7in-context learning · 1.7permutation testing · 0.9factor graph · 0.9multi-task learning · 0.8dynamic loss balancing · 0.8active learning · 0.6probabilistic programming · 0.4markov chain monte carlo · 0.4dynamic bayesian network · 0.4task weighting · 0.4
YearPublicationVenuePosition
2025 Assay2Mol: Large Language Model-based Drug Design Using BioAssay Context
abstract
Scientific databases aggregate vast amounts of quantitative data alongside descriptive text. In biochemistry, molecule screening assays evaluate candidate molecules' functional responses against disease targets. Unstructured text that describes the biological mechanisms through which these targets operate, experimental screening protocols, and other attributes of assays offer rich information for drug discovery campaigns but has been untapped because of that unstructured format. We present Assay2Mol, a large language model-based workflow that can capitalize on the vast existing biochemical screening assays for early-stage drug discovery. Assay2Mol retrieves existing assay records involving targets similar to the new target and generates candidate molecules using in-context learning with the retrieved assay screening data. Assay2Mol outperforms recent machine learning approaches that generate candidate ligand molecules for target protein structures, while also promoting more synthesizable molecule generation.
Spencer S. Ericksen, Anthony Gitter
EMNLP3
2025 MPAC: a computational framework for inferring pathway activities from multi-omic data
abstract
MOTIVATION: Fully capturing cellular state requires examining genomic, epigenomic, transcriptomic, proteomic, and other assays for a biological sample and comprehensive computational modeling to reason with the complex and sometimes conflicting measurements. Modeling these so-called multi-omic data is especially beneficial in disease analysis, where observations across omic data types may reveal unexpected patient groupings and inform clinical outcomes and treatments. RESULTS: We present Multi-omic Pathway Analysis of Cells (MPAC), a computational framework that interprets multi-omic data through prior knowledge from biological pathways. MPAC leverages network relationships encoded in pathways through a factor graph to infer consensus activity levels for proteins and associated pathway entities from multi-omic data, runs permutation testing to eliminate spurious activity predictions, and groups biological samples by pathway activities to allow identifying and prioritizing proteins with potential clinical relevance, e.g. associated with patient prognosis. Using DNA copy number alteration and RNA-seq data from head and neck squamous cell carcinoma patients from The Cancer Genome Atlas as an example, we demonstrate that MPAC predicts a patient subgroup related to immune responses not identified by analysis with either input omic data type alone. Key proteins identified via this subgroup have pathway activities related to clinical outcome as well as immune cell composition. Our MPAC R package enables similar multi-omic analyses on new datasets. AVAILABILITY AND IMPLEMENTATION: The MPAC package is available at Bioconductor https://bioconductor.org/packages/MPAC.
David Page, Paul Ahlquist, Irene M. Ong, Anthony Gitter
Bioinform.5
2022 An approachable, flexible and practical machine learning workshop for biologists
abstract
SUMMARY: The increasing prevalence and importance of machine learning in biological research have created a need for machine learning training resources tailored towards biological researchers. However, existing resources are often inaccessible, infeasible or inappropriate for biologists because they require significant computational and mathematical knowledge, demand an unrealistic time-investment or teach skills primarily for computational researchers. We created the Machine Learning for Biologists (ML4Bio) workshop, a short, intensive workshop that empowers biological researchers to comprehend machine learning applications and pursue machine learning collaborations in their own research. The ML4Bio workshop focuses on classification and was designed around three principles: (i) emphasizing preparedness over fluency or expertise, (ii) necessitating minimal coding and mathematical background and (iii) requiring low time investment. It incorporates active learning methods and custom open-source software that allows participants to explore machine learning workflows. After multiple sessions to improve workshop design, we performed a study on three workshop sessions. Despite some confusion around identifying subtle methodological flaws in machine learning workflows, participants generally reported that the workshop met their goals, provided them with valuable skills and knowledge and greatly increased their beliefs that they could engage in research that uses machine learning. ML4Bio is an educational tool for biological researchers, and its creation and evaluation provide valuable insight into tailoring educational resources for active researchers in different domains. AVAILABILITY AND IMPLEMENTATION: Workshop materials are available at https://github.com/carpentries-incubator/ml4bio-workshop and the ml4bio software is available at https://github.com/gitter-lab/ml4bio. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chris S. Magnano, Fangzhou Mu, Rosemary S. Russ, Milica Cvetkovic, Debora Treu, Anthony Gitter
Bioinform.6
2022 Ten quick tips for deep learning in biology
abstract
Machine learning is a modern approach to problem-solving and task automation.In particular, machine learning is concerned with the development and applications of algorithms that
Benjamin D. Lee, Anthony Gitter, Casey S. Greene, Sebastian Raschka, Finlay Maguire, Alexander J. Titus, Michael D. Kessler, Alexandra Lee, Marc G. Chevrette, Paul Allen Stewart, Thiago Britto-Borges, Evan M. Cofer, Kun-Hsing Yu, Juan Jose Carmona, Elana J. Fertig, Alexandr A. Kalinin, Brandon Signal, Benjamin J. Lengerich, Timothy J. Triche Jr., Simina M. Boca
PLoS Comput. Biol.2
2020 Inferring signaling pathways with probabilistic programming
abstract
MOTIVATION: Cells regulate themselves via dizzyingly complex biochemical processes called signaling pathways. These are usually depicted as a network, where nodes represent proteins and edges indicate their influence on each other. In order to understand diseases and therapies at the cellular level, it is crucial to have an accurate understanding of the signaling pathways at work. Since signaling pathways can be modified by disease, the ability to infer signaling pathways from condition- or patient-specific data is highly valuable. A variety of techniques exist for inferring signaling pathways. We build on past works that formulate signaling pathway inference as a Dynamic Bayesian Network structure estimation problem on phosphoproteomic time course data. We take a Bayesian approach, using Markov Chain Monte Carlo to estimate a posterior distribution over possible Dynamic Bayesian Network structures. Our primary contributions are (i) a novel proposal distribution that efficiently samples sparse graphs and (ii) the relaxation of common restrictive modeling assumptions. RESULTS: We implement our method, named Sparse Signaling Pathway Sampling, in Julia using the Gen probabilistic programming language. Probabilistic programming is a powerful methodology for building statistical models. The resulting code is modular, extensible and legible. The Gen language, in particular, allows us to customize our inference procedure for biological graphs and ensure efficient sampling. We evaluate our algorithm on simulated data and the HPN-DREAM pathway reconstruction challenge, comparing our performance against a variety of baseline methods. Our results demonstrate the vast potential for probabilistic programming, and Gen specifically, for biological network inference. AVAILABILITY AND IMPLEMENTATION: Find the full codebase at https://github.com/gitter-lab/ssps. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
David Merrell, Anthony Gitter
Bioinform.2
2020 Lag penalized weighted correlation for time series clustering
abstract
BACKGROUND: The similarity or distance measure used for clustering can generate intuitive and interpretable clusters when it is tailored to the unique characteristics of the data. In time series datasets generated with high-throughput biological assays, measurements such as gene expression levels or protein phosphorylation intensities are collected sequentially over time, and the similarity score should capture this special temporal structure. RESULTS: We propose a clustering similarity measure called Lag Penalized Weighted Correlation (LPWC) to group pairs of time series that exhibit closely-related behaviors over time, even if the timing is not perfectly synchronized. LPWC aligns time series profiles to identify common temporal patterns. It down-weights aligned profiles based on the length of the temporal lags that are introduced. We demonstrate the advantages of LPWC versus existing time series and general clustering algorithms. In a simulated dataset based on the biologically-motivated impulse model, LPWC is the only method to recover the true clusters for almost all simulated genes. LPWC also identifies clusters with distinct temporal patterns in our yeast osmotic stress response and axolotl limb regeneration case studies. CONCLUSIONS: LPWC achieves both of its time series clustering goals. It groups time series with correlated changes over time, even if those patterns occur earlier or later in some of the time series. In addition, it refrains from introducing large shifts in time when searching for temporal patterns by applying a lag penalty. The LPWC R package is available at https://github.com/gitter-lab/LPWC and CRAN under a MIT license.
Thevaa Chandereng, Anthony Gitter
BMC Bioinform.2
2019 Loss-Balanced Task Weighting to Reduce Negative Transfer in Multi-Task Learning
abstract
In settings with related prediction tasks, integrated multi-task learning models can often improve performance relative to independent single-task models. However, even when the average task performance improves, individual tasks may experience negative transfer in which the multi-task model’s predictions are worse than the single-task model’s. We show the prevalence of negative transfer in a computational chemistry case study with 128 tasks and introduce a framework that provides a foundation for reducing negative transfer in multitask models. Our Loss-Balanced Task Weighting approach dynamically updates task weights during model training to control the influence of individual tasks.
Shengchao Liu, Yingyu Liang, Anthony Gitter
AAAI3
2019 Open collaborative writing with Manubot
abstract
Open, collaborative research is a powerful paradigm that can immensely strengthen the scientific process by integrating broad and diverse expertise. However, traditional research and multi-author writing processes break down at scale. We present new software named Manubot, available at https://manubot.org, to address the challenges of open scholarly writing. Manubot adopts the contribution workflow used by many large-scale open source software projects to enable collaborative authoring of scholarly manuscripts. With Manubot, manuscripts are written in Markdown and stored in a Git repository to precisely track changes over time. By hosting manuscript repositories publicly, such as on GitHub, multiple authors can simultaneously propose and review changes. A cloud service automatically evaluates proposed changes to catch errors. Publication with Manubot is continuous: When a manuscript's source changes, the rendered outputs are rebuilt and republished to a web page. Manubot automates bibliographic tasks by implementing citation by identifier, where users cite persistent identifiers (e.g. DOIs, PubMed IDs, ISBNs, URLs), whose metadata is then retrieved and converted to a user-specified style. Manubot modernizes publishing to align with the ideals of open science by making it transparent, reproducible, immediate, versioned, collaborative, and free of charge.
Daniel S. Himmelstein, Vincent Rubinetti, David R. Slochower, Dongbo Hu, Venkat S. Malladi, Casey S. Greene, Anthony Gitter
PLoS Comput. Biol.7
2019 Predicting kinase inhibitors using bioactivity matrix derived informer sets
abstract
Prediction of compounds that are active against a desired biological target is a common step in drug discovery efforts. Virtual screening methods seek some active-enriched fraction of a library for experimental testing. Where data are too scarce to train supervised learning models for compound prioritization, initial screening must provide the necessary data. Commonly, such an initial library is selected on the basis of chemical diversity by some pseudo-random process (for example, the first few plates of a larger library) or by selecting an entire smaller library. These approaches may not produce a sufficient number or diversity of actives. An alternative approach is to select an informer set of screening compounds on the basis of chemogenomic information from previous testing of compounds against a large number of targets. We compare different ways of using chemogenomic data to choose a small informer set of compounds based on previously measured bioactivity data. We develop this Informer-Based-Ranking (IBR) approach using the Published Kinase Inhibitor Sets (PKIS) as the chemogenomic data to select the informer sets. We test the informer compounds on a target that is not part of the chemogenomic data, then predict the activity of the remaining compounds based on the experimental informer data and the chemogenomic data. Through new chemical screening experiments, we demonstrate the utility of IBR strategies in a prospective test on three kinase targets not included in the PKIS.
Huikun Zhang, Spencer S. Ericksen, Ching-pei Lee, Gene E. Ananiev, Nathan Wlodarchak, Julie C. Mitchell, Anthony Gitter, Stephen J. Wright 0001, F. Michael Hoffmann, Scott A. Wildman, Michael A. Newton
PLoS Comput. Biol.8
2018 Network inference reveals novel connections in pathways regulating growth and defense in the yeast salt response
abstract
Cells respond to stressful conditions by coordinating a complex, multi-faceted response that spans many levels of physiology. Much of the response is coordinated by changes in protein phosphorylation. Although the regulators of transcriptome changes during stress are well characterized in Saccharomyces cerevisiae, the upstream regulatory network controlling protein phosphorylation is less well dissected. Here, we developed a computational approach to infer the signaling network that regulates phosphorylation changes in response to salt stress. We developed an approach to link predicted regulators to groups of likely co-regulated phospho-peptides responding to stress, thereby creating new edges in a background protein interaction network. We then use integer linear programming (ILP) to integrate wild type and mutant phospho-proteomic data and predict the network controlling stress-activated phospho-proteomic changes. The network we inferred predicted new regulatory connections between stress-activated and growth-regulating pathways and suggested mechanisms coordinating metabolism, cell-cycle progression, and growth during stress. We confirmed several network predictions with co-immunoprecipitations coupled with mass-spectrometry protein identification and mutant phospho-proteomic analysis. Results show that the cAMP-phosphodiesterase Pde2 physically interacts with many stress-regulated transcription factors targeted by PKA, and that reduced phosphorylation of those factors during stress requires the Rck2 kinase that we show physically interacts with Pde2. Together, our work shows how a high-quality computational network model can facilitate discovery of new pathway interactions during osmotic stress.
Matthew E. MacGilvray, Evgenia Shishkova, Deborah Chasman, Michael Place, Anthony Gitter, Joshua J. Coon, Audrey P. Gasch
PLoS Comput. Biol.5
2016 Network-Based Interpretation of Diverse High-Throughput Datasets through the Omics Integrator Software Package
abstract
High-throughput, 'omic' methods provide sensitive measures of biological responses to perturbations. However, inherent biases in high-throughput assays make it difficult to interpret experiments in which more than one type of data is collected. In this work, we introduce Omics Integrator, a software package that takes a variety of 'omic' data as input and identifies putative underlying molecular pathways. The approach applies advanced network optimization algorithms to a network of thousands of molecular interactions to find high-confidence, interpretable subnetworks that best explain the data. These subnetworks connect changes observed in gene expression, protein abundance or other global assays to proteins that may not have been measured in the screens due to inherent bias or noise in measurement. This approach reveals unannotated molecular pathways that would not be detectable by searching pathway databases. Omics Integrator also provides an elegant framework to incorporate not only positive data, but also negative evidence. Incorporating negative evidence allows Omics Integrator to avoid unexpressed genes and avoid being biased toward highly-studied hub proteins, except when they are strongly implicated by the data. The software is comprised of two individual tools, Garnet and Forest, that can be run together or independently to allow a user to perform advanced integration of multiple types of high-throughput data as well as create condition-specific subnetworks of protein interactions that best connect the observed changes in various datasets. It is available at http://fraenkel.mit.edu/omicsintegrator and on GitHub at https://github.com/fraenkel-lab/OmicsIntegrator.
Nurcan Tuncbag, Sara J. C. Gosline, Amanda J. Kedaigle, Anthony R. Soltis, Anthony Gitter, Ernest Fraenkel
PLoS Comput. Biol.5
2014 Multitask Learning of Signaling and Regulatory Networks with Application to Studying Human Response to Flu
abstract
Reconstructing regulatory and signaling response networks is one of the major goals of systems biology. While several successful methods have been suggested for this task, some integrating large and diverse datasets, these methods have so far been applied to reconstruct a single response network at a time, even when studying and modeling related conditions. To improve network reconstruction we developed MT-SDREM, a multi-task learning method which jointly models networks for several related conditions. In MT-SDREM, parameters are jointly constrained across the networks while still allowing for condition-specific pathways and regulation. We formulate the multi-task learning problem and discuss methods for optimizing the joint target function. We applied MT-SDREM to reconstruct dynamic human response networks for three flu strains: H1N1, H5N1 and H3N2. Our multi-task learning method was able to identify known and novel factors and genes, improving upon prior methods that model each condition independently. The MT-SDREM networks were also better at identifying proteins whose removal affects viral load indicating that joint learning can still lead to accurate, condition-specific, networks. Supporting website with MT-SDREM implementation: http://sb.cs.cmu.edu/mtsdrem.
Siddhartha Jain 0001, Anthony Gitter, Ziv Bar-Joseph
PLoS Comput. Biol.2
2013 Identifying proteins controlling key disease signaling pathways
abstract
MOTIVATION: Several types of studies, including genome-wide association studies and RNA interference screens, strive to link genes to diseases. Although these approaches have had some success, genetic variants are often only present in a small subset of the population, and screens are noisy with low overlap between experiments in different labs. Neither provides a mechanistic model explaining how identified genes impact the disease of interest or the dynamics of the pathways those genes regulate. Such mechanistic models could be used to accurately predict downstream effects of knocking down pathway members and allow comprehensive exploration of the effects of targeting pairs or higher-order combinations of genes. RESULTS: We developed methods to model the activation of signaling and dynamic regulatory networks involved in disease progression. Our model, SDREM, integrates static and time series data to link proteins and the pathways they regulate in these networks. SDREM uses prior information about proteins' likelihood of involvement in a disease (e.g. from screens) to improve the quality of the predicted signaling pathways. We used our algorithms to study the human immune response to H1N1 influenza infection. The resulting networks correctly identified many of the known pathways and transcriptional regulators of this disease. Furthermore, they accurately predict RNA interference effects and can be used to infer genetic interactions, greatly improving over other methods suggested for this task. Applying our method to the more pathogenic H5N1 influenza allowed us to identify several strain-specific targets of this infection. AVAILABILITY: SDREM is available from http://sb.cs.cmu.edu/sdrem. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Anthony Gitter, Ziv Bar-Joseph
Bioinform.1
2012 A Network-based Approach for Predicting Missing Pathway Interactions
abstract
Embedded within large-scale protein interaction networks are signaling pathways that encode response cascades in the cell. Unfortunately, even for well-studied species like S. cerevisiae, only a fraction of all true protein interactions are known, which makes it difficult to reason about the exact flow of signals and the corresponding causal relations in the network. To help address this problem, we introduce a framework for predicting new interactions that aid connectivity between upstream proteins (sources) and downstream transcription factors (targets) of a particular pathway. Our algorithms attempt to globally minimize the distance between sources and targets by finding a small set of shortcut edges to add to the network. Unlike existing algorithms for predicting general protein interactions, by focusing on proteins involved in specific responses our approach homes-in on pathway-consistent interactions. We applied our method to extend pathways in osmotic stress response in yeast and identified several missing interactions, some of which are supported by published reports. We also performed experiments that support a novel interaction not previously reported. Our framework is general and may be applicable to edge prediction problems in other domains.
Saket Navlakha, Anthony Gitter, Ziv Bar-Joseph
PLoS Comput. Biol.2