Dario Ghersi

dblp:33/7244 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-0630-0843ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 TRain: T-cell receptor automated immunoinformatics
abstract
BACKGROUND: The scarcity of available structural data makes characterizing the binding of T-cell receptors (TCRs) to peptide-Major Histocompatibility Complexes (pMHCs) very challenging. The recent surge in sequencing data makes TCRs an ideal target for protein structure modeling. Through these 3D models, researchers can potentially identify key motifs on the TCR's binding regions. Furthermore, computational methods can be employed to pair a TCR structure with a pMHC, leading to predictions of docked TCRpMHC structures. However, going from sequence to predicted 3D TCRpMHC complexes requires a non-trivial amount of steps and specialized immunoinformatics expertise. RESULTS: We developed a Python tool named TRain (T-cell Receptor Automated ImmunoiNformatics) to streamline this process by: (1) converting single-cell sequencing data into full TCR amino acid sequences; (2) efficiently submitting TCR amino acid sequences to existing TCR-specific modeling pipelines; (3) pairing modeled TCR structures with existing crystal structures of pMHC complexes in a non-biased manner before docking; (3) automating the preparation and submission process of TCRs and pMHCs for docking using the RosettaDock tool; and (4) providing scripts to analyze the predicted TCRpMHC interface. We illustrate the basic functionality of TRain with a case study, while further information can be found in a dedicated manual. CONCLUSIONS: We introduced an open-source tool that streamlines going from full TCR sequence information to predicted 3D TCRpMHC complexes, using well-established tools. Analyzing these predicted complexes can provide deeper insights into the binding properties of TCRs, and can help shed light on one of the key steps in adaptive immune responses.
Austin Seamann, Maia Bennett-Boehm, Ryan Ehrlich, Anna Gil, Liisa K. Selin, Dario Ghersi
BMC Bioinform.6
2022 Variant calling enhances the identification of cancer cells in single-cell RNA sequencing data
abstract
Single-cell RNA-sequencing is an invaluable research tool that allows for the investigation of gene expression in heterogeneous cancer cell populations in ways that bulk RNA-seq cannot. However, normal (i.e., non tumor) cells in cancer samples have the potential to confound the downstream analysis of single-cell RNA-seq data. Existing methods for identifying cancer and normal cells include copy number variation inference, marker-gene expression analysis, and expression-based clustering. This work aims to extend the existing approaches for identifying cancer cells in single-cell RNA-seq samples by incorporating variant calling and the identification of putative driver alterations. We found that putative driver alterations can be detected in single-cell RNA-seq data obtained with full-length transcript technologies and noticed that a subset of cells in tumor samples are enriched for putative driver alterations as compared to normal cells. Furthermore, we show that the number of putative driver alterations and inferred copy number variation are not correlated in all samples. Taken together, our findings suggest that augmenting existing cancer-cell filtering methods with variant calling and analysis can increase the number of tumor cells that can be confidently included in downstream analyses of single-cell full-length transcript RNA-seq datasets.
William Gasper, Francesca Rossi 0003, Matteo Ligorio, Dario Ghersi
PLoS Comput. Biol.4
2021 Automatic Extension of Medical Subject Headings (MeSH) Thesaurus to Emerging Research
abstract
The proliferation of information technology infrastructure in recent decades has allowed for unprecedented ease of access to centrally-aggregated scholarly literature and scientific knowledge. This massive aggregation of knowledge requires an information retrieval infrastructure, to include formalized ontologies, that is engineered with careful consideration. A number of domains benefit from the use of hierarchical controlled vocabularies, which may be used to provide a rich set of descriptive terms for characterizing entities in a consistent manner. There are clear benefits to the creation and maintenance of these ontologies: search and retrieval is made easier and analyses of the contained entities are enabled that would not otherwise be possible. However, there may be the opportunity to decrease the manual burden of ontology creation and maintenance with automated methods that leverage natural language processing and other computational techniques. This work presents an automated ontology creation methodology, adapted and expanded from prior work [1], that can produce a topic hierarchy from natural language and may be used to assist in the creation of a novel ontology or the expansion of existing ontologies. The effectiveness of the proposed method is studied using two examples: immunology, an established biomedical domain and a prominent topic in MeSH, and graphene, from the 2D materials domain with wide-ranging biomedical applications, which also has a sparse presence in MeSH
William Gasper, Dario Ghersi, Etienne Z. Gnimpieba, Venkataramana Gadhamshetty, Parvathi Chundi
BIBM3
2021 SwarmTCR: a computational approach to predict the specificity of T cell receptors
abstract
BACKGROUND: With more T cell receptor sequence data becoming available, the need for bioinformatics approaches to predict T cell receptor specificity is even more pressing. Here we present SwarmTCR, a method that uses labeled sequence data to predict the specificity of T cell receptors using a nearest-neighbor approach. SwarmTCR works by optimizing the weights of the individual CDR regions to maximize classification performance. RESULTS: We compared the performance of SwarmTCR against another nearest-neighbor method and showed that SwarmTCR performs well both with bulk sequencing data and with single cell data. In addition, we show that the weights returned by SwarmTCR are biologically interpretable. CONCLUSIONS: Computationally predicting the specificity of T cell receptors can be a powerful tool to shed light on the immune response against infectious diseases and cancers, autoimmunity, cancer immunotherapy, and immunopathology. SwarmTCR is distributed freely under the terms of the GPL-3 license. The source code and all sequencing data are available at GitHub ( https://github.com/thecodingdoc/SwarmTCR ).
Ryan Ehrlich, Larisa Kamga, Anna Gil, Katherine Ruiz De Luzuriaga, Liisa K. Selin, Dario Ghersi
BMC Bioinform.6
2019 Identifying Structural Changes in Correlation Networks Models of Cancer Gene Expression by Stage
abstract
Gene expression analysis using correlation network modeling can help to identify systems-level cellular changes and cooperation among genes. Network modeling is a relatively novel method for comparing changes across stages of cancer, or between primary tumor tissue and the metastatic tissue. In this study, we develop a pipeline to identify the dynamic changes of cancer in gene expression level through time-dependent network analysis of cancer data from The Cancer Genome Atlas (TCGA). A total of 16 correlation networks were built from four (4) stages in four (4) different types of cancers: Thyroid Carcinoma, Colon Adenocarcinoma and Rectum Adenocarcinoma, Stomach Adenocarcinoma, and Kidney Renal Clear Cell Carcinoma. To identify the basic changes in network structure, we performed Jaccard similarity comparison of structurally relevant nodes. We employed mutation analysis to measure and present the time-based changes in mutation rate of genes that are specific to each cancer type. Finally, we present a case study to identify the gene expression changes among primary tumor tissue and metastatic tissue in skin cutaneous melanoma (SKCM).
Qianran Li, Dario Ghersi, Ishwor Thapa, Hesham Ali 0001, Kathryn M. Cooper
BIBM2
2019 FunSet: an open-source software and web server for performing and displaying Gene Ontology enrichment analysis
abstract
BACKGROUND: Gene Ontology enrichment analysis provides an effective way to extract meaningful information from complex biological datasets. By identifying terms that are significantly overrepresented in a gene set, researchers can uncover biological features shared by genes. In addition to extracting enriched terms, it is also important to visualize the results in a way that is conducive to biological interpretation. RESULTS: Here we present FunSet, a new web server to perform and visualize enrichment analysis. The web server identifies Gene Ontology terms that are statistically overrepresented in a target set with respect to a background set. The enriched terms are displayed in a 2D plot that captures the semantic similarity between terms, with the option to cluster terms via spectral clustering and identify a representative term for each cluster. FunSet can be used interactively or programmatically, and allows users to download the enrichment results both in tabular form and in graphical form as SVG files or in data format as JSON or csv. To enhance reproducibility of the analyses, users have access to historical data for the ontology and the annotations. The source code for the standalone program and the web server are made available with an open-source license.
Matthew L. Hale, Ishwor Thapa, Dario Ghersi
BMC Bioinform.3
2019 Uncovering and characterizing splice variants associated with survival in lung cancer patients
abstract
Splice variants have been shown to play an important role in tumor initiation and progression and can serve as novel cancer biomarkers. However, the clinical importance of individual splice variants and the mechanisms by which they can perturb cellular functions are still poorly understood. To address these issues, we developed an efficient and robust computational method to: (1) identify splice variants that are associated with patient survival in a statistically significant manner; and (2) predict rewired protein-protein interactions that may result from altered patterns of expression of such variants. We applied our method to the lung adenocarcinoma dataset from TCGA and identified splice variants that are significantly associated with patient survival and can alter protein-protein interactions. Among these variants, several are implicated in DNA repair through homologous recombination. To computationally validate our findings, we characterized the mutational signatures in patients, grouped by low and high expression of a splice variant associated with patient survival and involved in DNA repair. The results of the mutational signature analysis are in agreement with the molecular mechanism suggested by our method. To the best of our knowledge, this is the first attempt to build a computational approach to systematically identify splice variants associated with patient survival that can also generate experimentally testable, mechanistic hypotheses. Code for identifying survival-significant splice variants using the Null Empirically Estimated P-value method can be found at https://github.com/thecodingdoc/neep. Code for construction of Multi-Granularity Graphs to discover potential rewired protein interactions can be found at https://github.com/scwest/SINBAD.
Sean West, Surinder K. Batra, Hesham Ali 0001, Dario Ghersi
PLoS Comput. Biol.5
2018 Two critical positions in zinc finger domains are heavily mutated in three human cancer types
abstract
A major goal of cancer genomics is to identify somatic mutations that play a role in tumor initiation or progression. Somatic mutations within transcription factors are of particular interest, as gene expression dysregulation is widespread in cancers. The substantial gene expression variation evident across tumors suggests that numerous regulatory factors are likely to be involved and that somatic mutations within them may not occur at high frequencies across patient cohorts, thereby complicating efforts to uncover which ones are cancer-relevant. Here we analyze somatic mutations within the largest family of human transcription factors, namely those that bind DNA via Cys2His2 zinc finger domains. Specifically, to hone in on important mutations within these genes, we aggregated somatic mutations across all of them by their positions within Cys2His2 zinc finger domains. Remarkably, we found that for three classes of cancers profiled by The Cancer Genome Atlas (TCGA)-Uterine Corpus Endometrial Carcinoma, Colon and Rectal Adenocarcinomas, and Skin Cutaneous Melanoma-two specific, functionally important positions within zinc finger domains are mutated significantly more often than expected by chance, with alterations in 18%, 10% and 43% of tumors, respectively. Numerous zinc finger genes are affected, with those containing Krüppel-associated box (KRAB) repressor domains preferentially targeted by these mutations. Further, the genes with these mutations also have high overall missense mutation rates, are expressed at levels comparable to those of known cancer genes, and together have biological process annotations that are consistent with roles in cancers. Altogether, we introduce evidence broadly implicating mutations within a diverse set of zinc finger proteins as relevant for cancer, and propose that they contribute to the widespread transcriptional dysregulation observed in cancer cells.
Daniel Munro, Dario Ghersi, Mona Singh 0001
PLoS Comput. Biol.2
2017 In-silico analysis of the "memory anti-Naïve" effect in anti-viral cross-reactive responses
abstract
Of the examples of clonal competition for antigen among lymphocytes, the recently predicted “Memory anti-Naïve” phenomenon occurs when the challenging antigen is not identical to the priming, and will be consequently bound with lower avidity by preexisting memory cells. In this study we use computer modeling and a systematic schedule of viral injections to disentangle the complex relationship between different lineages of effector T cells in the presence of viruses. We measure the antiviral efficiency of memory cells as well as their dominance over naïve cells as a function of the antigenic distance between first and second infection. Our simulation show that at a critical range of antigenic distance memory cells, now unable to clear the infection, can however block the surge of naïve clones thus preventing an effective immune response. This finding motivate us to propose the Memory anti-Naïve phenomenon as the causative mechanism for the classic Original Antigenic Sin phenomenon described in the literature which occurs irregularly in returning pandemics, and also for the less glamorous, but certainly numerous and severe, cases of misfired vaccinations, and viral escapes.
Filippo Castiglione, Dario Ghersi, Franco Celada
BIBM2
2017 Analyzing T cell receptor alpha/beta usage in binding to the pMHC
abstract
T cells play a critical role in the adaptive immune response. They perform their function by recognizing infected cells presenting peptides on a specialized complex known as the MHC. The recognition process involves binding of the peptide-loaded MHC to the T cell receptor (TCR), a surface molecule comprised of an alpha and a beta chain. A large body of evidence suggests that T cells can respond to previously unseen pathogens, a phenomenon known as cross-reactivity. Cross-reactivity has important medical implications, as cross-reactive responses can be either protective or lead to disease. A possible mechanism that has been proposed to explain cross-reactivity is the differential usage of the alpha and beta chains, whereas one peptide can be recognized predominantly by the alpha chain and a different peptide by the beta chain. In this study we carry out a systematic analysis of a non-redundant set of 67 crystal structures, measuring TCR alpha/beta usage and its relationship with structural features of the interaction. Our results show a wide range of TCR alpha/beta usage in different complexes. Further, we find that alpha/beta usage significantly correlates with one of the docking angles between the TCR and the MHC.
Ryan Ehrlich, Dario Ghersi
BIBM2
2016 Forecasting the Spread of Mosquito-Borne Disease using Publicly Accessible Data: A Case Study in Chikungunya
Kathryn M. Cooper, Dhundy Bastola, Robin A. Gandhi, Dario Ghersi, Steven H. Hinrichs, Marsha Morien, Ann L. Fruhling
AMIA4
2015 A novel approach to identify shared fragments in drugs and natural products
abstract
Fragment-based approaches have now become an important component of the drug discovery process. At the same time, pharmaceutical chemists are more often turning to the natural world and its extremely large and diverse collection of natural compounds to discover new leads that can potentially be turned into drugs. In this study we introduce and discuss a computational pipeline to automatically extract statistically overrepresented chemical fragments in therapeutic classes, and search for similar fragments in a large database of natural products. By systematically identifying enriched fragments in therapeutic groups, we are able to extract and focus on few fragments that are likely to be active or structurally important as scaffolds. We show that several therapeutic classes (including antibacterial, antineoplastic, and drugs active on the cardiovascular system, among others) have enriched fragments that are also found in many natural compounds. Further, our method is able to detect fragments shared by a drug and a natural product even when the global similarity between the two molecules is generally low. A further development of this computational pipeline is to help predict putative therapeutic activities of natural compounds, and to help identify novel leads for drug discovery.
Akshay Balasubramanya, Ishwor Thapa, Dhundy Bastola, Dario Ghersi
BIBM4
2015 Genome-Wide Detection and Analysis of Multifunctional Genes
abstract
Many genes can play a role in multiple biological processes or molecular functions. Identifying multifunctional genes at the genome-wide level and studying their properties can shed light upon the complexity of molecular events that underpin cellular functioning, thereby leading to a better understanding of the functional landscape of the cell. However, to date, genome-wide analysis of multifunctional genes (and the proteins they encode) has been limited. Here we introduce a computational approach that uses known functional annotations to extract genes playing a role in at least two distinct biological processes. We leverage functional genomics data sets for three organisms--H. sapiens, D. melanogaster, and S. cerevisiae--and show that, as compared to other annotated genes, genes involved in multiple biological processes possess distinct physicochemical properties, are more broadly expressed, tend to be more central in protein interaction networks, tend to be more evolutionarily conserved, and are more likely to be essential. We also find that multifunctional genes are significantly more likely to be involved in human disorders. These same features also hold when multifunctionality is defined with respect to molecular functions instead of biological processes. Our analysis uncovers key features about multifunctional genes, and is a step towards a better genome-wide understanding of gene multifunctionality.
Yuri Pritykin, Dario Ghersi, Mona Singh 0001
PLoS Comput. Biol.2
2014 molBLOCKS: decomposing small molecule sets and uncovering enriched fragments
abstract
UNLABELLED: The chemical structures of biomolecules, whether naturally occurring or synthetic, are composed of functionally important building blocks. Given a set of small molecules-for example, those known to bind a particular protein-computationally decomposing them into chemically meaningful fragments can help elucidate their functional properties, and may be useful for designing novel compounds with similar properties. Here we introduce molBLOCKS, a suite of programs for breaking down sets of small molecules into fragments according to a predefined set of chemical rules, clustering the resulting fragments, and uncovering statistically enriched fragments. Among other applications, our software should be a great aid in large-scale chemical analysis of ligands binding specific targets of interest. AVAILABILITY AND IMPLEMENTATION: molBLOCKS is available as GPL C++ source code at http://compbio.cs.princeton.edu/molblocks.
Dario Ghersi, Mona Singh 0001
Bioinform.1
2009 EASYMIFS and SITEHOUND: a toolkit for the identification of ligand-binding sites in protein structures
abstract
UNLABELLED: SiteHound uses Molecular Interaction Fields (MIFs) produced by EasyMIFs to identify protein structure regions that show a high propensity for interaction with ligands. The type of binding site identified depends on the probe atom used in the MIF calculation. The input to EasyMIFs is a PDB file of a protein structure; the output MIF serves as input to SiteHound, which in turn produces a list of putative binding sites. Extensive testing of SiteHound for the detection of binding sites for drug-like molecules and phosphorylated ligands has been carried out. AVAILABILITY: EasyMIFs and SiteHound executables for Linux, Mac OS X, and MS Windows operating systems are freely available for download from http://sitehound.sanchezlab.org/download.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dario Ghersi, Roberto Sanchez
Bioinform.1