VLDB 2026 Research / reviewers in the wild / expert
Sameer Velankar
dblp:80/200 · also Samir S. Velankar
· DBLP profile ↗
10ranked-venue papers
0as first author
5since 2021 · last 2023
0000-0002-8439-5964ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | PDBImages: a command-line tool for automated macromolecular structure visualizationabstractSUMMARY: PDBImages is an innovative, open-source Node.js package that harnesses the power of the popular macromolecule structure visualization software Mol*. Designed for use by the scientific community, PDBImages provides a means to generate high-quality images for PDB and AlphaFold DB models. Its unique ability to render and save images directly to files in a browserless mode sets it apart, offering users a streamlined, automated process for macromolecular structure visualization. Here, we detail the implementation of PDBImages, enumerating its diverse image types, and elaborating on its user-friendly setup. This powerful tool opens a new gateway for researchers to visualize, analyse, and share their work, fostering a deeper understanding of bioinformatics. AVAILABILITY AND IMPLEMENTATION: PDBImages is available as an npm package from https://www.npmjs.com/package/pdb-images. The source code is available from https://github.com/PDBeurope/pdb-images. Adam Midlik, Sreenath Nair, Stephen Anyango, Mandar S. Deshpande, David Sehnal, Mihaly Varadi, Sameer Velankar |
Bioinform. | 7 |
| 2022 | Characterizing and explaining the impact of disease-associated mutations in proteins without known structures or structural homologsabstractMutations in human proteins lead to diseases. The structure of these proteins can help understand the mechanism of such diseases and develop therapeutics against them. With improved deep learning techniques, such as RoseTTAFold and AlphaFold, we can predict the structure of proteins even in the absence of structural homologs. We modeled and extracted the domains from 553 disease-associated human proteins without known protein structures or close homologs in the Protein Databank. We noticed that the model quality was higher and the Root mean square deviation (RMSD) lower between AlphaFold and RoseTTAFold models for domains that could be assigned to CATH families as compared to those which could only be assigned to Pfam families of unknown structure or could not be assigned to either. We predicted ligand-binding sites, protein-protein interfaces and conserved residues in these predicted structures. We then explored whether the disease-associated missense mutations were in the proximity of these predicted functional sites, whether they destabilized the protein structure based on ddG calculations or whether they were predicted to be pathogenic. We could explain 80% of these disease-associated mutations based on proximity to functional sites, structural destabilization or pathogenicity. When compared to polymorphisms, a larger percentage of disease-associated missense mutations were buried, closer to predicted functional sites, predicted as destabilizing and pathogenic. Usage of models from the two state-of-the-art techniques provide better confidence in our predictions, and we explain 93 additional mutations based on RoseTTAFold models which could not be explained based solely on AlphaFold models. Neeladri Sen, Ivan Anishchenko, Nicola Bordin, Ian Sillitoe, Sameer Velankar, David Baker 0001, Christine A. Orengo |
Briefings Bioinform. | 5 |
| 2021 | The impact of structural bioinformatics tools and resources on SARS-CoV-2 research and therapeutic strategiesabstractSARS-CoV-2 is the causative agent of COVID-19, the ongoing global pandemic. It has posed a worldwide challenge to human health as no effective treatment is currently available to combat the disease. Its severity has led to unprecedented collaborative initiatives for therapeutic solutions against COVID-19. Studies resorting to structure-based drug design for COVID-19 are plethoric and show good promise. Structural biology provides key insights into 3D structures, critical residues/mutations in SARS-CoV-2 proteins, implicated in infectivity, molecular recognition and susceptibility to a broad range of host species. The detailed understanding of viral proteins and their complexes with host receptors and candidate epitope/lead compounds is the key to developing a structure-guided therapeutic design. Since the discovery of SARS-CoV-2, several structures of its proteins have been determined experimentally at an unprecedented speed and deposited in the Protein Data Bank. Further, specialized structural bioinformatics tools and resources have been developed for theoretical models, data on protein dynamics from computer simulations, impact of variants/mutations and molecular therapeutics. Here, we provide an overview of ongoing efforts on developing structural bioinformatics tools and resources for COVID-19 research. We also discuss the impact of these resources and structure-based studies, to understand various aspects of SARS-CoV-2 infection and therapeutic development. These include (i) understanding differences between SARS-CoV-2 and SARS-CoV, leading to increased infectivity of SARS-CoV-2, (ii) deciphering key residues in the SARS-CoV-2 involved in receptor-antibody recognition, (iii) analysis of variants in host proteins that affect host susceptibility to infection and (iv) analyses facilitating structure-based drug and vaccine design against SARS-CoV-2. Vaishali P. Waman, Neeladri Sen, Mihaly Varadi, Antoine Daina, Shoshana J. Wodak, Vincent Zoete, Sameer Velankar, Christine A. Orengo |
Briefings Bioinform. | 7 |
| 2021 | PDBe aggregated API: programmatic access to an integrative knowledge graph of molecular structure dataabstractSUMMARY: The PDBe aggregated API is an open-access and open-source RESTful API that provides programmatic access to a wealth of macromolecular structural data and their functional and biophysical annotations through 80+ API endpoints. The API is powered by the PDBe graph database (https://pdbe.org/graph-schema), an open-access integrative knowledge graph that can be used as a discovery tool to answer complex biological questions. AVAILABILITY AND IMPLEMENTATION: The PDBe aggregated API provides up-to-date access to the PDBe graph database, which has weekly releases with the latest data from the Protein Data Bank, integrated with updated annotations from UniProt, Pfam, CATH, SCOP and the PDBe-KB partner resources. The complete list of all the available API endpoints and their descriptions are available at https://pdbe.org/graph-api. The source code of the Python 3.6+ API application is publicly available at https://gitlab.ebi.ac.uk/pdbe-kb/services/pdbe-graph-api. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sreenath Nair, Mihaly Varadi, Nurul Nadzirin, Lukás Pravda, Stephen Anyango, Saqib Mir, John M. Berrisford, David R. Armstrong, Aleksandras Gutmanas, Sameer Velankar |
Bioinform. | 10 |
| 2021 | PDBeCIF: an open-source mmCIF/CIF parsing and processing packageabstractBACKGROUND: Biomacromolecular structural data outgrew the legacy Protein Data Bank (PDB) format which the scientific community relied on for decades, yet the use of its successor PDBx/Macromolecular Crystallographic Information File format (PDBx/mmCIF) is still not widespread. Perhaps one of the reasons is the availability of easy to use tools that only support the legacy format, but also the inherent difficulties of processing mmCIF files correctly, given the number of edge cases that make efficient parsing problematic. Nevertheless, to fully exploit macromolecular structure data and their associated annotations such as multiscale structures from integrative/hybrid methods or large macromolecular complexes determined using traditional methods, it is necessary to fully adopt the new format as soon as possible. RESULTS: To this end, we developed PDBeCIF, an open-source Python project for manipulating mmCIF and CIF files. It is part of the official list of mmCIF parsers recorded by the wwPDB and is heavily employed in the processes of the Protein Data Bank in Europe. The package is freely available both from the PyPI repository ( http://pypi.org/project/pdbecif ) and from GitHub ( https://github.com/pdbeurope/pdbecif ) along with rich documentation and many ready-to-use examples. CONCLUSIONS: PDBeCIF is an efficient and lightweight Python 2.6+/3+ package with no external dependencies. It can be readily integrated with 3rd party libraries as well as adopted for broad scientific analyses. Glen van Ginkel, Lukás Pravda, Jose M. Dana, Mihaly Varadi, Peter A. Keller, Stephen Anyango, Sameer Velankar |
BMC Bioinform. | 7 |
| 2020 | The ELIXIR Core Data Resources: fundamental infrastructure for the life sciencesabstractSUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rachel Drysdale, Charles E. Cook, Robert Petryszak, Vivienne Baillie Gerritsen, Mary Barlow, Elisabeth Gasteiger, Franziska Gruhl, Jerry Lanfear, Rodrigo Lopez, Nicole Redaschi, Heinz Stockinger, Daniel Teixeira, Aravind Venkatesan, Alex Bateman, Alan J. Bridge, Guy Cochrane, Robert D. Finn, Frank Oliver Glöckner, Marc Hanauer, Thomas M. Keane, Luana Licata, Per Oksvold, Sandra E. Orchard, Christine A. Orengo, Helen E. Parkinson, Bengt Persson, Pablo Porras, Jordi Rambla De Argila, Ana Rath, Charlotte Rodwell, Ugis Sarkans, Dietmar Schomburg, Ian Sillitoe, J. Dylan Spalding, Mathias Uhlen, Sameer Velankar, Juan Antonio Vizcaíno, Kalle von Feilitzen, Christian von Mering, Andy Yates, Niklas Blomberg, Christine Durinx, Johanna R. McEntyre |
Bioinform. | 38 |
| 2020 | BinaryCIF and CIFTools - Lightweight, efficient and extensible macromolecular data managementabstract3D macromolecular structural data is growing ever more complex and plentiful in the wake of substantive advances in experimental and computational structure determination methods including macromolecular crystallography, cryo-electron microscopy, and integrative methods. Efficient means of working with 3D macromolecular structural data for archiving, analyses, and visualization are central to facilitating interoperability and reusability in compliance with the FAIR Principles. We address two challenges posed by growth in data size and complexity. First, data size is reduced by bespoke compression techniques. Second, complexity is managed through improved software tooling and fully leveraging available data dictionary schemas. To this end, we introduce BinaryCIF, a serialization of Crystallographic Information File (CIF) format files that maintains full compatibility to related data schemas, such as PDBx/mmCIF, while reducing file sizes by more than a factor of two versus gzip compressed CIF files. Moreover, for the largest structures, BinaryCIF provides even better compression-factor ten and four versus CIF files and gzipped CIF files, respectively. Herein, we describe CIFTools, a set of libraries in Java and TypeScript for generic and typed handling of CIF and BinaryCIF files. Together, BinaryCIF and CIFTools enable lightweight, efficient, and extensible handling of 3D macromolecular structural data. David Sehnal, Sebastian Bittrich, Sameer Velankar, Jaroslav Koca, Radka Svobodová Vareková, Stephen K. Burley, Alexander S. Rose |
PLoS Comput. Biol. | 3 |
| 2019 | Finding enzyme cofactors in Protein Data BankabstractMOTIVATION: Cofactors are essential for many enzyme reactions. The Protein Data Bank (PDB) contains >67 000 entries containing enzyme structures, many with bound cofactor or cofactor-like molecules. This work aims to identify and categorize these small molecules in the PDB and make it easier to find them. RESULTS: The Protein Data Bank in Europe (PDBe; pdbe.org) has implemented a pipeline to identify enzyme cofactor and cofactor-like molecules, which are now part of the PDBe weekly release process. AVAILABILITY AND IMPLEMENTATION: Information is made available on the individual PDBe entry pages at pdbe.org and programmatically through the PDBe REST API (pdbe.org/api). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Abhik Mukhopadhyay, Neera Borkakoti, Lukás Pravda, Jonathan D. Tyzack, Janet M. Thornton, Sameer Velankar |
Bioinform. | 6 |
| 2015 | The chemical component dictionary: complete descriptions of constituent molecules in experimentally determined 3D macromolecules in the Protein Data BankabstractUNLABELLED: The Chemical Component Dictionary (CCD) is a chemical reference data resource that describes all residue and small molecule components found in Protein Data Bank (PDB) entries. The CCD contains detailed chemical descriptions for standard and modified amino acids/nucleotides, small molecule ligands and solvent molecules. Each chemical definition includes descriptions of chemical properties such as stereochemical assignments, chemical descriptors, systematic chemical names and idealized coordinates. The content, preparation, validation and distribution of this CCD chemical reference dataset are described. AVAILABILITY AND IMPLEMENTATION: The CCD is updated regularly in conjunction with the scheduled weekly release of new PDB structure data. The CCD and amino acid variant reference datasets are hosted in the public PDB ftp repository at ftp://ftp.wwpdb.org/pub/pdb/data/monomers/components.cif.gz, ftp://ftp.wwpdb.org/pub/pdb/data/monomers/aa-variants-v1.cif.gz, and its mirror sites, and can be accessed from http://wwpdb.org. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. John D. Westbrook, Chenghua Shao, Zukang Feng, Marina Zhuravleva, Sameer Velankar, Jasmine Young |
Bioinform. | 5 |
| 2013 | BioJS: an open source JavaScript framework for biological data visualizationabstractSUMMARY: BioJS is an open-source project whose main objective is the visualization of biological data in JavaScript. BioJS provides an easy-to-use consistent framework for bioinformatics application programmers. It follows a community-driven standard specification that includes a collection of components purposely designed to require a very simple configuration and installation. In addition to the programming framework, BioJS provides a centralized repository of components available for reutilization by the bioinformatics community. AVAILABILITY AND IMPLEMENTATION: http://code.google.com/p/biojs/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. John Gómez, Leyla Jael Castro, Gustavo A. Salazar, Jose M. Villaveces, Swanand P. Gore, Alexander García Castro, Maria Jesus Martin, Guillaume Launay, Rafael Alcántara, Noemi del-Toro, Marine Sivade, Sandra E. Orchard, Sameer Velankar, Henning Hermjakob, Chenggong Zong, Peipei Ping, Manuel Corpas, Rafael C. Jiménez |
Bioinform. | 13 |