Yana Rose

dblp:290/0650 · also Yana Valasatava · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0003-1018-5718ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Multi-scale structural similarity embedding search across entire proteomes
abstract
MOTIVATION: The rapid expansion of three-dimensional (3D) biomolecular structure information, driven by breakthroughs in artificial intelligence/deep learning (AI/DL)-based structure predictions, has created an urgent need for scalable and efficient structure similarity search methods. Traditional alignment-based approaches, such as structural superposition tools, are computationally expensive and challenging to scale with the vast number of available macromolecular structures. RESULTS: Herein, we present a scalable structure similarity search strategy designed to navigate extensive repositories of experimentally determined structures and computed structure models predicted using AI/DL methods. Our approach leverages protein language models and a deep neural network architecture to transform 3D structures into fixed-length vectors, enabling efficient large-scale comparisons. Although trained to predict TM-scores between single-domain structures, our model generalizes beyond the domain level, accurately identifying 3D similarity for full-length polypeptide chains and multimeric assemblies. By integrating vector databases, our method facilitates efficient large-scale structure retrieval, addressing the growing challenges posed by the expanding volume of 3D biostructure information. AVAILABILITY AND IMPLEMENTATION: Source code available at https://github.com/bioinsilico/rcsb-embedding-search. Source code DOI: https://doi.org/10.6084/m9.figshare.30546698.v1. Benchmark datasets DOI: https://doi.org/10.6084/m9.figshare.30546650.v1. Web server prototype available at: http://embedding-search.rcsb.org/.
Joan Segura, Rubén Sánchez García, Sebastian Bittrich, Yana Rose, Stephen K. Burley, Jose M. Duarte
Bioinform.4
2024 RCSB protein Data Bank: exploring protein 3D similarities via comprehensive structural alignments
abstract
MOTIVATION: Tools for pairwise alignments between 3D structures of proteins are of fundamental importance for structural biology and bioinformatics, enabling visual exploration of evolutionary and functional relationships. However, the absence of a user-friendly, browser-based tool for creating alignments and visualizing them at both 1D sequence and 3D structural levels makes this process unnecessarily cumbersome. RESULTS: We introduce a novel pairwise structure alignment tool (rcsb.org/alignment) that seamlessly integrates into the RCSB Protein Data Bank (RCSB PDB) research-focused RCSB.org web portal. Our tool and its underlying application programming interface (alignment.rcsb.org) empowers users to align several protein chains with a reference structure by providing access to established alignment algorithms (FATCAT, CE, TM-align, or Smith-Waterman 3D). The user-friendly interface simplifies parameter setup and input selection. Within seconds, our tool enables visualization of results in both sequence (1D) and structural (3D) perspectives through the RCSB PDB RCSB.org Sequence Annotations viewer and Mol* 3D viewer, respectively. Users can effortlessly compare structures deposited in the PDB archive alongside more than a million incorporated Computed Structure Models coming from the ModelArchive and AlphaFold DB. Moreover, this tool can be used to align custom structure data by providing a link/URL or uploading atomic coordinate files directly. Importantly, alignment results can be bookmarked and shared with collaborators. By bridging the gap between 1D sequence and 3D structures of proteins, our tool facilitates deeper understanding of complex evolutionary relationships among proteins through comprehensive sequence and structural analyses. AVAILABILITY AND IMPLEMENTATION: The alignment tool is part of the RCSB PDB research-focused RCSB.org web portal and available at rcsb.org/alignment. Programmatic access is available via alignment.rcsb.org. Frontend code has been published at github.com/rcsb/rcsb-pecos-app. Visualization is powered by the open-source Mol* viewer (github.com/molstar/molstar and github.com/molstar/rcsb-molstar) plus the Sequence Annotations in 3D Viewer (github.com/rcsb/rcsb-saguaro-3d).
Sebastian Bittrich, Joan Segura, Jose M. Duarte, Stephen K. Burley, Yana Rose
Bioinform.5
2022 RCSB Protein Data Bank: improved annotation, search and visualization of membrane protein structures archived in the PDB
abstract
MOTIVATION: Membrane proteins are encoded by approximately one fifth of human genes but account for more than half of all US FDA approved drug targets. Thanks to new technological advances, the number of membrane proteins archived in the PDB is growing rapidly. However, automatic identification of membrane proteins or inference of membrane location is not a trivial task. RESULTS: We present recent improvements to the RCSB Protein Data Bank web portal (RCSB PDB, rcsb.org) that provide a wealth of new membrane protein annotations integrated from four external resources: OPM, PDBTM, MemProtMD and mpstruc. We have substantially enhanced the presentation of data on membrane proteins. The number of membrane proteins with annotations available on rcsb.org was increased by ∼80%. Users can search for these annotations, explore corresponding tree hierarchies, display membrane segments at the 1D amino acid sequence level, and visualize the predicted location of the membrane layer in 3D. AVAILABILITY AND IMPLEMENTATION: Annotations, search, tree data and visualization are available at our rcsb.org web portal. Membrane visualization is supported by the open-source Mol* viewer (molstar.org and github.com/molstar/molstar). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sebastian Bittrich, Yana Rose, Joan Segura, John D. Westbrook, Jose M. Duarte, Stephen K. Burley
Bioinform.2
2022 RCSB Protein Data Bank 1D3D module: displaying positional features on macromolecular assemblies
abstract
MOTIVATION: Mapping positional features from one-dimensional (1D) sequences onto three-dimensional (3D) structures of biological macromolecules is a powerful tool to show geometric patterns of biochemical annotations and provide a better understanding of the mechanisms underpinning protein and nucleic acid function at the atomic level. RESULTS: We present a new library designed to display fully customizable interactive views between 1D positional features of protein and/or nucleic acid sequences and their 3D structures as isolated chains or components of macromolecular assemblies. AVAILABILITY AND IMPLEMENTATION: https://github.com/rcsb/rcsb-saguaro-3d. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Joan Segura, Yana Rose, Sebastian Bittrich, Stephen K. Burley, Jose M. Duarte
Bioinform.2
2021 RCSB Protein Data Bank 1D tools and services
abstract
MOTIVATION: Interoperability between polymer sequences and structural data is essential for providing a complete picture of protein and gene features and helping to understand biomolecular function. RESULTS: Herein, we present two resources designed to improve interoperability between the RCSB Protein Data Bank, the NCBI and the UniProtKB data resources and visualize integrated data therefrom. The underlying tools provide a flexible means of mapping between the different coordinate spaces and an interactive tool allows convenient visualization of the 1-dimensional data over the web. AVAILABILITYAND IMPLEMENTATION: https://1d-coordinates.rcsb.org and https://rcsb.github.io/rcsb-saguaro. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Joan Segura, Yana Rose, John D. Westbrook, Stephen K. Burley, Jose M. Duarte
Bioinform.2
2019 BioJava 5: A community driven open-source bioinformatics library
abstract
BioJava is an open-source project that provides a Java library for processing biological data. The project aims to simplify bioinformatic analyses by implementing parsers, data structures, and algorithms for common tasks in genomics, structural biology, ontologies, phylogenetics, and more. Since 2012, we have released two major versions of the library (4 and 5) that include many new features to tackle challenges with increasingly complex macromolecular structure data. BioJava requires Java 8 or higher and is freely available under the LGPL 2.1 license. The project is hosted on GitHub at https://github.com/biojava/biojava. More information and documentation can be found online on the BioJava website (http://www.biojava.org) and tutorial (https://github.com/biojava/biojava-tutorial). All inquiries should be directed to the GitHub page or the BioJava mailing list (http://lists.open-bio.org/mailman/listinfo/biojava-l).
Aleix Lafita, Spencer Bliven, Andreas Prlic, Dmytro Guzenko, Peter W. Rose, Anthony R. Bradley, Paolo Pavan, Douglas Myers-Turnbull, Yana Rose, Michael L. Heuer, Matt Larson, Stephen K. Burley, Jose M. Duarte
PLoS Comput. Biol.9
2018 NGL viewer: web-based molecular graphics for large complexes
abstract
Motivation: The interactive visualization of very large macromolecular complexes on the web is becoming a challenging problem as experimental techniques advance at an unprecedented rate and deliver structures of increasing size. Results: We have tackled this problem by developing highly memory-efficient and scalable extensions for the NGL WebGL-based molecular viewer and by using Macromolecular Transmission Format (MMTF), a binary and compressed MMTF. These enable NGL to download and render molecular complexes with millions of atoms interactively on desktop computers and smartphones alike, making it a tool of choice for web-based molecular visualization in research and education. Availability and implementation: The source code is freely available under the MIT license at github.com/arose/ngl and distributed on NPM (npmjs.com/package/ngl). MMTF-JavaScript encoders and decoders are available at github.com/rcsb/mmtf-javascript.
Alexander S. Rose, Anthony R. Bradley, Yana Rose, Jose M. Duarte, Andreas Prlic, Peter W. Rose
Bioinform.3
2017 MMTF - An efficient file format for the transmission, visualization, and analysis of macromolecular structures
abstract
Recent advances in experimental techniques have led to a rapid growth in complexity, size, and number of macromolecular structures that are made available through the Protein Data Bank. This creates a challenge for macromolecular visualization and analysis. Macromolecular structure files, such as PDB or PDBx/mmCIF files can be slow to transfer, parse, and hard to incorporate into third-party software tools. Here, we present a new binary and compressed data representation, the MacroMolecular Transmission Format, MMTF, as well as software implementations in several languages that have been developed around it, which address these issues. We describe the new format and its APIs and demonstrate that it is several times faster to parse, and about a quarter of the file size of the current standard format, PDBx/mmCIF. As a consequence of the new data representation, it is now possible to visualize structures with millions of atoms in a web browser, keep the whole PDB archive in memory or parse it within few minutes on average computers, which opens up a new way of thinking how to design and implement efficient algorithms in structural bioinformatics. The PDB archive is available in MMTF file format through web services and data that are updated on a weekly basis.
Anthony R. Bradley, Alexander S. Rose, Antonín Pavelka, Yana Rose, Jose M. Duarte, Andreas Prlic, Peter W. Rose
PLoS Comput. Biol.4
2016 MetalPredator: a web server to predict iron-sulfur cluster binding proteomes
abstract
MOTIVATION: The prediction of the iron-sulfur proteome is highly desirable for biomedical and biological research but a freely available tool to predict iron-sulfur proteins has not been developed yet. RESULTS: We developed a web server to predict iron-sulfur proteins from protein sequence(s). This tool, called MetalPredator, is able to process complete proteomes rapidly with high recall and precision. AVAILABILITY AND IMPLEMENTATION: The web server is freely available at: http://metalweb.cerm.unifi.it/tools/metalpredator/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yana Rose, Antonio Rosato, Lucia Banci, Claudia Andreini
Bioinform.1