VLDB 2026 Research / reviewers in the wild / expert
Joan Segura
dblp:12/6080
· DBLP profile ↗
16ranked-venue papers
8as first author
5since 2021 · last 2026
0000-0001-5593-735XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 8 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-scale structural similarity embedding search across entire proteomesabstractMOTIVATION: The rapid expansion of three-dimensional (3D) biomolecular structure information, driven by breakthroughs in artificial intelligence/deep learning (AI/DL)-based structure predictions, has created an urgent need for scalable and efficient structure similarity search methods. Traditional alignment-based approaches, such as structural superposition tools, are computationally expensive and challenging to scale with the vast number of available macromolecular structures. RESULTS: Herein, we present a scalable structure similarity search strategy designed to navigate extensive repositories of experimentally determined structures and computed structure models predicted using AI/DL methods. Our approach leverages protein language models and a deep neural network architecture to transform 3D structures into fixed-length vectors, enabling efficient large-scale comparisons. Although trained to predict TM-scores between single-domain structures, our model generalizes beyond the domain level, accurately identifying 3D similarity for full-length polypeptide chains and multimeric assemblies. By integrating vector databases, our method facilitates efficient large-scale structure retrieval, addressing the growing challenges posed by the expanding volume of 3D biostructure information. AVAILABILITY AND IMPLEMENTATION: Source code available at https://github.com/bioinsilico/rcsb-embedding-search. Source code DOI: https://doi.org/10.6084/m9.figshare.30546698.v1. Benchmark datasets DOI: https://doi.org/10.6084/m9.figshare.30546650.v1. Web server prototype available at: http://embedding-search.rcsb.org/. Joan Segura, Rubén Sánchez García, Sebastian Bittrich, Yana Rose, Stephen K. Burley, Jose M. Duarte |
Bioinform. | 1 |
| 2024 | RCSB protein Data Bank: exploring protein 3D similarities via comprehensive structural alignmentsabstractMOTIVATION: Tools for pairwise alignments between 3D structures of proteins are of fundamental importance for structural biology and bioinformatics, enabling visual exploration of evolutionary and functional relationships. However, the absence of a user-friendly, browser-based tool for creating alignments and visualizing them at both 1D sequence and 3D structural levels makes this process unnecessarily cumbersome. RESULTS: We introduce a novel pairwise structure alignment tool (rcsb.org/alignment) that seamlessly integrates into the RCSB Protein Data Bank (RCSB PDB) research-focused RCSB.org web portal. Our tool and its underlying application programming interface (alignment.rcsb.org) empowers users to align several protein chains with a reference structure by providing access to established alignment algorithms (FATCAT, CE, TM-align, or Smith-Waterman 3D). The user-friendly interface simplifies parameter setup and input selection. Within seconds, our tool enables visualization of results in both sequence (1D) and structural (3D) perspectives through the RCSB PDB RCSB.org Sequence Annotations viewer and Mol* 3D viewer, respectively. Users can effortlessly compare structures deposited in the PDB archive alongside more than a million incorporated Computed Structure Models coming from the ModelArchive and AlphaFold DB. Moreover, this tool can be used to align custom structure data by providing a link/URL or uploading atomic coordinate files directly. Importantly, alignment results can be bookmarked and shared with collaborators. By bridging the gap between 1D sequence and 3D structures of proteins, our tool facilitates deeper understanding of complex evolutionary relationships among proteins through comprehensive sequence and structural analyses. AVAILABILITY AND IMPLEMENTATION: The alignment tool is part of the RCSB PDB research-focused RCSB.org web portal and available at rcsb.org/alignment. Programmatic access is available via alignment.rcsb.org. Frontend code has been published at github.com/rcsb/rcsb-pecos-app. Visualization is powered by the open-source Mol* viewer (github.com/molstar/molstar and github.com/molstar/rcsb-molstar) plus the Sequence Annotations in 3D Viewer (github.com/rcsb/rcsb-saguaro-3d). Sebastian Bittrich, Joan Segura, Jose M. Duarte, Stephen K. Burley, Yana Rose |
Bioinform. | 2 |
| 2022 | RCSB Protein Data Bank: improved annotation, search and visualization of membrane protein structures archived in the PDBabstractMOTIVATION: Membrane proteins are encoded by approximately one fifth of human genes but account for more than half of all US FDA approved drug targets. Thanks to new technological advances, the number of membrane proteins archived in the PDB is growing rapidly. However, automatic identification of membrane proteins or inference of membrane location is not a trivial task. RESULTS: We present recent improvements to the RCSB Protein Data Bank web portal (RCSB PDB, rcsb.org) that provide a wealth of new membrane protein annotations integrated from four external resources: OPM, PDBTM, MemProtMD and mpstruc. We have substantially enhanced the presentation of data on membrane proteins. The number of membrane proteins with annotations available on rcsb.org was increased by ∼80%. Users can search for these annotations, explore corresponding tree hierarchies, display membrane segments at the 1D amino acid sequence level, and visualize the predicted location of the membrane layer in 3D. AVAILABILITY AND IMPLEMENTATION: Annotations, search, tree data and visualization are available at our rcsb.org web portal. Membrane visualization is supported by the open-source Mol* viewer (molstar.org and github.com/molstar/molstar). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sebastian Bittrich, Yana Rose, Joan Segura, John D. Westbrook, Jose M. Duarte, Stephen K. Burley |
Bioinform. | 3 |
| 2022 | RCSB Protein Data Bank 1D3D module: displaying positional features on macromolecular assembliesabstractMOTIVATION: Mapping positional features from one-dimensional (1D) sequences onto three-dimensional (3D) structures of biological macromolecules is a powerful tool to show geometric patterns of biochemical annotations and provide a better understanding of the mechanisms underpinning protein and nucleic acid function at the atomic level. RESULTS: We present a new library designed to display fully customizable interactive views between 1D positional features of protein and/or nucleic acid sequences and their 3D structures as isolated chains or components of macromolecular assemblies. AVAILABILITY AND IMPLEMENTATION: https://github.com/rcsb/rcsb-saguaro-3d. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Joan Segura, Yana Rose, Sebastian Bittrich, Stephen K. Burley, Jose M. Duarte |
Bioinform. | 1 |
| 2021 | RCSB Protein Data Bank 1D tools and servicesabstractMOTIVATION: Interoperability between polymer sequences and structural data is essential for providing a complete picture of protein and gene features and helping to understand biomolecular function. RESULTS: Herein, we present two resources designed to improve interoperability between the RCSB Protein Data Bank, the NCBI and the UniProtKB data resources and visualize integrated data therefrom. The underlying tools provide a flexible means of mapping between the different coordinate spaces and an interactive tool allows convenient visualization of the 1-dimensional data over the web. AVAILABILITYAND IMPLEMENTATION: https://1d-coordinates.rcsb.org and https://rcsb.github.io/rcsb-saguaro. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Joan Segura, Yana Rose, John D. Westbrook, Stephen K. Burley, Jose M. Duarte |
Bioinform. | 1 |
| 2019 | BIPSPI: a method for the prediction of partner-specific protein-protein interfacesabstractMotivation: Protein-Protein Interactions (PPI) are essentials for most cellular processes and thus, unveiling how proteins interact is a crucial question that can be better understood by identifying which residues are responsible for the interaction. Computational approaches are orders of magnitude cheaper and faster than experimental ones, leading to proliferation of multiple methods aimed to predict which residues belong to the interface of an interaction. Results: We present BIPSPI, a new machine learning-based method for the prediction of partner-specific PPI sites. Contrary to most binding site prediction methods, the proposed approach takes into account a pair of interacting proteins rather than a single one in order to predict partner-specific binding sites. BIPSPI has been trained employing sequence-based and structural features from both protein partners of each complex compiled in the Protein-Protein Docking Benchmark version 5.0 and in an additional set independently compiled. Also, a version trained only on sequences has been developed. The performance of our approach has been assessed by a leave-one-out cross-validation over different benchmarks, outperforming state-of-the-art methods. Availability and implementation: BIPSPI web server is freely available at http://bipspi.cnb.csic.es. BIPSPI code is available at https://github.com/bioinsilico/BIPSPI. Docker image is available at https://hub.docker.com/r/bioinsilico/bipspi/. Supplementary information: Supplementary data are available at Bioinformatics online. Rubén Sánchez García, Carlos Oscar Sánchez Sorzano, José María Carazo, Joan Segura |
Bioinform. | 4 |
| 2019 | Validation of electron microscopy initial models via small angle X-ray scattering curvesabstractMOTIVATION: Cryo electron microscopy (EM) is currently one of the main tools to reveal the structural information of biological macromolecules. The re-construction of three-dimensional (3D) maps is typically carried out following an iterative process that requires an initial estimation of the 3D map to be refined in subsequent steps. Therefore, its determination is key in the quality of the final results, and there are cases in which it is still an open issue in single particle analysis (SPA). Small angle X-ray scattering (SAXS) is a well-known technique applied to structural biology. It is useful from small nanostructures up to macromolecular ensembles for its ability to obtain low resolution information of the biological sample measuring its X-ray scattering curve. These curves, together with further analysis, are able to yield information on the sizes, shapes and structures of the analyzed particles. RESULTS: In this paper, we show how the low resolution structural information revealed by SAXS is very useful for the validation of EM initial 3D models in SPA, helping the following refinement process to obtain more accurate 3D structures. For this purpose, we approximate the initial map by pseudo-atoms and predict the SAXS curve expected for this pseudo-atomic structure. The match between the predicted and experimental SAXS curves is considered as a good sign of the correctness of the EM initial map. AVAILABILITY AND IMPLEMENTATION: The algorithm is freely available as part of the Scipion 1.2 software at http://scipion.i2pc.es/. Amaya Jiménez, Slavica Jonic, Tomás Majtner, Joaquín Otón, Jose Luis Vilas, David Maluenda, Javier Mota, Erney Ramírez-Aportela, Marta Martínez, Yaiza Rancel, Joan Segura, Rubén Sánchez García, Roberto Melero, Laura del Cano, Pablo Conesa, Lars Skjærven, Roberto Marabini, José María Carazo, Carlos Oscar Sánchez Sorzano |
Bioinform. | 11 |
| 2019 | 3DBIONOTES v3.0: crossing molecular and structural biology data with genomic variationsabstractMOTIVATION: Many diseases are associated to single nucleotide polymorphisms that affect critical regions of proteins as binding sites or post translational modifications. Therefore, analysing genomic variants with structural and molecular biology data is a powerful framework in order to elucidate the potential causes of such diseases. RESULTS: A new version of our web framework 3DBIONOTES is presented. This version offers new tools to analyse and visualize protein annotations and genomic variants, including a contingency analysis of variants and amino acid features by means of a Fisher exact test, the integration of a gene annotation viewer to highlight protein features on gene sequences and a protein-protein interaction viewer to display protein annotations at network level. AVAILABILITY AND IMPLEMENTATION: The web server is available at https://3dbionotes.cnb.csic.es. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: Spanish National Institute for Bioinformatics (INB ELIXIR-ES) and Biocomputing Unit, National Centre of Biotechnology (CSIC)/Instruct Image Processing Centre, C/ Darwin nº 3, Campus of Cantoblanco, 28049 Madrid, Spain. Joan Segura, Rubén Sánchez García, Carlos Oscar Sánchez Sorzano, José María Carazo |
Bioinform. | 1 |
| 2017 | 3DBIONOTES v2.0: a web server for the automatic annotation of macromolecular structuresabstractMOTIVATION: Complementing structural information with biochemical and biomedical annotations is a powerful approach to explore the biological function of macromolecular complexes. However, currently the compilation of annotations and structural data is a feature only available for those structures that have been released as entries to the Protein Data Bank. RESULTS: To help researchers in assessing the consistency between structures and biological annotations for structural models not deposited in databases, we present 3DBIONOTES v2.0, a web application designed for the automatic annotation of biochemical and biomedical information onto macromolecular structural models determined by any experimental or computational technique. AVAILABILITY AND IMPLEMENTATION: The web server is available at http://3dbionotes-ws.cnb.csic.es. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Joan Segura, Rubén Sánchez García, Marta Martínez, Jesús Cuenca Alba, Daniel Tabas-Madrid, Carlos Oscar Sánchez Sorzano, José María Carazo |
Bioinform. | 1 |
| 2015 | Using neighborhood cohesiveness to infer interactions between protein domainsabstractMOTIVATION: In recent years, large-scale studies have been undertaken to describe, at least partially, protein-protein interaction maps, or interactomes, for a number of relevant organisms, including human. However, current interactomes provide a somehow limited picture of the molecular details involving protein interactions, mostly because essential experimental information, especially structural data, is lacking. Indeed, the gap between structural and interactomics information is enlarging and thus, for most interactions, key experimental information is missing. We elaborate on the observation that many interactions between proteins involve a pair of their constituent domains and, thus, the knowledge of how protein domains interact adds very significant information to any interactomic analysis. RESULTS: In this work, we describe a novel use of the neighborhood cohesiveness property to infer interactions between protein domains given a protein interaction network. We have shown that some clustering coefficients can be extended to measure a degree of cohesiveness between two sets of nodes within a network. Specifically, we used the meet/min coefficient to measure the proportion of interacting nodes between two sets of nodes and the fraction of common neighbors. This approach extends previous works where homolog coefficients were first defined around network nodes and later around edges. The proposed approach substantially increases both the number of predicted domain-domain interactions as well as its accuracy as compared with current methods. Joan Segura, Carlos Oscar Sánchez Sorzano, Jesús Cuenca Alba, Patrick Aloy, José María Carazo |
Bioinform. | 1 |
| 2014 | Frag'r'Us: knowledge-based sampling of protein backbone conformations for de novo structure-based protein designabstractMOTIVATION: The remodeling of short fragment(s) of the protein backbone to accommodate new function(s), fine-tune binding specificities or change/create novel protein interactions is a common task in structure-based computational design. Alternative backbone conformations can be generated de novo or by redeploying existing fragments extracted from protein structures i.e. knowledge-based. We present Frag'r'Us, a web server designed to sample alternative protein backbone conformations in loop regions. The method relies on a database of super secondary structural motifs called smotifs. Thus, sampling of conformations reflects structurally feasible fragments compiled from existing protein structures. Availability and implementation Frag'r'Us has been implemented as web application and is available at http://www.bioinsilico.org/FRAGRUS. Jaume Bonet, Joan Segura, Joan Planas-Iglesias, Baldomero Oliva, Narcis Fernandez-Fuentes |
Bioinform. | 2 |
| 2012 | Integrating human and murine anatomical gene expression data for improved comparisonsabstractMOTIVATION: Information concerning the gene expression pattern in four dimensions (species, genes, anatomy and developmental stage) is crucial for unraveling the roles of genes through time. There are a variety of anatomical gene expression databases, but extracting information from them can be hampered by their diversity and heterogeneity. RESULTS: aGEM 3.1 (anatomic Gene Expression Mapping) addresses the issues of diversity and heterogeneity of anatomical gene expression databases by integrating six mouse gene expression resources (EMAGE, GXD, GENSAT, Allen Brain Atlas data base, EUREXPRESS and BioGPS) and three human gene expression databases (HUDSEN, Human Protein Atlas and BioGPS). Furthermore, aGEM 3.1 provides new cross analysis tools to bridge these resources. AVAILABILITY AND IMPLEMENTATION: aGEM 3.1 can be queried using gene and anatomical structure. Output information is presented in a friendly format, allowing the user to display expression maps and correlation matrices for a gene or structure during development. An in-depth study of a specific developmental stage is also possible using heatmaps that relate gene expression with anatomical components. http://agem.cnb.csic.es CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Natalia Jiménez-Lozano, Joan Segura, José Ramón Macías, Juanjo Vega, José María Carazo |
Bioinform. | 2 |
| 2012 | A holistic in silico approach to predict functional sites in protein structuresabstractMOTIVATION: Proteins execute and coordinate cellular functions by interacting with other biomolecules. Among these interactions, protein-protein (including peptide-mediated), protein-DNA and protein-RNA interactions cover a wide range of critical processes and cellular functions. The functional characterization of proteins requires the description and mapping of functional biomolecular interactions and the identification and characterization of functional sites is an important step towards this end. RESULTS: We have developed a novel computational method, Multi-VORFFIP (MV), a tool to predicts protein-, peptide-, DNA- and RNA-binding sites in proteins. MV utilizes a wide range of structural, evolutionary, experimental and energy-based information that is integrated into a common probabilistic framework by means of a Random Forest ensemble classifier. While remaining competitive when compared with current methods, MV is a centralized resource for the prediction of functional sites and is interfaced by a powerful web application tailored to facilitate the use of the method and analysis of predictions to non-expert end-users. AVAILABILITY: http://www.bioinsilico.org/MVORFFIP Joan Segura, Pamela F. Jones, Narcis Fernandez-Fuentes |
Bioinform. | 1 |
| 2011 | Improving the prediction of protein binding sites by combining heterogeneous data and Voronoi DiagramsabstractBACKGROUND: Protein binding site prediction by computational means can yield valuable information that complements and guides experimental approaches to determine the structure of protein complexes. Predictions become even more relevant and timely given the current resolution of protein interaction maps, where there is a very large and still expanding gap between the available information on: (i) which proteins interact and (ii) how proteins interact. Proteins interact through exposed residues that present differential physicochemical properties, and these can be exploited to identify protein interfaces. RESULTS: Here we present VORFFIP, a novel method for protein binding site prediction. The method makes use of broad set of heterogeneous data and defined of residue environment, by means of Voronoi Diagrams that are integrated by a two-steps Random Forest ensemble classifier. Four sets of residue features (structural, energy terms, sequence conservation, and crystallographic B-factors) used in different combinations together with three definitions of residue environment (Voronoi Diagrams, sequence sliding window, and Euclidian distance) have been analyzed in order to maximize the performance of the method. CONCLUSIONS: The integration of different forms information such as structural features, energy term, evolutionary conservation and crystallographic B-factors, improves the performance of binding site prediction. Including the information of neighbouring residues also improves the prediction of protein interfaces. Among the different approaches that can be used to define the environment of exposed residues, Voronoi Diagrams provide the most accurate description. Finally, VORFFIP compares favourably to other methods reported in the recent literature. Joan Segura, Pamela F. Jones, Narcis Fernandez-Fuentes |
BMC Bioinform. | 1 |
| 2009 | aGEM: an integrative system for analyzing spatial-temporal gene-expression informationabstractAbstract Motivation: The work presented here describes the ‘anatomical Gene-Expression Mapping (aGEM)’ Platform, a development conceived to integrate phenotypic information with the spatial and temporal distributions of genes expressed in the mouse. The aGEM Platform has been built by extending the Distributed Annotation System (DAS) protocol, which was originally designed to share genome annotations over the WWW. DAS is a client-server system in which a single client integrates information from multiple distributed servers. Results: The aGEM Platform provides information to answer three main questions. (i) Which genes are expressed in a given mouse anatomical component? (ii) In which mouse anatomical structures are a given gene or set of genes expressed? And (iii) is there any correlation among these findings? Currently, this Platform includes several well-known mouse resources (EMAGE, GXD and GENSAT), hosting gene-expression data mostly obtained from in situ techniques together with a broad set of image-derived annotations. Availability: The Platform is optimized for Firefox 3.0 and it is accessed through a friendly and intuitive display: http://agem.cnb.csic.es Contact: [email protected] Supplementary information: Supplementary data are available at http://bioweb.cnb.csic.es/VisualOmics/aGEM/home.html and http://bioweb.cnb.csic.es/VisualOmics/index_VO.html and Bioinformatics online. Natalia Jiménez-Lozano, Joan Segura, José Ramón Macías, Juanjo Vega, José María Carazo |
Bioinform. | 2 |
| 2009 | Flexible structural protein alignment by a sequence of local transformationsabstractMOTIVATION: Throughout evolution, homologous proteins have common regions that stay semi-rigid relative to each other and other parts that vary in a more noticeable way. In order to compare the increasing number of structures in the PDB, flexible geometrical alignments are needed, that are reliable and easy to use. RESULTS: We present a protein structure alignment method whose main feature is the ability to consider different rigid transformations at different sites, allowing for deformations beyond a global rigid transformation. The performance of the method is comparable with that of the best ones from 10 aligners tested, regarding both the quality of the alignments with respect to hand curated ones, and the classification ability. An analysis of some structure pairs from the literature that need to be matched in a flexible fashion are shown. The use of a series of local transformations can be exported to other classifiers, and a future golden protein similarity measure could benefit from it. AVAILABILITY: A public server for the program is available at http://dmi.uib.es/ProtDeform/. SUPPLEMENTARY INFORMATION: All data used, results and examples are available at http://dmi.uib.es/people/jairo/bio/ProtDeform. Jairo Rocha, Joan Segura, Richard C. Wilson 0001, Swagata Dasgupta |
Bioinform. | 2 |