EDBT 2026 Demo / reviewers in the wild / expert
Aleksandra Gruca
dblp:89/7054
· DBLP profile ↗
12ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-2337-1894ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequencesabstractMOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653. Patryk Jarnot, Joanna Ziemska-Legiecka, Marcin Grynberg, Vasilis J. Promponas, Aleksandra Gruca |
Bioinform. | 5 |
| 2025 | Bottlenecks in advancing and applying multiomic data integration - common data resources as rate-limiting drivers - the high-impact use case of atherosclerotic cardiovascular diseaseabstractDespite striking successes in identifying novel biomarkers for improved patient stratification and predicting disease progression, numerous challenges remain in the effective integration and exploitation of multiomic data in biomedical applications beyond cancer, for which most bioinformatics strategies are developed and validated. That focus on cancer severely limits the effective development and advancement of algorithms in machine learning and artificial intelligence that do not suffer degraded out-of-domain performance. Generalizability and interpretability of models, however, are also required for robust insights that may translate into clinical practice. Work across different independent datasets is critical for establishing models robust towards unwanted variation in assays, protocols, and cohort populations. Disease-specific context like ethnicity, socioeconomic background, sex, lifestyle, disease phase, and tissue type also strongly affect molecular profiles. We here discuss atherosclerotic cardiovascular disease (ASCVD) as a high-impact non-cancer use case for the challenges remaining in the development and application of the latest bioinformatics approaches to multiomics data integration. ASCVD remains the leading cause of death globally. Disease aetiology, progression, and therapy outcome depend on a complex interplay of genetic, environmental, and lifestyle factors. Integrating these diverse data types effectively remains a challenge but holds transformative potential for personalized medicine. Discovery and access to data of sufficient diversity and extent form key bottlenecks. We here compile a first comprehensive overview of key data sets in ASCVD to complement the established cancer-focused resources as a foundation for future effective development and application of state-of-the-art bioinformatics tools for multiomic data integration. Stephanie Bezzina Wettinger, Kanita Karaduzovic-Hadziabdic, Ritienne Attard, Rosienne Farrugia, Brooke N. Wolford, Marco Chierici, Giuseppe Jurman, Panagiotis Alexiou, José L. Peñalvo, Rafael S. Costa, José Basilio, Frantisek Sabovcik, Rui Vitorino, Johannes A. Schmid, Rajesh Shigdel, Baiba Vilne, Artemis G. Hatzigeorgiou, Miron Sopic, Yvan Devaux, Paolo Magni, Maria Tellez-Plaza, David P. Kreil, Aleksandra Gruca |
Briefings Bioinform. | 23 |
| 2022 | Insights from analyses of low complexity regions with canonical methods for protein sequence comparisonabstractLow complexity regions are fragments of protein sequences composed of only a few types of amino acids. These regions frequently occur in proteins and can play an important role in their functions. However, scientists are mainly focused on regions characterized by high diversity of amino acid composition. Similarity between regions of protein sequences frequently reflect functional similarity between them. In this article, we discuss strengths and weaknesses of the similarity analysis of low complexity regions using BLAST, HHblits and CD-HIT. These methods are considered to be the gold standard in protein similarity analysis and were designed for comparison of high complexity regions. However, we lack specialized methods that could be used to compare the similarity of low complexity regions. Therefore, we investigated the existing methods in order to understand how they can be applied to compare such regions. Our results are supported by exploratory study, discussion of amino acid composition and biological roles of selected examples. We show that existing methods need improvements to efficiently search for similar low complexity regions. We suggest features that have to be re-designed specifically for comparing low complexity regions: scoring matrix, multiple sequence alignment, e-value, local alignment and clustering based on a set of representative sequences. Results of this analysis can either be used to improve existing methods or to create new methods for the similarity analysis of low complexity regions. Patryk Jarnot, Joanna Ziemska-Legiecka, Marcin Grynberg, Aleksandra Gruca |
Briefings Bioinform. | 4 |
| 2022 | MAINE: a web tool for multi-omics feature selection and rule-based data explorationabstractSUMMARY: Patient multi-omics datasets are often characterized by a high dimensionality; however, usually only a small fraction of the features is informative, that is change in their value is directly related to the disease outcome or patient survival. In medical sciences, in addition to a robust feature selection procedure, the ability to discover human-readable patterns in the analyzed data is also desirable. To address this need, we created MAINE-Multi-omics Analysis and Exploration. The unique functionality of MAINE is the ability to discover multidimensional dependencies between the selected multi-omics features and event outcome prediction as well as patient survival probability. Learned patterns are visualized in the form of interpretable decision/survival trees and rules. AVAILABILITY AND IMPLEMENTATION: MAINE is freely available at maine.ibemag.pl as an online web application. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Aleksandra Gruca, Joanna Badura, Iwona Kostorz, Tomasz Steclik, Lukasz Wróbel, Marek Sikora |
Bioinform. | 1 |
| 2021 | High-resolution multi-channel weather forecasting - First insights on transfer learning from the Weather4cast Competitions 2021abstractWeather forecasting is both a high impact application as well as a complex Big Data modelling challenge. Recent advances in machine learning have already demonstrated the power of non-physical modelling approaches for the prediction of rainfall. The Weather4cast competitions now provide a unique multi-channel benchmark for the prediction of up to 8 hours of weather with high temporal and spatial resolutions (15 min, 4 km) for a diverse set of large regions across Earth. This diversity, for the first time, also permits a meaningful spatial transfer learning challenge in weather forecasting.Weather4cast introduces multi-channel weather ‘movies’ that encode temperature, rainfall, cloud properties, and turbulence as derived from the meteorological satellites by the EUMETSAT NWC SAF. Inspired by the Traffic4cast competitions at the NeurIPS conferences in 2019 and 2020, weather forecasting is thus presented as a video frame prediction task. As then, the U-Net based models developed for photographic image analysis intriguingly did well on these artificial videos. In contrast, however, the winning submission did not employ a U-Net but a recurrent convolutional network with residual units.Weather4cast introduces the first spatial transfer learning challenge in weather forecasting: only one-hour short snippets from spatial regions never seen before were provided as input to models. We can thus now present first insights from submissions to this spatial transfer learning challenge. Notably, models with better core prediction performance also generalized better. Moreover, the two top-ranked models – one RCN based, one U-Net based – were further ahead of the remaining top-ranked submissions for spatial transfer learning (+6%) than in the core prediction challenge (+1%).While submissions tested varying strategies for input data selection and training, there remains a wide range of additional complementary approaches to be explored in future analyses. The competition and its leaderboards remain available and open for new submissions on the weather4cast.ai website. Pedro Herruzo, Aleksandra Gruca, Llorenç Lliso, Xavier Calbet, Pilar Rípodas, Sepp Hochreiter, Michael Kopp 0001, David P. Kreil |
IEEE BigData | 2 |
| 2021 | CDCEO'21 - First Workshop on Complex Data Challenges in Earth ObservationabstractHigh-resolution remote sensing technology for Earth Observation (EO) has radically changed how we monitor the state of our planet around the clock. An effective interpretation of the resulting complex large-scale time series adopts the best machine learning techniques from signal processing, computer vision, pattern recognition, and artificial intelligence. The First Workshop on Complex Data Challenges in Earth Observation was open to both method development and advanced applications in a wide range of related topics, including image and signal processing, gap-filling, data fusion, feature extraction, prediction of spatio-temporal features, and the detection of rules underlying the observed state transitions and causal relationships. The full agenda, featuring keynotes and a selection of high quality contributed talks is available online at www.iarai.ac.at/cdceo21 Aleksandra Gruca, Pedro Herruzo, Pilar Rípodas, Andrzej Kucik, Christian Briese, Michael Kopp 0001, Sepp Hochreiter, Pedram Ghamisi, David P. Kreil |
CIKM | 1 |
| 2021 | Common low complexity regions for SARS-CoV-2 and human proteomes as potential multidirectional risk factor in vaccine developmentabstractBACKGROUND: The rapid spread of the COVID-19 demands immediate response from the scientific communities. Appropriate countermeasures mean thoughtful and educated choice of viral targets (epitopes). There are several articles that discuss such choices in the SARS-CoV-2 proteome, other focus on phylogenetic traits and history of the Coronaviridae genome/proteome. However none consider viral protein low complexity regions (LCRs). Recently we created the first methods that are able to compare such fragments. RESULTS: We show that five low complexity regions (LCRs) in three proteins (nsp3, S and N) encoded by the SARS-CoV-2 genome are highly similar to regions from human proteome. As many as 21 predicted T-cell epitopes and 27 predicted B-cell epitopes overlap with the five SARS-CoV-2 LCRs similar to human proteins. Interestingly, replication proteins encoded in the central part of viral RNA are devoid of LCRs. CONCLUSIONS: Similarity of SARS-CoV-2 LCRs to human proteins may have implications on the ability of the virus to counteract immune defenses. The vaccine targeted LCRs may potentially be ineffective or alternatively lead to autoimmune diseases development. These findings are crucial to the process of selection of new epitopes for drugs or vaccines which should omit such regions. Aleksandra Gruca, Joanna Ziemska-Legiecka, Patryk Jarnot, Elzbieta Sarnowska, Tomasz J. Sarnowski, Marcin Grynberg |
BMC Bioinform. | 1 |
| 2020 | Disentangling the complexity of low complexity proteinsabstractThere are multiple definitions for low complexity regions (LCRs) in protein sequences, with all of them broadly considering LCRs as regions with fewer amino acid types compared to an average composition. Following this view, LCRs can also be defined as regions showing composition bias. In this critical review, we focus on the definition of sequence complexity of LCRs and their connection with structure. We present statistics and methodological approaches that measure low complexity (LC) and related sequence properties. Composition bias is often associated with LC and disorder, but repeats, while compositionally biased, might also induce ordered structures. We illustrate this dichotomy, and more generally the overlaps between different properties related to LCRs, using examples. We argue that statistical measures alone cannot capture all structural aspects of LCRs and recommend the combined usage of a variety of predictive tools and measurements. While the methodologies available to study LCRs are already very advanced, we foresee that a more comprehensive annotation of sequences in the databases will enable the improvement of predictions and a better understanding of the evolution and the connection between structure and function of LCRs. This will require the use of standards for the generation and exchange of data describing all aspects of LCRs. SHORT ABSTRACT: There are multiple definitions for low complexity regions (LCRs) in protein sequences. In this critical review, we focus on the definition of sequence complexity of LCRs and their connection with structure. We present statistics and methodological approaches that measure low complexity (LC) and related sequence properties. Composition bias is often associated with LC and disorder, but repeats, while compositionally biased, might also induce ordered structures. We illustrate this dichotomy, plus overlaps between different properties related to LCRs, using examples. Pablo Mier, Lisanna Paladin, Stella Tamana, Sophia Petrosian, Borbála Hajdu-Soltész, Annika Urbanek, Aleksandra Gruca, Dariusz Plewczynski, Marcin Grynberg, Pau Bernadó, Zoltán Gáspári, Christos A. Ouzounis, Vasilis J. Promponas, Andrey V. Kajava, John M. Hancock, Silvio C. E. Tosatto, Zsuzsanna Dosztányi, Miguel A. Andrade-Navarro |
Briefings Bioinform. | 7 |
| 2014 | Soft Approach to Identification of Cohesive Clusters in Two Gene RepresentationsabstractThe approach to identify clusters of genes represented both by expression values and Gene Ontology annotations, where cluster membership should not be in conflict with any of the representations is presented in the paper. The method enables to identify the genes that are differently clustered in different representations, what can lead to further analysis and interesting conclusions. The approach is based on the fuzzy clustering algorithms and the notion of proximity as the aggregation operation at the higher level than similarity matrices is performed. The approach is verified on two datasets: a small synthetic and real-world gene dataset. Michal Kozielski, Aleksandra Gruca |
KES | 2 |
| 2011 | Efficient System for Clustering of Dynamic Document Database
Pawel Foszner, Aleksandra Gruca, Andrzej Polanski |
CDVE | 2 |
| 2011 | Efficient Algorithm for Microarray Probes Re-annotation
Pawel Foszner, Aleksandra Gruca, Andrzej Polanski, Michal Marczyk, Roman Jaksik, Joanna Polanska |
ICCCI (2) | 2 |
| 2011 | Induction and selection of the most interesting Gene Ontology based multiattribute rules for descriptions of gene groups
Marek Sikora, Aleksandra Gruca |
Pattern Recognit. Lett. | 2 |