EDBT 2026 Demo / reviewers in the wild / expert
Aristidis G. Vrahatis
dblp:133/0585
· DBLP profile ↗
9ranked-venue papers in the field
1as first author
5since 2021 · last 2023
0000-0003-1892-0000ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 9 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Graph-Based Approach to Integrate Large-Scale Drug and Protein Data for Alzheimer's Disease Drug RepurposingabstractAlzheimer’s Disease (AD) remains a formidable challenge in neurodegenerative research, necessitating innovative approaches to uncover novel therapeutic strategies. This study presents a graph-based approach to integrate large-scale drug and protein data, aiming to identify potential drug repurposing candidates for AD. Our methodology constructs a comprehensive graph incorporating protein-protein interactions and drug-protein relations, providing a multifaceted view of the intricate relationships within the biological and pharmacological landscape. By leveraging this graph, we conduct an in-depth analysis to explore various drug repurposing possibilities, focusing on the alignment of AD-related single-cell transcriptomic data. Our approach enables the identification of promising drug candidates by examining the connectivity and interaction patterns within the graph, revealing potential therapeutic targets and drug synergies that may be beneficial for AD treatment. The integration of diverse data types allows for a more holistic understanding of the underlying molecular mechanisms and drug interactions. Through this graph-based analysis, we uncover several promising drug repurposing candidates, providing a foundation for further experimental validation and clinical investigation. This study underscores the potential of leveraging large-scale data and graph-based methodologies in drug repurposing efforts for neurodegenerative diseases, contributing to the advancement of therapeutic research in AD. Our findings illuminate the importance of integrative data analysis in biomedical research, paving the way for the development of more effective and targeted therapeutic interventions for AD. Georgios N. Dimitrakopoulos, Konstantinos Lazaros, Marios G. Krokidis, Themis P. Exarchos, Aristidis G. Vrahatis, Panayiotis M. Vlamos |
IEEE Big Data | 5 |
| 2023 | Advanced Big Data Analysis for Deciphering the Role of Protein Misfolding and Interactions in the Pathogenesis of Alzheimer's DiseaseabstractPredicting the three-dimensional structure of proteins directly from their sequence of amino acids remains a challenge in biomedical research. Protein functionality depends not only on its sequence but also on the precise folding that occurs during the process of developing their tertiary structure. Misfolded proteins may lead to the generation of entities that are inherently toxic to the organism, such as the formation of amyloid fibrils in the context of Alzheimer’s disease. Herein, the structural conformation of specific missense mutations in proteins involved in Alzheimer’s disease was performed through computational analysis and prediction of the binding mode and multiplicity of them was further assessed. Our findings reveal direct sequence-to-structure motifs from single polypeptides and the received domains with the proper fold. Nevertheless, the positions with the particular deviations are most commonly accompanied by limited downward spikes in pLDDT value, suggesting lower prediction confidence and potential disorder. Marios G. Krokidis, Georgios N. Dimitrakopoulos, Themis P. Exarchos, Aristidis G. Vrahatis, Panayiotis M. Vlamos |
IEEE Big Data | 4 |
| 2022 | Feature Selection For High Dimensional Data Using Supervised Machine Learning TechniquesabstractIn recent years, feature selection has become an increasingly active field of data science and machine learning research. Most of the datasets that are being used nowadays for various machine learning tasks consist of thousands of features (columns), which make them extremely complex and difficult to work with. In this paper, we propose a feature selection methodological pipeline that can be used to reduce the complexity of high dimensional datasets through the elimination of redundant and/or non-informative features as well as to improve the performance of machine learning models which are trained on high dimensional datasets. The proposed method has been applied to high-dimensional biomedical data and compared against a classic filter-based feature selection algorithm. Specifically, the method was applied to gene expression profiles of a single-cell RNA-seq dataset from healthy and infected by covid-19 human samples. Konstantinos Lazaros, Sotiris K. Tasoulis, Aristidis G. Vrahatis, Vassilis P. Plagianakos |
IEEE Big Data | 3 |
| 2021 | RLAC: Random Line Approximation ClusteringabstractWe explore how Random Projections can be used as an Approximate method for Projection Pursuit Clustering in high dimensional data. Traditional data transformations such as PCA for dimensionality reduction have been shown to be beneficial in clustering. However, their objective is not always relevant to the cluster structure producing undesirable results. On the other hand, Projection Pursuit methods present promising results in finding different "interesting" directions while being easily modified, though they came with high computational costs. In an attempt to provide a lightweight and simplified approach for Projection Pursuit clustering, we designed and implemented the Random Line Approximation Clustering (RLAC), a hierarchical divisive clustering algorithm that incorporates attributes from the Random Projection method. Petros T. Barbas, Aristidis G. Vrahatis, Sotiris K. Tasoulis |
IEEE BigData | 2 |
| 2021 | Recent Dimensionality Reduction Techniques for Visualizing High-Dimensional Parkinson's Disease Omics DataabstractOne challenge facing Systems Biology is the conversion of vast amounts of data into systematic and organized knowledge through automated processes. Improvements in experimental technologies have created an enormous pool of heterogeneous omics data such as genomics, proteomics and metabolomics. Exporting insights of these datasets can lead to important discoveries for complex pathologies such as age-related neurodegenerative disorders and specifically Parkinson's disease (PD). However, such data are characterized by huge dimensionality which increases their complexity for various data analyses as well for data mining processes such as clustering and classification. In this perspective, we implemented state-of-the-art Dimensionality Reduction Techniques for Visualizing Omics High-Dimensional Parkinson’s Disease Data. Our study highlights the cutting-edge dimensionality reduction techniques for 2D data visualization and their contribution the deeper interpretation of PD data. Approaches in this direction can provide a deeper understanding of biological dynamics and enable integrative multilayered diagnostic assessment of complex disorders such as neurodegenerative diseases. Marios G. Krokidis, Georgios N. Dimitrakopoulos, Aristidis G. Vrahatis, Themis P. Exarchos, Panayiotis M. Vlamos |
IEEE BigData | 3 |
| 2020 | Approximate kNN Classification for Biomedical DataabstractWe are in the era where the Big Data analytics has changed the way of interpreting the various biomedical phenomena, and as the generated data increase, the need for new machine learning methods to handle this evolution grows. An indicative example is the single-cell RNA-seq (scRNA-seq), an emerging DNA sequencing technology with promising capabilities but significant computational challenges due to the large-scaled generated data. Regarding the classification process for scRNA-seq data, an appropriate method is the k Nearest Neighbor (kNN) classifier since it is usually utilized for large-scale prediction tasks due to its simplicity, minimal parameterization, and model-free nature. However, the ultra-high dimensionality that characterizes scRNA-seq impose a computational bottleneck, while prediction power can be affected by the "Curse of Dimensionality". In this work, we proposed the utilization of approximate nearest neighbor search algorithms for the task of kNN classification in scRNA-seq data focusing on a particular methodology tailored for high dimensional data. We argue that even relaxed approximate solutions will not affect the prediction performance significantly. The experimental results confirm the original assumption by offering the potential for broader applicability. Panagiotis Anagnostou, Petros T. Barbas, Aristidis G. Vrahatis, Sotiris K. Tasoulis |
IEEE BigData | 3 |
| 2019 | Single-cell regulatory network inference and clustering from high-dimensional sequencing dataabstractWe are in the big data era which has affected several domains including biomedicine and healthcare. This revolution driven by the explosion of biomedical data offers the potential for better understanding of biology and human diseases. An illustrative example is the emerging single-cell sequencing technologies, which isolate and measure each cell individually, taking a step beyond the traditional techniques where consider their measurements from a bulk of cell. Although big single-cell RNA sequencing (scRNA-seq) data promises valuable insights into the cellular level, their volume poses several challenges related to the ultra-high dimensionality. Furthermore, to further elucidate the potential of these data, more insight into gene regulatory networks (GRN) is required. Network-based approaches can tackle part of the inherent complexity of human diseases, however, the challenges related to the ultra-high dimensionality are increased. Towards this direction, we propose the NIRP, an algorithm that copes with the high dimensionality of scRNA-data using a workflow based on fast multiple random projections and a radius-based nearest neighbors search. NIRP infers a gene regulatory network (GRN) from big scRNA-seq data by transforming the original data space to a lower dimensions space and capturing the similarities among gene expressions. The network is further analyzed using a random walk approach in order to achieve dense subgraphs, active to the case under study. The performance of NIRP is evaluated in a real single-cell experimental study among three well-established GRN tools. Our results make NIRP a reliable tool, able to handle big single-cell data with ultra-high dimensionality and complexity. he main advantage of this method is that it is not affected by the volume, as much as it increases, since it transforms the data space to a specific low dimensional space. Aristidis G. Vrahatis, Georgios N. Dimitrakopoulos, Sotiris K. Tasoulis, Spiros V. Georgakopoulos, Vassilis P. Plagianakos |
IEEE BigData | 1 |
| 2018 | Biomedical Data Ensemble Classification using Random ProjectionsabstractBiomedicine is undergoing a revolution driven by the explosion of biomedical data, which are generated by emerged medical imaging, sensor technologies and high-throughput technologies. An indicative example is the single cell sequencing technology which concerns the genome sequencing examination of hundreds of separate cells in a single tumor. Consequently, open challenges arising from this emerged technology and generally from the evolution of biomedical technologies under the big data perspective. Also, given the fact that approaches based on high-performance computing require high computing resources and advanced developers, solutions that reduce the problem complexity remain very attractive. Following this direction, in this paper a classification scheme based on Multiple Random Projections and Voting is presented. Random Projections offer a platform not only for a low computational time analysis by significantly reducing the data dimensionality, but also for an accurate analysis which may well exceed classical classification approaches. The proposed method was applied on real biomedical high dimensional data and compared against well-known classification schemes as to Random Projection-based cutting-edge methods. Specifically, we applied it on expression profiles for single-cell RNA-seq data from non-diabetic and type 2 diabetic human samples. Experimental results showed that based on simplistic tools we can create a computationally fast, simple, yet effective approach for biomedical Big Data analysis and knowledge discovery. Sotiris K. Tasoulis, Aristidis G. Vrahatis, Spiros V. Georgakopoulos, Vassilis P. Plagianakos |
IEEE BigData | 2 |
| 2018 | Visualizing High-dimensional single-cell RNA-sequencing data through multiple Random ProjectionsabstractRecent sequencing technology breakthroughs have resulted in a dramatic increase in the amount of available sequencing data, enabling major scientific advances in biology and medicine. Nowadays, sequencing transcriptome data of single cells (scRNA-seq) are growing rapidly, posing new challenges in their analysis, mostly due to their high dimensionality. In this paper, we study the problem of visualizing such high-dimensional scRNA-seq data. A new visualization scheme is presented based on a customized distance matrix retrieved by applying independently Nearest Neighbors search through multiple Random Projections. The proposed method is compared against well-known dimensionality reduction and visualization techniques showing its capabilities and performance. Sotiris K. Tasoulis, Aristidis G. Vrahatis, Spiros V. Georgakopoulos, Vassilis P. Plagianakos |
IEEE BigData | 2 |