EDBT 2026 Demo / reviewers in the wild / expert
Dariusz Plewczynski
dblp:56/812
· DBLP profile ↗
26ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-3840-7610ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The challenge of chromatin model comparison and validation: A project from the first international 4D Nucleome HackathonabstractThe computational modeling of chromatin structure is highly complex due to the hierarchical organization of chromatin, which reflects its diverse biophysical principles, as well as inherent dynamism, which underlies its complexity. Chromatin structure modeling can be based on diverse approaches and assumptions, making it essential to determine how different methods influence the modeling outcomes. We conducted a project at the NIH-funded 4D Nucleome Hackathon on March 18-21, 2024, at The University of Washington in Seattle, USA. The hackathon provided an amazing opportunity to gather an international, multi-institutional and unbiased group of experts to discuss, understand and undertake the challenges of chromatin model comparison and validation. Here we give an overview of the current state of the 3D chromatin field and discuss our efforts to run and validate the models. We used distance matrices to represent chromatin models and we calculated Spearman correlation coefficients to estimate differences between models, as well as between models and experimental data. In addition, we discuss challenges in chromatin structure modeling that include: 1) different aspects of chromatin biophysics and scales complicate model comparisons, 2) large diversity of experimental data (e.g., population-based, single-cell, protein-specific) that differ in mathematical properties, heatmap smoothness, noise and resolutions complicates model validation, 3) expertise in biology, bioinformatics, and physics is necessary to conduct comprehensive research on chromatin structure, 4) bioinformatic software, which is often developed in academic settings, is characterized by insufficient support and documentation. We also emphasize the importance of establishing guidelines for software development and standardization. Jedrzej Kubica, Sevastianos Korsak, Krzysztof H. Banecki, Dvir Schirman, Anurupa Devi Yadavalli, Ariana Brenner Clerkin, David Kouril, Michal Kadlof, Ben Busby, Dariusz Plewczynski |
PLoS Comput. Biol. | 10 |
| 2023 | BERTrand - peptide:TCR binding prediction using Bidirectional Encoder Representations from Transformers augmented with random TCR pairingabstractMOTIVATION: The advent of T-cell receptor (TCR) sequencing experiments allowed for a significant increase in the amount of peptide:TCR binding data available and a number of machine-learning models appeared in recent years. High-quality prediction models for a fixed epitope sequence are feasible, provided enough known binding TCR sequences are available. However, their performance drops significantly for previously unseen peptides. RESULTS: We prepare the dataset of known peptide:TCR binders and augment it with negative decoys created using healthy donors' T-cell repertoires. We employ deep learning methods commonly applied in Natural Language Processing to train part a peptide:TCR binding model with a degree of cross-peptide generalization (0.69 AUROC). We demonstrate that BERTrand outperforms the published methods when evaluated on peptide sequences not used during model training. AVAILABILITY AND IMPLEMENTATION: The datasets and the code for model training are available at https://github.com/SFGLab/bertrand. Alexander Myronov, Giovanni Mazzocco, Paulina Król, Dariusz Plewczynski |
Bioinform. | 4 |
| 2023 | cudaMMC: GPU-enhanced multiscale Monte Carlo chromatin 3D modellingabstractMOTIVATION: Investigating the 3D structure of chromatin provides new insights into transcriptional regulation. With the evolution of 3C next-generation sequencing methods like ChiA-PET and Hi-C, the surge in data volume has highlighted the need for more efficient chromatin spatial modelling algorithms. This study introduces the cudaMMC method, based on the Simulated Annealing Monte Carlo approach and enhanced by GPU-accelerated computing, to efficiently generate ensembles of chromatin 3D structures. RESULTS: The cudaMMC calculations demonstrate significantly faster performance with better stability compared to our previous method on the same workstation. cudaMMC also substantially reduces the computation time required for generating ensembles of large chromatin models, making it an invaluable tool for studying chromatin spatial conformation. AVAILABILITY AND IMPLEMENTATION: Open-source software and manual and sample data are freely available on https://github.com/SFGLab/cudaMMC. Michal Wlasnowolski, Pawel Z. Grabowski, Damian Roszczyk, Krzysztof Kaczmarski, Dariusz Plewczynski |
Bioinform. | 5 |
| 2022 | MCIBox: a toolkit for single-molecule multi-way chromatin interaction visualization and micro-domains identificationabstractThe emerging ligation-free three-dimensional (3D) genome mapping technologies can identify multiplex chromatin interactions with single-molecule precision. These technologies not only offer new insight into high-dimensional chromatin organization and gene regulation, but also introduce new challenges in data visualization and analysis. To overcome these challenges, we developed MCIBox, a toolkit for multi-way chromatin interaction (MCI) analysis, including a visualization tool and a platform for identifying micro-domains with clustered single-molecule chromatin complexes. MCIBox is based on various clustering algorithms integrated with dimensionality reduction methods that can display multiplex chromatin interactions at single-molecule level, allowing users to explore chromatin extrusion patterns and super-enhancers regulation modes in transcription, and to identify single-molecule chromatin complexes that are clustered into micro-domains. Furthermore, MCIBox incorporates a two-dimensional kernel density estimation algorithm to identify micro-domains boundaries automatically. These micro-domains were stratified with distinctive signatures of transcription activity and contained different cell-cycle-associated genes. Taken together, MCIBox represents an invaluable tool for the study of multiple chromatin interactions and inaugurates a previously unappreciated view of 3D genome structure. Simon Zhongyuan Tian, Guoliang Li 0002, Duo Ning, Kai Jing, Yewen Xu, Yang Yang 0145, Melissa Jane Fullwood, Pengfei Yin, Guangyu Huang, Dariusz Plewczynski, Jixian Zhai, Ziwei Dai, Meizhen Zheng |
Briefings Bioinform. | 10 |
| 2022 | ConsensuSV - from the whole-genome sequencing data to the complete variant listabstractSUMMARY: The detection of the structural variants (SVs) using Illumina sequencing of human DNA is not an easy task. Multiple approaches have been proposed; however, all the methods have their limitations. In this article, we present ConsensuSV pipeline that aids the research in complex variant detection. By using consensus meta-approach, eight independent SV callers are being used to identify a uniform set of high-quality SVs. The pipeline works using raw sequencing data and performs all the necessary steps automatically, significantly reducing the researchers' time required for processing the data. The output files contain SVs, single nucleotide polymorphisms and Indels. The pipeline uses luigi framework, allowing the software to be run efficiently and parallelly using the high-performance computing infrastructure. We strongly believe that the software is useful to the scientific community interested in the germline variant detection. AVAILABILITY AND IMPLEMENTATION: https://github.com/SFGLab/ConsensuSV-pipeline. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mateusz Chilinski, Dariusz Plewczynski |
Bioinform. | 2 |
| 2022 | JUPPI: A Multi-Level Feature Based Method for PPI Prediction and a Refined Strategy for Performance AssessmentabstractOver the years, several methods have been proposed for the computational PPI prediction with different performance evaluation strategies. While attempting to benchmark performance scores, most of these methods often suffer with ill-treated cross-validation strategies, adhoc selection of positive/negative samples etc. To address these issues, in our proposed multi-level feature based PPI prediction approach (JUPPI), using sequence, domain and GO information as features, a refined evaluation strategy has been introduced. During the evaluation process, we first extract high quality negative data using three-stage filtering, and then introduce a pair-input based cross validation strategy with three difficulty levels for test-set predictions. Our proposed evaluation strategy reduces the component-level overlapping issue in test sets. Performance of JUPPI is compared with those of the state-of-the-art approaches in this domain and tested on six independent PPI datasets. In almost all the datasets, JUPPI outperforms the state-of-the-art not only at human proteome level for PPI prediction, but also for prediction of interactors for intrinsic disordered human proteins. https://figshare.com/projects/JUPPI_A_Multi-level_Feature_Based_Method_for_PPI_Prediction_and_a_Refined_Strategy_for_Performance_Assessment/81656 JUPPI tool and the developed datasets (JUPPId) are available in public domain for academic use along with supplementary materials, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/10.1109/TCBB.2020.3004970. Anup Kumar Halder, Soumyendu Sekhar Bandyopadhyay, Piyali Chatterjee, Mita Nasipuri, Dariusz Plewczynski, Subhadip Basu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2021 | PartSeg: a tool for quantitative feature extraction from 3D microscopy images for dummiesabstractBACKGROUND: Bioimaging techniques offer a robust tool for studying molecular pathways and morphological phenotypes of cell populations subjected to various conditions. As modern high-resolution 3D microscopy provides access to an ever-increasing amount of high-quality images, there arises a need for their analysis in an automated, unbiased, and simple way. Segmentation of structures within the cell nucleus, which is the focus of this paper, presents a new layer of complexity in the form of dense packing and significant signal overlap. At the same time, the available segmentation tools provide a steep learning curve for new users with a limited technical background. This is especially apparent in the bulk processing of image sets, which requires the use of some form of programming notation. RESULTS: In this paper, we present PartSeg, a tool for segmentation and reconstruction of 3D microscopy images, optimised for the study of the cell nucleus. PartSeg integrates refined versions of several state-of-the-art algorithms, including a new multi-scale approach for segmentation and quantitative analysis of 3D microscopy images. The features and user-friendly interface of PartSeg were carefully planned with biologists in mind, based on analysis of multiple use cases and difficulties encountered with other tools, to offer an ergonomic interface with a minimal entry barrier. Bulk processing in an ad-hoc manner is possible without the need for programmer support. As the size of datasets of interest grows, such bulk processing solutions become essential for proper statistical analysis of results. Advanced users can use PartSeg components as a library within Python data processing and visualisation pipelines, for example within Jupyter notebooks. The tool is extensible so that new functionality and algorithms can be added by the use of plugins. For biologists, the utility of PartSeg is presented in several scenarios, showing the quantitative analysis of nuclear structures. CONCLUSIONS: In this paper, we have presented PartSeg which is a tool for precise and verifiable segmentation and reconstruction of 3D microscopy images. PartSeg is optimised for cell nucleus analysis and offers multi-scale segmentation algorithms best-suited for this task. PartSeg can also be used for the bulk processing of multiple images and its components can be reused in other systems or computational experiments. Grzegorz Bokota, Jacek Sroka, Subhadip Basu, Nirmal Das, Pawel Trzaskoma, Yana Yushkevich, Agnieszka Grabowska, Adriana Magalska, Dariusz Plewczynski |
BMC Bioinform. | 9 |
| 2020 | Disentangling the complexity of low complexity proteinsabstractThere are multiple definitions for low complexity regions (LCRs) in protein sequences, with all of them broadly considering LCRs as regions with fewer amino acid types compared to an average composition. Following this view, LCRs can also be defined as regions showing composition bias. In this critical review, we focus on the definition of sequence complexity of LCRs and their connection with structure. We present statistics and methodological approaches that measure low complexity (LC) and related sequence properties. Composition bias is often associated with LC and disorder, but repeats, while compositionally biased, might also induce ordered structures. We illustrate this dichotomy, and more generally the overlaps between different properties related to LCRs, using examples. We argue that statistical measures alone cannot capture all structural aspects of LCRs and recommend the combined usage of a variety of predictive tools and measurements. While the methodologies available to study LCRs are already very advanced, we foresee that a more comprehensive annotation of sequences in the databases will enable the improvement of predictions and a better understanding of the evolution and the connection between structure and function of LCRs. This will require the use of standards for the generation and exchange of data describing all aspects of LCRs. SHORT ABSTRACT: There are multiple definitions for low complexity regions (LCRs) in protein sequences. In this critical review, we focus on the definition of sequence complexity of LCRs and their connection with structure. We present statistics and methodological approaches that measure low complexity (LC) and related sequence properties. Composition bias is often associated with LC and disorder, but repeats, while compositionally biased, might also induce ordered structures. We illustrate this dichotomy, plus overlaps between different properties related to LCRs, using examples. Pablo Mier, Lisanna Paladin, Stella Tamana, Sophia Petrosian, Borbála Hajdu-Soltész, Annika Urbanek, Aleksandra Gruca, Dariusz Plewczynski, Marcin Grynberg, Pau Bernadó, Zoltán Gáspári, Christos A. Ouzounis, Vasilis J. Promponas, Andrey V. Kajava, John M. Hancock, Silvio C. E. Tosatto, Zsuzsanna Dosztányi, Miguel A. Andrade-Navarro |
Briefings Bioinform. | 8 |
| 2019 | Identification of Epigenetic Biomarkers with the use of Gene Expression and DNA Methylation for Breast Cancer SubtypesabstractBreast cancer is one of the most deadly cancers. It has four subtypes: Luminal A (LA), Luminal B (LB), HER2-enriched (HER2-E) and Basal-like (BL). For the cause of breast cancer subtypes, there are different genetic and epigenetic factors involved in its progression and susceptibility. Thus, the identification of genetic and/or epigenetic biomarkers can be helpful to understand the biological mechanisms better and to improve the diagnostic processes of this disease and its subtypes. Hence, this fact motivated us to investigate the epigenetic factor, such as DNA Methylation, with the integration of gene expression in order to find epigenetic biomarkers for breast cancer subtypes. In this regard, we have identified set of up and down regulated genes for each subtype using differential analysis. Thereafter, regression based feature ranking problem is formed in order to find the DNA Methylation site that is mostly responsible for the change in expression of a gene, which is considered as an epigenetic biomarker. A bagging integrated ensemble of decision trees is used for the same. The results of top ten up and down regulated genes and their corresponding most significant DNA Methylation sites are reported for breast cancer subtypes. Moreover, these genes are validated visually by means of survival and expression plots, showing TF-Gene-DNA Methylation interactions, Protein-Protein interaction network, KEGG pathway and GO enrichment analysis. The results show that top differentially expressed up and down regulated genes viz. MMP11, NUF2, EXO1, HJURP, HOXA4, SYNM, CAV1 and COL4A3BP in breast cancer subtypes may change their expression because of DNA Methylation sites viz. cg22418565, cg26029744, cg24741598, cg04550103, cg25952581, cg02109162, cg18498156 and cg04985097 respectively. The code, datasets and supplementary material are present online11http://www.nitttrkol.ac.in/indrajit/projects/epigenetic-mrna-breastcancer-subtypes/. Indrajit Saha, Somnath Rakshit, Michal Wlasnowolski, Dariusz Plewczynski |
TENCON | 4 |
| 2019 | 3gClust: Human Protein Cluster AnalysisabstractWe present a human protein cluster analysis by combining: 1) n-gram based amino acid frequency features, 2) optimal feature selection, 3) hierarchical clustering, and 4) advanced partitioning techniques. Our method qualitatively and quantitatively groups proteins with increasing sequence similarity into similar clusters by calculating the frequency model of amino acids using n-grams. We experiment with n = 1, i.e., unigrams, n = 2, i.e., bigrams, and finally n = 3, i.e., trigrams for optimal selection of features to design the 3gClust algorithm. The benchmarking results on 20,105 manually curated human proteins show that 3gClust ensures better cluster compactness in the case of proteins with similar functional groups, biological processes, structural alignment, and shared domains (e.g., aquaporins, keratins). Quantitative analysis of non singleton clusters shows significant improvement in their compactness in comparison to other state-of-the art methodologies. 3gClust is available at https://sites.google.com/site/bioinfoju/projects/3gclust for academic use along with supplementary materials, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/10.1109/TCBB.2018.2840996, and datasets. Anup Kumar Halder, Piyali Chatterjee, Mita Nasipuri, Dariusz Plewczynski, Subhadip Basu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2018 | Deep Learning for Integrated Analysis of Breast Cancer Subtype Specific Multi-omics DataabstractBreast cancer is a deadly disease which commonly occurs all over the world and has been found to be the largest cause of cancer in females. Its detection is still a major challenge, both from a computational and biological point of views. Next Generation Sequencing (NGS) techniques have accelerated the mapping of human genomes rapidly. Involvement of advanced NGS techniques reveals that multiple genetic molecules are responsible for the cause of breast cancer and its subtypes. However, the high volume of data that is produced by the NGS techniques is difficult to study because of their high dimensionality and complexity. Thus, the integrated study of multi-omics data is one of the major challenges in medical science. This fact motivated us to study the NGS based high throughput expression data of miRNAs and mRNAs as well as Beta values of DNA Methylation of the corresponding mRNAs. In this regard, first, these datasets, together consisting of 33564 features of 305 patients in five classes viz. Luminal A, Luminal B, HER2-enriched, Basal-like and Control, are analysed in an integrated fashion using deep learning technique to classify the breast cancer subtypes properly. Second, the results of the deep learning technique are further analysed in order to identify the deeply connected features, i.e. either miRNA or mRNA or DNA Methylation, which are pivotal in the classification of breast cancer subtypes as well as play a crucial role in its formation. For this purpose, a deep learning technique, called stacked autoencoder is used to encode/transform the features into a low dimensional space, which is then fed to the five well known classifiers for classification. Moreover, the same encoded data is used to select the potential features after performing multiplication with the original data and Bonferroni correction on the p-values produced by the one-sample t-test. The results have been validated quantitatively and through biological significance analysis where oncogene TP53 and tumor suppression gene BRCA1 have been found. These genes are known to play a crucial role in breast cancer. The datasets, code and supplementary materials of this work are provided online at http://www.nitttrkol.ac.in/indrajit/projects/integrated-analysis-breastcancer-subtypes/. Somnath Rakshit, Indrajit Saha, Subha Shankar Chakraborty, Dariusz Plewczynski |
TENCON | 4 |
| 2017 | Multi-levels 3D Chromatin Interactions Prediction Using Epigenomic Profiles
Ziad Al Bkhetan, Dariusz Plewczynski |
ISMIS | 2 |
| 2016 | 2dSpAn: semiautomated 2-d segmentation, classification and analysis of hippocampal dendritic spine plasticityabstractMOTIVATION: Accurate and effective dendritic spine segmentation from the dendrites remains as a challenge for current neuroimaging research community. In this article, we present a new method (2dSpAn) for 2-d segmentation, classification and analysis of structural/plastic changes of hippocampal dendritic spines. A user interactive segmentation method with convolution kernels is designed to segment the spines from the dendrites. Formal morphological definitions are presented to describe key attributes related to the shape of segmented spines. Spines are automatically classified into one of four classes: Stubby, Filopodia, Mushroom and Spine-head Protrusions. RESULTS: The developed method is validated using confocal light microscopy images of dendritic spines from dissociated hippocampal cultures for: (i) quantitative analysis of spine morphological changes, (ii) reproducibility analysis for assessment of user-independence of the developed software and (iii) accuracy analysis with respect to the manually labeled ground truth images, and also with respect to the available state of the art. The developed method is monitored and used to precisely describe the morphology of individual spines in real-time experiments, i.e. consequent images of the same dendritic fragment. AVAILABILITY AND IMPLEMENTATION: The software and the source code are available at https://sites.google.com/site/2dspan/ under open-source license for non-commercial use. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Subhadip Basu, Dariusz Plewczynski, Satadal Saha, Matylda Roszkowska, Marta Magnowska, Ewa Baczynska, Jakub Wlodarczyk |
Bioinform. | 2 |
| 2016 | Highlights from the 11th ISCB Student Council Symposium 2015: Dublin, Ireland. 10 July 2015abstractTable of contents A1 Highlights from the eleventh ISCB Student Council Symposium 2015 Katie Wilkins, Mehedi Hassan, Margherita Francescatto, Jakob Jespersen, R. Gonzalo Parra, Bart Cuypers, Dan DeBlasio, Alexander Junge, Anupama Jigisha, Farzana Rahman O1 Prioritizing a drug’s targets using both gene expression and structural similarity Griet Laenen, Sander Willems, Lieven Thorrez, Yves Moreau O2 Organism specific protein-RNA recognition: A computational analysis of protein-RNA complex structures from different organisms Nagarajan Raju, Sonia Pankaj Chothani, C. Ramakrishnan, Masakazu Sekijima; M. Michael Gromiha O3 Detection of Heterogeneity in Single Particle Tracking Trajectories Paddy J Slator, Nigel J Burroughs O4 3D-NOME: 3D NucleOme Multiscale Engine for data-driven modeling of three-dimensional genome architecture Przemysław Szałaj, Zhonghui Tang, Paul Michalski, Oskar Luo, Xingwang Li, Yijun Ruan, Dariusz Plewczynski O5 A novel feature selection method to extract multiple adjacent solutions for viral genomic sequences classification Giulia Fiscon, Emanuel Weitschek, Massimo Ciccozzi, Paola Bertolazzi, Giovanni Felici O6 A Systems Biology Compendium for Leishmania donovani Bart Cuypers, Pieter Meysman, Manu Vanaerschot, Maya Berg, Hideo Imamura, Jean-Claude Dujardin, Kris Laukens O7 Unravelling signal coordination from large scale phosphorylation kinetic data Westa Domanova, James R. Krycer, Rima Chaudhuri, Pengyi Yang, Fatemeh Vafaee, Daniel J. Fazakerley, Sean J. Humphrey, David E. James, Zdenka Kuncic Katie Wilkins, Mehedi Hassan, Margherita Francescatto, Jakob B. Jespersen, R. Gonzalo Parra, Bart Cuypers, Dan F. DeBlasio, Alexander Junge, Anupama Jigisha, Farzana Rahman, Griet Laenen, Sander Willems, Lieven Thorrez, Yves Moreau, Raju Nagarajan, Sonia P. Chothani, C. Ramakrishnan, Masakazu Sekijima, M. Michael Gromiha, Paddy Slator, Nigel J. Burroughs, Przemyslaw Szalaj, Zhonghui Tang, Paul J. Michalski, Oskar Luo, Xingwang Li 0004, Yijun Ruan, Dariusz Plewczynski, Giulia Fiscon, Emanuel Weitschek, Massimo Ciccozzi, Paola Bertolazzi, Giovanni Felici, Pieter Meysman, Manu Vanaerschot, Maya Berg, Hideo Imamura, Jean-Claude Dujardin, Kris Laukens, Westa Domanova, James R. Krycer, Rima Chaudhuri, Pengyi Yang, Fatemeh Vafaee, Daniel J. Fazakerley, Sean J. Humphrey, David E. James, Zdenka Kuncic |
BMC Bioinform. | 28 |
| 2013 | Information-Sharing in Three Interacting Minds Solving a Simple Perceptual Task
Michal Denkiewicz, Joanna Raczaszek-Leonardi, Piotr Migdal, Dariusz Plewczynski |
CogSci | 4 |
| 2012 | SVMeFC: SVM Ensemble Fuzzy Clustering for Satellite Image SegmentationabstractThe problem of unsupervised image segmentation of a satellite image in a number of homogeneous regions can be viewed as the task of clustering the pixels in the intensity space. This letter presents an approach that exploits the capability of some recently proposed fuzzy clustering techniques, as well as support vector machine (SVM) classifiers, to yield improved solutions. All the fuzzy clustering techniques are first used to produce a set of different clustering solutions. Each such solution has been improved by a novel technique based on an SVM classifier. Thereafter, the cluster-based similarity partition algorithm is used to create the final clustering solution from all improved ensemble solutions. Results demonstrating the effectiveness of the proposed technique are provided for numeric remote sensing data described in terms of feature vectors. Moreover, a remotely sensed image of Calcutta City has been segmented using the proposed technique to establish its utility. In addition, the additional information of this letter is given as supplementary at http://sysbio.icm.edu.pl/indra/SVMeFC.html. Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2011 | Brainstorming - Agent based Meta-learning Approach
Dariusz Plewczynski |
ICAART (2) | 1 |
| 2011 | PMAFC: A New Probabilistic Memetic Algorithm Based Fuzzy Clustering
Indrajit Saha, Ujjwal Maulik, Dariusz Plewczynski |
ISMIS | 3 |
| 2011 | Protein-protein interaction and pathway databases, a graphical reviewabstractThe amount of information regarding protein-protein interactions (PPI) at a proteomic scale is constantly increasing. This is paralleled with an increase of databases making information available. Consequently there are diverse ways of delivering information about not only PPIs but also regarding the databases themselves. This creates a time consuming obstacle for many researchers working in the field. Our survey provides a valuable tool for researchers to reduce the time necessary to gain a broad overview of PPI-databases and is supported by a graphical representation of data exchange. The graphical representation is made available in cooperation with the team maintaining www.pathguide.org and can be accessed at http://www.pathguide.org/interactions.php in a new Cytoscape web implementation. The local copy of Cytoscape cys file can be downloaded from http://bio.icm.edu.pl/~darman/ppi web page. Tomas Klingström, Dariusz Plewczynski |
Briefings Bioinform. | 2 |
| 2011 | Improvement of new automatic differential fuzzy clustering using SVM classifier for microarray analysis
Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski |
Expert Syst. Appl. | 4 |
| 2011 | Unsupervised and Supervised Learning Approaches Together for Microarray AnalysisabstractIn this article, a novel concept is introduced by using both unsupervised and supervised learning. For unsupervised learning, the problem of fuzzy clustering in microarray data as a multiobjective optimization is used, which simultaneously optimizes two internal fuzzy cluster validity indices to yield a set of Pareto-optimal clustering solutions. In this regards, a new multiobjective differential evolution based fuzzy clustering technique has been proposed. Subsequently, for supervised learning, a fuzzy majority voting scheme along with support vector machine is used to integrate the clustering information from all the solutions in the resultant Pareto-optimal set. The performances of the proposed clustering techniques have been demonstrated on five publicly available benchmark microarray data sets. A detail comparison has been carried out with multiobjective genetic algorithm based fuzzy clustering, multiobjective differential evolution based fuzzy clustering, single objective versions of differential evolution and genetic algorithm based fuzzy clustering as well as well known fuzzy c-means algorithm. While using support vector machine, comparative studies of the use of four different kernel functions are also reported. Statistical significance test has been done to establish the statistical superiority of the proposed multiobjective clustering approach. Finally, biological significance test has been carried out using a web based gene annotation tool to show that the proposed integrated technique is able to produce biologically relevant clusters of coexpressed genes. Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski |
Fundam. Informaticae | 4 |
| 2010 | Real-coded differential crisp clustering for MRI brain image segmentationabstractIn this paper, a segmentation technique of multi-spectral magnetic resonance image of the brain using a new differential evolution based crisp clustering is proposed. Real-coded encoding of the cluster centres is used for this purpose. Here assignments of points to different clusters are made based on the Euclidean distance. The proposed method is applied on several simulated T1-weighted, T2-weighted and proton density for normal and MS lesion magnetic resonance brain images. Superiority of the proposed method over genetic algorithm based crisp clustering, simulated annealing based crisp clustering, K-means and average linkage are demonstrated quantitatively. Segmentation obtained by differential evolution based crisp clustering technique is also compared with the available ground truth information. Also statistical analysis has been conducted to judge the effectiveness. Matlab version of the software is available at http://bio.icm.edu.pl/~darman/MRI. Indrajit Saha, Ujjwal Maulik, Sanghamitra Bandyopadhyay, Dariusz Plewczynski |
IEEE Congress on Evolutionary Computation | 4 |
| 2010 | AMS 3.0: prediction of post-translational modificationsabstractBACKGROUND: We present here the recent update of AMS algorithm for identification of post-translational modification (PTM) sites in proteins based only on sequence information, using artificial neural network (ANN) method. The query protein sequence is dissected into overlapping short sequence segments. Ten different physicochemical features describe each amino acid; therefore nine residues long segment is represented as a point in a 90 dimensional space. The database of sequence segments with confirmed by experiments post-translational modification sites are used for training a set of ANNs. RESULTS: The efficiency of the classification for each type of modification and the prediction power of the method is estimated here using recall (sensitivity), precision values, the area under receiver operating characteristic (ROC) curves and leave-one-out tests (LOOCV). The significant differences in the performance for differently optimized neural networks are observed, yet the AMS 3.0 tool integrates those heterogeneous classification schemes into the single consensus scheme, and it is able to boost the precision and recall values independent of a PTM type in comparison with the currently available state-of-the art methods. CONCLUSIONS: The standalone version of AMS 3.0 presents an efficient way to identify post-translational modifications for whole proteomes. The training datasets, precompiled binaries for AMS 3.0 tool and the source code are available at http://code.google.com/p/automotifserver under the Apache 2.0 license scheme. Subhadip Basu, Dariusz Plewczynski |
BMC Bioinform. | 2 |
| 2006 | PDB-UF: database of predicted enzymatic functions for unannotated protein structures from structural genomicsabstractBACKGROUND: The number of protein structures from structural genomics centers dramatically increases in the Protein Data Bank (PDB). Many of these structures are functionally unannotated because they have no sequence similarity to proteins of known function. However, it is possible to successfully infer function using only structural similarity. RESULTS: Here we present the PDB-UF database, a web-accessible collection of predictions of enzymatic properties using structure-function relationship. The assignments were conducted for three-dimensional protein structures of unknown function that come from structural genomics initiatives. We show that 4 hypothetical proteins (with PDB accession codes: 1VH0, 1NS5, 1O6D, and 1TO0), for which standard BLAST tools such as PSI-BLAST or RPS-BLAST failed to assign any function, are probably methyltransferase enzymes. CONCLUSION: We suggest that the structure-based prediction of an EC number should be conducted having the different similarity score cutoff for different protein folds. Moreover, performing the annotation using two different algorithms can reduce the rate of false positive assignments. We believe, that the presented web-based repository will help to decrease the number of protein structures that have functions marked as "unknown" in the PDB file. AVAILABILITY: http://paradox.harvard.edu/PDB-UF and http://bioinfo.pl/PDB-UF. Marcin von Grotthuss, Dariusz Plewczynski, Krzysztof Ginalski, Leszek Rychlewski, Eugene I. Shakhnovich |
BMC Bioinform. | 2 |
| 2005 | AutoMotif server: prediction of single residue post-translational modifications in proteinsabstractUNLABELLED: The AutoMotif Server allows for identification of post-translational modification (PTM) sites in proteins based only on local sequence information. The local sequence preferences of short segments around PTM residues are described here as linear functional motifs (LFMs). Sequence models for all types of PTMs are trained by support vector machine on short-sequence fragments of proteins in the current release of Swiss-Prot database (phosphorylation by various protein kinases, sulfation, acetylation, methylation, amidation, etc.). The accuracy of the identification is estimated using the standard leave-one-out procedure. The sensitivities for all types of short LFMs are in the range of 70%. AVAILABILITY: The AutoMotif Server is available free for academic use at http://automotif.bioinfo.pl/ Dariusz Plewczynski, Adrian Tkacz, Lucjan Stanislaw Wyrwicz, Leszek Rychlewski |
Bioinform. | 1 |
| 2004 | Integrated web service for improving alignment quality based on segments comparisonabstractBACKGROUND: Defining blocks forming the global protein structure on the basis of local structural regularity is a very fruitful idea, extensively used in description, and prediction of structure from only sequence information. Over many years the secondary structure elements were used as available building blocks with great success. Specially prepared sets of possible structural motifs can be used to describe similarity between very distant, non-homologous proteins. The reason for utilizing the structural information in the description of proteins is straightforward. Structural comparison is able to detect approximately twice as many distant relationships as sequence comparison at the same error rate. RESULTS: Here we provide a new fragment library for Local Structure Segment (LSS) prediction called FRAGlib which is integrated with a previously described segment alignment algorithm SEA. A joined FRAGlib/SEA server provides easy access to both algorithms, allowing a one stop alignment service using a novel approach to protein sequence alignment based on a network matching approach. The FRAGlib used as secondary structure prediction achieves only 73% accuracy in Q3 measure, but when combined with the SEA alignment, it achieves a significant improvement in pairwise sequence alignment quality, as compared to previous SEA implementation and other public alignment algorithms. The FRAGlib algorithm takes approximately 2 min. to search over FRAGlib database for a typical query protein with 500 residues. The SEA service align two typical proteins within circa approximately 5 min. All supplementary materials (detailed results of all the benchmarks, the list of test proteins and the whole fragments library) are available for download on-line at http://ffas.ljcrf.edu/darman/results/. CONCLUSIONS: The joined FRAGlib/SEA server will be a valuable tool both for molecular biologists working on protein sequence analysis and for bioinformaticians developing computational methods of structure prediction and alignment of proteins. Dariusz Plewczynski, Leszek Rychlewski, Yuzhen Ye, Lukasz Jaroszewski, Adam Godzik |
BMC Bioinform. | 1 |