VLDB 2026 Research / reviewers in the wild / expert
Burcu Bakir-Gungor
dblp:81/7242 · also Burcu Bakir, Burcu Bakir-Güngör
· DBLP profile ↗
10ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-2272-6270ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GenShare: A Blockchain-Based Genomic Data Sharing PlatformabstractEvery day, hundreds of gigabytes of data are produced due to the exponential growth of next-generation sequencing and omics technologies. By combining omics data with other data types, such as electronic health record data, panomics research is actively attempting to uncover novel and potentially useful biomarkers. For the effective analysis of high-throughput-derived omics data, it is imperative to establish robust and reliable platforms that prioritize ethical considerations while effectively managing privacy, ownership concerns, and the responsible sharing of data. The GenShare model was proposed to provide an efficient platform that fits these needs. GenShare is a hybrid platform that utilizes blockchain technology. Paillier’s homomorphic encryption scheme in tandem with Intel Software Guard Extension (SGX) serves to enable the sharing of genomic data, execution of count queries, and statistical analysis of genomic data while preserving privacy and avoiding compromise of sensitive information. The objective of this paradigm is to confront security and privacy concerns through the integration of homomorphic encryption and SGX, addressing additional challenges associated with Hyperledger Fabric and Ethereum. In pursuit of this objective, the implementation of the system involved establishing the Hyperledger Fabric network, with various workloads employed to assess the network’s efficiency. Consequently, it was hypothesized that the new GenShare model would enhance the data collection and dissemination cycle and serve as a proficient platform catering to the needs of its users. Beyhan Adanur Dedeturk, Ahmet Soran, Burcu Bakir-Gungor |
Distributed Ledger Technol. Res. Pract. | 3 |
| 2025 | Breast Cancer Detection Using a New Parallel Hybrid Logistic Regression Model Trained by Particle Swarm Optimization and Clonal Selection AlgorithmsabstractABSTRACT Breast cancer is one of the most widespread kinds of cancer, especially in women, and it has a high mortality rate. With the help of technology, it is possible to develop a computer‐aided method for the diagnosis of breast cancer, which is crucial for effective treatment. Recent breast cancer diagnosis studies utilizing numerous machine learning models were efficient and innovative. However, it has been observed that they may have problems such as long training times and low accuracy rates. To this end, in this study, we present a new classifier that utilizes a hybrid of the clonal selection algorithm (CSA) and the particle swarm optimization (PSO) algorithm for the training of the logistic regression (LR) model, which is named CSA‐PSO‐LR. The proposed method is evaluated using two publicly accessible breast cancer datasets, that is, the Wisconsin Diagnostic Breast Cancer (WDBC) database and the Wisconsin Breast Cancer Database (WBCD), with 10‐fold cross‐validation and Bayesian hyperparameter optimization techniques. Additionally, a CPU parallelization method is applied, which substantially shortens the training time of the model. The efficacy of the CSA‐PSO‐LR classifier is compared with state‐of‐the‐art machine learning algorithms and related studies in the literature. Performance analysis indicates that the proposed method achieves 98.75% accuracy and 98.27% F1‐score on the WDBC dataset, and 97.94% accuracy and 97.35% F1‐score on the WBCD dataset. These results demonstrate the potential of the proposed method as an effective approach for improving breast cancer diagnosis. Mustafa Etcil, Bilge Kagan Dedeturk, Burak Kolukisa, Burcu Bakir-Gungor, Vehbi C. Gungor |
Concurr. Comput. Pract. Exp. | 4 |
| 2024 | SEMANT - Feature Group Selection Utilizing FastText-Based Semantic Word Grouping, Scoring, and Modeling Approach for Text Classification
Daniel Voskergian, Burcu Bakir-Gungor, Malik Yousef |
DEXA (2) | 2 |
| 2024 | Novel Antimicrobial Peptide Design Using Motif Match Score RepresentationabstractAntimicrobial peptides (AMPs) have drawn the interest of the researchers since they offer an alternative to the traditional antibiotics in the fight against antibiotic resistance and they exhibit additional pharmaceutically significant properties. Recently, computational approaches attemp to reveal how antibacterial activity is determined from a machine learning perspective and they aim to search and find the biological cues or characteristics that control antimicrobial activity via incorporating motif match scores. This study is dedicated to the development of a machine learning framework aimed at devising novel antimicrobial peptide (AMP) sequences potentially effective against Gram-positive /Gram-negative bacteria. In order to design newly generated sequences classified as either AMP or non-AMP, various classification models were trained. These novel sequences underwent validation utilizing the "DBAASP:strain-specific antibacterial prediction based on machine learning approaches and data on AMP sequences" tool. The findings presented herein represent a significant stride in this computational research, streamlining the process of AMP creation or modification within wet lab environments. Ümmü Gülsüm Söylemez, Malik Yousef, Zülal Kesmen, Burcu Bakir-Gungor |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | miRcorrNetPro: Unraveling Algorithmic Insights through Cross-Validation in Multi-Omics Integration for Comprehensive Data AnalysisabstractHigh throughput -omics technologies facilitate the investigation of regulatory mechanisms of complex diseases. Along this line, scientists develop promising tools and methods to extend our understanding at the molecular and functional levels. To this end, miRcorrNet tool performs integrative analysis of microRNA (miRNA) and gene expression profiles via machine learning (ML) approach to identify significant miRNA groups and their associated target genes. In this study, we propose miRcorrNetPro tool, which extends miRcorrNet by tracking group scoring, ranking and other information through the cross-validation iterations. Heatmap visualizations enable deep novel insights into the collective behavior of clusters of groups in cellular signaling and hence facilitate detection of potential biomarkers for the disease under investigation. Although miRcorrNetPro is designed as a generic tool, here we present our findings and potential miRNA biomarkers for Breast Cancer (BRCA). The miRcorrNetPro tool and all other supplementary files are available at https://github.com/Miray-Unlu/miRcorrNetPro. Miray Unlu Yazici, Malik Yousef, J. S. Marron, Burcu Bakir-Gungor |
BIBM | 4 |
| 2023 | PriPath: identifying dysregulated pathways from differential gene expression via grouping, scoring, and modeling with an embedded feature selection approachabstractBACKGROUND: Cell homeostasis relies on the concerted actions of genes, and dysregulated genes can lead to diseases. In living organisms, genes or their products do not act alone but within networks. Subsets of these networks can be viewed as modules that provide specific functionality to an organism. The Kyoto encyclopedia of genes and genomes (KEGG) systematically analyzes gene functions, proteins, and molecules and combines them into pathways. Measurements of gene expression (e.g., RNA-seq data) can be mapped to KEGG pathways to determine which modules are affected or dysregulated in the disease. However, genes acting in multiple pathways and other inherent issues complicate such analyses. Many current approaches may only employ gene expression data and need to pay more attention to some of the existing knowledge stored in KEGG pathways for detecting dysregulated pathways. New methods that consider more precompiled information are required for a more holistic association between gene expression and diseases. RESULTS: PriPath is a novel approach that transfers the generic process of grouping and scoring, followed by modeling to analyze gene expression with KEGG pathways. In PriPath, KEGG pathways are utilized as the grouping function as part of a machine learning algorithm for selecting the most significant KEGG pathways. A machine learning model is trained to differentiate between diseases and controls using those groups. We have tested PriPath on 13 gene expression datasets of various cancers and other diseases. Our proposed approach successfully assigned biologically and clinically relevant KEGG terms to the samples based on the differentially expressed genes. We have comparatively evaluated the performance of PriPath against other tools, which are similar in their merit. For each dataset, we manually confirmed the top results of PriPath in the literature and found that most predictions can be supported by previous experimental research. CONCLUSIONS: PriPath can thus aid in determining dysregulated pathways, which applies to medical diagnostics. In the future, we aim to advance this approach so that it can perform patient stratification based on gene expression and identify druggable targets. Thereby, we cover two aspects of precision medicine. Malik Yousef, Fatma Ozdemir, Amhar Jaber, Jens Allmer, Burcu Bakir-Gungor |
BMC Bioinform. | 5 |
| 2022 | The Determination of Distinctive Single Nucleotide Polymorphism Sets for the Diagnosis of Behçet's DiseaseabstractBehçet's Disease (BD) is a multi-system inflammatory disorder in which the etiology remains unclear. The most probable hypothesis is that genetic tendency and environmental factors play roles in the development of BD. In order to find the essential reasons, genetic changes on thousands of genes should be analyzed. Besides, there is a need for extra analysis to find out which genetic factor affects the disease. Machine learning approaches have high potential for extracting the knowledge from genomics and selecting the representative Single Nucleotide Polymorphisms (SNPs) as the most effective features for the clinical diagnosis process. In this study, we have attempted to identify representative SNPs using feature selection methods, incorporating biological information and aimed to develop a machine-learning model for diagnosing Behçet's disease. By combining biological information and machine learning classifiers, up to 99.64 percent accuracy of disease prediction is achieved using only 13,611 out of 311,459 SNPs. In addition, we revealed the SNPs that are most distinctive by performing repeated feature selection in cross-validation experiments. Yunus Emre Isik, Yasin Görmez, Zafer Aydin, Burcu Bakir-Gungor |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2018 | Evaluation of Classification Algorithms, Linear Discriminant Analysis and a New Hybrid Feature Selection Methodology for the Diagnosis of Coronary Artery DiseaseabstractAccording to the World Health Organization (WHO), 31% of the world's total deaths in 2016 (17.9 million) was due to cardiovascular diseases (CVD). With the development of information technologies, it has become possible to predict whether people have heart diseases or not by checking certain physical and biochemical values at a lower cost. In this study, we have evalated a set of different classification algorithms, linear discriminant analysis and proposed a new hybrid feature selection methodology for the diagnosis of coronary heart diseases (CHD). Throughout this research effort, using three publicly available Heart Disease diagnosis datasets (UCI Machine Learning Repository), we have conducted comparative performance evaluations in terms of accuracy, sensitivity, specificity, F-measure, AUC and running time. Burak Kolukisa, Hilal Hacilar, Gokhan Goy, Mustafa Kus, Burcu Bakir-Gungor, Atilla Aral, Vehbi C. Gungor |
IEEE BigData | 5 |
| 2014 | PANOGA: a web server for identification of SNP-targeted pathways from genome-wide association study dataabstractAbstract Summary: Genome-wide association studies (GWAS) have revolutionized the search for the variants underlying human complex diseases. However, in a typical GWAS, only a minority of the single-nucleotide polymorphisms (SNPs) with the strongest evidence of association is explained. One possible reason of complex diseases is the alterations in the activity of several biological pathways. Here we present a web server called Pathway and Network-Oriented GWAS Analysis to devise functionally important pathways through the identification of SNP-targeted genes within these pathways. The strength of our methodology stems from its multidimensional perspective, where we combine evidence from the following five resources: (i) genetic association information obtained through GWAS, (ii) SNP functional information, (iii) protein–protein interaction network, (iv) linkage disequilibrium and (v) biochemical pathways. Availability: PANOGA web server is freely available at: http://panoga.sabanciuniv.edu/. The source code is available to academic users ‘as is’ on request. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Burcu Bakir-Gungor, Ece Egemen, Osman Ugur Sezerman |
Bioinform. | 1 |
| 2014 | HomSI: a homozygous stretch identifier from next-generation sequencing dataabstractUNLABELLED: In consanguineous families, as a result of inheriting the same genomic segments through both parents, the individuals have stretches of their genomes that are homozygous. This situation leads to the prevalence of recessive diseases among the members of these families. Homozygosity mapping is based on this observation, and in consanguineous families, several recessive disease genes have been discovered with the help of this technique. The researchers typically use single nucleotide polymorphism arrays to determine the homozygous regions and then search for the disease gene by sequencing the genes within this candidate disease loci. Recently, the advent of next-generation sequencing enables the concurrent identification of homozygous regions and the detection of mutations relevant for diagnosis, using data from a single sequencing experiment. In this respect, we have developed a novel tool that identifies homozygous regions using deep sequence data. Using *.vcf (variant call format) files as an input file, our program identifies the majority of homozygous regions found by microarray single nucleotide polymorphism genotype data. AVAILABILITY AND IMPLEMENTATION: HomSI software is freely available at www.igbam.bilgem.tubitak.gov.tr/softwares/HomSI, with an online manual. Zeliha Gormez, Burcu Bakir-Gungor, Mahmut Samil Sagiroglu |
Bioinform. | 2 |