VLDB 2026 Research / reviewers in the wild / expert
Mehmet Baysan
dblp:32/5745
· DBLP profile ↗
12ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-7359-2965ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 3 · 2 first-authorTheory of computation · 2Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From ideal to practical: Heterogeneity of student-generated variant lists highlights hidden reproducibility gapsabstractNext-generation sequencing (NGS) technologies offer detailed and inexpensive identification of the genetic structure of living organisms. The massive data volume necessitates the utilization of advanced computational resources for analyses. However, the rapid accumulation of data and the urgent need for analysis tools have caused the development of imperfect software solutions. Given their immense potential in clinical applications and the recent reproducibility crisis discussions in science and technology, these tools must be thoroughly examined. Typically, NGS data analysis tools are benchmarked under homogeneous conditions, with well-trained personnel and ideal hardware and data environments. However, in the real world, these analyses are done under heterogeneous conditions in terms of computing environments and experience levels. This difference is mostly overlooked, therefore studies that examine NGS workflows generated under various conditions would be highly valuable. Moreover, a detailed assessment of the difficulties faced by the trainees would allow for improved educational programs for better NGS analysis training. Considering these needs, we designed an elective undergraduate bioinformatics course project for computer engineering students at Istanbul Technical University. Students were tasked to perform and compare 12 different somatic variant calling pipelines on the recently published SEQC2 dataset. Upon examining the results, we have realized that despite seeming correct, the final variant lists created by different student groups display a high level of heterogeneity. Notably, the operating systems and installation methods were the most influential factors in variant-calling performance. Here, we present detailed evaluations of our case study and provide insights for better bioinformatics training. Rumeysa Aslihan Erturk, Abdullah Asim Emül, Busra Nur Darendeli Kiraz, Fatma Zehra Sari, Mehmet Arif Ergun, Mehmet Baysan |
PLoS Comput. Biol. | 6 |
| 2024 | Improving somatic exome sequencing performance by biological replicatesabstractBACKGROUND: Next-generation sequencing (NGS) technologies offer fast and inexpensive identification of DNA sequences. Somatic sequencing is among the primary applications of NGS, where acquired (non-inherited) variants are based on comparing diseased and healthy tissues from the same individual. Somatic mutations in genetic diseases such as cancer are tightly associated with genomic instability. Genomic instability increases heterogenity, complicating sequencing efforts further, a task already challenged by the presence of short reads and repetitions in human DNA. This leads to low concordance among studies and limits reproducibility. This limitation is a significant problem since identified mutations in somatic sequencing are major biomarkers for diagnosis and the primary input of targeted therapies. Benchmarking studies were conducted to assess the error rates and increase reproducibility. Unfortunately, the number of somatic benchmarking sets is very limited due to difficulties in validating true somatic variants. Moreover, most NGS benchmarking studies are based on relatively simpler germline (inherited) sequencing. Recently, a comprehensive somatic sequencing benchmarking set was published by Sequencing Quality Control Phase 2 (SEQC2). We chose this dataset for our experiments because it is a well-validated, cancer-focused dataset that includes many tumor/normal biological replicates. Our study has two primary goals. First goal is to determine how replicate-based consensus approaches can improve the accuracy of somatic variant detection systems. Second goal is to develop highly predictive machine learning (ML) models by employing replicate-based consensus variants as labels during the training phase. RESULTS: Ensemble approaches that combine alternative algorithms are relatively common; here, as an alternative, we study the performance enhancement potential of biological replicates. We first developed replicate-based consensus approaches that utilize the biological replicates available in this study to improve variant calling performance. Subsequently, we trained ML models using these biological replicates and achieved performance comparable to optimal ML models, those trained using high-confidence variants identified in advance. CONCLUSIONS: Our replicate-based consensus approach can be used to improve variant calling performance and develop efficient ML models. Given the relative ease of obtaining biological replicates, this strategy allows for the development of efficient ML models tailored to specific datasets or scenarios. Yunus Emre Cebeci, Rumeysa Aslihan Erturk, Mehmet Arif Ergun, Mehmet Baysan |
BMC Bioinform. | 4 |
| 2024 | Correction: Improving somatic exome sequencing performance by biological replicatesabstracthttps://doi.org/10.1186/s12859-024-05828-0 Yunus Emre Cebeci, Rumeysa Aslihan Erturk, Mehmet Arif Ergun, Mehmet Baysan |
BMC Bioinform. | 4 |
| 2024 | VCF observer: a user-friendly software tool for preliminary VCF file analysis and comparisonabstractBACKGROUND: Advancements over the past decade in DNA sequencing technology and computing power have created the potential to revolutionize medicine. There has been a marked increase in genetic data available, allowing for the advancement of areas such as personalized medicine. A crucial type of data in this context is genetic variant data which is stored in variant call format (VCF) files. However, the rapid growth in genomics has presented challenges in analyzing and comparing VCF files. RESULTS: In response to the limitations of existing tools, this paper introduces a novel web application that provides a user-friendly solution for VCF file analyses and comparisons. The software tool enables researchers and clinicians to perform high-level analysis with ease and enhances productivity. The application's interface allows users to conveniently upload, analyze, and visualize their VCF files using simple drag-and-drop and point-and-click operations. Essential visualizations such as Venn diagrams, clustergrams, and precision-recall plots are provided to users. A key feature of the application is its support for metadata-based file grouping, accomplished through flexible data matrix uploads, streamlining organization and analysis of user-defined categories. Additionally, the application facilitates standardized benchmarking of VCF files by integrating user-provided ground truth regions and variant lists. CONCLUSIONS: By providing a user-friendly interface and supporting essential visualizations, this software enhances the accessibility of VCF file analysis and assists researchers and clinicians in their scientific inquiries. Abdullah Asim Emül, Mehmet Arif Ergun, Rumeysa Aslihan Erturk, Omer Cinal, Mehmet Baysan |
BMC Bioinform. | 5 |
| 2024 | COSAP: Comparative Sequencing Analysis PlatformabstractBACKGROUND: Recent improvements in sequencing technologies enabled detailed profiling of genomic features. These technologies mostly rely on short reads which are merged and compared to reference genome for variant identification. These operations should be done with computers due to the size and complexity of the data. The need for analysis software resulted in many programs for mapping, variant calling and annotation steps. Currently, most programs are either expensive enterprise software with proprietary code which makes access and verification very difficult or open-access programs that are mostly based on command-line operations without user interfaces and extensive documentation. Moreover, a high level of disagreement is observed among popular mapping and variant calling algorithms in multiple studies, which makes relying on a single algorithm unreliable. User-friendly open-source software tools that offer comparative analysis are an important need considering the growth of sequencing technologies. RESULTS: Here, we propose Comparative Sequencing Analysis Platform (COSAP), an open-source platform that provides popular sequencing algorithms for SNV, indel, structural variant calling, copy number variation, microsatellite instability and fusion analysis and their annotations. COSAP is packed with a fully functional user-friendly web interface and a backend server which allows full independent deployment for both individual and institutional scales. COSAP is developed as a workflow management system and designed to enhance cooperation among scientists with different backgrounds. It is publicly available at https://cosap.bio and https://github.com/MBaysanLab/cosap/ . The source code of the frontend and backend services can be found at https://github.com/MBaysanLab/cosap-webapi/ and https://github.com/MBaysanLab/cosap_frontend/ respectively. All services are packed as Docker containers as well. Pipelines that combine algorithms can be customized and new algorithms can be added with minimal coding through modular structure. CONCLUSIONS: COSAP simplifies and speeds up the process of DNA sequencing analyses providing commonly used algorithms for SNV, indel, structural variant calling, copy number variation, microsatellite instability and fusion analysis as well as their annotations. COSAP is packed with a fully functional user-friendly web interface and a backend server which allows full independent deployment for both individual and institutional scales. Standardized implementations of popular algorithms in a modular platform make comparisons much easier to assess the impact of alternative pipelines which is crucial in establishing reproducibility of sequencing analyses. Mehmet Arif Ergun, Omer Cinal, Berkant Bakisli, Abdullah Asim Emül, Mehmet Baysan |
BMC Bioinform. | 5 |
| 2013 | Batching and delivery in semi-online distribution systems
Igor Averbakh, Mehmet Baysan |
Discret. Appl. Math. | 2 |
| 2012 | Polynomial time solution to minimum forwarding set problem in wireless networks under disk coverage model
Mehmet Baysan, Kamil Saraç, Ramaswamy Chandrasekaran |
Ad Hoc Networks | 1 |
| 2011 | On a labeling problem in graphs
Ramaswamy Chandrasekaran, Milind Dawande, Mehmet Baysan |
Discret. Appl. Math. | 3 |
| 2009 | A Polynomial Time Solution to Minimum Forwarding Set Problem in Wireless Networks under Unit Disk Coverage ModelabstractNetwork-wide broadcast (simply broadcast) is a frequently used operation in wireless ad hoc networks (WANETs). One promising practical approach for energy-efficient broadcast is to use localized algorithms to minimize the number of nodes involved in the propagation of the broadcast messages. In this context, the minimum forwarding set problem (MFSP) (also known as multipoint relay (MPR) problem) has received a considerable attention in the research community. Even though the general form of the problem is shown to be NP-complete, the complexity of the problem has not been known under the practical application context of ad hoc networks. In this paper, we present a polynomial time algorithm to solve the MFSP for wireless network under unit disk coverage model. We prove the existence of some geometrical properties for the problem and then propose a polynomial time algorithm to build an optimal solution based on these properties. To the best of our knowledge, our algorithm is the first polynomial time solution to the MFSP under the unit disk coverage model. We believe that the work presented in this paper will have an impact on the design and development of new algorithms for several wireless network applications including energy-efficient multicast, broadcast, and topology control protocols for WANETs and sensor networks. Mehmet Baysan, Kamil Saraç, Ramaswamy Chandrasekaran, Sergey Bereg |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2008 | Variable power broadcast using local information in ad hoc networks
Avinash Chiganmi, Mehmet Baysan, Kamil Saraç, Ravi Prakash 0001 |
Ad Hoc Networks | 2 |
| 2006 | Cluster based approaches for end-to-end complete feedback collection in multicastabstractIn this paper we study the end-to-end complete feedback collection (ECFC) problem in large scale multi-cast applications. We consider the case where each receiver is expected to send feedback in a timely manner without causing implosion at the source site. To address the scalability problem and improve timely feedback collection, we introduce the use of clustering algorithms for feedback collection. Our simulation based comparisons show that the clustering based approaches outperform the existing pure (without clustering) multi-round probabilistic and pure (without clustering) delayed feedback collection approaches both in terms of collection delay and message overhead Mehmet Baysan, Kamil Saraç |
IPCCC | 1 |
| 2004 | Improving Energy Savings in Power Adaptive Broadcasting in MANETsabstractNetwork wide broadcast is a frequently used operation in mobile ad hoc networks (MANETs). Many unicast protocols including dynamic source routing (DSR) and ad hoc on-demand distance vector (AODV) protocol use broadcast to discover the unicast routes toward destinations. In addition, depending on the application, broadcast can be used to deliver actual data packets to all the nodes in the network. Nodes in MANETs work with limited battery power and the efficient utilization of this power is important for increasing the lifetime of the individual nodes as well as the overall network. As a result, it is important to utilize energy efficient algorithms in achieving network wide broadcast in MANETs. In this poster paper, the authors improve the energy efficiency in power adaptive broadcast using local information. In this proposal, the authors introduce an improved algorithm for deciding the transmission power level. The paper also attempts to further increase the energy savings by reducing the redundant transmission by including an efficient forward node set at the selection algorithm. Mehmet Baysan, Saipriya Gowdamachandran, Kamil Saraç |
BROADNETS | 1 |