EDBT 2026 Demo / reviewers in the wild / expert
Gurhan Gunduz
dblp:13/723
· DBLP profile ↗
9ranked-venue papers
1as first author
4since 2021 · last 2022
0000-0002-0719-2688ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Computer networks · 1 · 1 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Stochastic gradient descent-based support vector machines training optimization on Big Data and HPC frameworksabstractSummary Support vector machines (SVM) is a widely used machine learning algorithm. With the increasing amount of research data nowadays, understanding how to do efficient training is more important than ever. This article discusses the performance optimizations and benchmarks related to providing high‐performance support for SVM training. In this research, we have focused on a highly scalable gradient descent‐based approach to implementing the core SVM algorithm. In providing a scalable solution, we have designed optimized high‐performance computing and dataflow‐oriented SVM implementations. A high‐performance computing approach means the algorithm is implemented with the bulk synchronous parallel (BSP) model. In addition, we analyzed the language level optimizations and math kernel optimizations on a prominent HPC modeling programming language (C++) and dataflow modeling programming language (Java). In the experiments, we compared the performance of classic HPC models, classic dataflow models, and hybrid models designed on classic HPC and dataflow programming models. Our research illustrates a scientific approach in designing the SVM algorithm at scale in classic HPC, dataflow, and hybrid systems. Vibhatha Abeykoon, Geoffrey C. Fox, Saliya Ekanayake, Supun Kamburugamuve, Kannan Govindarajan, Pulasthi Wickramasinghe, Niranda Perera, Chathura Widanage, Ahmet Uyar, Gurhan Gunduz, Selahatin Akkas |
Concurr. Comput. Pract. Exp. | 11 |
| 2022 | A framework for investigating search engines' stemming mechanisms: A case study on BingabstractSummary Big data attracts the attention of governments and a lot of companies today. The developments in technology and the Internet make it one of the important sources of big data. It is easy to get lost in the enormous amount of information contained on the Internet if there were no search engines. Knowing how the search engines work will be helpful to access the desired information. This work aims to be a guide for accessing the right information and also to help to understand search engine stemming and indexing algorithm for interested parties. In this article, we have developed a framework that could be used to investigate the stemming mechanisms of search engines. Our framework also uses Word2vec to analyze semantic relations. We have used our framework to investigate the stemming algorithm of the search engine Bing for English language. In order to achieve that we have used this framework to select words, create queries, send them to Bing, and finally analyze the millions of returned results. We have discussed the results in the context of our article. The results indicate that our framework is useful for analyzing the stemming mechanisms of search engines. Fatmana Sentürk, Gurhan Gunduz |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | Twister2 Cross-platform resource scheduler for big dataabstractAbstract Twister2 is an open‐source big data hosting environment designed to process both batch and streaming data at scale. Twister2 runs jobs in both high‐performance computing (HPC) and big data clusters. It provides a cross‐platform resource scheduler to run jobs in diverse environments. Twister2 is designed with a layered architecture to support various clusters and big data problems. In this paper, we present the cross‐platform resource scheduler of Twister2. We identify required services and explain implementation details. We present job startup delays for single jobs and multiple concurrent jobs in Kubernetes and OpenMPI clusters. We compare job startup delays for Twister2 and Spark at a Kubernetes cluster. In addition, we compare the performance of terasort algorithm on Kubernetes and bare metal clusters at AWS cloud. Ahmet Uyar, Gurhan Gunduz, Supun Kamburugamuve, Pulasthi Wickramasinghe, Chathura Widanage, Kannan Govindarajan, Niranda Perera, Vibhatha Abeykoon, Selahattin Akkas, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | High-performance iterative dataflow abstractions in Twister2: TSetabstractSummary The dataflow model is gradually becoming the de facto standard for big data applications. While many popular frameworks are built around this model, very little research has been done on understanding its inner workings, which in turn has led to inefficiencies in existing frameworks. It is important to note that understanding the relationship between dataflow and high performance computing (HPC) building blocks allows us to address and alleviate many of these fundamental inefficiencies by learning from the extensive research literature in the HPC community. In this article, we present TSets, the dataflow abstraction of Twister2, which is a big data framework designed for high‐performance dataflow and iterative computations. We discuss the dataflow model adopted by TSets and the rationale behind implementing iteration handling at the worker level. Finally, we evaluate TSets to show the performance of the framework and the importance of the worker level iteration model. Pulasthi Wickramasinghe, Niranda Perera, Supun Kamburugamuve, Kannan Govindarajan, Vibhatha Abeykoon, Chathura Widanage, Ahmet Uyar, Gurhan Gunduz, Selahattin Akkas, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 8 |
| 2019 | Streaming Machine Learning Algorithms with Big Data SystemsabstractDesigning low latency applications that can process large volumes data with higher efficiency is a challenging problem. With the limited time to process data, usage of online algorithms are becoming important in the big-data applications. Stream processing is a well-known area that has been studied for a long time. In this research, our objective is to use state of the art big-data analytic engines to implement online algorithms and compare the strengths and weaknesses in each system. We use a streaming version of Support Vector Machines (SVM) and KMeans to do the analysis. Apache Flink, Apache Storm and Twister2 streaming frameworks are used to implement these algorithms. Our study focuses on the efficiency of online training of these algorithms and the results show higher performance in Twister2 framework for these algorithms. Vibhatha Abeykoon, Gregor von Laszewski, Supun Kamburugamuve, Kannan Govindarajan, Pulasthi Wickramasinghe, Chathura Widanage, Niranda Perera, Ahmet Uyar, Gurhan Gunduz, Selahattin Akkas |
IEEE BigData | 9 |
| 2018 | Twister: Net - Communication Library for Big Data Processing in HPC and Cloud EnvironmentsabstractStreaming processing and batch data processing are the dominant forms of big data analytics today, with numerous systems such as Hadoop, Spark, and Heron designed to process the ever-increasing explosion of data. Generally, these systems are developed as single projects with aspects such as communication, task management, and data management integrated together. By contrast, we take a component-based approach to big data by developing the essential features of a big data system as independent components with polymorphic implementations to support different requirements. Consequently, we recognize the requirements of both dataflow used in popular Apache Systems and the Bulk Synchronous Processing communication style common in High-Performance Computing (HPC) for different applications. Message Passing Interface (MPI) implementations are dominant in HPC but there are no such standard libraries available for big data. Twister:Net is a stand-alone, highly optimized dataflow style parallel communication library which can be used by big data systems or advanced users. Twister:Net can work both in cloud environments using TCP or HPC environments using MPI implementations. This paper introduces Twister:Net and compares it with existing systems to highlight its design and performance. Supun Kamburugamuve, Pulasthi Wickramasinghe, Kannan Govindarajan, Ahmet Uyar, Gurhan Gunduz, Vibhatha Abeykoon, Geoffrey C. Fox |
IEEE CLOUD | 5 |
| 2016 | Popularity-based scalable peer-to-peer topology growth
Gurhan Gunduz, Murat Yuksel |
Comput. Networks | 1 |
| 2015 | POSN: A Personal Online Social Network
Esra Erdin, Eric Klukovich, Gurhan Gunduz, Mehmet Hadi Gunes |
SEC | 3 |
| 2002 | Grid services for earthquake scienceabstractAbstract We describe an information system architecture for the ACES (Asia–Pacific Cooperation for Earthquake Simulation) community. It addresses several key features of the field—simulations at multiple scales that need to be coupled together; real‐time and archival observational data, which needs to be analyzed for patterns and linked to the simulations; a variety of important algorithms including partial differential equation solvers, particle dynamics, signal processing and data analysis; a natural three‐dimensional space (plus time) setting for both visualization and observations; the linkage of field to real‐time events both as an aid to crisis management and to scientific discovery. We also address the need to support education and research for a field whose computational sophistication is rapidly increasing and spans a broad range. The information system assumes that all significant data is defined by an XML layer which could be virtual, but whose existence ensures that all data is object‐based and can be accessed and searched in this form. The various capabilities needed by ACES are defined as grid services, which are conformant with emerging standards and implemented with different levels of fidelity and performance appropriate to the application. Grid Services can be composed in a hierarchical fashion to address complex problems. The real‐time needs of the field are addressed by high‐performance implementation of data transfer and simulation services. Further, the environment is linked to real‐time collaboration to support interactions between scientists in geographically distant locations. Copyright © 2002 John Wiley & Sons, Ltd. Geoffrey C. Fox, Sung Hoon Ko, Marlon E. Pierce, Ozgur Balsoy, Jake Kim, Sangmi Lee, Kangseok Kim, Sangyoon Oh 0001, Xi Rao, Mustafa Varank, Hasan Bulut, Gurhan Gunduz, Xiaohong Qiu, Shrideep Pallickara, Ahmet Uyar, Choon-Han Youn |
Concurr. Comput. Pract. Exp. | 12 |