EDBT 2026 Demo / reviewers in the wild / expert
Ahmet Uyar
dblp:81/1693
· DBLP profile ↗
11ranked-venue papers
1as first author
4since 2021 · last 2022
0000-0002-7247-7110ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Stochastic gradient descent-based support vector machines training optimization on Big Data and HPC frameworksabstractSummary Support vector machines (SVM) is a widely used machine learning algorithm. With the increasing amount of research data nowadays, understanding how to do efficient training is more important than ever. This article discusses the performance optimizations and benchmarks related to providing high‐performance support for SVM training. In this research, we have focused on a highly scalable gradient descent‐based approach to implementing the core SVM algorithm. In providing a scalable solution, we have designed optimized high‐performance computing and dataflow‐oriented SVM implementations. A high‐performance computing approach means the algorithm is implemented with the bulk synchronous parallel (BSP) model. In addition, we analyzed the language level optimizations and math kernel optimizations on a prominent HPC modeling programming language (C++) and dataflow modeling programming language (Java). In the experiments, we compared the performance of classic HPC models, classic dataflow models, and hybrid models designed on classic HPC and dataflow programming models. Our research illustrates a scientific approach in designing the SVM algorithm at scale in classic HPC, dataflow, and hybrid systems. Vibhatha Abeykoon, Geoffrey C. Fox, Saliya Ekanayake, Supun Kamburugamuve, Kannan Govindarajan, Pulasthi Wickramasinghe, Niranda Perera, Chathura Widanage, Ahmet Uyar, Gurhan Gunduz, Selahatin Akkas |
Concurr. Comput. Pract. Exp. | 10 |
| 2022 | Twister2 Cross-platform resource scheduler for big dataabstractAbstract Twister2 is an open‐source big data hosting environment designed to process both batch and streaming data at scale. Twister2 runs jobs in both high‐performance computing (HPC) and big data clusters. It provides a cross‐platform resource scheduler to run jobs in diverse environments. Twister2 is designed with a layered architecture to support various clusters and big data problems. In this paper, we present the cross‐platform resource scheduler of Twister2. We identify required services and explain implementation details. We present job startup delays for single jobs and multiple concurrent jobs in Kubernetes and OpenMPI clusters. We compare job startup delays for Twister2 and Spark at a Kubernetes cluster. In addition, we compare the performance of terasort algorithm on Kubernetes and bare metal clusters at AWS cloud. Ahmet Uyar, Gurhan Gunduz, Supun Kamburugamuve, Pulasthi Wickramasinghe, Chathura Widanage, Kannan Govindarajan, Niranda Perera, Vibhatha Abeykoon, Selahattin Akkas, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 1 |
| 2022 | High-performance iterative dataflow abstractions in Twister2: TSetabstractSummary The dataflow model is gradually becoming the de facto standard for big data applications. While many popular frameworks are built around this model, very little research has been done on understanding its inner workings, which in turn has led to inefficiencies in existing frameworks. It is important to note that understanding the relationship between dataflow and high performance computing (HPC) building blocks allows us to address and alleviate many of these fundamental inefficiencies by learning from the extensive research literature in the HPC community. In this article, we present TSets, the dataflow abstraction of Twister2, which is a big data framework designed for high‐performance dataflow and iterative computations. We discuss the dataflow model adopted by TSets and the rationale behind implementing iteration handling at the worker level. Finally, we evaluate TSets to show the performance of the framework and the importance of the worker level iteration model. Pulasthi Wickramasinghe, Niranda Perera, Supun Kamburugamuve, Kannan Govindarajan, Vibhatha Abeykoon, Chathura Widanage, Ahmet Uyar, Gurhan Gunduz, Selahattin Akkas, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 7 |
| 2021 | HPTMT: Operator-Based Architecture for Scalable High-Performance Data-Intensive FrameworksabstractData-intensive applications impact many domains, and their steadily increasing size and complexity demands highperformance, highly usable environments. We integrate a set of ideas developed in various data science and data engineering frameworks. They employ a set of operators on specific data abstractions that include vectors, matrices, tensors, graphs, and tables. Our key concepts are inspired from systems like MPI, HPF (High-Performance Fortran), NumPy, Pandas, Spark, Modin, PyTorch, TensorFlow, RAPIDS(NVIDIA), and OneAPI (Intel). Further, it is crucial to support different languages in everyday use in the Big Data arena, including Python, R, C++, and Java. We note the importance of Apache Arrow and Parquet for enabling language agnostic high performance and interoperability. In this paper, we propose High-Performance Tensors, Matrices and Tables (HPTMT), an operator-based architecture for data-intensive applications, and identify the fundamental principles needed for performance and usability success. We illustrate these principles by a discussion of examples using our software environments, Cylon and Twister2 that embody HPTMT. Supun Kamburugamuve, Chathura Widanage, Niranda Perera, Vibhatha Abeykoon, Ahmet Uyar, Thejaka Amila Kanewala, Gregor von Laszewski, Geoffrey C. Fox |
CLOUD | 5 |
| 2020 | A Fast, Scalable, Universal Approach For Distributed Data AggregationsabstractIn the current era of Big Data, data engineering has transformed into an essential field of study across many branches of science. Advancements in Artificial Intelligence (AI) have broadened the scope of data engineering and opened up new applications in both enterprise and research communities. Aggregations (also termed reduce in functional programming) are an integral functionality in these applications. They are traditionally aimed at generating meaningful information on large data-sets, and today, they are being used for engineering more effective features for complex AI models. Aggregations are usually carried out on top of data abstractions such as tables/ arrays and are combined with other operations such as grouping of values. There are frameworks that excel in the said domains individually. But, we believe that there is an essential requirement for a data analytics tool that can universally integrate with existing frameworks, and thereby increase the productivity and efficiency of the entire data analytics pipeline. Cylon endeavors to fulfill this void. In this paper, we present Cylon's fast and scalable aggregation operations implemented on top of a distributed in-memory table structure that universally integrates with existing frameworks. Niranda Perera, Vibhatha Abeykoon, Chathura Widanage, Supun Kamburugamuve, Thejaka Amila Kanewala, Pulasthi Wickramasinghe, Ahmet Uyar, Hasara Maithree, Damitha Lenadora, Geoffrey C. Fox |
IEEE BigData | 7 |
| 2019 | Streaming Machine Learning Algorithms with Big Data SystemsabstractDesigning low latency applications that can process large volumes data with higher efficiency is a challenging problem. With the limited time to process data, usage of online algorithms are becoming important in the big-data applications. Stream processing is a well-known area that has been studied for a long time. In this research, our objective is to use state of the art big-data analytic engines to implement online algorithms and compare the strengths and weaknesses in each system. We use a streaming version of Support Vector Machines (SVM) and KMeans to do the analysis. Apache Flink, Apache Storm and Twister2 streaming frameworks are used to implement these algorithms. Our study focuses on the efficiency of online training of these algorithms and the results show higher performance in Twister2 framework for these algorithms. Vibhatha Abeykoon, Gregor von Laszewski, Supun Kamburugamuve, Kannan Govindarajan, Pulasthi Wickramasinghe, Chathura Widanage, Niranda Perera, Ahmet Uyar, Gurhan Gunduz, Selahattin Akkas |
IEEE BigData | 8 |
| 2018 | Twister: Net - Communication Library for Big Data Processing in HPC and Cloud EnvironmentsabstractStreaming processing and batch data processing are the dominant forms of big data analytics today, with numerous systems such as Hadoop, Spark, and Heron designed to process the ever-increasing explosion of data. Generally, these systems are developed as single projects with aspects such as communication, task management, and data management integrated together. By contrast, we take a component-based approach to big data by developing the essential features of a big data system as independent components with polymorphic implementations to support different requirements. Consequently, we recognize the requirements of both dataflow used in popular Apache Systems and the Bulk Synchronous Processing communication style common in High-Performance Computing (HPC) for different applications. Message Passing Interface (MPI) implementations are dominant in HPC but there are no such standard libraries available for big data. Twister:Net is a stand-alone, highly optimized dataflow style parallel communication library which can be used by big data systems or advanced users. Twister:Net can work both in cloud environments using TCP or HPC environments using MPI implementations. This paper introduces Twister:Net and compares it with existing systems to highlight its design and performance. Supun Kamburugamuve, Pulasthi Wickramasinghe, Kannan Govindarajan, Ahmet Uyar, Gurhan Gunduz, Vibhatha Abeykoon, Geoffrey C. Fox |
IEEE CLOUD | 4 |
| 2005 | Performance of a possible Grid message infrastructureabstractAbstract In this paper we present the results pertaining to the NaradaBrokering middleware infrastructure. NaradaBrokering is designed to run on a large network of cooperating broker nodes. NaradaBrokering capabilities include, among other things, support for a wide variety of transport protocols, Java Message Service compliance, support for routing JXTA interactions, support for audio/video conferencing applications and, finally, support for multiple constraint specification formats such as XPath, SQL and regular expression queries. This paper demonstrates the suitability of NaradaBrokering to a wide variety of applications and scenarios. Copyright © 2005 John Wiley & Sons, Ltd. Shrideep Pallickara, Geoffrey C. Fox, Ahmet Uyar, Xi Rao, David W. Walker, Beytullah Yildiz |
Concurr. Pract. Exp. | 3 |
| 2004 | Global multimedia collaboration systemabstractAbstract In order to build an integrated collaboration system over heterogeneous collaboration technologies, we propose a global multimedia collaboration system (Global‐MMCS) based on XGSP A/V Web‐Services framework. This system can integrate multiple A/V services, and support various collaboration clients and communities. Now the prototype is being developed and deployed across many universities in U.S.A. and China. Copyright © 2004 John Wiley & Sons, Ltd. Geoffrey C. Fox, Wenjun Wu 0001, Ahmet Uyar, Hasan Bulut, Shrideep Pallickara |
Concurr. Pract. Exp. | 3 |
| 2003 | Integration of SIP VoIP and Messaging Systems with AccessGrid and H.323
Wenjun Wu 0001, Ahmet Uyar, Hasan Bulut, Geoffrey C. Fox |
ICWS | 2 |
| 2002 | Grid services for earthquake scienceabstractAbstract We describe an information system architecture for the ACES (Asia–Pacific Cooperation for Earthquake Simulation) community. It addresses several key features of the field—simulations at multiple scales that need to be coupled together; real‐time and archival observational data, which needs to be analyzed for patterns and linked to the simulations; a variety of important algorithms including partial differential equation solvers, particle dynamics, signal processing and data analysis; a natural three‐dimensional space (plus time) setting for both visualization and observations; the linkage of field to real‐time events both as an aid to crisis management and to scientific discovery. We also address the need to support education and research for a field whose computational sophistication is rapidly increasing and spans a broad range. The information system assumes that all significant data is defined by an XML layer which could be virtual, but whose existence ensures that all data is object‐based and can be accessed and searched in this form. The various capabilities needed by ACES are defined as grid services, which are conformant with emerging standards and implemented with different levels of fidelity and performance appropriate to the application. Grid Services can be composed in a hierarchical fashion to address complex problems. The real‐time needs of the field are addressed by high‐performance implementation of data transfer and simulation services. Further, the environment is linked to real‐time collaboration to support interactions between scientists in geographically distant locations. Copyright © 2002 John Wiley & Sons, Ltd. Geoffrey C. Fox, Sung Hoon Ko, Marlon E. Pierce, Ozgur Balsoy, Jake Kim, Sangmi Lee, Kangseok Kim, Sangyoon Oh 0001, Xi Rao, Mustafa Varank, Hasan Bulut, Gurhan Gunduz, Xiaohong Qiu, Shrideep Pallickara, Ahmet Uyar, Choon-Han Youn |
Concurr. Comput. Pract. Exp. | 15 |