VLDB 2026 Research / reviewers in the wild / expert
Selahattin Akkas
dblp:252/7008
· DBLP profile ↗
9ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0001-8121-9300ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 4 first-author · 3 since 2021Systems, architecture and hardware · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainable Graph Sparsification with Shapley ValuesabstractWe consider the problem of graph sparsification while preserving the test accuracy of Graph Neural Networks (GNNs). Prior work in this area is often motivated by the Lottery Ticket Hypothesis, which aims to prune redundant edges. However, these sparsification approaches typically operate as black boxes and provide no justification for which edges are removed. In contrast, edge importance scores obtained from GNN explanation methods provide a principled and interpretable basis for sparsification. In particular, we show that Shapley value–based explainers such as GNNShap enable effective sparsification, allowing up to 80% of edges to be removed without degrading model accuracy. We show that Shapley values are well-suited for this task due to their robustness in identifying less influential edges, resulting in sparse yet faithful subgraphs that are efficient for downstream applications. Selahattin Akkas, Ariful Azad |
WWW | 1 |
| 2026 | Scalable GNN Explanations with Distributed Shapley Values
Selahattin Akkas, Aditya Devarakonda, Ariful Azad |
Proc. VLDB Endow. | 1 |
| 2024 | GNNShap: Scalable and Accurate GNN Explanation using Shapley ValuesabstractGraph neural networks (GNNs) are popular machine learning models for graphs with many applications across scientific domains. However, GNNs are considered black box models, and it is challenging to understand how the model makes predictions. Game theoric Shapley value approaches are popular explanation methods in other domains but are not well-studied for graphs. Some studies have proposed Shapley value based GNN explanations, yet they have several limitations: they consider limited samples to approximate Shapley values; some mainly focus on small and large coalition sizes, and they are an order of magnitude slower than other explanation methods, making them inapplicable to even moderate-size graphs. In this work, we propose GNNShap, which provides explanations for edges since they provide more natural explanations for graphs and more fine-grained explanations. We overcome the limitations by sampling from all coalition sizes, parallelizing the sampling on GPUs, and speeding up model predictions by batching. GNNShap gives better fidelity scores and faster explanations than baselines on real-world datasets. The code is available at https://github.com/HipGraph/GNNShap. Selahattin Akkas, Ariful Azad |
WWW | 1 |
| 2022 | Twister2 Cross-platform resource scheduler for big dataabstractAbstract Twister2 is an open‐source big data hosting environment designed to process both batch and streaming data at scale. Twister2 runs jobs in both high‐performance computing (HPC) and big data clusters. It provides a cross‐platform resource scheduler to run jobs in diverse environments. Twister2 is designed with a layered architecture to support various clusters and big data problems. In this paper, we present the cross‐platform resource scheduler of Twister2. We identify required services and explain implementation details. We present job startup delays for single jobs and multiple concurrent jobs in Kubernetes and OpenMPI clusters. We compare job startup delays for Twister2 and Spark at a Kubernetes cluster. In addition, we compare the performance of terasort algorithm on Kubernetes and bare metal clusters at AWS cloud. Ahmet Uyar, Gurhan Gunduz, Supun Kamburugamuve, Pulasthi Wickramasinghe, Chathura Widanage, Kannan Govindarajan, Niranda Perera, Vibhatha Abeykoon, Selahattin Akkas, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 9 |
| 2022 | High-performance iterative dataflow abstractions in Twister2: TSetabstractSummary The dataflow model is gradually becoming the de facto standard for big data applications. While many popular frameworks are built around this model, very little research has been done on understanding its inner workings, which in turn has led to inefficiencies in existing frameworks. It is important to note that understanding the relationship between dataflow and high performance computing (HPC) building blocks allows us to address and alleviate many of these fundamental inefficiencies by learning from the extensive research literature in the HPC community. In this article, we present TSets, the dataflow abstraction of Twister2, which is a big data framework designed for high‐performance dataflow and iterative computations. We discuss the dataflow model adopted by TSets and the rationale behind implementing iteration handling at the worker level. Finally, we evaluate TSets to show the performance of the framework and the importance of the worker level iteration model. Pulasthi Wickramasinghe, Niranda Perera, Supun Kamburugamuve, Kannan Govindarajan, Vibhatha Abeykoon, Chathura Widanage, Ahmet Uyar, Gurhan Gunduz, Selahattin Akkas, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 9 |
| 2021 | Rank Position Forecasting in Car RacingabstractRank position forecasting in car racing is a challenging problem when using a Deep Learning-based model over time-series data. It is featured with highly complex global dependency among the racing cars, with uncertainty resulted from existing and external factors; and it is also a problem with data scarcity. Existing methods, including statistical models, machine learning regression models, and several state-of-the-art deep forecasting models all perform not well on this problem. By an elaborate analysis of pit stop events, we find it critical to decompose the cause-and-effect relationship and model the rank position and pit stop events separately. In choosing a sub-model from different neural network models, we find the model with weak assumptions on the global dependency structure performs the best. Based on these observations, we propose RankNet, a combination of the encoder-decoder network and a separate Multilayer Perception network that is capable of delivering probabilistic forecasting to model the pit stop events and rank position in car racing. Further with the help of feature optimizations, RankNet demonstrates a significant performance improvement, where MAE improves 19% in two laps forecasting task and 7% in the stint forecasting task over the best baseline and is also more stable when adapting to unseen new data. Details of the model optimizations and performance profiling are presented. It is promising to provide useful interactions of neural networks in forecasting racing cars and shine a light on solutions to similar challenging issues in general forecasting problems. Bo Peng 0011, Selahattin Akkas, Takuya Araki, Ohno Yoshiyuki, Judy Qiu |
IPDPS | 3 |
| 2019 | Streaming Machine Learning Algorithms with Big Data SystemsabstractDesigning low latency applications that can process large volumes data with higher efficiency is a challenging problem. With the limited time to process data, usage of online algorithms are becoming important in the big-data applications. Stream processing is a well-known area that has been studied for a long time. In this research, our objective is to use state of the art big-data analytic engines to implement online algorithms and compare the strengths and weaknesses in each system. We use a streaming version of Support Vector Machines (SVM) and KMeans to do the analysis. Apache Flink, Apache Storm and Twister2 streaming frameworks are used to implement these algorithms. Our study focuses on the efficiency of online training of these algorithms and the results show higher performance in Twister2 framework for these algorithms. Vibhatha Abeykoon, Gregor von Laszewski, Supun Kamburugamuve, Kannan Govindarajan, Pulasthi Wickramasinghe, Chathura Widanage, Niranda Perera, Ahmet Uyar, Gurhan Gunduz, Selahattin Akkas |
IEEE BigData | 10 |
| 2019 | A Fast Video Image Detection using TensorFlow Mobile Networks for Racing CarsabstractWith the growth of the Internet of Things, we see an increase in the importance of analysis of data from the edge, often with the results needed in real-time. Indy Car series is one of the well-known racing series in North America. All cars are equipped with multiple cameras. The video streams captured by these cameras can be used for detection and predictive tasks to increase race safety and develop better strategies to win the race. Moreover, the data can be used together with the telemetry data to provide better analysis and predictions for the drivers and the teams. In a lot of video analytics tasks, the tasks begin with object detection as its foundation. The existing pretrained object detection models are inadequate to detect IndyCar race cars. Therefore, we have created a new dataset and have compared three different Single Shot Multibox Detector models from TensorFlow Detection Model Zoo. We run experiments on CPU and GPU. Since transferring the data from edge devices to a server, running inference, and sending the result back is time and resource consuming, we also test mobile detection models on an Edge TPU, which is a Google Coral Dev Board. Our initial results show that the Edge TPU gives the best inference time, and it is more suitable for a real-time machine learning task. Selahattin Akkas, Sahaj Singh Maini, Judy Qiu |
IEEE BigData | 1 |
| 2019 | HarpGBDT: Optimizing Gradient Boosting Decision Tree for Parallel EfficiencyabstractGradient Boosting Decision Tree (GBDT) is a widely used machine learning algorithm, whose training involves both irregular computation and random memory access and is challenging for system optimizations. In this paper, we conduct a comprehensive performance analysis of two state-of-the-art systems, XGBoost and LightGBM. They represent two typical parallel implementations for GBDT; one is data parallel and the other one is parallel over features. Substantial thread synchronization overhead, as well as the inefficiency of random memory access, is identified. We propose HarpGBDT, a new GBDT system designed from the perspective of parallel efficiency optimization. Firstly, we adopt a new tree growth method that selects the top K candidates of tree nodes to enable the use of more levels of parallelism without sacrificing the algorithm's accuracy. Secondly, we organize the training data and model data in blocks and propose a block-wise approach as a general model that enables the exploration of various parallelism options. Thirdly, we propose a mixed mode to utilize the advantages of a different mode of parallelism in different phases of training. By changing the configuration of the block size and parallel mode, HarpGBDT is able to attain better parallel efficiency. By extensive experiments on four datasets with different statistical characteristics on the Intel(R) Xeon(R) E5-2699 server, HarpGBDT on average performs 8x faster than XGBoost and 2.6x faster than LightGBM. Bo Peng 0011, Judy Qiu, Langshi Chen, Selahattin Akkas, Egor Smirnov, Ruslan Israfilov, Sergey Khekhnev, Andrey Nikolaev |
CLUSTER | 6 |