Gunduz Vehbi Demirci

dblp:72/10543 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 2 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 66% High-performance computing · 14% Electronic design automation · 14%
Artificial intelligence
1 paper
Graph learning · 50% Efficient and distributed learning · 50%
Databases, data mining, and information retrieval
1 paper
Graph data management · 50% Distributed and cloud data management · 50%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
0.612022
Scalable Graph Convolutional Network Training on Distributed-Memory Systems · Proc. VLDB Endow. 2022
Machine learning › Graph learning › graph neural network
graph convolutional network
0.612022
Scalable Graph Convolutional Network Training on Distributed-Memory Systems · Proc. VLDB Endow. 2022
Parallel and multicore computing › parallel algorithms
distributed-memory parallel algorithms
0.612022
Scalable Graph Convolutional Network Training on Distributed-Memory Systems · Proc. VLDB Endow. 2022
Parallel and multicore computing
graph partitioning
0.612022
Scalable Graph Convolutional Network Training on Distributed-Memory Systems · Proc. VLDB Endow. 2022
Parallel and multicore computing › data distribution
cartesian partitioning
0.412020
Cartesian Partitioning Models for 2D and 3D Parallel SpGEMM Algorithms · IEEE Trans. Parallel Distributed Syst. 2020
Parallel and multicore computing › parallelization strategies
distributed-memory parallelization
0.412020
Cartesian Partitioning Models for 2D and 3D Parallel SpGEMM Algorithms · IEEE Trans. Parallel Distributed Syst. 2020
Electronic design automation › physical design › circuit partitioning
hypergraph partitioning
0.412020
Cartesian Partitioning Models for 2D and 3D Parallel SpGEMM Algorithms · IEEE Trans. Parallel Distributed Syst. 2020
High-performance computing › sparse linear algebra
sparse matrix multiplication
0.412020
Cartesian Partitioning Models for 2D and 3D Parallel SpGEMM Algorithms · IEEE Trans. Parallel Distributed Syst. 2020
Graph data management
graph partitioning
0.412019
Cascade-aware partitioning of large graph databases · VLDB J. 2019
Distributed systems
communication optimization
0.212022
Scalable Graph Convolutional Network Training on Distributed-Memory Systems · Proc. VLDB Endow. 2022

Methods — techniques the papers use, named apart from their topics

hypergraph partitioning · 1.6non-blocking point-to-point communication · 1.1randomized estimation · 0.4
YearPublicationVenuePosition
2024 Foresight plus: serverless spatio-temporal traffic forecasting
abstract
Abstract Building a real-time spatio-temporal forecasting system is a challenging problem with many practical applications such as traffic and road network management. Most forecasting research focuses on achieving (often marginal) improvements in evaluation metrics such as MAE/MAPE on static benchmark datasets, with less attention paid to building practical pipelines which achieve timely and accurate forecasts when the network is under heavy load. Transport authorities also need to leverage dynamic data sources such as roadworks and vehicle-level flow data, while also supporting ad-hoc inference workloads at low cost. Our cloud-based forecasting solution Foresight, developed in collaboration with Transport for the West Midlands (TfWM), is able to ingest, aggregate and process streamed traffic data, enhanced with dynamic vehicle-level flow and urban event information, to produce regularly scheduled forecasts with high accuracy. In this work, we extend Foresight with several novel enhancements, into a new system which we term Foresight Plus. New features include an efficient method for extending the forecasting scale, enabling predictions further into the future. We also augment the inference architecture with a new, fully serverless design which offers a more cost-effective solution and which seamlessly handles sporadic inference workloads over multiple forecasting scales. We observe that Graph Neural Network (GNN) forecasting models are robust to extensions of the forecasting scale, achieving consistent performance up to 48 hours ahead. This is in contrast to the 1 hour forecasting periods popularly considered in this context. Further, our serverless inference solution is shown to be more cost-effective than provisioned alternatives in corresponding use-cases. We identify the optimal memory configuration of serverless resources to achieve an attractive cost-to-performance ratio.
Joe Oakley 0001, Chris Conlan, Gunduz Vehbi Demirci, Alexandros Sfyridis, Hakan Ferhatosmanoglu
GeoInformatica3
2022 Scalable Graph Convolutional Network Training on Distributed-Memory Systems
abstract
Graph Convolutional Networks (GCNs) are extensively utilized for deep learning on graphs. The large data sizes of graphs and their vertex features make scalable training algorithms and distributed memory systems necessary. Since the convolution operation on graphs induces irregular memory access patterns, designing a memory- and communication-efficient parallel algorithm for GCN training poses unique challenges. We propose a highly parallel training algorithm that scales to large processor counts. In our solution, the large adjacency and vertex-feature matrices are partitioned among processors. We exploit the vertex-partitioning of the graph to use non-blocking point-to-point communication operations between processors for better scalability. To further minimize the parallelization overheads, we introduce a sparse matrix partitioning scheme based on a hypergraph partitioning model for full-batch training. We also propose a novel stochastic hypergraph model to encode the expected communication volume in mini-batch training. We show the merits of the hypergraph model, previously unexplored for GCN training, over the standard graph partitioning model which does not accurately encode the communication costs. Experiments performed on real-world graph datasets demonstrate that the proposed algorithms achieve considerable speedups over alternative solutions. The optimizations achieved on communication costs become even more pronounced at high scalability with many processors. The performance benefits are preserved in deeper GCNs having more layers as well as on billion-scale graphs.
Gunduz Vehbi Demirci, Aparajita Haldar, Hakan Ferhatosmanoglu
Proc. VLDB Endow.1
2021 Partitioning sparse deep neural networks for scalable training and inference
abstract
The state-of-the-art deep neural networks (DNNs) have significant computational and data management requirements. The size of both training data and models continue to increase. Sparsification and pruning methods are shown to be effective in removing a large fraction of connections in DNNs. The resulting sparse networks present unique challenges to further improve the computational efficiency of training and inference in deep learning. Both the feedforward (inference) and backpropagation steps in stochastic gradient descent (SGD) algorithm for training sparse DNNs involve consecutive sparse matrix-vector multiplications (SpMVs). We first introduce a distributed-memory parallel SpMV-based solution for the SGD algorithm to improve its scalability. The parallelization approach is based on row-wise partitioning of weight matrices that represent neuron connections between consecutive layers. We then propose a novel hypergraph model for partitioning weight matrices to reduce the total communication volume and ensure computational load-balance among processors. Experiments performed on sparse DNNs demonstrate that the proposed solution is highly efficient and scalable. By utilizing the proposed matrix partitioning scheme, the performance of our solution is further improved significantly.
Gunduz Vehbi Demirci, Hakan Ferhatosmanoglu
ICS1
2020 Scaling sparse matrix-matrix multiplication in the accumulo database
Gunduz Vehbi Demirci, Cevdet Aykanat
Distributed Parallel Databases1
2020 Cartesian Partitioning Models for 2D and 3D Parallel SpGEMM Algorithms
abstract
The focus is distributed-memory parallelization of sparse-general-matrix-multiplication (SpGEMM). Parallel SpGEMM algorithms are classified under one-dimensional (1D), 2D, and 3D categories denoting the number of dimensions by which the 3D sparse workcube representing the iteration space of SpGEMM is partitioned. Recently proposed successful 2D- and 3D-parallel SpGEMM algorithms benefit from upper bounds on communication overheads enforced by 2D and 3D cartesian partitioning of the workcube on 2D and 3D virtual processor grids, respectively. However, these methods are based on random cartesian partitioning and do not utilize sparsity patterns of SpGEMM instances for reducing the communication overheads. We propose hypergraph models for 2D and 3D cartesian partitioning of the workcube for further reducing the communication overheads of these 2D- and 3D- parallel SpGEMM algorithms. The proposed models utilize two- and three-phase partitioning that exploit multi-constraint hypergraph partitioning formulations. Extensive experimentation performed on 20 SpGEMM instances by using upto 900 processors demonstrate that proposed partitioning models significantly improve the scalability of 2D and 3D algorithms. For example, in 2D-parallel SpGEMM algorithm on 900 processors, the proposed partitioning model respectively achieves 85 and 42 percent decrease in total volume and total number of messages, leading to 1.63 times higher speedup compared to random partitioning, on average.
Gunduz Vehbi Demirci, Cevdet Aykanat
IEEE Trans. Parallel Distributed Syst.1
2019 Locality-aware and load-balanced static task scheduling for MapReduce
Oguz Selvitopi, Gunduz Vehbi Demirci, Ata Turk, Cevdet Aykanat
Future Gener. Comput. Syst.2
2019 Cascade-aware partitioning of large graph databases
abstract
Graph partitioning is an essential task for scalable data management and analysis. The current partitioning methods utilize the structure of the graph, and the query log if available. Some queries performed on the database may trigger further operations. For example, the query workload of a social network application may contain re-sharing operations in the form of cascades. It is beneficial to include the potential cascades in the graph partitioning objectives. In this paper, we introduce the problem of cascade-aware graph partitioning that aims to minimize the overall cost of communication among parts/servers during cascade processes. We develop a randomized solution that estimates the underlying cascades, and use it as an input for partitioning of large-scale graphs. Experiments on 17 real social networks demonstrate the effectiveness of the proposed solution in terms of the partitioning objectives.
Gunduz Vehbi Demirci, Hakan Ferhatosmanoglu, Cevdet Aykanat
VLDB J.1