EDBT 2026 Demo / reviewers in the wild / expert
Buse Yilmaz
dblp:37/8981
· DBLP profile ↗
8ranked-venue papers
4as first author
2since 2021 · last 2023
0000-0001-5529-7188ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 70% GPUs and heterogeneous computing · 23% Parallel and multicore computing · 7% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
CPU-GPU heterogeneous computing |
0.5 | 1 | 2021 | A Split Execution Model for SpTRSV · IEEE Trans. Parallel Distributed Syst. 2021 |
High-performance computing
sparse linear algebra |
0.5 | 1 | 2021 | A Split Execution Model for SpTRSV · IEEE Trans. Parallel Distributed Syst. 2021 |
High-performance computing › sparse linear solver
sparse triangular solve |
0.5 | 1 | 2021 | A Split Execution Model for SpTRSV · IEEE Trans. Parallel Distributed Syst. 2021 |
Compilers and program optimization
autotuning |
0.2 | 1 | 2016 | Autotuning Runtime Specialization for Sparse Matrix-Vector Multiplication · ACM Trans. Archit. Code Optim. 2016 |
Compilers and program optimization › program specialization
runtime specialization |
0.2 | 1 | 2016 | Autotuning Runtime Specialization for Sparse Matrix-Vector Multiplication · ACM Trans. Archit. Code Optim. 2016 |
High-performance computing › sparse linear algebra
sparse matrix computation |
0.2 | 1 | 2016 | Autotuning Runtime Specialization for Sparse Matrix-Vector Multiplication · ACM Trans. Archit. Code Optim. 2016 |
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication |
0.2 | 1 | 2016 | Autotuning Runtime Specialization for Sparse Matrix-Vector Multiplication · ACM Trans. Archit. Code Optim. 2016 |
Methods — techniques the papers use, named apart from their topics
heuristics-based split-point selection · 0.5feature-based prediction · 0.5code generation · 0.5auto-tuning · 0.5DAG analysis · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A novel graph transformation strategy for optimizing SpTRSV on CPUsabstractSummary Sparse triangular solve (SpTRSV) is an extensively studied computational kernel. An important obstacle in parallel SpTRSV implementations is that in some parts of a sparse matrix the computation is serial. By transforming the dependency graph, it is possible to increase the parallelism of the parts that lack it. In this work, we present a novel graph transformation strategy to increase the parallelism degree of a sparse matrix and compare it to our previous strategy. It is seen that our transformation strategy can provide a speedup as high as . Buse Yilmaz |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | A Split Execution Model for SpTRSVabstractSparse Triangular Solve (SpTRSV) is an important and extensively used kernel in scientific computing. Parallelism within SpTRSV depends upon matrix sparsity pattern and, in many cases, is non-uniform from one computational step to the next. In cases where the SpTRSV computational steps have contrasting parallelism characteristics- some steps are more parallel, others more sequential in nature, the performance of an SpTRSV algorithm may be limited by the contrasting parallelism characteristics. In this work, we propose a split-execution model for SpTRSV to automatically divide SpTRSV computation into two sub-SpTRSV systems and an SpMV, such that one of the sub-SpTRSVs has more parallelism than the other. Each sub-SpTRSV is then computed using different SpTRSV algorithms, which are possibly executed on different platforms (CPU or GPU). By analyzing the SpTRSV Directed Acyclic Graph (DAG) and matrix sparsity features, we use a heuristics-based approach to (i) automatically determine the suitability of an SpTRSV for split-execution, (ii) find the appropriate split-point, and (iii) execute SpTRSV in a split fashion using two SpTRSV algorithms while managing any required inter-platform communication. Experimental evaluation of the execution model on two CPU-GPU machines with a matrix dataset of 327 matrices from the SuiteSparse Matrix Collection shows that our approach correctly selects the fastest SpTRSV method (split or unsplit) for 88 percent of matrices on the Intel Xeon Gold (6148) + NVIDIA Tesla V100 and 83 percent on the Intel Core I7 + NVIDIA G1080 Ti platform achieving speedups up to 10x and 6.36x respectively. Najeeb Ahmad, Buse Yilmaz, Didem Unat |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | A Prediction Framework for Fast Sparse Triangular Solves
Najeeb Ahmad, Buse Yilmaz, Didem Unat |
Euro-Par | 2 |
| 2020 | Adaptive Level Binning: A New Algorithm for Solving Sparse Triangular SystemsabstractSparse triangular solve (SpTRSV) is an important scientific kernel used in several applications such as preconditioners for Krylov methods. Parallelizing SpTRSV on multi-core systems is challenging since it exhibits limited parallelism due to computational dependencies and introduces high parallelization overhead due to finegrained and unbalanced nature of workloads. We propose a novel method, named Adaptive Level Binning (ALB), that addresses these challenges by eliminating redundant synchronization points and adapting the work granularity with an efficient load balancing strategy. Similar to the commonly used level-set methods for solving SpTRSV, ALB constructs level-sets of rows, where each level can be computed in parallel. Differently, ALB bins rows to levels adaptively and reduces redundant dependencies between rows. On an Intel® Xeon® Gold 6148 processor and NVIDIA® Tesla V100 GPU, ALB obtains 1.83x speedup on average and up to 5.28x speedup over Intel MKL and, over NVIDIA cuSPARSE, an average speedup of 2.80x and a maximum speedup of 39.40x for 29 matrices selected from Suite Sparse Matrix Collection. Buse Yilmaz, Bugrra Sipahiogrlu, Najeeb Ahmad, Didem Unat |
HPC Asia | 1 |
| 2016 | Autotuning Runtime Specialization for Sparse Matrix-Vector MultiplicationabstractRuntime specialization is used for optimizing programs based on partial information available only at runtime. In this paper we apply autotuning on runtime specialization of Sparse Matrix-Vector Multiplication to predict a best specialization method among several. In 91% to 96% of the predictions, either the best or the second-best method is chosen. Predictions achieve average speedups that are very close to the speedups achievable when only the best methods are used. By using an efficient code generator and a carefully designed set of matrix features, we show the runtime costs can be amortized to bring performance benefits for many real-world cases. Buse Yilmaz, Baris Aktemur, María Jesús Garzarán, Samuel N. Kamin, Furkan Kiraç |
ACM Trans. Archit. Code Optim. | 1 |
| 2014 | Optimization by runtime specialization for sparse matrix-vector multiplicationabstractRuntime specialization optimizes programs based on partial information available only at run time. It is applicable when some input data is used repeatedly while other input data varies. This technique has the potential of generating highly efficient codes. In this paper, we explore the potential for obtaining speedups for sparse matrix-dense vector multiplication using runtime specialization, in the case where a single matrix is to be multiplied by many vectors. We experiment with five methods involving runtime specialization, comparing them to methods that do not (including Intel's MKL library). For this work, our focus is the evaluation of the speedups that can be obtained with runtime specialization without considering the overheads of the code generation. Our experiments use 23 matrices from the Matrix Market and Florida collections, and run on five different machines. In 94 of those 115 cases, the specialized code runs faster than any version without specialization. If we only use specialization, the average speedup with respect to Intel's MKL library ranges from 1.44x to 1.77x, depending on the machine. We have also found that the best method depends on the matrix and machine; no method is best for all matrices and machines. Samuel N. Kamin, María Jesús Garzarán, Baris Aktemur, Danqing Xu, Buse Yilmaz, Zhongbo Chen |
GPCE | 5 |
| 2011 | Hybrid local search algorithms on Graph Coloring ProblemabstractHybridization of local search algorithms yield promising algorithms for combinatorial optimization problems such as Graph Coloring Problem (GCP). This paper presents a new meta-heuristic Simulated Annealing with Backtracking (SABT) and shows the effect of hill climber and tabu search on SABT for solving GCP. The algorithm proposed merges the power of simulated annealing approach and backtracking mechanism. Some hill climbers are integrated for fine tuning and also tabu search is integrated for avoiding from redundant search. Several tests are run on a collection of benchmarks from DIMACS challenge suite and promising results are obtained. A comparison of SABT framework with some other state-of-the-art algorithms is presented along with an analysis of the performance of the algorithm. Cagri Yesil, Buse Yilmaz, Emin Erkan Korkmaz |
HIS | 2 |
| 2010 | Representation issue in graph coloringabstractLinear Linkage Encoding (LLE) is a powerful encoding scheme utilized when genetic algorithms (GAs) are applied to grouping problems. It discards the redundancy of other traditional encoding schemes. However, some genetic operators are quite costly in terms of computational time when LLE is utilized. In this study, two supplementary encoding schemes Linear Linkage Encoding with Ending Node Links (LLE-e) and Linear Linkage Encoding with Backward Links (LLE-b) are designed and used together with LLE. When a genetic operator is costly in LLE, the operation is carried out on one of the supplementary encodings and the result is reflected back to LLE. The algorithm is implemented in a Multi Objective Genetic Algorithm (MOGA) framework and various tests are carried out on Graph coloring Problem (GCP) instances obtained from DIMACS Challenge Suite. A performance improvement has been obtained when both supplementary encoding schemes are included in the algorithm. Buse Yilmaz, Emin Erkan Korkmaz |
ISDA | 1 |