Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Orhan Kislal

dblp:61/11411 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 40% Memory systems · 35% Interconnection networks and networks-on-chip · 17%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning and data management
in-database machine learning
0.512021
Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches · Proc. VLDB Endow. 2021
Parallel and multicore computing › task allocation
computation-to-core mapping
0.312018
Enhancing computation-to-core assignment with physical location information · PLDI 2018
Compilers and program optimization › memory optimization
data movement optimization
0.312017
Data movement aware computation partitioning · MICRO 2017
Compilers and program optimization
parallelizing compiler
0.312017
Data movement aware computation partitioning · MICRO 2017
Memory systems › processing-in-memory
near-data processing
0.312017
Data movement aware computation partitioning · MICRO 2017
Parallel and multicore computing › parallel algorithms › parallel algorithm design
parallel algorithm mapping
0.312017
Data movement aware computation partitioning · MICRO 2017
Memory systems
processing-in-memory
0.312017
Data movement aware computation partitioning · MICRO 2017
Cloud and datacenter computing
cluster resource management and scheduling
0.112021
Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches · Proc. VLDB Endow. 2021
Parallel and multicore computing › parallel computation models
parallel execution models
0.112021
Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches · Proc. VLDB Endow. 2021
Memory systems › memory hierarchy › cache hierarchy
last-level cache
0.112018
Enhancing computation-to-core assignment with physical location information · PLDI 2018

Methods — techniques the papers use, named apart from their topics

model hopper parallelism · 1.0empirical benchmarking · 1.0analytical comparison · 1.0loop nest partitioning · 0.6compiler scheduling · 0.6compiler-guided mapping · 0.3
YearPublicationVenuePosition
2021 Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches
abstract
Deep learning (DL) is growing in popularity for many data analytics applications, including among enterprises. Large business-critical datasets in such settings typically reside in RDBMSs or other data systems. The DB community has long aimed to bring machine learning (ML) to DBMS-resident data. Given past lessons from in-DBMS ML and recent advances in scalable DL systems, DBMS and cloud vendors are increasingly interested in adding more DL support for DB-resident data. Recently, a new parallel DL model selection execution approach called Model Hopper Parallelism (MOP) was proposed. In this paper, we characterize the particular suitability of MOP for DL on data systems, but to bring MOP-based DL to DB-resident data, we show that there is no single "best" approach, and an interesting tradeoff space of approaches exists. We explain four canonical approaches and build prototypes upon Greenplum Database, compare them analytically on multiple criteria (e.g., runtime efficiency and ease of governance) and compare them empirically with large-scale DL workloads. Our experiments and analyses show that it is non-trivial to meet all practical desiderata well and there is a Pareto frontier; for instance, some approaches are 3x-6x faster but fare worse on governance and portability. Our results and insights can help DBMS and cloud vendors design better DL support for DB users. All of our source code, data, and other artifacts are available at https://github.com/makemebitter/cerebro-ds.
Frank Mcquillan, Nandish Jayaram, Nikhil Kak, Ekta Khanna, Orhan Kislal, Domino Valdano, Arun Kumar 0001
Proc. VLDB Endow.6
2018 Quantifying and Optimizing Data Access Parallelism on Manycores
abstract
The following topics are dealt with: storage management; cache storage; pattern clustering; cloud computing; optimisation; flash memories; resource allocation; scheduling; parallel processing; data mining.
Jihyun Ryoo, Orhan Kislal, Xulong Tang, Mahmut T. Kandemir
MASCOTS2
2018 Enhancing computation-to-core assignment with physical location information
abstract
Going beyond a certain number of cores in modern architectures requires an on-chip network more scalable than conventional buses. However, employing an on-chip network in a manycore system (to improve scalability) makes the latencies of the data accesses issued by a core non-uniform. This non-uniformity can play a significant role in shaping the overall application performance. This work presents a novel compiler strategy which involves exposing architecture information to the compiler to enable an optimized computation-to-core mapping. Specifically, we propose a compiler-guided scheme that takes into account the relative positions of (and distances between) cores, last-level caches (LLCs) and memory controllers (MCs) in a manycore system, and generates a mapping of computations to cores with the goal of minimizing the on-chip network traffic. The experimental data collected using a set of 21 multi-threaded applications reveal that, on an average, our approach reduces the on-chip network latency in a 6×6 manycore system by 38.4% in the case of private LLCs, and 43.8% in the case of shared LLCs. These improvements translate to the corresponding execution time improvements of 10.9% and 12.7% for the private LLC and shared LLC based systems, respectively.
Orhan Kislal, Jagadish Kotra, Xulong Tang, Mahmut T. Kandemir, Myoungsoo Jung
PLDI1
2018 Data access skipping for recursive partitioning methods
Orhan Kislal, Mahmut T. Kandemir
Comput. Lang. Syst. Struct.1
2017 POSTER: Location-Aware Computation Mapping for Manycore Processors
abstract
Employing an on-chip network in a manycore system (to improve scalability) makes the latencies of data accesses issued by a core non-uniform, which significant impact application performance. This paper presents a compiler strategy which involves exposing architecture information to the compiler to enable optimized computation-to-core mapping. Our scheme takes into account the relative positions of (and distances between) cores, last-level caches (LLCs) and memory controllers (MCs) in a manycore system, and generates a mapping of computations to cores with the goal of minimizing the on-chip network traffic. Our experiments of 12 multi-threaded applications reveal that, on average, our approach reduces the on-chip network latency in a 6x6 manycore system by 49.5% in the case of private LLCs and 52.7% in the case of shared LLCs. These improvements translate to the corresponding execution time improvements of 14.8% and 15.2% for the private LLC and shared LLC based systems.
Orhan Kislal, Jagadish Kotra, Xulong Tang, Mahmut T. Kandemir, Myoungsoo Jung
PACT1
2017 Data movement aware computation partitioning
abstract
Data access costs dominate the execution times of most parallel applications and they are expected to be even more important in the future. To address this, recent research has focused on Near Data Processing (NDP) as a new paradigm that tries to bring computation to data, instead of bringing data to computation (which is the norm in conventional computing). This paper explores the potential of compiler support in exploiting NDP in the context of emerging manycore systems. To that end, we propose a novel compiler algorithm that partitions the computations in a given loop nest into subcomputations and schedules the resulting subcomputations on different cores with the goal of reducing the distance-to-data on the on-chip network. An important characteristic of our approach is that it exploits NDP while taking advantage of data locality. Our experiments with 12 multithreaded applications running on a state-of-the-art commercial manycore system indicate that the proposed compiler-based approach significantly reduces data movements on the on-chip network by taking advantage of NDP, and these benefits lead to an average execution time improvement of 18.4%.
Xulong Tang, Orhan Kislal, Mahmut T. Kandemir, Mustafa Karaköy
MICRO2