EDBT 2026 Demo / reviewers in the wild / expert
Mehmet Esat Belviranli
dblp:98/9423 · also Mehmet E. Belviranli
· DBLP profile ↗
22ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0001-9434-9833ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Closing the Spatial Execution Gap in Digital Whiteboards via Verifiable Reinforcement LearningabstractWhile multi-modal large language models such as GPT-5 demonstrate exceptional general understanding, they suffer from a fundamental Spatial Execution Gap, failing to translate visual semantics into precise, schema-valid coordinate operations in interactive environments.In this work, we show that model scale alone cannot close this gap; instead, verifiable structured reasoning provides the key to spatial precision.We present a comprehensive pipeline that leverages Group Relative Policy Optimization to enforce a strict Identify-Reason-Verify protocol, effectively shifting the computational burden from parameters to test-time reasoning.By utilizing a multi-agent system to distill optimal reasoning schemas and training on execution-verifiable rewards, our specialized 3B agent achieves 100% format coherence and 81.12% operation accuracy on digital whiteboard tasks.Crucially, our approach outperforms a state-of-the-art frontier model, GPT-5, by 16.75% in operation accuracy.The results suggest that for complex user interface manipulation, small, RL-aligned models with dedicated reasoning protocols are superior to generalist frontier models, offering a promising direction for building reliable web agents. Chang Liu 0122, Benjamin Wagley, Mehmet Esat Belviranli, Bo Wu 0002 |
ACL (1) | 4 |
| 2026 | Hibiscus: End-to-end Architectural Simulation Framework for Hybrid SFQ/CMOS-Memory Compute SystemsabstractAs conventional CMOS technology approaches power and performance limits, superconducting single flux quantum (SFQ) logic offers a path to high-speed, energy-efficient computing. However, SFQ circuits require cryogenic temperatures, introducing complex challenges in memory integration and data movement between thermal zones. This paper presents an end-to-end simulation framework for hybrid SFQ/CMOS-memory systems that accurately models processor, memory, and interconnect behavior across cryogenic $(4 \mathrm{~K}, 77 \mathrm{~K})$ and room temperatures $(300 \mathrm{~K})$. The framework integrates gate-level pipelined Rapid SFQ (RSFQ) RISC-V processors, temperature-aware CryoMEM memory models, and physically grounded interconnect latency models. The simulator facilitates cross-layer design space exploration across diverse parameters such as cache placement, interconnect stack selection, and granularity. These features allow the community to identify technological gaps and re-evaluate the bottlenecks in memory-compute throughput. Our evaluations highlight the critical interplay between processor frequency and memory bandwidth, demonstrate the speedup potential of 4K SFQ caches, and quantify the impact of cryostat cabling choices on system performance. Ryan Marsala, Yerzhan Mustafa, Prabhath Tangella, Mohammad Sonji, George Michelogiannakis, Selçuk Köse, Adwait Jog, Mehmet Esat Belviranli |
ISPASS | 8 |
| 2025 | ${MC}^{3}$: Memory Contention-Based Covert Channel Communication on Shared DRAM System-on-ChipsabstractShared memory system-on-chips (SM-SoCs) are ubiquitously employed by a wide range of computing platforms, including edge/IoT devices, autonomous systems, and smartphones. In SM-SoCs, system-wide shared memory enables a convenient and cost-effective mechanism for making data accessible across dozens of processing units (PUs), such as CPU cores and domain-specific accelerators. Due to the diverse computational characteristics of the PUs they embed, SM-SoCs often do not employ a shared last-level cache (LLC). Although covert channel attacks have been widely studied in shared memory systems, high-throughput communication has previously been feasible only by relying on an LLC or by possessing privileged or physical access to the shared memory subsystem. In this study, we introduce a new memory-contention-based covert communication attack,$\boldsymbol{MC}^{3}$, which specifically targets shared system memory in mobile SoCs. Unlike existing attacks, our approach achieves high-throughput communication without the need for an LLC or elevated access to the system. We explore the effectiveness of our methodology by demonstrating the tradeoff between the channel transmission rate and the robustness of the communication. We evaluate$\boldsymbol{MC}^{3}$on NVIDIA Orin AGX, NX, and Nano platforms and achieve transmission rates up to 6.4 Kbps with less than 1% error rate. Ismet Dagli, James Crea, Soner Seçkiner, Yuanchao Xu 0001, Selçuk Köse, Mehmet Esat Belviranli |
DATE | 6 |
| 2025 | HARNESS: Holistic Resource Management for Diversely Scaled Edge Cloud SystemsabstractComputing systems are evolving to be more ubiquitous, heterogeneous, and dynamic.Many emerging domains, such as Internet of Things (IoT), federated learning, and smart buildings, rely on a diverse edge-to-cloud continuum where the execution of applications spans various tiers of systems with significantly different computational capabilities.Computing resources in each tier, such as processing units inside of in-the-field edge devices and high-performance servers in datacenters, are handled in isolation due to scalability and resource segregation.This practice results in task mappings limited to only a subset of all available processing units, preventing an efficient overall utilization of the system.In this paper, we propose a holistic approach to capture diverse computational characteristics of edge-cloud systems with arbitrary topologies and to efficiently manage computational resources with the whole continuum in the scope.Our approach is built upon a multi-layer graph-based hardware (HW) representation and a modular performance modeling interface that can capture interactions and interference between computational resources in the system.We introduce an orchestrator mechanism that leverages the graph-based HW representation to hierarchically locate processing units to which a given set of tasks can be mapped while respecting the isolation between the computational tiers of an edge-cloud system.We demonstrate the utility of our approach on two distinct edge-cloud systems deployed in the field, improving the latency up to 47% over the best baseline with less than 2% scheduling overhead and reducing the average prediction error rate from 27.4% to 3.2%. Ismet Dagli, Justin Davis, Mehmet Esat Belviranli |
ICS | 3 |
| 2024 | Context-aware Multi-Model Object Detection for Diversely Heterogeneous Compute SystemsabstractIn recent years, deep neural networks (DNNs) have gained widespread adoption for continuous mobile object detection (OD) tasks, particularly in autonomous systems. However, a prevalent issue in their deployment is the one-size-fits-all approach, where a single DNN is used, resulting in inefficient utilization of computational resources. This inefficiency is particularly detrimental in energy-constrained systems, as it degrades overall system efficiency. We identify that, the contextual information embedded in the input data stream (e.g., the frames in the camera feed that the OD models are run on) could be exploited to allow a more efficient multi-model-based OD process. In this paper, we propose SHIFT which continuously selects from a variety of DNN-based OD models depending on the dynamically changing contextual information and computational constraints. During this selection, SHIFT uniquely considers multi-accelerator execution to better optimize the energy-efficiency while satisfying the latency constraints. Our proposed methodology results in improvements of up to 7.5x in energy usage and 2.8x in latency compared to state-of-the-art GPU-based single model OD approaches. Justin Davis, Mehmet Esat Belviranli |
DATE | 2 |
| 2024 | Constraint-Aware Resource Management for Cyber-Physical SystemsabstractCyber-physical systems (CPS) such as robots and self-driving cars pose strict physical requirements to avoid failure. Scheduling choices impact these requirements. This presents a challenge: how do we find efficient schedules for CPS with heterogeneous processing units, such that the schedules are resource-bounded to meet the physical requirements? We propose the creation of a structured system, the Constrained Autonomous Workload Scheduler, which determines scheduling decisions with direct relations to the environment. By using a representation language (AuWL), Timed Petri nets, and mixed-integer linear programming, our scheme offers novel capabilities to represent and schedule many types of CPS workloads, real world constraints, and optimization criteria. Justin McGowen, Ismet Dagli, Neil Dantam, Mehmet Esat Belviranli |
DATE | 4 |
| 2024 | Scheduling for Cyber-Physical Systems with Heterogeneous Processing Units under Real-World ConstraintsabstractCyber-physical systems (CPS) such as robots and self-driving cars pose strict physical requirements to avoid failure. The scheduling choices impact these requirements. This presents a challenge: How do we find efficient schedules for CPS with heterogeneous processing units, such that the schedules are resource-bounded to meet the physical requirements? For example, tasks that require significant computation time in a self-driving car can delay reaction, decreasing available braking time. Heterogeneous computing systems — containing CPUs, GPUs, and other types of domain-specific accelerators — offer effective capabilities to reduce computation time or energy consumption and expand such operating conditions. However, doing so under physical requirements presents several challenges that existing scheduling solutions fail to address. Justin McGowen, Ismet Dagli, Neil Dantam, Mehmet Esat Belviranli |
ICS | 4 |
| 2024 | Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-ChipsabstractTwo distinguishing features of state-of-the-art mobile and autonomous systems are: 1) There are often multiple workloads, mainly deep neural network (DNN) inference, running concurrently and continuously. 2) They operate on shared memory System-on-Chips (SoC) that embed heterogeneous accelerators tailored for specific operations. State-of-the-art systems lack efficient performance and resource management techniques necessary to either maximize total system throughput or minimize end-to-end workload latency. In this work, we propose HaX-CoNN, a novel scheme that characterizes and maps layers in concurrently executing DNN inference workloads to a diverse set of accelerators within an SoC. Our scheme uniquely takes per-layer execution characteristics, shared memory (SM) contention, and inter-accelerator transitions into account to find optimal schedules. We evaluate HaX-CoNN on NVIDIA Orin, NVIDIA Xavier, and Qualcomm Snapdragon 865 SoCs. Our experimental results indicate that HaX-CoNN can minimize memory contention by up to 45% and improve total latency and throughput by up to 32% and 29%, respectively, compared to the state-of-the-art. Ismet Dagli, Mehmet Esat Belviranli |
PPoPP | 2 |
| 2022 | Optimizing Regular Expressions via Rewrite-Guided SynthesisabstractRegular expressions are pervasive in modern systems. Many real-world regular expressions are inefficient, sometimes to the extent that they are vulnerable to complexity-based attacks, and while much research has focused on detecting inefficient regular expressions or accelerating regular expression matching at the hardware level, we investigate automatically transforming regular expressions to remove inefficiencies. We reduce this problem to general expression optimization, an important task necessary in a variety of domains even beyond compilers, e.g., digital logic design, etc. Syntax-guided synthesis (SyGuS) with a cost function can be used for this purpose, but ordered enumeration through a large space of candidate expressions can be prohibitively expensive. Equality saturation is an alternative approach which allows efficient construction and maintenance of expression equivalence classes generated by rewrite rules, but the procedure may not reach saturation, meaning global minimality cannot be confirmed. We present a new approach called rewrite-guided synthesis (ReGiS), in which a unique interplay between SyGuS and equality saturation-based rewriting helps to overcome these problems, resulting in an efficient, scalable framework for expression optimization. Jedidiah McClurg, Miles Claver, Jackson Garner, Jake Vossen, Jordan Schmerge, Mehmet Esat Belviranli |
PACT | 6 |
| 2022 | AxoNN: energy-aware execution of neural network inference on multi-accelerator heterogeneous SoCsabstractThe energy and latency demands of critical workload execution, such as object detection, in embedded systems vary based on the physical system state and other external factors. Many recent mobile and autonomous System-on-Chips (SoC) embed a diverse range of accelerators with unique power and performance characteristics. The execution flow of the critical workloads can be adjusted to span into multiple accelerators so that the trade-off between performance and energy fits to the dynamically changing physical factors. Ismet Dagli, Alexander Cieslewicz, Jedidiah McClurg, Mehmet Esat Belviranli |
DAC | 4 |
| 2021 | PCCS: Processor-Centric Contention-aware Slowdown Model for Heterogeneous System-on-ChipsabstractMany slowdown models have been proposed to characterize memory interference of workloads co-running on heterogeneous System-on-Chips (SoCs). But they are mostly for post-silicon usage. How to effectively consider memory interference in the SoC design stage remains an open problem. This paper presents a new approach to this problem, consisting of a novel processor-centric slowdown modeling methodology and a new three-region interference-conscious slowdown model. The modeling process needs no measurement of co-running of various combinations of applications, but the produced slowdown models can be used to estimate the co-run slowdowns of arbitrary workloads on various SoC designs that embed a newer generation of accelerators, such as deep learning accelerators (DLA), in addition to CPUs and GPUs. The new method reduces average prediction errors of the state-of-art model from 30.3% to 8.7% on GPU, from 13.4% to 3.7% on CPU, from 20.6% to 5.6% on DLA and demonstrates much improved efficacy in guiding SoC designs. Yuanchao Xu 0001, Mehmet Esat Belviranli, Xipeng Shen, Jeffrey S. Vetter |
MICRO | 2 |
| 2021 | A computational-graph partitioning method for training memory-constrained DNNs
Fareed Qararyah, Mohamed Wahib, Doga Dikbayir, Mehmet Esat Belviranli, Didem Unat |
Parallel Comput. | 4 |
| 2020 | MEPHESTO: Modeling Energy-Performance in Heterogeneous SoCs and Their Trade-OffsabstractIntegrated shared memory heterogeneous architectures are pervasive because they satisfy the diverse needs of mobile, autonomous, and edge computing platforms. Although specialized processing units (PUs) that share a unified system memory improve performance and energy efficiency by reducing data movement, they also increase contention for this memory since the PUs interact with each other. Prior work has investigated performance degradation due to memory contention, but few have studied the relationship of power and energy to memory contention. Moreover, a comprehensive solution that models memory contention for kernel placement on contemporary heterogeneous systems on chip (SoCs) in response to energy and performance has been largely unaddressed. Mohammad Alaul Haque Monil, Mehmet Esat Belviranli, Seyong Lee, Jeffrey S. Vetter, Allen D. Malony |
PACT | 2 |
| 2018 | Juggler: a dependence-aware task-based execution framework for GPUsabstractScientific applications with single instruction, multiple data (SIMD) computations show considerable performance improvements when run on today's graphics processing units (GPUs). However, the existence of data dependences across thread blocks may significantly impact the speedup by requiring global synchronization across multiprocessors (SMs) inside the GPU. To efficiently run applications with interblock data dependences, we need fine-granular task-based execution models that will treat SMs inside a GPU as stand-alone parallel processing units. Such a scheme will enable faster execution by utilizing all internal computation elements inside the GPU and eliminating unnecessary waits during device-wide global barriers. Mehmet Esat Belviranli, Seyong Lee, Jeffrey S. Vetter, Laxmi N. Bhuyan |
PPoPP | 1 |
| 2018 | DRAGON: breaking GPU memory capacity limits with direct NVM access
Pak Markthub, Mehmet Esat Belviranli, Seyong Lee, Jeffrey S. Vetter, Satoshi Matsuoka |
SC | 2 |
| 2017 | Wireframe: supporting data-dependent parallelism through dependency graph execution in GPUsabstractGPUs lack fundamental support for data-dependent parallelism and synchronization. While CUDA Dynamic Parallelism signals progress in this direction, many limitations and challenges still remain. This paper introduces Wireframe, a hardware-software solution that enables generalized support for data-dependent parallelism and synchronization. Wireframe enables applications to naturally express execution dependencies across different thread blocks through a dependency graph abstraction at run-time, which is sent to the GPU hardware at kernel launch. At run-time, the hardware enforces the dependencies specified in the dependency graph through a dependency-aware thread block scheduler. Overall, Wireframe is able to improve total execution time up to 65.20% with an average of 45.07%. AmirAli Abdolrashidi, Devashree Tripathy, Mehmet Esat Belviranli, Laxmi N. Bhuyan, Daniel Wong 0001 |
MICRO | 3 |
| 2016 | CuMAS: Data Transfer Aware Multi-Application Scheduling for Shared GPUsabstractRecent generations of GPUs and their corresponding APIs provide means for sharing compute resources among multiple applications with greater efficiency than ever. This advance has enabled the GPUs to act as shared computation resources in multi-user environments, like supercomputers and cloud computing. Recent research has focused on maximizing the utilization of GPU computing resources by simultaneously executing multiple GPU applications (i.e., concurrent kernels) via temporal or spatial partitioning. However, they have not considered maximizing the utilization of the PCI-e bus which is equally important as applications spend a considerable amount of time on data transfers. Mehmet Esat Belviranli, Farzad Khorasani, Laxmi N. Bhuyan, Rajiv Gupta 0001 |
ICS | 1 |
| 2015 | Stadium Hashing: Scalable and Flexible Hashing on GPUsabstractHashing is one of the most fundamental operations that provides a means for a program to obtain fast access to large amounts of data. Despite the emergence of GPUs as many-threaded general purpose processors, high performance parallel data hashing solutions for GPUs are yet to receive adequate attention. Existing hashing solutions for GPUs not only impose restrictions (e.g., inability to concurrently execute insertion and retrieval operations, limitation on the size of key-value data pairs) that limit their applicability, their performance does not scale to large hash tables that must be kept out-of-core in the host memory. In this paper we present Stadium Hashing (Stash) that is scalable to large hash tables and practical as it does not impose the aforementioned restrictions. To support large out-of-core hash tables, Stash uses a compact data structure named ticket-board that is separate from hash table buckets and is held inside GPU global memory. Ticket-board locally resolves significant portion of insertion and lookup operations and hence, by reducing accesses to the host memory, it accelerates the execution of these operations. Split design of the ticket-board also enables arbitrarily large keys and values. Unlike existing methods, Stash naturally supports concurrent insertions and retrievals due to its use of double hashing as the collision resolution strategy. Furthermore, we propose Stash with collaborative lanes (clStash) that enhances GPU's SIMD resource utilization for batched insertions during hash table creation. For concurrent insertion and retrieval streams, Stadium hashing can be up to 2 and 3 times faster than GPU Cuckoo hashing for in-core and out-of-core tables respectively. Farzad Khorasani, Mehmet Esat Belviranli, Rajiv Gupta 0001, Laxmi N. Bhuyan |
PACT | 2 |
| 2015 | PeerWave: Exploiting Wavefront Parallelism on GPUs with Peer-SM SynchronizationabstractNested loops with regular iteration dependencies span a large class of applications ranging from string matching to linear system solvers. Wavefront parallelism is a well-known technique to enable concurrent processing of such applications and is widely being used on GPUs to benefit from their massively parallel computing capabilities. Wavefront parallelism on GPUs uses global barriers between processing of tiles to enforce data dependencies. However, such diagonal-wide synchronization causes load imbalance by forcing SMs to wait for the completion of the SM with longest computation. Moreover, diagonal processing causes loss of locality due to elements that border adjacent tiles. Mehmet Esat Belviranli, Laxmi N. Bhuyan, Rajiv Gupta 0001, Qi Zhu 0002 |
ICS | 1 |
| 2013 | Thermal prediction and scheduling of network applications on multicore processorsabstractAs processor power density increases, chip/core temperature control becomes critical for building multicore systems. This paper addresses the problem of inter-core thermal coupling and periodic thermal variation while executing multi-threaded network applications in a multicore architecture. Chih-Hsun Chou, Mehmet Esat Belviranli, Laxmi N. Bhuyan |
ANCS | 2 |
| 2013 | A dynamic self-scheduling scheme for heterogeneous multiprocessor architecturesabstractToday's heterogeneous architectures bring together multiple general-purpose CPUs and multiple domain-specific GPUs and FPGAs to provide dramatic speedup for many applications. However, the challenge lies in utilizing these heterogeneous processors to optimize overall application performance by minimizing workload completion time. Operating system and application development for these systems is in their infancy. In this article, we propose a new scheduling and workload balancing scheme, HDSS, for execution of loops having dependent or independent iterations on heterogeneous multiprocessor systems. The new algorithm dynamically learns the computational power of each processor during an adaptive phase and then schedules the remainder of the workload using a weighted self-scheduling scheme during the completion phase. Different from previous studies, our scheme uniquely considers the runtime effects of block sizes on the performance for heterogeneous multiprocessors. It finds the right trade-off between large and small block sizes to maintain balanced workload while keeping the accelerator utilization at maximum. Our algorithm does not require offline training or architecture-specific parameters. We have evaluated our scheme on two different heterogeneous architectures: AMD 64-core Bulldozer system with nVidia Fermi C2050 GPU and Intel Xeon 32-core SGI Altix 4700 supercomputer with Xilinx Virtex 4 FPGAs. The experimental results show that our new scheduling algorithm can achieve performance improvements up to over 200% when compared to the closest existing load balancing scheme. Our algorithm also achieves full processor utilization with all processors completing at nearly the same time which is significantly better than alternative current approaches. Mehmet Esat Belviranli, Laxmi N. Bhuyan, Rajiv Gupta 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2013 | CiSE: A Circular Spring Embedder Layout AlgorithmabstractWe present a new algorithm for automatic layout of clustered graphs using a circular style. The algorithm tries to determine optimal location and orientation of individual clusters intrinsically within a modified spring embedder. Heuristics such as reversal of the order of nodes in a cluster and swap of neighboring node pairs in the same cluster are employed intermittently to further relax the spring embedder system, resulting in reduced inter-cluster edge crossings. Unlike other algorithms generating circular drawings, our algorithm does not require the quotient graph to be acyclic, nor does it sacrifice the edge crossing number of individual clusters to improve respective positioning of the clusters. Moreover, it reduces the total area required by a cluster by using the space inside the associated circle. Experimental results show that the execution time and quality of the produced drawings with respect to commonly accepted layout criteria are quite satisfactory, surpassing previous algorithms. The algorithm has also been successfully implemented and made publicly available as part of a compound and clustered graph editing and layout tool named CHISIO. Ugur Dogrusoz, Mehmet Esat Belviranli, Alptug Dilek |
IEEE Trans. Vis. Comput. Graph. | 2 |