EDBT 2026 Demo / reviewers in the wild / expert
Ján Veselý
dblp:173/9812
· DBLP profile ↗
8ranked-venue papers
3as first author
2since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Memory systems · 37% GPUs and heterogeneous computing · 30% Cloud and datacenter computing · 18% | |
| Software engineering, system software, and programming languages
3 papers |
Operating systems · 59% Program analysis · 41% |
Topics — the 17 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU memory management |
0.8 | 1 | 2024 | SUV: Static Analysis Guided Unified Virtual Memory · MICRO 2024 |
Cloud and datacenter computing › resource management › datacenter memory management
memory oversubscription |
0.8 | 1 | 2024 | SUV: Static Analysis Guided Unified Virtual Memory · MICRO 2024 |
GPUs and heterogeneous computing › GPU memory management
unified virtual memory |
0.8 | 1 | 2024 | SUV: Static Analysis Guided Unified Virtual Memory · MICRO 2024 |
Memory systems › memory management
virtual memory |
0.5 | 2 | 2017 | Hardware Translation Coherence for Virtualized Systems · ISCA 2017 Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015 |
Embedded and real-time systems
brain-computer interface |
0.4 | 1 | 2020 | Hardware-Software Co-Design for Brain-Computer Interfaces · ISCA 2020 |
Operating systems › resource management › memory management
virtual memory |
0.3 | 1 | 2018 | LATR: Lazy Translation Coherence · ASPLOS 2018 |
Memory systems › memory management › virtual memory
address translation |
0.3 | 2 | 2018 | Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015 LATR: Lazy Translation Coherence · ASPLOS 2018 |
Cloud and datacenter computing
virtualization |
0.3 | 2 | 2017 | Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015 Hardware Translation Coherence for Virtualized Systems · ISCA 2017 |
Memory systems
cache coherence |
0.3 | 1 | 2017 | Hardware Translation Coherence for Virtualized Systems · ISCA 2017 |
Memory systems › virtual memory management
translation coherence |
0.3 | 1 | 2017 | Hardware Translation Coherence for Virtualized Systems · ISCA 2017 |
Program analysis › memory analysis
memory access pattern analysis |
0.2 | 1 | 2024 | SUV: Static Analysis Guided Unified Virtual Memory · MICRO 2024 |
Program analysis
static analysis |
0.2 | 1 | 2024 | SUV: Static Analysis Guided Unified Virtual Memory · MICRO 2024 |
Memory systems › memory management › virtual memory
huge pages |
0.2 | 1 | 2015 | Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015 |
Memory systems › memory management
memory deduplication |
0.2 | 1 | 2015 | Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015 |
Memory systems
memory management |
0.2 | 1 | 2015 | Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015 |
Memory systems › memory management › virtual memory › address translation
TLB |
0.2 | 1 | 2015 | Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015 |
Cloud and datacenter computing › virtualization
virtual machine migration |
0.1 | 1 | 2015 | Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015 |
Methods — techniques the papers use, named apart from their topics
static analysis · 1.5software prefetching · 1.5page migration · 1.5OS kernel modification · 0.7heterogeneous processing elements · 0.4RISC-V microcontroller · 0.4speculative translation grouping · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SUV: Static Analysis Guided Unified Virtual MemoryabstractUnified Virtual Memory (UVM) eases GPU programming and enables oversubscription of the limited GPU memory capacity. Unfortunately, UVM may cause significant slowdowns due to thrashing of the GPU memory and overheads of page faults, especially under memory oversubscription. We propose to leverage high-level memory access patterns of CUDA applications to reduce the overheads of UVM. We create SUV, a hybrid framework that leverages compiler-inferred (static analysis) memory access semantics to make proactive memory management decisions where possible and selectively leverages runtime page migrations where needed. It can pin data structures entirely or partially on the GPU memory or CPU's DRAM based on their inferred usefulness and automatically issue software prefetches at kernel boundaries. It also selectively lets parts of data structures migrate on demand onto reserved HBM capacity at runtime. SUV reduces execution times of a variety of applications by 74% over UVM under memory oversubscription. Pratheek B, Guilherme Cox, Ján Veselý, Arkaprava Basu |
MICRO | 3 |
| 2022 | Distill: Domain-Specific Compilation for Cognitive ModelsabstractComputational models of cognition enable a better understanding of the human brain and behavior, psychiatric and neurological illnesses, clinical interventions to treat illnesses, and also offer a path towards human-like artificial intelligence. Cognitive models are also, however, laborious to develop, requiring composition of many types of computational tasks, and suffer from poor performance as they are generally designed using high-level languages like Python. In this work, we present Distill, a domain-specific compilation tool to accelerate cognitive models while continuing to offer cognitive scientists the ability to develop their models in flexible high-level languages. Distill uses domain-specific knowledge to compile Python-based cognitive models into LLVM IR, carefully stripping away features like dynamic typing and memory management that add performance overheads without being necessary for the underlying computation of the models. The net effect is an average of 27 × performance improvement in model execution over state-of-the-art techniques using Pyston and PyPy. Distill also repurposes classical compiler data flow analyses to reveal properties about data flow in cognitive models that are useful to cognitive scientists. Distill is publicly available, integrated in the PsyNeuLink cognitive modeling environment, and is already being used by researchers in the brain sciences. Ján Veselý, Raghavendra Pradyumna Pothukuchi, Ketaki Joshi, Samyak Gupta, Jonathan D. Cohen 0003, Abhishek Bhattacharjee |
CGO | 1 |
| 2020 | Hardware-Software Co-Design for Brain-Computer InterfacesabstractBrain-computer interfaces (BCIs) offer avenues to treat neurological disorders, shed light on brain function, and interface the brain with the digital world. Their wider adoption rests, however, on achieving adequate real-time performance, meeting stringent power constraints, and adhering to FDA-mandated safety requirements for chronic implantation. BCIs have, to date, been designed as custom ASICs for specific diseases or for specific tasks in specific brain regions. General-purpose architectures that can be used to treat multiple diseases and enable various computational tasks are needed for wider BCI adoption, but the conventional wisdom is that such systems cannot meet necessary performance and power constraints. We present HALO (Hardware Architecture for LOw-power BCIs), a general-purpose architecture for implantable BCIs. HALO enables tasks such as treatment of disorders (e.g., epilepsy, movement disorders), and records/processes data for studies that advance our understanding of the brain. We use electrophysiological data from the motor cortex of a non-human primate to determine how to decompose HALO's computational capabilities into hardware building blocks. We simplify, prune, and share these building blocks to judiciously use available hardware resources while enabling many modes of brain-computer interaction. The result is a configurable heterogeneous array of hardware processing elements (PEs). The PEs are configured by a low-power RISC-V micro-controller into signal processing pipelines that meet the target performance and power constraints necessary to deploy HALO widely and safely. Ioannis Karageorgos, Karthik Sriram, Ján Veselý, Marc Powell, David A. Borton, Rajit Manohar, Abhishek Bhattacharjee |
ISCA | 3 |
| 2018 | LATR: Lazy Translation CoherenceabstractWe propose LATR-lazy TLB coherence-a software-based TLB shootdown mechanism that can alleviate the overhead of the synchronous TLB shootdown mechanism in existing operating systems. By handling the TLB coherence in a lazy fashion, LATR can avoid expensive IPIs which are required for delivering a shootdown signal to remote cores, and the performance overhead of associated interrupt handlers. Therefore, virtual memory operations, such as free and page migration operations, can benefit significantly from LATR's mechanism. For example, LATR improves the latency of munmap() by 70.8% on a 2-socket machine, a widely used configuration in modern data centers. Real-world, performance-critical applications such as web servers can also benefit from LATR: without any application-level changes, LATR improves Apache by 59.9% compared to Linux, and by 37.9% compared to ABIS, a highly optimized, state-of-the-art TLB coherence technique. Mohan Kumar, Steffen Maass, Sanidhya Kashyap, Ján Veselý, Zi Yan, Taesoo Kim, Abhishek Bhattacharjee, Tushar Krishna |
ASPLOS | 4 |
| 2018 | Generic System Calls for GPUsabstractGPUs are becoming first-class compute citizens and increasingly support programmability-enhancing features such as shared virtual memory and hardware cache coherence. This enables them to run a wider variety of programs. However, a key aspect of general-purpose programming where GPUs still have room for improvement is the ability to invoke system calls. We explore how to directly invoke system calls from GPUs. We examine how system calls can be integrated with GPGPU programming models, where thousands of threads are organized in a hierarchy of execution groups. To answer questions on GPU system call usage and efficiency, we implement Genesys, a generic GPU system call interface for Linux. Numerous architectural and OS issues are considered and subtle changes to Linux are necessary, as the existing kernel assumes that only CPUs invoke system calls. We assess the performance of Genesys using micro-benchmarks and applications that exercise system calls for signals, memory management, filesystems, and networking. Ján Veselý, Arkaprava Basu, Abhishek Bhattacharjee, Gabriel H. Loh, Mark Oskin, Steven K. Reinhardt |
ISCA | 1 |
| 2017 | Hardware Translation Coherence for Virtualized SystemsabstractTo improve system performance, operating systems (OSes) often undertake activities that require modification of virtual-to-physical address translations. For example, the OS may migrate data between physical pages to manage heterogeneous memory devices. We refer to such activities as page remappings. Unfortunately, page remappings are expensive. We show that a big part of this cost arises from address translation coherence, particularly on systems employing virtualization. In response, we propose hardware translation invalidation and coherence or HATRIC, a readily implementable hardware mechanism to piggyback translation coherence atop existing cache coherence protocols. We perform detailed studies using KVM-based virtualization, showing that HATRIC achieves up to 30% performance and 10% energy benefits, for per-CPU area overheads of 0.2%. We also quantify HATRIC's benefits on systems running Xen and find up to 33% performance improvements. Zi Yan, Ján Veselý, Guilherme Cox, Abhishek Bhattacharjee |
ISCA | 2 |
| 2016 | Observations and opportunities in architecting shared virtual memory for heterogeneous systemsabstractComputing is becoming increasingly heterogeneous with accelerators like GPUs being tightly integrated with CPUs on the same die. Extending the CPU's virtual addressing mechanism to these accelerators is a key step in making accelerators easily programmable. In this work, we analyze, using real-system measurements, shared virtual memory across the CPU and an integrated GPU. We make several key observations and highlight consequent research opportunities: (1) servicing a TLB miss from the GPU can be an order of magnitude slower than that from the CPU and consequently it is imperative to enable many concurrent TLB misses to hide this larger latency; (2) divergence in memory accesses impacts the GPU's address translation more than the rest of the memory hierarchy, and research in designing address translation mechanisms tolerant to this effect is imperative; and (3) page faults from the GPU are considerably slower than that from the CPU and software-hardware co-design is essential for efficient implementation of page faults from throughput-oriented accelerators like GPUs. We present a detailed measurement study of a commercially available integrated APU that illustrates these effects and motivates future research opportunities. Ján Veselý, Arkaprava Basu, Mark Oskin, Gabriel H. Loh, Abhishek Bhattacharjee |
ISPASS | 1 |
| 2015 | Large pages and lightweight memory management in virtualized environments: can you have it both ways?abstractLarge pages have long been used to mitigate address translation overheads on big-memory systems, particularly in virtualized environments where TLB miss overheads are severe. We show, however, that far from being a panacea, large pages are used sparingly by modern virtualization software. This is because large pages often preclude lightweight memory management, which can outweigh their Translation Lookaside Buffer (TLB) benefits. For example, they reduce opportunities to deduplicate memory among virtual machines in overcommitted systems, interfere with lightweight memory monitoring, and hamper the agility of virtual machine (VM) migrations. While many of these problems are particularly severe in overcommitted systems with scarce memory resources, they can (and often do) exist generally in cloud deployments. In response, virtualization software often (though it doesn't have to) splinters guest operating system (OS) large pages into small system physical pages, sacrificing address translation performance for overall system-level benefits. We introduce simple hardware that bridges this fundamental conflict, using speculative techniques to group contiguous, aligned small page translations such that they approach the address translation performance of large pages. Our Generalized Large-page Utilization Enhancements (GLUE) allow system hypervisors to splinter large pages for agile memory management, while retaining almost all of the TLB performance of unsplintered large pages. Binh Pham 0003, Ján Veselý, Gabriel H. Loh, Abhishek Bhattacharjee |
MICRO | 2 |