Ján Veselý

dblp:173/9812 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
2since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Memory systems · 37% GPUs and heterogeneous computing · 30% Cloud and datacenter computing · 18%
Software engineering, system software, and programming languages
3 papers
Operating systems · 59% Program analysis · 41%

Topics — the 17 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU memory management
0.812024
SUV: Static Analysis Guided Unified Virtual Memory · MICRO 2024
Cloud and datacenter computing › resource management › datacenter memory management
memory oversubscription
0.812024
SUV: Static Analysis Guided Unified Virtual Memory · MICRO 2024
GPUs and heterogeneous computing › GPU memory management
unified virtual memory
0.812024
SUV: Static Analysis Guided Unified Virtual Memory · MICRO 2024
Memory systems › memory management
virtual memory
0.522017
Hardware Translation Coherence for Virtualized Systems · ISCA 2017
Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015
Embedded and real-time systems
brain-computer interface
0.412020
Hardware-Software Co-Design for Brain-Computer Interfaces · ISCA 2020
Operating systems › resource management › memory management
virtual memory
0.312018
LATR: Lazy Translation Coherence · ASPLOS 2018
Memory systems › memory management › virtual memory
address translation
0.322018
Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015
LATR: Lazy Translation Coherence · ASPLOS 2018
Cloud and datacenter computing
virtualization
0.322017
Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015
Hardware Translation Coherence for Virtualized Systems · ISCA 2017
Memory systems
cache coherence
0.312017
Hardware Translation Coherence for Virtualized Systems · ISCA 2017
Memory systems › virtual memory management
translation coherence
0.312017
Hardware Translation Coherence for Virtualized Systems · ISCA 2017
Program analysis › memory analysis
memory access pattern analysis
0.212024
SUV: Static Analysis Guided Unified Virtual Memory · MICRO 2024
Program analysis
static analysis
0.212024
SUV: Static Analysis Guided Unified Virtual Memory · MICRO 2024
Memory systems › memory management › virtual memory
huge pages
0.212015
Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015
Memory systems › memory management
memory deduplication
0.212015
Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015
Memory systems
memory management
0.212015
Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015
Memory systems › memory management › virtual memory › address translation
TLB
0.212015
Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015
Cloud and datacenter computing › virtualization
virtual machine migration
0.112015
Large pages and lightweight memory management in virtualized environments: can you have it both ways? · MICRO 2015

Methods — techniques the papers use, named apart from their topics

static analysis · 1.5software prefetching · 1.5page migration · 1.5OS kernel modification · 0.7heterogeneous processing elements · 0.4RISC-V microcontroller · 0.4speculative translation grouping · 0.2
YearPublicationVenuePosition
2024 SUV: Static Analysis Guided Unified Virtual Memory
abstract
Unified Virtual Memory (UVM) eases GPU programming and enables oversubscription of the limited GPU memory capacity. Unfortunately, UVM may cause significant slowdowns due to thrashing of the GPU memory and overheads of page faults, especially under memory oversubscription. We propose to leverage high-level memory access patterns of CUDA applications to reduce the overheads of UVM. We create SUV, a hybrid framework that leverages compiler-inferred (static analysis) memory access semantics to make proactive memory management decisions where possible and selectively leverages runtime page migrations where needed. It can pin data structures entirely or partially on the GPU memory or CPU's DRAM based on their inferred usefulness and automatically issue software prefetches at kernel boundaries. It also selectively lets parts of data structures migrate on demand onto reserved HBM capacity at runtime. SUV reduces execution times of a variety of applications by 74% over UVM under memory oversubscription.
Pratheek B, Guilherme Cox, Ján Veselý, Arkaprava Basu
MICRO3
2022 Distill: Domain-Specific Compilation for Cognitive Models
abstract
Computational models of cognition enable a better understanding of the human brain and behavior, psychiatric and neurological illnesses, clinical interventions to treat illnesses, and also offer a path towards human-like artificial intelligence. Cognitive models are also, however, laborious to develop, requiring composition of many types of computational tasks, and suffer from poor performance as they are generally designed using high-level languages like Python. In this work, we present Distill, a domain-specific compilation tool to accelerate cognitive models while continuing to offer cognitive scientists the ability to develop their models in flexible high-level languages. Distill uses domain-specific knowledge to compile Python-based cognitive models into LLVM IR, carefully stripping away features like dynamic typing and memory management that add performance overheads without being necessary for the underlying computation of the models. The net effect is an average of 27 × performance improvement in model execution over state-of-the-art techniques using Pyston and PyPy. Distill also repurposes classical compiler data flow analyses to reveal properties about data flow in cognitive models that are useful to cognitive scientists. Distill is publicly available, integrated in the PsyNeuLink cognitive modeling environment, and is already being used by researchers in the brain sciences.
Ján Veselý, Raghavendra Pradyumna Pothukuchi, Ketaki Joshi, Samyak Gupta, Jonathan D. Cohen 0003, Abhishek Bhattacharjee
CGO1
2020 Hardware-Software Co-Design for Brain-Computer Interfaces
abstract
Brain-computer interfaces (BCIs) offer avenues to treat neurological disorders, shed light on brain function, and interface the brain with the digital world. Their wider adoption rests, however, on achieving adequate real-time performance, meeting stringent power constraints, and adhering to FDA-mandated safety requirements for chronic implantation. BCIs have, to date, been designed as custom ASICs for specific diseases or for specific tasks in specific brain regions. General-purpose architectures that can be used to treat multiple diseases and enable various computational tasks are needed for wider BCI adoption, but the conventional wisdom is that such systems cannot meet necessary performance and power constraints. We present HALO (Hardware Architecture for LOw-power BCIs), a general-purpose architecture for implantable BCIs. HALO enables tasks such as treatment of disorders (e.g., epilepsy, movement disorders), and records/processes data for studies that advance our understanding of the brain. We use electrophysiological data from the motor cortex of a non-human primate to determine how to decompose HALO's computational capabilities into hardware building blocks. We simplify, prune, and share these building blocks to judiciously use available hardware resources while enabling many modes of brain-computer interaction. The result is a configurable heterogeneous array of hardware processing elements (PEs). The PEs are configured by a low-power RISC-V micro-controller into signal processing pipelines that meet the target performance and power constraints necessary to deploy HALO widely and safely.
Ioannis Karageorgos, Karthik Sriram, Ján Veselý, Marc Powell, David A. Borton, Rajit Manohar, Abhishek Bhattacharjee
ISCA3
2018 LATR: Lazy Translation Coherence
abstract
We propose LATR-lazy TLB coherence-a software-based TLB shootdown mechanism that can alleviate the overhead of the synchronous TLB shootdown mechanism in existing operating systems. By handling the TLB coherence in a lazy fashion, LATR can avoid expensive IPIs which are required for delivering a shootdown signal to remote cores, and the performance overhead of associated interrupt handlers. Therefore, virtual memory operations, such as free and page migration operations, can benefit significantly from LATR's mechanism. For example, LATR improves the latency of munmap() by 70.8% on a 2-socket machine, a widely used configuration in modern data centers. Real-world, performance-critical applications such as web servers can also benefit from LATR: without any application-level changes, LATR improves Apache by 59.9% compared to Linux, and by 37.9% compared to ABIS, a highly optimized, state-of-the-art TLB coherence technique.
Mohan Kumar, Steffen Maass, Sanidhya Kashyap, Ján Veselý, Zi Yan, Taesoo Kim, Abhishek Bhattacharjee, Tushar Krishna
ASPLOS4
2018 Generic System Calls for GPUs
abstract
GPUs are becoming first-class compute citizens and increasingly support programmability-enhancing features such as shared virtual memory and hardware cache coherence. This enables them to run a wider variety of programs. However, a key aspect of general-purpose programming where GPUs still have room for improvement is the ability to invoke system calls. We explore how to directly invoke system calls from GPUs. We examine how system calls can be integrated with GPGPU programming models, where thousands of threads are organized in a hierarchy of execution groups. To answer questions on GPU system call usage and efficiency, we implement Genesys, a generic GPU system call interface for Linux. Numerous architectural and OS issues are considered and subtle changes to Linux are necessary, as the existing kernel assumes that only CPUs invoke system calls. We assess the performance of Genesys using micro-benchmarks and applications that exercise system calls for signals, memory management, filesystems, and networking.
Ján Veselý, Arkaprava Basu, Abhishek Bhattacharjee, Gabriel H. Loh, Mark Oskin, Steven K. Reinhardt
ISCA1
2017 Hardware Translation Coherence for Virtualized Systems
abstract
To improve system performance, operating systems (OSes) often undertake activities that require modification of virtual-to-physical address translations. For example, the OS may migrate data between physical pages to manage heterogeneous memory devices. We refer to such activities as page remappings. Unfortunately, page remappings are expensive. We show that a big part of this cost arises from address translation coherence, particularly on systems employing virtualization. In response, we propose hardware translation invalidation and coherence or HATRIC, a readily implementable hardware mechanism to piggyback translation coherence atop existing cache coherence protocols. We perform detailed studies using KVM-based virtualization, showing that HATRIC achieves up to 30% performance and 10% energy benefits, for per-CPU area overheads of 0.2%. We also quantify HATRIC's benefits on systems running Xen and find up to 33% performance improvements.
Zi Yan, Ján Veselý, Guilherme Cox, Abhishek Bhattacharjee
ISCA2
2016 Observations and opportunities in architecting shared virtual memory for heterogeneous systems
abstract
Computing is becoming increasingly heterogeneous with accelerators like GPUs being tightly integrated with CPUs on the same die. Extending the CPU's virtual addressing mechanism to these accelerators is a key step in making accelerators easily programmable. In this work, we analyze, using real-system measurements, shared virtual memory across the CPU and an integrated GPU. We make several key observations and highlight consequent research opportunities: (1) servicing a TLB miss from the GPU can be an order of magnitude slower than that from the CPU and consequently it is imperative to enable many concurrent TLB misses to hide this larger latency; (2) divergence in memory accesses impacts the GPU's address translation more than the rest of the memory hierarchy, and research in designing address translation mechanisms tolerant to this effect is imperative; and (3) page faults from the GPU are considerably slower than that from the CPU and software-hardware co-design is essential for efficient implementation of page faults from throughput-oriented accelerators like GPUs. We present a detailed measurement study of a commercially available integrated APU that illustrates these effects and motivates future research opportunities.
Ján Veselý, Arkaprava Basu, Mark Oskin, Gabriel H. Loh, Abhishek Bhattacharjee
ISPASS1
2015 Large pages and lightweight memory management in virtualized environments: can you have it both ways?
abstract
Large pages have long been used to mitigate address translation overheads on big-memory systems, particularly in virtualized environments where TLB miss overheads are severe. We show, however, that far from being a panacea, large pages are used sparingly by modern virtualization software. This is because large pages often preclude lightweight memory management, which can outweigh their Translation Lookaside Buffer (TLB) benefits. For example, they reduce opportunities to deduplicate memory among virtual machines in overcommitted systems, interfere with lightweight memory monitoring, and hamper the agility of virtual machine (VM) migrations. While many of these problems are particularly severe in overcommitted systems with scarce memory resources, they can (and often do) exist generally in cloud deployments. In response, virtualization software often (though it doesn't have to) splinters guest operating system (OS) large pages into small system physical pages, sacrificing address translation performance for overall system-level benefits. We introduce simple hardware that bridges this fundamental conflict, using speculative techniques to group contiguous, aligned small page translations such that they approach the address translation performance of large pages. Our Generalized Large-page Utilization Enhancements (GLUE) allow system hypervisors to splinter large pages for agile memory management, while retaining almost all of the TLB performance of unsplintered large pages.
Binh Pham 0003, Ján Veselý, Gabriel H. Loh, Abhishek Bhattacharjee
MICRO2