Djordje Jevdjic

dblp:97/11028 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0002-3341-9364ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2024 DNA Storage Toolkit: A Modular End-to-End DNA Data Storage Codec and Simulator
abstract
With the amount of data being generated every year increasing exponentially, figuring out where and how to store it efficiently and inexpensively is becoming a larger problem every day. The rapid improvement in performance and cost of DNA synthesis and sequencing methods has led to an increased interest in the use of DNA as a durable and compact medium for data storage. Today, we have a large spectrum of available chemical tools that enable efficient data access and manipulation of in-DNA data. While several DNA storage architectures have been proposed, there is no open-source codec or simulator that implements all of the required components of the DNA-based data storage pipeline for research and development. We present an open-source end-to-end DNA data storage toolkit that can take an input file through the entire DNA storage pipeline. Our work contains implementations of the state-of-the-art techniques for each step of the pipeline, including our own algorithms for each step. These steps include encoding data into DNA strands, simulating the wetlab processes of synthesis, storage and sequencing of those DNA strands, clustering of the sequenced results, reconstruction of DNA strands from noisy clusters, and decoding the initially encoded file with support for error-correction mechanisms. Each module can be used individually or combined to form an entire pipeline. We hope that our toolkit will be useful to researchers and developers who seek to experiment with the new and promising storage technology.
Puru Sharma, Gary Goh Yipeng, Longshen Ou, Dehui Lin, Djordje Jevdjic
ISPASS7
2024 Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Pengfei Zuo
USENIX ATC5
2024 Scalable and Effective Page-table and TLB management on NUMA Systems
Qingxuan Kang, Hao-Wei Tee, Kyle Timothy Ng Chu, Alireza Sanaee, Djordje Jevdjic
USENIX ATC6
2023 Efficiently Enabling Block Semantics and Data Updates in DNA Storage
abstract
We propose a novel and flexible DNA-storage architecture, which divides the storage space into fixed-size units (blocks) that can be independently and efficiently accessed at random for both read and write operations, and further allows efficient sequential access to consecutive data blocks. In contrast to prior work, in our architecture a pair of random-access PCR primers of length 20 does not define a single object, but an independent storage partition, which is internally blocked and managed independently of other partitions. We expose the flexibility and constraints with which the internal address space of each partition can be managed, and incorporate them into our design to provide rich and functional storage semantics, such as block-storage organization, efficient implementation of data updates, and sequential access. To leverage the full power of the prefix-based nature of PCR addressing, we define a methodology for transforming the internal addressing scheme of a partition into an equivalent that is PCR-compatible. This allows us to run PCR with primers that can be variably elongated to include a desired part of the internal address, and thus narrow down the scope of the reaction to retrieve a specific block or a range of blocks within the partition with sufficiently high accuracy. Our wetlab evaluation demonstrates the practicality of the proposed ideas and a 140x reduction in sequencing cost and latency for retrieval of individual blocks within the partition.
Puru Sharma, Cheng-Kai Lim, Dehui Lin, Yash Pote, Djordje Jevdjic
MICRO5
2022 Managing reliability skew in DNA storage
abstract
DNA is emerging as an increasingly attractive medium for data storage due to a number of important and unique advantages it offers, most notably the unprecedented durability and density. While the technology is evolving rapidly, the prohibitive cost of reads and writes, the high frequency and the peculiar nature of errors occurring in DNA storage pose a significant challenge to its adoption.
Dehui Lin, Yasamin Tabatabaee, Yash Pote, Djordje Jevdjic
ISCA4
2022 OS-level Implications of Using DRAM Caches in Memory Disaggregation
abstract
Memory disaggregation has attracted great attention recently due to its benefits in resource utilization efficiency, isolation of failures, and easier reconfiguration of memory hardware. However, applications running on a system with disaggregated memory are expected to suffer from performance degradation due to increased remote memory access latency and network contention. The performance gap is meant to be bridged using DRAM caches on the processor side, which would filter out most of the network traffic.This work examines the overheads of the disaggregated memory abstraction. By experimenting with both micro-benchmarks and production applications, we observe severe degradation in memory access latency and potential bottlenecks within the OS kernel. These bottlenecks could potentially be avoided through low-level optimizations in memory management tailored for memory disaggregation.
Hao-Wei Tee, Alireza Sanaee, Soh Boon Jun, Djordje Jevdjic
ISPASS5
2022 Simulating Noisy Channels in DNA Storage
abstract
Compared to conventional storage mediums, DNA-based data storage offers benefits such as durability, high density and low energy consumption. With increased demand for DNA data storage, it has become important to quickly evaluate proposed approaches. However, experiments that involves reading and writing synthetic DNA are costly and time-consuming, thus requiring cheap and fast simulation prior to experimentation. DNA sequencing technologies such as Nanopore and Illumina have highly characteristic error profiles, and simulating them is challenging. We propose a DNA simulator for Nanopore data that improves on existing simulators by incorporating key parameters; our simulator better converges to error profiles of real data on most parameters.We show that the spatial distribution of errors within a strand is a key determinant of trace reconstruction accuracy; which is a factor that had not been considered by existing simulators.
Mayank Keoliya, Puru Sharma, Djordje Jevdjic
ISPASS3
2017 Near-Memory Address Translation
abstract
Memory and logic integration on the same chip is becoming increasingly cost effective, creating the opportunity to offload data-intensive functionality to processing units placed inside memory chips. The introduction of memory-side processing units (MPUs) into conventional systems faces virtual memory as the first big showstopper: without efficient hardware support for address translation MPUs have highly limited applicability. Unfortunately, conventional translation mechanisms fall short of providing fast translations as contemporary memories exceed the reach of TLBs, making expensive page walks common. In this paper, we are the first to show that the historically important flexibility to map any virtual page to any page frame is unnecessary in today's servers. We find that while limiting the associativity of the virtual-to-physical mapping incurs no penalty, it can break the translate-then-fetch serialization if combined with careful data placement in the MPU's memory, allowing for translation and data fetch to proceed independently and in parallel. We propose the Distributed Inverted Page Table (DIPTA), a near-memory structure in which the smallest memory partition keeps the translation information for its data share, ensuring that the translation completes together with the data fetch. DIPTA completely eliminates the performance overhead of translation, achieving speedups of up to 3.81× and 2.13× over conventional translation using 4KB and 1GB pages respectively.
Javier Picorel, Djordje Jevdjic, Babak Falsafi
PACT2
2017 Approximate Storage of Compressed and Encrypted Videos
abstract
The popularization of video capture devices has created strong storage demand for encoded videos. Approximate storage can ease this demand by enabling denser storage at the expense of occasional errors. Unfortunately, even minor storage errors, such as bit flips, can result in major visual damage in encoded videos. Similarly, video encryption, widely employed for privacy and digital rights management, may create long dependencies between bits that show little or no tolerance to storage errors.
Djordje Jevdjic, Karin Strauss, Luis Ceze, Henrique S. Malvar
ASPLOS1
2017 Clustering Billions of Reads for DNA Data Storage
abstract
Storing data in synthetic DNA offers the possibility of improving information density and durability by several orders of magnitude compared to current storage technologies. However, DNA data storage requires a computationally intensive process to retrieve the data. In particular, a crucial step in the data retrieval pipeline involves clustering billions of strings with respect to edit distance. Datasets in this domain have many notable properties, such as containing a very large number of small clusters that are well-separated in the edit distance metric space. In this regime, existing algorithms are unsuitable because of either their long running time or low accuracy. To address this issue, we present a novel distributed algorithm for approximately computing the underlying clusters. Our algorithm converges efficiently on any dataset that satisfies certain separability properties, such as those coming from DNA data storage systems. We also prove that, under these assumptions, our algorithm is robust to outliers and high levels of noise. We provide empirical justification of the accuracy, scalability, and convergence of our algorithm on real and synthetic data. Compared to the state-of-the-art algorithm for clustering DNA sequences, our algorithm simultaneously achieves higher accuracy and a 1000x speedup on three real datasets.
Cyrus Rashtchian, Konstantin Makarychev, Miklós Z. Rácz, Siena Ang, Djordje Jevdjic, Sergey Yekhanin, Luis Ceze, Karin Strauss
NIPS5
2014 Unison Cache: A Scalable and Effective Die-Stacked DRAM Cache
abstract
Recent research advocates large die-stacked DRAM caches in many core servers to break the memory latency and bandwidth wall. To realize their full potential, die-stacked DRAM caches necessitate low lookup latencies, high hit rates and the efficient use of off-chip bandwidth. Today's stacked DRAM cache designs fall into two categories based on the granularity at which they manage data: block-based and page-based. The state-of-the-art block-based design, called Alloy Cache, collocates a tag with each data block (e.g., 64B) in the stacked DRAM to provide fast access to data in a single DRAM access. However, such a design suffers from low hit rates due to poor temporal locality in the DRAM cache. In contrast, the state-of-the-art page-based design, called Footprint Cache, organizes the DRAM cache at page granularity (e.g., 4KB), but fetches only the blocks that will likely be touched within a page. In doing so, the Footprint Cache achieves high hit rates with moderate on-chip tag storage and reasonable lookup latency. However, multi-gigabyte stacked DRAM caches will soon be practical and needed by server applications, thereby mandating tens of MBs of tag storage even for page-based DRAM caches. We introduce a novel stacked-DRAM cache design, Unison Cache. Similar to Alloy Cache's approach, Unison Cache incorporates the tag metadata directly into the stacked DRAM to enable scalability to arbitrary stacked-DRAM capacities. Then, leveraging the insights from the Footprint Cache design, Unison Cache employs large, page-sized cache allocation units to achieve high hit rates and reduction in tag overheads, while predicting and fetching only the useful blocks within each page to minimize the off-chip traffic. Our evaluation using server workloads and caches of up to 8GB reveals that Unison cache improves performance by 14% compared to Alloy Cache due to its high hit rate, while outperforming the state-of-the art page-based designs that require impractical SRAM-based tags of around 50MB.
Djordje Jevdjic, Gabriel H. Loh, Cansu Kaynak, Babak Falsafi
MICRO1
2013 From A to E: analyzing TPC's OLTP benchmarks: the obsolete, the ubiquitous, the unexplored
abstract
Introduced in 2007, TPC-E is the most recently standardized OLTP benchmark by TPC. Even though TPC-E has already been around for six years, it has not gained the popularity of its predecessor TPC-C: all the published results for TPC-E use a single database vendor's product. TPC-E is significantly different than its predecessors. Some of its distinguishing characteristics are the non-uniform input creation, longer-running and more complicated transactions, more difficult partitioning etc. These factors slow down the adoption of TPC-E. In turn, there is little knowledge in the community about how TPC-E behaves micro-architecturally and within the database engine.
Pinar Tözün, Ippokratis Pandis, Cansu Kaynak, Djordje Jevdjic, Anastasia Ailamaki
EDBT4
2013 Die-stacked DRAM caches for servers: hit ratio, latency, or bandwidth? have it all with footprint cache
abstract
Recent research advocates using large die-stacked DRAM caches to break the memory bandwidth wall. Existing DRAM cache designs fall into one of two categories --- block-based and page-based. The former organize data in conventional blocks (e.g., 64B), ensuring low off-chip bandwidth utilization, but co-locate tags and data in the stacked DRAM, incurring high lookup latency. Furthermore, such designs suffer from low hit ratios due to poor temporal locality. In contrast, page-based caches, which manage data at larger granularity (e.g., 4KB pages), allow for reduced tag array overhead and fast lookup, and leverage high spatial locality at the cost of moving large amounts of data on and off the chip.
Djordje Jevdjic, Stavros Volos, Babak Falsafi
ISCA1
2012 Clearing the clouds: a study of emerging scale-out workloads on modern hardware
abstract
Emerging scale-out workloads require extensive amounts of computational resources. However, data centers using modern server hardware face physical constraints in space and power, limiting further expansion and calling for improvements in the computational density per server and in the per-operation energy. Continuing to improve the computational resources of the cloud while staying within physical constraints mandates optimizing server efficiency to ensure that server hardware closely matches the needs of scale-out workloads.
Michael Ferdman, Almutaz Adileh, Yusuf Onur Koçberber, Stavros Volos, Mohammad Alisafaee, Djordje Jevdjic, Cansu Kaynak, Adrian Daniel Popescu, Anastasia Ailamaki, Babak Falsafi
ASPLOS6
2012 Thermal characterization of cloud workloads on a power-efficient server-on-chip
abstract
We propose a power-efficient many-core server-on-chip system with 3D-stacked Wide I/O DRAM targeting cloud workloads in datacenters. The integration of 3D-stacked Wide I/O DRAM on top of a logic die increases available memory bandwidth by using dense and fast Through-Silicon Vias (TSVs) instead of off-chip IOs, enabling faster data transfers at much lower energy per bit. We demonstrate a methodology that includes full-system microarchitectural modeling and rapid virtual physical prototyping with emphasis on the thermal analysis. Our findings show that while executing CPU-centric benchmarks (e.g. SPECInt and Dhrystone), the temperature in the server-on-chip (logic+DRAM) is in the range of 175-200°C at a power consumption of less than 20W, exceeding the reliable operating bounds without any cooling solutions, even with embedded cores. However, with real cloud workloads, the power density in the server-on-chip remains much below the temperatures reached by the CPU-centric workloads as a result of much lower power burnt by memory-intensive cloud workloads. We show that such a server-on-chip system is feasible with a low-cost passive heat sink eliminating the need for a high-cost active heat sink with an attached fan, creating an opportunity for overall cost and energy savings in datacenters.
Dragomir Milojevic, Sachin Idgunji, Djordje Jevdjic, Emre Ozer 0001, Pejman Lotfi-Kamran, Andreas Panteli, Andreas Prodromou, Chrysostomos Nicopoulos, Damien Hardy, Babak Falsafi, Yiannakis Sazeides
ICCD3
2012 Scale-out processors
abstract
Scale-out datacenters mandate high per-server throughput to get the maximum benefit from the large TCO investment. Emerging applications (e.g., data serving and web search) that run in these datacenters operate on vast datasets that are not accommodated by on-die caches of existing server chips. Large caches reduce the die area available for cores and lower performance through long access latency when instructions are fetched. Performance on scale-out workloads is maximized through a modestly-sized last-level cache that captures the instruction footprint at the lowest possible access latency. In this work, we introduce a methodology for designing scalable and efficient scale-out server processors. Based on a metric of performance-density, we facilitate the design of optimal multi-core configurations, called pods. Each pod is a complete server that tightly couples a number of cores to a small last-level cache using a fast interconnect. Replicating the pod to fill the die area yields processors which have optimal performance density, leading to maximum per-chip throughput. Moreover, as each pod is a stand-alone server, scale-out processors avoid the expense of global (i.e., interpod) interconnect and coherence. These features synergistically maximize throughput, lower design complexity, and improve technology scalability. In 20nm technology, scaleout chips improve throughput by 5x-6.5x over conventional and by 1.6x-1.9x over emerging tiled organizations.
Pejman Lotfi-Kamran, Boris Grot, Michael Ferdman, Stavros Volos, Yusuf Onur Koçberber, Javier Picorel, Almutaz Adileh, Djordje Jevdjic, Sachin Idgunji, Emre Ozer 0001, Babak Falsafi
ISCA8
2012 Quantifying the Mismatch between Emerging Scale-Out Applications and Modern Processors
abstract
Emerging scale-out workloads require extensive amounts of computational resources. However, data centers using modern server hardware face physical constraints in space and power, limiting further expansion and calling for improvements in the computational density per server and in the per-operation energy. Continuing to improve the computational resources of the cloud while staying within physical constraints mandates optimizing server efficiency to ensure that server hardware closely matches the needs of scale-out workloads. In this work, we introduce CloudSuite, a benchmark suite of emerging scale-out workloads. We use performance counters on modern servers to study scale-out workloads, finding that today’s predominant processor microarchitecture is inefficient for running these workloads. We find that inefficiency comes from the mismatch between the workload needs and modern processors, particularly in the organization of instruction and data memory systems and the processor core microarchitecture. Moreover, while today’s predominant microarchitecture is inefficient when executing scale-out workloads, we find that continuing the current trends will further exacerbate the inefficiency in the future. In this work, we identify the key microarchitectural needs of scale-out workloads, calling for a change in the trajectory of server processors that would lead to improved computational density and power efficiency in data centers.
Michael Ferdman, Almutaz Adileh, Yusuf Onur Koçberber, Stavros Volos, Mohammad Alisafaee, Djordje Jevdjic, Cansu Kaynak, Adrian Daniel Popescu, Anastasia Ailamaki, Babak Falsafi
ACM Trans. Comput. Syst.6