EDBT 2026 Demo / reviewers in the wild / expert
Dejan Vucinic
dblp:144/6240
· DBLP profile ↗
7ranked-venue papers
1as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Computer networks · 2Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 53% Memory systems · 28% Hardware reliability and fault tolerance · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Indexing and storage engines · 100% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › flash and SSD › flash memory management › flash translation layer
address mapping |
0.7 | 1 | 2023 | LeaFTL: A Learning-Based Flash Translation Layer for Solid-State Drives · ASPLOS (2) 2023 |
Storage systems › flash and SSD › flash memory management
flash translation layer |
0.7 | 1 | 2023 | LeaFTL: A Learning-Based Flash Translation Layer for Solid-State Drives · ASPLOS (2) 2023 |
Storage systems › flash and SSD
solid-state drive |
0.7 | 1 | 2023 | LeaFTL: A Learning-Based Flash Translation Layer for Solid-State Drives · ASPLOS (2) 2023 |
Memory systems
non-volatile memory |
0.5 | 2 | 2018 | Consensus for Non-volatile Main Memory · ICNP 2018 DC express: shortest latency protocol for reading phase change memory over PCI express · FAST 2014 |
Hardware reliability and fault tolerance
memory reliability |
0.3 | 1 | 2018 | Consensus for Non-volatile Main Memory · ICNP 2018 |
Memory systems › non-volatile memory
storage class memory |
0.3 | 1 | 2018 | Consensus for Non-volatile Main Memory · ICNP 2018 |
Indexing and storage engines
learned index |
0.2 | 1 | 2023 | LeaFTL: A Learning-Based Flash Translation Layer for Solid-State Drives · ASPLOS (2) 2023 |
Interconnection networks and networks-on-chip › high-speed interconnect
PCIe interconnect |
0.2 | 1 | 2014 | DC express: shortest latency protocol for reading phase change memory over PCI express · FAST 2014 |
Memory systems › non-volatile memory
phase change memory |
0.2 | 1 | 2014 | DC express: shortest latency protocol for reading phase change memory over PCI express · FAST 2014 |
Distributed systems
consensus |
0.1 | 1 | 2018 | Consensus for Non-volatile Main Memory · ICNP 2018 |
Distributed systems › replication
replication and fault tolerance |
0.1 | 1 | 2018 | Consensus for Non-volatile Main Memory · ICNP 2018 |
Methods — techniques the papers use, named apart from their topics
learning-based optimization · 1.3software memory controller emulation · 0.3data replication · 0.3consensus protocol · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DISCO: Distributed Inference with Sparse CommunicationsabstractDeep neural networks (DNNs) have great potential to solve many real-world problems, but they usually require an extensive amount of computation and memory. It is of great difficulty to deploy a large DNN model to a single resource-limited device with small memory capacity. Distributed computing is a common approach to reduce single-node memory consumption and to accelerate the inference of DNN models. In this paper, we explore the "within-layer model parallelism", which distributes the inference of each layer into multiple nodes. In this way, the memory requirement can be distributed to many nodes, making it possible to use several edge devices to infer a large DNN model. Due to the dependency within each layer, data communications between nodes during this parallel inference can be a bottleneck when the communication bandwidth is limited. We propose a framework to train DNN models for Distributed Inference with Sparse Communications (DISCO). We convert the problem of selecting which subset of data to transmit between nodes into a model optimization problem, and derive models with both computation and communication reduction when each layer is inferred on multiple nodes. We show the benefit of the DISCO framework on a variety of CV tasks such as image classification, object detection, semantic segmentation, and image super resolution. The corresponding models include important DNN building blocks such as convolutions and transformers. For example, each layer of a ResNet-50 model can be distributively inferred across two nodes with 5x less data communications, almost half overall computations and less than half memory requirement for a single node, and achieve comparable accuracy to the original ResNet-50 model. Minghai Qin, Jaco Hofmann, Dejan Vucinic |
WACV | 4 |
| 2023 | LeaFTL: A Learning-Based Flash Translation Layer for Solid-State DrivesabstractIn modern solid-state drives (SSDs), the indexing of flash pages is a critical component in their storage controllers. It not only affects the data access performance, but also determines the efficiency of the precious in-device DRAM resource. A variety of address mapping schemes and optimizations have been proposed. However, most of them were developed with human-driven heuristics. Jinghan Sun, Shaobo Li 0005, Yunxin Sun 0001, Dejan Vucinic, Jian Huang 0006 |
ASPLOS (2) | 5 |
| 2019 | Garbage Collection Algorithms for Meta Data Updates in NAND FlashabstractGarbage collection (GC) is ubiquitously used to reclaim useful space during the data update process in NAND flash memories. Conventional GC algorithms keep track of a table storing the number of valid pages in each block and select ones with the smallest number of valid pages to erase. We notice that NAND flash not only stores user data but also keeps updating meta data frequently and one of the most important requirements of GC for meta data updates is latency predictability, which means the latency of GC should not only be low on average and on tail distributions, but also independent of workload, i.e., the sequence of meta data updates. It is also desirable to have all NAND pages/blocks written the same number of times to avoid any latency incurred by wear leveling. In this paper, we propose GC algorithms that do not require any look-up tables, which avoids the latency for reading the table and sorting the number of valid pages, and intrinsically enables all pages to be worn equally. Several GC algorithms are proposed to make trade-offs between over-provisioning and write amplification. We also provide sufficient and necessary conditions that the proposed GC algorithms will not encounter a deadlock, i.e., the algorithms can run continuously on all meta data update sequences without data loss. Minghai Qin, Robert Mateescu, Qingbo Wang, Cyril Guyot, Dejan Vucinic, Zvonimir Bandic |
ICC | 5 |
| 2018 | Consensus for Non-volatile Main MemoryabstractTraditionally, computer storage has been separated into a hierarchy based on response time, volatility, and cost of media. This tiering is undergoing a significant upheaval as a new breed of memory technologies, termed Storage Class Memories (SCM), now make it feasible to replace several tiers of the hierarchy with a single, cost-effective, uniform type of memory/storage. To make large-scale SCM deployments practical, however, memory system designers will first need to solve the problem of how to guard against unavoidable storage wear-out and failures-problems traditionally absent from "main memory" and handled by software at leisurely timescales in the domain of storage. In this paper, we propose a novel approach to providing fault tolerance in SCM-based main memory. Our key insight is to treat memory as a distributed storage system and rely on data replication and a consensus protocol to keep the replicas consistent. Separate memory instances store replicated copies of the data, and we use a programmable network interconnect to provide fast consensus between the memory instances. Our initial experiments using software memory controller emulation demonstrate reasonable overhead over local memory reads and show great promise as scalable main memory. Huynh Tu Dang, Jaco Hofmann, Marjan Radi, Dejan Vucinic, Robert Soulé, Fernando Pedone |
ICNP | 5 |
| 2015 | A Parallel and Pipelined Architecture for Accelerating Fingerprint Computation in High Throughput Data StoragesabstractRabin fingerprints are short tags for large objects that can be used in a wide range of applications, such as data deduplication, web querying, packet routing, and caching. We present a pipelined hardware architecture for computing Rabin fingerprints on data being transferred on a high throughput bus. The design conducts real-time fingerprinting with short latencies, and can be tuned for optimized clock rate with "split fresh" technique. A pipelined sampling logic selects fingerprints based on the Minwise theory and adds only a few clock cycles of latency before returning the final results. The design can be replicated to work in parallel for higher throughput data traffic. This architecture is implemented on a Xilinx Virtex-6 FPGA, and is tested on a storage prototyping platform. The implementation shows that the design can achieve clock rates above 300 MHz with an order of magnitude improvement in latency over prior software implementations, while consuming little hardware resource. The scheme is extensible to other types of fingerprints and CRC computations, and is readily applicable to primary storages and caches in hybrid storage systems. Qing Yang 0001, Qingbo Wang, Cyril Guyot, Ashwin Narasimha, Dejan Vucinic, Zvonimir Bandic |
FCCM | 6 |
| 2015 | Hardware accelerator for similarity based data dedupeabstractData deduplication has proven important in backup storage systems as large amount of identical or similar data chunks exist. Recent studies have shown the great potential of data deduplication in primary storage and storage caches. Deduplications in these environments require high speed processing not to drag down production performance. This paper presents a hardware accelerator for similarity based data deduplication. It implements three compute-intensive kernel modules to improve throughput and latency in dedupe systems: sketch computation for data blocks, index searching for reference block, and delta encoding over similar blocks. Adopting pipelined computation and parallel data lookup across multiple hardware modules, our HW design is capable of processing high throughput data traffic by working on multiple data units concurrently, thus enabling wire speed dedupe for data stream where similar blocks present. Using a PC host system connected to the FPGA-based accelerator through a PCIe Gen 2×4 interface, our experiments show that the similarity based data dedupe performs 30% better in data reduction ratio than conventional dedupe techniques that look at identical blocks only. By comparing the hardware implementation with its software counterpart, the experimental results show that our preliminary FPGA implementation with maximum clock speed of 250MHz achieves at least 6 times improvement in latency over the software implementation running on state-of-art servers. Qingbo Wang, Cyril Guyot, Ashwin Narasimha, Dejan Vucinic, Zvonimir Bandic, Qing Yang 0001 |
NAS | 5 |
| 2014 | DC express: shortest latency protocol for reading phase change memory over PCI express
Dejan Vucinic, Qingbo Wang, Cyril Guyot, Robert Mateescu, Filip Blagojevic, Luiz Franca-Neto, Damien Le Moal, Trevor Bunker, Jian Xu 0012, Steven Swanson, Zvonimir Bandic |
FAST | 1 |