EDBT 2026 Demo / reviewers in the wild / expert
Daniel G. Waddington
dblp:119/1659 · also Dan G. Waddington, Daniel Giles Waddington
· DBLP profile ↗
18ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0001-8758-910XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speeding up Model Loading with FastsafetensorsabstractThe rapid increases in model parameter sizes introduces new challenges in pre-trained model loading. Currently, machine learning code often deserializes each parameter as a tensor object in host memory before copying it to device memory. We found that this approach underutilized storage throughput and significantly slowed down loading large models with a widely-used model file formats, safetensors. In this work, we present fastsafetensors, a Python library designed to optimize the deserialization of tensors in safetensors files. Our approach first copies groups of on-disk parameters to device memory, where they are directly instantiated as tensor objects. This design enables further optimization in low-level I/O and high-level tensor preprocessing, including parallelized copying, peer-to-peer DMA, and GPU offloading. Experimental results show performance improvements of 4.8x to 7.5x in loading models such as Llama (7, 13, and 70 billion parameters), Falcon (40 billion parameters), and the Bloom (176 billion parameters). Takeshi Yoshimura, Tatsuhiro Chiba, Manish Sethi, Daniel G. Waddington, Swaminathan Sundararaman |
CLOUD | 4 |
| 2025 | DPUF: DPU-accelerated Near-storage Secure FilteringabstractQuerying data stored in cloud object stores often leads to network bottlenecks, particularly when large datasets need to be transferred over wide area networks (WANs) for processing. Encryption further complicates this challenge by requiring entire encrypted objects to be fetched from the object store before analysis. To address this, we push down filtering and perform secure computing near storage using a Data Processing Unit (DPU) integrated into the cloud server. Narangerelt Batsoyol, Daniel G. Waddington, Swaminathan Sundararaman, Steven Swanson |
SYSTOR | 2 |
| 2024 | Dictionary Based Cache Line CompressionabstractActive-standby mechanisms for VM high-availability demand frequent synchronization of memory and CPU state, involving the identification and transfer of "dirty" memory pages to a standby target. Building upon the granularity offered by CXL-enabled memory devices, as discussed by Waddington et al. [21], this paper proposes a dictionary-based compression method operating on 64-byte cache lines to minimize snapshot volume and synchronization latency. The method aims to transmit only necessary information required to reconstruct the memory state at the standby machine, augmented by byte grouping and cache-line partitioning techniques. We assess the compression benefits on memory access patterns across 20 benchmarks snapshots and compare our approach to standard off-the-shelf compression methods. Our findings reveal significant improvements across nearly all benchmarks, with some experiencing over a twofold enhancement compared to standard compression, while others show more moderate gains. We conduct an in-depth experimental analysis on the contribution of each method and examine the nature of the benchmarks. We ascertain that the repeating nature of cache lines across snapshots (caused by transient memory changes) and their concise representation contributes most to the size reduction, accounting for 92% of the gains. Our work paves the way for further reduction in the data transferred to standby machines, thereby enhancing VM high-availability and reducing synchronization latency. Sarel Cohen, Dalit Naor, Daniel G. Waddington, Moshe Hershcovitch |
HotStorage | 4 |
| 2023 | Fast Feature Selection with Fairness ConstraintsabstractWe study the fundamental problem of selecting optimal features for model construction. This problem is computationally challenging on large datasets, even with the use of greedy algorithm variants. To address this challenge, we extend the adaptive query model, recently proposed for the greedy forward selection for submodular functions, to the faster paradigm of Orthogonal Matching Pursuit for non-submodular functions. The proposed algorithm achieves exponentially fast parallel run time in the adaptive query model, scaling much better than prior work. Furthermore, our extension allows the use of downward-closed constraints, which can be used to encode certain fairness criteria into the feature selection process. We prove strong approximation guarantees for the algorithm based on standard assumptions. These guarantees are applicable to many parametric models, including Generalized Linear Models. Finally, we demonstrate empirically that the proposed algorithm competes favorably with state-of-the-art techniques for feature selection, on real-world and synthetic datasets. Francesco Quinzan, Rajiv Khanna, Moshe Hershcovitch, Sarel Cohen, Daniel G. Waddington, Tobias Friedrich 0001, Michael W. Mahoney |
AISTATS | 5 |
| 2023 | Cache Line Deltas CompressionabstractSynchronization of replicated data and program state is an essential aspect of application fault-tolerance. Current solutions use virtual memory mapping to identify page writes and replicate them at the destination. This approach has limitations because the granularity is restricted to a minimum of 4KiB per page, which may result in more data being replicated. Motivated by the emerging CXL hardware, we expand on the work Waddington, et al. [SoCC 22] by evaluating popular compression algorithms on VM snapshot data at cache line granularity. We measure the compression ratio vs. the compression time and present our conclusions. Sarel Cohen, Dalit Naor, Daniel G. Waddington, Moshe Hershcovitch |
SYSTOR | 4 |
| 2022 | A case for using cache line deltas for high frequency VM snapshottingabstractActive-standby schemes for Virtual Machine (VM) high availability require periodic synchronization of memory and CPU state. The most common approach to synchronization is to use page tables and software to identify "dirty" memory pages at the source and in turn copy them to the target via a network or interconnect. However, this approach results in significanct page table traversal and data copying overhead, resulting in considerable VM downtime. A principal contributor to this overhead is that many applications using this approach incur data copy-amplification as a result of copying more data than is necessary; this arises because of the processor's virtual memory system design in which memory pages are 4KiB or larger. Daniel G. Waddington, Moshe Hershcovitch, Swaminathan Sundararaman, Clem Dickey |
SoCC | 1 |
| 2022 | Elastic Indexes: Dynamic Space vs. Query Efficiency Tuning for In-Memory Database Indexing
Moshe Hershcovitch, Artem Khyzha, Daniel G. Waddington, Adam Morrison 0001 |
EDBT | 3 |
| 2022 | System-level crash safe sorting on persistent memoryabstractSorting is a fundamental operation in software systems. An example for that is a prepossessing phase before executing analytics operations. Omri Arad, Yoav Ben Shimon, Ron Zadicario, Daniel G. Waddington, Moshe Hershcovitch, Adam Morrison 0001 |
SYSTOR | 4 |
| 2021 | PyMM: Heterogeneous Memory Programming for Python Data ScienceabstractWhile persistent memory (PMEM) is a promising technology, leveraging it with legacy applications is non-trivial. This is primarily because legacy applications assume all memory is volatile and there is no notion of crash-consistency or state recovery. As new types of persistent and intelligent memory emerge, propelled by the CXL standard, the problem of integration and adoption remains. Daniel G. Waddington, Moshe Hershcovitch, Clem Dickey |
PLOS@SOSP | 1 |
| 2020 | Evaluating Intel 3D-Xpoint NVDIMM Persistent Memory in the Context of a Key-Value StoreabstractIntel's 3D-Xpoint NVDIMM product1 is now in general availability. This technology, herein termed Optane-PM, is a first-in-breed persistent memory technology, designed for the enterprise storage domain. Optane-PM is attached to the memory-bus and is load-store addressable. It provides higher capacity and density than DRAM, while also enabling non-volatility - data is persistent and durable across power-cycles. This paper presents a detailed evaluation of adopting Optane-PM in key-value stores. To achieve this goal, we designed and implemented a high-performance, memory-centric key-value store (MCKVS), that directly leverages persistent memory for data and meta-data storage. Multiple storage back-ends, with different software stacks, are used to provide a comparative analysis and help us understand how Optane-PM is positioned on the performance landscape. We also conducted comparisons against popular in-memory KV stores based on DRAM. Daniel G. Waddington, Clem Dickey, Luna Xu, Travis Janssen, Jantz Tran, Kshitij A. Doshi |
ISPASS | 1 |
| 2017 | A single-node datastore for high-velocity multidimensional sensor dataabstractSources of multidimensional data are becoming more prevalent, partly due to the rise of the Internet of Things (IoT), and so is the need to ingest and analyze data streams at rates higher than before. Some industrial IoT applications require ingesting millions of records per second, while processing queries on recently ingested and historical data. Unfortunately, existing database systems targeting multidimensional data exhibit low per-node ingestion performance, and even if they can scale horizontally in distributed settings, they require large number of nodes to meet such ingest demands. For this reason, in this paper we present a single-node datastore able to ingest multidimensional sensor data at very high rates. Its design centers around a two-level indexing structure, wherein the global index is an in-memory R*-tree and the local indices are serialized kd-trees. This study is confined to records with numerical indexing fields and range queries, and covers ingest throughput, query response time, and storage footprint. We show that the adopted design streamlines data ingestion and offers ingress rates two orders of magnitude higher than those of a selection of open-source database systems, namely Percona Server, SQLite, and Druid. Our prototype also reports query response times comparable to or better than those of Percona Server and Druid, and compares favorably in terms of storage footprint. We believe the experience reported here is valuable to researchers and practitioners interested in building database systems for high-velocity multidimensional sensor data. Juan A. Colmenares, Reza Dorrigiv, Daniel G. Waddington |
IEEE BigData | 3 |
| 2015 | Techniques for fast and scalable time series traffic generationabstractMany IoT applications ingest and process time series data with emphasis on 5Vs (Volume, Velocity, Variety, Value and Veracity). To design and test such systems, it is desirable to have a high-performance traffic generator specifically designed for time series data, preferably using archived data to create a truly realistic workload. However, most existing traffic generator tools either are designed for generic network applications, or only produce synthetic data based on certain time series models. In addition, few have raised their performance bar to millions-packets-per-second level with minimum time violations. In this paper, we design, implement and evaluate a highly efficient and scalable time series traffic generator for IoT applications. Our traffic generator stands out in the following four aspects: 1) it generates time-conforming packets based on high-fidelity reproduction of archived time series data; 2) it leverages an open-source Linux Exokernel middleware and a customized userspace network subsystem; 3) it includes a scalable 10G network card driver and uses "absolute" zero-copy in stack processing; and 4) it has an efficient and scalable application-level software architecture and threading model. We have conducted extensive experiments on both a quad-core Intel workstation and a 20-core Intel server equipped with Intel X540 10G network cards and Samsung's NVMe SSDs. Compared with a stock Linux baseline and a traditional mmap-based file I/O approach, we observe that our traffic generator significantly outperforms other alternatives in terms of throughput (10X), scalability (3.6X) and time violations (46.2X). Jilong Kuang, Daniel G. Waddington, Changhui Lin |
IEEE BigData | 2 |
| 2013 | Towards a Scalable Microkernel Personality for Multicore Processors
Jilong Kuang, Daniel G. Waddington, Chen Tian 0005 |
Euro-Par | 2 |
| 2012 | A Scalable Physical Memory Allocation Scheme for L4 MicrokernelabstractL4 microkernel family has become very successful on mobile devices. However, with the rapid shift from uniprocessor to multicore and manycore processor, many critical OS functions including physical memory allocator (PMA) must be re-designed in order to achieve better system throughput. While research and engineering efforts have been made for PMA in monolithic kernels such as Linux, not much work can be found for L4 microkernels. Due to the the design difference, the PMA in L4 microkernels is part of user level page fault handler (a.k.a. pager), which is executed as a stand-alone server in the least privilege mode. Memory allocation and free requests are handled through inter-process communication (IPC) rather than normal system or kernel function calls. In this work, we first study the scalability issue of the PMA implementation in L4 microkernels, and propose our solution in the context of Fiasco.OC, a state-of-the-art L4 microkernel implementation. We also discuss how to leverage the L4 microkernel design advantages to implement a PMA with more advanced features, such as load balancing, customizability and NUMA-awareness. Finally, we conduct experiments to verify the scalability result of our solution. The experiment is conducted on a 48-core AMD magny-cours server. Chen Tian 0005, Daniel G. Waddington, Jilong Kuang |
COMPSAC | 2 |
| 2012 | Load Balancing Aware Real-Time Task Partitioning in Multicore SystemsabstractReal-time applications of future IT will continue to drive the demand for performance scaling in devices ranging from sensors to servers. Parallel processing in the form of mul-ticore and manycore architectures will also continue to be the principal route to unleashing next generation performance capabilities. To fully exploit multicore processors, real-time applications are expected to provide a large degree of parallel-ism, where real-time tasks can utilize multiple cores at the same time. Guaranteeing real-time performance, while making efficient use of multicore resources, requires a scheduling method that offers both high schedulability and effective load balancing. Many existing real-time scheduling methods for multicore systems focus on schedulability or load balancing, but not both -- each coming at the expense of the other. In this work we develop an efficient scheduling algorithm that not only guarantees real-time performance but also demonstrates effective distribution of tasks across cores. Experimental re-sults show that our method significantly outperforms state-of-the-art approaches in terms of load balancing while still providing good schedulability. We also show the benefits with respect to energy reduction that result from balanced load. Jaeyeon Kang, Daniel G. Waddington |
RTCSA | 2 |
| 2007 | High-fidelity C/C++ code transformation
Daniel G. Waddington |
Sci. Comput. Program. | 1 |
| 2003 | Topology Inference in the Presence of Anonymous RoutersabstractMany topology discovery systems rely on traceroute to discover path information in public networks. However, for some routers, traceroute detects their existence but not their address; we term such routers anonymous routers. This paper considers the problem of inferring the network topology in the presence of anonymous routers. We illustrate how obvious approaches to handle anonymous routers lead to incomplete, inflated, or inaccurate topologies. We formalize the topology inference problem and show that producing both exact and approximate solutions is intractable. Two heuristics are proposed and evaluated through simulation. These heuristics have been used to infer the topology of the 6Bone, and could be incorporated into existing tools to infer more comprehensive and accurate topologies. Ramesh Viswanathan, Fangzhe Chang, Daniel G. Waddington |
INFOCOM | 4 |
| 1997 | A Distributed Multimedia Component ArchitectureabstractA new framework is required for the consistent and coordinated construction and configuration of multimedia components across heterogeneous distributed object environments. The paper presents a distributed multimedia component architecture which extends beyond the current well developed distributed object models, such as the Common Object Request Broker Architecture (CORBA) and the Microsoft Distributed Component Object Model (DCOM) models, to support additional mechanisms and abstractions for continuous media stream interactions aimed particularly at distributed multimedia applications. This framework additionally includes distributed resource management and dynamic Quality of Service (QoS) monitoring and adaptation in order to support the implementation of complex distributed multimedia applications over heterogeneous networks and end-systems. The multimedia component architecture (MCA) presented is being used as the middleware component of a comprehensive resource management architecture for distributed applications (Waddington, 1997). Daniel G. Waddington, Geoff Coulson |
EDOC | 1 |