EDBT 2026 Demo / reviewers in the wild / expert
Martha A. Kim
dblp:29/10949
· DBLP profile ↗
24ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0001-6243-5753ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 3 since 2021Software engineering, systems software and programming languages · 12 · 1 since 2021Theory of computation · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automated Power Domain Insertion and Control in Dataflow CircuitsabstractEnergy efficiency is a key concern in circuit design. With the end of Dennard scaling, voltages no longer scale with transistor size, making full-capacity operation unsustainable due to heat and power limits. Power gating offers a solution, but controlling independent power domains is challenging. Large domains are manageable but inefficient; small domains are efficient but require intricate control. Martha Barker, Mark Santolucito, Stephen A. Edwards, Martha A. Kim |
MEMOCODE | 4 |
| 2025 | MINDFUL: Safe, Implantable, Large-Scale Brain-Computer Interfaces from a System-Level Design Perspective
Guy Eichler, Yatin Gilhotra, Nanyu Zeng, Martha A. Kim, Kenneth L. Shepard, Luca P. Carloni |
MICRO | 4 |
| 2022 | Synthesized Garbage Collection for FPGA AcceleratorsabstractSpeed and ease of accelerator design is a growing need. High level programming languages have provided significant gains in the software world, but lag in the hardware realm. We present a hardware implementation of a garbage collector, which automates memory management, one of the major conveniences of modern software languages. Our garbage collector is integrated with a Haskell to hardware HLS flow. The collector runs concurrently with the application, using already-idle memory slots to do its work with little to no impact on performance. To achieve this, our collector exploits rapid synchronization that is straightforward in hardware but very difficult in software. With this synchronization, the collector is able to safely pop in and out of very fine pockets of idleness on a cycle by cycle, heap by heap basis. In most cases our collector incurred negligible overhead, and slowed the application only when the heaps were so tiny that the application was unable to proceed with its operation until the collector had freed more space. Our experiments further show that our concurrent collector performs best under an eager collection policy that collects garbage well-before the application exhausts the available memory. Although this eager strategy performs more collection operations than strictly necessary, the application never pauses and the collector operates entirely in the background. Martha Barker, Stephen A. Edwards, Martha A. Kim |
FPGA | 3 |
| 2022 | Synthesized In-BramGarbage Collection for Accelerators with Immutable MemoryabstractSpeed and ease of accelerator design is a growing need. High level programming languages have provided significant gains in the software world, but lag for hardware. We present a hardware implementation of a garbage collector that automates memory management, one of the major conveniences of modern software languages. Our garbage collector runs concurrently with the application it serves, using already-idle memory slots to do its work with little to no impact on performance. To achieve this, our collector exploits rapid synchronization that is straightforward in hardware but difficult in software. This synchronization enables the collector to interleave with fine pockets of idleness on a per-cycle, per-heap basis. Our collector typically incurred negligible overhead, only slowing the application to wait for collection to free memory when the heaps were so small that the collector could not keep pace with allocation. We also found that our concurrent collector performs best under an eager collection policy that collects garbage well before the application exhausts the available memory. Although this eager strategy performs more collection operations than strictly necessary, the application never pauses because our collector operates entirely in the background. Martha Barker, Stephen A. Edwards, Martha A. Kim |
FPL | 3 |
| 2019 | Master of none acceleration: a comparison of accelerator architectures for analytical query processingabstractHardware accelerators are one promising solution to contend with the end of Dennard scaling and the slowdown of Moore's law. For mature workloads that are regular and have high compute per byte, hardening an application into one or more hardware modules is a standard approach. However, for some applications, we find that a programmable homogeneous architecture is preferable. Andrea Lottarini, Joao Pedro Cerqueira, Thomas J. Repetti, Stephen A. Edwards, Kenneth A. Ross, Mingoo Seok, Martha A. Kim |
ISCA | 7 |
| 2019 | Compositional Dataflow CircuitsabstractWe present a technique for implementing dataflow networks as compositional hardware circuits. We first define an abstract dataflow model with unbounded buffers that supports data-dependent blocks (mux, demux, and nondeterministic merge); we then show how to faithfully implement such networks with bounded buffers and handshaking. Handshaking admits compositionality: our circuits can be connected with or without buffers, and combinational cycles arise only from a completely unbuffered cycle. While bounding buffer sizes can cause the system to deadlock prematurely, the system is guaranteed to produce the same, correct, data before then. Thus, unless the system deadlocks, inserting or removing buffers only affects its performance. We demonstrate how this enables design space to be explored. Stephen A. Edwards, Richard Townsend, Martha Barker, Martha A. Kim |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2018 | vbench: Benchmarking Video Transcoding in the CloudabstractThis paper presents vbench, a publicly available benchmark for cloud video services. We are the first study, to the best of our knowledge, to characterize the emerging video-as-a-service workload. Unlike prior video processing benchmarks, vbench's videos are algorithmically selected to represent a large commercial corpus of millions of videos. Reflecting the complex infrastructure that processes and hosts these videos, vbench includes carefully constructed metrics and baselines. The combination of validated corpus, baselines, and metrics reveal nuanced tradeoffs between speed, quality, and compression. We demonstrate the importance of video selection with a microarchitectural study of cache, branch, and SIMD behavior. vbench reveals trends from the commercial corpus that are not visible in other video corpuses. Our experiments with GPUs under vbench's scoring scenarios reveal that context is critical: GPUs are well suited for live-streaming, while for video-on-demand shift costs from compute to storage and network. Counterintuitively, they are not viable for popular videos, for which highly compressed, high quality copies are required. We instead find that popular videos are currently well-served by the current trajectory of software encoders. Andrea Lottarini, Alex Ramírez, Joel Coburn, Martha A. Kim, Parthasarathy Ranganathan, Daniel Stodolsky, Mark Wachsler |
ASPLOS | 4 |
| 2017 | From functional programs to pipelined dataflow circuits
Richard Townsend, Martha A. Kim, Stephen A. Edwards |
CC | 2 |
| 2017 | Network Synthesis for Database Processing UnitsabstractWe explore on-chip network topologies for the Q100, an analytic query accelerator for relational databases. In such data-centric accelerators, interconnects play a critical role by moving large volumes of data. In this paper we show that various interconnect topologies can trade a factor of 2.5x in performance for 3.3x area. Moreover, standard topologies (e.g., ring or mesh) are not optimal. Andrea Lottarini, Stephen A. Edwards, Kenneth A. Ross, Martha A. Kim |
DAC | 4 |
| 2017 | Deadlock-free joins in DB-mesh, an asynchronous systolic array acceleratorabstractPrevious database accelerator proposals such as the Q100 provide a fixed set of database operators, chosen to support a target query workload. Some queries may not be well-supported by a fixed accelerator, typically because they need more resources/operators of a particular kind than the accelerator provides. By Amdahl's law, these queries become relatively more expensive as they are not fully accelerated. We propose a second-level accelerator, DB-Mesh, to take up some of this workload. DB-Mesh is an asynchronous systolic array that is more generic than the Q100, and can be configured to run a variety of operators with configurable parameters such as record widths. We demonstrate DB-Mesh applied to nested loops joins, an operator that is not directly supported on the Q100. We show that a naïve implementation has the potential for deadlock, and show how to avoid deadlock with a careful design. We also demonstrate how the data flow policy used in the array influences system throughput. Bingyi Cao, Kenneth A. Ross, Stephen A. Edwards, Martha A. Kim |
DaMoN | 4 |
| 2017 | Hotspot monitoring and Temperature Estimation with miniature on-chip temperature sensorsabstractThis paper presents analysis and evaluation of the impact of size and voltage scalability of on-chip temperature sensor on the accuracy of hotspot monitoring and temperature estimation in dynamic thermal management of high performance microprocessors. The analysis is based on both the layout level and the system level across state-of-the-art sensors in terms of accuracy, voltage-scalability, and silicon footprint. Our analysis shows that a sensor having compact footprint and good voltage scalability can be placed on exact hotspot locations, typically among digital cells, significantly improving accuracy in tracking hotspots and estimating temperature of microarchitecture blocks, as compared to two other sensors that have higher sensor-circuit accuracy, large footprint and little voltage scalability limiting flexible placement. Pavan Kumar Chundi, Yini Zhou, Martha A. Kim, Eren Kursun, Mingoo Seok |
ISLPED | 3 |
| 2017 | Compositional dataflow circuitsabstractWe present a technique for implementing dataflow networks as compositional hardware circuits. We first define an abstract dataflow model with unbounded buffers that supports data-dependent blocks (mux, demux, and nondeterministic merge); we then show how to faithfully implement such networks with bounded buffers and handshaking. Handshaking admits compositionality: our circuits can be connected with or without buffers and still compute the same function without introducing spurious combinational cycles. As such, inserting or removing buffers affects the performance but not the functionality of our networks, which we demonstrate through experiments that show how design space can be explored. Stephen A. Edwards, Richard Townsend, Martha A. Kim |
MEMOCODE | 3 |
| 2017 | Pipelining a triggered processing elementabstractProgrammable spatial architectures composed of ensembles of autonomous fixed-ISA processing elements offer a compelling design point between the flexibility of an FPGA and the compute density of a GPU or shared-memory many-core. The design regularity of spatial architectures demands examination of the processing element microarchitecture early in the design process to optimize overall efficiency. Thomas J. Repetti, Joao Pedro Cerqueira, Martha A. Kim, Mingoo Seok |
MICRO | 3 |
| 2016 | NRG-loops: adjusting power from within applicationsabstractNRG-Loops are source-level abstractions that allow an application to dynamically manage its power and energy through adjustments to functionality, performance, and accuracy. The adjustments, which come in the form of truncated, adapted, or perforated loops, are conditionally enabled as runtime power and energy constraints dictate. NRG-Loops are portable across different hardware platforms and operating systems and are complementary to existing system-level efficiency techniques, such as DVFS and idle states. Using a prototype C library supported by commodity hardware energy meters (and with no modifications to the compiler or operating system), this paper demonstrates four NRG-Loop applications that in 2-6 lines of source code changes can save up to 55% power and 90% energy, resulting in up to 12X better energy efficiency than system-level techniques Melanie Kambadur, Martha A. Kim |
CGO | 2 |
| 2015 | Implementing latency-insensitive dataflow blocksabstractTo simplify the implementation of dataflow systems in hardware, we present a technique for designing latency- insensitive dataflow blocks. We provide buffering with backpressure, resulting in blocks that compose into deep, high-speed pipelines without introducing long combinational paths. Our input and output buffers are easy to assemble into simple unit- rate dataflow blocks, arbiters, and blocks for Kahn networks. We prove the correctness of our buffers, illustrate how they can be used to assemble arbitrary dataflow blocks, discuss pitfalls, and present experimental results that suggest our pipelines can operate at a high clock rate independent of length. Bingyi Cao, Kenneth A. Ross, Martha A. Kim, Stephen A. Edwards |
MEMOCODE | 3 |
| 2014 | Q100: the architecture and design of a database processing unitabstractIn this paper, we propose Database Processing Units, or DPUs, a class of domain-specific database processors that can efficiently handle database applications. As a proof of concept, we present the instruction set architecture, microarchitecture, and hardware implementation of one DPU, called Q100. The Q100 has a collection of heterogeneous ASIC tiles that process relational tables and columns quickly and energy-efficiently. The architecture uses coarse grained in- structions that manipulate streams of data, thereby maximizing pipeline and data parallelism, and minimizing the need to time multiplex the accelerator tiles and spill inter- mediate results to memory. This work explores a Q100 de- sign space of 150 configurations, selecting three for further analysis: a small, power-conscious implementation, a high- performance implementation, and a balanced design that maximizes performance per Watt. We then demonstrate that the power-conscious Q100 handles the TPC-H queries with three orders of magnitude less energy than a state of the art software DBMS, while the performance-oriented design out- performs the same DBMS by 70X. Lisa Wu Wills, Andrea Lottarini, Timothy K. Paine, Martha A. Kim, Kenneth A. Ross |
ASPLOS | 4 |
| 2014 | ParaShares: Finding the Important Basic Blocks in Multithreaded Programs
Melanie Kambadur, Kui Tang, Martha A. Kim |
Euro-Par | 3 |
| 2014 | An experimental survey of energy management across the stackabstractModern demand for energy-efficient computation has spurred research at all levels of the stack, from devices to microarchitecture, operating systems, compilers, and languages. Unfortunately, this breadth has resulted in a disjointed space, with technologies at different levels of the system stack rarely compared, let alone coordinated. Melanie Kambadur, Martha A. Kim |
OOPSLA | 2 |
| 2014 | Energy Analysis of Hardware and Software Range PartitioningabstractData partitioning is a critical operation for manipulating large datasets because it subdivides tasks into pieces that are more amenable to efficient processing. It is often the limiting factor in database performance and represents a significant fraction of the overall runtime of large data queries. This article measures the performance and energy of state-of-the-art software partitioners, and describes and evaluates a hardware range partitioner that further improves efficiency. The software implementation is broken into two phases, allowing separate analysis of the partition function computation and data shuffling costs. Although range partitioning is commonly thought to be more expensive than simpler strategies such as hash partitioning, our measurements indicate that careful data movement and optimization of the partition function can allow it to approach the throughput and energy consumption of hash or radix partitioning. For further acceleration, we describe a hardware range partitioner, or HARP, a streaming framework that offers a seamless execution environment for this and other streaming accelerators, and a detailed analysis of a 32nm physical design that matches the throughput of four to eight software threads while consuming just 6.9% of the area and 4.3% of the power of a Xeon core in the same technology generation. Lisa Wu Wills, Orestis Polychroniou, Raymond J. Barker, Martha A. Kim, Kenneth A. Ross |
ACM Trans. Comput. Syst. | 4 |
| 2013 | Navigating big data with high-throughput, energy-efficient data partitioningabstractThe global pool of data is growing at 2.5 quintillion bytes per day, with 90% of it produced in the last two years alone [24]. There is no doubt the era of big data has arrived. This paper explores targeted deployment of hardware accelerators to improve the throughput and energy efficiency of large-scale data processing. In particular, data partitioning is a critical operation for manipulating large data sets. It is often the limiting factor in database performance and represents a significant fraction of the overall runtime of large data queries. Lisa Wu Wills, Raymond J. Barker, Martha A. Kim, Kenneth A. Ross |
ISCA | 3 |
| 2013 | Parallel scaling properties from a basic block viewabstractAs software scalability lags behind hardware parallelism, understanding scaling behavior is more important than ever. This paper demonstrates how to use Parallel Block Vector (PBV) profiles to measure the scaling properties of multithreaded programs from a new perspective: the basic block's view. Through this lens, we guide users through quick and simple methods to produce high-resolution application scaling analyses. This method requires no manual program modification, new hardware, or lengthy simulations, and captures the impact of architecture, operating systems, threading models, and inputs. We apply these techniques to a set of parallel benchmarks, and, as an example, demonstrate that when it comes to scaling, functions in an application do not behave monolithically. Melanie Kambadur, Kui Tang, Joshua Lopez, Martha A. Kim |
SIGMETRICS | 4 |
| 2012 | Harmony: Collection and analysis of parallel block vectorsabstractEfficient execution of well-parallelized applications is central to performance in the multicore era. Program analysis tools support the hardware and software sides of this effort by exposing relevant features of multithreaded applications. This paper describes parallel block vectors, which uncover previously unseen characteristics of parallel programs. Parallel block vectors provide block execution profiles per concurrency phase (e.g., the block execution profile of all serial regions of a program). This information provides a direct and fine-grained mapping between an application's runtime parallel phases and the static code that makes up those phases. This paper also demonstrates how to collect parallel block vectors with minimal application perturbation using Harmony. Harmony is an instrumentation pass for the LLVM compiler that introduces just 16-21% overhead on average across eight Parsec benchmarks. We apply parallel block vectors to uncover several novel insights about parallel applications with direct consequences for architectural design. First, that the serial and parallel phases of execution used in Amdahl's Law are often composed of many of the same basic blocks. Second, that program features, such as instruction mix, vary based on the degree of parallelism, with serial phases in particular displaying different instruction mixes from the program as a whole. Third, that dynamic execution frequencies do not necessarily correlate with a block's parallelism. Melanie Kambadur, Kui Tang, Martha A. Kim |
ISCA | 3 |
| 2012 | Measuring interference between live datacenter applicationsabstractApplication interference is prevalent in datacenters due to contention over shared hardware resources. Unfortunately, understanding interference in live datacenters is more difficult than in controlled environments or on simpler architectures. Most approaches to mitigating interference rely on data that cannot be collected efficiently in a production environment. This work exposes eight specific complexities of live datacenters that constrain measurement of interference. It then introduces new, generic measurement techniques for analyzing interference in the face of these challenges and restrictions. We use the measurement techniques to conduct the first large-scale study of application interference in live production datacenter workloads. Data is measured across 1000 12-core Google servers observed to be running 1102 unique applications. Finally, our work identifies several opportunities to improve performance that use only the available data; these opportunities are applicable to any datacenter. Melanie Kambadur, Tipp Moseley, Rick Hank, Martha A. Kim |
SC | 4 |
| 2011 | Retinal Oximetry Based on Nonsimultaneous Image Acquisition Using a Conventional Fundus CameraabstractTo measure the retinal arteriole and venule oxygen saturation (SO(2)) using a conventional fundus camera, retinal oximetry based on nonsimultaneous image acquisition was developed and evaluated. Two retinal images were sequentially acquired using a conventional fundus camera with two bandpass filters (568 nm: isobestic, 600 nm: nonisobestic wavelength), one after another, instead of a built-in green filter. The images were registered to compensate for the differences caused by eye movements during the image acquisition. Retinal SO(2) was measured using two wavelength oximetry. To evaluate sensitivity of the proposed method, SO(2) in the arterioles and venules before and after inhalation of 100% O(2) were compared, respectively, in 11 healthy subjects. After inhalation of 100% O(2), SO(2) increased from 96.0 ±6.0% to 98.8% ±7.1% in the arterioles (p=0.002) and from 54.0 ±8.0% to 66.7% ±7.2% in the venules (p=0.005) (paired t-test, n=11). Reproducibility of the method was 2.6% and 5.2% in the arterioles and venules, respectively (average standard deviation of five measurements, n=11). Sun Kwon Kim, Dong Myung Kim, Min Hee Suh, Martha A. Kim, Hee Chan Kim |
IEEE Trans. Medical Imaging | 4 |