VLDB 2026 Research / reviewers in the wild / expert
Michal Karczmarek
dblp:46/5869
· DBLP profile ↗
6ranked-venue papers
3as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 2 first-authorSystems, architecture and hardware · 3 · 1 first-authorTheory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 66% Embedded and real-time systems · 25% Processor architecture and microarchitecture · 10% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel programming models
stream programming |
0.1 | 2 | 2005 | Teleport messaging for distributed stream programs · PPoPP 2005 A stream compiler for communication-exposed architectures · ASPLOS 2002 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2005 | Teleport messaging for distributed stream programs · PPoPP 2005 |
Embedded and real-time systems › model-based design › dataflow modeling
synchronous dataflow |
0.1 | 1 | 2005 | Teleport messaging for distributed stream programs · PPoPP 2005 |
Compilers and program optimization
parallelizing compiler |
0.0 | 1 | 2002 | A stream compiler for communication-exposed architectures · ASPLOS 2002 |
Compilers and program optimization › domain-specific compilation
stream compiler |
0.0 | 1 | 2002 | A stream compiler for communication-exposed architectures · ASPLOS 2002 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2002 | A stream compiler for communication-exposed architectures · ASPLOS 2002 |
Processor architecture and microarchitecture
tiled architecture |
0.0 | 1 | 2002 | A stream compiler for communication-exposed architectures · ASPLOS 2002 |
Methods — techniques the papers use, named apart from their topics
stream dependence function · 0.1static scheduling · 0.1graph layout · 0.1fission and fusion transformations · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | A new synthesis procedure for atomic rules containing multi-cycle function blocksabstractA new method for hardware synthesis from atomic rules where rules can take unknown number of cycles is presented. Some complex functions, especially the ones involving data-dependent control, are more easily expressed as loops and take much less area when implemented as multi-cycle folded circuits. Multicycle rules also provide a high-level method for the designer to deal with the timing-closure problem by treating a long combinational path as a multicycle path. Our synthesis procedure uses minimal extra storage and executes all rules eagerly. It resets all those rules whose read sets are affected by a committing rule. It also makes use of rule reservations to avoid a short rule from resetting a long rule repeatedly. This technique automatically takes advantage of different timings of different conditional branches. Our syntax-directed synthesis procedure composes signals indicating when a computation is done and when inputs to a computation have changed. Preliminary results from our implementation look very promising. Michal Karczmarek, Arvind 0001, Muralidaran Vijayaraghavan |
MEMOCODE | 1 |
| 2008 | Synthesis from multi-cycle atomic actions as a solution to the timing closure problemabstractOne solution to the timing closure problem is to perform infrequent operations in more than one cycle. Despite simplicity of the solution statement, it is not easily considered because it requires changes in RTL, which, in turn, exacerbates the verification problem. We offer a timing closure solution guaranteed to preserve functional correctness of designs expressed using atomic actions or rules. We exploit the fact that the semantics of atomic actions are untimed, that is, the time to execute an action is not specified. The current hardware synthesis technique from atomic actions assumes that each rule takes one clock cycle to complete its computation. Consequently, the rule with the longest combinational path determines the clock cycle of the entire design, often leading to needlessly slow circuits. We present a synthesis procedure for a system where the combinational circuits embodied in a rule can take multiple cycles without changing the semantics of the original design. We also present preliminary results based on an experimental compiler which uses the Bluespec (BSV) compiler front end and generates Verilog. The results show that the clock speed and the performance of circuits can be improved substantially by allowing slow paths to complete over multiple cycles. Our technique is orthogonal to solutions based on multiple clock domains. Michal Karczmarek, Arvind 0001 |
ICCAD | 1 |
| 2005 | Teleport messaging for distributed stream programsabstractIn this paper, we develop a new language construct to address one of the pitfalls of parallel programming: precise handling of events across parallel components. The construct, termed teleport messaging, uses data dependences between components to provide a common notion of time in a parallel system. Our work is done in the context of the Synchronous Dataflow (SDF) model, in which computation is expressed as a graph of independent components (or actors) that communicate in regular patterns over data channels. We leverage the static properties of SDF to compute a stream dependence function, SDEP, that compactly describes the ordering constraints between actor executions.Teleport messaging utilizes SDEP to provide powerful and precise event handling. For example, an actor A can specify that an event should be processed by a downstream actor B as soon as B sees the "effects" of the current execution of A. We argue that teleport messaging improves readability and robustness over existing practices. We have implemented messaging as part of the StreamIt compiler, with a backend for a cluster of workstations. As teleport messaging exposes optimization opportunities to the compiler, it also results in a 49% performance improvement for a software radio benchmark. William Thies, Michal Karczmarek, Janis Sermulins, Rodric M. Rabbah, Saman P. Amarasinghe |
PPoPP | 2 |
| 2003 | Phased scheduling of stream programsabstractAs embedded DSP applications become more complex, it is increasingly important to provide high-level stream abstractions that can be compiled without sacrificing efficiency. In this paper, we describe scheduler support for StreamIt, a high-level language for signal processing applications. A StreamIt program consists of a set of autonomous filters that communicate with each other via FIFO queues. As in Synchronous Dataflow (SDF), the input and output rates of each filter are known at compile time. However, unlike SDF, the stream graph is represented using hierarchical structures, each of which has a single input and a single output.We describe a scheduling algorithm that leverages the structure of StreamIt to provide a flexible tradeoff between code size and buffer size. The algorithm describes the execution of each hierarchical unit as a set of phases. A complete cycle through the phases represents a single steady-state execution. By varying the granularity of a phase, our algorithm provides a continuum between single appearance schedules and minimum latency schedules. We demonstrate that a minimal latency schedule is effective in decreasing buffer requirements for some applications, while the phased representation mitigates the associated increase in code size. Michal Karczmarek, William Thies, Saman P. Amarasinghe |
LCTES | 1 |
| 2002 | A stream compiler for communication-exposed architecturesabstractWith the increasing miniaturization of transistors, wire delays are becoming a dominant factor in microprocessor performance. To address this issue, a number of emerging architectures contain replicated processing units with software-exposed communication between one unit and another (e.g., Raw, SmartMemories, TRIPS). However, for their use to be widespread, it will be necessary to develop compiler technology that enables a portable, high-level language to execute efficiently across a range of wire-exposed architectures.In this paper, we describe our compiler for StreamIt: a high-level, architecture-independent language for streaming applications. We focus on our backend for the Raw processor. Though StreamIt exposes the parallelism and communication patterns of stream programs, some analysis is needed to adapt a stream program to a software-exposed processor. We describe a partitioning algorithm that employs fission and fusion transformations to adjust the granularity of a stream graph, a layout algorithm that maps a stream graph to a given network topology, and a scheduling strategy that generates a fine-grained static communication pattern for each computational element.We have implemented a fully functional compiler that parallelizes StreamIt applications for Raw, including several load-balancing transformations. Using the cycle-accurate Raw simulator, we demonstrate that the StreamIt compiler can automatically map a high-level stream abstraction to Raw without losing performance. We consider this work to be a first step towards a portable programming model for communication-exposed architectures. Michael I. Gordon, William Thies, Michal Karczmarek, Jasper Lin, Ali S. Meli, Andrew A. Lamb, Chris Leger, Jeremy Wong, Henry Hoffmann, David Maze, Saman P. Amarasinghe |
ASPLOS | 3 |
| 2002 | StreamIt: A Language for Streaming Applications
William Thies, Michal Karczmarek, Saman P. Amarasinghe |
CC | 2 |