EDBT 2026 Demo / reviewers in the wild / expert
Tiberiu Chelcea
dblp:33/5391
· DBLP profile ↗
9ranked-venue papers
5as first author
0since 2021 · last 2007
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Integrated circuit design · 33% Electronic design automation · 25% Processor architecture and microarchitecture · 15% |
Topics — the 10 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
high-level synthesis |
0.1 | 2 | 2007 | Hardware compilation of application-specific memory-access interconnect · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 Global Critical Path: A Tool for System-Level Timing Analysis · DAC 2007 |
Integrated circuit design
asynchronous circuit design |
0.1 | 1 | 2007 | Self-Resetting Latches for Asynchronous Micro-Pipelines · DAC 2007 |
Integrated circuit design
low-power circuit design |
0.1 | 1 | 2007 | Self-Resetting Latches for Asynchronous Micro-Pipelines · DAC 2007 |
Integrated circuit design › asynchronous circuit design
micropipeline |
0.1 | 1 | 2007 | Self-Resetting Latches for Asynchronous Micro-Pipelines · DAC 2007 |
Reconfigurable computing and FPGAs › reconfigurable architecture
reconfigurable fabric |
0.1 | 1 | 2006 | Tartan: evaluating spatial computation for whole program execution · ASPLOS 2006 |
Electronic design automation › logic synthesis
asynchronous circuit synthesis |
0.0 | 1 | 2002 | Resynthesis and peephole transformations for the optimization of large-scale asynchronous systems · DAC 2002 |
Interconnection networks and networks-on-chip › flow control
latency-insensitive protocols |
0.0 | 1 | 2001 | Robust Interfaces for Mixed-Timing Systems with Application to Latency-Insensitive Protocols · DAC 2001 |
Interconnection networks and networks-on-chip
on-chip communication |
0.0 | 1 | 2001 | Robust Interfaces for Mixed-Timing Systems with Application to Latency-Insensitive Protocols · DAC 2001 |
Energy-efficient computing › power-performance tradeoff
energy-delay tradeoff |
0.0 | 1 | 2006 | Tartan: evaluating spatial computation for whole program execution · ASPLOS 2006 |
Energy-efficient computing
energy-efficient architecture |
0.0 | 1 | 2004 | Spatial computation · ASPLOS 2004 |
Methods — techniques the papers use, named apart from their topics
hardware profiling · 0.1energy-delay metric analysis · 0.1discrete-event model · 0.1static configuration · 0.1static analysis · 0.1packet-switched routing · 0.1interconnect topology optimization · 0.1dynamic scheduling · 0.1automatic partitioning · 0.1compilation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2007 | Self-Resetting Latches for Asynchronous Micro-PipelinesabstractAsynchronous circuits are increasingly attractive as low power or high-performance replacements to synchronous designs. A key part of these circuits are asynchronous micropipelines; unfortunatelly, the existing micropipeline styles either improve performance or decrease power consumption, but not both. Very often, the pipeline register plays a crucial role in these cost metrics. In this paper we introduce a new register design, called self-resetting latches, for asynchronous micropipelines which bridges the gap between fast, but power hungry, latch-based designs and slow, but low power, flip-flop designs. The energy-delay metric for large asynchronous systems implemented with self-resetting latches is, on average, 41% better than latch-based designs and 15% better than flip-flop designs. Tiberiu Chelcea, Girish Venkataramani, Seth Copen Goldstein |
DAC | 1 |
| 2007 | Global Critical Path: A Tool for System-Level Timing AnalysisabstractAn effective method for focusing optimization effort on the most important parts of a design is to examine those elements on the critical path. Traditionally, the critical path is defined at the RTL level, as the longest path in the combinational logic between clocked registers. In this paper, we present a system-level timing analysis technique to define the concept of a global critical path (GCP), for predicting system-level performance. We show how the GCP can be used as a theoretical and practical tool for understanding, summarizing and optimizing the behavior of highly concurrent self-timed circuits. We formally define the GCP and show how it can be constructed using a discrete event model and hardware profiling techniques. The GCP provides valuable insight into the control-path behavior of circuits and in finding system-level bottlenecks. We have incorporated the GCP construction and analysis framework into a high-level synthesis and simulation toolchain, thus enabling complete automation in modeling, analysis and optimization. Girish Venkataramani, Mihai Budiu, Tiberiu Chelcea, Seth Copen Goldstein |
DAC | 3 |
| 2006 | Tartan: evaluating spatial computation for whole program executionabstractSpatial Computing (SC) has been shown to be an energy-efficient model for implementing program kernels. In this paper we explore the feasibility of using SC for more than small kernels. To this end, we evaluate the performance and energy efficiency of entire applications on Tartan, a general-purpose architecture which integrates a reconfigurable fabric (RF) with a superscalar core. Our compiler automatically partitions and compiles an application into an instruction stream for the core and a configuration for the RF. We use a detailed simulator to capture both timing and energy numbers for all parts of the system.Our results indicate that a hierarchical RF architecture, designed around a scalable interconnect, is instrumental in harnessing the benefits of spatial computation. The interconnect uses static configuration and routing at the lower levels and a packet-switched, dynamically-routed network at the top level. Tartan is most energyefficient when almost all of the application is mapped to the RF, indicating the need for the RF to support most general-purpose programming constructs. Our initial investigation reveals that such a system can provide, on average, an order of magnitude improvement in energy-delay compared to an aggressive superscalar core on single-threaded workloads. Mahim Mishra, Timothy J. Callahan, Tiberiu Chelcea, Girish Venkataramani, Seth Copen Goldstein, Mihai Budiu |
ASPLOS | 3 |
| 2006 | Hardware compilation of application-specific memory-access interconnectabstractA major obstacle to successful high-level synthesis (HLS) of large-scale application-specified integrated circuit systems is the presence of memory accesses to a shared-memory subsystem. The latency to access memory is often not statically predictable, which creates problems for scheduling operations dependent on memory reads. More fundamental is that dependences between accesses may not be statically provable (e.g., if the specification language permits pointers), which introduces memory-consistency problems. Addressing these issues with static scheduling results in overly conservative circuits, and thus, most state-of-the-art HLS tools limit memory systems to those that have predictable latencies and limit programmers to specifications that forbid arbitrary memory-reference patterns. A new HLS framework for the synthesis and optimization of memory accesses (SOMA) is presented. SOMA enables specifications to include arbitrary memory references (e.g., pointers) and allows the memory system to incorporate features that might cause the latency of a memory access to vary dynamically. This results in raising the level of abstraction in the input specification, enabling faster design times. SOMA synthesizes a memory access network (MAN) architecture that facilitates dynamic scheduling and ordering of memory accesses. The paper describes a basic MAN construction technique that illustrates how dynamic ordering helps in efficiently maintaining memory consistency and how dynamic scheduling helps alleviate the variable-latency problem. Then, it is shown how static analysis of the access patterns can be used to optimize the MAN. One optimization changes the MAN interconnect topology to increase concurrence. A second optimization reduces the synchronization overhead necessary to maintain memory consistency. Postlayout experiments demonstrate that SOMA's application-specific MAN construction significantly improves power and performance for a range of benchmarks. Girish Venkataramani, Tobias Bjerregaard, Tiberiu Chelcea, Seth Copen Goldstein |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2004 | Spatial computationabstractThis paper describes a computer architecture, Spatial Computation (SC), which is based on the translation of high-level language programs directly into hardware structures. SC program implementations are completely distributed, with no centralized control. SC circuits are optimized for wires at the expense of computation units.In this paper we investigate a particular implementation of SC: ASH (Application-Specific Hardware). Under the assumption that computation is cheaper than communication, ASH replicates computation units to simplify interconnect, building a system which uses very simple, completely dedicated communication channels. As a consequence, communication on the datapath never requires arbitration; the only arbitration required is for accessing memory. ASH relies on very simple hardware primitives, using no associative structures, no multiported register files, no scheduling logic, no broadcast, and no clocks. As a consequence, ASH hardware is fast and extremely power efficient.In this work we demonstrate three features of ASH: (1) that such architectures can be built by automatic compilation of C programs; (2) that distributed computation is in some respects fundamentally different from monolithic superscalar processors; and (3) that ASIC implementations of ASH use three orders of magnitude less energy compared to high-end superscalar processors, while being on average only 33% slower in performance (3.5x worst-case). Mihai Budiu, Girish Venkataramani, Tiberiu Chelcea, Seth Copen Goldstein |
ASPLOS | 3 |
| 2004 | Robust interfaces for mixed-timing systemsabstractThis paper presents several low-latency mixed-timing FIFO (first-in-first-out) interfaces designs that interface systems on a chip working at different speeds. The connected systems can be either synchronous or asynchronous. The designs are then adapted to work between systems with very long interconnect delays, by migrating a single-clock solution by Carloni et al. (1999, 2000, and 2001) (for "latency-insensitive" protocols) to mixed-timing domains. The new designs can be made arbitrarily robust with regard to metastability and interface operating speeds. Initial simulations for both latency and throughput are promising. Tiberiu Chelcea, Steven M. Nowick |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2002 | Resynthesis and peephole transformations for the optimization of large-scale asynchronous systemsabstractSeveral approaches have been proposed for the syntax-directed compilation of asynchronous circuits from high-level specification languages, such as Balsa and Tangram. Both compilers have been successfully used in large real-world applications; however, in practice, these methods suffer from significant performance overheads due to their reliance on straightforward syntax-directed translation.This paper introduces a powerful new set of transformations, and an extended channel-based language to support them, which can be used an optimizing back-end for Balsa. The transforms described in this paper fall into two categories: resynthesis and peephole. The proposed optimization techniques have been fully integrated into a comprehensive asynchronous CAD package, Balsa. Experimental results on several substantial design examples indicate significant performance improvements. supported by NSF ITR Award No. NSF-CCR-0086036 and NSF Award No. CCR-99-88241, and by a grant from the New York State Microelectronics Design Center. Tiberiu Chelcea, Steven M. Nowick |
DAC | 1 |
| 2002 | A Burst-Mode Oriented Back-End for the Balsa Synthesis SystemabstractThis paper introduces several new component clustering techniques for the optimization of asynchronous systems. In particular, novel "burst-mode aware" restriction.; are imposed to limit the cluster sizes and to ensure synthesizability. A new control specification language, CH, is also introduced which facilitates the manipulation and optimization of handshake control components. The new method has been fully integrated into a comprehensive asynchronous synthesis package, Balsa. Experimental results on several substantial design examples, including a 32-bit microprocessor core, indicate significant performance improvements for the optimized circuits. Tiberiu Chelcea, Steven M. Nowick, Andrew Bardsley, Doug A. Edwards |
DATE | 1 |
| 2001 | Robust Interfaces for Mixed-Timing Systems with Application to Latency-Insensitive ProtocolsabstractThis paper presents several low-latency mixed-timing FIFO designs that interface systems on a chip working at different speeds. The connected systems can be either synchronous or asynchronous. The design are then adapted to work between systems with very long interconnection delays, by migrating a single-clock solution by Carloni et al. (for “latency-insensitive” protocols) to mixed-timing domains. The new designs can be made arbitrarily robust with regard to metastability and interface operating speeds. Initial simulations for both latency and throughput are promising. Tiberiu Chelcea, Steven M. Nowick |
DAC | 1 |