Daniel Gajski

dblp:g/DanielDGajski · also Daniel D. Gajski · DBLP profile ↗
← Back
131ranked-venue papers
16as first author
0since 2021 · last 2010
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 121 · 14 first-authorSoftware engineering, systems software and programming languages · 17 · 3 first-authorTheory of computation · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
58 papers
Electronic design automation · 83% Embedded and real-time systems · 5% Processor architecture and microarchitecture · 4%
Software engineering, system software, and programming languages
4 papers
Compilers and program optimization · 99% Program analysis · 1%

Topics — the 30 heaviest of 79, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
high-level synthesis
0.4212010
What input-language is the best choice for high level synthesis (HLS)? · DAC 2010
C-based design flow: a case study on G.729A for voice over internet protocol (VoIP) · DAC 2008
Specify-explore-refine (SER): from specification to implementation · DAC 2008
Electronic design automation
system-level design
0.472009
Electronic System-Level Synthesis Methodologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009
Specify-explore-refine (SER): from specification to implementation · DAC 2008
Automatic Layer-Based Generation of System-On-Chip Bus Communication Models · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Electronic design automation
design space exploration
0.262008
Specify-explore-refine (SER): from specification to implementation · DAC 2008
Automatic Layer-Based Generation of System-On-Chip Bus Communication Models · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Retargetable profiling for rapid, early system-level design space exploration · DAC 2004
Electronic design automation › system-level design
system-level synthesis
0.232009
Electronic System-Level Synthesis Methodologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009
Automatic generation of equivalent architecture model from functional specification · DAC 2004
Automatic communication refinement for system level design · DAC 2003
Electronic design automation
design methodology
0.112009
Electronic System-Level Synthesis Methodologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009
Electronic design automation
hardware/software co-design
0.112009
Electronic System-Level Synthesis Methodologies · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2009
Electronic design automation › system-level design
communication synthesis
0.122007
Automatic Layer-Based Generation of System-On-Chip Bus Communication Models · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2007
Protocol Generation for Communication Channels · DAC 1994
Processor architecture and microarchitecture › microprocessor design › processor core design
datapath design
0.112008
Automatic architecture refinement techniques for customizing processing elements · DAC 2008
Electronic design automation
logic synthesis
0.1111999
IP-based Design Methodology · DAC 1999
Component synthesis from functional descriptions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1993
High-Level Transformations for Minimizing Syntactic Variances · DAC 1993
Compilers and program optimization
code size reduction
0.112007
FPGA-friendly code compression for horizontal microcoded custom IPs · FPGA 2007
Embedded and real-time systems › embedded processor
code compression
0.112007
FPGA-friendly code compression for horizontal microcoded custom IPs · FPGA 2007
Electronic design automation › system-level design › high-level system modeling
transaction-level modeling
0.112007
TLM: Crossing Over From Buzz To Adoption · DAC 2007
Embedded and real-time systems
embedded system design
0.022008
C-based design flow: a case study on G.729A for voice over internet protocol (VoIP) · DAC 2008
Specify-explore-refine (SER): from specification to implementation · DAC 2008
Electronic design automation › high-level synthesis
scheduling
0.041999
Soft Scheduling in High Level Synthesis · DAC 1999
A transformation-based method for loop folding · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1994
Hypertool: A Programming Aid for Message-Passing Systems · IEEE Trans. Parallel Distributed Syst. 1990
Electronic design automation
hardware description language
0.022001
Panel: The Next HDL: If C++ is the Answer, What was the Question? · DAC 2001
SpecCharts: a VHDL front-end for embedded systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1995
Electronic design automation
physical design
0.081992
Partitioning algorithms for layout synthesis from register-transfer netlists · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1992
Layout placement for sliced architecture · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1992
A technique for pull-up transistor folding · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1988
Electronic design automation › hardware/software co-design
hardware/software partitioning
0.021998
System-level exploration with SpecSyn · DAC 1998
Hardware/Software Partitioning and Pipelining · DAC 1997
Electronic design automation › high-level synthesis › behavioral transformation
behavioral synthesis
0.061990
Chippe: a system for constraint driven behavioral synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1990
An Intermediate Representation for Behavioral Synthesis · DAC 1990
VHDL Synthesis Using Structured Modeling · DAC 1989
Electronic design automation › hardware verification and test
hardware verification
0.022007
TLM: Crossing Over From Buzz To Adoption · DAC 2007
Flow graph representation · DAC 1986
Energy-efficient computing › low-power design
low-power processor design
0.012008
Automatic architecture refinement techniques for customizing processing elements · DAC 2008
Integrated circuit design
system-on-chip
0.012007
TLM: Crossing Over From Buzz To Adoption · DAC 2007
Electronic design automation › logic synthesis › formal synthesis
functional synthesis
0.021993
Component synthesis from functional descriptions · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1993
Functional Synthesis Using Area and Delay Optimization · DAC 1992
Electronic design automation › physical design
layout synthesis
0.031992
Partitioning algorithms for layout synthesis from register-transfer netlists · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1992
LES: a layout expert system · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1988
LES: A Layout Expert System · DAC 1987
Electronic design automation
hardware verification and test
0.041995
Interfacing Incompatible Protocols Using Interface Process Generation · DAC 1995
An Intermediate Representation for Behavioral Synthesis · DAC 1990
Design of Testable Structures Defined by Simple Loops · IEEE Trans. Computers 1981
Processor architecture and microarchitecture
pipelining
0.011997
Hardware/Software Partitioning and Pipelining · DAC 1997
Electronic design automation › high-level synthesis › folding
loop folding
0.011994
A transformation-based method for loop folding · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1994
Electronic design automation › high-level synthesis › scheduling
resource-constrained scheduling
0.021994
Percolation Based Synthesis · DAC 1990
A transformation-based method for loop folding · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1994
Electronic design automation
expert systems
0.021988
LES: a layout expert system · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1988
LES: A Layout Expert System · DAC 1987
Integrated circuit design › system-on-chip
system-on-chip design
0.012001
Panel: The Next HDL: If C++ is the Answer, What was the Question? · DAC 2001
Electronic design automation › logic synthesis › circuit optimization
area-delay optimization
0.011992
Functional Synthesis Using Area and Delay Optimization · DAC 1992

Methods — techniques the papers use, named apart from their topics

dictionary-based code compression · 0.1design space exploration · 0.1classification · 0.1virtual prototyping · 0.1trimming · 0.1specc · 0.1nanocode-based architecture · 0.1C-to-RTL synthesis · 0.1systemc · 0.1automatic model generation · 0.1threaded schedule · 0.0soft scheduling · 0.0performance estimation · 0.0heuristic cost minimization · 0.0bit-serial tuple-parallel processing · 0.0VLSI design · 0.0
YearPublicationVenuePosition
2010 What input-language is the best choice for high level synthesis (HLS)?
abstract
As of 2010, over 30 of the world's top semiconductor / systems companies have adopted HLS. In 2009, SOCs tape-outs containing IPs developed using HLS exceeded 50 for the first time. Now that the practicality and value of HLS is established, engineers are turning to the question of "what input-language works best?" The answer is critical because it drives key decisions regarding the tool/methodology infrastructure companies will create around this new flow. ANSI-C/C++ advocates cite ease-of-learning, simulation speed. SystemC advocates make similar claims, and point to SystemC's hardware-oriented features. Proponents of BSV (Bluespec SystemVerilog) claim that language enhances architectural transparency and control. To maximize the benefits of HLS, companies must consider many factors and tradeoffs.
Daniel Gajski, Todd M. Austin, Steve Svoboda
DAC1
2010 Accurate timed RTOS model for transaction level modeling
abstract
In this paper, we present an accurate timed RTOS model within transaction level models (TLMs). Our RTOS model, implemented on top of system level design language (SLDL), incorporates two key features: RTOS behavior model and RTOS overhead model. The RTOS behavior model provides dynamic scheduling, inter-process communication (IPC), and external communication for timing annotated user applications. While the RTOS behavior model is running, all RTOS events, such as context switch and interrupt handling, are passed to RTOS over-head model to adopt the overhead during system execution. Our RTOS overhead model has processor- and RTOS-specific pre-characterized overhead information to provide cycle approximate estimation. We demonstrate the applicability of our model using a multi-core platform executing a JPEG encoder. Experimental results show that the proposed RTOS model provides the high accuracy, 7% off compared to on-board measurements while simulating at speeds close to the reference C code.
Yonghyun Hwang, Gunar Schirner, Samar Abdi, Daniel Gajski
DATE4
2010 Design exploration and automatic generation of MPSoC platform TLMs from Kahn Process Network applications
abstract
With increasingly more complex Multi-Processor Systems on Chip (MPSoC) and shortening time-to-market projections, Transaction Level Modeling and Platform Aware Design are seen as promising >approaches to efficient MPSoC design.In this paper, we present an automatized 3-phase process of Platform Aware Design and apply it to Kahn Process Networks (KPN) applications, a widely used model of computation for data-flow applications. We start with the KPN application and an abstract platform template and automatically generate an executable TLM with estimated timing that accurately reflects the system platform. We support homogeneous and heterogeneous multi-master platform models with shared memory or direct communication paradigm. The communication in heterogeneous platform modules is enabled with the transducer unit (TX) for protocol translation. TX units also act as message routers to support Network on Chip (NoC) communication.We evaluate our approach with the case study of the H.264 Encoder design process, in which the specification compliant design was reached from the KPN application in less than 2 hours. The example demonstrates that automatic generation of platform aware TLMs enables a fast, efficient and error resilient design process.
Ines Viskic, Lochi Yu, Daniel Gajski
LCTES3
2009 Hardware-dependent software synthesis for many-core embedded systems
abstract
This paper presents synthesis of hardware dependent software (HdS) for multicore and many-core designs using embedded system environment (ESE). ESE is a tool set, developed at UC Irvine, for transaction level design of multicore embedded systems. HdS synthesis is a key component of ESE back-end design flow. We follow a design process that starts with an application model consisting of C processes communicating via abstract message passing channels. The application model is mapped to a platform net-list of SW and HW cores, buses and buffers. A high speed transaction level model (TLM) is generated to validate abstract communication between processes mapped to different cores. The TLM is further refined into a pin-cycle accurate model (PCAM) for board implementation. The PCAM includes C code for all the HdS layers including routing, packeting, synchronization and bus transfer. The generated HdS methods provide a library of application level services to the C processes on individual SW cores. Therefore, the application developer does not need to write low level HdS for board implementation. Synthesis results for an multi-core MP3 decoder design, using ESE, show that the HdS is generated in order of seconds, compared to hours of manual coding. The quality of synthesized code is comparable to manually written code in terms of performance and code size.
Samar Abdi, Gunar Schirner, Ines Viskic, Hansu Cho, Yonghyun Hwang, Lochi Yu, Daniel Gajski
ASP-DAC7
2009 Electronic System-Level Synthesis Methodologies
abstract
With ever-increasing system complexities, all major semiconductor roadmaps have identified the need for moving to higher levels of abstraction in order to increase productivity in electronic system design. Most recently, many approaches and tools that claim to realize and support a design process at the so-called electronic system level (ESL) have emerged. However, faced with the vast complexity challenges, in most cases at best, only partial solutions are available. In this paper, we develop and propose a novel classification for ESL synthesis tools, and we will present six different academic approaches in this context. Based on these observations, we can identify such common principles and needs as they are leading toward and are ultimately required for a true ESL synthesis solution, covering the whole design process from specification to implementation for complete systems across hardware and software boundaries.
Andreas Gerstlauer, Christian Haubelt, Andy D. Pimentel, Todor P. Stefanov, Daniel Gajski, Jürgen Teich
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2008 Specify-explore-refine (SER): from specification to implementation
abstract
Driven by increasing complexity and reliability demands, the Japanese Aerospace Exploration Agency (JAXA) in 2004 commissioned development of ELEGANT, a complete SpecC-based environment for electronic system-level (ESL) design of space and satellite electronics. As integral part of ELEGANT, the Center for Embedded Computer System (CECS) has developed and supplied the SER tool set. Following a Specify-Explore-Refine methodology, SER supports system-level design space exploration, interactive platform development and automatic model refinement and model generation. The SER engine has been successfully integrated into ELEGANT. With SER at its core, ELEGANT provides a seamless tool chain for modeling verification and synthesis from top-level specification down to embedded HW/SW implementation. ELEGANT and SER have been successfully delivered to JAXA and its suppliers. Tools are currently being deployed in companies like NEC Toshiba Space Systems. Evaluation results prove the feasibility of the approach for design space exploration, rapid virtual prototyping and system synthesis resulting in tremendous productivity and reliability gains. In addition, ELEGANT has been commercialized for general market availability. The SER component has been licensed to InterDesign Technologies, Inc. (IDT) and it is available from, sold and supported by IDT.
Andreas Gerstlauer, Junyu Peng, Dongwan Shin, Daniel Gajski, Atsushi Nakamura, Dai Araki, Yuuji Nishihara
DAC4
2008 Automatic architecture refinement techniques for customizing processing elements
abstract
In this paper, we propose an approach for designing high-performance energy-efficient processing elements (PEs) using statically-scheduled nanocode-based architectures. Our approach is based on bottom-up refinement/trimming techniques that optimize a given datapath irrespective of whether it was designed manually or generated automatically. The optimizations can also preserve parts of the netlist specified by the designers, and hence, allow reuse of design efforts and can lead to predictable convergence. In this paper, we show that trimming unused and underutilized resources of typical general-purpose datapaths can lead to 30-40% average energy savings, without any performance loss. However, general-purpose architectures often compromise parallelism to make the design implementable. With our trimming approach, we can afford to have a base architecture that is not intended for implementation and has more parallelism, and then apply refinement to make it implementable. For our benchmarks, we achieved up to 1.8 times (avg. 25%) and 2.6 times (avg. 40%) performance improvement, compared to two general-purpose architectures (i.e. a 4-issue VLIW and a DLX), respectively. Additionally, the energy consumption is reduced by up to 5 times (avg. 2 times) compared to the trimmed general-purpose architectures.
Bita Gorjiara, Daniel Gajski
DAC2
2008 C-based design flow: a case study on G.729A for voice over internet protocol (VoIP)
abstract
In this paper we present the design of a G. 729a codec in a C-based design flow. The codec is used in VoIP applications for sending speech over internet protocol. We started from the standard reference C implementation and generated several customized designs using the NISCT C-to-RTL toolset. Our final designs could run at very low clock frequencies (11 MHz for the decoder and 30 MHz for the coder) while meeting the timing requirements of the standard. We present these designs and the corresponding C-based design flow in this paper.
Mehrdad Reshadi, Bita Gorjiara, Daniel Gajski
DAC3
2008 Cycle-approximate Retargetable Performance Estimation at the Transaction Level
abstract
This paper presents a novel cycle-approximate performance estimation technique for automatically generated transaction level models (TLMs) for heterogeneous multi-core designs. The inputs are application C processes and their mapping to processing units in the platform. The processing unit model consists of pipelined datapath, memory hierarchy and branch delay model. Using the processing unit model, the basic blocks in the C processes are analyzed and annotated with estimated delays. This is followed by a code generation phase where delay-annotated C code is generated and linked with a SystemC wrapper consisting of inter-process communication channels. The generated TLM is compiled and executed natively on the host machine. Our key contribution is that the estimation technique is close to cycle-accurate, it can be applied to any multi-core platform and it produces high-speed native compiled TLMs. For experiments, timed TLMs for industrial scale designs such as MP3 decoder were automatically generated for 4 heterogeneous multi-processor platforms with up to 5 PEs under 1 minute. Each TLM simulated under 1 second, compared to 3-4 hrs of instruction set simulation (ISS) and 15-18 hrs of RTL simulation. Comparison to on-board measurement showed only 8 % error on average in estimated number of cycles.
Yonghyun Hwang, Samar Abdi, Daniel Gajski
DATE3
2008 Merged Dictionary Code Compression for FPGA Implementation of Custom Microcoded PEs
abstract
Horizontal Microcoded Architecture (HMA) is a paradigm for designing programmable high-performance processing elements (PEs). However, it suffers from large code size, which can be addressed by compression. In this article, we study the code size of one of the new HMA-based technologies called No-Instruction-Set Computer (NISC). We show that NISC code size can be several times larger than a typical RISC processor, and we propose several low-overhead dictionary-based code compression techniques to reduce its code size. Our compression algorithm leverages the knowledge of “don't care” values in the control words and can reduce the code size by 3.3 times, on average. Despite such good results, as shown in this article, these compression techniques lead to poor FPGA implementations because they require many on-chip RAMs. To address this issue, we introduce an FPGA-aware dictionary-based technique that uses the dual-port feature of on-chip RAMs to reduce the number of utilized block RAMs by half. Additionally, we propose cascading two-levels of dictionaries for code size and block RAM reduction of large programs. For an MP3 application, a merged, cascaded, three-dictionary implementation reduces the number of utilized block RAMs by 4.3 times (76%) compared to a NISC without compression. This corresponds to 20% additional savings over the best single level dictionary-based compression.
Bita Gorjiara, Mehrdad Reshadi, Daniel Gajski
ACM Trans. Reconfigurable Technol. Syst.3
2008 An Interactive Design Environment for C-Based High-Level Synthesis of RTL Processors
abstract
Abstract—Much effort in register transfer level (RTL) design has been devoted to developing “push-button ” types of tools. However, given the highly complex nature, and lack of control on RTL design, push-button type synthesis is not accepted by many designers. Interactive design with assistance of algorithms and tools can be more effective if it provides control to the steps of synthesis. In this paper, we propose an interactive RTL design environment which enables designers to control the design steps and to integrate hardware components into a system. Our design environment is targeting a generic RTL processor architecture and supporting pipelining, multicycling, and chaining. Tasks in the RTL design process include clock definition, component allocation, scheduling, binding, and validation. In our interactive environment, the user can control the design process at every stage, observe the effects of design decisions, and manually override synthesis decisions at will. We present a set of experimental results that demonstrate the benefits of our approach. Our combination of automated tools and interactive control by the designer results in quickly generated RTL designs with better performance than fully-automatic results, comparable to fully manually optimized designs. Index Terms—Embedded systems, high level synthesis, interactive design environment, register transfer level (RTL) processor, system-on-chip (SoC). I.
Dongwan Shin, Andreas Gerstlauer, Rainer Dömer, Daniel Gajski
IEEE Trans. Very Large Scale Integr. Syst.4
2007 TLM: Crossing Over From Buzz To Adoption
abstract
Transaction-level modeling --- originally used decades ago in the development of telecommunications network architecture --- is now widely used in SoC design. Why? Because the modern SoC is now so complex that systematic modeling and analysis are required to devise the optimal chip architecture. The architectural model is the essential platform that kick-starts two other key tasks --- verification testbench development and software development. In addition, the interoperability imperatives of SoC design and verification, IP reuse, software development, and system evaluation and integration have driven the replacement of proprietary transaction-level modeling methodologies by a TLM standard that leverages the power of SystemC. The standard --- devised by the Open SystemC Initiative (OSCI) in collaboration with the Open Core Protocol International Partnership (OCP-IP) --- covers the multiple levels of abstraction required for all of the foregoing tasks.
Francine Bacchini, Daniel Gajski, Laurent Maillet-Contoz, Haruhisa Kashiwagi, Jack Donovan, Tommi Mäkeläinen, Jack Greenbaum, Rishiyur S. Nikhil
DAC2
2007 Interrupt and low-level programming support for expanding the application domain of statically-scheduled horizontal-microcoded architectures in embedded systems
Mehrdad Reshadi, Daniel Gajski
DATE2
2007 FPGA-friendly code compression for horizontal microcoded custom IPs
abstract
Shrinking time-to-market and high demand for productivity has driven traditional hardware designers to use design methodologies that start from high-level languages. However, meeting timing constraints of automatically generated IPs is often a challenging and time-consuming task that must be repeated every time the specification is modified. To address this issue, a new generation of IP-design technologies that is capable of generating custom datapaths as well as programming an existing one is developed. These technologies are often based on Horizontal Microcoded Architectures. Large code size is a well-know problem in HMAs, and is referred to as "code bloating" problem.In this paper, we study the code size of one of the new HMA-based technologies called NISC. We show that NISC code size can be several times larger than a typical RISC processor, and we propose several low-overhead dictionary-based code compression techniques to reduce the code size. Our compression algorithm leverages the knowledge of "don't care" values in the control words to better compress the content of dictionary memories. Our experiments show that by selecting proper memory architectures the code size of NISC can be reduced by 70% (i.e. 3.3 times) at cost of only 9% performance degradation. We also show that some code compression techniques may increase number of utilized block RAMs in FPGA-based implementations. To address this issue, we propose combining dictionaries and implementing them using embedded dual-port memories.
Bita Gorjiara, Daniel Gajski
FPGA2
2007 A novel profile-driven technique for simultaneous power and code-size optimization of microcoded IPs
abstract
Microcoded customized IPs have significantly better performance, yet larger code size, compared to similarly-sized instruction-based processors. Storing wide microcodes on-chip requires wide memory-blocks that occupy a large area and consume high leakage power. Therefore, addressing the code size of microcoded IPs is very important. In this paper, we introduce compression techniques that along with careful resolution of ldquodonpsilat carerdquo values (denoted by dasiaXpsila) in microcode can address the code size issue. We observed that dasiaXpsila values can be used for improving either dynamic power of IPs or their compression. However, achieving the efficiency of both is challenging. In this paper, we propose a profile-guided dasiaXpsila-resolution technique that can achieve both power and compression efficiency. Using our technique, the code size of microcoded IPs is reduced by 2.7 times, while saving 20% dynamic power, on average.
Bita Gorjiara, Daniel Gajski
ICCD2
2007 Interface synthesis for heterogeneous multi-core systems from transaction level models
abstract
This paper presents a tool for automatic synthesis of RTL interfaces for heterogeneous MPSoC from transaction level models (TLMs). The tool captures the communication parameters in the platform and generates interface modules called universal bridges between buses in the design. The design and configuration of the bridges depend on several platform components including heterogeneity of the components, traffic on the bus, size of messages and so on. We define these parameters and show how the synthesizable RTL code for the bridge can be automatically derived based on these parameters. We use industrial strength design drivers such as an MP3 decoder to test our automatically generated bridges for a variety of platforms and compare them to manually designed bridges on different quality metrics. Our experimental results show that performance of automatically generated bridges are within 5% of manual design for simple platforms but surpasses them for more complex platforms. The area and RTL code size is consistently better than manual design while giving 5 orders of improvement in development time.
Hansu Cho, Samar Abdi, Daniel Gajski
LCTES3
2007 Automatic generation of embedded communication SW for heterogeneous MPSoC platforms
abstract
This paper addresses the problem of long design cycle of MPSoCs communication SW with automatic synthesis. The tool we propose takes as input a transaction level model (TLM) of MPSoC communication and outputs pin and cycle-accurate (PCA) bus drivers that can be linked to the synthesizable PCA model (PCAM).
Ines Viskic, Samar Abdi, Daniel Gajski
LCTES3
2007 Automatic Layer-Based Generation of System-On-Chip Bus Communication Models
abstract
With growing market pressures and rising system complexities, automated system-level communication design with efficient design space exploration capabilities is becoming increasingly important. At the same time, customized network-oriented communication architectures become necessary in enabling a high-performance communication among the system components. To this end, corresponding communication design flows that are supported by efficient design automation techniques need to be developed. In this paper, we present a system-level design environment for the generation of bus-based system-on-chip architectures. Our approach supports a two-stage design flow using automated model refinement toward custom heterogeneous communication networks. Starting from an abstract specification of the desired communication channels, our environment automatically generates tailored network models at various levels of abstraction. At its core, an automatic layer-based refinement approach is utilized. We have applied our approach to a set of industrial-strength examples with a wide range of target architectures. Our experimental results show significant productivity gains over a traditional communication design, allowing early and rapid design space exploration.
Andreas Gerstlauer, Dongwan Shin, Junyu Peng, Rainer Dömer, Daniel Gajski
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2006 Design and implementation of transducer for ARM-TMS communication
abstract
Communication between components, with different interface protocols, requires an extra component that must translate one protocol to another. This component is referred to as a transducer. In this paper we describe the design and implementation of a transducer between AMBA bus and TMS DSP bus. The transducer allows system designers to send data from AMBA compliant components to TMS compliant ones, and vice versa. The transducer was modeled in Verilog and implemented on Xilinx VirtexII FPGA board
Hansu Cho, Samar Abdi, Daniel Gajski
ASP-DAC3
2006 Designing a custom architecture for DCT using NISC technology
abstract
This paper presents design of a custom architecture for discrete cosine transform (DCT) using no-instruction-set computer (NISC) technology that is developed for fast processor customization. Using several software transformations and hardware customization, we achieved up to 10 times performance improvement, 2 times power reduction, 12.8 times energy reduction, and 3 times area reduction compared to an already-optimized soft-core MIPS implementation
Bita Gorjiara, Mehrdad Reshadi, Daniel Gajski
ASP-DAC3
2006 A Graph Based Algorithm for Data Path Optimization in Custom Processors
abstract
The rising complexity, customization and short time to market of modern digital systems requires automatic methods for generation of high performance architectures for such systems. This paper presents algorithms to automatically create custom data path for a given application that optimizes both resource utilization and performance. The inputs to the architecture generator include application source code, operation execution frequency obtained by the profile run and a component library (consisting of ALUs, busses, multiplexers etc.). The output is the application specific data path specified as the set of resource instances and their connections. The algorithm starts with a dense architecture and iteratively refines it until an efficient architecture is derived. The key optimization goal is to keep performance within given boundaries while maximizing resource utilization. Our experimental results show that generated architectures are comparable to manual designs, but can be obtained in a matter of few seconds, thereby leading to significant productivity gains
Jelena Trajkovic, Mehrdad Reshadi, Bita Gorjiara, Daniel Gajski
DSD4
2006 Aspect-Oriented Architecture Description for Retargetable Compilation, Simulation and Synthesis of Application-Specific Pipelined Datapaths
abstract
Constraints of embedded systems and the shrinking time-to-market have elevated the importance of designer productivity and design predictability more than ever. To improve productivity, in ASIP approaches the system is designed with software and executed on a customized processor. In ASIP design flow, the processor is described in an Architecture Description Language (ADL) and the toolset is generated from that ADL automatically. However, in these approaches design predictability is low because the designer has little or no control over the quality of the final implementation. In this paper, we present a new design approach where the target processor or Intellectual Property (IP) does not have any predefined instruction-set and its datapath component netlist is described in a Generic Netlist Representation (GNR). The GNR is used by the toolset to generate the controller of the IP and the RTL of the design. The GNR is an order of magnitude shorter than state-of-the-art ADLs with RTL generation capabilities and yet can capture any structural details that affect the implementation quality. We have also developed a web-based interface for our toolset, so that users can upload and evaluate new IPs described in GNR.
Bita Gorjiara, Mehrdad Reshadi, Daniel Gajski
ICCD3
2005 A formalism for functionality preserving system level transformations
abstract
With the rise in complexity of modern systems, designers are spending a significant time on modeling at the system level of abstraction. This paper introduces Model Algebra, a formalism built on top of system level design languages, that can be used for implementing functionality preserving transformations on system level models. Such transformations enable us to implement high level design decisions without having to write new models for each design decision. Moreover, since these transformations preserve functionality, the transformed models do not need to be re-verified. We present the definition of Model Algebra and show how system level models can be represented as expressions in this formalism. The laws of Model Algebra are use to define correct model transformations. We show a system level design scenario, where design decisions gradually refine the functional model of the system to an architectural model with components and communication structure. The refinement can be performed using the correct model transformations in our formalism.
Samar Abdi, Daniel Gajski
ASP-DAC2
2005 Multi-metric and multi-entity characterization of applications for early system design exploration
abstract
At system level, intensively analyzing the system application will produce a variety of useful characteristics and provide designers valuable exploration indications. In this paper, we present such an analysis approach based on the instrumentation-based profiling. The proposed approach analyzes complex system application and generates multi-metric and multi-entity characteristics. Experimental results show the applicability of the approach for efficient early design space exploration.
Lukai Cai, Andreas Gerstlauer, Daniel Gajski
ASP-DAC3
2005 System-level communication modeling for network-on-chip synthesis
abstract
As we are entering the network-on-chip era and system communication is becoming a dominating factor, communication abstraction and synthesis are becoming the integral part of system design flows. The key to the success of any design flow are well-defined abstraction levels and models, which enable automation of early validation, synthesis and verification. In this paper, we define system communication abstraction layers and corresponding design models that support successive, stepwise refinement from abstract message-passing down to a cycle-accurate, bus-functional implementation. Experimental results show the benefits of our definitions and design flow.
Andreas Gerstlauer, Dongwan Shin, Rainer Dömer, Daniel Gajski
ASP-DAC4
2005 A clustering technique to optimize hardware/software synchronization
abstract
In this paper we present a scheme for reducing the amount of synchronization overhead needed between components, after HW/SW partitioning, to preserve the original control flow of the specification. Since traffic between components is expensive, our scheme can significantly enhance the performance of the system implementation. Our optimization technique dynamically groups the tasks in the specification such that synchronization for different tasks can be shared. The grouping depends on the partitioning decision, and hence, is performed during the generation of the partitioned model. We apply our grouping algorithm for various partitions on system level models of industry standard designs. The experimental results show significant reduction in synchronization overhead compared to the unoptimized model.
Junyu Peng, Samar Abdi, Daniel Gajski
ASP-DAC3
2005 Functional Validation of System Level Static Scheduling
abstract
Increase in system level modeling has given rise to a need for efficient functional validation of models above cycle accurate level. This paper presents a technique for comparing system level models, before and after the static scheduling of tasks on processing elements of the architecture. We derive a graph representation from models written in system level design languages (SLDLs) and define their execution semantics. The notion of functional equivalence of system level models is established using these graphs. We then present well defined rules for reduction of such graphs to a normal form. Finally, we show how to check for functional equivalence of two system level models by isomorphism of their normal graph representations. A checker built on the above concept is used to automatically validate the functional correctness of the static scheduling step. As a result, the models generated for various scheduling decisions do not have to be reverified using costly simulations.
Samar Abdi, Daniel Gajski
DATE2
2005 Defining an Enhanced RTL Semantics
abstract
In this paper, we formally define an enhanced RTL semantics. This is intended to elevate the RTL design abstraction level and help bridge the HDL semantic gap among synthesis, simulation and formal verification tools. We define the enhanced semantics based on a new RTL++ language that supports pipelined operations using a new pipelined register variable concept. The execution semantics of RTL++ is specified in a structural operational semantics style aimed at forming the basis for related simulation and formal verification algorithm development. An RFSM model is defined to support natively the synthesis semantics of RTL++. We also present an example of extending SystemC to support the notion of a pipelined register variable.
Shuqing Zhao, Daniel Gajski
DATE2
2005 Utilizing Horizontal and Vertical Parallelism with a No-Instruction-Set Compiler for Custom Datapaths
abstract
Performance of programs can be improved by utilizing their horizontal and vertical parallelism. In some processors (VLIW based), compiler can utilize horizontal parallelism by controlling the schedule of independent operations. Vertical parallelism is utilized through pipelining. However, in all processors, structure of pipeline is fixed and compiler has no control over it. In application-specific-instruction set-processors (ASIPs), pipeline structure can be customized and utilized in the program through custom instructions. Practical constraints on the instruction decoder limit the number and complexity of custom instructions in ASIPs. Detecting the frequent and beneficial custom instructions and incorporating them in the compiler are complex and sometimes very time consuming tasks. In this paper, we present an architecture that does not limit the number of custom functionalities that can be implemented on its datapath. Instead of using custom instructions and then relying on the decoder in hardware to generate the control signals, we generate the control signal values in compiler. Since there are no predefined instructions in this architecture, we call it no-instruction-set-computer (NISC). The NISC compiler maps the application directly on the datapath. It has complete fine grain control over datapath and hence can very well utilize resources in the hardware as well as horizontal and vertical parallelism in the program. We also explain the algorithm for mapping the CDFG of a program on a given datapath in NISC. Using our algorithm and a NISC architecture with the datapath of a MIPS, we achieved up to 70% speedup over the traditional MIPS compiler. In another experiment, we started from a base architecture and customized it by adding resources and interconnect to increase its horizontal and vertical parallelism. The algorithm achieved up to 15.5 times speedup by utilizing the available parallelism in the program and the datapath.
Mehrdad Reshadi, Bita Gorjiara, Daniel Gajski
ICCD3
2005 System design extreme makeover
abstract
With complexities of systems-on-chip (SOCs) rising almost daily, the design community has been searching for a new methodology that can handle given complexities with increased productivity and decreased time-to-market. In order to find a solution for the system-level design flow, we must look at the system gap between SW and HW designs and then try to bridge this gap by developing a design flow that is based on common principles applicable to software and hardware. In order to achieve this design flow we can define design process by using the concepts found in standard algebras which, in turn, allows us to define design models more formally with clean unambiguous semantics. Such clean semantics allows automatic model generation, simplifies synthesis algorithms and verification techniques.
Daniel Gajski
MEMOCODE1
2005 Structural operational semantics for supporting multi-cycle operations in RTL HDLs
abstract
In this paper we formally define an operational semantics framework RTL++ for modeling behavioral RTL hardware IP. The semantics we define is neutral to existing HDLs and extends traditional sense RTL by natively supporting pipelined and multi-cycled operations with a unified registervariable type. We believe this formalization help to guide the design of new HDLs or extensions of existing HDLs in terms of elevating RTL design abstraction level and also bridging the current HDL semantic gap among synthesis, simulation and formal verification tools. The intra-module and inter-module execution of RTL++ semantics are specified in Plotkin-style structural operational semantics framework. An example of implementing the RTL++ extension of SystemC is presented along with experimental results showing the benefit of modeling in RTL++.
Shuqing Zhao, Daniel Gajski
MEMOCODE2
2004 On deriving equivalent architecture model from system specification
Samar Abdi, Daniel Gajski
ASP-DAC2
2004 A novel memory size model for variable-mapping in system level design
Lukai Cai, Haobo Yu, Daniel Gajski
ASP-DAC3
2004 Automatic generation of bus functional models from transaction level models
Dongwan Shin, Samar Abdi, Daniel Gajski
ASP-DAC3
2004 Embedded software generation from system level design languages
Haobo Yu, Rainer Dömer, Daniel Gajski
ASP-DAC3
2004 Automatic generation of equivalent architecture model from functional specification
abstract
This paper presents an algorithm for automatic generation of an architecture model from a functional specification, and proves its correctness. The architecture model is generated by distributing the intended system functionality over various components in the platform architecture. We then define simple transformations that preserve the execution semantics of system level models. Finally, the model generation algorithm is proved correct using our transformations. As a result, we have an automated path from a functional model of the system to an architectural one and we need to debug and verify only the functional specification model, which is smaller and simpler than the architecture model. Our experimental results show significant savings in both the modeling and the validation effort.
Samar Abdi, Daniel Gajski
DAC2
2004 Retargetable profiling for rapid, early system-level design space exploration
abstract
Fast and accurate estimation is critical for exploration of any design space in general. As we move to higher levels of abstraction, estimation of complete system designs at each level of abstraction is needed. Estimation should provide a variety of useful metrics relevant to design tasks in different domains and at each stage in the design process.In this paper, we present such a system-level estimation approach based on a novel combination of dynamic profiling and static retargeting. Co-estimation of complete system implementations is fast while accurately reflecting even dynamic effects. Furthermore, retargetable profiling is supported at multiple levels of abstraction, providing multiple design quality metrics at each level. Experimental results show the applicability of the approach for efficient design space exploration.
Lukai Cai, Andreas Gerstlauer, Daniel Gajski
DAC3
2004 Were the good old days all that good?: EDA then and now
abstract
A long, long time ago, in a laboratory far, far away, EDA researchers and developers used paper tape instead of Linux, rubylith instead of GDS II, yellow wires instead of ten levels of metal. Sitting around a potbellied stove, in their rocking chairs, practitioners of that era (and this) will offer insight into why some great ideas were immediately put into practice while others stayed on the drawing board or in the ivory tower. They will share remembrances of things past, of simpler days when foundries made steel, when options meant CMOS or bipolar, when real parts were measured instead of benchmarks touted. Their stories of what it was like, what has changed, and whether the "good old days" were then or now will be followed by questions and, possibly, answers.
Shishpal Rawat, William H. Joyner Jr., John A. Darringer, Daniel Gajski, Pat O. Pistilli, Hugo De Man, Carl Harris, James Solomon
DAC4
2003 Automatic communication refinement for system level design
abstract
This paper presents a methodology and algorithms for automatic communication refinement. The communication refinement task in system-level synthesis transforms abstract data-transfer between components to its actual bus level implementation. The input model of the communication refinement is a set of concurrently executing components, communicating with each other through abstract communication channels. The refined model reflects the actual communication architecture. Choosing a good communication architecture in system level designs requires sufficient exploration through evaluation of various architectures. However, this would not be possible with manually refining the system model for each communication architecture. For one, manual refinement is tedious and error-prone. Secondly, it wastes substantial amount of precious designer time. We solve this problem with automatic model refinement. We also present a set of experimental results to demonstrate how the proposed approach works on a typical system level design.
Samar Abdi, Dongwan Shin, Daniel Gajski
DAC3
2003 RTOS Modeling for System Level Design
Andreas Gerstlauer, Haobo Yu, Daniel Gajski
DATE3
2003 Transaction Based Design: Another Buzzword or the Solution to a Design Problem?
Heinz-Josef Schlebusch, Gary Smith 0001, Donatella Sciuto, Daniel Gajski, Carsten Mielenz, Christopher K. Lennard, Frank Ghenassia, Stuart Swan, Joachim Kunkel
DATE4
2002 Top-Down System Level Design Methodology Using SpecC, VCC and SystemC
abstract
In this paper we suggest a top-down methodology from C to silicon. In our methodology, we focus on methods to make the design flow smooth, efficient, and easy. The proposed methodology is a pure top-down methodology. We developed our design methodology by using SpecC, VCC, and SystemC. We choose SpecC, VCC and SystemC because they are all C-related and each have strong support in at least one field of design. Our proposal for a methodology is based on our experiences of attempting to model the JPEG encoder with SpecC, SystemC and VCC, and one internal project, attempting to implement architecture exploration for MPEG encoding and decoding using VCC.
Lukai Cai, Daniel Gajski, Paul Kritzinger, Mike Olivarez
DATE2
2002 Seamless approach for the design of control systems for power electronics and electric drives
abstract
Today, the shortest time-to-market in the electric drives industries is being a pressing requirement, consequently development time of new algorithms and new control systems and debugging them must be minimized. This requirement can be satisfied only by using a well-defined System-level design methodology and by reducing the migration time between the algorithm development language and the hardware specification language. In this paper, we propose to apply the SpecC methodology to the the design of control systems for power electronics and electric drives. We first begin with an executable specification model of the control device. Then, we describe the different steps and transformations used to convert this model to a communication model, which can be then transformed to an implementation model ready for manufacturing.
Slim Ben Saoud, Daniel Gajski, Andreas Gerstlauer
SMC2
2002 An ultra-fast instruction set simulator
abstract
In this paper, we present new techniques which further improve the static compilation-based instruction set architecture (ISA) simulation by the aggressive utilization of the host machine resources. Such utilization is achieved by defining a low-level code-generation interface specialized for ISA simulation, rather than the traditional approaches which use C as a code-generation interface. We are able to perform the simulation at a speed of up to 10/sup 2/ millions of simulated instructions per second (MIPS) on a 270 MHz Ultra-5 workstation. This result is only on average 1.6 times slower than the native execution on the host machine, the fastest to the best of our knowledge.
Jianwen Zhu, Daniel Gajski
IEEE Trans. Very Large Scale Integr. Syst.2
2001 Compiling SpecC for simulation
abstract
Systems-on-chip (SOC) design calls for the use of executable system design language (SLDL). SpecC is a C-based SLDL designed to embrace the IP-centric design methodology. In this paper, we present a SpecC-language-based approach for system level simulation of SOC. Pros and cons of this approach is compared against the existing library-based approaches. Furthermore, we discuss in detail the various design considerations for the SpecC simulation API as well as our reference implementation.
Jianwen Zhu, Daniel Gajski
ASP-DAC2
2001 Panel: The Next HDL: If C++ is the Answer, What was the Question?
abstract
The focus of this panel is on issues surrounding the use of C++ in modeling, integration of silicon IP and system-on-chip designs. In the last two years there have been several announcements promoting C++ based solutions and of multiple consortia (SystemC, Cynapps, Accellera, SpecC) that represent increasing commercial interest both from tool vendors as well as perhaps expression of genuine needs from the design houses. There are, however, serious questions about what value proposition does a C++ based design methodology bring to the IC or system designer? What has changed in the modeling technology (and/or available tools) that gives a new capability? Is synthesis the right target? or VAlidation? Tester modeling or testbench generation? This panel brings together advocates and opponents from the user community to highlight the achievements and the challenges that remain in use C++ for use in microelectronic circuits and systems.
Rajesh K. Gupta 0001, Shishpal Rawat, Ingrid Verbauwhede, Gérard Berry, Ramesh Chandra, Daniel Gajski, Kris Konigsfeld, Patrick Schaumont
DAC6
2001 C/C++: progress or deadlock in system-level specification
Daniel Gajski, Eugenio Villar, Wolfgang Rosenstiel, Vassilios Gerousis, D. Barton, Jonas Plantin, S. E. Ericsson, Patrizia Cavalloro, Gjalt G. de Jong
DATE1
2001 Performance-constrained hierarchical pipelining for behaviors, loops, and operations
abstract
Behavioral specifications of DSP systems generally contain a number of nested loops. In order to obtain high date rates for such systems, it is necessary to pipeline the system within the behavior, within the loop bodies, and also within the operations. In order to hierarchically pipeline a performance-constrained system, an important step consists of distributing the performance constraint among the loops in such a manner that the constraint is satisfied and design cost is minimized. This paper presents an algorithm for propagating constraints and hierarchically pipelining a given throughput-constrained system. Along with pipelining, the algorithm schedules the operations within the loop bodies and selects components for them, with the aim of minimizing cost while satisfying the constraint propagated to the loop body. Results demonstrate the necessity of pipelining across the three granularity levels in order to obtain high performance designs. They also demonstrate the feasibility and quality of our approach, the indicate that it may be efficiently used for synthesizing or estimating within system-level design.
Smita Bakshi, Daniel Gajski
ACM Trans. Design Autom. Electr. Syst.2
2000 Reuse and protection of intellectual property in the SpecC system
abstract
No abstract available.
Rainer Dömer, Daniel Gajski
ASP-DAC2
2000 Usage-based characterization of complex functional blocks for reuse in behavioral synthesis
abstract
Abstract — This paper presents a novel usage-basedcharacterization method to capture pre-designed complex functional blocks for automatic reuse in behavioral synthesis. We identify attributes necessary for reuse of such complex components and illustrate how the attributes are captured into a design database. A complex componentMT X MULT 8X8, which computes product of two 8 8 matrices, is captured with the proposed method, and its reuse in behavioral synthesis is demonstrated with design of a DCT example. Feasibility of the method for capturing various components is demonstrated as well. I.
Nong Fan, Viraphol Chaiyakul, Daniel Gajski
ASP-DAC3
2000 Embedded tutorial: essential issues for IP reuse
abstract
Article Free Access Share on Embedded tutorial: essential issues for IP reuse Authors: Daniel D. Gajski University of California, Irvine University of California, IrvineView Profile , Allen C.-H. Wu Tsing Hua University, Taiwan, ROC Tsing Hua University, Taiwan, ROCView Profile , Viraphol Chaiyakul Y Explorations Inc. Y Explorations Inc.View Profile , Shojiro Mori Toshiba Corp., Japan Toshiba Corp., JapanView Profile , Tom Nukiyama NEC Corp., Japan NEC Corp., JapanView Profile , Pierre Bricaud Mentor Graphics Corp. Mentor Graphics Corp.View Profile Authors Info & Claims ASP-DAC '00: Proceedings of the 2000 Asia and South Pacific Design Automation ConferenceJanuary 2000 Pages 37–42https://doi.org/10.1145/368434.368504Published:28 January 2000Publication History 19citation399DownloadsMetricsTotal Citations19Total Downloads399Last 12 Months37Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Daniel Gajski, Allen C.-H. Wu, Viraphol Chaiyakul, Shojiro Mori, Tom Nukiyama, Pierre Bricaud
ASP-DAC1
2000 One language or more?: how can we design an SoC at a system level?
abstract
No abstract available.
Masaharu Imai, Gary Smith 0001, Steven Schulz, Karen Bartleson, Daniel Gajski, Wolfgang Rosenstiel, Peter Flake, Hiroto Yasuura
ASP-DAC5
1999 IP-based Design Methodology
abstract
No abstract available.
Daniel Gajski
DAC1
1999 Soft Scheduling in High Level Synthesis
abstract
In this paper, we establish a theoretical framework for a new concept of scheduling called soft scheduling.In contrasts to the traditional schedulers referred as hard schedulers, soft schedulers make soft decisions at a time, or decisions that can be adjusted later.Soft scheduling has a potential to alleviate the phase coupling problem that has plagued traditional high level synthesis (HLS), HLS for deep submicron design and VLIW code generation.We then develop a specific soft scheduling formulation, called threaded schedule, under which a linear, optimal (in the sense of online optimality) algorithm is guaranteed.
Jianwen Zhu, Daniel Gajski
DAC2
1999 OpenJ: An Extensible System Level Design Language
abstract
There is an increasing research interest in system level design languages which can carry designers from specification to implementation of a system-on-a-chip. Unfortunately two of the most important goals in designing such a language, are at odds with each other: heterogeneity requires components of the system to be captured precisely by domain specific models to simplify analysis and synthesis; simplicity requires a consistent notation to avoid confusion. In this paper, we focus on our effort in resolving this dilemma in an extensible language called OpenJ. In contrast to the conventional monolithic languages, OpenJ has a layered structure consisting of the kernel layer which is essentially an object oriented language designed to be simple, modular and polymorphic; the open layer which exports parameterizable language constructs; the domain layer which precisely captures the computational models essential for embedded systems. The domain layer can be provided by vendors via a common protocol defined by an open layer which enables the supersetting or/and subsetting of the kernel. A compiler has been built for this language and experiments are conducted for popular models such as synchronous, discrete event and dataflow.
Jianwen Zhu, Daniel Gajski
DATE2
1999 Partitioning and pipelining for performance-constrained hardware/software systems
abstract
In order to satisfy cost and performance requirements, digital signal processing and telecommunication systems are generally implemented with a combination of different components, from custom-designed chips to off-the-shelf processors. These components vary in their area, performance, programmability and so on, and the system functionality is partitioned amongst the components to best utilize this tradeoff. However, for performance critical designs, it is not sufficient to only implement the critical sections as custom-designed high-performance hardware, but it is also necessary to pipeline the system at several levels of granularity. We present a design flow and an algorithm to first allocate software and hardware components, and then partition and pipeline a throughput-constrained specification amongst the selected components. This is performed to best satisfy the throughput constraint at minimal application-specific integrated-circuit cost. Our ability to incorporate partitioning with pipelining at several levels of granularity enables us to attain high throughput designs, and also distinguishes this paper from previously proposed hardware/software partitioning algorithms.
Smita Bakshi, Daniel Gajski
IEEE Trans. Very Large Scale Integr. Syst.2
1998 System-level exploration with SpecSyn
abstract
We present the SpecSyn system-level design environment supporting the specify-explore-re#ne #SER# design paradigm. This three-step approach includes precise speci#cation of system functionality, rapid exploration of numerous systemlevel design options, and re#nement of the speci#cation into one re#ecting the chosen option. A system-level design option consists of an allocation of system components like standard and custom processors, and a partitioning of functionality among those components. Focusing on SpecSyn's exploration techniques, we emphasize its two-phase estimation approach and highlight experiments using SpecSyn. 1 Introduction The focus of design e#ort on higher abstraction levels, driven by increasing system complexity, shorter design times, and migration of entire systems onto a single chip, demands a system-level design methodology and supporting tools. We can isolate three tasks in such a methodology. First, wemust specify the system's functionality and constraints. S...
Daniel Gajski, Frank Vahid, Sanjiv Narayan
DAC1
1998 Hierarchical pipelining for behaviors, loops, and operations
abstract
Behavioral specifications of DSP systems generally contain a number of nested loops. In order to obtain high date rates for such systems, it is necessary to pipeline them across their tasks, loops, and operations. This paper presents an algorithm for pipelining a given throughput-constrained system at these three different levels of granularity, while at the same time, scheduling the operations within the loop bodies and selecting components for them. Results demonstrate the feasibility and quality of our approach, and also indicate that it may be used for synthesis or estimation purposes in system-level design.
Smita Bakshi, Daniel Gajski
ICCD2
1998 SpecSyn: an environment supporting the specify-explore-refine paradigm for hardware/software system design
abstract
System-level design issues are gaining increasing attention, as behavioral synthesis tools and methodologies mature. We present the SpecSyn system-level design environment, which supports the new specify-explore-refine (SER) design paradigm. This three-step approach to design includes precise specification of system functionality, rapid exploration of numerous system-level design options, and refinement of the specification into one reflecting the chosen option. A system-level design option consists of an allocation of system components, such as standard and custom processors, memories, and buses, and a partitioning of functionality among those components. After refinement, the functionality assigned to each component can then he synthesized to hardware or compiled to software. We describe the issues and approaches for each part of the SpecSyn environment. The new paradigm and environment are expected to lead to a more than ten times reduction in design time, and our experiments support this expectation.
Daniel Gajski, Frank Vahid, Sanjiv Narayan
IEEE Trans. Very Large Scale Integr. Syst.1
1997 A quantitative analysis for optimizing memory allocation
abstract
Memory allocation problem has two independent goals: minimization of number of memories and minimization of number of registers in one memory. Our concern is the ordering of bindings during memory allocation. We formulate and analyze three different memory allocation algorithms by changing their binding order. It is shown that when we combine these subtasks and solve them simultaneously by heuristic cost function significant savings (up to 20%) can be obtained in the total area of memories.
Youn-Sik Hong, Choong-Hee Cho, Daniel Gajski
ASP-DAC3
1997 Hardware/Software Partitioning and Pipelining
abstract
For a given throughput constrained system-level specification,we present a design flow and an algorithm to select software(general purpose processors) and hardware components,and then partition and pipeline the specification amongstthe selected components.This is done so as to beat satisfythe throughput constraint at minimal hardware cost.Ourability to pipeline the design at several levels, enables us toattain high throughput designs, and also distinguishes ourwork from previously proposed hardware/software partitioning algorithms.
Smita Bakshi, Daniel Gajski
DAC2
1997 Model refinement for hardware-software codesign
abstract
Hardware-software codesign, which implements a given specification with a set of system components such as ASICs and processors, includes several key tasks such as system component allocation, functional partitioning, quality metrics estimation, and model refinement. In this work, we focus on the model refinement task which transforms a specification from an original functional model to a refined implementation model. First, we categorize several commonly used implementation models and describe a set of refinement procedures to transform a specification to each of these implementation models. We also present a set of experimental results to compare the implementation models and to demonstrate how the proposed approach can be used to explore different implementation styles.
Daniel Gajski, Smita Bakshi
ACM Trans. Design Autom. Electr. Syst.2
1996 Clock-driven performance optimization in interactive behavioral synthesis
abstract
In interactive behavioral synthesis, the designer can control the design process at every stage, including modifying the schedule of the design to improve its performance. In this paper, we present a methodology for performance optimization in interactive behavioral synthesis. Also proposed are several quality metrics and hints that can assist the user in utilizing the proposed methodology. When the user is optimizing the performance of the design, one important decision is the selection of a clock period. To facilitate clock selection by the user, we have developed an algorithm to estimate the effect of different clock periods on the execution time of the design. We have tested our methodology on several benchmarks. The experimental results support the proposed methodology by demonstrating an average improvement of 46.2% in design performance.
Hsiao-Ping Juan, Daniel Gajski, Viraphol Chaiyakul
ICCAD2
1996 Opportunities and pitfalls in HDL-based system design
abstract
This panel discusses the complexities of system designs using textual Hardware Description Languages (HDLs) such as Verilog and VHDL. As the proliferation of circuits and systems design using HDLs continues questions arise as to whether HDL-based programming provides any real productivity gains in the design of complex integrated hardware. Are there are any alternatives, such as graphical or visual formalisms that are perhaps better suited for the task? We start the discussion by examining the relevant features of today's complex hardware systems and the requirements these impose on the modeling language and methodologies.
Rajesh K. Gupta 0001, Daniel Gajski, Randy Allen, Yatin Trivedi
ICCD2
1996 An optimal clock period selection method based on slack minimization criteria
abstract
An important decision in synthesizing a hardware implementation from a behavioral description is selecting the clock period to schedule the datapath operations into control steps. Prior to scheduling, most existing behavioral synthesis systems either require the designer to specify the clock period explicitly or require that the delays of the operators used in the design be specified in multiples of the clock period. An unfavorable choice of clock period could result in operations being idle for a large portion of the clock period and, consequently, affect the performance of the synthesized design. In this article, we demonstrate the effect of clock slack on the performance of designs and present an algorithm to find a slack-minimal clock period. We prove the optimality of our method and apply it to several examples to demonstrate its effectiveness in maximizing design performance.
En-Shou Chang, Daniel Gajski, Sanjiv Narayan
ACM Trans. Design Autom. Electr. Syst.2
1996 Component selection for high-performance pipelines
abstract
The use of a realistic component library with multiple implementations of operators results in cost-efficient designs; slow components can then be used on noncritical paths and the more expensive components on only the critical paths. This paper presents a cost-optimized algorithm for selecting components and pipelining a data-flow graph, given such a library, and throughput and latency constraints. Experimental results on several large examples indicate the importance of component selection as a parameter in design exploration.
Smita Bakshi, Daniel Gajski
IEEE Trans. Very Large Scale Integr. Syst.2
1996 System design methodologies: aiming at the 100 h design cycle
abstract
As methodologies and tools for chip-level design mature, design effort becomes focused on increasingly higher levels of abstraction. We present a tutorial on a design methodology for chip and system design and present a test case that justifies the future goal of a 100 h design cycle.
Daniel Gajski, Sanjiv Narayan, Loganath Ramachandran, Frank Vahid, Peter Fung
IEEE Trans. Very Large Scale Integr. Syst.1
1995 Interfacing Incompatible Protocols Using Interface Process Generation
abstract
During system design, one or more portions of the system may be implemented with standard components that have a fixed pin structure and communication protocol. This paper described a new technique, interface process generation, for interfacing standard components that have incompatible protocols. Given an HDL description of the two protocols, we present a method to generate an interface process that allows the two protocols to communicate with each other.
Sanjiv Narayan, Daniel Gajski
DAC2
1995 SpecCharts: a VHDL front-end for embedded systems
abstract
VHDL and other hardware description languages are commonly used as specification languages during system design. However, the underlying model of those languages does not directly support the specification of embedded systems, making the task of specifying such systems tedious and error-prone. We introduce a new conceptual model, called Program-State Machines (PSM), that caters to embedded systems. We describe SpecCharts, a VHDL extension that supports capture of the PSM model. The extensions we describe can also be applied to other languages. SpecCharts can be easily incorporated into a VHDL design environment using automatic translation to VHDL. We highlight several experiments that demonstrate the advantages of significantly reduced specification time, fewer errors, and improved specification readability.>
Frank Vahid, Sanjiv Narayan, Daniel Gajski
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
1995 Performance evaluation for application-specific architectures
abstract
Performance evaluation is critical for the minimization of design cost. It consists of two parts: modeling the underlying hardware engine and evaluating the performance of the application code for the model developed in the first part. In this paper, we propose a new parameterized model for application-specific architectures and present a retargetable scheduler for performance evaluation. The model, different from those proposed previously, reflects comprehensive architectural characteristics that affect hardware parallelism. The scheduler, distinguished from previous ones, takes into account not only functional and storage unit resources but also interconnect resources during the performance evaluation. The new architecture model, together with the retargetable scheduler, enables designers to accurately evaluate the performance of a variety of ASIC and ASIP architectures.
Daniel Gajski, Alexandru Nicolau
IEEE Trans. Very Large Scale Integr. Syst.2
1995 Incremental hardware estimation during hardware/software functional partitioning
abstract
To aid in the functional partitioning of a system into interacting hardware and software components, fast yet accurate estimations of hardware size are necessary. We introduce a technique for obtaining such estimates in two orders of magnitude less time than previous approaches without sacrificing substantial accuracy, by incrementally updating a design model for a changed partition rather than re-estimating entirely.>
Frank Vahid, Daniel Gajski
IEEE Trans. Very Large Scale Integr. Syst.2
1994 Protocol Generation for Communication Channels
abstract
System-level partitioning groups processes and variables in the system speci cation into modules representing chips and memories.Communication between the modules is represented by abstract communication channels, which are merged and implemented a s a b u s to minimize interconnect.Given a set of channels, bus generation synthesizes the bus structure, by trading o the the width of the bus and the performance of the processes communicating over it.For each channel, we describe a method to generate protocols that specify the mechanism of data transfer over the bus.Protocol generation presented in this paper results in a re ned system speci cation that is simulatable.Both busgeneration and protocol-generation are demonstrated on detailed examples.
Sanjiv Narayan, Daniel Gajski
DAC2
1994 Design exploration for high-performance pipelines
Smita Bakshi, Daniel Gajski
ICCAD2
1994 Condition graphs for high-quality behavioral synthesis
Hsiao-Ping Juan, Viraphol Chaiyakul, Daniel Gajski
ICCAD3
1994 A transformation-based method for loop folding
abstract
We propose a transformation-based scheduling algorithm for the problem given a loop construct, a target initiation interval and a set of resource constraints, schedule the loop in a pipelined fashion such that the iteration time of executing an iteration of the loop is minimized. The iteration time is an important quality measure of a data path design because it affects both storage and control costs. Our algorithm first performs an As Soon As Possible Pipelined (ASAPp) scheduling regardless the resource constraint. It then resolves resource constraint violations by rescheduling some operations. The software system implementing the proposed algorithm, called Theda.Fold, can deal with behavioral loop descriptions that contain chained, multicycle and/or structural pipelined operations as well as those having data dependencies across iteration boundaries. Experiment on a number of benchmarks is reported.>
Tsing-Fa Lee, Allen C.-H. Wu, Youn-Long Lin, Daniel Gajski
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
1993 High-Level Transformations for Minimizing Syntactic Variances
abstract
Most synthesissystems generate designs from hardware de-
Viraphol Chaiyakul, Daniel Gajski, Loganath Ramachandran
DAC2
1993 Component synthesis from functional descriptions
abstract
The authors point out that in functional modeling, the functionalities of one or more components, like arithmetic/logic units, memories, and counters, are described as separate concurrent blocks. An algorithm, called the functional synthesis algorithm (FSA), for synthesis from these functional descriptions is presented. The algorithm automatically synthesizes components needed to implement a functional description while minimizing hardware costs and performance. Since a functional description uses standard operators in the hardware description language, a mismatch between the operators of the language and the functionalities provided by library components arises. FSA solves this functionality mismatch problem by pattern matching between the description and a library of function patterns. In addition, FSA clusters functions to maximally match components from a given library. Experimental results show that automated functional synthesis produces designs that are comparable to those produced by human designers.>
Elke A. Rundensteiner, Daniel Gajski, Lubomir F. Bic
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1992 Functional Synthesis Using Area and Delay Optimization
Elke A. Rundensteiner, Daniel Gajski
DAC2
1992 Specification Partitioning for System Design
Frank Vahid, Daniel Gajski
DAC2
1992 An effective methodology for functional pipelining
abstract
The problem of scheduling a loop in a pipelined fashion such that the iteration time (turnaround time) is minimized, given a loop behavior, a target initiation interval, and resource constraints, is considered. The iteration time is an important quality measure of a data path design because of its direct correlation with both the storage and the control costs. The scheduler starts with performing as-soon-as-possible-pipelined (ASAP/sub p/) scheduling without regard to the resource constraint. It then resolves the resource constraint violations, if there are any, by repeatedly rescheduling some operations.>
Tsing-Fa Lee, Allen C.-H. Wu, Daniel Gajski, Youn-Long Lin
ICCAD3
1992 Accurate layout area and delay modeling for system level design
abstract
The problem of estimating design quality measures to accurately reflect design tradeoffs and efficiently explore the design space is discussed. Specifically, interest is centered on predicting the layout area and delay of a given structural RT level design. Clearly, current RT level cost measures are highly simplified and do not reflect the real physical design. In order to establish a more realistic assessment of layout effects, a layout model which accurately and efficiently accounts for the effects of wiring and floorplanning on the area and performance layout of RT level designs is proposed. Benchmarking has shown that this model is quite accurate.>
Champaka Ramachandran, Fadi J. Kurdahi, Daniel Gajski, Allen C.-H. Wu, Viraphol Chaiyakul
ICCAD3
1992 An efficient multi-view design model for real-time interactive synthesis
abstract
An efficient multiview design model for real-time interactive synthesis of behavioral descriptions into layout data is described. A hybrid data structure which combines all of the design data needed throughout multiple levels of abstraction, including behavior, structure, and floorplan, into a single unified view is presented. A detailed time and space complexity analysis of the proposed design model is also given, showing that it provides fast updating capabilities for incremental design changes but does not require an exorbitant amount of memory space. These features make this design model ideal for user-controlled synthesis systems that support incremental design and redesign tasks. Furthermore, the simplicity of the data structure allows easy implementation, maintenance, and extensibility.>
Allen C.-H. Wu, Tedd Hadley, Daniel Gajski
ICCAD3
1992 Layout placement for sliced architecture
abstract
The authors define a new, sliced layout architecture for compilation of arbitrary schematics (netlists) into layout for CMOS technology. This sliced architecture uses over-the-cell routing on the second metal layer. The authors define three different architectures with simple folding, interleaved folding, and unrestricted folding and give algorithms for optimizing the layout area for several variants of the selected architecture. A proof demonstrating that the architecture with interleaved folding is as good as the architecture with unrestricted folding with respect to area minimization of the total layout is given. The authors also present results of random benchmarks as well as several real benchmarks.>
Lawrence L. Larmore, Daniel Gajski, Allen C.-H. Wu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1992 Partitioning algorithms for layout synthesis from register-transfer netlists
abstract
A partitioning methodology that exploits the regularity of register-transfer components is presented, and partitioning algorithms that are used to generate the floor plan are described. The partitioning algorithms not only select the layout style best suited for each component, but also consider critical paths, I/O pin locations, and connections between components. This approach improves the overall area utilization and minimizes the wire length on the critical paths.>
Allen C.-H. Wu, Daniel Gajski
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1991 System Specification and Synthesis with the SpecCharts Language
abstract
There is a need for capturing behavioral specifications of entire systems and obtaining multi-chip designs from those specifications. The authors discuss system level specification and synthesis issues, along with the unique requirements they place on a specification language. Since no current language meets those requirements, the SpecCharts language was created on top of VHDL (VHSIC Hardware Description Language). The SpecCharts language permits concise, understandable, and accurate specification of systems while supporting the concept of behavioral hierarchy, which considerably aided the specification of hardware systems modeled by the authors. Its constructs aid system level synthesis tasks such as partitioning, estimation, interface synthesis, and bus merging by permitting high level communication, and maintaining information and permitting modification at the level at which most modelers think at.>
Sanjiv Narayan, Frank Vahid, Daniel Gajski
ICCAD3
1991 An Algorithm for Component Selection in Performance Optimized Scheduling
abstract
The authors describe a novel algorithm that combines the hardware scheduling and component selection phases for high level synthesis. The algorithm improves on previous work in scheduling, by being able to simultaneously select components from a given library. This enlarges the design space, resulting in better optimized designs. Experimental results on the elliptic filter benchmark demonstrate that exploiting all available components in the library results in designs with smaller area compared to designs produced by scheduling with a single implementation for each component type.>
Loganath Ramachandran, Daniel Gajski
ICCAD2
1991 Obtaining Functionally Equivalent Simulations using VHDL and a Time-Shift Transformation
abstract
It is pointed out that many translation schemes from domain-specific languages to supposedly functionally equivalent VDHL (VHSIC hardware description language) have been developed as an approach to simulation. However, due to a subtle theoretical limitation to this approach, functionally equivalent VHDL cannot be created for the general case, making such translations an unsound technique. The authors propose an alternative approach which strives instead for functionally equivalent simulation, while still taking advantage of VHDL simulators. This method uses a novel time-shift transformation in conjunction with any translation scheme, making correct simulations easily obtainable. This bridges the gap to a sound and advantageous use of VHDL as a tool for simulating domain-specific languages.>
Frank Vahid, Daniel Gajski
ICCAD2
1991 Layout-Area Models for High-Level Synthesis
abstract
The authors propose a novel layout area model for quality measures in high-level synthesis. The model is proposed for two commonly used datapath and control layout architectures. Except for macrocells (PLAs), the proposed models formulate layout area as a function of transistors and routing tracks which can be computed in O(n log n) time complexity, where n is the number of nets in the netlist. This allows one to explore design space in high-level synthesis rapidly and efficiently. The authors have tested their layout models on the widely used elliptic-filter benchmark. The results show that these models can more accurately predict layout areas than models based on the number and size of registers and multiplexers.>
Allen C.-H. Wu, Viraphol Chaiyakul, Daniel Gajski
ICCAD3
1990 An Intelligent Component Database for Behavioral Synthesis
abstract
This paper describes an intelligent component database system that delivers components to synthesis tools when given a set of attributes and constraints. Requirements of a component server are defined and an implementation is described. Our experiments demonstrate that such a component sever can replace component libraries and component catalogs with hundreds of pages.
Gwo-Dong Chen, Daniel Gajski
DAC2
1990 An Intermediate Representation for Behavioral Synthesis
abstract
This paper describes an intermediate representation for behavioral and structural designs that is based on annotated state tables. It facilitates user control of the synthesis process by allowing specification of partially design structures, and a mixture of behavior, structure and user specified bindings between the abstract behavior and the structure. The format's general model allows the capture of synchronous and asynchronous behavior, and permits hierarchical descriptions with concurrency. The format is easily translated to VHDL for simulation at each stage of the design process. It therefore complements a good simulation language (VHDL) by providing an excellent input path for behavioral and register-transfer synthesis. The format's simple and uniform syntax allows it to be used both as an intermediate exchange format for various behavioral synthesis tools, and as a graphical tabular interface for the user, thereby allowing a natural medium for automatic or manual refinement of the design.
Nikil Dutt, Tedd Hadley, Daniel Gajski
DAC3
1990 Percolation Based Synthesis
abstract
A new approach called Percolation Based Synthesis for the scheduling phase of High Level Synthesis (HLS) is presented. We discuss some new techniques (which are implemented in our tools) for compaction of flow graphs beyond basic blocks limits, which can produce order of magnitude speed ups versus serial execution. Our algorithm applies to programs with conditional jumps, loops and multicycle pipelined operations. In order to schedule under resource constraints we start by first finding the optimal schedule (without constraints) and then add heuristics to map the optimal schedule onto the given system. We argue that starting from an optimal schedule is one of the most important factors in scheduling because it offers the user flexibility to tune the heuristics and gives him a good bound for the resource constrained schedule. This scheduling algorithm is integrated with synthesis tool which uses VHDL as input description and produces a structural netlist of generic register-transfer components and a unit based control table as output. We show that our algorithm obtains better results than previously published algorithms.
Roni Potasman, Joseph Lis, Alexandru Nicolau, Daniel Gajski
DAC4
1990 The Component Sythesis Algorithm: Technology Mapping for Register Transfer Descriptions
abstract
In functional modeling, one or more micro-architecture components are described as separate concurrent blocks. An algorithm, called the component synthesis algorithm, is presented that automatically synthesizes micro-architecture components for a functional description. Experimental results show that the automated functional synthesis is comparable to the performance of human designers.>
Elke A. Rundensteiner, Daniel Gajski, Lubomir F. Bic
ICCAD2
1990 Partitioning Algorithms for Layout Synthesis from Register-Transfer Netlists
abstract
A sliced-layout architecture is presented to alleviate the problems of the general bit-sliced layouts. Also described are partitioning algorithms that are used to generate the floorplan for this layout architecture. The partitioning algorithms not only select the best suited layout style for each component, but also consider critical paths, I/O pin locations, and connections between logic blocks. This approach improves the overall area utilization and minimizes the total wire length.>
Allen C.-H. Wu, Daniel Gajski
ICCAD2
1990 The Role of Learning in Logic Synthesis
abstract
The goal of logic synthesis is to obtain high-quality designs from specifications. Current approaches to logic synthesis often trade off design quality for technology independence. In this paper, we present a model of logic synthesis that uses technology-specific design rules and extends rule-based search to functional decomposition and technology mapping. While this model improves design quality by taking advantage of the target technology, it is not robust to technology changes. To improve robustness, we augment the model with two learning components: one for acquiring rules that make use of physical cells in a technology library, and another for acquiring rules that make use of appropriate design styles. These components are related to work in the learning of macro-operators and explanation-based learning.
James R. Kipps, Daniel Gajski
Int. J. Pattern Recognit. Artif. Intell.2
1990 Chippe: a system for constraint driven behavioral synthesis
abstract
The Chippe system for constrained behavioural architecture synthesis uses a novel closed-loop design iteration technique in which the present state of the design is analyzed with respect to the goals and then modified for the next iteration. In this way the design state is iteratively driven towards meeting the global constraints imposed by the designer. The design synthesis is performed by a set of algorithmic tools specially constructed to permit the imposition of a wide variety of local constraints, and by a rule-based system which makes design analysis and modification decisions to set these local constraints. Key to these decisions is a design evaluator which examines the present state of the design and interactively reports its findings to the rule-based system. The closed-loop iteration strategy, the interaction between the rule base and the tools, and the evaluation performed to support the design decisions are detailed. Also presented are results from sample designs, including designs for the TMS320 digital signal processor chip.>
Forrest Brewer, Daniel Gajski
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1990 Hypertool: A Programming Aid for Message-Passing Systems
abstract
Programming assistance, automation concepts, and their application to a message-passing system program development tool called Hypertool are discussed. Hypertool performs scheduling and handles the communication primitive insertion automatically, thereby increasing productivity and eliminating synchronization errors. Two algorithms, based on the critical-path method, are presented for scheduling processes statically. Hypertool also generates the performance estimates and other program quality measures to help programmers improve their algorithms and programs.>
Min-You Wu, Daniel Gajski
IEEE Trans. Parallel Distributed Syst.2
1989 Designer Controlled Behavioral Synthesis
abstract
This paper describes features of EXEL, a graphic language that gives the designer control over the behavioral synthesis process. Control is achieved by allowing the designer to partially specify the structural design into which the description is going to be compiled, or by binding desired variables and operators to particular components or connections, and binding desired operations to particular states of the final design. EXEL's compiler runs on SUN-3 workstations and is written in C and SUNVIEW.
Nikil Dutt, Daniel Gajski
DAC2
1989 VHDL Synthesis Using Structured Modeling
abstract
This paper describes the use of VHDL in a behavioral synthesis system. A structured modeling methodology is presented which suggests standard practices for writing VHDL descriptions which span a variety of design models. The VHDL Synthesis System (VSS) processes each of these input descriptions and produces a structural description of generic components.
Joseph Lis, Daniel Gajski
DAC2
1989 Hypertool: A Programming Aid for Multicomputers
Min-You Wu, Daniel Gajski
ICPP (2)2
1989 Power routing in channelless floorplan layouts
Chidchanok Lursinsap, Daniel Gajski
Integr.2
1989 Computer-aided programming for message-passing systems: problems and solutions
abstract
As the number of processors and the complexity of problems to be solved increase, programming multiprocessing systems becomes more difficult and error prone. Program development tools are necessary since programmers are not able to develop complex parallel programs efficiently. Parallel models of computation, parallelization problems, and tools for computer-aided programming (CAP) are discussed. As an example, a CAP tool that performs scheduling and inserts communication primitives automatically is described. It also generates the performance estimates and other program quality measures to help programmers in improving their algorithms and programs.>
Min-You Wu, Daniel Gajski
Proc. IEEE2
1988 MILO: A Microarchitecture and Logic Optimizer
Nels Vander Zanden, Daniel Gajski
DAC2
1988 Synthesis from VHDL
abstract
The VHDL Synthesis System (VSS) uses VHDL dataflow or behavioral descriptions as input and outputs a structural description of generic components. This structural description is converted into a schematic and captured by the microarchitecture and logic optimization system for technology mapping and constraint-driven optimization. VSS allows a designer to modify the compiled design by changing the input description, selecting optimization and mapping strategies, or graphically changing the generated design schematic. Redesign to new technologies can be accomplished by changing only the component library.>
Joseph Lis, Daniel Gajski
ICCD2
1988 LES: a layout expert system
abstract
The LES expert system for layout generation of random logic modules in a hierarchical CMOS VLSI design system is described. It applies a combination of rule- and algorithmic-based techniques on a novel layout style. The layout style utilizes silicon area more efficiently than a previously developed style. Experimental results have demonstrated the superiority of this expert system against various standard-cell systems and its competitiveness with human designers.>
Youn-Long Lin, Daniel Gajski
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1988 A technique for pull-up transistor folding
abstract
The authors consider the constraint limiting multiple folding in programmable logic array (PLA) layouts imposed by the layout architecture which positions pull-up transistors on the boundary of the cell and uses another metal layer to connect pull-ups to terms inside the PLA. The general PLA architecture supports only input- and output-port folding by sharing them in either the same column or the same row. Term folding is allowed only with drastic changes in the layout architecture. Folding optimization algorithms, generally, have not considered a pull-up transistor placement as a constraint. A layout architecture is introduced and a technique for pull-up transistor folding based on a weighted-graph model is presented. The architecture supports also both I/O and term foldings. In comparison with other architectures the described architecture allows significant area improvement.>
Chidchanok Lursinsap, Daniel Gajski
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1988 A programming aid for hypercube architectures
Min-You Wu, Daniel Gajski
J. Supercomput.2
1987 Knowledge Based Control in Micro-Architecture Design
abstract
This paper describes the principles and implementation of design-process control in a micro-architecture compiler. The knowledge-base relies on both local and global evaluations to determine strategies to achieve global goals and then implements those strategies by manipulating hardware allocations and search heuristics. A system overview and annotated sample run are presented.
Forrest Brewer, Daniel Gajski
DAC2
1987 LES: A Layout Expert System
abstract
In this paper we describe an expert system for layout generation in a hierarchical VLSI design system. It applies a combination of rule- and algorithmic-based techniques on a new layout style. Experimental results have demonstrated the superiority of this expert system against various standard-cell systems and its competitiveness with human designers.
Youn-Long Lin, Daniel Gajski
DAC2
1987 Improving a PLA Area by Pull-Up Transistor Folding
abstract
The constraints limiting multiple PLA folding are the positions of pull-up transistors (Tpu) and the layout architecture which supports only I/O folding. In this paper, we present a novel architecture supporting multiple I/O, term, and pull-up foldings with through-the-cell net routing. The pull-up folding allows implementation of multiple level Boolean functions. Our architecture model achieves a better densities in comparison with PLA model.
Chidchanok Lursinsap, Daniel Gajski
DAC2
1987 Design Tools for Intelligent Silicon Compilation
abstract
This paper describes behavioral compilation tools built for use in an intelligent silicon compiler. These tools allow the user or an expert system to compile behavioral descriptions to a register transfer level under user-imposed constraints. A flexible design model offers a combination of design features previously unavailable in behavioral compilers such as multicycle, chained, and pipelined function units along with the ability to choose between bus- and mux-based connectivity models. Furthermore, we present a new design strategy that allows easy exploratory design and we describe algorithms for state synthesis and connectivity binding that achieved higher quality designs than previous systems on selected benchmarks. The code for this project is run under 42 BSD Unix on a VAX 11 780 and is written in C.
Barry M. Pangrle, Daniel Gajski
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1986 An expert-system paradigm for design
abstract
Article Free Access Share on An expert-system paradigm for design Authors: Forrest D. Brewer Dept. of Computer Science, University of Illinois at Urbana-Champaign, 1304 West Springfield Ave., Urbana, Illinois Dept. of Computer Science, University of Illinois at Urbana-Champaign, 1304 West Springfield Ave., Urbana, IllinoisView Profile , Daniel D. Gajski Dept. of Computer Science, University of Illinois at Urbana-Champaign, 1304 West Springfield Ave., Urbana, Illinois Dept. of Computer Science, University of Illinois at Urbana-Champaign, 1304 West Springfield Ave., Urbana, IllinoisView Profile Authors Info & Claims DAC '86: Proceedings of the 23rd ACM/IEEE Design Automation ConferenceJuly 1986 Pages 62–68Published:02 July 1986Publication History 10citation250DownloadsMetricsTotal Citations10Total Downloads250Last 12 Months4Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Forrest Brewer, Daniel Gajski
DAC2
1986 Flow graph representation
Alex Orailoglu, Daniel Gajski
DAC2
1986 CAMP: A Programming Aide for Multiprocessors
Jih-Kwon Peir, Daniel Gajski
ICPP2
1986 A Heuristic for Suffix Solutions
abstract
The suffix problem has appeared in solutions of recurrence systems for parallel and pipelined machines and more recently in the design of gate and silicon compilers. In this paper we present two algorithms. The first algorithm generates parallel suffix solutions with minimum cost for a given length, time delay, availability of initial values, and fanout. This algorithm generates a minimal solution for any length n and depth range from log2 n to n. The second algorithm reduces the size of the solutions generated by the first algorithm.
Avinoam Bilgory, Daniel Gajski
IEEE Trans. Computers2
1985 Decomposition of logic networks into silicon
abstract
This paper describes a module compiler for decomposing arbitrary functional units of any complexity into abstract cells for customized VLSI layouts. The compiler takes the description of a functional unit as input and builds a dependence graph representation. The graph is then partitioned and the nodes are packed into abstract cell output descriptions. The algorithm will tailor the design to a given area and aspect ratio. Routing is done automatically through the cells.
Steven T. Healey, Daniel Gajski
DAC2
1985 Comparison of five multiprocessor systems
Daniel Gajski, Jih-Kwon Peir
Parallel Comput.1
1984 Silicon compilers and expert systems for VLSI
Daniel Gajski
DAC1
1984 Cell compilation with constraints
Chidchanok Lursinsap, Daniel Gajski
DAC2
1984 Microprocessor synthesis
Vijay K. Raj, Barry M. Pangrle, Daniel Gajski
DAC3
1984 Fast Execution of Loops With IF Statements
abstract
In this paper we show how to execute in parallel loops containing IF statements. We give an architectural model of parallel computation and describe the design of a hardware Boolean Recurrence Solver. Our method of handling such loops is then compared with those used by some of the existing supercomputers.
Utpal Banerjee, Daniel Gajski
ISCA2
1984 A Parallel Pipelined Relational Query Processor: An Architectural Overview
abstract
This paper outlines the overall architecture of a query processor for relational queries and describes the design and control of its major processing modules. The query processor consists of only four processing modules and a number of random-access memory modules. Each processing module processes tuples of relations in a bit-serial, tuple-parallel manner for each of the primitive database operations which comprise a complex relational query. The query processor is designed to be manufacturable using existing VLSI technology, and to support in a uniform manner both the numeric and nonnumeric processing requirements a high-level query language like SQL presents.
Daniel Gajski, Won Kim 0001, Shinya Fushimi
ISCA1
1984 Fast Execution of Loops with IF Statements
abstract
A parallel method of execution for a certain class of loops containing IF statements is described. We replace a given loop by an equivalent set of five loops, four of which are vectorizable; the fifth loop is executed in hardware as a Boolean recurrence. The proposed architecture handles all loops that produce recurrences with order ≤m, a hardware parameter.
Utpal Banerjee, Daniel Gajski
IEEE Trans. Computers2
1984 A Parallel Pipelined Relational Query Processor
abstract
This paper presents the design of a relational query processor. The query processor consists of only four processing PIPEs and a number of random-access memory modules. Each PIPE processes tuples of relations in a bit-serial, tuple-parallel manner for each of the primitive database operations which comprise a complex relational query. The design of the query processor meets three major objectives: the query processor must be manufacturable using existing and near-term LSI (VLSI) technology; it must support in a uniform manner both the numeric and nonnumeric processing requirements a high-level user interface like SQL presents; and it must support the query-processing strategy derived in the query optimizer to satisfy certain system-wide performance optimality criteria.
Won Kim 0001, Daniel Gajski, David J. Kuck
ACM Trans. Database Syst.2
1983 Cedar : A Large Scale Multiprocessor
Daniel Gajski, David J. Kuck, Duncan H. Lawrie, Ahmed H. Sameh
ICPP1
1982 Iterative algorithms for tridiagonal matrices on a WSI-multiprocessor
Daniel Gajski, Ahmed H. Sameh, John A. Wisniewski
ICPP1
1981 Automatic generation of cells for recurrence structures
Avinoam Bilgory, Daniel Gajski
DAC2
1981 Design of Testable Structures Defined by Simple Loops
abstract
A methodology is given for generating combinational structures from high-level descriptions (using assignment statements, "if" statements, and single-nested loops) of register-transfer (RT) level operators. The generated structures are cellular, and are interconnected in a tree structure. A general algorithm is given to test cellular tree structures with a test length which grows only linearly with the size of the tree. It is proved that this test length is optimal to within a constant factor. Ways of making the structures self-checking are also indicated.
Jacob A. Abraham, Daniel Gajski
IEEE Trans. Computers2
1981 An Algorithm for Solving Linear Recurrence Systems on Parallel and Pipelined Machines
abstract
A new algorithm for the solution of linear recurrence systems on parallel or pipelined computers is described. Time bounds, speed-up and efficiency for SIMD and MIMD computers with fixed number of arithmetic elements (AE's), as well as for pipelined computers with fixed number of stages per operation, are obtained. The model of each computer is discussed in detail to explain better performance of the pipelined model. A simple modification in the design of AE's for parallel computers makes parallel model superior.
Daniel Gajski
IEEE Trans. Computers1
1980 Automatic design with dependence graphs
abstract
A design automation system for the design of digital systems from a high-level algorithmic description is proposed. The definition of the data-dependence graph and techniques for performing transformations that lead to optimization of hardware are described. The system can be used on several levels of design with the VLSI layout level given particular emphasis.
Albert E. Casavant, Daniel Gajski, David J. Kuck
DAC2
1980 Parallel Compressors
abstract
A subclass of generalized parallel counters, called parallel compressors, is introduced in this correspondence. Under present-day packaging technology, parallel compressors with their higher compression ratio and fewer input/output pins are more efficient in multiple operand addition and multiplication than parallel counters. Cost and time bounds are obtained for schemes using parallel compressors for reduction of N summands to m summands. Furthermore, a method for synthesizing large parallel counters using only one type of parallel compressor is given.
Daniel Gajski
IEEE Trans. Computers1
1978 Design of arithmetic elements for Burroughs Scientific Processor
abstract
The design criteria and implementation of the Arithmetic Element (AE) of the Burroughs Scientific Processor, a vector machine intended for scientific computation requiring speed of up to 50 million floating-point operations per second, is discussed. An array of 16 AEs operate in lockstep mode, executing the same instruction on 16 sets of data. The 16 AEs are one stage in a pipeline which consists of 17 memory modules, an input alignment network, and an output alignment network. The AE itself is not pipelined. It can perform over one hundred different operations including a floating-point addition, subtraction and multiplication, division, square root, among the others. Eight registers are provided for the storage of intermediate values and results. Modulo 3 residue arithmetic is used for checking hardware failures.
Daniel Gajski, Louis P. Rubinfield
IEEE Symposium on Computer Arithmetic1