Ruchir Puri

dblp:84/1742 · DBLP profile ↗
← Back
65ranked-venue papers
24as first author
5since 2021 · last 2025
0009-0006-8803-7079ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 57 · 22 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
29 papers
Electronic design automation · 34% Integrated circuit design · 24% Memory systems · 14%
Artificial intelligence
2 papers
Reinforcement learning · 70% Trustworthy machine learning · 21% Deep learning architectures and training · 9%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Theoretical computer science
3 papers
Algorithms and data structures · 99% Mathematical optimization · 1% Graph algorithms and graph theory · 1%

Topics — the 30 heaviest of 60, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
agent evaluation
0.912025
ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks · ICML 2025
Electronic design automation
logic synthesis
0.8102016
Polynomial Time Algorithm for Area and Power Efficient Adder Synthesis in High-Performance Designs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Towards Optimal Performance-Area Trade-Off in Adders by Synthesis of Parallel Prefix Structures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Towards optimal performance-area trade-off in adders by synthesis of parallel prefix structures · DAC 2013
Program synthesis and code generation
code generation from natural language
0.712023
Invited: Automated Code generation for Information Technology Tasks in YAML through Large Language Models · DAC 2023
Integrated circuit design › digital circuit design › arithmetic circuit design
adder design
0.632016
Polynomial Time Algorithm for Area and Power Efficient Adder Synthesis in High-Performance Designs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Towards Optimal Performance-Area Trade-Off in Adders by Synthesis of Parallel Prefix Structures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Towards optimal performance-area trade-off in adders by synthesis of parallel prefix structures · DAC 2013
Integrated circuit design › digital arithmetic circuits › parallel adder
parallel prefix adder
0.422016
Polynomial Time Algorithm for Area and Power Efficient Adder Synthesis in High-Performance Designs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Towards Optimal Performance-Area Trade-Off in Adders by Synthesis of Parallel Prefix Structures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.412019
The Next Generation of Deep Learning Hardware: Analog Computing · Proc. IEEE 2019
Memory systems › processing-in-memory
non-volatile memory in-memory computing
0.412019
The Next Generation of Deep Learning Hardware: Analog Computing · Proc. IEEE 2019
Memory systems
processing-in-memory
0.412019
The Next Generation of Deep Learning Hardware: Analog Computing · Proc. IEEE 2019
Electronic design automation
physical design
0.342010
History-based VLSI legalization using network flow · DAC 2010
Track Routing and Optimization for Yield · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
TROY: Track Router with Yield-driven Wire Planning · DAC 2007
Machine learning › Trustworthy machine learning
AI safety
0.312025
ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks · ICML 2025
Electronic design automation
design for manufacturability
0.232008
Fast Dummy-Fill Density Analysis With Coupling Constraints · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Is Your Layout-Density Verification Exact? - A Fast Exact Deep Submicrometer Density Calculation Algorithm · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
TROY: Track Router with Yield-driven Wire Planning · DAC 2007
Emerging computing paradigms › neuromorphic computing
associative memory
0.212015
Emerging Trends in Design and Applications of Memory-Based Computing and Content-Addressable Memories · Proc. IEEE 2015
Memory systems
content-addressable memory
0.212015
Emerging Trends in Design and Applications of Memory-Based Computing and Content-Addressable Memories · Proc. IEEE 2015
Parallel and multicore computing › parallel algorithms › sorting
parallel sorting
0.212015
PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort · Proc. VLDB Endow. 2015
Algorithms and data structures › sequence algorithms › sorting › integer sorting
radix sort
0.212015
PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort · Proc. VLDB Endow. 2015
Algorithms and data structures › sequence algorithms
sorting
0.212015
PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort · Proc. VLDB Endow. 2015
Integrated circuit design
digital circuit design
0.222013
Towards optimal performance-area trade-off in adders by synthesis of parallel prefix structures · DAC 2013
SOI Digital CMOS VLSI - a Design Perspective · DAC 1999
Electronic design automation
physical verification
0.222008
Fast Dummy-Fill Density Analysis With Coupling Constraints · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Is Your Layout-Density Verification Exact? - A Fast Exact Deep Submicrometer Density Calculation Algorithm · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Electronic design automation › physical design › routing › detailed routing
track routing
0.222008
Track Routing and Optimization for Yield · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
TROY: Track Router with Yield-driven Wire Planning · DAC 2007
Integrated circuit design
low-power circuit design
0.152008
Keeping hot chips cool · DAC 2005
Pushing ASIC performance in a power envelope · DAC 2003
Keeping hot chips cool: are IC thermal problems hot air? · DAC 2008
Integrated circuit design
technology scaling
0.122011
Moore's Law: another casualty of the financial meltdown? · DAC 2009
Design, CAD and technology challenges for future processors: 3D perspectives · DAC 2011
Energy-efficient computing
leakage power reduction
0.122006
Gain-based technology mapping for minimum runtime leakage under input vector uncertainty · DAC 2006
Keeping hot chips cool · DAC 2005
Electronic design automation › physical design
legalization
0.112010
History-based VLSI legalization using network flow · DAC 2010
Electronic design automation › physical design
placement
0.112010
History-based VLSI legalization using network flow · DAC 2010
Electronic design automation › design for manufacturability › design for yield
yield enhancement
0.122008
TROY: Track Router with Yield-driven Wire Planning · DAC 2007
Fast Dummy-Fill Density Analysis With Coupling Constraints · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2008
Energy-efficient computing › power management › system-level power management
system-level power optimization
0.112009
From milliwatts to megawatts: system level power challenge · DAC 2009
Integrated circuit design
3d integration
0.132011
Design, CAD and technology challenges for future processors: 3D perspectives · DAC 2011
Moore's Law: another casualty of the financial meltdown? · DAC 2009
Interconnects in the Third Dimension: Design Challenges for 3D ICs · DAC 2007
Energy-efficient computing
thermal management
0.112008
Keeping hot chips cool: are IC thermal problems hot air? · DAC 2008
Interconnection networks and networks-on-chip › die-to-die interconnect
3d interconnect
0.112007
Interconnects in the Third Dimension: Design Challenges for 3D ICs · DAC 2007
Electronic design automation › physical design
routing
0.112007
TROY: Track Router with Yield-driven Wire Planning · DAC 2007

Methods — techniques the papers use, named apart from their topics

large language model · 2.2benchmarking · 1.7non-volatile memory · 0.8approximate computing · 0.8analog computing · 0.8distribution-adaptive load balancing · 0.4prefix graph synthesis · 0.4prefix node cloning · 0.2polynomial-time algorithm · 0.2pareto-optimal design space exploration · 0.2speculative permutation · 0.2recursive clique extraction · 0.0prime compatibles · 0.0lower bound pruning · 0.0local search · 0.0graph partitioning · 0.0fail-first heuristic · 0.0
YearPublicationVenuePosition
2025 Agent Trajectory Explorer: Visualizing and Providing Feedback on Agent Trajectories
abstract
Agentic systems interleave large language model (LLM) reasoning, tool usage, and tool observations over multiple iterations to tackle complex tasks. The raw data from an agent's problem-solving process (the agents' trajectory) is not an ideal format for human analysis and oversight. There is a need for tooling that converts this primary data into an easily navigable and understandable visual format for better human feedback. To address this opportunity, we developed the Agent Trajectory Explorer, a tool designed to help AI developers and researchers visualize, annotate, and demonstrate agent behavior.
Michael Desmond, Ibrahim Ibrahim, James M. Johnson, Avirup Sil, Justin MacNair, Ruchir Puri
AAAI7
2025 ITBench: Evaluating AI Agents across Diverse Real-World IT Automation Tasks
abstract
Realizing the vision of using AI agents to automate critical IT tasks depends on the ability to measure and understand effectiveness of proposed solutions. We introduce ITBench, a framework that offers a systematic methodology for benchmarking AI agents to address real-world IT automation tasks. Our initial release targets three key areas: Site Reliability Engineering (SRE), Compliance and Security Operations (CISO), and Financial Operations (FinOps). The design enables AI researchers to understand the challenges and opportunities of AI agents for IT automation with push-button workflows and interpretable metrics. IT-Bench includes an initial set of 102 real-world scenarios, which can be easily extended by community contributions. Our results show that agents powered by state-of-the-art models resolve only 11.4% of SRE scenarios, 25.2% of CISO scenarios, and 25.8% of FinOps scenarios (excluding anomaly detection). For FinOps-specific anomaly detection (AD) scenarios, AI agents achieve an F1 score of 0.35. We expect ITBench to be a key enabler of AI-driven IT automation that is correct, safe, and fast. IT-Bench, along with a leaderboard and sample agent implementations, is available at https://github.com/ibm/itbench.
Saurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa, Yinfang Chen, Jackson Clark, Bhavya, Mudit Verma, Hirokuni Kitahara, Noah Zheutlin, Saki Takano, Divya Pathak, Felix George, Xinbo Wu, Bekir O. Turkkan, Gerard Vanloo, Michael Nidd, Oishik Chatterjee, Pranjal Gupta, Suranjana Samanta, Pooja Aggarwal, Rong Lee, Jae-wook Ahn, Debanjana Kar, Amit M. Paradkar, Yu Deng 0004, Pratibha Moogi, Prateeti Mohapatra, Naoki Abe, Chandrasekhar Narayanaswami 0001, Tianyin Xu, Lav R. Varshney, Ruchi Mahindru, Anca Sailer, Larisa Shwartz, Daby M. Sow, Nicholas C. Fuller, Ruchir Puri
ICML40
2024 Engineering the Future of IC Design with AI
abstract
Software and Semiconductors are two fundamental technologies that have become woven into every aspect of our society, and it will be fair to say that "Software and Semiconductors have eaten the world". More recently, advances in AI are starting to transform every aspect of our society as well. These are three tectonic forces of transformation - "AI", -Software", and "Semiconductors" which are colliding together resulting in a seismic shift - a future where both software and semiconductor chips themselves will be designed, optimized, and operated by AI - pushing us towards a future where "Computers can program themselves!". In this talk, we will discuss these forces of "AI for Chips and Code" and how the future of Semiconductor chip design and software engineering is being redefined by AI.
Ruchir Puri
ISPD1
2023 Invited: Automated Code generation for Information Technology Tasks in YAML through Large Language Models
abstract
The recent improvement in code generation capabilities due to the use of large language models has mainly benefited general purpose programming languages. Domain specific languages, such as the ones used for IT Automation, received far less attention, despite involving many active developers and being an essential component of modern cloud platforms. This work focuses on the generation of Ansible YAML, a widely used markup language for IT Automation. We present Ansible Wisdom, a natural-language to Ansible YAML code generation tool, aimed at improving IT automation productivity. Results show that Ansible Wisdom can accurately generate Ansible script from natural language prompts with performance comparable or better than existing state of the art code generation models.
Saurabh Pujar, Luca Buratti, Nicolas Dupuis, Burn L. Lewis, Sahil Suneja, Atin Sood, Ganesh Nalawade, Alessandro Morari, Ruchir Puri
DAC11
2021 Engineering the Future of AI for the Enterprises : Keynote 4
abstract
Summary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. Recent advances in AI are starting to transform every aspect of our society from healthcare, manufacturing, environment, and beyond. Future of AI for enterprises will be engineered with success along three foundational dimensions. We will dive deeper along these dimensions - Automation of AI; Trust of AI; and Scaling of AI - and conclude with the opportunities and challenges of AI for businesses.
Ruchir Puri
SERVICES1
2019 Bias Mitigation Post-processing for Individual and Group Fairness
abstract
Whereas previous post-processing approaches for increasing the fairness of predictions of biased classifiers address only group fairness, we propose a method for increasing both individual and group fairness. Our novel framework includes an individual bias detector used to prioritize data samples in a bias mitigation algorithm aiming to improve the group fairness measure of disparate impact. We show superior performance to previous work in the combination of classification accuracy, individual fairness and group fairness on several real-world datasets in applications such as credit, employment, and criminal justice.
Pranay Lohia, Karthikeyan Natesan Ramamurthy, Manish Bhide, Diptikalyan Saha, Kush R. Varshney, Ruchir Puri
ICASSP6
2019 The Next Generation of Deep Learning Hardware: Analog Computing
abstract
Initially developed for gaming and 3-D rendering, graphics processing units (GPUs) were recognized to be a good fit to accelerate deep learning training. Its simple mathematical structure can easily be parallelized and can therefore take advantage of GPUs in a natural way. Further progress in compute efficiency for deep learning training can be made by exploiting the more random and approximate nature of deep learning work flows. In the digital space that means to trade off numerical precision for accuracy at the benefit of compute efficiency. It also opens the possibility to revisit analog computing, which is intrinsically noisy, to execute the matrix operations for deep learning in constant time on arrays of nonvolatile memories. To take full advantage of this in-memory compute paradigm, current nonvolatile memory materials are of limited use. A detailed analysis and design guidelines how these materials need to be reengineered for optimal performance in the deep learning space shows a strong deviation from the materials used in memory applications.
Wilfried Haensch, Tayfun Gokmen, Ruchir Puri
Proc. IEEE3
2017 Memory-Centric Reconfigurable Accelerator for Classification and Machine Learning Applications
abstract
Big Data refers to the growing challenge of turning massive, often unstructured datasets into meaningful, organized, and actionable data. As datasets grow from petabytes to exabytes and beyond, it becomes increasingly difficult to run advanced analytics, especially Machine Learning (ML) applications, in a reasonable time and on a practical power budget using traditional architectures. Previous work has focused on accelerating analytics readily implemented as SQL queries on data-parallel platforms, generally using off-the-shelf CPUs and General Purpose Graphics Processing Units (GPGPUs) for computation or acceleration. However, these systems are general-purpose and still require a vast amount of data transfer between the storage devices and computing elements, thus limiting the system efficiency. As an alternative, this article presents a reconfigurable memory-centric advanced analytics accelerator that operates at the last level of memory and dramatically reduces energy required for data transfer. We functionally validate the framework using an FPGA-based hardware emulation platform and three representative applications: Naïve Bayesian Classification, Convolutional Neural Networks, and k-Means Clustering. Results are compared with implementations on a modern CPU and workstation GPGPU. Finally, the use of in-memory dataset decompression to further reduce data transfer volume is investigated. With these techniques, the system achieves an average energy efficiency improvement of 74× and 212× over GPU and single-threaded CPU, respectively, while dataset compression is shown to improve overall efficiency by an additional 1.8× on average.
Robert Karam, Somnath Paul, Ruchir Puri, Swarup Bhunia
ACM J. Emerg. Technol. Comput. Syst.3
2016 Polynomial Time Algorithm for Area and Power Efficient Adder Synthesis in High-Performance Designs
abstract
Adders are the most fundamental arithmetic units, and often on the timing critical paths of microprocessors. Among various adder configurations, parallel prefix adders provide the best performance vs. power/area trade-off, especially for higher bit-widths. With aggressive technology scaling, the performance of a parallel prefix adder, in addition to the dependence on the logic-level, is determined by wire-length and congestion which can be mitigated by adjusting fan-out. This paper proposes a polynomial-time algorithm to synthesize n bit parallel prefix adders targeting the minimization of the size of the prefix graph with log2n logic level and any arbitrary fan-out restriction. A structure aware prefix node cloning is then applied to the resultant prefix adder solutions to further optimize the size of the prefix graphs. The design space exploration by our approach provides a set of pareto-optimal solutions for delay vs. power trade-off, and these pareto-optimal solutions can be used in high-performance designs instead of picking from a fixed library (Kogge-Stone, Sklansky, etc.). Experimental results demonstrate that our approach: 1) excels highly competitive industry standard Synopsys design compiler adder, regular adders such as Sklansky adder and Kogge-Stone adder, and a highly runtime/memory intensive recent algorithm in 32 nm technology node and 2) improves performance/area over even 64 bit custom designed adders targeting 22 nm technology library and implemented in an industrial high-performance design.
Subhendu Roy, Mihir R. Choudhury, Ruchir Puri, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2016 Energy-Efficient Adaptive Hardware Accelerator for Text Mining Application Kernels
abstract
Text mining is a growing field of applications, which enables the analysis of large text data sets using statistical methods. In recent years, exponential increase in the size of these data sets has strained existing systems, requiring more computing power, server hardware, networking interconnects, and power consumption. For practical reasons, this trend cannot continue in the future. Instead, we propose a reconfigurable hardware accelerator designed for text analytics systems, which can simultaneously improve performance and reduce power consumption. Situated near the last level of memory, it mitigates the need for high-bandwidth processor-to-memory connections, instead capitalizing on close data proximity, massively parallel operation, and analytic-inspired functional units to maximize energy efficiency, while remaining flexible to easily map common text analytic kernels. A field-programmable gate array-based emulation framework demonstrates the functional correctness of the system, and a full eight-core accelerator is synthesized for power, area, and delay estimates. The accelerator can achieve two to three orders of magnitude improvement in energy efficiency versus CPU and general-purpose graphics processing unit (GPU) for various text mining kernels. As a case study, we demonstrate how indexing performance of Lucene, a popular text search and analytics platform, can be improved by an average of 70% over CPU and GPU while significantly reducing data transfer energy and latency.
Robert Karam, Ruchir Puri, Swarup Bhunia
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Polynomial time algorithm for area and power efficient adder synthesis in high-performance designs
abstract
Adders are the most fundamental arithmetic units, and often on the timing critical paths of microprocessors. Among various adder configurations, parallel prefix structures provide the high performance adders for higher bit-widths. With aggressive technology scaling, the performance of a parallel prefix adder, in addition to the dependence on the logic-level, is determined by wire-length and congestion which can be mitigated by adjusting fan-out. This paper proposes a polynomial-time algorithm to synthesize n bit parallel prefix adders targeting the minimization of the size of the prefix graph with log2n logic level and any arbitrary fan-out restriction. The design space exploration by our algorithm provides a set of Pareto-optimal solutions for delay vs. power trade-off, and these Pareto-optimal solutions can be used in high-performance designs instead of picking from a fixed library (Kogge Stone, Sklansky etc.). Experimental results demonstrate that our approach (i) excels highly competitive industry standard Synopsys Design Compiler adder (128 bit) in performance (2%), area (25%) and power (13.3%) in 32nm technology node, and (ii) improves performance/area over even 64 bit custom designed adders targeting 22nm technology library and implemented in an industrial high-performance design.
Subhendu Roy, Mihir R. Choudhury, Ruchir Puri, David Z. Pan
ASP-DAC3
2015 Message from the program chairs
abstract
It is our great pleasure to welcome you to the 2015 ACM/IEEE International Symposium on Low Power Electronics and Design - ISLPED'15, in the “eternal city” of Rome, Italy. This year's symposium continues its two decade long tradition of being the premier forum for presentation of research results and industrial experience reports on leading-edge issues in low power design. ISLPED has always been unique in the sense that it brings together researchers and practitioners interested in various aspects of low power design at a single venue and provides them an opportunity to share their perspectives with each other.
Ruchir Puri, Vijay Raghunathan
ISLPED1
2015 Emerging Trends in Design and Applications of Memory-Based Computing and Content-Addressable Memories
abstract
Content-addressable memory (CAM) and associative memory (AM) are types of storage structures that allow searching by content as opposed to searching by address. Such memory structures are used in diverse applications ranging from branch prediction in a processor to complex pattern recognition. In this paper, we review the emerging challenges and opportunities in implementing different varieties of CAM/AM structures. Beyond-CMOS silicon and nonsilicon memory technologies hold significant promise in implementing dense, fast, and energy-efficient CAM/AM structures. We describe circuit/architecture level implementations of CAM/AM using these technologies, as well as novel applications in different domains, including informatics, text analytics, data mining, and reconfigurable computing platforms.
Robert Karam, Ruchir Puri, Swaroop Ghosh, Swarup Bhunia
Proc. IEEE2
2015 PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort
abstract
In-place radix sort is a popular distribution-based sorting algorithm for short numeric or string keys due to its linear run-time and constant memory complexity. However, efficient parallelization of in-place radix sort is very challenging for two reasons. First, the initial phase of permuting elements into buckets suffers read-write dependency inherent in its in-place nature. Secondly, load balancing of the recursive application of the algorithm to the resulting buckets is difficult when the buckets are of very different sizes, which happens for skewed distributions of the input data. In this paper, we present a novel parallel in-place radix sort algorithm, PARADIS, which addresses both problems: a) "speculative permutation" solves the first problem by assigning multiple non-continuous array stripes to each processor. The resulting shared-nothing scheme achieves full parallelization. Since our speculative permutation is not complete, it is followed by a "repair" phase, which can again be done in parallel without any data sharing among the processors. b) "distribution-adaptive load balancing" solves the second problem. We dynamically allocate processors in the context of radix sort, so as to minimize the overall completion time. Our experimental results show that PARADIS offers excellent performance/scalability on a wide range of input data sets.
Minsik Cho, Daniel Brand, Rajesh Bordawekar, Ulrich Finkler, Vincent KulandaiSamy, Ruchir Puri
Proc. VLDB Endow.6
2014 Energy-efficient hardware acceleration through computing in the memory
abstract
Energy-efficiency has emerged as a major barrier to performance scalability for modern processors. We note that significant part of processor's energy requirement is contributed by processor-memory communication. To address the energy issue in processors, we propose a novel hardware accelerator framework that transforms high-density memory array into a configurable computing resource to accelerate variety of tasks - both compute- and data-intensive. It exploits the block-based architecture of nanoscale memory to create a spatially connected array of lightweight processors, each of which uses a memory block as its local memory. The proposed framework provides some unique advantages for hardware acceleration compared to conventional accelerators: 1) memory array provides large set of parallel resources with high bandwidth, which can be configured to perform computing in spatio/temporal manner leading to dramatic reduction in processor-memory traffic; 2) it brings the computing engine close to the data, thus drastically minimizing the von Neumann bottleneck; 3) finally, it exploits the advances in memory technologies and integration approaches e.g. 3D integration to achieve better technology scalability compared to alternative reconfigurable accelerator platforms. Simulation results for several data-intensive applications show that the proposed computing approach provides significant improvement in energy-efficiency compared to software while achieving significantly lower hardware overhead.
Somnath Paul, Robert Karam, Swarup Bhunia, Ruchir Puri
DATE4
2014 Application driven high level design in the era of heterogeneous computing
abstract
Summary form only given. Escalating costs of semiconductor technology and its lagging performance relative to historic trends is motivating acceleration and specialization as more impactful means to increase system value. Targeted specialization is being increasingly pursued as an important way to achieve dramatic improvements in workload acceleration. This requires a broad understanding of workloads, system structures, and algorithms to determine what to accelerate / specialize, and how, i.e., via SW?; via HW?; or via SW+HW? which presents many choices, necessitating co-optimization of SW and HW. In this talk, we will focus on an application driven approach to high level design for software and system co-optimization, based on inventing new software algorithms, that have strong affinity to hardware acceleration.
Ruchir Puri
ICCAD1
2014 Bridging high performance and low power in processor design
abstract
The design complexity of modern high performance processors calls for innovative design techniques and methodologies for achieving time-to-market goals. New design techniques are also needed to curtail power increases that inherently arise from ever increasing performance targets. This paper describes new processor design and optimization approaches that bridge the gap between high performance and low power. These techniques are flexible as they rely on automated synthesis-centric optimizations to enable power reduction without sacrificing performance. These methodology innovations contributed to the industry leading performance of the POWER8 processor.
Ruchir Puri, Mihir R. Choudhury, Haifeng Qian, Matthew M. Ziegler
ISLPED1
2014 Towards Optimal Performance-Area Trade-Off in Adders by Synthesis of Parallel Prefix Structures
abstract
This paper proposes an efficient algorithm to synthesize prefix graph structures that yield adders with the best performance-area trade-off. For designing a parallel prefix adder of a given bit-width, our approach generates prefix graph structures to optimize an objective function such as size of prefix graph subject to constraints like bit-wise output logic level. Given bit-width n and level (L) restriction, our algorithm excels the existing algorithms in minimizing the size of the prefix graph. We also prove its size-optimality when n is a power of two and L= log2n. Besides prefix graph size optimization and having the best performance-area trade-off, our approach, unlike existing techniques, can 1) handle more complex constraints such as maximum node fanout or wire-length that impact the performance/area of a design and 2) generate several feasible solutions that minimize the objective function. Generating several size-optimal solutions provides the option to choose adder designs that mitigate constraints such as wire congestion or power consumption that are difficult to model as constraints during logic synthesis. Experimental results demonstrate that our approach improves performance by 3% and area by 9% over even a 64-bit full custom designed adder implemented in an industrial high-performance design.
Subhendu Roy, Mihir R. Choudhury, Ruchir Puri, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 Towards optimal performance-area trade-off in adders by synthesis of parallel prefix structures
abstract
This paper proposes an efficient algorithm to synthesize prefix graph structures that yield adders with the best performance-area trade-off. For designing a parallel prefix adder of a given bit-width, our approach generates prefix graph structures to optimize an objective function such as size of prefix graph subject to constraints like bit-wise output logic level. Besides having the best performance-area trade-off our approach, unlike existing techniques, can (i) handle more complex constraints such as maximum node fanout or wire-length that impact the performance/area of a design and (ii) generate several feasible solutions that minimize the objective function. Generating several optimal solutions provides the option to choose adder designs that mitigate constraints such as wire congestion or power consumption that are difficult to model as constraints during logic synthesis. Experimental results demonstrate that our approach improves performance by 3% and area by 9% over even a 64-bit full custom designed adder implemented in an industrial high-performance design.
Subhendu Roy, Mihir R. Choudhury, Ruchir Puri, David Z. Pan
DAC3
2013 Intuitive ECO synthesis for high performance circuits
abstract
In the IC industry, chip design cycles are becoming more compressed, while designs themselves are growing in complexity. These trends necessitate efficient methods to handle late-stage engineering change orders (ECOs) to the functional specification, often in response to errors discovered after much of the implementation is finished. Past ECO synthesis algorithms have typically treated ECOs as functional errors and applied error diagnosis techniques to solve them. However, error diagnosis methods are primarily geared towards finding a single change, and moreover, tend to be computationally complex. In this paper, we propose a unique methodology that can systematically incorporate human intuition into the ECO process. Our methodology involves finding a set of directly substitutable points known as functional correspondences between the original implementation and the new specification by using name-preserving synthesis and user hints, to diminish the size of the ECO problem. On average, our approach can reduce the size of logic changes by 94% from those reported in current literature. We then incorporate our logic ECO changes into an incremental physical synthesis flow to demonstrate its usability in an industrial setting. Our ECO synthesis methodology is evaluated on high-performance industrial designs. Results indicate that post-ECO worst negative slack (WNS) improved 14% and total negative slack (TNS) improved 46% over pre-ECO.
Haoxing Ren, Ruchir Puri, Lakshmi N. Reddy, Smita Krishnaswamy, Cindy Washburn, Joel Earl, Joachim Keinert
DATE2
2013 LatchPlanner: latch placement algorithm for datapath-oriented high-performance VLSI designs
abstract
In this paper, we present a novel algorithm for latch placement, LatchPlanner which enables a placement engine to deliver high quality placement for datapath-oriented design. Datapath-oriented VLSI designs are in general hand-crafted by human at high cost, as understanding and capturing datapath structure is critical for the performance. The conventional placement algorithms by itself cannot exploit the underlying datapath due to lack of logic structure recognition and inaccurate/approximated wirelength estimation. LatchPlanner addresses such drawbacks by placing and fixing latches in the datapath context, a key element in datapath structure. By taking placed/fixed latches as constraints, a placer can find a more datapath-friendly placement effectively, which results in higher-quality hardware. LatchPlanner begins latch clustering/sizing/ordering to prepare the following steps, a) global latch placement based on linear programming to place latch clusters, and b) local latch placement based on network flow optimization to place latches within each cluster. Experimental results on eighteen industrial benchmarks show that LatchPlanner improves total wirelength by 32%, total negative slack by 25%, and area by 3% without CPU overhead over a commercial placement engine, and delivers near semi-custom-quality solutions.
Minsik Cho, Hua Xiang 0001, Haoxing Ren, Matthew M. Ziegler, Ruchir Puri
ICCAD5
2013 Depth controlled symmetric function fanin tree restructure
abstract
A symmetric-function fanin tree (SFFT) is a fanout-free cone of logic that computes a symmetric function such as AND, OR and XOR. These trees are usually created during logic synthesis, when there is no knowledge of the tree gate locations. Because of this, large SFFTs present a challenge to placement algorithms. The consequence is that the tree placements are generally far from optimal, leading to wiring congestion, excess buffering, and timing problems. [10] proposed a fanin-tree restructure algorithm to reduce the SFFT wirelength. However, [10] was based on Steiner trees and might cause serious timing problems due to the high Steiner tree depth. In this paper, we extend the SFFT tree identification algorithm to allow both positive and negative tree inputs. Contrary to the Steiner-tree based approach, we propose a new tree restructure flow to build SFFTs from bottom to top level by level at the physical design stage. The tree restructure algorithm is in a transaction mode so that only improved trees are accepted, and the new tree won't cause any placement legal issue. A new partitioning algorithm is proposed to serve for gate creation. In addition, various optimization techniques are developed to reduce tree wirelength On tested designs, the total tree wirelength is reduced by 31% with similar tree gates and tree depths.
Hua Xiang 0001, Lakshmi N. Reddy, Louise Trevillyan, Ruchir Puri
ICCAD4
2013 Opportunities and challenges for high performance microprocessor designs and design automation
abstract
With end of an era of classical technology scaling and exponential frequency increases, high end microprocessor designs and design automation methodologies are at an inflection point. With power and current demands reaching breaking points, and significant challenges in application software stack, we are also reaching diminishing returns from simply adding more cores. In design methodologies for high end microprocessors, although chip physical design efficiency has seen tremendous improvements, strong indications are emerging for maturing of those gains as well. In order to continue the cost-performance scaling in systems in light of these maturing trends, we must innovate up the design stack, moving focus from technology and physical design implementation to new IP and methodologies at logic, architecture, and at the boundary of hardware and software, solving key bottlenecks through application acceleration. This new era of innovation, which moves the focus up the design stack presents new challenges and opportunities to the design and design automation communities. This talk will motivate these trends and focus on challenges for high performance microprocessor design and design automation in the years to come.
Ruchir Puri
ISPD1
2013 Network flow based datapath bit slicing
abstract
In deep sub-micro designs, more functions are integrated into one chip, and datapath has become a critical part of the design. Typical datapath consists an array of bit slices. The inherent high degree regularity of datapaths is especially attractive to the placement and routing to achieve regular layout with high density and high performance. However, the current design methodology may generate inferior datapath designs because the datapath regularity cannot be well understood by the traditional design tools. In previous works, several techniques are proposed to preserve/re-identify datapath structures. However, they either restrict the datapath optimization or have little tolerance on bit slice difference.
Hua Xiang 0001, Minsik Cho, Haoxing Ren, Matthew M. Ziegler, Ruchir Puri
ISPD5
2011 Design, CAD and technology challenges for future processors: 3D perspectives
abstract
Technology scaling has provided the semiconductor industry a recipe to successfully meet the application demands for performance for over three decades. This computational capacity was further fueled by the success of circuit and architecture-level innovation, which provided performance improvement in each processor generation. However, future processor designs face a number of key challenges in sustaining the growth trends. The dawn of the 22nm node, and beyond, marks an era of new trends and challenges; where the cost and complexity associated with each technology node is increasing at much faster rate than the device performance gains. Novel tools and design methodologies are needed to not only compensate for these challenges but also to leverage emerging technologies to achieve the desired performance in future processor architectures. Technology alternatives such as 3D integration have attracted significant interest as an additional way of sustaining the density scaling and performance growth.
Jeff Burns, Gary Carpenter, Eren Kursun, Ruchir Puri, James D. Warnock, Michael Scheuermann
DAC4
2010 History-based VLSI legalization using network flow
abstract
In VLSI placement, legalization is an essential step where the overlaps between gates/macros must be removed. In this paper, we introduce a history-based legalization algorithm with min-cost network flow optimization. We find a legal solution with the minimum deviation from a given placement to fully honor/preserve the initial placement, by solving a gate-centric network flow formulation in an iterative manner. In order to realize a flow into gate movements, we develop efficient techniques which solve an approximated Subset-sum problem. Over the iterations, we factor into our formulation the history which captures a set of likely-to-fail gate movements. Such a history-based scheme enables our algorithm to intelligently legalize highly complex designs. Experimental results on over 740 real cases show that our approach is significantly superior to the existing algorithms in terms of failure rate (no failure) as well as quality of results (55% less max-deviation).
Minsik Cho, Haoxing Ren, Hua Xiang 0001, Ruchir Puri
DAC4
2010 EDA challenges and options: investing for the future
abstract
As the overall economy and semiconductor industry emerges from one of the worst recessions in years, it is time to take stock of EDA challenges and its future. This panel will focus on which challenges will surge and dominate EDA over the course of next several years and which challenges we can sell short.
Ruchir Puri, William H. Joyner, Raj Jammy, Ahmed Jerraya, Jan M. Rabaey, Walden C. Rhines, Leon Stok
DAC1
2010 Novel binary linear programming for high performance clock mesh synthesis
abstract
Clock mesh is popular in high performance VLSI design because it is more robust against variations than clock tree at a cost of higher power consumption. In this paper, we propose novel techniques based on binary linear programming for clock mesh synthesis for the first time in the literature. The proposed approach can explore both regular and irregular mesh configurations, adapting to non-uniform load capacitance distribution. Our synthesis consists of two steps: mesh construction to minimize total capacitance and skew, and balanced sink assignment to improve slew/skew characteristics. We first show that mesh construction can be analytically formulated as binary polynomial programming (a class of nonlinear discrete optimization), then apply a compact linearization technique to transform into binary linear programming, significantly reducing computational overhead. Second, our balanced sink assignment enables a sink to tap the least loaded mesh segment (not the nearest one) with another binary linear programming which reduces both slew and skew. Experiments show that our techniques improve the worst skew and total capacitance by 14% and 15% over the state-of-the-art clock mesh algorithm on ISPD09 benchmarks.
Minsik Cho, David Z. Pan, Ruchir Puri
ICCAD3
2010 Logical and physical restructuring of fan-in trees
abstract
A symmetric-function fan-in tree (SFFT) is a fanout-free cone of logic that computes a symmetric function, so that all of the leaf nets in its support set are commutative. Such trees are frequently found in designs, especially when the design originated as two-level logic.These trees are usually created during logic synthesis, when there is no knowledge of the locations of the tree root or of the source gates of the leaf nets. Because of this, large SFFTs present a challenge to placement algorithms. The result is that the tree placements are generally far from optimal, leading to wiring congestion, excess buffering, and timing problems. Restructuring such trees can produce a more placeable and wire-efficient design.In this paper, we propose algorithms to identify and to restructure SFFTs during physical design. The key feature of an SFFT is that it can be implemented with various structures of a uniform set of gates with commutative inputs, i.e. AND, OR, or XOR. Drawing on the flexibility of SFFT logic structures, the proposed tree restructuring algorithm uses existing placement information to rebuild the SFFTs with reduced tree wire lengths. The experimental results demonstrate the efficiency and effectiveness of the algorithms.
Hua Xiang 0001, Haoxing Ren, Louise Trevillyan, Lakshmi N. Reddy, Ruchir Puri, Minsik Cho
ISPD5
2009 CAD challenges for 3D ICs
abstract
A fundamental shift in the technology has occurred at 90 nm CMOS and beyond where the interconnect resistance has been increasing so much that the distance a clock cycle can reach has been dwindling as a fraction of the dimension of the chip. to cause a repeater explosion problem. This problem translates into an explosion of repeaters which not only added significant overhead in area but also power, as repeaters are major contributors to leakage. By reaching out to the vertical dimension, 3D technology has the potential of easing repeater explosion (Figure 1), reducing latency between units, increasing memory bandwidth and integrating heterogeneous technologies. However, in order to exploit the full potential of 3D technology, new challenges in the area of system level design and analysis, physical design, thermal analysis need to be addressed.
David S. Kung 0001, Ruchir Puri
ASP-DAC2
2009 Moore's Law: another casualty of the financial meltdown?
abstract
Given the exponential increase of fabrication costs, the global recession and credit crunch, one may ask if Moore's Law is financially viable beyond 22nm node. Can we justify the return-of-investment (ROI) for continuous scaling beyond 22nm? Shall we consider other alternatives for integration, such as silicon-in-a-package (SiP) or 3D integrations?
Jason Cong, N. S. Nagaraj, Ruchir Puri, William H. Joyner, Jeff Burns, Moshe Gavrielov, Riko Radojcic, Peter Rickert, Hans Stork
DAC3
2009 From milliwatts to megawatts: system level power challenge
abstract
This panel discusses power optimization at the system level. What are the needs and opportunities? What are examples of successful practices? Can system-level power optimizations ever be automated? Can power optimizations be the enabling event for a methodology change in system-level design?
Ruchir Puri, Eshel Haritan, Stan Krolikoski, Jason Cong, Tim Kogel, Bradley D. McCredie, John Shen, Andrés Takach
DAC1
2009 DeltaSyn: An efficient logic difference optimizer for ECO synthesis
abstract
During the IC design process, functional specifications are often modified late in the design cycle, after placement and routing are completed. However, designers are left either to manually process such modifications by hand or to restart the design process from scratch---a very costly option. In order to address this issue, we present DeltaSyn, a method for generating a highly optimized logic difference between a modified high-level specification and an implemented design. DeltaSyn has the ability to locate boundaries in implemented logic within which changes can be confined. DeltaSyn demarcates the boundary in two phases. The first phase employs fast functional and structural analysis techniques to identify equivalent signals forming the input-side boundary of the changes. The second phase locates the outputside boundary of the changes through a novel dynamic algorithm that detects matching logic downstream from the changes required by the ECO. Experiments on industrial designs show that together these techniques successfully implement ECOs while preserving an average of 97% of the existing logic. Unlike previous approaches, the use of bitparallel logic simulation and fast SAT solvers enables high performance and scalability. DeltaSyn can process and verify a typical ECO for a design of around 10K gates in about 200 seconds or less.
Smita Krishnaswamy, Haoxing Ren, Nilesh Modi, Ruchir Puri
ICCAD4
2009 Will 22nm be our catch 22!: design and cad challenges
abstract
With the impact of cost crunch in every industry including semi-conductor industry across the world, complexities of 22nm CMOS may appear to be way out in the distance future. However, more than the technology complexity, the new era of financial restraint will impose unprecedented productivity requirements on future design flows. This presents a significant opportunity as well as a complex challenge for the EDA industry. In addition, with semiconductor industry's march towards 32nm CMOS technology and introduction of new technologies such as 3D ICs in sight for 22nm nodes, it is crucial for the IC design and CAD community to understand the challenges posed by these potential technology changes. This talk will cover these challenges and outline the opportunities including those in physical design domain. The advanced device and interconnect structures and materials including 3D technology that could be introduced at 22nm node have significant impact on the direction of the CAD industry. We will also discuss the design methodology and CAD implications of these technology changes.
Ruchir Puri
ISPD1
2008 Custom is from Venus and synthesis from Mars
abstract
Due to ever increasing cost of doing design, design productivity and more specifically, cost of design has become a major bottleneck in large scale design projects. Due to the cost crunch, automated synthesis techniques have been gaining ground on manual and cost intensive (although potentially yielding higher quality of results) custom techniques. The push towards system level performance with multiple lower performance cores as opposed to single high performance core is also making the pursuit of frequency at any cost meaningless. Synthesis vs Custom has traditionally been a lively debate in microprocessor companies but now it is becoming more widespread with the desire of fabless companies to more fully utilize the technology node in order to compete with IDMs by utilizing more custom techniques. Custom hotshots argue that we cannot afford to leave performance or power on table when you are spending 3billion $s in fabs. Industry needs to get maximum benefits possible. Synthesis fanatics will argue that designers should get over the last few ps and last few nW, as time to market, cost, and system performance are the new drivers and synthesis can not only meet the challenge but beat it many times with lower power solutions. This panel will present this internal debate in a public forum.
Ruchir Puri, William H. Joyner, Shekhar Borkar, Ty Garibay, Jonathan Lotz, Robert K. Montoye
DAC1
2008 Keeping hot chips cool: are IC thermal problems hot air?
abstract
Thermal issues are becoming more important but is the hype getting the better of the facts? Does this deserve more attention than for some niche designs and technologies such as 3D ICs.? Does the broader design community need to worry about it at 32nm and beyond or it will only impact a small segment of designs? In short, does the severity of power issues coupled with packaging complexity translate into a thermal crisis in future? This is an educational panel with a little bit of controversy that will address the thermal issue in IC design. When will this issue be emerging as a crucial concern if at all? What are the solutions to resolve this potential crisis?
Ruchir Puri, Devadas Varma, Darvin Edwards, Alan J. Weger, Paul D. Franzon, Stephen V. Kosonocky
DAC1
2008 Track Routing and Optimization for Yield
abstract
In this paper, we propose track routing and optimization for yield (TROY), the first track router for the optimization of yield loss due to random defects. As the probability of failure (POF), which is an integral of the critical area and the defect size distribution, strongly depends on wire ordering, sizing, and spacing, track routing can play a key role in effective wire planning for yield optimization. However, a straightforward formulation of yield-driven track routing can be shown to be integer nonlinear programming, which is a nondeterministic polynomial-time complete problem. TROY overcomes the computational complexity by combining two effective techniques, i.e., the minimum Hamiltonian path (MHP) from graph theory and the second-order cone programming (SOCP) from mathematical optimization. First, TROY performs wire ordering to minimize the critical area for short defects by finding an MHP. Then, TROY carries out optimal wire sizing/spacing through SOCP optimization based on the given wire order. Since the SOCP can be optimally solved in near linear time, TROY efficiently achieves globally optimal wire sizing/spacing for the minimal POF.
Minsik Cho, Hua Xiang 0001, Ruchir Puri, David Z. Pan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2008 Is Your Layout-Density Verification Exact? - A Fast Exact Deep Submicrometer Density Calculation Algorithm
abstract
As the device shapes keep shrinking, the designs are more sensitive to manufacturing processes. In order to improve performance predictability and yield, mask-layout uniformity/evenness is highly desired, and it is usually measured by the feature densities within defined feasible ranges determined by the manufacturing-process design rules. To address the density-control problem, one fundamental problem is how to calculate density accurately and efficiently. In this paper, we propose a fast exact algorithm to identify the maximum/minimum density for a given layout. Compared with the existing exact algorithms, our algorithm reduces the running time from days/long hours to a few minutes/seconds. Moreover, it is even faster than the existing approximate algorithms in the literature.
Hua Xiang 0001, Kai-Yuan Chao, Ruchir Puri, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2008 Fast Dummy-Fill Density Analysis With Coupling Constraints
abstract
In modern very large scale integration manufacturing processes, dummy fills are widely used to adjust local metal density in order to improve layout uniformity and yield optimization. However, the introduction of a large amount of dummy features also affects wire electrical properties. In this paper, we propose the first coupling-constrained dummy-fill analysis algorithm which identifies feasible locations for dummy fills such that the fill-induced coupling capacitance can be bounded within the given coupling threshold of each wire segment. A speedup approach is presented based on the cache concept. The algorithm also makes efforts to maximize ground dummy fills, which are more robust and predictable. The output of the algorithm can be treated as the upper bound for dummy-fill insertion, and it can be easily adopted in density models to guide dummy-fill insertion without disturbing the existing design.
Hua Xiang 0001, Liang Deng, Ruchir Puri, Kai-Yuan Chao, Martin D. F. Wong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2007 Interconnects in the Third Dimension: Design Challenges for 3D ICs
abstract
Despite generation upon generation of scaling, computer chips have until now remained essentially 2-dimensional. Improvements in on-chip wire delay and in the maximum number of I/O per chip have not been able to keep up with transistor performance growth; it has become steadily harder to hide the discrepancy. 3D chip technologies come in a number of flavors, but are expected to enable the extension of CMOS performance. Designing in three dimensions, however, forces the industry to look at formerly-two- dimensional integration issues quite differently, and requires the re-fitting of multiple existing EDA capabilities.
Kerry Bernstein, Paul S. Andry, Jerome Cann, Philip G. Emma, David Greenberg, Wilfried Haensch, Mike Ignatowski, Steven J. Koester, John Magerlein, Ruchir Puri, Albert M. Young
DAC10
2007 TROY: Track Router with Yield-driven Wire Planning
abstract
In this paper, we propose TROY, the first track router with yield-driven wire planning to optimize yield loss due to random defects. As the probability of failure (POF) computed from critical area analysis and defect size distribution strongly depends on wire ordering, sizing, and spacing, track routing plays a key role in effective wire planning for yield optimization. TROY formulates wire ordering into a preference-aware minimum Hamiltonian path problem. For simultaneous wire sizing and spacing optimization, TROY solves it optimally by formulating the problems into a second order conic programming (SOCP). Experimental results show that TROY can reduce the random-defect yield loss by 18% on average without any overhead in wirelength, compared with the widely used greedy approach.
Minsik Cho, Hua Xiang 0001, Ruchir Puri, David Z. Pan
DAC3
2007 Making Manufacturing Work For You
Srikanth Venkataraman, Ruchir Puri, Steve Griffith, Ankush Oberai, Robert Madge, Greg Yeric, Walter Ng, Yervant Zorian
DAC2
2007 Is your layout density verification exact?: a fast exact algorithm for density calculation
abstract
As the device shapes keep shrinking, the designs are more sensitive to manufacturing processes. In order to improve performance predictability and yield, mask layout uniformity/evenness is highly desired, and it is usually measured by the feature density with defined feasible range in manufacture process design rules. To address the density control problem, one fundamental problem is how to calculate density accurately and efficiently. In this paper, we propose a fast exact algorithm to identify the maximum density for a given layout. Compared with the existing exact algorithms, our algorithm reduces the running time from days/hours to a few minutes/seconds. And it is even faster than the existing approximate algorithms in literature.
Hua Xiang 0001, Kai-Yuan Chao, Ruchir Puri, Martin D. F. Wong
ISPD3
2007 Dummy fill density analysis with coupling constraints
abstract
In modern VLSI manufacturing processes, dummy fills are widely used to adjust local metal density in order to improve layout uniformity and yield optimization. However, the introduction of a large amount of dummy features also affects wire electrical properties. In this paper, we propose the first Coupling constrained Dummy Fill (CDF) analysis algorithm which identifies feasible locations for dummy fills such that the fill induced coupling capacitance can be bounded within the given coupling threshold of each wire segment. The algorithm also makes efforts to maximize ground dummy fills, which are more robust and predictable. The output of the algorithm can be treated as the upper bound for dummy fill insertion, and it can be easily adopted in density models to guide dummy fill insertion without disturbing the existing design.
Hua Xiang 0001, Liang Deng, Ruchir Puri, Kai-Yuan Chao, Martin D. F. Wong
ISPD3
2006 Gain-based technology mapping for minimum runtime leakage under input vector uncertainty
abstract
The gain-based technology mapping paradigm has been successfully employed for finding minimum delay and minimum area mappings. However, existing gain-based technology mappers fail to find circuits with minimal leakage power. In this paper, we introduce algorithms and modeling strategies that enable efficient gain-based technology mapping for minimum leakage power. The proposed algorithm is probability-aware and can rigorously take into account input state probability distribution to generate a circuit mapping with minimum leakage at a given percentile. Minimizing leakage at high percentiles is essential for minimizing peak leakage, which strongly influences the cooling limits and packaging costs.The algorithms have been tested on the ISCAS85 benchmark suite. Results indicate that the mappings produced by the new algorithm consume, on average 14% lesser leakage power at the 99% percentile with 1% delay penalty when compared with the approaches used in previous gain-based mappers [2]. Also, compared to a dominant-state mapper, our approach produces mappings with 15% lesser mean value of leakage. The new algorithm also reduces leakage at high quantiles by 12.8% on average, compared to a dominant state leakage minimizing mapper and the maximum savings can be as high as 21.49% across the benchmarks. Compared to the bin based mapper [10], the runtime of the algorithm is 15X faster.
Ashish Kumar Singh, Murari Mani, Ruchir Puri, Michael Orshansky
DAC3
2006 Wire density driven global routing for CMP variation and timing
abstract
In this paper, we propose the first wire density driven global routing that considers CMP variation and timing. To enable CMP awareness during global routing, we propose a compact predictive CMP model with dummy fill, and validate it with extensive industry data. While wire density has some correlation and similarity to the conventional congestion metric, they are indeed different in the global routing context. Therefore, wire density rather than congestion should be a unified metric to improve both CMP variation and timing. The proposed wire density driven global routing is implemented in a congestion-driven global router [5] for CMP and timing optimization. The new global router utilizes several novel techniques to reduce the wire density of CMP and timing hotspots. Our experimental results are very encouraging. The proposed algorithm improves CMP variation and timing by over 7% with negligible overhead in wirelength and even slightly better routability, compared to the pure congestion-driven global router [5].
Minsik Cho, David Z. Pan, Hua Xiang 0001, Ruchir Puri
ICCAD4
2006 Design and CAD challenges in 45nm CMOS and beyond
abstract
With semiconductor industry's aggressive march towards 45nm CMOS technology and introduction of new materials and device structures in sight for 32nm and 22nm nodes, it is crucial for the IC design and CAD community to understand the challenges posed by these potential technology changes. This tutorial will focus on these challenges starting from front end of line (devices) to the back end of line (interconnects) and finally the impact on CAD. We will discuss the impact of various device technology options/improvements, such as high-k, metal gate, low temperature operation, increased mobility and reduced variability, on the overall chip performance in the context of power-constrained technology optimization. This will show that power constraints limit, but do not eliminate, the performance improvements available from new technology. The integration issues related to low-k materials for interconnects in 45nm and beyond will be examined in the context of advanced IC design. Ultra low-k materials, evolution of etch and chemical mechanical polishing (CMP), and techniques to limit damage during processing and their impact on design performance will be discussed in detail. These advanced device and interconnect structures and materials including 3D technology have tremendous impact on the direction of the CAD industry. We will discuss the design methodology and CAD implications of these imminent technology changes.
David J. Frank, Ruchir Puri, Dorel Toma
ICCAD2
2005 Keeping hot chips cool
abstract
With 90nm CMOS in production and 65nm testing in progress, power has been pushed to the forefront of design metrics. This paper will outline practical techniques that are used to reduce both leakage as well as active power in a standard-cell library based high-performance design flow. We will discuss the design and cost issues for using different power saving techniques such as: power gating to reduce leakage, multiple and hybrid threshold libraries for leakage reduction and multiple supply voltage based design. In addition techniques to reduce clock tree power will be presented as power consumed in clocks accounts for a significant portion of total chip power. Practical aspects of implementing these techniques will also be discussed.
Ruchir Puri, Leon Stok, Subhrajit Bhattacharya
DAC1
2003 Pushing ASIC performance in a power envelope
abstract
Power dissipation is becoming the most challenging design constraint in nanometer technologies. Among various design implementation schemes, standard cell ASICs offer the best power efficiency for high-performance applications. The flexibility of ASICs allow for the use of multiple voltages and multiple thresholds to match the performance of critical regions to their timing constraints, and minimize the power everywhere else. We explore the trade-off between multiple supply voltages and multiple threshold voltages in the optimization of dynamic and static power. The use of multiple supply voltages presents some unique physical and electrical challenges. Level shifters need to be introduced between the various voltage regions. Several level shifter implementations will be shown. The physical layout needs to be designed to ensure the efficient delivery of the correct voltage to various voltage regions. More flexibility can be gained by using appropriate level shifters. We will discuss optimization techniques such as clock skew scheduling which can be effectively used to push performance in a power neutral way.
Ruchir Puri, Leon Stok, John M. Cohn, David S. Kung 0001, David Z. Pan, Dennis Sylvester, Ashish Srivastava, Sarvesh H. Kulkarni
DAC1
2003 Design and CAD Challenges in sub-90nm CMOS Technologies
Kerry Bernstein, Ching-Te Chuang, Rajiv V. Joshi, Ruchir Puri
ICCAD4
2002 Fast and accurate wire delay estimation for physical synthesis of large ASICs
abstract
Interconnect delays represent an increasingly dominant portion of overall circuit delays. During timing-driven physical synthesis process, timing analysis is repeatedly performed over several hundred thousand components. Thus, fast and accurate estimation of interconnect delays is crucial. Traditionally, lumped and elmore delay models have been widely used for computing interconnect delays in physical synthesis due to their computational efficiency. However, these delay models are known to be inaccurate since they ignore slew and resistive shielding effects. In this paper, we propose a new iterative refinement based delay estimation approach that considers resistive shielding along with driver slew. Experimental results show that the proposed approach gives not only highly accurate results for far end RC-line delays but also compares very favorably to more difficult to match source end delays and source slews. In addition, use of the proposed delay model in physical synthesis yields significant performance improvement on several large industrial ASICs.
Ruchir Puri, David S. Kung 0001, Anthony D. Drumm
ACM Great Lakes Symposium on VLSI1
2000 Combinatorial cell design for CMOS libraries
Frederik Beeftink, Prabhakar Kudva, David S. Kung 0001, Ruchir Puri, Leon Stok
Integr.4
1999 SOI Digital CMOS VLSI - a Design Perspective
abstract
Abstract|This paper reviews the recent advances of SOI for digital CMOS VLSI applications with particular emphasis on the design issues and advantages resulting from the unique SOI device structure. The technology/device requirements and design issues/challenges for high-performance, generalpurpose microprocessor applications are di erentiated with respect to low-power portable applications. Particular emphases are placed on the impact of oating-body in partiallydepleted devices on the circuit operation, stability, and functionality. Unique SOI design aspects such as parasitic bipolar e ect and hysteretic VT variation are addressed. Circuit techniques to improve the noise immunity and global design issues are discussed. 1.
Ching-Te Chuang, Ruchir Puri
DAC2
1999 Optimal P/N width ratio selection for standard cell libraries
abstract
The effectiveness of logic synthesis to satisfy increasingly tight timing constraints in deep-submicron high-performance circuits heavily depends on the range and variety of logic gates available in the standard cell library. Primarily, research in the design of high-performance standard cell libraries has been focused on drive strength selection of various logic gates. Since CMOS logic circuit delays not only depend on the drive strength of each gate but also on its PM width ratio, it is crucial to provide good PM width ratios for each cell. The main contribution of this paper is the development of a theoretical framework through which library designers can determine "optimal" PM width ratio for each logic gate in their high-performance standard cell library. This theoretical framework utilizes new gate delay models that explicitly represent the dependence of delay on P/N width ratio and load. These delay models yield highly accurate delay for CMOS gates in a 0.12 /spl mu/m L/sub eff/ deep-submicron technology.
David S. Kung 0001, Ruchir Puri
ICCAD2
1999 Hysteresis effect in floating-body partially-depleted SOI CMOS domino circuits
abstract
This paper investigates the basic mechanisms of hysteretic delay and noise margin variations for floating-body Partially-Depleted SOI CMOS domino circuits in detail.Three cases, basedon whether the input signals are "domino input signals" from other domino circuits; "static input signals" from static circuits or latches; or a combination of "domino and static input signals" are examined and differentiated.It is shown that hysteretic delay variation is larger and noise margin worse for the later case with "mixed domino and static input signals."Although the delay and noise margin disparities between the three types of input signals are significant at beginning of the clock cycles, they converge as the circuit approaches steadystate.
Ruchir Puri, Ching-Te Chuang
ISLPED1
1998 Design issues in mixed static-domino circuit implementations
abstract
Due to its performance advantage, domino logic has been extensively used to implement critical paths in advanced microprocessor designs. Since domino circuits are employed in a mainly static environment, static logic imposes strict timing constraints on domino logic and its use. In this paper, we present the design issues involved in implementing mixed domino and static logic circuits in high performance microprocessors designs. These issues are addressed with design guidelines in a custom design methodology. Simulations of some designs with mixed static-domino implementations show that static-domino interface can severely constrain their performance. To alleviate this bottleneck, we reanalyze the interaction between static and domino signals in detail and a new relaxed static-domino interface constraint is derived. This new relaxed constraint can result in a significant performance improvement of mixed static-domino implementations.
Ruchir Puri
ICCD1
1996 Logic optimization by output phase assignment in dynamic logic synthesis
abstract
Domino logic is one of the most popular dynamic circuit configurations for implementing high-performance logic designs. Since domino logic is inherently noninverting, it presents a fundamental constraint of implementing logic functions without any intermediate inversions. Removal of intermediate inverters requires logic duplication for generating both the negative and positive signal phases, which results in significant area overhead. This area overhead can be substantially reduced by selecting an optimal output phase assignment, which results in a minimum logic duplication penalty for obtaining inverter-free logic. In this paper, we present this previously unaddressed problem of output phase assignment for minimum area duplication in dynamic logic synthesis. We give both optimal and heuristic algorithms for minimizing logic duplication.
Ruchir Puri, Andrew Bjorksten, Thomas E. Rosser
ICCAD1
1995 Asynchronous circuit synthesis with Boolean satisfiability
abstract
Asynchronous circuits are widely used in many real time applications such as digital communication and computer systems. The design of complex asynchronous circuits is a difficult and error-prone task. An adequate synthesis method will significantly simplify the design and reduce errors. In this paper, we present a general and efficient partitioning approach to the synthesis of asynchronous circuits from general Signal Transition Graph (STG) specifications. The method partitions a large signal transition graph into smaller and manageable subgraphs which significantly reduces the complexity of asynchronous circuit synthesis. Experimental results of our partitioning approach with large number of practical industrial asynchronous circuit benchmarks are presented. They show that, compared to the existing asynchronous circuit synthesis techniques, this partitioning approach achieves many orders of magnitude of performance improvements in terms of computing time, in addition to the reduced circuit implementation area. This lends itself well to practical asynchronous circuit synthesis from general STG specifications.>
Ruchir Puri
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1994 A Modular Partitioning Approach for Asynchronous Circuit Synthesis
abstract
Asynchronous circuits are widely used in many real time applications such as digital communication and computer systems. The design of complex asynchronous circuits is a di cult and error-prone task. An adequate synthesis method will signi cantly simplify the design and reduce errors. In this paper, we present a general and e cient partitioning approach to the synthesis of asynchronous circuits from general Signal Transition Graph (STG) speci cations. The method partitions a large signal transition graph into smaller and manageable subgraphs which signi cantly reduces the complexity of asynchronous circuit synthesis. Experimental results of our partitioning approach with large number of practical industrial asynchronous circuit benchmarks are presented. They show that, compared to the existing asynchronous circuit synthesis techniques, this partitioning approach achieves many orders of magnitude of performance improvements in terms of computing time, in addition to the reduced circuit implementation area. This lends itself well to practical asynchronous circuit synthesis from general STG speci cations.
Ruchir Puri
DAC1
1994 Area Efficient Synthesis of Asynchronous Interface Circuits
abstract
Asynchronous circuits are widely used in many real time applications such as digital communication and computer systems. The design of complex asynchronous interface circuits is a difficult and error-prone task. We present an area and time efficient synthesis algorithm for general signal transition graph (STG) specifications. It utilizes a divide-and-conquer approach to significantly reduce the number of design constraints. We present a BDD constraint satisfaction algorithm that exploits the don't cares for area efficient synthesis. Experimental results with a large number of practical signal transition graph benchmarks are presented. These results show that compared to the existing techniques, the divide-and-conquer BDD technique is capable of achieving an average of 20% reduction in implementation area.>
Ruchir Puri
ICCD1
1993 Signal Transition Graph Constraints for Speed-independent Ciruit Synthesis
Ruchir Puri
ISCAS1
1993 An efficient algorithm to search for minimal closed covers in sequential machines
abstract
The removal of redundant states in a finite state machine (FSM) is essential to reducing the complexity of a sequential circuit. An efficient algorithm for state minimization in incompletely specified state machines is presented. This algorithm employs a tight lower bound and a fail-first heuristic and generates a relatively small search space from the prime compatibles. It utilizes efficient pruning rules to further reduce the search space and finds a minimal closed cover. The technique guarantees the elimination of all the redundant states in a very short execution time. Experimental results with a large number of FSMs including the MCNC FSM benchmarks are presented. The results are compared with other work in this area.>
Ruchir Puri
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1993 Microword length minimization in microprogrammed controller synthesis
abstract
The problem of microword length minimization is crucial to the synthesis of microprogrammed controllers in digital systems. Unfortunately, this problem is NP-hard. Although various enumerative and heuristic methods have been developed, usually they cannot provide fast and efficient solutions to a large size problem. Here, the problem is formulated into a graph partitioning problem. An efficient graph partitioning algorithm was developed that works by recursively extracting large size cliques from the graph. Furthermore, a local search approach is used to reduce the microword length. This yields an efficient algorithm that outperforms any technique to solve the problem. The algorithm has been tested with practical microcodes. Experimental results are compared with other methods.>
Ruchir Puri
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1992 An Efficient algorithm for Microword Length Minimization
Ruchir Puri
DAC1
1991 Searching for a minimal finite state automaton (FSA)
abstract
An efficient algorithm to search for a minimal finite state automaton (FSA) is presented. This algorithm eliminates all the redundant states in a given FSA and is guaranteed to produce an optimal solution for a reduced FSA. The performance is achieved because of the application of a fail-first heuristics in search tree level and nodal ordering and taking advantage of an efficient search space pruning criterion in search tree generation and in the search process.>
Ruchir Puri
ICTAI1