EDBT 2026 Demo / reviewers in the wild / expert
Baris Taskin
dblp:90/4900
· DBLP profile ↗
81ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-7631-5696ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 79 · 6 first-author · 11 since 2021Software engineering, systems software and programming languages · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ROA-Based Subharmonic Injection Locking for Oscillator-Based Ising MachinesabstractThis paper introduces on-chip integrated rotary traveling wave oscillators (RTWOs) organized into rotary oscillator array (ROA) bricks as an external perturbation to induce subharmonic injection locking (SHIL) in oscillator-based Ising machines (OIMs). The implementation of SHILs on chip is challenging, as the frequency of SHILs must be multiples of the operating frequency of the OIM nodes, with on-chip variations affecting the phase, degrading the SHIL process. This impedes the scaling of OIM implementations, regardless of the topology of Ising nodes, coupling or graph mapping mechanisms. The ROA brick topology implementation of RTWOs generates high frequency signals that are shown to provide a stable 2.31 GHz SHIL signal under process, voltage, and temperature (PVT) variations. Under PVT variations, distributed ring oscillator-based SHILs (ROSC-SHIL) fail to perform injection locking while the proposed ROA brick-based SHIL (ROA-SHIL) preserve 93% to 97% accuracy (the same accuracy of an ideal SHIL signal) in the OIM solutions of a sample 324-node max-cut problem. The driving strength and floorplan of the ROA brick are also shown to be amenable for scaling with an energy-to-solution impact of 2.49 nJ for the proposed ROA-SHIL. Nicholas Sica, Baris Taskin |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | An ASIC Emulated Oscillator Ising/Potts Machine Solving Combinatorial Optimization Problems
Yilmaz Ege Gonul, Baris Taskin |
ISCAS | 2 |
| 2025 | A Multi-Stage Potts Machine Based on Coupled CMOS Ring OscillatorsabstractThis work presents a multi-stage coupled ring oscillator based Potts machine, designed with phase-shifted Sub-Harmonic-Injection-Locking (SHIL) to represent multivalued Potts spins at different solution stages with oscillator phases. The proposed Potts machine is able to solve a certain class of combinatorial optimization problems that natively require multivalued spins with a divide-and-conquer approach, facilitated through the alternating phase-shifted SHILs acting on the oscillators. The proposed architecture eliminates the need for any external intermediary mappings or usage of external memory, as the influence of SHIL allows oscillators to act as both memory and computation units. Planar 4-coloring problems of sizes up to 2116 nodes are mapped to the proposed architecture. Simulations demonstrate that the proposed Potts machine provides exact solutions for smaller problems (e.g. 49 nodes) and generates solutions reaching up to 97% accuracy for larger problems (e.g. 2116 nodes). Yilmaz Ege Gonul, Baris Taskin |
DATE | 2 |
| 2025 | GPU-Accelerated Simulated Oscillator Ising/Potts Machine Solving Combinatorial Optimization Problems
Yilmaz Ege Gonul, Ceyhun Efe Kayan, Ilknur Mustafazade, Nagarajan Kandasamy, Baris Taskin |
ACM Great Lakes Symposium on VLSI | 5 |
| 2024 | Multi-phase Coupled CMOS Ring Oscillator based Potts MachineabstractThis paper presents a coupled ring oscillator based Potts machine to solve NP-hard combinatorial optimization problems (COPs). Potts model is a generalization of the Ising model, capturing multivalued spins in contrast to the binary-valued spins allowed in the Ising model. Similar to recent literature on Ising machines, the proposed architecture of Potts machines implements the Potts model with interacting spins represented by coupled ring oscillators. Unlike Ising machines which are limited to two spin values, Potts machines model COPs that require a larger number of spin values. A major novelty of the proposed Potts machine is the utilization of the N-SHIL (Sub-Harmonic Injection Locking) mechanism, where multiple stable phases are obtained from a single (i.e. ring) oscillator. In evaluation, 3-coloring problems from the DIMACS SATBLIB benchmark and two randomly generated larger problems are mapped to the proposed architecture. The proposed architecture is demonstrated to solve problems of varying size with 89% to 92% accuracy averaged over multiple iterations. The simulation results show that there is no degradation in accuracy, no significant increase in solution time, and only a linear increase in power dissipation with increasing problem sizes up to 2000 nodes. Yilmaz Ege Gonul, Baris Taskin |
ICCAD | 2 |
| 2024 | Design Automation for Charge Recovery LogicabstractThis paper introduces a novel design automation methodology for charge recovery logic (CRL). The proposed methodology combines a novel logic compression algorithm with automatic schematic generation to automate the design process of CRL, enabling power and performance simulations for a large number and variety of CRL circuits. As a measure of the effectiveness of the proposed design flow, automated implementations of CRL equivalents of the LGSynth’91 combinational benchmark circuits are compared with their CMOS counterparts. The results demonstrate a trade-off in power for area: Automatically generated CRL circuits dissipate 51.3% less power on average compared to CMOS equivalents, occupying 54.9% larger area. Yilmaz Ege Gonul, Leo Filippini, Junghoon Oh, Ragh Kuttappa, Scott Lerner, Mineo Kaneko, Baris Taskin |
ISCAS | 7 |
| 2024 | High-Speed Phase-Based ComputingabstractThis work presents the utilization of rotary traveling wave oscillators (RTWOs) to implement an Ising machine. Ising machines utilizing ring oscillators have recently been demonstrated on silicon, for instance, for the solution of a max-cut problem. Rotary traveling wave oscillators scale better in frequency compared to ring oscillators, but have increased power consumption. Phase-based computing principles, implemented with the proposed RTWO-based Ising machines, are prime for high speed phase-based computation. The experiments reveal the proposed RTWO-based Ising machines provide significant reduction (5x) in runtime in the solution of the max-cut problem. The power dissipation is two orders of magnitude higher than the minuscule, low power ring-oscillators but RTWO-based Ising machines sub-linear increase with the demonstrated frequency increase from 2GHz to 32GHz for high speed phase-based computing. The accuracy of the solution is significantly improved as well, as demonstrated with respect to two of the D-Wave solvers (tabu and simulated annealing) acting as the baseline for ring oscillator and RTWO based Ising machines. Nicholas Sica, Ragh Kuttappa, Vinayak Honkote, Baris Taskin |
ISCAS | 4 |
| 2022 | A 0.45 pJ/bit 20 Gb/s/Wire Parallel Die-to-Die Interface with Rotary Traveling Wave OscillatorsabstractIn this work, a die-to-die communication architecture with the integration of resonant clocking is presented. The novelty of the architecture are the rotary traveling wave oscillators, designed across the interposer of a 2. 5D multi-die system to provide a synchronous high frequency clock to all chiplets simultaneously. The transmitter and receiver interface circuits of the architecture benefit from the use of the low power, low skew, multiple phase clock signals across the chiplets. In experimentation, a channel length of 4 mm between transceivers is investigated over a 5 mm $\times 5$ mm silicon interposer. SPICE based simulations with post-layout, parasitic extracted models are performed. The proposed architecture demonstrates 20 Gb/s operation at 0.45 pJ/bit over a 4mm channel at a nominal 1 V supply voltage. The overall clock power of the proposed architecture is 56% lower than prior works at 20 Gb/s. Ragh Kuttappa, Baris Taskin |
ISCAS | 2 |
| 2022 | Resonant Rotary Clock Synchronization with Active and Passive Silicon InterposerabstractRotary traveling wave oscillators (RTWO) are designed to provide a high frequency clock signal through the silicon interposer to multiple chiplets in a heterogeneous 2.5D system. In particular, two different RTWO synchronization topologies are presented: 1) Active interposer RTWO and 2) passive interposer RTWO. The proposed topologies are evaluated across a silicon interposer with a dimension of 42 mm × 20 mm. Each topology is implemented with post-layout, parasitic extracted models for a clock frequency of ≈8 GHz. The performance metrics are presented for clock period, skew, rise time, fall time, and oscillation start-up and settling times across the multi-die system (MDS) with SPICE based simulations. Ragh Kuttappa, Baris Taskin, Vinayak Honkote, Satish Yada, Jainaveen Sundaram, Dileep Kurian, Tanay Karnik, Anuradha Srinivasan |
ISCAS | 2 |
| 2022 | Multiphase Digital Low-Dropout RegulatorsabstractIn this work, multiphase digital low-dropout (MP-DLDO) regulators are designed with resonant rotary clocks (ReRoCs) in order to improve on the tradeoff of conventional DLDOs between current efficiency and transient response speed. The proposed DLDOs are multiphased, coined MP-DLDOs, designed with a clock-gated control technique to provide high current efficiencies along with transient response improvements at GHz frequency levels. The multiple phases within the MP-DLDO are served with ReRoCs that provide: 1) a robust high-speed low-power resonant clock distribution solution for the synchronous elements in the multiphase DLDO architecture and 2) improve the transient response characteristics [dynamic voltage scaling (DVS) speed and voltage ripple] while saving power in the controller circuitry. The proposed MP-DLDOs are distributed across the chip to achieve low voltage ripple. SPICE simulations are performed on post-layout, parasitic-extracted models to evaluate the MP-DLDO architecture on open-source digital cores, with performance metrics that include the voltage ripple reduction, transient response speed improvement, and power savings in the control logic. The proposed MP-DLDO architecture, evaluated on an RISC-V design, demonstrates a DVS speed of 6.5 V/$\mu \text{s}$and an output voltage ripple of 21.1 mV (38% reduction when compared to a conventional DLDO) with a sampling frequency of 2 GHz. Ragh Kuttappa, Selçuk Köse, Baris Taskin |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | Resonant Clock Synchronization With Active Silicon Interposer for Multi-Die SystemsabstractThis paper presents the integration of resonant clocking to multi-die architectures to synchronize individual chiplets connected through an active silicon interposer. The proposed inter-chiplet synchronization through the active silicon interposer rotary oscillator array (ASI-ROA) provides a unitary clock domain to the multiple die (i.e. multiple chiplets) in the package with a very low design overhead. System performance analysis is performed with parasitics-extracted, post-layout simulation models of two different sizes of representative heterogeneous multi-die architectures, each with varying number of RISC-V cores per die. Each RISC-V core of the multi-die package belongs to the unitary clock domain, designed with ASI-ROA to operate at a frequency of 2 GHz. The proposed architecture is investigated for robustness in frequency and skew across the multi-die system (MDS) with SPICE based simulations of post layout models, demonstrating variations of only 80 MHz for a 2 GHz target frequency. The power savings are upto 41% for the overall MDS, compared to an equivalent implementation with a contemporary ADPLL used to synchronize the multiple chiplets over the active interposer. The average clock skew of the completely resonant architecture presented in this work is 8.2 ps. Ragh Kuttappa, Baris Taskin, Scott Lerner, Vasil Pano |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | SnackNoC: Processing in the Communication LayerabstractIn this work, we propose and evaluate a Network-on-Chip (NoC) augmented with light-weight processing elements to provide a lean dataflow-style system. We show that contemporary NoC routers can frequently experience long periods of idle time, with less than 10% link utilization in HPC applications. By repurposing the temporal and spatial slack of the NoC, the proposed platform, SnackNoC, is able to compute linear algebra kernels efficiently within the communication layer with minimal additional resource costs. SnackNoC 'Snack' application kernels are programmed with a producer-consumer data model that uses the NoC slack to store and transmit intermediate data between processing elements. SnackNoC is demonstrated in a multi-program environment that continually executes linear algebra kernels on the NoC simultaneously with chip multiprocessor (CMP) applications on the processor cores. Linear algebra kernels are computed up to 14.2x faster on SnackNoC compared to an Intel Haswell EPx86 processing core. The cost of executing 'snack' kernels in parallel to the CMP applications is a minimal runtime impact of 0.01% to 0.83% due to higher link utilization, and an uncore area overhead of 1.1%. Karthik Sangaiah, Michael Lui, Ragh Kuttappa, Baris Taskin, Mark Hempstead |
HPCA | 4 |
| 2020 | Comprehensive Low Power Adiabatic Circuit Design with Resonant Power ClockingabstractIn this paper, the first comprehensive methodology is presented for design of low power adiabatic circuits inclusive of the adiabatic core design and the power-clock generation. Prior works have focused on either designing adiabatic cores or the power clock generation circuit, only. These non-comprehensive views can misrepresent the performance savings and fail to address the opportunities at integration. In this work, a comprehensive solution is presented that also features a unique innovation for the power clock generation circuit in step-charged circuits designed with rotary traveling wave oscillators (RTWO) and adiabatic frequency dividers. In experimentation, SPICE based simulations are performed at 416 MHz and 330 MHz in the 90 nm technology node and compared to CMOS based implementations, as well as other known power-clock generation techniques. A 32-bit CMOS adder consumes 3.5× more power when compared to the proposed 32-bit ECRL adder operating at a frequency of 416 MHz. Furthermore, 1000 32-bit CMOS adders in parallel consumes 3.4× more power when compared to 1000 32-bit ECRL adders in parallel designed with the proposed architecture at a frequency of 416 MHz. Ragh Kuttappa, Steven Khoa, Leo Filippini, Vasil Pano, Baris Taskin |
ISCAS | 5 |
| 2020 | FinFET - Based Low Swing Rotary Traveling Wave OscillatorsabstractFinFET based, low swing clocking with rotary traveling wave oscillators (RTWO) is presented in this paper. It is shown that the low-swing clock signal generation by RTWOs is very effective, thanks to FinFETs accommodating high frequency operation and voltage scaling better than planar CMOS transistors. Low swing clocks are aimed at lowering the power dissipation of the clock networks, while maintaining the full voltage operation of non-clock components (such as logic and memory). In this work shows that robust low swing (LS) RTWOs are designed with FinFET based technologies. To this end, SPICE simulations are performed on the ISPD'10 clock benchmark circuits operating at 2.25 GHz and 3 GHz in the 16 nm FinFET technology node. LS-RTWO based designs are compared to an all digital phase locked loop (ADPLL) based designs operating at the same target frequency. At 3 GHz, the LS-RTWO consumes 36% lower power with 42.7dB better phase noise @10 MHz on comparison to corresponding ADPLL based designs. Ragh Kuttappa, Baris Taskin |
ISCAS | 2 |
| 2019 | Low Voltage Clock Tree Synthesis with Local Gate ClustersabstractIn this paper, a novel local clock gate cluster-aware low voltage clock tree synthesis methodology is introduced. In low voltage/swing clocking, timing closure is a challenging problem due to tight skew and slew constraints. The clock gating makes this problem more challenging due to the high delay mismatch between the gated and the non-gated sinks. The proposed methodology preserves the power savings of the clock gating and exploits low swing clocking to further reduce the power consumption, while maintaining the same skew and slew constraints as the full swing counterpart. Experimental results performed on the large circuits of ISCAS'89 benchmarks operating at 1.5GHz in the 45nm technology node demonstrate that the proposed methodology can provide 38% power savings as compared to a full swing gated clock tree, achieving an additional 12% savings as compared to a low swing non-gated clock tree. Can Sitik, Baris Taskin, Emre Salman |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | Low Swing - Low Frequency Rotary Traveling Wave OscillatorsabstractThis paper presents the design and implementation of low-swing, low-frequency rotary traveling wave oscillators (RT-WOs). Low-swing rotary clocks are designed to operate with low-swing D flip-flops. A methodology is proposed to perform dynamic frequency scaling for integer division ratios of 3 to n to generate low-frequency rotary clocks at target frequencies. SPICE-based experiments are performed on the ISPD'10 benchmark circuits operating at 100 MHz, 200 MHz, and 500 MHz. The low-swing, low-frequency resonant rotary clock based designs are compared against full-swing and low-swing traditional designs operating at the same frequency and voltage level. The results show that the low-swing rotary clock based ISPD'10 designs at 500 MHz consume 61% and 43% lower power on comparison to the traditional full-swing and low-swing designs, respectively, where the traditional designs are implemented with a PLL, and a bounded-skew tree. Ragh Kuttappa, Scott Lerner, Leo Filippini, Baris Taskin |
ISCAS | 4 |
| 2019 | Robust Low Power Clock Synchronization for Multi-Die SystemsabstractA novel clock generation and distribution network is proposed for multi-die architectures connected through an active silicon interposer. The proposed clock network generates and distributes a resonant clock through the active silicon interposer between dies, with each die served through resonant local clock trees. The proposed active silicon interposer rotary oscillator array (AI-ROA) serves to establish a unitary clock domain, providing constant phase and magnitude clock sources to the multiple die (i.e. multiple chiplets) in the package. Analysis is performed with multiple ARM CORTEX M0 cores per die of a homogeneous multi-die package architecture. Each M0 core of the multi-die package belongs to the unitary clock domain, designed with AI-ROA to operate at a frequency of 1 GHz. The multiple die are designed in the 28 nm technology node and the active interposer is designed in the 65 nm technology node. SPICE based simulations of post-layout models provides analysis and evaluation of the proposed architecture for performance metrics under process, voltage, and temperature variations. In particular, performance metrics are reported for 1) power consumption in comparison to PLL based architectures designed and synthesized with an industrial tool, 2) robustness against process variations, and 3) clock skew across the cores throughout the multiple die. Ragh Kuttappa, Baris Taskin, Scott Lerner, Vasil Pano, Ioannis Savidis |
ISLPED | 2 |
| 2019 | 3D NoCs with active interposer for multi-die systemsabstractAdvances in interconnect technologies for system-in-package manufacturing have re-introduced multi-chip module (MCM) architectures as an alternative to the current monolithic approach. MCMs or multi-die systems implement multiple smaller chiplets in a single package. These MCMs are connected through various package interconnect technologies, such as current industry solutions in AMD's Infinity Fabric, Intel's Foveros active interposer, and Marvell's Mochi Interconnect. Although MCMs improve manufacturing yields and are cost-effective, additional challenges on the Network-on-Chip (NoC) within a single chiplet and across multiple chiplets need to be addressed. These challenges include routing, scalability performance, and resource allocation. This work introduces a scalable MCM 3D interconnect infrastructure called "MCM-3D-NoC" with multiple 3D chiplets connected through an active interposer. System-level simulations of MCM-3D-NoC are performed to validate the proposed architecture and provide performance evaluation of network latency, throughput, and EDP. Vasil Pano, Ragh Kuttappa, Baris Taskin |
NOCS | 3 |
| 2019 | Slew Merging Region Propagation for Bounded Slew and Skew Clock Tree SynthesisabstractBuilding clock trees for tight skew constraints of clock delivery networks is standard in the industry. Tight slew constraints of high-performance designs require post-processing techniques to satisfy slew constraints after clock tree synthesis (CTS). Post-processing adversely impacts the power dissipation. This paper proposes slew merging region CTS (SMRcts); a novel algorithm to satisfy bounded slew and skew constraints simultaneously during synthesis. Experimental results performed on International Symposium on Physical Design (ISPD) 2010 benchmarks using a 20-nm FinFET technology show an average reduction of 15% power over a bounded skew approach. Comparison to the ISPD 2010 CTS contest solutions in the literature shows SMRcts producing a 51% improvement in a utility metric. Scalability of SMRcts is demonstrated on ISPD 2013 benchmarks with up to 100k sinks. Scott Lerner, Baris Taskin |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | Custard: ASIC Workload-Aware Reliable Design for Multicore IoT ProcessorsabstractIn the typical application-specified integrated circuit (ASIC) design flow, reliability-driven performance loss is computed, in part, with switching activity files. However, for ASIC designs of multicore processors, the typical switching activity files lack multithreaded software workload information. An accurate switching activity for a multicore design can be generated using a logic simulator. However, the logic simulator process suffers from long runtimes when dealing with real workloads. This paper analyzes the effects of scaling multithreaded workloads and proposes Custard, a hardware methodology for lifetime improvement of multicore processors by obtaining multithreaded switching activity signatures in a short period of time using a performance simulator (gem5), logic simulator (VCS), and thermal simulator (HotSpot). Custard is particularly important for multicore, Internet of Things processors as the runtime feedback-based reliability mechanisms used on current multicore processors incur area and power overhead that could be prohibitive for smaller form factors and power budgets. Experiments are performed with Custard using real workloads on an OpenSPARC T1 design with two, four, and eight cores that are fully synthesized and routed. The default-sized T1 core is improved to have a reliability increase of 4.1×, with 0.08% and 1.57% increase on average in cell area and switching power, respectively. Scott Lerner, Isikcan Yilmaz, Baris Taskin |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | SLECTS: Slew-Driven Clock Tree SynthesisabstractA slew-driven clock tree synthesis (SLECTS) methodology is proposed for nanoscale technologies where the interconnect resistance dominates device resistance, thereby increasing the challenge of satisfying the slew constraint. This issue is exacerbated at lower voltages due to degraded drive ability of the clock buffers. A paradigm shift from the traditional delay (and skew)-driven approaches to the proposed slew-driven methodology is therefore required. SLECTS is developed in this paper to satisfy tight slew constraints, which can be costly or infeasible with delay (skew)-driven methodologies and reduce the power dissipation of the clock tree, since the slew and skew constraints are simultaneously and methodically considered. Experimental results performed on an industrial circuit with more than 1M gates designed in 28-nm technology demonstrate that clock power is reduced by approximately 15% as compared to a commercial clock tree synthesis tool under similar slew and skew constraints. Can Sitik, Emre Salman, Baris Taskin, Savithri Sundareswaran, Benjamin Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | A 900 MHz Charge Recovery Comparator With 40 fJ per ConversionabstractThe idea of recycling part of the charge used to drive a load is well understood in digital circuits, and falls under the umbrella term of charge recovery logic (CRL). By recovering part of the charge from the load, these circuits achieve lower energy consumption with respect to static CMOS. Recently, a comparator that uses the principles of charge recovery was presented, introducing these energy advantages to the world of mixed-signal circuits. The original design has a maximum operating frequency of 1 kHz, and thus is limited to niche applications. In this work, an improved charge recovery comparator is introduced, operating at up to 900 MHz. Post-layout simulations in 65 nm technology show an energy consumption of 40 fJ per conversion, and an input offset voltage of 32 mV. Leo Filippini, Baris Taskin |
ISCAS | 2 |
| 2018 | Low Frequency Rotary Traveling Wave OscillatorsabstractA methodology is presented in order to automate the selection of a design style for low frequency operation of ultra low power rotary traveling wave oscillators (RTWOs). The methodology systemically analyzes the power dissipation profiles of design styles with two priori known resonant frequency dividers, a static and a dynamic frequency divider. The proposed methodology is very effective in leading to a unique design style solution for systems with one frequency target. For systems with multiple frequency targets, the proposed methodology not only leads to a unique solution but also provides guidance on re-selection of frequency targets for improved performance. A 1.2 GHz RTWO is used to generate a target frequency of 135 MHz, while consuming 0.82× the power of an RTWO for that frequency using a static frequency divider. A dynamic frequency divider is used to generate two target frequencies of 137 MHz and 204 MHz using a master clock of 1.2 GHz, while consuming 0.75× and 0.8× the power respectively for that RTWO frequency. Ragh Kuttappa, Baris Taskin |
ISCAS | 2 |
| 2018 | NoC Router Lifetime Improvement using Per-Port Router UtilizationabstractNetwork-on-chip (NoC) routers have a non-uniform utilization based on the number of links, location, and the workload running. Uneven utilization can lead to reliability issues, namely Negative Bias Temperature Instability (NBTI), that results in a reduced lifetime for gates. To address this, a physical design-based solution using workload signatures and cell sizing is proposed to improve the lifetime of NoCs. Using real workloads in conjunction with logic simulation and physical design, this paper analyzes the utilization of each port of NoC routers and, based on a desired lifetime, resizes the physical design. Results using SPLASH2 benchmarks show that the NoC router lifetime can be increased by 3.4× over the default sized router with an average increase of 5.58% and 5.27% for area and power, respectively. Scott Lerner, Vasil Pano, Baris Taskin |
ISCAS | 3 |
| 2018 | Workload-Aware Routing (WAR) for Network-on-Chip Lifetime ImprovementabstractThe emergence of Network-on-Chip (NoCs) as a scalable interconnection infrastructure for Chip Multi-Processors (CMPs) creates the need to analyze the degradation and lifetime repercussions incurred by network traffic from exa-scale and distributed workloads. Reliability methods and redundancies are in place to circumvent defective parts or to roll back to a previous safe state. These reliability techniques are reactive in nature and do not focus effort in avoiding degradation which may result in full system failure. Current methods for improving degradation and lifetime focus on design and runtime optimizations. This paper proposes a workload-aware routing algorithm, which complements known techniques and optimizes the overall network traffic balance. The end product is a methodology that utilizes workload signatures and priority-based routing to improve port and router utilization across the network. The proposed WAR algorithm stays within ≈2% of the latency and average hop count of a NoC using XY routing. Simulation results show an improvement in network traffic balance upto ≈17% (average improvement of ≈8.6%) with WAR on NoCs for lifetime improvement. Vasil Pano, Scott Lerner, Isikcan Yilmaz, Michael Lui, Baris Taskin |
ISCAS | 5 |
| 2018 | Towards Cross-Framework Workload Analysis via Flexible Event-Driven InterfacesabstractHardware/software co-design and software profiling rest on the ability to perform a range of workload analyses. State-of-the-art tools and methods used in such analyses utilize either custom solutions or complex frameworks. There are two problems with this approach: 1) duplicated development work when moving to new and unsupported frameworks or platforms, and 2) the additional burden of in-depth knowledge required to develop the analysis tools. This work presents a methodology to solve these inefficiencies by decoupling workload analysis from the underlying techniques used to observe the workload. The interface is designed to be cross-platform and presents workloads as a set of configurable events with scalable levels-of-detail. An implementation of the methodology, PRISM, is presented which leverages two popular dynamic binary instrumentation tools, Valgrind and DynamoRIO, and additionally Intel PT via Linux perf. The goals of the methodology are three-fold: modularity, flexibility, and productivity. Three analyses are conducted using PRISM to demonstrate these properties: 1) discrepancies are assessed between workloads generated with Valgrind, DynamoRIO, and perf, 2) scalability of a complex Valgrind trace generation tool is improved, and 3) prototyping of a new dynamic loop detection and data-dependence tool is demonstrated. The average overhead of PRISM compared to in-framework analysis is 33% in the worse case and under 1% during typical analysis. Michael Lui, Karthik Sangaiah, Mark Hempstead, Baris Taskin |
ISPASS | 4 |
| 2018 | SynchroTrace: Synchronization-Aware Architecture-Agnostic Traces for Lightweight Multicore Simulation of CMP and HPC WorkloadsabstractTrace-driven simulation of chip multiprocessor (CMP) systems offers many advantages over execution-driven simulation, such as reducing simulation time and complexity, allowing portability, and scalability. However, trace-based simulation approaches have difficulty capturing and accurately replaying multithreaded traces due to the inherent nondeterminism in the execution of multithreaded programs. In this work, we present SynchroTrace, a scalable, flexible, and accurate trace-based multithreaded simulation methodology. By recording synchronization events relevant to modern threading libraries (e.g., Pthreads and OpenMP) and dependencies in the traces, independent of the host architecture, the methodology is able to accurately model the nondeterminism of multithreaded programs for different hardware platforms and threading paradigms. Through capturing high-level instruction categories, the SynchroTrace average CPI trace Replay timing model offers fast and accurate simulation of many-core in-order CMPs. We perform two case studies to validate the SynchroTrace simulation flow against the gem5 full-system simulator: (1) a constraint-based design space exploration with traditional CMP benchmarks and (2) a thread-scalability study with HPC-representative applications. The results from these case studies show that (1) our trace-based approach with trace filtering has a peak speedup of up to 18.7× over simulation in gem5 full-system with an average of 9.6× speedup, (2) SynchroTrace maintains the thread-scaling accuracy of gem5 and can efficiently scale up to 64 threads, and (3) SynchroTrace can trace in one platform and model any platform in early stages of design. Karthik Sangaiah, Michael Lui, Radhika Jagtap, Stephan Diestelhorst, Siddharth Nilakantan, Ankit More, Baris Taskin, Mark Hempstead |
ACM Trans. Archit. Code Optim. | 7 |
| 2018 | Vertical Arbitration-Free 3-D NoCsabstractThe vertical interlayer communication channel plays a critical role in defining the performance of a 3-D networkon-chip (NoC). In this paper, an arbitration-free design for the shared vertical channels is proposed. The proposed vertical arbitration-free 3-D NoC is compared with other 3-D NoC architectures using traditional synthetic traffic patterns and Rentian traffic emulating applications for chip multiprocessors. The results of the analysis show comparable performance in throughput, energy, and latency compared to a symmetric 3-D NoC with savings up to ≈ 20% in area. The proposed NoC is superior in performance to a 3-D NoC utilizing vertical arbitration with a similar area footprint. Ankit More, Vasil Pano, Baris Taskin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Reconfigurable threshold logic gates using optoelectronic capacitorsabstractThis paper investigates the integration of optoelectronic devices with CMOS to implement reconfigurable threshold logic gates for Boolean functions. The weight of the optoelectronic device can be altered by changing the optical power which is used to reconfigure the threshold logic (TL) gate. These novel reconfigurable gates are called the optoelectronic capacitor based TL (OECTL) gates. The OECTL gates are designed for i) simplistic AND/NAND gates and OR/NOR gates with large fan-in and ii) linearly separable Boolean functions that can be reconfigured to other linearly separable Boolean functions, constrained in reconfiguration by the specifics of TL operation. SPICE simulations in 65nm bulk CMOS technology with a Verilog-A model for the optoelectronic capacitor demonstrate i) AND/NAND gates and OR/NOR gates are 2× faster as fan-in goes above 3 and consumes low power and ii) Boolean functions can be reconfigured with 0.58× smaller delay and 0.46× less power of standard CMOS design. Ragh Kuttappa, Lunal Khuon, Bahram Nabet, Baris Taskin |
DATE | 4 |
| 2017 | Stability of Rotary Traveling Wave Oscillators under process variations and NBTIabstractResonant rotary clocking is a low-power clocking technology for multi-phase clock generation in GHz frequency range. In this paper, Rotary Traveling Wave Oscillators (RTWOs) are analyzed under process variations and negative bias temperature instability (NBTI) at the 90nm technology node. The analysis is focused on 1) variations in the physical geometries of the rotary ring, 2) inter and intra-die transistor variations, 3) power supply fluctuation and 4) NBTI. Monte-Carlo based analysis are performed to study the effects of process variations and transistor aging on the operating frequency and power consumption of the rotary ring at a temperature of 110° C. SPICE based simulations show natural robustness against process variations, and NBTI. Ragh Kuttappa, Leo Filippini, Scott Lerner, Baris Taskin |
ISCAS | 4 |
| 2017 | Special issue on IEEE/ACM System Level Interconnect Prediction (SLIP) Workshop 2016
Tsung-Yi Ho, Baris Taskin |
Integr. | 2 |
| 2016 | Wireless Network-on-Chip analysis of propagation technique for on-chip communicationabstractNetwork-on-Chip (NoC) is a communication paradigm capable of facilitating a scalable interconnection infrastructure for multi core processors. Wireless NoCs have been introduced to improve the communication performance over long-distance processing nodes. Current on-chip antennas used in wireless NoCs communicate predominantly through surface waves, where the efficacy of the wireless nodes is partially determined by the radiation efficiency and transmission gain limited due to the conductivity loss of the silicon substrate. Recently, an on-chip propagation technique of radio waves was introduced, through the un-doped silicon layer as opposed to surface-waves prevalent in literature. The through-substrate propagation waves provide a unique solution to overcome the challenge of long-distance communication between processing nodes. In this work, overall improvements are shown compared to traditional wireless NoCs with the placement of antennas on undoped silicon (i.e. communicating through surface waves), simulated in NoC architectures across performance metrics of area, power consumption and latency. Vasil Pano, Isikcan Yilmaz, Yuqiao Liu 0001, Baris Taskin, Kapil R. Dandekar |
ICCD | 4 |
| 2016 | Energy aware routing of multi-level Network-on-Chip trafficabstractThe emergence of Network-on-Chip (NoC) as a communication paradigm for Multi-Processor System-on-Chips (MPSoCs) significantly exacerbates the need to provide a methodology that optimizes the energy consumption of the overall system. This is especially important when factoring in current Network-on-Chip advances which have multiple communication media such as on-chip wireless or nano-photonics links, hybrid with traditional wired links. All of these media have different energy profiles, and if not taken into consideration the system will incur a higher power consumption throughout the runtime of the application. In this work, the case for EDP (energy-delay product) optimization between different levels of a multi-level Network-on-Chip is presented. Using a dynamic, energy aware algorithm, the EDP improvement is compared to a multi-level Network-on-Chip using a statically optimized routing. The proposed routing algorithm handles the different types of energy-delay profiles of multiple links. The end product is a methodology that lowers the overall energy consumption by optimizing the energy profile of the Network-on-Chip while also minimizing the network delay. Vasil Pano, Isikcan Yilmaz, Ankit More, Baris Taskin |
ICCD | 4 |
| 2016 | Charge recovery logic for thermal harvesting applicationsabstractThis paper investigates the substitution of CMOS or near-threshold CMOS with Charge Recovery Logic (CRL) in applications where energy is thermally harvested. By doing so, it is possible to eliminate the bulky DC/DC stage needed to provide the supply voltage for CMOS operation. Instead, a simple LC-tank oscillator is used to generate a power-clock suitable for CRL operation. Simulation results of a 256-stage inverter chain designed in Efficient Charge Recovery Logic (ECRL) are presented. Two additional novelties are presented i) using ECRL at a near-threshold voltage and ii) generating the four-phase power-clock by means of a quadrature oscillator. The traditional, full-swing CMOS system dissipates 18.2× the power dissipated by the proposed TP-ECRL system. Leo Filippini, Baris Taskin |
ISCAS | 2 |
| 2016 | Exploiting useful skew in gated low voltage clock treesabstractLow swing/voltage clocking is a well-studied approach to reduce dynamic power consumption in clock networks. It is, however, challenging to maintain the same performance at scaled clock voltages due to timing degradation in the Enable paths that are required for clock gating, another highly popular method to reduce dynamic power. A useful skew methodology is proposed in this paper to increase the timing slack of the Enable paths when the clock network is operating at a lower swing voltage. The skew schedule is determined via linear programming. The methodology is evaluated on five largest IS-CAS'89 benchmark circuits. The results demonstrate an average 47% increase in the timing slack of the Enable path, thereby facilitating low swing operation without degrading performance. Emre Salman, Can Sitik, Baris Taskin |
ISCAS | 4 |
| 2016 | Design Methodology for Voltage-Scaled Clock Distribution NetworksabstractA low-voltage/swing clocking methodology is developed through both circuit and algorithmic innovations. The primary objective is to significantly reduce the power consumed by the clock network while maintaining the circuit performance the same. The methodology consists of two primary components: a novel D-flip-flop (DFF) cell that maximizes power savings by enabling low-voltage/swing operation throughout the entire clock network and a novel clock tree synthesis algorithm to ensure that the same timing constraints (i.e., clock frequency, skew, and slew) are satisfied. The proposed methodology is integrated within an industrial design flow. Experimental results on ISCAS'89 benchmark circuits demonstrate that the overall power consumed by the clock tree can be reduced by up to 27% and 44% in, respectively, 32- and 45-nm technologies while satisfying the same timing constraints. Furthermore, the proposed low-swing DFF cell maintains the clock-to-Q delay the same while achieving up to 32% and 15% power savings in the overall flip-flop power of the benchmark circuits at, respectively, 1- and 1.5-GHz clock frequencies. Can Sitik, Baris Taskin, Emre Salman |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Clock Skew Scheduling in the Presence of Heavily Gated Clock NetworksabstractClock skew scheduling is a common and well known technique to improve the performance of sequential circuits by exploiting the mismatches in the data path delays. Existing clock skew scheduling techniques, however, cannot effectively consider heavily gated clock networks where a local clock tree exists between clock gating cells and registers. A methodology is proposed in this paper to efficiently achieve clock skew scheduling in circuits with gated clock networks. The methodology is implemented via both linear programming and constraint graph based approaches, and evaluated using the largest ISCAS'89 benchmark circuits with clock gating. The results demonstrate up to approximately 21% reduction in clock period while maintaining the power savings achieved by clock gating. A conventional design flow is used for the experiments, demonstrating the applicability of the proposed algorithms to automation. Emre Salman, Can Sitik, Baris Taskin |
ACM Great Lakes Symposium on VLSI | 4 |
| 2015 | A Novel Static D-Flip-Flop Topology for Low Swing ClockingabstractLow swing clocking is a well known technique to reduce dynamic power consumption of a clock network. A novel static D flip-flop topology is proposed that can reliably operate with a low swing clock signal (down to 50% of the VDD) despite the full swing data and output signals. The proposed topology enables low swing signals within the entire clock network, thereby maximizing the power saved by low swing operation. The proposed flip-flop is compared with existing low swing flip-flops using a 45 nm technology node at a clock frequency of 1.5 GHz. The results demonstrate an average reduction of 38.1% and 44.4% in, respectively, power consumption and power-delay product. The sensitivity of each circuit to clock swing is investigated. The robustness of the proposed topology is also demonstrated by ensuring reliable operation at various process, voltage, and temperature corners. Mallika Rathore, Emre Salman, Can Sitik, Baris Taskin |
ACM Great Lakes Symposium on VLSI | 5 |
| 2015 | Uncore RPD: Rapid Design Space Exploration of the Uncore via Regression ModelingabstractA regression-based design space exploration methodology is proposed that models the impacts of the memory hierarchy and the network-on-chip (NoC) on the overall chip multiprocessor (CMP) performance. Designers cannot explore all possible designs for a NoC without considering interactions with the rest of the uncore, in particular the cache configuration and memory hierarchy which determine the amount and pattern of the traffic on the NoC. The proposed regression model is able to capture the salient design points of the uncore for a comprehensive design space exploration by designing memory and NoC-specific regression models and leveraging recent advances in uncore simulation. To show the utility of our methodology, Uncore RPD, two case studies are presented: i) analyzing and refining regression models for an 8-core CMP and ii) performing a rapid design space exploration to find best performing designs of a NoC-based CMP given area-constraints for CMPs of up to 64 cores. Through these case studies, it is shown that i) simultaneous consideration of the memory and NoC parameters in the NoC design space exploration can refine uncore-based regression models, ii) sampling techniques must consider the dynamic design space of the uncore, and iii) overall, the proposed regression models reduce the amount of simulations required to characterize the NoC design space by up to four orders of magnitude. Karthik Sangaiah, Mark Hempstead, Baris Taskin |
ICCAD | 3 |
| 2015 | A wirelessly powered system with charge recovery logicabstractIn this paper, charge recovery logic is proposed as an alternative to traditional or near-threshold CMOS logic for high-performance systems where the power is wirelessly delivered, e.g. bio-implantable devices. This approach has two primary, complementary advantages in i) providing a wirelessly transmitted sine-wave as the power clock source to the charge recovery logic and ii) eliminating the AC/DC power stage required to provide a stable supply voltage needed in CMOS circuits. The paper presents solutions to the main obstacles of this method and shows simulation results of a simple logic load designed in Efficient Charge Recovery Logic (ECRL) as part of a wireless powered system. The designed wirelessly powered ECRL (coined WP-ECRL) system i) consumes 15.2 × less power than full-swing CMOS and ii) operates at higher frequencies than near-threshold CMOS. These comparative trends in power dissipation are for the computing circuit only, and do not include the bulky AC/DC stage that would be necessary for CMOS implementations. In terms of resilience, it is shown that logic functionality is preserved even when the coupling coefficient of the wireless link is decreased by 60% from the nominal value or when coils with very poor quality factor (down to Q = 0.1) are used. Leo Filippini, Emre Salman, Baris Taskin |
ICCD | 3 |
| 2015 | Enhanced level shifter for multi-voltage operationabstractA novel level-up shifter with dual supply voltage is proposed. The proposed design significantly reduces the short circuit current in conventional cross-coupled topology, improving the transient power consumption. Compared with the bootstrapping technique, the proposed circuit consumes significantly less area, making it more practical for ICs with a large number of supply voltages. The minimum power-delay product (PDP) for each level shifter is analyzed and compared. Worst-case corner analysis is performed for transient power, delay, and leakage power. The dependence of power and delay on input supply voltage level is also investigated for each topology. Simulation results demonstrate 43% and 36% reduction in, respectively, transient power and leakage power as compared to cross-coupled level shifter, while consuming 9.5% and 79.5% less physical area than, respectively, cross-coupled and bootstrapping techniques. Emre Salman, Can Sitik, Baris Taskin |
ISCAS | 4 |
| 2015 | Synchrotrace: synchronization-aware architecture-agnostic traces for light-weight multicore simulationabstractTrace-driven simulation of chip multiprocessor (CMP) systems offers many advantages over execution-driven simulation, such as reducing simulation time and complexity, and allowing portability, and scalability. However, trace-based simulation approaches have encountered difficulty capturing and accurately replaying multi-threaded traces due to the inherent non-determinism in the execution of multi-threaded programs. In this work, we present SynchroTrace, a scalable, flexible, and accurate trace-based multi-threaded simulation methodology. The methodology captures synchronization- and dependency-aware, architectureagnostic, multi-threaded traces and uses a replay mechanism that plays back these traces correctly. By recording synchronization events and dependencies in the traces, independent of the host architecture, the methodology is able to accurately model the non-determinism of multi-threaded programs for different platforms. We validate the SynchroTrace simulation flow by successfully achieving the equivalent results of a constraint-based design space exploration with the Gem5 Full-System simulator. The results from simulating benchmarks from PARSEC 2.1 and Splash-2 show that our trace-based approach with trace filtering has a peak speedup of up to 18.4ξ over simulation in Gem5 Full-System with an average of about 7.5ξ speedup. We are also able to compress traces up to 74% of their original size with almost no impact on accuracy. Siddharth Nilakantan, Karthik Sangaiah, Ankit More, Giordano Salvador, Baris Taskin, Mark Hempstead |
ISPASS | 5 |
| 2015 | FinFET-Based Low-Swing ClockingabstractA low-swing clocking methodology is introduced to achieve low-power operation at 20nm FinFET technology. Low-swing clock trees are used in existing methodologies in order to decrease the dynamic power consumption in a trade-off for 3 issues: (1) the effect of leakage power consumption, which is becoming more dominant when the process scales sub-32nm; (2) the increase in insertion delay, resulting in a high clock skew; and (3) the difficulty in driving the existing DFF sinks with a low-swing clock signal without a timing violation. In this article, a FinFET-based low-swing clocking methodology is introduced to preserve the dynamic power savings of low-swing clocking while minimizing these three negative effects, facilitated through an efficient use of FinFET technology. At scaled performance constraints, the proposed methodology at 20nm FinFET leads to 42% total power savings (clock network+DFF) compared to a FinFET-based full-swing counterpart at the same frequency (3 GHz), thanks to the dynamic power savings of low-swing clocking and 3% power savings compared to a CMOS-based low-swing implementation running at the half frequency (1.5 GHz), thanks to the leakage power savings of FinFET technology. Can Sitik, Emre Salman, Leo Filippini, Sung-Jun Yoon, Baris Taskin |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2015 | Locality-Aware Network Utilization Balancing in NoCsabstractHierarchical and multi-network networks-on-chip (NoCs) have been proposed in the literature to improve the energy- and performance-efficient scalability of the traditional flat-mesh NoC architecture. Theoretically, based on a small-world network-based analysis, traditional hierarchical NoCs are expected to provide good scalability. However, the traditional theoretical analysis (e.g. for small-worldness) does not take into account the congestion phenomenon experienced in such networks. Counterintuitively, as shown in this work, breaking the hierarchy in traditional hierarchical NoCs and utilizing the proposed locality-aware network utilization (NU) balancing technique performs better. This improvement in performance is observed through experimental analysis, which is contrasted with the theoretical analysis that does not account for congestion. In addition to the novelties for hierarchical networks, the application of the proposed locality-aware NU balancing scheme is extended to multi-network NoC topologies (with already separated networks). Results of the analysis show the superiority of applying the locality-aware NU balancing technique for a throughput and energy-efficient scaling of the multi-network NoC architectures, much like those of the hierarchical NoCs. For instance, for a NoC with 1024 nodes, the proposed NU balancing technique provides up to 95% higher throughput efficiency and consumes up to 29% less energy per flit compared to the best NoC topology without the NU balancing technique. The analysis also helps to render the choice of a NoC topology for traffic patterns varying in locality and nonlocality on exascale computing CMPs. Ankit More, Baris Taskin |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2015 | ROA-Brick Topology for Low-Skew Rotary Resonant Clock Network DesignabstractThis paper presents a topology-based solution for a low-skew rotary oscillator array (ROA) clock distribution network design. An ROA-brick structure is proposed that limits the traveling wave oscillation to only two uniform ring rotation directions in the ROA-brick: all the rings in clockwise (CW) direction or all the rings in counter CW direction. An ROA built from the ROA-bricks has the following advantages: 1) similar to the ROA-brick, only two uniform ring rotation directions are feasible in the ROA; 2) the same phase tapping points of all the rings in the ROA are identifiable; and 3) these same phase tapping points of the ROA are independent from the two possible rotation directions. It is mathematically proved that the ROA-brick is the only ROA structure, which can limit the ring rotation direction combinations so as to guarantee the generation of same phase clock signals. The proposed brick-based ROA clock generation and distribution networks are designed for ISPD 10 clock benchmarks demonstrating the gigahertz operation with the low-skew clock generation and distributions through HSPICE. Ying Teng, Baris Taskin |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Frequency-centric resonant rotary clock distribution network designabstractA frequency-centric methodology is proposed for the selection of the physical parameters of the resonant rotary clock for a target frequency. This proposed methodology is performed once for each target frequency on a semiconductor technology, in order to create a cell library of resonant rotary clock design components. A case study is performed to demonstrate the efficiency of the proposed methodology in accuracy and run-time. Simulation results show that the frequency-centric design provides a good approximation for the resonant clock distribution network. At target frequencies between 4GHz and 6GHz, the frequency difference is less than 0.20% and the run-time is reduced by approximately 70% compared to the traditional HSPICE simulation of an entire SROA network without the proposed simplifications for run-time improvement. Ying Teng, Baris Taskin |
ICCAD | 2 |
| 2014 | Static thread mapping for NoCs via binary instrumentation tracesabstractA novel methodology is proposed for thread mapping on a chip-multiprocessor (CMP) system with a network-on-chip (NoC). This novel mapping leverages multi-threaded traces produced by a binary instrumentation tool, which classifies the communication and computation events for each thread of a multi-threaded program application. Processing these binary instrumentation traces after profiling, a static thread mapping is computed to improve the NoC performance. Giordano Salvador, Siddharth Nilakantan, Baris Taskin, Mark Hempstead, Ankit More |
ICCD | 3 |
| 2014 | Timing characterization of clock buffers for clock tree synthesisabstractIt is formidable to embed iterative simulations into the clock tree synthesis process to verify the skew and slew constraints. Instead, accurate and simple timing models for clock buffers are traditionally used so as to perform clock tree synthesis with sufficient accuracy. Two-pole RC and/or piecewise linear models accurately models the gate delay without a waveform dependency for a wide range of waveform properties. However, they unnecessarily complicate the problem for the time modeling of clock buffers where, unlike logic gates, the input and output waveform properties are similar. Look-up table-based approaches are traditionally used in order to obtain the clock buffer timing with inputs being the input slew and the output capacitance. However, the effective capacitance estimation of the highly resistive wires of sub-45nm technologies is a challenge, making it hard to identify the output capacitance. Also, the multiple or dynamically-scaled voltage levels of the current designs necessitate a costly LUT-based pre-characterization process. In this work, a timing estimation scheme for clock buffers is proposed which models both the delay and the slew as linear equations, bypassing the costly LUT characterization process. The experimental results performed with SAED 32nm buffer library show that the proposed timing model can achieve a maximum absolute value error of ≈5ps to ≈10ps for the buffer timing compared to SPICE simulations. Furthermore, the proposed timing model provides an error from 0.2% to 4.6% at different timing constraints and operating voltage levels, when used for insertion delay computation. Can Sitik, Scott Lerner, Baris Taskin |
ICCD | 3 |
| 2014 | Iterative skew minimization for low swing clocks
Can Sitik, Baris Taskin |
Integr. | 2 |
| 2013 | Sparse-rotary oscillator array (SROA) design for power and skew reductionabstractThis paper presents a unique rotary oscillator array (ROA) topology—the sparse-ROA (SROA). The SROA eliminates the need for redundant rings in a typical, mesh-like rotary topology optimizing the global distribution network of the resonant clocking technology. To this end, a design methodology is proposed for SROA construction based on the distribution of the synchronous components. The methodology eliminates the redundant rings of the ROA and reduces the tapping wirelength, which leads to a power saving of 32.1%. Furthermore, a skew control function is implemented into the SROA design methodology as a part of the optimization of the connections among tapping points and subtree roots. This control function leads to a clock skew reduction of 47.1% compared to a square-shaped ROA network design, which is verified through HSPICE. Ying Teng, Baris Taskin |
DATE | 2 |
| 2013 | Skew-bounded low swing clock tree optimizationabstractThis paper introduces a methodology that optimizes the performance of a low swing clock tree under a skew bound. Low-swing clock trees are preferred for a reduction in the clock switching power, with an expected trade-off in clock slew and skew. In this paper, a heuristic optimization process is introduced that keeps the clock skew under the same skew budget of the originating full-swing clock tree. In this low swing clock optimization, the low power consumption property is preserved. The effect of slew on the logic timing, which is naturally degraded due to low-swing operation, is analyzed within timing slack of some paths in order to highlight the effectiveness of the low swing clock trees in lowering power consumption with limited impact on timing constraints. The experiments performed with the 4 largest ISCAS'89 benchmark circuits operating at 500~MHz, 90~nm technology and 4 different Vdd levels show that the optimized low swing clock tree can achieve an average of upto 11.0% reduction in the power consumption with no more than a skew degradation of 0.5% of the clock period (i.e. within the practical skew budget). Can Sitik, Baris Taskin |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | Multi-corner multi-voltage domain clock mesh designabstractThis paper introduces a novel multi-voltage domain clock mesh design methodology that is effective under multiple process corners. In multi-voltage designs, a single clock mesh that spans multiple voltage domains is infeasible due to the incompatibility of voltage levels of the clock drivers on the electrically-shorted mesh-each voltage domain requires a separate mesh. The skew among these isolated meshes need to be matched and a novel premesh tree synthesis is required to tolerate the impact of PVT variations exacerbated due to the separation of clock meshes for multiple voltage levels. The experiments performed with the largest three ISCAS'89 benchmark circuits operating at 500 MHz, 90 nm technology and 3 process corners show that: 1) The multi-voltage domain clock mesh can achieve up to 42% lower power on average with 39.04 ps skew, as low as ≈1.95% of the clock period, on average over a typical single voltage domain clock mesh, and 2) multi-corner optimized multi-voltage domain clock mesh can decrease the global skew of all corners by 190.42 ps on average over a multi-voltage domain clock mesh optimized for a single corner with a 15% degradation in power consumption necessary for variation-tolerance. Can Sitik, Baris Taskin |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | Rotary traveling wave oscillator frequency division at nanoscale technologiesabstractNo abstract available. Ying Teng, Baris Taskin |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | Resonant frequency divider design methodology for dynamic frequency scalingabstractA rotary traveling wave oscillator (RTWO) frequency divider design methodology is proposed for dynamic frequency scaling. The proposed methodology can be used for designing dividers for integer division ratios of 3 to 9 within one circuit topology. HSPICE-based experiments are performed to test the electrical characteristics of the RTWO frequency dividers. The simulation results show that the power consumption of a frequency divider is as low as approximately 5mW for different frequency division ratios. Ying Teng, Baris Taskin |
ICCD | 2 |
| 2012 | Synchronization scheme for brick-based rotary oscillator arraysabstractIn this paper, a brick-based rotary oscillator array (ROA) synchronization scheme is proposed, which directs all the rotary traveling wave oscillators (RTWOs) in the ROA to rotate in a pre-determined direction. This synchronization scheme increases the speed of the ROA synchronization process by eliminating the repetitive start-up trials due to start-ups from incorrect points on the oscillatory array. Simulation results confirm the effectiveness of the ROA synchronization scheme. Furthermore, the synchronization scheme is applied to an ROA-based clock generation and distribution network designed for an ISPD 10 clock benchmark in order to demonstrate its application at a larger scale. Ying Teng, Baris Taskin |
ACM Great Lakes Symposium on VLSI | 2 |
| 2012 | High-Performance, Low-Power Resonant Clocking: Embedded tutorialabstractClock distribution networks consume a significant portion of on-chip power. Traditional buffered clock distribution power is limited by frequency, capacitance, and activity rates. Resonant clock distributions can reduce this power by "recycling" energy on-chip and reducing the overall clock power. This tutorial introduces recent techniques for distributed-LC, traveling wave, and standing wave resonant clock distributions. In particular, the tutorial discusses the recent developments and open research problems. The tutorial covers both circuits, computer-aided design algorithms and methodologies for resonant clocking. Matthew R. Guthaus, Baris Taskin |
ICCAD | 2 |
| 2012 | Clock mesh synthesis with gated local trees and activity driven register clusteringabstractA clock mesh network synthesis method is proposed which enables clock gating on the local sub-trees in order to reduce the clock power dissipation. Clock gating is performed with a register clustering strategy that considers both i) the similarity of switching activities between registers in a local area and ii) the timing slack on every local data path in the design. This is the first work known in literature that encapsulates the efficient implementation of the gated local trees and activity driven register clustering with timing slack awareness for clock mesh synthesis. Experimental results show that with gated local tree and activity driven register clustering, the switching capacitance on the mesh network can be reduced by 22% with limited skew degradation. The proposed method has two synthesis modes as low power mode and high performance mode to serve different design purposes. Jianchao Lu, Xiaomi Mao, Baris Taskin |
ICCAD | 3 |
| 2012 | Multi-voltage domain clock mesh designabstractThis paper investigates the effectiveness of a multi-voltage clock network design that is built using the mesh topology. Unlike a clock tree, a single clock mesh that spans multiple voltage domains is infeasible due to the incompatibility of voltage levels of the clock drivers on the electrically-shorted mesh - each voltage domain requires a separate mesh. These disjoint meshes need to be matched in clock skew between the domains. In addition, the additional power dissipation of the level shifters in the logic needs to be compared against the power savings of multi-voltage domain implementation. The case study performed with the largest ISCAS'89 benchmark circuits operating at 500 MHz, 90 nm technology concludes two important results that highlight the benefits of multi-voltage clock mesh design: 1) The multi-voltage domain clock mesh can achieve 37.14% lower power with a 9 ps increase in clock skew over the single-voltage domain clock mesh, and 2) The multi-voltage domain clock mesh achieves 66 ps less skew with a 20.92% increase in power dissipation over a multi-voltage domain clock tree. Can Sitik, Baris Taskin |
ICCD | 2 |
| 2012 | Clock mesh synthesis method using the Earth Mover's Distance under transformationsabstractA novel clock mesh generation method is proposed based on the EMDg(Earth Mover's Distance under transformations) algorithm. A bottom-up approach is adopted in creating local-level tree clusters to drive the generation of a regional-level uniform clock mesh. The EMDgmethod incrementally moves the regional-level uniform clock mesh closer to the register cluster roots in order to reduce the total stub wirelength. Post-EMDgmesh reduction, the redundant mesh wires are eliminated from the initial uniform mesh in order to reduce the mesh wirelength, preserving the stub wire connections and the integrity of the clock mesh. The optimization results show that the proposed method can achieve an average total wirelength saving of 20.1% and power savings of 12.2% on a suite of ISCAS'89 benchmark circuits compared to the previous clock mesh generation methods. Ying Teng, Baris Taskin |
ICCD | 2 |
| 2012 | A unified design methodology for a hybrid wireless 2-D NoCabstractHybrid wireless NoCs are proposed to improve the communication throughput and energy compared to flat mesh-based NoCs. However, the multi-faceted design of a hybrid wireless NoC presents a conundrum amongst the different design paradigms. In this work, a methodology with minimal design complexity is presented for the design of hybrid wireless NoCs encompassing the different paradigms. To illustrate one use of the methodology, path loss models for the wireless channel implemented with various operating frequencies and substrate types are presented which are to be used in the design space exploration of the NoC. Ankit More, Baris Taskin |
ISCAS | 2 |
| 2012 | Integrated Clock Mesh Synthesis With Incremental Register PlacementabstractA clock mesh planning and synthesis method is proposed which significantly reduces the power dissipation on the network while considering the power density and timing slack simultaneously. The proposed method is performed at the postplacement stage and consists of three major steps: 1) feasible moving region construction of each register considering timing slack; 2) mesh grid wire generation and placement; and 3) incremental register placement for stub wire minimization considering power density and timing slack. The advantages of the proposed method are the reduced power dissipation-28% on average on the benchmark circuits-the optimized power density, and the guaranteed nonnegative timing slack. These advantages are possible through a decreased timing slack (1.1% of the clock period) and change in the logic wirelength (+5.9%) on the benchmark circuits. Jianchao Lu, Xiaomi Mao, Baris Taskin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2012 | ZeROA: Zero Clock Skew Rotary Oscillatory ArrayabstractResonant rotary clocking is a clocking technology for high frequency clock generation and distribution at a low power dissipation rate. It is commonly conceived that the multiple phases on the rings of the rotary oscillatory array (ROA) necessitate a non-zero clock skew operation. In this paper, the feasibility of zero clock skew synchronization with the rotary clocking technology implemented on the ROA is shown. Design automation experiments are performed to demonstrate that the zero clock skew operation can be achieved with minimal change in the performance of rotary clock operation. In particular, a marginal ±1.5% change in the tapping wirelength and a negligible 0.38% average skew mismatch are reported in experiments on R1-R5 and ISPD 2010 benchmark circuits. Vinayak Honkote, Baris Taskin |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | A Reconfigurable Clock Polarity Assignment Flow for Clock Gated DesignsabstractThis paper presents a clock polarity assignment flow which permits post-silicon reconfigurability. The proposed method inserts xor gates at one level of the clock tree to facilitate the polarity assignment. The polarity of the xor gates can be reconfigured for different modes of clock gating (sleep mode, busy mode, etc.) such that a mode-specific reduction of the peak current can be achieved. Experimental results show that the worst case peak current on a clock tree can be reduced by 33.3% by assigning polarity to xor gates at the sink level of the clock tree. An additional 12.8% reduction in the worst case peak current can be achieved by reconfiguring the polarity assignment based on the clock gating information. The proposed flow increases the area by 7.1% but reduces both the total power consumption by 23.8% and the global skew increase (due to polarity assignment) from 19.3 to 8.8 ps. The insertion of xor gates at the non-sink nodes is also studied to further reduce the global skew increase and the area overhead. Jianchao Lu, Ying Teng, Baris Taskin |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Steiner tree based rotary clock routing with bounded skew and capacitive load balancingabstractA novel rotary clock network routing method is proposed for the low-power resonant rotary clocking technology which guarantees: 1. The balanced capacitive load driven by each of the tapping points on the rotary rings, 2. Customized bounded clock skew among all the registers on chip, 3. A sub-optimally minimized total wirelength of the clock wire routes. In the proposed method, a forest of steiner trees is first created which connects the registers so as to achieve zero skew and greedily balance the total capacitance of each tree. Then, a balanced assignment of the steiner trees to the tapping points is performed to guarantee a balanced capacitive load on the rotary network. The proposed routing method is tested with the ISPD clock network contest and IBM r1-r5 benchmarks. The experimental results show that the capacitive load imbalance is very limited. The total wirelength is reduced by 64.2% compared to the best previous work known in literature through the combination of steiner tree routing and the assignment of trees to the tapping points. The average clock skew simulated using HSPICE is only 8.8ps when the bounded skew target is set to 10.0ps. Jianchao Lu, Vinayak Honkote, Baris Taskin |
DATE | 4 |
| 2011 | EM and circuit co-simulation of a reconfigurable hybrid wireless NoC on 2D ICsabstractThe feasibility of the dynamic reconfigurability of the network layer of a hybrid wireless network-on-chip (NoC) that uses on-chip antennas for the wireless network layer and metal interconnects for the wired network layer is studied. The reconfigurability of the NoC is analyzed using a circuit co-simulation technique with a 3D finite element method (FEM) based full-wave electro-magnetic analysis of the antennas. The die and the circuits are modeled according to a typical complementary metal oxide semiconductor (CMOS) technology. It is shown that, it is possible to have 1) at least two different frequency domains for the signal sources and 2) the dynamic switching of the signal sinks between the two frequency domains, with minimal design and area overhead. When implemented, the proposed reconfigurable hybrid network architecture can reduce the latency and increase the network throughput. Ankit More, Baris Taskin |
ICCD | 2 |
| 2011 | ROA-brick topology for rotary resonant clocksabstractThis paper presents a topology design-based solution that addresses one of the major challenges in the design of Rotary Traveling Wave Oscillator (RTWO) based clock networks-the direction of oscillation. A “rotary oscillator array (ROA) brick” structure is proposed that guarantees the consistency of the rotation direction of the traveling signals on all the RTWO rings in an ROA. The ROA built from ROA bricks has the following advantages: (1) The same phase point of all the RTWO rings in the array can easily be tracked, (2) The same phase points of the ROA are independent of the specific rotation direction of the traveling signals on the ROA. SPICE simulations demonstrate these advantages of the brick-based ROA circuit design in establishing the directional consistency of the RTWO rings. Ying Teng, Jianchao Lu, Baris Taskin |
ICCD | 3 |
| 2011 | Register On MEsh (ROME): A novel approach for clock mesh network synthesisabstractA clock mesh network synthesis and optimization flow is proposed which entails the optimal mesh size selection, incremental register placement, mesh reduction and buffer driver insertion. The proposed method is based on incrementally placing the registers on a mesh, which gives the method its name "Register on MEsh (ROME)". The primary objectives of ROME are low global clock skew and power dissipation, which are achieved through a sparse mesh implementation with registers mesh. Experimental results show that the total wirelength on the clock mesh (grid wires and stub wires) is reduced by 36.1% with a 2.8ps clock skew improvement. The total power consumption of the experimented circuits is reduced by 14.1% on average. Jianchao Lu, Yusuf Aksehir, Baris Taskin |
ISCAS | 3 |
| 2011 | Reconfigurable clock polarity assignment for peak current reduction of clock-gated circuitsabstractThis paper presents a novel clock polarity assignment method to reduce the peak current on the vdd/gnd rails of a clock-gated integrated circuit. The proposed method inserts XOR gates at one level of the clock tree to facilitate the polarity assignment with limited skew degradation. The polarity of clock buffers are configured during runtime such that a maximal peak current reduction is obtained after clock gating. The method is integrated into an industrial design flow to study the practicality. Experimental results show that the worst case peak current on a clock tree can be reduced by 33.3% and 33.9% by inserting XOR gates at the sink level and non-sink level of the clock tree, respectively. Additional 12.8% and 12.9% reductions in the worst case peak current for a clock tree with XOR gates inserted at the sink and non-sink level, respectively, can be achieved by reconfiguring the polarity assignment during runtime based on the clock gating information. Jianchao Lu, Baris Taskin |
ISCAS | 2 |
| 2011 | Timing slack aware incremental register placement with non-uniform grid generation for clock mesh synthesisabstractA novel clock mesh network synthesis approach is proposed in this paper which generates an improved mesh size with registers placed incrementally considering the timing slack on the data paths and the non-uniform grid wire placement. The primary objective of the method is to reduce the power dissipation without a global skew degradation, which is achieved through a sparse and non-uniform mesh implementation with registers incrementally placed in close vicinity to the mesh grids. The incremental register placement is based on the timing information in order to preserve the timing slack of the circuit. Experimental results show that the total wirelength (mesh grid wires and stub wires) as well as the power dissipation is reduced significantly on the clock mesh network. Specifically, the wirelength of the mesh network and the power dissipation of the clock network are reduced by 52% and 48% on average, respectively. Moreover, the global clock skew and the non-negative timing slack are preserved. Jianchao Lu, Xiaomi Mao, Baris Taskin |
ISPD | 3 |
| 2011 | Clock buffer polarity assignment with skew tuningabstractA clock polarity assignment method is proposed that reduces the peak current on the vdd/gnd rails of an integrated circuit. The impacts of (i) the output capacitive load on the peak current drawn by the sink-level clock buffers, and (ii) the buffer/inverter replacement scheme of polarity assignment on timing accuracy are considered in the formulation. The proposed sink-level-only polarity assignment is performed by a lexi-search algorithm in order to balance the peak current on the clock tree. Most of the previous polarity assignment methods that do not include clock tree resynthesis lead to an undesirable increase in the worst corner clock skew. Hence, a skew-tuning scheme is proposed that reduces the clock skew through polarity refinement and not through clock tree resynthesis. The proposed polarity assignment method with the skew-tuning scheme is implemented within an industrial design flow for practicality. Experimental results show that the worst-case peak current drawn by the clock tree can be reduced by an average of 36.5%. The worst corner clock skew is increased from 60.7ps to 76.2ps by applying the proposed polarity assignment method. The proposed skew-tuning scheme reduces the worst-case clock skew from 76.2ps to 61.5ps, on average, with a limited degradation in the peak current improvement (36.5% to 31.2%, on average). Jianchao Lu, Baris Taskin |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2011 | CROA: Design and Analysis of the Custom Rotary Oscillatory ArrayabstractRotary clocking is a resonant clocking technology for clock network design and distribution in high performance digital VLSI circuits. Rotary clocking technology offers an attractive alternative to the conventional clocking with high frequency clock signal generation at a low power dissipation rate. Traditionally, rotary clocking has been implemented using a regular array (grid) topology called rotary oscillatory arrays (ROA). In this paper, a custom rotary oscillatory array (CROA) topology is proposed for the generation and distribution of rotary clocking. The issues related to timing closure are addressed and the simulation-based analysis of the custom rotary rings is presented. The CROA design methodology is tested on the IBM R1-R5 benchmark circuits. Compared to the traditional ROA, custom ROA results in 39.25% of tapping wirelength savings. The parasitic effects due to the customization of the topology - computed with partial element equivalent circuit (PEEC) analysis - are incorporated and the CROA topologies are simulated in SPICE. The simulation results show that, with additional parasitics due to the topological factors, the resultant clock frequency is observed to be 8.79% slower (assuming the tapping wirelength remains the same) than the expected frequency of operation without considering the topological factors. Vinayak Honkote, Baris Taskin |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Electromagnetic interaction of on-chip antennas and CMOS metal layers for wireless IC interconnectsabstractThe electromagnetic interaction of on-chip antennas and metal interconnects modeled in a 250 nm complementary metal-oxide semiconductor (CMOS) technology is investigated. A finite element method (FEM) based 3-D full-wave solver is used to perform the electromagnetic field analysis. It is shown that there can be significant signal coupling between the on-chip transmitting antenna and the metal interconnects on a die (-12.09 dB for a 1.6 mm long, 2 um wide interconnect at a distance of 1 um from the antenna). Design considerations for metal interconnects in the presence of on-chip antennas are presented in order to limit the undesirable electromagnetic coupling. Ankit More, Baris Taskin |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Skew-aware capacitive load balancing for low-power zero clock skew rotary oscillatory arrayabstractRotary clocking is a traveling wave based high-speed resonant clocking technology with low-power and controllable-skew properties. Capacitive load balance and bounded clock skew are identified as the primary requirements to maintain a stable oscillation frequency across the rings and to achieve timing closure, respectively, in the rotary oscillatory array (ROA). Towards this end, two methodologies are proposed to achieve balanced capacitive loads across the rings of the ROA with a bounded skew constraint. Experiments performed on IBM R1-R5 benchmark circuits show a 5.62X improved capacitive balance and a 3.67% improved clock skew to a total skew of 6.55% of the clock period at 1.8GHz. SPICE simulations show that the frequency variation across the rings of the ROA is reduced from 10.14% to 2.12% as well. Power dissipated with the proposed optimization methodologies are within ±1.5% of the conventional design automation techniques for rotary synchronization. Vinayak Honkote, Baris Taskin |
ICCD | 2 |
| 2010 | PEEC based parasitic modeling for power analysis on custom rotary ringsabstractResonant rotary clocking is a low power-high speed clock distribution technology for the modern VLSI circuits. Alternative topological implementations of rotary clocking with non-regular custom rings have been proposed in literature. In this paper, the impact of parasitics of the non-regular topological geometries on the rotary operating characteristics is presented. In particular, partial element equivalent circuit (PEEC) analysis is used to show that the corner geometry in a custom ring increases the mutual inductance approximately by 80%. Also, SPICE simulations are performed where the parasitics due to the topological factors are incorporated for an 8% increased accuracy in simulation. Further, the power dissipation on the rotary ring is analyzed with varying number of corners. When tested with the IBM R1-R5 benchmark circuits, the total power dissipated on a custom ring (corners between 4 and 12) is within ±5% of the total power dissipated on a regular ring(4 corners). Vinayak Honkote, Baris Taskin |
ISLPED | 2 |
| 2009 | A shift-register-based QCA memory architectureabstractA quantum-dot cellular automata (QCA) design of an nxm -bit, shift-register-based memory architecture is presented. The architecture maintains data at a stable conformation, which is contrary to traditional data in-motion concept for QCA architectures. The memory architecture is based on an existing dual-phase-synchronized, line-based, one-bit QCA memory cell building block that provides size and latency improvements over other known one-bit memory cells through its novel clocking scheme. Read/write latencies up to ∼2X lower than the existing tile-based architecture with three-phase, line-based memory cells are obtained. Simulations with QCADesigner and HDLQ are performed on a sample 4 x 8 bit memory architecture implementation. Baris Taskin, Andy Chiu, Jonathan Salkind, Daniel Venutolo |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2009 | Custom topology rotary clock router with tree subnetworksabstractIncreasing demands on computing power have spurred the development of faster, higher-density Integrated Circuits (ICs), compounding power and complexity concerns in design budgets. The clock distribution network is a significant contributor to such power and complexity concerns. Resonant rotary clocking is a relatively new technology that realizes several benefits over current clocking methods, including power, frequency, and variation tolerance, yet lacks the automation tools to promote increased use. Towards this end, an automated rotary clock routing methodology is presented that generates custom topology rotary ring routes with tree subnetworks. In addition to the benefits of adiabatic clocking, the presented custom topology router permits 38.6% shorter wirelengths on average for register tapping, compared to traditional prescribed skew, binary tree routing. Baris Taskin, Joseph Demaio, Owen Farell, Michael Hazeltine, Ryan Ketner |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2008 | Custom rotary clock routerabstractTiming closure and power envelopes for contemporary multi-core chips with high speed clock networks make the clock distribution design a challenging task. Resonant rotary clocking is a novel clocking technology for multi-gigahertz rate clock generation that provides minimal power dissipation. Rotary clocking implementations can easily provide independent synchronization of multiple cores as well. The traditional rotary clock design involves a regular array topology of oscillatory rings. In this paper, the rotary clock networks are designed and implemented using a custom ring topology. Custom ring topologies are advantageous as they reduce the total tapping wirelength for the registers tapping onto the oscillatory rings. A maze router based algorithm is developed for the implementation of custom topology rotary rings. In experiments performed on UCLA IBM R1-R5 benchmark circuits with the Elmore delay model, an improvement of 11.04% for register tapping wirelength is achieved on average. Vinayak Honkote, Baris Taskin |
ICCD | 2 |
| 2008 | Improving Line-Based QCA Memory Cell Design Through Dual Phase ClockingabstractThis paper describes a line-based, quantum-dot cellular automata (QCA) memory cell design that is synchronized by a dual-phase clocking scheme. In line-based QCA memory cells, data bits are stored oscillating along QCA lines. The best known line-based memory cell implementation requires three new clocking zones in addition to the four clocking zones defined by the conventional QCA clocking scheme and utilizes three parallel clocking zones per cell. The proposed memory cell requires only two new clocking zones and utilizes two parallel clock zones per memory cell; permitting less CMOS circuity for clock design and denser QCA system implementations. Furthermore, read throughput is improved to one operation per clock cycle (from one read per two clock cycles). Simulations with the QCADesigner simulator are performed to verify the functionality of the proposed QCA memory cell. Baris Taskin |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2006 | Delay Insertion Method in Clock Skew SchedulingabstractThis paper describes a delay insertion method that improves the efficiency of clock skew scheduling. It is shown that reconvergent paths limit the improvement of circuit performance achievable through clock skew scheduling. A delay insertion method is proposed such that the optimal clock period achievable through clock skew scheduling is improved by mitigating the limitations caused by reconvergent paths. Experimental results demonstrate that reconvergent paths are limiting for 34% (41% for level sensitive) of the selected suite of ISCAS'89 benchmark circuits. Through the application of clock skew scheduling with delay insertion, an average improvement of 10% shorter clock periods (9% for level sensitive) is observed for ISCAS'89 benchmark circuits compared to the results of conventional clock skew scheduling. Baris Taskin, Ivan S. Kourtev |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2005 | Delay insertion method in clock skew schedulingabstractThis paper describes a delay insertion method that improves the efficiency of clock skew scheduling. Clock skew scheduling is performed on synchronous circuits in order to improve the performance of a circuit; most often by permitting the circuit to operate at a lower clock period or by increasing the tolerance of the circuit against secondary order effects and process parameter variations. With clock skew scheduling, the original circuit topology is preserved while the clock distribution network is modified to satisfy an optimal clock schedule (set of clock signal arrival delays). The work presented here studies a circuit modification technique requiring systematic delay insertion within the circuit logic (delay insertion method) in order to improve the minimum clock period achieved through clock skew scheduling. The proposed delay insertion method is defined and demonstrated on both edge-triggered and level-sensitive synchronous circuits leading to average clock period improvements of 9% and 10%, respectively, over standard clock skew scheduling algorithms. Overall, the clock period improvements over zero clock skew, flip-flop based circuits are improved to 34% on average, both for the edge-triggered and level-sensitive designs of ISCAS'89 benchmark circuits. Baris Taskin, Ivan S. Kourtev |
ISPD | 1 |
| 2004 | Linearization of the timing analysis and optimization of level-sensitive digital synchronous circuitsabstractThis paper describes a linear programming (LP) problem formulation applicable to the static-timing analysis of large scale synchronous circuits with level-sensitive latches. Specifically, an LP formulation for the clock period minimization problem is developed. In order to minimize the clock period of level-sensitive circuits, the simultaneous effects of time borrowing and nonzero clock skew scheduling are considered. The clock period minimization problem is formulated for both single-phase and multi-phase clocking schemes. The ISCAS'89 benchmark circuits are used to derive experimental results. LP minimization problems for these benchmark circuits are generated using the modified big M (MBM) method and the generated problems are solved using the industrial LP solver CPLEX . The experimental results demonstrate up to 63% improvements in minimum clock period compared to flip-flop based circuits with zero clock skew. Baris Taskin, Ivan S. Kourtev |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |