VLDB 2026 Research / reviewers in the wild / expert
Seda Ogrenci Memik
dblp:m/SedaOgrenciMemik · also Seda Ogrenci, Seda Ogrenci-Memik
· DBLP profile ↗
94ranked-venue papers
8as first author
8since 2021 · last 2026
0000-0001-8327-9585ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 92 · 8 first-author · 8 since 2021Software engineering, systems software and programming languages · 10Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Introduction to the Special Issue on Open Source Tools for Reconfigurable Devices and Systems
Seda Ogrenci Memik, Stephen Neuendorffer, Tobias Grosser, Fredrik Kjolstad, Marco Santambrogio, Evangeline F. Y. Young |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2026 | hls4ml: A Flexible, Open Source Platform for Deep Learning Acceleration on Reconfigurable HardwareabstractWe present hls4ml , a free and open source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this article, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results. Jan-Frederik Schulte, Benjamin Ramhorst, Jovan Mitrevski, Nicolò Ghielmetti, Enrico Lupi, Dimitrios Danopoulos, Vladimir Loncar, Javier M. Duarte, David Burnette, Lauri Laatu, Stylianos Tzelepis, Konstantinos Axiotis, Quentin Berthet, Haoyan Wang, Suleyman Demirsoy, Marco Colombo, Thea Aarrestad, Sioni Summers, Maurizio Pierini, Giuseppe Di Guglielmo, Jennifer Ngadiuba, Javier Campos, Benjamin Hawks, Abhijith Gandrakota, Farah Fahim, George A. Constantinides, Zhiqiang Que, Wayne Luk, Alexander D. Tapper, Duc Hoang, Noah Paladino, Philip C. Harris, Bo-Cheng Lai, Manuel Valentin, Ryan Forelli, Seda Ogrenci Memik, Lino Gerlach, Rian Brooks Flynn, Mia Liu, Daniel Diaz 0003, Elham E Khoda, Melissa Quinnan, Russell Solares, Santosh Parajuli, Mark S. Neubauer, Christian Herwig, Ho Fung Tsoi, Dylan S. Rankin, Shih-Chieh Hsu, Scott Hauck |
ACM Trans. Reconfigurable Technol. Syst. | 39 |
| 2026 | A Physics-Informed Neural Network Surrogate for Runtime PDN and Dynamic Droop Prediction in 2.5-D Chiplet Integration
Xi Chen 0099, Yuhao Ju, Seda Ogrenci Memik, Jie Gu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | Toward Reconfigurable In-Pixel Computing: A Fault-Tolerant Design Flow for Machine Learning Accelerators: (Invited Paper)abstractIn-pixel computing has emerged as a promising paradigm for edge machine learning, enabling computations directly within sensors to achieve low latency and real-time performance. However, current hardware designs generated by high-level synthesis (HLS) tools often lack flexibility and adaptability to new tasks. Additionally, robustness against single-event upsets (SEUs) in harsh radiation environments remains a critical requirement. To overcome these limitations, we propose a fault-tolerant framework that evaluates bit-level susceptibility of neural network weights, selectively applying Triple Modular Redundancy (TMR) to critical components. Our methodology integrates custom memory modules into the HLS workflow, enabling post-deployment reconfiguration for task updates and adaptive fault protection. Moreover, we introduce a TMR-aware back-end synthesis flow that dynamically allocates redundancy according to user-defined area and power constraints, achieving scalable and efficient hardware implementations. Houxuan Guo, Manuel Blanco Valentin, Xiuyuan He, Seda Ogrenci Memik |
FCCM | 4 |
| 2025 | SPRING: Systematic Profiling of Randomly Interconnected Neural Networks Generated by HLSabstractProfiling is important for performance optimization by providing real-time observations and measurements of important parameters of hardware execution. Existing profiling tools for High-Level Synthesis (HLS) IPs running on FPGAs are far less mature compared with those developed for fixed CPU and GPU architectures and they still lag behind mainly due to their dynamic architecture. This limitation is reflected in the typical approach of extracting monitoring signals off of an FPGA device individually from dedicated ports, using one BRAM per signal for temporary information storage, or embedding vendor specific primitives to manually analyze the waveform. In this paper, we propose a systematic profiling method tailored to the dynamic nature of FPGA systems, particularly suitable for streaming accelerators. Instead of relying on signal extraction, the proposed profiling stream flows alongside the actual data, dynamically splitting and merging in synchrony with the data stream, and is ultimately directed to the processing system (PS) side. We conducted a preliminary evaluation of this method on randomly interconnected neural networks (RINNs) using the FIFO fullness metric, with co-simulation results for validation. Seda Ogrenci Memik |
VTS | 2 |
| 2024 | A High Level Synthesis Methodology for Dynamic Monitoring of FPGA ML AcceleratorsabstractIn this paper, we present concepts towards a HLS-driven dynamic monitoring and debugging framework. Traditionally, in-situ debugging and dynamic monitoring is accessible during the early design stages through costly co-simulation cycles and through invasive tools and interfaces. We propose a methodology where dynamic monitoring is embedded into the high level synthesis description of machine learning (ML) accelerators within the open source hls4ml tool. We discuss the usage of the framework for monitoring FIFO channel utilization, which is a critical structure utilized to implement streaming based ML accelerators on FPGAs. Ryan F. Forelli, Seda Ogrenci Memik, Joshua Agar |
VTS | 3 |
| 2023 | A Cryogenic Readout IC with 100 KSPS in-Pixel ADC for Skipper CCD-in-CMOS SensorsabstractThe Skipper CCD-in-CMOS Parallel Read-Out Circuit (SPROCKET) is a mixed-signal front-end design for the readout of Skipper CCD-in-CMOS image sensors. SPROCKET is fabricated in a 65 nm CMOS process and each pixel occupies a$50\ \mu \mathrm{m}\times 50\ \mu \mathrm{m}$footprint. SPROCKET is intended to be heterogeneously integrated with a Skipper-in-CMOS sensor array, such that one readout pixel is connected to a multiplexed array of nine Skipper-in-CMOS pixels to enable massively parallel readout. The front-end includes a variable gain preamplifier, a correlated double sampling circuit, and a 10-bit serial successive approximation register (SAR) ADC. The circuit achieves a sample rate of 100 ksps with$0.48\ \mathrm{e}_{\text{rms}}^{-}$equivalent noise at the input to the ADC. SPROCKET achieves a maximum dynamic range of$9,000\ e^{-}$at the lowest gain setting (or$900\ e^{-}$at the lowest noise setting). The circuit operates at 100 Kelvin with a power consumption of$40\ \mu W$per pixel. A SPROCKET test chip was submitted in September 2022, and test results will be presented at the conference. Adam Quinn, Manuel Blanco Valentin, Tom Zimmerman 0002, Davide Braga, Seda Ogrenci Memik, Farah Fahim |
ISCAS | 5 |
| 2021 | Thermal Management for FPGA Nodes in HPC SystemsabstractThe integration of FPGAs into large-scale computing systems is gaining attention. In these systems, real-time data handling for networking, tasks for scientific computing, and machine learning can be executed with customized datapaths on reconfigurable fabric within heterogeneous compute nodes. At the same time, thermal management, particularly battling the cooling cost and guaranteeing the reliability, is a continuing concern. The introduction of new heterogeneous components into HPC nodes only adds further complexities to thermal modeling and management. The thermal behavior of multi-FPGA systems deployed within large compute clusters is less explored. In this article, we first show that the thermal behaviors of different FPGAs of the same generation can vary due to their physical locations in a rack and process variation, even though they are running the same tasks. We present a machine learning–based model to capture the thermal behavior of each individual FPGA in the cluster. We then propose two thermal management strategies guided by our thermal model. First, we mitigate thermal variation and hotspots across the cluster by proactive thermal-aware task placement. Under the tested system and benchmarks, we achieve up to 26.4° C and on average 13.3° C system temperature reduction with no performance penalty. Second, we utilize this thermal model to guide HLS parameter tuning at the task design stage to achieve improved thermal response after deployment. Yingyi Luo, Joshua Zhao 0001, Arnav Aggarwal, Seda Ogrenci Memik, Kazutomo Yoshii |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2020 | A Low-Power, High-Speed Readout for Pixel Detectors Based on an Arbitration TreeabstractIn this article, a low-power, high-speed arbitration tree for pixel detector readout is presented. The synchronized, binary tree priority encoder establishes a position-dependent priority list at the start of every time frame. Pixels that indicate the presence of data for readout are sequentially granted access to a shared bus for data transfer to the periphery, without the use of an additional global strobe signal. It can be used for either full frame imaging or zero-suppressed readout, in which case it can simultaneously generate the pixel address. To increase the readout frame rate, the pixel array is subdivided into two halves, which allow interleaved latching of data at the output serializer. The design was implemented in a 65-nm LP-CMOS process for the readout of a 64×64 pixel array. Measurement results demonstrate a deadtimeless, full frame imaging rate of ~50 kfps, achieved with a dedicated output for every (32×32) 1024 pixels and for a pixel data packet of 11 bits, with no bit errors detected over 1000 frames. The measured energy per bit is 0.94 pJ. Farah Fahim, Siddhartha Joshi, Seda Ogrenci Memik, Hooman Mohseni |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Minimizing Thermal Variation in Heterogeneous HPC Systems with FPGA NodesabstractThe presence of FPGAs in data centers has been growing due to their superior performance as accelerators. Thermal management, particularly battling the cooling cost in these high performance systems, is a primary concern. Introduction of new heterogeneous components only adds further complexities to thermal modeling and management. The thermal behavior of multi-FPGA systems deployed within large compute clusters is little explored. In this paper, we first show that the thermal behaviors of different FPGAs of the same generation can vary due to their physical locations in a rack and process variation, even though they are running the same tasks. We present a machine learning based model to capture the thermal behavior of a multi-node FPGA cluster. We then propose to mitigate thermal variation and hotspots across the cluster by proactive task placement guided by our thermal model. Our experiments show that through proper placement of tasks on the multi-FPGA system, we can reduce the peak temperature by up to 11.50°C with no impact on performance. Yingyi Luo, Seda Ogrenci Memik, Gokhan Memik, Kazutomo Yoshii, Pete Beckman |
ICCD | 3 |
| 2018 | Machine Learning-Based Temperature Prediction for Runtime Thermal Management Across System ComponentsabstractElevated temperatures limit the peak performance of systems because of frequent interventions by thermal throttling. Non-uniform thermal states across system nodes also cause performance variation within seemingly equivalent nodes leading to significant degradation of overall performance. In this paper we present a framework for creating a lightweight thermal prediction system suitable for run-time management decisions. We pursue two avenues to explore optimized lightweight thermal predictors. First, we use feature selection algorithms to improve the performance of previously designed machine learning methods. Second, we develop alternative methods using neural network and linear regression-based methods to perform a comprehensive comparative study of prediction methods. We show that our optimized models achieve improved performance with better prediction accuracy and lower overhead as compared with the Gaussian process model proposed previously. Specifically we present a reduced version of the Gaussian process model, a neural network-based model, and a linear regression-based model. Using the optimization methods, we are able to reduce the average prediction errors in the Gaussian process from 4.2°C to 2.9°C. We also show that the newly developed models using neural network and Lasso linear regression have average prediction errors of 2.9°C and 3.8°C respectively. The prediction overheads are 0.22, 0.097, and 0.026 ms per prediction for reduced Gaussian process, neural network, and Lasso linear regression models, respectively, compared with 0.57 ms per prediction for the previous Gaussian process model. We have implemented our proposed thermal prediction models on a two-node system configuration to help identify the optimal task placement. The task placement identified by the models reduces the average system temperature by up to 11.9°C without any performance degradation. Furthermore, these models respectively achieve 75, 82.5, and 74.17 percent success rates in correctly pointing to those task placements with better thermal response, compared with 72.5 percent success for the original model in achieving the same objective. Finally, we extended our analysis to a 16-node system and we were able to train models and execute them in real time to guide task migration and achieve on average 17 percent reduction in the overall system cooling power. Akhil Guliani, Seda Ogrenci Memik, Gokhan Memik, Kazutomo Yoshii, Rajesh Sankaran, Pete Beckman |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Adaptive Thermal Management for 3D ICs with Stacked DRAM CachesabstractWe describe an adaptive thermal management system for 3D-ICs with stacked DRAM cache memories. We present a detailed analysis of the impact of 3D-IC hotspot aggregation on the refresh behavior of the stacked DRAM-based L3 cache. We also present the consequence of the refresh variation on the overall system performance and cache energy consumption. Our analysis demonstrates that memory intensive applications are influenced more strongly by the DRAM refresh variation. We show that there is an optimal operating point where, with a reduced clock frequency, processor cores would actually recover any performance loss induced by DRAM refresh and at the same time the cache energy consumption could be optimized. We propose a low overhead run-time method that can identify the best CPU frequency modulation factor to cool the system to minimize accelerated refresh rates in the DRAM caches. Our system can provide a customizable trade-off between performance of the processor and energy savings of the memory. Akhil Guliani, Seda Ogrenci Memik |
DAC | 4 |
| 2017 | Evaluating irregular memory access on OpenCL FPGA platforms: A case study with XSBenchabstractFPGAs are becoming an attractive choice as a heterogeneous computing unit for scientific computing because FPGA vendors are adding floating-point-optimized architectures to their product lines. Additionally, high-level synthesis (HLS) tools such as Altera OpenCL SDK are emerging, which could potentially break the FPGA programming wall and provide a streamlined flow for domain experts in scientific computing. On the other hand, providing high performance in the presence of irregular memory access patterns to off-chip memory remains a challenge for the automated synthesis flows. In this paper, we study the performance/energy characteristics of OpenCL-generated FPGA designs on irregular memory access patterns, targeting XSBench, a memory-intensive Monte Carlo simulation code, as a case study. To complete our study, we implement XSBench in OpenCL and study optimization strategies for FPGAs. We observe that our OpenCL implantation of XSBench achieves 50 % higher energy efficiency on an Intel Arria10-based FPGA platform than that on an Intel Xeon 8-core CPU while trading off 35 % of performance. Yingyi Luo, Xianshan Wen, Kazutomo Yoshii, Seda Ogrenci Memik, Gokhan Memik, Hal Finkel, Franck Cappello |
FPL | 4 |
| 2017 | Cell-to-array thermal-aware analysis of stacked RRAMabstractThe crossbar resistive random access memory (RRAM) has been studied extensively due to its low-power, low-cost, high density and nonvolatile characteristics. However, the dependence of RRAM performance parameters on temperature, from cell to array level, is less explored. Particularly, thin-film based RRAM that is integrated into 3D ICs is subject to severe thermal conditions. Hence, temperature dependence of RRAM behavior needs to be well understood in order to construct the most effective thermal management strategy for these systems. In this paper, a detailed RRAM device thermal model is proposed. Our experiments show that the temperature of the surrounding environment critically impacts the RRAM readout margin, which in turn, jeopardizes its promised advantages for high-density stacked integration. Our thermal model can be utilized at the architectural level to predict the latency variation in the RRAM memory as a function of the thermal environment and drive thermal management policies for the 3D IC. Yingyi Luo, Seda Ogrenci Memik, Jie Gu 0001 |
ISCAS | 2 |
| 2017 | End-to-End Analysis of Integration for Thermocouple-Based Sensors Into 3-D ICsabstractSolutions to the integration challenges of a new thermal sensor technology into 3-D integrated circuits (ICs) will be discussed in this paper. Our proposed architecture uses bimetallic thin-film thermocouples, which are thermally linked to points of measurement throughout the 3-D stack with dedicated vias. These vias will be similar to thermal through-silicon vias (TSVs) in structure, yet different in functionality. We propose a low-overhead design methodology by linking the sensor placement task with the existing thermal TSV planning phase for 3-D ICs. A fraction of thermal TSV resources is decoupled from their original use and repurposed for the temperature sensing infrastructure. Tradeoffs concerning the reduction of the thermal TSV resources are investigated. Furthermore, we present an end-to-end system, including the physical realization of the sensor network as well as its analog interface circuitry with the sensor data sampling unit. We demonstrate the operation and correctness of this interface with transistor-level simulations. Next, through thermal modeling and simulation using a state-of-the-art tool (FloTHERM), we demonstrate that we can achieve high accuracy (1 °C error) in temperature tracking while still maintaining the effectiveness of the thermal TSVs in heat management (conforming to a peak temperature constraint of 95 °C). Siddhartha Joshi, Seda Ogrenci Memik |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Lazy Pipelines: Enhancing quality in approximate computing
Georgios Tziantzioulis, Ali Murat Gok, S. M. Faisal, Nikos Hardavellas, Seda Ogrenci Memik, Srinivasan Parthasarathy 0001 |
DATE | 5 |
| 2016 | A Partial Carry-Save On-the-Fly Correction Multispeculative MultiplierabstractFunctional Units that are designed to receive inputs and produce outputs using a non-redundant format typically exhibit an inferior performance. In order to overcome this limitation, the carry-save and partial carry-save formats have been proposed. Both approaches are very suitable when implementing addition trees. Nevertheless, if there are multiplications in the datapath, the inputs to the multiplier must be reduced to a non-redundant form, to avoid applying the distributive property. In this paper we present a multiplier able to receive two numbers in partial carry-save format, and produce a result in partial carry-save format as well. This is done by modifying the Booth encoder and leveraging the generate and propagate group signals that are available because of the partial carry-save format. Hence, this can allow to fully implement datapaths without additional penalty cycles due to reductions to non-redundant forms. Experiments show that our proposed multiplier has 15 percent shorter delay with respect to a conventional Booth radix-4 multiplier. Moreover, when combining it with partial carry-save adders it is possible to reduce 36 percent execution time on average for several benchmarks, achieving a 32.7 percent reduction in the Energy Delay Product at the same time. Alberto A. Del Barrio, Román Hermida, Seda Ogrenci Memik |
IEEE Trans. Computers | 3 |
| 2015 | Edge importance identification for energy efficient graph processingabstractModern graphs are large, often containing billions of nodes and edges that demand huge amount of processing for analysis purposes. The algorithms processing these graphs often run for long time and consume substantial amount of energy. However, not all edges in the graphs are equally important. Some edges play critical role in maintaining the community and other interesting structures in the graph, while the rest are less important for analysis. Identifying edges as important and unimportant allows one to apply elastic fidelity computing when processing edges of low importance, hence saving significant amount of energy while processing large graphs. In this paper we propose a novel technique for identifying important edges in a graph using a fast method that exploits locality sensitive hashing. We then propose a framework for energy-efficient computing that applies elastic fidelity computing when processing edges of low importance and applies full fidelity computing when processing important edges. This allows the framework to deliver good results while saving energy when processing a large number of low-importance edges. Our proposed technique reduces the power consumption by 3-30% while still producing results that are within acceptable range of the full-accuracy results. S. M. Faisal, Georgios Tziantzioulis, Ali Murat Gok, Nikos Hardavellas, Seda Ogrenci Memik, Srinivasan Parthasarathy 0001 |
IEEE BigData | 5 |
| 2015 | b-HiVE: a bit-level history-based error model with value correlation for voltage-scaled integer and floating point unitsabstractExisting timing error models for voltage-scaled functional units ignore the effect of history and correlation among outputs, and the variation in the error behavior at different bit locations. We propose b-HiVE, a model for voltage-scaling-induced timing errors that incorporates these attributes and demonstrates their impact on the overall model accuracy. On average across several operations, b-HiVE's estimation is within 1--3% of comprehensive analog simulations, which corresponds to 5--17x higher accuracy (6--10x on average) than error models currently used in approximate computing research. To the best of our knowledge, we present the first bit-level error models of arithmetic units, and the first error models for voltage scaling of bitwise logic operations and floating-point units. Georgios Tziantzioulis, Ali Murat Gok, S. M. Faisal, Nikos Hardavellas, Seda Ogrenci Memik, Srinivasan Parthasarathy 0001 |
DAC | 5 |
| 2015 | User-specific skin temperature-aware DVFS for smartphones
Begum Egilmez, Gokhan Memik, Seda Ogrenci Memik, Oguz Ergin |
DATE | 3 |
| 2015 | A methodology for power characterization of associative memoriesabstractContent Addressable Memories (CAM) have become increasingly more important in applications requiring high speed memory search due to their inherent massively parallel processing architecture. We present a complete power analysis methodology for CAM systems to aid the exploration of their power-performance trade-offs in future systems. Our proposed methodology uses detailed transistor level circuit simulation of power behavior and a handful of input data types to simulate full chip power consumption. Furthermore, we applied our power analysis methodology on a custom designed associative memory test chip. This chip was developed by Fermilab for the purpose of developing high performance real-time pattern recognition on high volume data produced by a future large-scale scientific experiment. We applied our methodology to configure a power model for this test chip. Our model is capable of predicting the total average power within 4% of actual power measurements. Our power analysis methodology can be generalized and applied to other CAM-like memory systems and accurately characterize their power behavior. Siddhartha Joshi, Seda Ogrenci Memik, James Hoff, Sergo Jindariani, Tiehui Liu 0001, Jamieson Olsen |
ICCD | 3 |
| 2015 | Minimizing Thermal Variation Across System ComponentsabstractThermal overheating is a serious concern in modern supercomputing systems. Elevated temperature levels reduce the reliability and the lifetime of the underlying hardware and increase their power consumption. Previous studies on mitigating thermal hotspots at the hardware and run-time system levels have typically used approaches that trade off performance for reduced operating temperatures. In this paper, we first show that in a large-scale system, physical attributes cause an uneven temperature distribution. We then develop a model to characterize the thermal behaviour of a complex system using various machine learning methods. We propose to improve application placement by incorporating thermal awareness into the decision-making process. Specifically, our system predicts the thermal condition of the system based on application mapping and uses these predictions to mitigate thermal hotspots without any performance loss. We provide two versions of our prediction mechanism. On a two-node configuration, these models achieve 72.5% and 78.8% success rates in their predictions, respectively. In other words, the scheduling decisions of our models result in a task placement that has a lower maximum average temperature. Overall, the more aggressive scheme reduces the average peak temperature by up to 11.9°C (2.3°C on average) without any performance degradation. Seda Ogrenci Memik, Gokhan Memik, Kazutomo Yoshii, Rajesh Sankaran, Pete Beckman |
IPDPS | 2 |
| 2015 | On-chip integration of thermoelectric energy harvesting in 3D ICsabstractWe present a full system integration of a thermoelectric energy harvesting system as an on-chip component into a 3D IC. Our system incorporates a lithographically patterned bi-metallic thin-film thermocouple network with a switched capacitor power converter and a charge buffer capacitor to harvest thermal energy produced by temperature gradients in typical 3D IC structures. Through heat transfer and transistor-level circuit simulations we demonstrate the energy harvesting potential of our system to power a low energy circuit component. Our proposed thin film based harvester does not require package re-design, since it is integrated on-chip using low cost CMOS compatible material. We evaluated integration of our proposed system into a 3D stacking of processor cores and DRAM memory. Even when operating at a conservative thermal bound of 84°C sufficient energy is harvested to continuously sustain a low-power adder for 29,640 cycles of single bit additions or 463 cycles of 64-bit additions with 12usec charging delay. Effectively we can run the adder continuously with less than 0.80% delay between bursts of operations. Seda Ogrenci Memik, Lawrence J. Henschen |
ISCAS | 2 |
| 2013 | Multispeculative additive trees in high-level synthesisabstractMultispeculative Functional Units (MSFUs) are arithmetic functional units that operate using several predictors for the carry signal. The carry prediction helps to shorten the critical path of the functional unit. The average performance of these units is determined by the hit rate of the prediction. In spite of utilizing more than one predictor, none or only one additional cycle is enough for producing the correct result in the majority of the cases. In this paper we present multispeculation as a way of increasing the performance of tree structures with a negligible area penalty. By judiciously introducing these structures into computation trees, it will only be necessary to predict in certain selected nodes, thus minimizing the number of operations that can potentially mispredict. Hence, the average latency will be diminished and thus performance will be increased. Our experiments show that it is possible to improve on average 24% and 38% execution time, when considering logarithmic and linear modules, respectively. Alberto A. Del Barrio, Román Hermida, Seda Ogrenci Memik, Jose Manuel Mendias, María C. Molina |
DATE | 3 |
| 2013 | Exploring the energy efficiency of Multispeculative AddersabstractVariable Latency Adders are attracting strong interest for increasing performance at a low cost. However, most of the literature is focused on achieving a good area-delay tradeoff. In this paper we consider multispeculation as an alternative for designing adders with low energy consumption, while offering better performance than the corresponding non-speculative ones. Instead of introducing more logic to accelerate the computation, the adder is split into several fragments which operate in parallel, and whose carry-in signals are provided by predictor units. On the one hand, the critical path of the module is shortened, and on the other hand the frequent useless glitches produced in the carry propagation structure are diminished. Hence, this will be translated into an overall energy reduction. Several experiments have been performed with linear and logarithmic adders, and results show energy savings by up to 90% and 70%, respectively, while achieving an additional execution time decrease. Furthermore, when utilized in whole datapaths with current control techniques, it is possible to reduce execution time by 24.5% (34% best case) and energy by 32% (48% best case) on average. Alberto A. Del Barrio, Román Hermida, Seda Ogrenci Memik |
ICCD | 3 |
| 2013 | Integrating thermocouple sensors into 3D ICsabstractIn this paper, we present a novel architecture for embedding bi-metallic thermocouple based temperature sensors into 3D IC stacks. To the best of our knowledge this is the first work addressing this specific integration problem. Our architecture uses dedicated vias to thermally couple sensors in the metal layer with the hotspots to be monitored in the active layer throughout the multi-stack structures. We propose a low cost solution by leveraging a fraction of existing thermal TSVs for this purpose. Through thermal modeling and simulation using a state-of-the-art tool (FloTHERM), we demonstrate that we can achieve high accuracy (less than 1°C error) in temperature tracking while still maintaining the effectiveness of the thermal TSVs in heat management (conforming to a fixed peak temperature threshold of 95°C). Seda Ogrenci Memik |
ICCD | 3 |
| 2013 | A fragmentation aware High-Level Synthesis flow for low power heterogenous datapaths
Alberto A. Del Barrio, Seda Ogrenci Memik, María C. Molina, Jose Manuel Mendias, Román Hermida |
Integr. | 2 |
| 2013 | Theory and Analysis for Optimization of On-Chip Thermoelectric Cooling SystemsabstractWe established a novel theoretical analysis framework for optimizing the cooling system configuration of chips employing thermoelectric cooling (TEC) elements by extending the theory of inverse-positive matrices and the eigenvalue/eigenvector theory in linear algebra. In this brief, we present a new theorem and its formal proof, which is the key enabler to achieving a provably optimal solution for configuring bias current levels of TEC devices. Jieyi Long, Seda Ogrenci Memik, Semail Ülgen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2012 | Multispeculative Addition Applied to Datapath SynthesisabstractAddition is the key arithmetic operation in most digital circuits and processors. Therefore, their performance and other parameters, such as area and power consumption, are highly dependent on the adders' features. In this paper, we present multispeculation as a way of increasing adders' performance with a low area penalty. In our proposed design, dividing an adder into several fragments and predicting the carry-in of each fragment enables computing every addition in two very short cycles at the most, with 99% or higher probability. Furthermore, based on multispeculation principles, we propose a new strategy for implementing addition chains and hiding most of the penalty cycles due to mispredictions, while keeping at the same time the resource sharing capabilities that are sought in high-level synthesis. Our results show that it is possible to build linear and logarithmic adders more than$4.7\times$and$1.7\times$faster than the nonspeculative case, respectively. Moreover, this is achieved with a low area penalty (38% for linear adders) or even an area reduction (${-}8\%$for logarithmic adders). Finally, applying multispeculation principles to signal processing benchmarks that use addition chains will result in 25% execution time reduction, with an additional 3% decrease in datapath area with respect to implementations with logarithmic fast adders. Alberto A. Del Barrio, Román Hermida, Seda Ogrenci Memik, Jose Manuel Mendias, María C. Molina |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2011 | Power optimization in heterogenous datapathsabstractHeterogenous datapaths maximize the utilization of functional units (FUs) by customizing their widths individually through fragmentation of wide operands. In comparison, slices in large functional units in a homogenous datapath could be spending many cycles not performing actual useful work. Various fragmentation techniques demonstrated benefits in minimizing the total functional unit area. Upon a closer look at fragmentation techniques, we observe that the area savings achieved by heterogenous datapaths can be traded-off for power optimization. Our specific approach is to introduce choices for functional units with power/area trade-offs for different fragmentation and allocation choices, for reducing power consumption while satisfying the area constraint imposed on the heterogenous datapath. As low power FUs in literature produce an area penalty, a methodology must be developed in order to introduce them in the HLS flow while complying with the area constraint. We propose an allocation and module selection algorithms that pursue a trade-off between area and power consumption for fragmented datapaths under a total area constraint. Results show that it is possible to reduce power by 37% on average (49% in the best case). Moreover latency and cycle time will be equal or nearly the same as in the baseline case, which will lead to an energy reduction, too. Alberto A. Del Barrio, Seda Ogrenci Memik, María C. Molina, Jose Manuel Mendias, Román Hermida |
DATE | 2 |
| 2011 | Hardware/software techniques for DRAM thermal managementabstractThe performance of the main memory is an important factor on overall system performance. To improve DRAM performance, designers have been increasing chip densities and the number of memory modules. However, these approaches increase power consumption and operating temperatures: temperatures in existing DRAM modules can rise to over 95°C. Another important property of DRAM temperature is the large variation in DRAM chip temperatures. In this paper, we present our analysis collected from measurements on a real system indicating that temperatures across DRAM chips can vary by over 10°C. This work aims to minimize this variation as well as the peak DRAM temperature. We first develop a thermal model to estimate the temperature of DRAM chips and validate this model against real temperature measurements. We then propose three hardware and software schemes to reduce peak temperatures. The first technique introduces a new cache line replacement policy that reduces the number of accesses to the overheating DRAM chips. The second technique utilizes a Memory Write Buffer to improve the access efficiency of the overheated chips. The third scheme intelligently allocates pages to relatively cooler ranks of the DIMM. Our experiments show that in a high performance memory system, our schemes reduce the peak DRAM chip temperature by as much as 8.39°C over 10 workloads (5.36°C on average). Our schemes also improve performance mainly due to reduction in thermal emergencies: for a baseline system with memory bandwidth throttling scheme, the IPC is improved by as much as 15.8% (4.1% on average). Brian Leung, Alexander Neckar, Seda Ogrenci Memik, Gokhan Memik, Nikos Hardavellas |
HPCA | 4 |
| 2011 | A Comprehensive Tapered buffer optimization algorithm for unified design metricsabstractTapered buffers are widely used in CMOS integrated circuits to drive large capacitive loads. During the design of a tapered buffer, there are several design objectives to consider including delay, area, and power consumption. Existing methods produce suboptimal solutions considering multiple metrics, largely because they decouple different metrics during the design phase, which restricts the solution space for the combined metric. In this paper, a new algorithm is proposed to derive the optimal solution for a unified metric through comprehensive exploration of the solution space. Compared with existing methods, our method yields as high as 18.8% (9.0% on average) improvement in a unified design metric optimizing area, delay, and power simultaneously. The proposed algorithm also reduces buffer power consumption under delay constraints. The power reduction over existing alternatives is as much as 48.1% (28.7% on average). Seda Ogrenci Memik, Yehea I. Ismail |
ISCAS | 2 |
| 2011 | A Distributed Controller for Managing Speculative Functional Units in High Level SynthesisabstractSpeculative functional units (SFUs) are arithmetic functional units that operate using a predictor for the carry signal. The carry prediction helps to shorten the critical path of the functional unit. The average case performance of these units is determined by the hit rate of the prediction. In case of mispredictions, the SFUs need to be coordinated by the datapath control mechanism to perform corrections and to maintain the datapath in the correct state. Devising a control mechanism for correcting mispredictions without adversely impacting overall performance is the most important challenge. In this paper, we present techniques for designing a datapath controller for seamless deployment of SFUs in high level synthesis. We have developed two techniques based on two main control paradigms: centralized and distributed control. The centralized approach stops the execution of the entire datapath for each misprediction and resumes execution once the correct value of the carry is known. The distributed approach decouples the functional unit suffering from the misprediction from the rest of the datapath. Hence, it allows the remainder of the functional units to carry on execution and be at different scheduling states at different times. We tested datapaths utilizing both linear structures and logarithmic structures for speculative arithmetic functional units. Our results show that it is possible to reduce execution time by as much as 38% (33% on average) for linear structures and by as much as 37.2% (25% on average) for logarithmic structures. Alberto A. Del Barrio, Seda Ogrenci Memik, María C. Molina, Jose Manuel Mendias, Román Hermida |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | A framework for optimizing thermoelectric active cooling systemsabstractThin-film thermoelectric cooling is a promising technology for mitigating heat dissipation in high performance chips. In this paper, we present an optimization framework for an active cooling system that is comprised of an array of thin-film thermoelectric coolers. We observe a set of constraints of the cooling system design. Firstly, integrating an excessive amount of coolers increases the chip package cost. Moreover, thermoelectric coolers are active devices, which dissipate heat in the chip package when they are in operation. Hence, setting the supply current level to operate the cooler improperly can actually lead to overheating of the chip package. Besides, the supply current needs to be delivered to the integrated cooler devices via dedicated pins. However, extra pins available on high-performance chip packages are limited. Observing these constraints, we propose an optimization framework for configuring the active cooling system, which minimizes the maximum silicon temperature. This includes determining the amount of coolers to deploy and their locations, the mapping of supply pins to the coolers, and determining the current levels of each pin. We propose algorithms to tackle the optimal configuration problem. We found that only a small portion of the silicon die needs to be covered by TEC devices (18% on average). Our experiments show that our algorithms are able to reduce the temperatures of the hot spots by as much as 10.6 °C (compared to the cases without integrated thermoelectric coolers). The average temperature reduction is 8.6 °C when 4 dedicated pins are available on the package. The total power consumption of the resulting active cooling system is reasonably small (~ 2 W). Our experiments also reveal that our framework maximizes the efficiency of the cooling devices. In the ideal case where hundreds of pins are available to tune the supply level of each individual cooler, the additional average reduction of the hot spot temperature is only 0.3 °C. Jieyi Long, Seda Ogrenci Memik |
DAC | 2 |
| 2010 | Using Speculative Functional Units in high level synthesisabstractSpeculative Functional Units (SFUs) enable a new execution paradigm for High Level Synthesis (HLS). SFUs are arithmetic functional units that operate using a predictor for the carry signal, which reduces the critical path delay. The performance of these units is determined by the success in the prediction of the carry value, i.e. the hit rate of the prediction. Hence SFUs reduce critical path at a low cost, but they cannot be used in HLS with the current techniques. In order to use them, it is necessary to include hardware support to recover from mispredictions of the carry signals. In this paper, we present techniques for designing a datapath controller for seamless deployment of SFUs in HLS. We have developed two techniques for this goal. The first approach stops the execution of the entire datapath for each misprediction and resumes execution once the correct value of the carry is known. The second approach decouples the functional unit suffering from the misprediction from the rest of the datapath. Hence, it allows the rest of the SFUs to carry on execution and be at different scheduling states at different times. Experiments show that it is possible to reduce execution time by as much as 38% and by 33% on average. Alberto A. Del Barrio, María C. Molina, Jose Manuel Mendias, Román Hermida, Seda Ogrenci Memik |
DATE | 5 |
| 2010 | Optimization of the bias current network for accurate on-chip thermal monitoringabstractMicroprocessor chips employ increasingly larger number of thermal sensing devices. These devices are networked by an underlying infrastructure, which provides bias currents to sensing devices and collects measurements. In this work, we address the optimization of the bias current distribution network utilized by the sensing devices. We show that the choice between two fundamental topologies (the 2-wire and the 4-wire measurement) for this network has a non-negligible impact on the precision of the monitoring system. We also show that the 4-wire measurement principle supports the remote sensing technique better. However, it requires more routing resources. We thus propose a novel routing algorithm to minimize its routing cost. We also present a detailed evaluation of the quality of the resulting system in presence of process and thermal variations. Our Monte Carlo simulations using the IBM 10SF 65nm SPICE models show that the monitoring accuracies can be as high as 0.6°C under considerable amount of process and temperature variation. Moreover, by adopting a customized routing approach for the current mirror network, the total wire length of the bias current network can be reduced by as much as 42.74% and by 27.65% on average. Jieyi Long, Seda Ogrenci Memik |
DATE | 2 |
| 2010 | Inversed Temperature Dependence aware clock skew scheduling for sequential circuitsabstractWe present an Inversed Temperature Dependence (ITD) aware clock skew scheduling framework. Specifically, we demonstrate how our framework can assist dual-Vthassignment in preventing timing violations arising due to ITD effect. We formulate the ITD aware synthesis problem and prove that it is NP-Hard. Then, we propose an algorithm for synergistic temperature aware clock skew scheduling and dual-Vthassignment. Experiments on ISCAS89 benchmarks reveal that several circuits synthesized by the traditional high-temperature corner based flow with a commercial tool exhibit timing violations in the low temperature range while all circuits generated using our methodology for the same timing constraints have guaranteed timing. Jieyi Long, Seda Ogrenci Memik |
DATE | 2 |
| 2010 | Optimization of an on-chip active cooling system based on thin-film thermoelectric coolersabstractIn this paper, we explore the design and optimization of an on-chip active cooling system based on thin-film thermoelectric coolers (TEC). We start our investigation by establishing the compact thermal model for the chip package with integrated thin-film TEC devices. We observe that deploying an excessive number of TEC devices and/or providing the TEC devices with an improper supply current might adversely result in the overheating of the chip, rendering the cooling system ineffective. A large amount of supply current could even cause the thermal runaway of the system. Motivated by this observation, we formulate the deployment of the integrated TEC devices and their supply current setting as a system-level design problem. We propose a greedy algorithm to determine the deployment of TEC devices and a convex programming based scheme for setting the supply current levels. Leveraging the theory of inverse-positive matrix, we provide an optimality condition for the current setting algorithm. We have tested our algorithms on various benchmarks. We observe that our algorithms are able to determine the proper deployment and supply current level of the TEC devices which reduces the temperatures of the hot spots by as much as 7.5 °C compared to the cases without integrated TEC devices. Jieyi Long, Seda Ogrenci Memik, Matthew Grayson |
DATE | 2 |
| 2010 | An Interior Point Optimization Solver for Real Time Inter-frame Collision Detection: Exploring Resource-Accuracy-Platform TradeoffsabstractWe present and compare implementations of an affine interior-point algorithm for real-time collision detection on a GPGPU and an FPGA. This particular interior-point algorithm is distinguished from other collision detection methods by its ability to perform detection between pairs of objects undergoing fast rotational and translational movement. This enables inter-frame collision detection, i.e. collision that might occur during the transition from one frame to another. In our design for the FPGA, we implemented the algorithm both in single-precision floating point and 32-bit fixed point and analyzed the trade-off between resource usage, data accuracy/precision, and system efficiency. Then, we compare them to a floating point implementation on a GPGPU using CUDA. With an object resolution of 45 vertices (45 vertices representing each polyhedral object), our FPGA implementation processes 1562 frames/sec for floating point and 1350 frames/second for fixed point and offers an 11× speedup over the GPGPU implementation. With object resolutions greater than 242 vertices, our GPGPU implementation outperforms our FPGA implementations. Brian Leung, Chih-Hung Wu, Seda Ogrenci Memik, Sanjay Mehrotra |
FPL | 3 |
| 2010 | Placement and Floorplanning in Dynamically Reconfigurable FPGAsabstractThe aim of this article is to describe a complete partitioning and floorplanning algorithm tailored for reconfigurable architectures deployable on FPGAs and considering communication infrastructure feasibility. This article proposes a novel approach for resource- and reconfiguration- aware floorplanning. Different from existing approaches, our floorplanning algorithm takes specific physical constraints such as resource distribution and the granularity of reconfiguration possible for a given FPGA device into account. Due to the introduction of constraints typical of other problems like partitioning and placement, the proposed approach is named floorplacer in order to underline the great differences with respect to traditional floorplanners. These physical constraints are typically considered at the later placement stage. Different aspects of the problems have been described, focusing particularly on the FPGAs resource heterogeneity and the temporal dimension typical of reconfigurable systems. Once the problem is introduced a comparison among related works has been provided and their limits have been pointed out. Experimental results proved the validity of the proposed approach. Alessio Montone, Marco D. Santambrogio, Donatella Sciuto, Seda Ogrenci Memik |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2010 | An Approach for Adaptive DRAM Temperature and Power ManagementabstractHigh-performance DRAMs are providing increasing memory access bandwidth to processors, which is leading to high power consumption and operating temperature in DRAM chips. In this paper, we propose a customized low-power technique for high-performance DRAM systems to improve DRAM page hit rate by buffering write operations that may incur page misses. This approach reduces DRAM system power consumption and temperature without any performance penalty. We combine the throughput-aware page-hit-aware write buffer (TAP) with low-power-state-based techniques for further power and temperature reduction, namely, TAP-low. Our experiments show that a system with TAP-low could reduce the total DRAM power consumption by up to 68.6% (19.9% on average). The steady-state temperature can be reduced by as much as 7.84$^{\circ}\hbox{C}$and 2.55$^{\circ}\hbox{C}$on average across eight representative workloads. Seda Ogrenci Memik, Gokhan Memik |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | SACTA: A Self-Adjusting Clock Tree Architecture for Adapting to Thermal-Induced Delay VariationabstractAggressive technology scaling down and low-power design techniques lead to uneven distributed power density, which translates into heat flow in the chips, causing significant temperature variations in both spatial and temporal terms. In order to mitigate the negative impacts of temperature variations on circuit timing, we propose SACTA, a self-adjusting clock tree architecture, which performs temperature-dependent dynamic clock skew scheduling to prevent timing violations in a pipelined circuit. The dynamic and adaptive features of SACTA are enabled by our proposed automatic temperature-adjustable skew buffers and temperature-insensitive skew buffers. These special delay elements are carefully tuned to ensure resilience of the entire circuit against temperature variation. To determine their configurations, we proposed an efficient and general clock tree design and optimization framework. Furthermore, we show that SACTA is applicable across a wide spectrum of circuits, including multi-${V}_{\rm dd}/{V}_{\rm th}$designs. Experimental results show that a pipeline supported by SACTA is able to prevent thermal-induced timing violations within a significantly larger range of operating temperatures (on average, the violation-free range can be enhanced by over 15$^{\circ}\hbox {C}$). Jieyi Long, Ja Chun Ku, Seda Ogrenci Memik, Yehea I. Ismail |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | A Fast Heuristic Algorithm for Multidomain Clock Skew SchedulingabstractIn the most general form, clock skew scheduling (CSS) generates a dedicated clock delay for each individual sequential component in the clock distribution network in order to minimize the clock period. Multidomain CSS (MDCSS) relieves this requirement. Instead, sequential components are grouped into several clusters (called clock domains), each of which has a uniform clock delay for all registers within that domain. The skew values of clock domains are provided by a set of deskew buffers with electrically programmable phase shifts and injected after the chip is manufactured. This technique is attractive since, due to process variations, it is becoming overwhelmingly difficult to create precise clock network delays for all sequential elements in a design globally. In this paper, we present a fast algorithm for determining the minimum number of clock domains to be used by MDCSS. The exact solution to this problem cannot be found within a reasonable time if the number of clock domains increases beyond three domains. We show that, even with a small-size circuit, in order to obtain the minimum clock period, more than three clock domains may be required. Therefore, a fast heuristic algorithm is needed to identify these domains. To the best of our knowledge, we present the first efficient heuristic algorithm for this problem. For large benchmark circuits, we solve the problem within 14.7 min on average (as high as 31.7 min for the worst case), while a commercial mixed-integer linear program solver cannot finish in over 5 h. Furthermore, our results show that, for 19 out of 21 small- and medium-size benchmarks, our algorithm yields the optimal solution. Min Ni, Seda Ogrenci Memik |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | FPGA Implementation of the Interior-Point Algorithm with Applications to Collision DetectionabstractThe interior-point algorithm is a powerful method for solving a linear program (LP). A variety of optimization problems can be formulated as LPs. Often times the limiting factor of deploying an algorithm to solve LPs in a high performance system is the run-time efficiency. In this paper, we present the FPGA implementation of an affine interior-point algorithm that is designed to solve LPs. Specifically, we present the application of this algorithm to solving the LP for the real-time collision detection. The most important feature that distinguishes this particular algorithm from other collision detection methods is its superior ability to perform detection between pairs of objects undergoing fast rotational and translational motions. Chih-Hung Wu, Seda Ogrenci Memik, Sanjay Mehrotra |
FCCM | 2 |
| 2008 | A power and temperature aware DRAM architectureabstractTechnological advances enable modern processors to utilize increasingly larger DRAMs with rising access frequencies. This is leading to high power consumption and operating temperature in DRAM chips. As a result, temperature management has become a real and pressing issue in high performance DRAM systems. Traditional low power techniques are not suitable for high performance DRAM systems with high bandwidth. In this paper, we propose and evaluate a customized DRAM low power technique based on Page Hit Aware Write Buffer (PHA-WB). Our proposed approach reduces DRAM system power consumption and temperature without any performance penalty. Our experiments show that a system with a 64-entry PHA-WB could reduce the total DRAM power consumption by up to 22.0% (9.6% on average). The peak and average temperature reductions are 6.1°C and 2.1°C, respectively. Seda Ogrenci Memik, Gokhan Memik |
DAC | 2 |
| 2008 | Automated design of self-adjusting pipelinesabstractWe propose a self-adjusting pipeline structure to enhance chip performance and robustness considering the effects of process variations. We achieve this by introducing delay sensors to monitor internal timing violations within a pipeline stage and variable clock skew buffers to adjust the timing of the pipeline stage based on the feedback from the delay sensors. Furthermore, we formulate the delay sensor insertion and variable clock skew configuration problem as a stochastic mixed-integer programming problem and propose a simulated-annealing based algorithm to solve it. A comparison between the designs with and without the self-adjusting enhancement reveals that, we are able to improve the average performance of a batch of chips by 9.5%. Jieyi Long, Seda Ogrenci Memik |
DAC | 2 |
| 2008 | Leakage power-aware clock skew scheduling: converting stolen time into leakage power reductionabstractClock skew scheduling has been traditionally considered as a tool for improving the clock period in a sequential circuit. Timing slack is "stolen" from fast combinational blocks to be used by slower blocks to meet a more stringent clock cycle time. Instead, we can leverage on the borrowed time to achieve leakage power reduction during gate sizing and/or dual Vth assignment. In this paper, we present the first approach to the best of our knowledge for integrating clock skew scheduling, threshold voltage assignment, and gate sizing into one optimization formulation. Over 29 circuits in the ISCAS89 benchmark suite, this integrated approach can reduce leakage power by as much as 55.83% and by 18.79% on average, compared to using combinational circuit based power optimization on each combinational block without considering clock skews. Using a 65nm dual Vth technology library, this corresponds to a 23.87% peak reduction (6.15% on average) in total power at the ambient operating temperature. The average total power reduction further increases to 9.83% if the high temperature library of the same process technology is used. Min Ni, Seda Ogrenci Memik |
DAC | 2 |
| 2008 | Towards an "early neural circuit simulator": A FPGA implementation of processing in the rat whisker systemabstractWe have constructed a FPGA-based “early neural circuit simulator” to model the first two stages of stimulus encoding and processing in the rat whisker system. Rats use tactile input from their whiskers to extract object features such as size and shape. We use the simulator to examine the plausibility of the hypothesis that neural circuits in the rat’s brain compute gradients of radial distance across the whisker array to make predictions about the environment. This prediction could be a component of a feed-forward signal that guides the navigation behavior of the rat. The use of a FPGA is highly suitable for such an application, because the computation involved in this system is a massively parallel problem. For our applications, we determined that a Cyclone II FPGA could simulate up to 14 neurons in parallel in just 265 ns achieving a 386-fold speedup over the software implementation of the same model. Brian Leung, Yan Pan 0010, Christopher L. Schroeder, Seda Ogrenci Memik, Gokhan Memik, Mitra J. Z. Hartmann |
FPL | 4 |
| 2008 | An approach for adaptive DRAM temperature and power managementabstractWith rising capacities and higher accessing frequencies, high-performance DRAMs are providing increasing memory access bandwidth to the processors. However, the increasing DRAM performance comes with the price of higher power consumption and temperature in DRAM chips. Traditional low power approaches for DRAM systems focus on utilizing low power modes, which is not always suitable for high performance systems. Existing DRAM temperature management techniques, on the other hand, utilize generic temperature management methods inherited from those applied on processor cores. These methods reduce DRAM temperature by controlling the number of DRAM accesses, similar to throttling the processor core, which incurs significant performance penalty. In this paper, we propose a customized low power technique for high performance DRAM systems, namely the Page Hit Aware Write Buffer (PHA-WB). The PHA-WB improves DRAM page hit rate by buffering write operations that may incur page misses. This approach reduces DRAM system power consumption and temperature without any performance penalty. Our proposed Throughput-Aware PHA-WB (TAP) dynamically configures the write buffer for different applications and workloads, thus achieves the best trade off between DRAM power reduction and buffer power overhead. Our experiments show that a system with TAP could reduce the total DRAM power consumption by up to 18.36% (8.64% on average). The steady-state temperature can be reduced by as much as 5.10°C and by 1.93°C on average across eight representative workloads. Seda Ogrenci Memik, Gokhan Memik |
ICS | 2 |
| 2008 | An O(nlogn) edge-based algorithm for obstacle-avoiding rectilinear steiner tree constructionabstractObstacle-avoiding Steiner tree construction is a fundamental problem in VLSI physical design. In this paper, we provide a new approach for rectilinear Steiner tree construction in the presence of obstacles. We propose a novel algorithm, which generates sparse obstacle-avoiding spanning graphs efficiently. We design a fast algorithm for the minimum terminal spanning tree construction, which is the bottleneck step of several existing approaches in terms of running time. We adopt an edge-based heuristic, which enables us to perform both local and global refinement, leading to Steiner trees with small lengths. The time complexity of our algorithm is O(nlogn). Hence, our technique is the most efficient one to the best of our knowledge. Experimental results on various benchmarks show that our algorithm achieves 25.8 times speedup on average, while the average length of the resulting obstacle-avoiding rectilinear Steiner trees is only 1.58% larger than the best existing solution Jieyi Long, Hai Zhou 0001, Seda Ogrenci Memik |
ISPD | 3 |
| 2008 | Thermal monitoring mechanisms for chip multiprocessorsabstractWith large-scale integration and increasing power densities, thermal management has become an important tool to maintain performance and reliability in modern process technologies. In the core of dynamic thermal management schemes lies accurate reading of on-die temperatures. Therefore, careful planning and embedding of thermal monitoring mechanisms into high-performance systems becomes crucial. In this paper, we propose three techniques to create sensor infrastructures for monitoring the maximum temperature on a multicore system. Initially, we extend a nonuniform sensor placement methodology proposed in the literature to handle chip multiprocessors (CMPs) and show its limitations. We then analyze a grid-based approach where the sensors are placed on a static grid covering each core and show that the sensor readings can differ from the actual maximum core temperature by as much as 12.6°C when using 16 sensors per core. Also, as large as 10.6% of the thermal emergencies are not captured using the same number of sensors. Based on this observation, we first develop an interpolation scheme, which estimates the maximum core temperature through interpolation of the readings collected at the static grid points. We show that the interpolation scheme improves the measurement accuracy and emergency coverage compared to grid-based placement when using the same number of sensors. Second, we present a dynamic scheme where only a subset of the sensor readings is collected to predict the maximum temperature of each core. Our results indicate that, we can reduce the number of active sensors by as much as 50%, while maintaining similar measurement accuracy and emergency coverage compared to the case where the entire sensor set on the grid is sampled at all times. Jieyi Long, Seda Ogrenci Memik, Gokhan Memik, Rajarshi Mukherjee |
ACM Trans. Archit. Code Optim. | 2 |
| 2008 | EBOARST: An Efficient Edge-Based Obstacle-Avoiding Rectilinear Steiner Tree Construction AlgorithmabstractObstacle-avoiding Steiner routing has arisen as a fundamental problem in the physical design of modern VLSI chips. In this paper, we present EBOARST, an efficient four-step algorithm to construct a rectilinear obstacle-avoiding Steiner tree for a given set of pins and a given set of rectilinear obstacles. Our contributions are fourfold. First, we propose a novel algorithm, which generates sparse obstacle-avoiding spanning graphs efficiently. Second, we present a fast algorithm for the minimum terminal spanning tree construction step, which dominates the running time of several existing approaches. Third, we present an edge-based heuristic, which enables us to perform both local and global refinements, leading to Steiner trees with small lengths. Finally, we discuss a refinement technique called segment translation to further enhance the quality of the trees. The time complexity of our algorithm isO(nlogn). Experimental results on various benchmarks show that our algorithm achieves 16.56 times speedup on average, while the average length of the resulting obstacle-avoiding rectilinear Steiner trees is only 0.46% larger than the best existing solution. Jieyi Long, Hai Zhou 0001, Seda Ogrenci Memik |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | Presynthesis Area Estimation of Reconfigurable Streaming AcceleratorsabstractIn this paper, we propose algorithms for presynthesis estimation of hardware cost of a streaming accelerator. Our proposed estimation method helps to accelerate the design-space-exploration phase by orders of magnitude by eliminating the need to perform logic and physical synthesis in each iteration. We present algorithms to perform early cost estimation of resources that are specific to a streaming accelerator, and we evaluate our techniques using an industrial tool flow and a set of streaming benchmarks. For the register-queue sizes, our estimations are in the range of 28%-9% of actual synthesis results on average, depending on the given resource constraints, while the datapath area estimations are within 14%. A typical estimation requires less than a minute, while generating the configuration bitstream of a streaming accelerator can take as much as 30 min according to our experiments. Considering several repetitions of the synthesis stage for the design space exploration, our estimation framework yields an order of magnitude speedup. Seda Ogrenci Memik, Nikolaos Bellas, Somsubhra Mondal |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2008 | Optimizing Thermal Sensor Allocation for MicroprocessorsabstractHigh-performance microprocessor families employ dynamic-thermal-management techniques to cope with the increasing thermal stress resulting from peaking power densities. These techniques operate on feedback generated from on-die thermal sensors. The allocation and the placement of thermal-sensing elements directly impact the effectiveness of the dynamic management mechanisms. In this paper, we propose systematic techniques for determining the optimal locations for thermal sensors to provide high-fidelity thermal monitoring of a complex microprocessor system. Our strategies can be divided into two main categories: uniform sensor allocation and nonuniform sensor allocation. In the uniform approach, the sensors are placed on a regular grid. The nonuniform allocation identifies an optimal physical location for each sensor such that the sensor's attraction toward steep thermal gradients is maximized, which can result in uneven concentrations of sensors on different locations of the chip. We also present a hybrid algorithm that shows the tradeoffs associated with number of sensors and expected accuracy. Our experimental results show that our uniform approach using interpolation can detect the chip temperature with a maximum error of 5.47degC and an average maximum error of 1.05degC . On the other hand, our nonuniform strategy is able to create a sensor distribution for a given microprocessor architecture, providing thermal measurements with a maximum error of 3.18degC and an average maximum error of 1.63degC across a wide set of applications. Seda Ogrenci Memik, Rajarshi Mukherjee, Min Ni, Jieyi Long |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2008 | A high-level clustering algorithm targeting dual Vdd FPGAsabstractRecent advanced power optimizations deployed in commercial FPGAs, laid out a roadmap towards FPGA devices that can be integrated into ultra low power systems. In this article, we present a high-level design tool to support the process of mapping an application onto a FPGA device with dual supply voltages. Our main contribution in this paper is an algorithm, which creates voltage scaling ready clusters by utilizing the timing slack available in the designs. We propose to first create clusters of CLBs within a given CLB-level netlist. This clustering algorithm intends to group chains of CLBs possessing similar amounts of timing slack along their critical path together. Once these clusters are identified, they are placed onto respective V dd partitions on the device. We have evaluated different dual V dd fabrics and the potential gain in power consumption is explored. When a subset of the logic blocks on the device can be driven by low V dd levels (either with a dedicated low V dd supply or with a programmable selection between low and high V dd levels for these blocks) this affects placement and routing. As a result the maximum frequency of the designs may be affected. In order to evaluate the overall impact of creating voltage islands, we measured the Energy-Delay Product for our benchmark designs. We observed that the Energy-Delay product can be decreased by 26.9% when the placement of the designs into different voltage levels is guided by our clustering algorithm. Rajarshi Mukherjee, Seda Ogrenci Memik, Somsubhra Mondal |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2007 | Self-heating-aware optimal wire sizing under Elmore delay model
Min Ni, Seda Ogrenci Memik |
DATE | 2 |
| 2007 | A self-adjusting clock tree architecture to cope with temperature variationsabstractEnsuring resilience against environmental variations is becoming one of the great challenges of chip design. In this paper, we propose a self adjusting clock tree architecture, SACTA, to improve chip performance and reliability in the presence of on-chip temperature variations. SACTA performs temperature dependent dynamic clock skew scheduling to prevent timing violations in a pipelined circuit. We present an automatic temperature adjustable skew buffer design, which enables the adaptive feature of SACTA. Furthermore, we propose an efficient and general optimization framework to determine the configuration of these special delay elements. Experimental results show that a pipeline supported by SACTA is able to prevent thermal induced timing violations within a significantly larger range of operating temperatures (enhancing the violation-free range by as much as 45°C). Jieyi Long, Ja Chun Ku, Seda Ogrenci Memik, Yehea I. Ismail |
ICCAD | 3 |
| 2007 | Early planning for clock skew scheduling during register bindingabstractDesign decisions made during high-level synthesis usually have great impacts on the later design stages. In this paper, We present a general framework, which plans for the clock skew-scheduling in physical design stages during register binding in high-level synthesis. Our proposed technique pursues the optimality of the native objective functions of the register binding problem. At the same time, it ensures not invalidating the subsequent clock skew scheduling for optimizing the clock period. We use the switching power as the native objective of our register binding problem. The problem is first formulated as a MILP problem. An acceleration scheme based on the concept of weakly compatible edge set (WCES) is proposed to speed up the MILP solver to obtain the optimal solution. Then, we present our heuristic algorithm to reduce the running time further. The experimental results show that on average our acceleration scheme can speed up the solver by 8.6 times, and our heuristic is 70 times faster than the solver with a 5.25% degradation of the native objective. The minimum and maximum degradation among our benchmark set are 0.82% and 12.2%, respectively. Min Ni, Seda Ogrenci Memik |
ICCAD | 2 |
| 2007 | A novel SoC design methodology combining adaptive software and reconfigurable hardwareabstractReconfigurable hardware is becoming a prominent component in a large variety of SoC designs. Reconfigurability allows for efficient hardware acceleration and virtually unlimited adaptability. On the other hand, overheads associated with reconfiguration and interfaces with the software component need to be evaluated carefully during the exploration phase. The aim of this paper is to identify the best trade-off considering application-specific features in software, which can lend itself to software-based acceleration and lead to a revision of the view that certain computationally intensive tasks can only be accelerated through hardware. In order to validate the effectiveness of our proposed techniques, we built an extensive development and experimental setup, bringing together the MLTon-based programming environment and physical mapping of the software and hardware onto a real dynamically reconfigurable SoC system. Marco D. Santambrogio, Seda Ogrenci Memik, Vincenzo Rana, Umut A. Acar, Donatella Sciuto |
ICCAD | 2 |
| 2006 | Systematic temperature sensor allocation and placement for microprocessorsabstractModern high performance processors employ advanced techniques for thermal management, which rely on accurate readings of on-die thermal sensors. As the importance of thermal effects on reliability and performance of integrated circuits increases careful planning and embedding of thermal monitoring mechanisms into these systems will be crucial. Systematic tools for analysis of thermal behavior and determination of best allocation and placement of thermal sensing elements is therefore a highly relevant problem. In this paper, we propose novel optimization techniques for determining the optimal locations and allocations for thermal sensors to provide a high fidelity thermal profile of a complex microprocessor system. Our algorithm identifies an optimal physical location for each sensor such that the sensor's the attraction towards steep thermal gradient is maximized. We also present a hybrid allocation and placement strategy showing the trade-offs associated with number of sensors used and expected accuracy. Our results show that our tool is able to create a sensor distribution for a given microprocessor architecture providing thermal measurements with maximum error of 3.18°C and average maximum error of 1.63°C across a wide set of applications. Rajarshi Mukherjee, Seda Ogrenci Memik |
DAC | 2 |
| 2006 | Pre-synthesis Queue Size Estimation of Streaming Data Flow GraphsabstractIn this paper the authors propose a pre-synthesis register queue size estimation technique for an unscheduled streaming DFG (sDFG) for pipelined synthesis. Our estimation method first designates a minimum queue size to each communication edge of the sDFG based on the ALAP value of the source node and ASAP value of the sink node of that edge. Our aim is to further refine this initial minimum queue size estimation. Our main tool is based on the likelihood estimation that the source node may actually be producing data before its ALAP time, and likewise, the sink node may actually be consuming data after its ASAP time Somsubhra Mondal, Seda Ogrenci Memik, Nikolaos Bellas |
FCCM | 2 |
| 2006 | Adaptive Metrics for System-Level Functional Partitioning
Giovanni Agosta, Marco D. Santambrogio, Seda Ogrenci Memik |
FDL | 3 |
| 2006 | Power Optimization Techniques for SRAM-Based FPGAsabstractIn this work, the authors aim to improve the power efficiency of FPGAs by proposing two power reduction techniques: the authors present a low-penalty optimization technique to reduce leakage power consumption in FPGA logic blocks by exploiting the variance in LUT utilization across different designs (Mondal and Memik, (2005)), and presents a dual-Vdd-dual-Vtrouting architecture to reduce interconnect power consumption by using two levels of Vddand Vt(Mondal and Memik, (2005)) Somsubhra Mondal, Seda Ogrenci Memik |
FPL | 2 |
| 2006 | Pre-Synthesis Area Estimation of Reconfigurable Streaming AcceleratorsabstractOne of the major challenges in automated synthesis of reconfigurable accelerators is to create efficient designs that conform to the resource capacity of the target device. This work concerns estimation of the hardware cost before actually attempting the synthesis of streaming accelerators on reconfigurable platforms. Specifically, our proposed framework tackles the problem of presynthesis estimation of data queuing cost, while incorporating the potential impact of resource constraints on the final implementation. We present a probabilistic push-and-pull approach for register queue size estimation of a streaming data flow graph. We evaluated our techniques using an industrial tool/low. For the register queue sizes our estimations are within the range of -14.4% to 12.4% on an average, for various resource constraints on a set of multimedia applications Somsubhra Mondal, Seda Ogrenci Memik, Nikolaos Bellas |
FPL | 2 |
| 2006 | Combining hardware reconfiguration and adaptive computation for a novel SoC design methodologyabstractIn the face of dominant communication overheads and reconfiguration cost of programmable hardware often deployed in SoC environments, a new paradigm is necessary to revisit the partitioning and allocation problems. Our aim is to integrate generalized performance models into codesign to explore the gray area between hardware and software effectively. We propose to use the adaptive computation approach. Adaptivity implies that due to input changes the output of the system is updated only re-evaluating those portions of the program affected by the changes. We study the impact of our model onto a SoC architecture consisting of embedded processors and dynamically reconfigurable hardware. We present an image processing application mapped onto this architecture as a case study Vincenzo Rana, Marco D. Santambrogio, Seda Ogrenci Memik, Donatella Sciuto |
FPT | 3 |
| 2006 | Physical aware frequency selection for dynamic thermal management in multi-core systemsabstractIn order to maintain performance per Watt in microprocessors, there is a shift towards the chip level multiprocessing paradigm. Microprocessor manufacturers are experimenting with tens of cores, forecasting the arrival of hundreds of cores per single processor die in the near future. With such large-scale integration and increasing power densities, thermal management continues to be a significant design effort to maintain performance and reliability in modern process technologies. In this paper, we present two mechanisms to perform frequency scaling as part of Dynamic Frequency and Voltage Scaling (DVFS) to assist Dynamic Thermal Management (DTM). Our frequency selection algorithms incorporate the physical interaction of the cores on a large-scale system onto the emergency intervention mechanisms for temperature reduction of the hotspot, while aiming to minimize the performance impact of frequency scaling on the core that is in thermal emergency. Our results show that our algorithm consistently succeeds in maximizing the operating frequency of the most critical core while successfully relieving the thermal emergency of the core. A comparison of our two alternative techniques reveals that our physical aware criticality-based algorithm results in 11.7 % faster clock frequencies compared to our aggressive scaling algorithm. We also show that our technique is extremely fast and is suited for real time thermal management Rajarshi Mukherjee, Seda Ogrenci Memik |
ICCAD | 2 |
| 2006 | Thermal sensor allocation and placement for reconfigurable systemsabstractTemperature monitoring using thermal sensors is an essential tool for evaluating the thermal behavior and sustaining the reliable operation in high-performance and high-power systems. With current technology scaling and integration trends timely and accurate detection of localized heating will be evermore important. In this work, we address the creation of a resource efficient sensor infrastructure for computing systems that are of regular nature, such as logic array-based computing platforms. We propose algorithms to embed thermal sensors into a regular structure to minimize the number of sensors and determine sensor locations required to maintain a given accuracy in temperature sensing for a given design. Our algorithms are tailored for minimal usage of thermal sensors to suit a variety of architectural conditions. For programmable logic arrays the highly application-specific usage of the hardware resources leads to unpredictable thermal profiles. As a result, post-manufacture instantiation of thermal sensors is desired, which in turn demands the use of native hardware resources, which can be scarce. We demonstrate that using our techniques the number of sensors required to monitor a set of hotspots is reduced by 75 % on an average, across different sizes of logic arrays for different hotspot distributions compared to a uniform distribution of sensors throughout the fabrics. Rajarshi Mukherjee, Somsubhra Mondal, Seda Ogrenci Memik |
ICCAD | 3 |
| 2006 | Thermal-induced leakage power optimization by redundant resource allocationabstractTraditionally, at early design stages, leakage power is associated with the number of transistors in a design. Hence, intuitively an implementation with minimum resource usage would be best for low leakage. Such an allocation would generally be followed by switching optimal resource binding to achieve a low power design. This treatment of leakage power is unaware of operating conditions such as temperature. In this paper, we propose a technique to reduce the total leakage power of a design by identifying the optimal number of resources during allocation and binding. We demonstrate that, contrary to the general tendency to minimize the number of resources, the best solution can actually be achieved if a certain degree of redundancy is allowed. This is due to the fact that leakage is strongly dependent on the on-chip temperature profile. Distributing activity over a higher number of resources can reduce power density, remove potential hotspots and subsequently minimize thermal induced leakage. On the other hand, using an arbitrarily high number of resources will not yield the best solution. In this paper, we show that there is a power density, hence, temperature, at which the total leakage power will reach its optimal value. Such an optimal resource number can be a better starting point for the subsequent switching-driven low power binding. We also present a high-level power density-aware leakage model. Based on the estimates by this model, we optimize the total leakage power by 53.8% on average compared to the minimum resource binding, and 35.7% on average compared to a temperature-aware resource binding technique. Min Ni, Seda Ogrenci Memik |
ICCAD | 2 |
| 2006 | Fine-grain thermal profiling and sensor insertion for FPGAsabstractIncreasing logic densities and clock frequencies on FPGAs lead to rapid increase in power density, which translates to higher on-chip temperature. In this paper, we investigate the thermal behavior of general applications on fine-grain reconfigurable fabrics and we introduce the pre-mapping sensor insertion problem for thermal monitoring. Our study shows that on average the maximum temperature on the chip is 19.5degC higher than the ambient temperature for a transition density of 0.5 at the primary inputs. For fine-grain reconfigurable devices targeted for general applications it is difficult to predict the locations of potential hotspots a priori. However, programmability presents a unique opportunity for effective thermal monitoring. It would allow us to perform a thermal simulation on a given design first and obtain the locations of potential points of interest in a design. Then, in the pre-mapping stage the design can be updated with insertion of thermal sensors. Given a set of expected hot spots in a design we aim to determine the minimum number of sensors and their locations in order to monitor these locations with a given sensitivity requirement. Since the thermal sensors are implemented using unused CLBs on the fabric it is essential to use the logic resources efficiently. We propose an efficient algorithm to solve the sensor placement problem addressing this optimization goal Somsubhra Mondal, Rajarshi Mukherjee, Seda Ogrenci Memik |
ISCAS | 3 |
| 2006 | An Integrated Approach to Thermal Management in High-Level SynthesisabstractThermal effects are becoming an important factor in the design of integrated circuits due to the adverse impact of temperature on performance, reliability, leakage, and chip packaging costs. Making all phases of the design flow aware of this physical phenomenon helps in reaching faster design closure. In this paper, we present an integrated approach to thermal management in architectural synthesis. Our synthesis flow combines temperature-aware scheduling and binding based on feedback from thermal simulation. We show that our flow is effective in preventing hotspot formation and creating an even thermal profile of the resources. Our integrated thermal management technique on average reduces the peak temperature of the resources by 7.34 degC when compared to a thermal unaware flow without increasing the number of resources across our set of benchmarks Rajarshi Mukherjee, Seda Ogrenci Memik |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Resource sharing in pipelined CDFG synthesisabstractEfficient use of limited available resources on an FPGA remains a crucial problem for synthesizing pipelined designs. Resource sharing addresses this challenge. In this paper, we propose resource sharing techniques that can be incorporated into an automated synthesis flow to generate pipelined designs. Given a synthesized pipelined design, we create a direct relationship between available time slack on modules and the multiplexing overhead due to sharing. This flexibility is maximally exploited without violating any throughput constraints. We propose different techniques to address resource sharing problems of varying restrictions. Specifically, we propose an optimal algorithm for Constant-Slack Resource Sharing and a heuristic for the general Intra-Pipeline Stage Resource Sharing. On an average the demand on arithmetic functional units can be reduced by 39.5% for a set of benchmarks from the multimedia domain using our resource sharing technique. Somsubhra Mondal, Seda Ogrenci Memik |
ASP-DAC | 2 |
| 2005 | Evaluation of dual VDD fabrics for low power FPGAsabstractPower efficiency is becoming an increasingly important design aspect for FPGAs. Recently it has been shown that well-known power minimization techniques in the ASICs such as creating supply voltage (Vdd) scalable islands of different granularity can be applied to FPGAs. However, the discrete routing architecture of FPGAs amplifies any constraint imposed on the placement stage. In this work, we evaluate the overheads of voltage scaling schemes in relation to FPGA architectures and design flows in terms of critical path delay, channel-width and area/delay product. We present a detailed evaluation of the impact of alternative realizations of voltage scaling schemes onto the physical design flow of FPGAs and show that as high as 47% dynamic power gain is possible with 17% area/delay product penalty and 30% power gain is possible with as low as 6% area/delay product penalty for different voltage island configurations. Rajarshi Mukherjee, Seda Ogrenci Memik |
ASP-DAC | 2 |
| 2005 | Temperature-aware resource allocation and binding in high-level synthesisabstractPhysical phenomena such as temperature have an increasingly important role in performance and reliability of modern process technologies. This trend will only strengthen with future generations. Attempts to minimize the design effort required for reaching closure in reliability and performance constraints are agreeing on the fact that higher levels of design abstractions need to be made aware of lower level physical phenomena. In this paper, we investigated techniques to incorporate temperature-awareness into high-level synthesis. Specifically, we developed two temperature-aware resource allocation and binding algorithms that aim to minimize the maximum temperature that can be reached by a resource in a design. Such a control scheme will have an impact on the prevention of hot spots, which in turn is one of the major hurdles in front of reliability for future integrated circuits. Our algorithms are able to reduce the maximum attained temperature by any module in a design by up to 19.6°C compared to a binding that optimizes switching power. Rajarshi Mukherjee, Seda Ogrenci Memik, Gokhan Memik |
DAC | 2 |
| 2005 | Hierarchical LUT structures for leakage power reduction (abstract only)abstractReconfigurable technologies have made remarkable progress in the last few years. However, due to the overhead of programmability along with improved system capabilities, power dissipation has become one of the major concerns for designers. Moreover, with the trends in technology scaling, overall leakage power consumption is increasing alarmingly. In this work, we propose a leakage power optimization technique using look-up tables (LUTs) for SRAM-based FPGAs. Analysis of a set of 20 MCNC benchmarks using 4-input LUTs show that on an average only 53% of these LUTs use all their inputs. Based on this observation, we shut down SRAM cells and certain transistors associated with the unused inputs by employing hierarchical look-up tables, which can yield LUTs with varying number of inputs within the same logic block. We effectively utilize a Vdd gating scheme to cut-off power supply to one half or three quarters of the original 4-input LUT, which in turn yields a 3-input LUT or a 2-input LUT from a 4-input LUT respectively. Our experiments show that for 180 nm technology logic block leakage power savings is 22.9% on an average for the set of 20 MCNC benchmarks. This saving will be even higher for technologies 90 nm and smaller. Somsubhra Mondal, Seda Ogrenci Memik, Debasish Das |
FPGA | 2 |
| 2005 | Real-Time Feature Extraction for High Speed NetworksabstractWith the onset of Gigabit networks, current generation networking components will soon be insufficient for numerous reasons: most notably because existing methods cannot support high performance demands. Feature extraction (or flow monitoring), an essential component in anomaly detection, summarizes network behavior from a packet stream. This information is fed into intrusion detection methods such as association rule mining, outlier analysis, and classification algorithms in order to characterize network behavior. However, current feature extraction methods based on per-flow analysis are expensive, not scalable, and thus prohibitive for large-scale networks. In this paper, we propose an accurate and scalable feature extraction module (FEM) based on sketches. We present the details of the FEM design on an FPGA and show that using FPGAs we can achieve significantly better performance compared to existing software and ASIC implementations. Specifically, the optimal FEM configuration achieves 21.25 Gbps throughput and 97.61% accuracy. Gokhan Memik, Seda Ogrenci Memik, Alok N. Choudhary |
FPL | 3 |
| 2005 | Fine-grain leakage optimization in SRAM based FPGAsabstractFPGAs are evolving at a rapid pace with improved performance and logic density. At the same time, trends in technology scaling makes leakage power a serious concern for designers. In this paper, we propose a hierarchical look-up table (LUT) structure for FPGAs to improve leakage power consumption. We present a detailed analysis on the number of inputs actually used by LUTs, and we observe that on an average 47% LUTs do not use one or more inputs. In the proposed hierarchical LUT structure depending on the number of inputs used by the LUTs we shut off certain SRAM cells and transistors associated with the unused LUT inputs. Based on this technique, for 180nm technology, we report an average savings of 22.94% (as high as 64.22%) in leakage power per LUT. The savings will be even greater for technologies as low as 90nm currently in use for FPGA production as well as for future technologies. Somsubhra Mondal, Seda Ogrenci Memik |
ACM Great Lakes Symposium on VLSI | 2 |
| 2005 | Peak temperature control and leakage reduction during binding in high level synthesisabstractTemperature is becoming a first rate design criterion in ASICs due to its negative impact on leakage power, reliability, performance, and packaging cost. Incorporating awareness of such lower level physical phenomenon in high level synthesis algorithms will help to achieve better designs. In this work, we developed a temperature aware binding algorithm. Switching power of a module correlates with its operating temperature. The goal of our binding algorithm is to distribute the activity evenly across functional units. This approach avoids steep temperature differences between modules on a chip, hence, the occurrence of hot spots. Starting with a switching optimal binding solution, our algorithm iteratively minimizes the maximum temperature reached by the hottest functional unit. Our algorithm does not change the number of resources used in the original binding. We have used HotSpot, a temperature modeling tool, to simulate temperature of a number ASIC designs. Our binding algorithm reduces temperature reached by the hottest resource by 12.21°C on average. Reducing the peak temperature has a positive impact on leakage as well. Our binding technique improves leakage power by 11.89%, and overall power by 3.32% on average at 130nm technology node compared to a switching optimal binding Rajarshi Mukherjee, Seda Ogrenci Memik, Gokhan Memik |
ISLPED | 2 |
| 2005 | On effective slack management in postscheduling phaseabstractIn this paper, we propose techniques for effective slack management in high-level synthesis. Our design methodology improves the usability of slack. This manifests itself in the form of relaxed latency constraints on resources. Relaxed latency constraints could be exploited to generate designs with better power, area, routability, and other measures. The slack-management engine has two key components: delay budgeting and resource binding. We propose a left edge traversal-based algorithm for delay budgeting. For resource binding, we developed an algorithm that applies a locally optimal binding procedure at each clock step. In order to demonstrate the effectiveness of our strategy, we built an experimental flow that integrated SUIF, Synopsys Design Compiler, Cadence Silicon Ensemble, and our own optimization tools. Experiments with the MediaBench suite shows that our methodology could generate designs with better quality than designs and faster design closure when compared with designs generated without slack management. Ankur Srivastava 0001, Seda Ogrenci Memik, Bo-Kyung Choi, Majid Sarrafzadeh |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2005 | A scheduling algorithm for optimization and early planning in high-level synthesisabstractComplexities of applications implemented on embedded and programmable systems grow with the advances in capacities and capabilities of these systems. Mapping applications onto them manually is becoming a very tedious task. This draws attention to using high-level synthesis within design flows. Meanwhile, it is essential to provide a flexible formulation of optimization objectives as well as to perform efficient planning for various design objectives early on in the design flow. In this work, we address these issues in the context of data flow graph (DFG) scheduling, which is an essential element within the high-level synthesis flow. We present an algorithm that schedules a chain of operations with data dependencies among consecutive operations at a single step. This local problem is repeated to generate the schedule for the whole DFG. The local problem is formulated as a maximum weight noncrossing bipartite matching. We use a technique from the computational geometry domain to solve the matching problem. This technique provides a theoretical guarantee on the solution quality for scheduling a single chain of operations. Although still being local, this provides a relatively wider perspective on the global scheduling objectives. In our experiments we compared the latencies obtained using our algorithm with the optimal latencies given by the exact solution to the integer linear programming (ILP) formulation of the problem. In 9 out of 14 DFGs tested, our algorithm found the optimal solution, while generating latencies comparable to the optimal solution in the remaining five benchmarks. The formulation of the objective function in our algorithm provides flexibility to incorporate different optimization goals. We present examples of how to exploit the versatility of our algorithm with specific examples of objective functions and experimental results on the ability of our algorithm to capture these objectives efficiently in the final schedules. Seda Ogrenci Memik, Ryan Kastner, Elaheh Bozorgzadeh, Majid Sarrafzadeh |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2004 | Power Management for FPGAs: Power-Driven Design PartitioningabstractIn order to enable efficient integration of FPGAs into cost effective and reliable high-performance systems as well potentially into low power mobile systems, their power efficiency needs to be improved. This paper proposes a power management scheme for FPGAs centered on the power-driven partitioning technique. Power-driven partitioner create clusters within the a design such that within individual clusters, power consumption can be improved via voltage scaling. The aim is to identify subgraphs/partitions in a design, such that the total power consumption is minimised while resource constraints associated with the partitioning problem are satisfied. Rajarshi Mukherjee, Seda Ogrenci Memik |
FCCM | 2 |
| 2004 | Power-Driven Design Partitioning
Rajarshi Mukherjee, Seda Ogrenci Memik |
FPL | 2 |
| 2003 | Global resource sharing for synthesis of control data flow graphs on FPGAsabstractIn this paper we discuss the global resource sharing problem during synthesis of control data flow graphs for FPGAs. We first define the Global Resource Sharing (GRS) problem. Then, we introduce the Global Inter Basic Block Resource Sharing (GIBBS) technique to solve the GRS problem. We developed five heuristics to solve the GRS problem. The first tries to minimize the number of connections between modules, the second considers the area gain, the third uses the criticality of operations assigned to resources as a measure for deciding on merging any given pair of resources, the fourth tries to capture common resource chains and overlap those to minimize both area and delay, and the fifth is the combination of these heuristics. While applying resource sharing, we also consider the execution frequency of the basic blocks. Using our techniques we synthesized several CDFGs representing applications from MediaBench suite. Our results show that, we can reduce the total area requirement by 44% on average (up to 59%) while increasing the execution time by 6% on average. Seda Ogrenci Memik, Gokhan Memik, Roozbeh Jafari, Eren Kursun |
DAC | 1 |
| 2003 | Achieving Design Closure Through Delay Relaxation Parameter
Ankur Srivastava 0001, Seda Ogrenci Memik, Bo-Kyung Choi, Majid Sarrafzadeh |
ICCAD | 2 |
| 2003 | Analysis and FPGA Implementation of Image Restoration under Resource ConstraintsabstractProgrammable logic is emerging as an attractive solution for many digital signal processing applications. In this work, we have investigated issues arising due to the resource constraints of FPGA-based systems. Using an iterative image restoration algorithm as an example we have shown how to manipulate the original algorithm to suit it to an FPGA implementation. Consequences of such manipulations have been estimated, such as loss of quality in the output image. We also present performance results from an actual implementation on a Xilinx FPGA. Our experiments demonstrate that, for different criteria, such as result quality or speed, the best implementation is different as well. Seda Ogrenci Memik, Aggelos K. Katsaggelos, Majid Sarrafzadeh |
IEEE Trans. Computers | 1 |
| 2002 | Design and Analysis of a Layer Seven Network Processor Accelerator Using Reconfigurable LogicabstractIn this paper, we present an accelerator that is designed to improve performance of network processing applications, particularly layer seven networking applications. The accelerator can easily be integrated in Network Processors. We present the design details of two different FPGA implementations: a design where each task is implemented in the accelerator and another one where the accelerator must be partially reconfigured for different tasks. We also present novel algorithms for important tasks such as tree lookup and pattern matching that utilize the accelerator. We show that the accelerator improves the overall execution time by as much as 20-times for these tasks. We show that the accelerator can improve the execution time of a representative layer seven application by an order of magnitude. Finally, we discuss the effects of reconfiguration time and frequency over the performance of the accelerator. Gokhan Memik, Seda Ogrenci Memik, William H. Mangione-Smith |
FCCM | 2 |
| 2002 | Accelerated SAT-based Scheduling of Control/Data Flow GraphsabstractIn this paper we present a satisfiability-based approach to the scheduling problem in high-level synthesis. We formulate the resource constrained scheduling as a satisfiability (SAT) problem. We present experimental results on the performance of the state-of-the-art SAT solver Chaff, and demonstrate techniques to reduce the SAT problem size by applying bounding techniques on the scheduling problem. In addition, we demonstrate the use of transformations on control data flow graphs such that the same lower bound techniques can operate on them as well. Our experiments show that Chaff is able to outperform the integer linear program (ILP) solver CPLEX in terms of CPU time by as much as 59 fold. Finally, we conclude that the satisfiability-based approach is a promising alternative for obtaining optimal solutions to NP-complete scheduling problem instances. Seda Ogrenci Memik, Farzan Fallah |
ICCD | 1 |
| 2002 | Early evaluation techniques for low power bindingabstractThis paper presents effective metrics to evaluate the power dissipation of scheduled data flow graphs (DFGs). This enables early evaluation of schedules without performing the computationally expensive resource-binding step. Our metrics correlate heavily (as high as 0.95 and > 0.75 for most test cases) with power dissipation values obtained after resource binding and rescheduling for power optimization steps. An experimental flow that integrates path-based scheduling, power optimal binding and power driven iterative rescheduling stages is constructed. The flow integrates commercial tools; like Synopsys, VSS and academic compilers like SUIF in a common optimization framework. Experimental results on DFGs from MediaBench suit also demonstrate the fact that metric evaluation is on average 42.6 times faster than performing optimal binding and iterative power improvement. Hence metric based evaluation enables fast design exploration at early stages. Eren Kursun, Ankur Srivastava 0001, Seda Ogrenci Memik, Majid Sarrafzadeh |
ISLPED | 3 |
| 2002 | Instruction generation for hybrid reconfigurable systemsabstractFuture computing systems need to balance flexibility, specialization, and performance in order to meet market demands and the computing power required by new applications. Instruction generation is a vital component for determining these trade-offs. In this work, we present theory and an algorithm for instruction generation. The algorithm profiles a dataflow graph and iteratively contracts edges to create the templates. We discuss how to target the algorithm toward the novel problem of instruction generation for hybrid reconfigurable systems. In particular, we target the Strategically Programmable System, which embeds complex computational units such as ALUs, IP blocks, and so on into a configurable fabric. We argue that an essential compilation step for these systems is instruction generation, as it is needed to specify the functionality of the embedded computational units. In addition, instruction generation can be used to create soft reconfigurable macros---tightly sequenced prespecified operations placed in the reconfigurable fabric. Ryan Kastner, Adam Kaplan, Seda Ogrenci Memik, Elaheh Bozorgzadeh |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2001 | RPack: routability-driven packing for cluster-based FPGAsabstractRouting tools consume a significant portion of the total design time. Considering routability at earlier steps of the CAD flow would both yield better quality and faster design process. In this paper we are presenting a routability-driven clustering method for cluster-based FPGAs. Our method packs LUTs into logic clusters while incorporating routability metrics into a cost function. The objective is to minimize this routability cost function . Our cost function is consistently able to indicate improved routability. Our method yields up to 50 % improvement over existing clustering methods in terms of the number of routing tracks required. The average improvement obtained is 16.5 %. Reduction in number of tracks yields reduced routing area. Elaheh Bozorgzadeh, Seda Ogrenci Memik, Majid Sarrafzadeh |
ASP-DAC | 2 |
| 2001 | Integrating Scheduling and Physical Design into a Coherent Compilation Cycle for Reconfigurable Computing Architectures
Kia Bazargan, Seda Ogrenci Memik, Majid Sarrafzadeh |
DAC | 2 |
| 2001 | Instruction Generation for Hybrid Reconfigurable SystemsabstractWe present an algorithm for simultaneous template generation and matching. The algorithm profiles the graph and iteratively contracts edges to create the templates. The algorithm is general and can be applied to any type of graph, including directed graphs and hypergraphs. We discuss how to target the algorithm towards the novel problem of instruction generation and selection for a hybrid (re)configurable systems. In particular, we target the strategically programmable system, which embeds complex computational units like ALUs, IP blocks, etc. into a configurable fabric. We argue that an essential compilation step for these systems is instruction generation, as it is needed to specify the functionality of the embedded computational units. Additionally, instruction generation can be used to create soft macros tightly sequenced pre-specified operations placed in the configurable fabric. Ryan Kastner, Seda Ogrenci Memik, Elaheh Bozorgzadeh, Majid Sarrafzadeh |
ICCAD | 2 |
| 2001 | A Super-Scheduler for Embedded Reconfigurable SystemsabstractEmerging reconfigurable systems attain high performance with embedded optimized cores. For mapping designs on such special architectures, synthesis tools, that are aware of the special capabilities of the underlying architecture are necessary. We propose an algorithm to perform simultaneous scheduling and binding, targeting embedded reconfigurable systems. The algorithm differs from traditional scheduling methods in its capability of efficiently utilizing embedded blocks within the reconfigurable system. The algorithm can be used to implement several other scheduling techniques, such as ASAP, ALAP, and list scheduling. Hence we refer to it as a super-scheduler. The algorithm is a path-based scheduling algorithm. At each step, an individual path from the input DFG is scheduled. The experiments with several DFGs extracted from MediaBench suite indicate promising results. The scheduler presents the capability to perform the trade-off between maximally utilizing the high-performance embedded blocks and exploiting parallelism in the schedule. Seda Ogrenci Memik, Elaheh Bozorgzadeh, Ryan Kastner, Majid Sarrafzadeh |
ICCAD | 1 |
| 2001 | Fast floorplanning for effective prediction and constructionabstractFloorplanning is a crucial phase in VLSI physical design. The subsequent placement and routing of the cells/modules are coupled very closely with the quality of the floorplan. A widely used technique for floorplanning is simulated annealing. It gives very good floorplanning results but has major limitation in terms of run time. For circuit sizes exceeding tens of modules simulated annealing is not practical. Floorplanning forms the core of many synthesis applications. Designers need faster prediction of system metrics to quickly evaluate the effects of design changes. Early prediction of metrics is imperative for estimating timing and routability. In this work we propose a constructive technique for predicting floorplan metrics. We show how to modify the existing top-down partitioning-based floorplanning to obtain a fast and accurate floorplan prediction. The prediction gets better as the number of modules and flexibility in the shapes increase. We also explore applicability of the traditional sizing theorem when combining two modules based on their sizes and interconnecting wirelength. Experimental results show that our prediction algorithm can predict the area/length cost function normally within 5-10% of the results obtained by simulated annealing and is, on average, 1000 times faster. Kia Bazargan, Seda Ogrenci Memik, Majid Sarrafzadeh |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2000 | A C to Hardware/Software CompilerabstractImprovements in FPGA technology have resulted in the introduction of reconfigurable computing machines, where the hardware adapts itself to the running application to gain speedup. We present a top-down compilation method, under development, for such systems. We compile a C program into hierarchical VHDL source files, and annotate them with the placement information of the hardware modules to be configured on the FPGA. Static scheduling combined with a fast, two-stage placement core reduces the compilation time of large programs to minutes. Kia Bazargan, Ryan Kastner, Seda Ogrenci Memik, Majid Sarrafzadeh |
FCCM | 3 |