EDBT 2026 Demo / reviewers in the wild / expert
Yusuf Leblebici
dblp:l/YusufLeblebici
· DBLP profile ↗
122ranked-venue papers
4as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 97 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 12Software engineering, systems software and programming languages · 12Graphics, computer vision, multimedia, augmented reality and games · 9Security and privacy · 2Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fixed-Point Implementation Analysis of MLSE Receiver DSP for High-Speed Wireline Transceivers
Dohyeon Kwon, Donggeon Kim, Yusuf Leblebici, Gain Kim |
ISCAS | 4 |
| 2021 | An 8-Bit 800 MS/s Loop-Unrolled SAR ADC With Common-Mode Adaptive Background Offset Calibration in 28 nm FDSOIabstractThis paper presents a low-power single-channel 8-bit loop-unrolled (LU) successive approximation register (SAR) analog-to-digital-converter (ADC) with a novel common-mode adaptive background comparator offset calibration scheme. LU-SAR ADCs use multiple comparators to reduce the SAR loop delay. Offset mismatch between the comparators severely degrades the effective resolution. This paper addresses the common-mode voltage variation in the LU-SAR architecture due to comparator kickback and the problems related to the common-mode dependency of the comparator offset. The proposed offset calibration scheme ensures that the comparators are calibrated at the same input common-mode voltage at which they each operate during the SAR conversion to prevent the common-mode dependent offset mismatch between them. Moreover, the proposed ADC design exploits the common-mode variation immunity of the proposed calibration scheme to optimize the figure-of-merit (FoM). The prototype ADC manufactured in 28nm FDSOI CMOS achieves 42.57dB signal-to-noise-and-distortion ratio and 22.8fJ/conv.-step FoM at 800MS/s with near Nyquist frequency input, and occupies an area of 0.0037mm2. Ayca Akkaya, Firat Celik, Yusuf Leblebici |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | A 32-Gb/s PAM-4 SST Transmitter With Four-Tap FFE Using High-Impedance Driver in 28-nm FDSOIabstractWith increasing modulation order, a larger number of parallel source-series-terminated (SST) segments are required to implement precise feedforward equalization (FFE) tap-weight control in SST transmitters (TX). Due to the data routing complexity of the segmented structure, the dynamic power consumption of SST TX has become much more significant causing degradation in energy efficiency. This article presents a 32-Gb/s quarter-rate four-level pulse-amplitude modulation (PAM-4) SST TX with four-tap FFE implemented in 28-nm FDSOI CMOS technology using a high-impedance driver technique to decrease the gate loading of the data path of the SST TX. The output of the whole TX is kept matched to the standard characteristic impedance of the system even though the output impedance of the SST driver alone is high. The high-impedance driver technique decreases the total power consumption by 20% compared to the conventional design by providing a significant reduction in the capacitive load. Our measurement results show that the prototype TX consumes 77.9 mW and achieves an energy efficiency of 2.4 pJ/bit at a 32-Gb/s data rate for PAM-4 signaling. Firat Celik, Ayca Akkaya, Armin Tajalli, Yusuf Leblebici |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | A Low-Power 9-Bit 222 MS/s Asynchronous SAR ADC in 65 nm CMOSabstractThis paper presents a 9-bit 222 MS/s low-power asynchronous single-bit/cycle successive approximation register (SAR) ADC. The SAR ADC combines techniques such as asynchronous clocking, binary-weighted custom-designed capacitive DAC with small unit capacitors, splitting monotonic capacitor switching, and dynamic SAR memory to optimize both power consumption and SAR loop delay using a single comparator. Measurement results show that the 9-bit SAR ADC achieves 47.6 dB SNDR and 29.6 fJ/conversion-step figure-of-merit (FoM) near Nyquist frequency at 222 MS/s, consuming 1.07 mA from a 1.2 V supply. The ADC does not consume static power and the current consumption scales down linearly with the sampling rate, keeping the low FoM value constant over a very wide sampling frequency range. The design occupies an active area of 92 μm × 180 μm (0.017 mm2) in 65 nm CMOS. Ayca Akkaya, Firat Celik, Yusuf Leblebici |
ISCAS | 3 |
| 2020 | Discrete Time Analysis of Phase Detector Linear Range Extension in Sub-Sampling PLLabstractThe discrete time analysis of the phase detector linear range extension in the first and the second order sub-sampling phase-locked loop (SSPLL) is presented. The aim is to understand how much the stability and the pull-in range are affected by the linear range of the sub-sampled phase detector. To change the linear range, sinusoidal, triangular and sawtooth signals are used as the voltage-controlled oscillator (VCO) output signals. The discrete time domain equations of the first and the second order loops are built and solved to determine the stability conditions. The pull-in ranges are examined numerically. The simulations are done in MATLAB, and it is shown that the linear range extension results in an extended pull-in range in both cases, but the initial condition of the loop filter affects the extended pull-in range for the second order loop. Duygu Kostak, Yusuf Leblebici |
ISCAS | 2 |
| 2019 | Real-Time Textureless-Region Tolerant High-Resolution Depth Estimation SystemabstractThis study presents a real-time depth estimation hardware system aiming to provide high-resolution depth data and to eliminate its noise on textureless regions without causing any interference problem by utilizing artificial pattern projection. The system generates up to 2K resolution depth data and reaches up to 256 disparity range which are configurable by the end-user owing to its parameterized design. It is capable of streaming depth data with 21 frames per second (fps) with 2K resolution and 128 pixel disparity range, and its throughput performance changes depending on configuration of the output resolution and the disparity range. Bilal Demir, Jean-Philippe Thiran, Yusuf Leblebici |
DSD | 3 |
| 2019 | Multi-ReRAM Synapses for Artificial Neural Network TrainingabstractMetal-oxide-based resistive memory devices (ReRAM) are being actively researched as synaptic elements of neuromorphic co-processors for training deep neural networks (DNNs). However, device-level non-idealities are posing significant challenges. In this work we present a multi-ReRAM-based synaptic architecture with a counter-based arbitration scheme that shows significant promise. We present a 32×2 crossbar array comprising Pt/HfO2/Ti/TiN-based ReRAM devices with multi-level storage capability and bidirectional conductance response. We study the device characteristics in detail and model the conductance response. We show through simulations that an in-situ trained DNN with a multi-ReRAM synaptic architecture can perform handwritten digit classification task with high accuracies, only 2% lower than software simulations using floating point precision, despite the stochasticity, nonlinearity and large conductance change granularity associated with the devices. Moreover, we show that a network can achieve accuracies > 80% even with just binary ReRAM devices with this architecture. Irem Boybat, Cecilia Giovinazzo, Elmira Shahrabi, Igor Krawczuk, Iason Giannopoulos, Christophe Piveteau, Manuel Le Gallo, Carlo Ricciardi, Abu Sebastian, Evangelos Eleftheriou, Yusuf Leblebici |
ISCAS | 11 |
| 2018 | An area and power efficient on-the-fly LBCS transformation for implantable neuronal signal acquisition systemsabstractA power and area efficient hardware encoding system tailored for wireless implantable applications is presented. Constant medical monitoring allowed by implantable devices is the most relevant alternative to current bulky monitoring systems, which, in case of severe mental diseases, require heavy surgery and long term hospitalization periods. In this work, the circuit design and the signal processing algorithm dovetail in order to allow real-time neuronal signal monitoring. Two main features must be met on the circuit level to facilitate the acceptance of the implant from the human body: small area and low power consumption. The presented work proposes a new compression scheme based on the Learning-Based Compressive Subsampling approach, which allows an area reduction with respect to recent published works, while allowing high signal reconstruction quality within low power requirements. The proposed method implements on-the-fly compression coefficients generation, which does not require large static memories. This new fully digital architecture handles the data compression of each individual neuronal acquisition channel with an area of 200 × 190μm in 0.18 μm CMOS technology, and a power dissipation of only 1.15μW. Cosimo Aprile, Johannes Wüthrich, Luca Baldassarre, Yusuf Leblebici, Volkan Cevher |
CF | 4 |
| 2018 | Online Feature Learning from a non-i.i.d. Stream in a Neuromorphic System with Synaptic CompetitionabstractNeuromorphic computing takes inspiration from how the brain works to design power- and area-efficient hardware architectures for learning systems. Recently, unsupervised feature learning neuromorphic architectures have been presented, including a concept of synaptic competition that promotes the engagement of the synapses in the learning beyond weight storage. However, it is common to train these neuromorphic systems following the classic machine learning assumption of i.i.d. dataset sampling, which may not hold for real world inputs. In this paper, we propose a more realistic dataset sampling technique and apply it for online learning in a neuromorphic system using phase-change memristors as synapses and implementing synaptic competition. Furthermore, we propose a novel formulation of synaptic competition that captures orthogonal features, alternatively to independent components. We experimentally demonstrate the operation of the system for a non-i.i.d. stream and compare the performance to the models of lateral inhibition and dendritic inhibition. The obtained results demonstrate online feature learning capabilities of the proposed system and robustness to non-i.i.d. inputs. Stanislaw Wozniak, Angeliki Pantazi, Yusuf Leblebici, Evangelos Eleftheriou |
IJCNN | 3 |
| 2018 | Hardware Implementation of a Smart Camera with Keypoint Detection and DescriptionabstractFeature detection and description constitute important steps of many computer vision applications such as object detection and panorama stitching. Since those steps are computationally heavy, they might occupy significant portion of the full operation. Although fast feature detection algorithms and resource-efficient binary description methods have been proposed and implemented, resource limited embedded devices and distributed camera systems still require more effective solutions. In this paper, we propose a novel smart camera architecture which finds the FAST keypoints and computes their FREAK descriptions by processing pixel stream. Thus, this smart camera system provides useful metadata associated with the pixel stream at the same time with no latency. Moreover, performance of this hardware reaches very high frame rates with power and area efficiency. With this approach, this costly operation is locally solved in the smart camera node, and this leads to meet timing and power constraints of the large camera networks. Selman Ergünay, Yusuf Leblebici |
ISCAS | 2 |
| 2018 | Parallel Implementation Technique of Digital Equalizer for Ultra-High-Speed Wireline ReceiverabstractThis paper presents a parallel implementation technique of digital equalizer for high-speed wireline serial link receiver (RX). In wireline RX, inter-symbol interference (ISI) is mitigated by continuous-time linear equalizer, and the remaining ISI is cancelled out by decision-feedback equalizer (DFE). However, due to the existence of feedback loop in DFE, there is no trivial way to parallelize it, making it difficult to be realized in digital circuits for wireline RX based on analog-to-digital converter (ADC) with ≥ 56 Gb/s data rate. In this work, convolution theorem is applied for achieving parallel digital equalizer implementation. The digital equalizer datapath consists of discrete Fourier transform (DFT) core, inverse-DFT (IDFT) core, complex multipliers between DFT and IDFT cores, and overlap-add circuit. Design considerations for low-area VLSI implementation of such architecture is discussed. Gain Kim, Lukas Kull, Danny Luu, Matthias Braendli, Christian Menolfi, Pier Andrea Francese, Cosimo Aprile, Thomas Morf, Marcel A. Kossel, Alessandro Cevrero, Ilter Özkaya, Thomas Toifl, Yusuf Leblebici |
ISCAS | 13 |
| 2018 | FPGA-Based Hardware Implementation of Real-Time Optical Flow CalculationabstractOptical flow calculation algorithms are hard to implement on the hardware level in real-time, due to their complexity and high computational load. Therefore, presented works in the literature focusing on the hardware implementation are limited. In this paper, we present a hierarchical block matching-based optical flow algorithm suitable for real-time hardware implementation. The algorithm estimates the initial optical flow with block matching and refines the vectors with local smoothness constraints in each level. We evaluate the proposed algorithm with novel data sets and provide results compared with the ground truth optical flow. Furthermore, we present a reconfigurable hardware architecture of the proposed algorithm for calculating the optical flow in real-time. The presented system can process $640\times 480$ resolution frames at 39 frames/s. Kerem Seyid, Andrea Richaud, Raffaele Capoccia, Yusuf Leblebici |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Unsupervised Learning Using Phase-Change Synapses and Complementary Patterns
Severin Sidler, Angeliki Pantazi, Stanislaw Wozniak, Yusuf Leblebici, Evangelos Eleftheriou |
ICANN (1) | 4 |
| 2017 | Neuromorphic system with phase-change synapses for pattern learning and feature extractionabstractNeuromorphic systems provide biologically inspired methods of computing, alternative to the classical von Neumann approach. In these systems, computation is performed by a network of spiking neurons controlled by the values of their synaptic weights, which are updated in the process of learning. Providing efficient synaptic learning rules, such as spike-timing-dependent plasticity (STDP), is a challenging task. These rules need to primarily use local information, but simultaneously develop a knowledge representation that is useful in the global context. From the implementation viewpoint, they also need to be suited for particular hardware technology. In this work, we propose a system with spiking neurons and synapses realized using phase-change devices. We design in a bottom-up manner an architecture for pattern learning and feature extraction. Experimental results from a prototype hardware platform demonstrate the capabilities of the proposed neuromorphic system. Stanislaw Wozniak, Angeliki Pantazi, Yusuf Leblebici, Evangelos Eleftheriou |
IJCNN | 3 |
| 2017 | Reducing circuit design complexity for neuromorphic machine learning systems based on Non-Volatile Memory arraysabstractMachine Learning (ML) is an attractive application of Non-Volatile Memory (NVM) arrays [1,2]. However, achieving speedup over GPUs will require minimal neuron circuit sharing and thus highly area-efficient peripheral circuitry, so that ML reads and writes are massively parallel and time-multiplexing is minimized [2]. This means that neuron hardware offering full `software-equivalent' functionality is impractical. We analyze neuron circuit needs for implementing back-propagation in NVM arrays and introduce approximations to reduce design complexity and area. We discuss the interplay between circuits and NVM devices, such as the need for an occasional RESET step, the number of programming pulses to use, and the stochastic nature of NVM conductance change. In all cases we show that by leveraging the resilience of the algorithm to error, we can use practical circuit approaches yet maintain competitive test accuracies on ML benchmarks. Pritish Narayanan, Lucas L. Sanches, Alessandro Fumarola, Robert M. Shelby, Stefano Ambrogio, Jun-Woo Jang, Hyunsang Hwang, Yusuf Leblebici, Geoffrey W. Burr |
ISCAS | 8 |
| 2017 | Thermal aware design and comparative analysis of a high performance 64-bit adder in FD-SOI and bulk CMOS technologies
Can Baltaci, Yusuf Leblebici |
Integr. | 2 |
| 2016 | Learning-Based Near-Optimal Area-Power Trade-offs in Hardware Design for Neural Signal AcquisitionabstractWireless implantable devices capable of monitoring the electrical activity of the brain are becoming an important tool for understanding and potentially treating mental diseases such as epilepsy and depression. While such devices exist, it is still necessary to address several challenges to make them more practical in terms of area and power dissipation. In this work, we apply Learning Based Compressive Subsampling (LBCS) to tackle the power and area trade-offs in neural wireless devices. To this end, we propose a low-power and area-efficient system for neural signal acquisition which yields state-of-art compression rates up to 64x with high reconstruction quality, as demonstrated on two human iEEG datasets. This new fully digital architecture handles one neural acquisition channel, with an area of 210x210μm in 90nm CMOS technology, and a power dissipation of only 1μW. Cosimo Aprile, Luca Baldassarre, Juhwan Yoo, Mahsa Shoaran, Yusuf Leblebici, Volkan Cevher |
ACM Great Lakes Symposium on VLSI | 6 |
| 2016 | Design and Implementation of Real-Time Multi-sensor Vision SystemsabstractImplementation of high performance multi-camera / multi-sensor imaging systems that are required to produce real-time video output pose a large number of unique challenges to conventional digital design based on general-purpose processors or GPUs. In potential application areas ranging from machine vision, automotive, and virtual reality, the need for real-time operation with very limited latency dictates customized architectures that are not readily implementable using conventional approaches. In this talk, we will discuss the algorithms that are utilized in such multi-sensor platforms, the specialized architectures that allow streaming video processing, and system-level integration issues. Problems such as light-field reconstruction, multi-band blending, pixel-level interpolation, and rectification are presented, with detailed circuit and system level examples. Yusuf Leblebici |
ACM Great Lakes Symposium on VLSI | 1 |
| 2016 | A fully-digital spectrum shaping signaling for serial-data transceiver with crosstalk and ISI reduction property in multi-drop memory interfacesabstractAn efficient signaling scheme for serial-data transceivers (TRXs) has been proposed, which can properly reduce inter-symbol interference (ISI) and crosstalk (Xtalk) in memory interfaces. The proposed architecture relies on fully-digital implementation rather than analog/multi-tone approach, which can offer a very power-efficient and versatile silicon implementation. Moreover, the Xtalk induced noise can be fairly reduced by applying the proposed signaling, and the whole TRX can customize to the communication link trough digital calibration, while the aggregate data rate is kept fixed. Kiarash Gharibdoust, Gain Kim, Armin Tajalli, Yusuf Leblebici |
ISCAS | 4 |
| 2016 | Block matching based real-time optical flow hardware implementationabstractOptical flow calculation algorithms are hard to implement on hardware level in real-time due to their complexity and high computational load. In this work, we present a novel hierarchical block matching based optical flow algorithm. The algorithm estimates the initial optical flow with block matching based methods, and refines the vectors with local smoothness constraints in each hierarchy level. We evaluate the proposed algorithm with novel datasets and provide results compared to ground truth optical flow. Furthermore, we present a hardware architecture of the proposed algorithm for calculating optical flow in real-time. The presented design can process 640×480 resolution at 26 frames per second (fps). Kerem Seyid, Andrea Richaud, Raffaele Capoccia, Yusuf Leblebici |
ISCAS | 4 |
| 2016 | Review of advances in neural networks: Neural design technology stack
Adela-Diana Almasi, Stanislaw Wozniak, Valentin Cristea, Yusuf Leblebici, Antonius P. J. Engbersen |
Neurocomputing | 4 |
| 2015 | A ultra-low-power FPGA based on monolithically integrated RRAMs
Pierre-Emmanuel Gaillardon, Xifan Tang, Jury Sandrini, Maxime Thammasack, Somayyeh Rahimian Omam, Davide Sacchetto, Yusuf Leblebici, Giovanni De Micheli |
DATE | 7 |
| 2015 | A low-power 490 mpixels/s hardware accelerator for pyramidal decomposition of imagesabstractThis paper introduces a pyramidal decomposition system suitable for high frame rate and real-time applications. The presented system's architecture omits the image transpose block used in standard separable filters, and implements internal downsampling to reduce number of computations. The decomposition is implemented in form of a field programmable gate array (FPGA) hardware accelerator and the presented results show the low resource utilization of the design. The internal downsampling reduces the power consumption by an order of magnitude compared to state-of-the-art, which makes this accelerator an excellent addition to co-processors on mobile platforms. Vladan Popovic, Yusuf Leblebici |
ICIP | 2 |
| 2015 | Live demonstration: Real-time free viewpoint synthesis using three-camera disparity estimation hardwareabstractLive results obtained from the first real-time high-resolution free viewpoint synthesis hardware that utilizes three-camera disparity estimation are presented. The proposed hardware generates high-quality free viewpoint video at 55 frames per second on a Virtex-7 FPGA at a 1024×768 XGA video resolution for any horizontally-aligned arbitrary camera positioned between the leftmost and rightmost physical cameras. Abdulkadir Akin, Raffaele Capoccia, Jonathan Narinx, Jonathan Masur, Alexandre Schmid, Yusuf Leblebici |
ISCAS | 6 |
| 2015 | Real-time free viewpoint synthesis using three-camera disparity estimation hardwareabstractThe recent development of high-quality free viewpoint synthesis algorithms and their implementations allows to realize glasses-free 3D perception. Although many algorithms have been developed in this domain, the real-time hardware realization of a free viewpoint synthesis for real-world images is challenging due to its high computational load and memory bandwidth requirements. In this paper, the first real-time high-resolution free viewpoint synthesis hardware utilizing three-camera disparity estimation is presented. The proposed hardware generates high-quality free viewpoint video at 55 frames per second using a Virtex-7 FPGA at a 1024×768 XGA video resolution for any horizontally-aligned arbitrary camera positioned between the leftmost and rightmost physical cameras. Abdulkadir Akin, Raffaele Capoccia, Jonathan Narinx, Jonathan Masur, Alexandre Schmid, Yusuf Leblebici |
ISCAS | 6 |
| 2015 | Design of high-temperature SRAM for reliable operation beyond 250°CabstractIn this paper, we analyze the 6T SRAM cell failures caused by temperature and supply voltage variations, and we explore the design of robust SRAM cells for high temperature operation. This integral SRAM reliability study is performed using 180nm SOI CMOS process transistor models. Three different operation regions are identified based on the temperature and supply voltage impact on failure rates. We show that in the super-threshold operation region failure rates increase with elevated temperatures while the opposite is true in the sub-threshold operation regions. We also provide physical interpretation of particularly interesting near-threshold operation region that demonstrated extremely high reliability and low failure rates. Further, we present reliability improvements of the 6T SRAM cell which lead to the fully-digital Latch based design. Silicon measurements demonstrate reliable, state-of-the-art, SRAM operation at 275°C (fMAX= 10MHz, PTOT= 400mW), that is by far the highest reported operating temperature for digital on-chip SRAM module. Radisav Cojbasic, Yusuf Leblebici |
ISCAS | 2 |
| 2015 | Low-voltage read/write circuit design for transistorless ReRAM crossbar arrays in 180nm CMOS technologyabstractThis paper presents a read-write design solution for passive ReRAM crossbar memory arrays to overcome the sneak current paths problem. The proposed circuitry includes an auto-calibration feature to overcome the sneak current effects during the READ operation, and a WRITE protocol to minimize the current at each row and column lines. The presented circuit has been designed in 180nm standard CMOS technology based on the electrical characteristics of fabricated ReRAM devices. Jury Sandrini, Tugba Demirci, Maxime Thammasack, Davide Sacchetto, Yusuf Leblebici |
ISCAS | 5 |
| 2015 | Jitter analysis and measurement in subthreshold source-coupled differential ring oscillatorsabstractThe jitter and the phase noise of ring oscillators utilizing subthreshold source-coupled logic (STSCL) style are analyzed in this paper. Closed-form equations are derived to predict the jitter and phase noise caused by white and flicker noise. Measurement results of a test chip fabricated in a standard CMOS 90 nm technology are presented to validate these expressions. The performed analysis shows that jitter in STSCL-based ring oscillator is independent of technology parameters, as opposed to its CMOS counterparts that depend on supply voltage and parameters of technology. Based on measured results, noise on current control line can dominate the total jitter of the oscillator. Design guidelines are proposed to limit the jitter effect of ring oscillators using STSCL logic. The proposed STSCL-based ring oscillator achieves an average RMS jitter as low as 0.24 % of the oscillation period at a 1.08 MHz/μA energy efficiency, which demonstrates its suitability for ultra-low-power applications. Mahsa Shoaran, Armin Tajalli, Massimo Alioto, Yusuf Leblebici |
ISCAS | 4 |
| 2015 | Design optimization of polyphase digital down converters for extremely high frequency wireless communicationsabstractIn this paper, an area-optimized polyphase digital down converter (DDC) architecture is introduced, where the mixers can be completely merged into the polyphase decimation filter under certain conditions. We also introduce an interface architecture, called synchronizer, between the back-end of an extremely high-speed time interleaved ADC (TI-ADC) and the front-end of a polyphase DDC. The synchronizer enables safe downsampling for a polyphase DDC, when the TI-ADC's sampling rate is above tens of GS/s. We show that the proposed interface architecture prevents any potential timing constraint violations that might occur in the interface between a TI-ADC and a polyphase DDC for extremely high frequency (EHF) wireless communication applications. Gain Kim, Raffaele Capoccia, Yusuf Leblebici |
VLSI-SoC | 3 |
| 2015 | Hardware implementation of real-time multiple frame super-resolutionabstractSuper-resolution reconstruction is a method for reconstructing higher resolution images from a set of low resolution observations. The sub-pixel differences among different observations of the same scene allow to create higher resolution images with better quality. In the last thirty years, many methods for creating high resolution images have been proposed. However, hardware implementations of such methods are limited. In this work, highly parallel and pipelined implementation for iterative back projection super-resolution algorithm is presented. The proposed hardware implementation is capable of reconstructing 512×512 sized images from set of 20 lower resolution observations, with real-time capabilities up to 25 frame per second (fps). Explained system has been synthesized and verified via Xilinx VC707 FPGAs. To the best of our knowledge, the system is currently the fastest super-resolution implementation based on FPGA. Kerem Seyid, Sebastien Blanc, Yusuf Leblebici |
VLSI-SoC | 3 |
| 2015 | A Real-Time Multiaperture Omnidirectional Visual Sensor Based on an Interconnected Network of Smart CamerasabstractCentralized and multilevel implementations of the Panoptic omnidirectional multiaperture visual system were previously presented by us, relying on the transmission of all camera outputs to a single central processing node for omnidirectional image and video reconstruction. In this paper, a novel distributed and parallel implementation of the omnidirectional vision reconstruction algorithm of the Panoptic system is presented. The parallel approach aims to overcome the scalability problems and memory bandwidth limitations of the centralized approach. The real-time hardware implementation is presented for camera modules with image processing, memory, and interconnectivity features. A methodology is introduced for the arrangement of camera modules with interconnectivity feature into a target interconnection network topology. A unique custom-made multiple-field-programmable gate array hardware platform is introduced for the implementation of an interconnected network of 49 camera prototype Panoptic system. A hardware architecture based on presented hardware platform enabling the real-time implementation of the blending algorithms is presented, along with the imaging results and resource utilization. The real-time implementation results of the implemented omnivision application on the mentioned prototype are demonstrated. Kerem Seyid, Vladan Popovic, Omer Cogal, Abdulkadir Akin, Hossein Afshari, Alexandre Schmid, Yusuf Leblebici |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2014 | An Immersive Telepresence System Using a Real-Time Omnidirectional Camera and a Virtual Reality Head-Mounted DisplayabstractCurrent telepresence systems are limited by the use of standard, narrow angle of view cameras. By using an Omni directional camera, an improved visual telepresence system is achieved. This work presents a more immersive visual experience using the Omni directional Panoptic camera in combination with a modern, low cost virtual reality head-mounted display. The camera system and its video data format are briefly described. The presented system is capable of broadcasting the Omni directional live stream via network connection. The system consists of the Panoptic camera as the acquisition device, a personal computer providing server support, and a head-mounted display connected to a client computer. The server and the client application that create a virtual environment from the video data are presented and the test setup is shown. Finally, a video example of the system's real-time operation is provided. Luis Manuel Gaemperle, Kerem Seyid, Vladan Popovic, Yusuf Leblebici |
ISM | 4 |
| 2014 | Reconfigurable forward homography estimation system for real-time applicationsabstractImage processing and computer vision algorithms extensively use projections, such as homography, as one of the processing steps. Systems for homography calculation usually observe homography as an inverse problem and provide an exact solution. However, the systems processing larger resolution images cannot meet inherently tight real-time constraints. Look-up table based systems provide an option for forward homography solutions, but they require large memory availability. Recent compressed look-up table methods reduce the memory requirements at the expense of lower peak signal-to-noise-ratio. In this work, we present a forward homography estimation algorithm which provides higher image quality than compressed look-up table methods. The algorithm is based on bounding the homography error, and neglecting the pixels out of the determined bound. The presented FPGA implementation of the estimation system requires a small amount of hardware, and no memory storage. The prototype system project an image frame onto a spherical surface at 295 Mpixels/s rate which is, up to our knowledge, currently the fastest homography system. Vladan Popovic, Yusuf Leblebici |
VLSI-SoC | 2 |
| 2014 | Dynamically adaptive real-time disparity estimation hardware using iterative refinement
Abdulkadir Akin, Ipek Baz, Alexandre Schmid, Yusuf Leblebici |
Integr. | 4 |
| 2013 | Towards structured ASICs using polarity-tunable Si nanowire transistorsabstractIn addition to scaling semiconductor devices down to their physical limit, novel devices show enhanced functionality compared to conventional CMOS. At advanced technology nodes, many devices exhibit ambipolar behavior, i.e., they show n- and p-type characteristics simultaneously. This phenomenon can be tamed using double-gate structures. In this paper, we present a complete framework relying on Double-Gate-all-around Vertically stacked NanoWire FETs (DG-NWFETs). Such device enables a compact realization of arithmetic logic functions and presents unprecedented interest for structured ASIC applications. Pierre-Emmanuel Gaillardon, Michele De Marchi, Luca G. Amarù, Shashikanth Bobba, Davide Sacchetto, Yusuf Leblebici, Giovanni De Micheli |
DAC | 6 |
| 2013 | Fast and accurate BER estimation methodology for I/O links based on extreme value theoryabstractThis paper introduces a novel approach towards the statistical analysis of modern high-speed I/O and similar communication links, which is capable of reliably to determine extremely low (∼10−12or lower) bit error rates (BER) by using techniques from extreme value theory (EVT). The new method requires only a small amount of voltage values at the received eye center, which can be generated by running circuit/system level simulations or measuring fabricated I/O circuits, to predict link BERs. Unlike conventional techniques, no simplifying assumptions on link noise and interference sources are required making this approach extremely portable to any communication system operating with very low BER. Our experimental results show that the BER estimates from the proposed methodology are on the same order of magnitude as traditional time domain, transient eye diagram simulations for links with BER of 10−6and 10−5operating at 9.6 and 10.1 Gbps respectively. Alessandro Cevrero, Nestoras E. Evmorfopoulos, Charalampos Antoniadis, Paolo Ienne, Yusuf Leblebici, Andreas Peter Burg, Georgios I. Stamoulis |
DATE | 5 |
| 2013 | Vertically-stacked double-gate nanowire FETs with controllable polarity: from devices to regular ASICsabstractVertically stacked nanowire FETs (NWFETs) with gate-all-around structure are the natural and most advanced extension of FinFETs. At advanced technology nodes, many devices exhibit ambipolar behavior, i.e., the device shows n- and p-type characteristics simultaneously. In this paper, we show that, by engineering of the contacts and by constructing independent double-gate structures, the device polarity can be electrostatically programmed to be either n- or p-type. Such a device enables a compact realization of XOR-based logic functions at the cost of a denser interconnect. To mitigate the added area/routing overhead caused by the additional gate, an approach for designing an efficient regular layout, called Sea-of-Tiles is presented. Then, specific logic synthesis techniques, supporting the higher expressive power provided by this technology, are introduced and used to showcase the performance of the controllable-polarity NWFETs circuits in comparison with traditional CMOS circuits. Pierre-Emmanuel Gaillardon, Luca G. Amarù, Shashikanth Bobba, Michele De Marchi, Davide Sacchetto, Yusuf Leblebici, Giovanni De Micheli |
DATE | 6 |
| 2013 | 3D-MMC: a modular 3D multi-core architecture with efficient resource poolingabstractThis paper demonstrates a fully functional hardware and software design for a 3D stacked multi-core system for the first time. Our 3D system is a low-power 3D Modular Multi-Core (3D-MMC) architecture built by vertically stacking identical layers. Each layer consists of cores, private and shared memory units, and communication infrastructures. The system uses shared memory communication and Through-Silicon-Vias (TSVs) to transfer data across layers. A serialization scheme is employed for inter-layer communication to minimize the overall number of TSVs. The proposed architecture has been implemented in HDL and verified on a test chip targeting an operating frequency of 400MHz with a vertical bandwidth of 3.2Gbps. The paper first evaluates the performance, power and temperature characteristics of the architecture using a set of software applications we have designed. We demonstrate quantitatively that the proposed modular 3D design improves upon the cost and performance bottlenecks of traditional 2D multi-core design. In addition, a novel resource pooling approach is introduced to efficiently manage the shared memory of the 3D stacked system. Our approach reduces the application execution time significantly compared to 2D and 3D systems with conventional memory sharing. Tiansheng Zhang, Alessandro Cevrero, Giulia Beanato, Panagiotis Athanasopoulos, Ayse K. Coskun, Yusuf Leblebici |
DATE | 6 |
| 2013 | A hardware-oriented dynamically adaptive disparity estimation algorithm and its real-time hardwareabstractThe computational complexity of disparity estimation algorithms and the need of large size and bandwidth for the external and internal memory make the real-time processing of disparity estimation challenging, especially for High Resolution (HR) images. This paper proposes a hardware-oriented adaptive window size disparity estimation (AWDE) algorithm and its real-time reconfigurable hardware implementation that targets HR video with high quality disparity results. The proposed algorithm is a hybrid solution involving the Sum of Absolute Differences and the Census cost computation methods to vote and select the best suitable disparity candidates. It utilizes a pixel intensity based refinement step to remove faulty disparity computations. The AWDE algorithm dynamically adapts the window size considering the local texture of the image to increase the disparity estimation quality. The proposed reconfigurable hardware of the AWDE algorithm enables handling 60 frames per second on Virtex-5 FPGA at a 1024×768 XGA video resolution for a 120 pixel disparity range.1 Abdulkadir Akin, Ipek Baz, Baris Atakan, Irem Boybat, Alexandre Schmid, Yusuf Leblebici |
ACM Great Lakes Symposium on VLSI | 6 |
| 2013 | High frame-rate low-power compressive sampling CMOS image sensor architecture: [extended abstract]abstractA novel compressive sampling scheme suitable for highly scalable hardware implementation is presented. The prototype design is implemented in a 0.18μm standard CMOS technology and utilizes compressed acquisition to achieve high frame rates and maintain low power consumption. Specialized pixels, convenient for Comparator-Based Switched Capacitor readout are developed for this purpose. A custom measurement matrix generation algorithm is implemented which reduces in-pixel hardware complexity and performs measurement matrix generation in a single clock cycle. Per-column Differential Cyclic-ADCs based on the Zero-Crossing Detection (ZCD) technique are used to convert the analog image measurements. Physical IC design issues such as the required dynamic range, device noise, mismatch and non-linearity, are analyzed and their effects on compressed image acquisition are presented and discussed. The final simulation results show that the proposed 256x256 pixels architecture consumes 1.45mW at 250fps and 26.2mW at 8000fps. The proposed architecture can easily be scaled towards newer technology nodes and higher image resolutions. Nikola Katic, Mahdad Hosseini Kamal, Mustafa Kilic, Alexandre Schmid, Pierre Vandergheynst, Yusuf Leblebici |
ACM Great Lakes Symposium on VLSI | 6 |
| 2013 | Compressive multichannel cortical signal recordingabstractThis paper presents a novel approach to acquire multichannel wireless intracranial neural data based on a compressive sensing scheme. The designed circuits are extremely compact and low-power which confirms the relevance of the proposed approach for multichannel high-density neural interfaces. The proposed compression model enables the acquisition system to record from a large number of channels by reducing the transmission power per channel. Our main contributions are the twofold. First, a CMOS compressive sensing system to realize multichannel intracranial neural recording is described. Second, we explain a joint sparse decoding algorithm to recover the multichannel neural data. The idea has been implemented at system as well as circuit levels. The simulation results reveal that the multichannel intracranial neural data can be acquired by compression ratios as high as four. Mahdad Hosseini Kamal, Mahsa Shoaran, Yusuf Leblebici, Alexandre Schmid, Pierre Vandergheynst |
ICASSP | 3 |
| 2013 | Real-time hardware implementation of multi-resolution image blendingabstractA novel real-time implementation of a multi-resolution image blending algorithm is presented in this paper. A multi-resolution decomposition of the input is used to blend multiple images at different scales. Processing time is shortened by designing a pipeline system. The proposed solution requires less hardware multipliers and is able to achieve very high operating frequencies, compared to the current designs. The presented hardware architecture is optimized to support multiple simultaneous video streams, and high frame rates at High-Definition (HD) resolutions. Vladan Popovic, Kerem Seyid, Alexandre Schmid, Yusuf Leblebici |
ICASSP | 4 |
| 2013 | Characterization of standard CMOS compatible photodiodes and pixels for Lab-on-Chip devicesabstractHigh quality CMOS image sensors are of great importance for LoC - Lab-on-Chip devices based on optical measurements. The main target in these devices is to minimize the cost and area while achieving a good resolution. The performance parameters of image sensor pixels and CMOS compatible photodiodes depend on the size, type and the geometry of the photodiode layout and varies for each technology. In this study, we present a comparative analysis of CMOS compatible photodiode types at different areas. The results have shown n-well/p-sub type photodiode with 5×5 μm2diffusion area achieves the highest sensitivity (69.81 × 1012V.s-1.cm-2/W.cm-2) and with 40 × 40 μm2diffusion area, highest SNR - Signal-to-Noise Ratio (72.26dB) at 630 nm, while the p+/n-well/p-sub type photodiode with 40 × 40 μm2diffusion area results in highest responsivity (0.466 A. cm-2/W.cm-2) at the same wavelength. Gozen Koklu, Ralph Etienne-Cummings, Yusuf Leblebici, Giovanni De Micheli, Sandro Carrara |
ISCAS | 3 |
| 2013 | A low-power area-efficient compressive sensing approach for multi-channel neural recordingabstractHigh-density wireless intracranial neural recording is a promising technology enabling the autonomous diagnosis and therapy of brain diseases. Increasing the number of recording channels is accompanied by the increased amount of data resulting in an unacceptable transmission power. A comprehensive study of possible compressed sensing methods in the context of neural signals has been done, and the compression of signals originating from different channels in the spatial domain has been implemented at the system and circuit levels. Results of the simulations in a UMC 0.18μm CMOS technology and subsequent reconstructions show the possibility of compressing with ratios as high as 2.6 with a recovery SNR of at least 10dB using extremely compact and low-power circuits. The power efficiency and limited area per channel confirm the relevance of the proposed approach for multi-channel high-density neural interfaces. Mahsa Shoaran, Mariazel Maqueda Lopez, Vijaya Sankara Rao Pasupureddi, Yusuf Leblebici, Alexandre Schmid |
ISCAS | 4 |
| 2013 | Spherical Panorama Construction Using Multi Sensor Registration Priors and Its Real-Time HardwareabstractIn this work, a novel method is presented to improve the quality of panoramic images on a spherically arranged multi sensor imaging system. The new method is composed of two parts. The first approach proposed is based on mapping the panorama generation problem onto a Markov Random Field (MRF) and then estimating posterior probabilities from initial likelihoods. The novelty of approach is based on extracting the prior evidence from the registration information of multiple cameras and estimating expected value on an undirected graph. The second part of the method is a geometrical approach targeting a better estimation for the initial priors, which is also not applied before. The aim of both approaches is to decrease the parallax errors and ghosting effects which occur due to the nature of multi camera systems. It is shown that instead of directly using independent intensity coefficients extracted from registration information, applying a neighborhood based local probability distribution for each pixel of panorama utilizing the registration information as prior gives better results. Visual comparisons are provided to show the achieved quality enhancement in terms of seamless and more natural panoramic image with less ghosting effects. Since the registration priors are used effectively with a single iteration step in a 4 connected neighborhood, the need for an intensity based loopy and iterative inference method is prohibited. Hence, the proposed methods are suitable for real-time hardware implementation. A hardware implementation of the method for real-time operation is proposed. Omer Cogal, Vladan Popovic, Yusuf Leblebici |
ISM | 3 |
| 2013 | Compressed look-up-table based real-time rectification hardwareabstractStereo image rectification is a pre-processing step of disparity estimation intended to remove image distortions and to enable stereo matching along an epipolar line. A real-time disparity estimation system needs to perform real-time rectification which requires solving the models of lens distortions, image translations and rotations. Look-up-table based rectification algorithms allow image rectification without demanding high complexity operations. However, they require an external memory to store large size look-up-tables. In this work, we present an intermediate solution that compresses the rectification information to fit the look-up-table into the on-chip memory of a Virtex-5 FPGA. The low-complexity decompression process requires a negligible amount of hardware resources for its real-time implementation. The proposed image rectification hardware consumes 0.28% of the DFF and 0.32% of the LUT resources of the Virtex-5 XCUVP-110T FPGA, it can process 347 frames per second for a 1024×768 pixels image resolution, and it does not need the availability of an external memory. Abdulkadir Akin, Ipek Baz, Luis Manuel Gaemperle, Alexandre Schmid, Yusuf Leblebici |
VLSI-SoC | 5 |
| 2012 | Physical synthesis onto a Sea-of-Tiles with double-gate silicon nanowire transistorsabstractWe have designed and fabricated double-gate ambipolar field-effect transistors, which exhibit p-type and n-type characteristics by controlling the polarity of the second gate. In this work, we present an approach for designing an efficient regular layout, called Sea-of-Tiles (SoTs). First, we address gate-level routing congestion by proposing compact layout techniques and novel symbolic-layout styles. Second, we design four logic tiles, which form the basic building block of the SoT fabric. We run extensive comparisons of mapping standard benchmarks on the SoT. Our study shows that SoT with TileG2 and TileG1h2, on an average, outperforms the one with TileG1 and TileG3 by 16% and 10% in area utilization, respectively. Shashikanth Bobba, Michele De Marchi, Yusuf Leblebici, Giovanni De Micheli |
DAC | 3 |
| 2012 | Design and Implementation of Multi-camera Systems Distributed over a Spherical Geometry
Hossein Afshari, Kerem Seyid, Alexandre Schmid, Yusuf Leblebici |
Diagrams | 4 |
| 2012 | Enhanced Omnidirectional Image Reconstruction Algorithm and Its Real-Time HardwareabstractOmnidirectional stereoscopy and depth estimation are complex problems of image processing to which the Panoptic camera offers a novel solution. The Panoptic camera is a biologically-inspired vision sensor made of multiple cameras. It is a polydioptric system mimicking the eyes of flying insects where multiple imagers, each with a distinct focal point, are distributed over a hemisphere. Recently, the omnidirectional image reconstruction algorithm (OIR) and its real-time hardware implementation have been proposed for the Panoptic camera. This paper presents an enhanced omnidirectional image reconstruction algorithm (EOIR) and its real-time implementation. The proposed EOIR algorithm provides improved realistic omnidirectional images and residuals compared to OIR. As a processing core of EOIR, 57% of the available slice resources in a Virtex 5 FPGA are consumed. The proposed platform provides the high bandwidth required to simultaneously process data originating from 40 cameras, and reconstruct omnidirectional images of 256x1024 pixels at 25 fps. This proposed hardware and algorithmic enhancements enable advanced real-time applications including omnidirectional image reconstruction, 3D model construction and depth estimation. Abdulkadir Akin, Elif Erdede, Hossein Afshari, Alexandre Schmid, Yusuf Leblebici |
DSD | 5 |
| 2012 | Quantitative comparison of commercial CCD and custom-designed CMOS camera for biological applicationsabstractIn biological applications and systems where even the smallest details have a meaning, CCD cameras are mostly preferred and they hold most of the market share despite their high costs. In this paper, we propose a custom-designed CMOS camera to compete with the default CCD camera of an inverted microscope for fluorescence imaging. The custom-designed camera includes a commercially available mid-performance CMOS image sensor and a Field-Programmable Gate Array (FPGA) based hardware platform (FPGA4U). The high cost CCD camera of the microscope is replaced by the custom-designed CMOS camera and the two are quantitatively compared for a specific application where an Estrogen Reception (ER) expression in breast cancer diagnostic samples that emits light at 665nm has been imaged by both cameras. The gray-scale images collected by both cameras show a very similar intensity distribution. In addition, normalized white pixels after thresholding resulted in 4.96% for CCD and 3.38% for CMOS. The results and images after thresholding show that depending on the application even a mid-performance CMOS camera can provide enough image quality when the target is localization of fluorescent stained biological details. Therefore the cost of the cameras can be drastically reduced while benefiting from the inherent advantages of CMOS devices plus adding more features and flexibility to the camera systems with FPGAs. Gozen Koklu, Julien Ghaye, Rene Beuchat, Giovanni De Micheli, Yusuf Leblebici, Sandro Carrara |
ISCAS | 5 |
| 2012 | 3D-LIN: A configurable low-latency interconnect for multi-core clusters with 3D stacked L1 memoryabstractAbstract—Shared L1 memories are of interest for tightlycoupled processor clusters in programmable accelerators as they provide a convenient shared memory abstraction while avoiding cache coherence overheads. The performance of a shared-L1 memory critically depends on the architecture of the low-latency interconnect between processors and memory banks, which needs to provide ultra-fast access to the largest possible L1 working set. The advent of 3D technology provides new opportunities to improve the interconnect delay and the form factor. In this paper we propose a network architecture, 3D-LIN, based on 3D integration technology. The network can be configured based on user specifications and technology constraints to provide fast access to L1 memories on multiple stacked dies. The extracted results from the physical synthesis of 3D-LIN permit to explore trade-offs between memory size and network latency from a planar design to multiple memory layers stacked on top of logic. In the case where the system memory requirements lead to a memory area that occupies 60 % of the chip, the form factor can be reduced by more than 60 % by stacking 2 memory layers on the logic. Latency reduction is also promising: the network itself, configured for connecting 16 processing elements to 128 memory banks on 2 memory layers is 24 % faster than the planar system. I. Giulia Beanato, Igor Loi, Giovanni De Micheli, Yusuf Leblebici, Luca Benini |
VLSI-SoC | 4 |
| 2012 | GMS: Generic memristive structure for non-volatile FPGAsabstractThe invention of the memristor enables new possibilities for computation and non-volatile memory storage. In this paper we propose a Generic Memristive Structure (GMS) for 3-D FPGA applications. The GMS cell is demonstrated to be utilized for steering logic useful for multiplexing signals, thus replacing the traditional pass-gates in FPGAs. Moreover, the same GMS cell can be utilized for programmable memories as a replacement for the SRAMs employed in the look-up tables of FPGAs. A fabricated GMS cell is presented and its use in FPGA architecture is demonstrated by the area and delay improvement for several architectural benchmarks. Pierre-Emmanuel Gaillardon, Davide Sacchetto, Shashikanth Bobba, Yusuf Leblebici, Giovanni De Micheli |
VLSI-SoC | 4 |
| 2012 | Multiterminal Memristive Nanowire Devices for Logic and Memory Applications: A ReviewabstractMemristive devices have the potential for a complete renewal of the electron devices landscape, including memory, logic, and sensing applications. This is especially true when considering that the memristive functionality is not limited to two-terminal devices, whose practical realization has been demonstrated within a broad range of different technologies. For electron devices, the memristive functionality can be generally attributed to a material state modification, whose dynamics can be engineered to target a specific application. In this review paper, we show that trap charging dynamics can explain some of the memristive effects previously reported for Schottky-barrier field-effect Si nanowire transistors (SB SiNW FETs). Moreover, the SB SiNW FETs do show additional memristive functionality due to trap charging at the metal/semiconductor surface. The combination of these two memristive effects into multiterminal metal-oxide-semiconductor field-effect transistor (MOSFET) devices gives rise to new opportunities for both memory and logic applications as well as new sensors based on the physical mechanism that originate memristance. In the special case of four-terminal memristive Si nanowire devices, which are presented for the first time in this paper, enhanced functionality is demonstrated. Finally, the multiterminal memristive devices presented here have the potential of a very high integration density, and they are suitable for hybrid complementary metal-oxide-semiconductor (CMOS) cofabrication with a CMOS-compatible process. Davide Sacchetto, Giovanni De Micheli, Yusuf Leblebici |
Proc. IEEE | 3 |
| 2011 | Power-gated MOS current mode logic (PG-MCML): a power aware DPA-resistant standard cell libraryabstractMOS Current Mode Logic (MCML) is one of the most promising logic style to counteract power analysis attacks. Unfortunately, the static power consumption of MCML standard cells is significantly higher compared to equivalent functions implemented using static CMOS logic. As a result, the use of such a logic style is very limited in portable devices. Paradoxically, these devices are the most sensitive to physical attacks, thus the ones which would benefit more from the adoption of MCML. Alessandro Cevrero, Francesco Regazzoni 0001, Micheal Schwander, Stéphane Badel, Paolo Ienne, Yusuf Leblebici |
DAC | 6 |
| 2011 | Towards thermally-aware design of 3D MPSoCs with inter-tier coolingabstractNew tendencies envisage 3D Multi-Processor System-On-Chip (MPSoC) design as a promising solution to keep increasing the performance of the next-generation high-performance computing (HPC) systems. However, as the power density of HPC systems increases with the arrival of 3D MPSoCs, supplying electrical power to the computing equipment and constantly removing the generated heat is rapidly becoming the dominant cost in any HPC facility. Thus, both power and thermal/cooling implications play a major role in the design of new HPC systems, given the energy constraints in our society. Therefore, EPFL, IBM and ETHZ have been working within the CMOSAIC Nano-Tera.ch program project in the last three years on the development of a holistic thermally-aware design. This paper presents the exploration in CMOSAIC of novel cooling technologies, as well as suitable thermal modeling and system-level design methods, which are all necessary to develop 3D MPSoCs with inter-tier liquid cooling systems. As a result, we develop energy-efficient run-time thermal control strategies to achieve energy-efficient cooling mechanisms to compress almost 1 Tera nano-sized functional units into one cubic centimeter with a 10 to 100 fold higher connectivity than otherwise possible. The proposed thermally-aware design paradigm includes exploring the synergies of hardware-, software- and mechanical-based thermal control techniques as a fundamental step to design 3D MPSoCs for HPC systems. More precisely, we target the use of inter-tier coolants ranging from liquid water and two-phase refrigerants to novel engineered environmentally friendly nano-fluids, as well as using specifically designed micro-channel arrangements, in combination with the use of dynamic thermal management at system-level to tune the flow rate of the coolant in each micro-channel to achieve thermally-balanced 3D-ICs. Our management strategy prevents the system from surpassing the given threshold temperature while achieving up to 67% reduction in cooling energy and up to 30% reduction in system-level energy in comparison to setting the flow rate at the maximum value to handle the worst-case temperature. Mohamed M. Sabry, Arvind Sridhar, David Atienza 0001, Yuksel Temiz, Yusuf Leblebici, S. Szczukiewicz, Navid Borhani, John Richard Thome, Thomas Brunschwiler, Bruno Michel |
DATE | 5 |
| 2011 | Hardware security in VLSIabstractThis special session addresses the increasingly critical area of hardware security in VLSI. As computing becomes ubiquitous, most systems need to implement various layers of security, all of which rely on the security of the underlying hardware. At the same time, persistent trends in VLSI are producing challenges and opportunities in the design of security mechanisms. Increased levels of integration, performance and power efficiency, allow the implementation of strong security protocols, however challenges arise due to the emergence of side-channel attacks and the ability to insert hardware Trojans that are difficult to detect. But on the other side, imperfections in the manufacturing process and on-chip noise can be leveraged to implement security primitives such as unique IDs and random numbers. The lessons learned from Hardware Security in advanced CMOS also have more general application. Statistical design and variation- and noise-aware methodologies are all current hot topics in VLSI that share themes with hardware security. Wayne P. Burleson, Yusuf Leblebici |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | Alternative design methodologies for the next generation logic switchabstractNext generation logic switch devices are expected to rely on radically new technologies mainly due to the increasing difficulties and limitations of state-of-the-art CMOS switches, which, in turn, will also require innovative design methodologies that are distinctly different from those used for CMOS technologies. In this paper, three alternative emerging technologies are showcased in terms of their requirements for design implementation and in terms of potential advantages. First, a CMOS evolutionary approach based on vertically-stacked gate-all-around Si nanowire FETs is discussed. Next, an alternative design methodology based on ambipolar carbon nanotube FETs is presented. Finally, a novel approach based on the recently discovered memristive devices is presented, offering the possibility of combining memory and logic functions. Davide Sacchetto, Michele De Marchi, Giovanni De Micheli, Yusuf Leblebici |
ICCAD | 4 |
| 2011 | Area, throughput, and energy-efficiency trade-offs in the VLSI implementation of LDPC decodersabstractLow-density parity-check (LDPC) codes are key ingredients for improving reliability of modern communication systems and storage devices. On the implementation side however, the design of energy-efficient and high-speed LDPC decoders with a sufficient degree of reconfigurability to meet the flexibility demands of recent standards remains challenging. This survey paper provides an overview of the state-of-the-art in the design of LDPC decoders using digital integrated circuits. To this end, we summarize available algorithms and characterize the design space. We analyze the different architectures and their connection to different codes and requirements. The advantages and disadvantages of the various choices are illustrated by comparing state-of-the-art LDPC decoder designs. Christoph Roth, Alessandro Cevrero, Christoph Studer, Yusuf Leblebici, Andreas Peter Burg |
ISCAS | 4 |
| 2010 | AVGS-Mux style: A novel technology and device independent technique for reducing power and compensating process variations in FPGA fabricsabstractThis work presents Adaptive Vgs Multiplexer (AVGS-Mux) Technique. Proposed method controls the transistor current by the source voltage. It can provide ±1.6X control on the delay and ±7X exponential control on sub-threshold and gate leakages in the switch-box, LUT, and interconnects. For equal leakage, it improves the speed 9%, reduces dynamic power 13%, and reduces random dopant fluctuations effect. AVGS-Mux is a good replacement of adaptive body biasing and adaptive supply voltage techniques in emerging Multi-Gate devices which have very small body effect and cannot tolerate voltages higher than nominal VDD due to reliability issues. Bahman Kheradmand Boroujeni, Christian Piguet, Yusuf Leblebici |
DATE | 3 |
| 2010 | Ultra-low power mixed-signal design platform using subthreshold source-coupled circuitsabstractThis article discusses system-level techniques to optimize the power-performance trade-off in subthreshold circuits and presents a uniform platform for implementing ultra-low power power-scalable analog and digital integrated circuits. The proposed technique is based on using subthreshold source-coupled or current-mode approach for both analog and digital circuits. In addition to possibility of operating with ultra-low power dissipation, because of similar basis for constructing analog and digital parts, a common power management unit could be used for optimizing the power-performance of the entire mixed-signal system. Some circuit examples have been provided to show the performance of the proposed circuits in practice. Armin Tajalli, Yusuf Leblebici |
DATE | 2 |
| 2010 | A (256×256) pixel 76.7mW CMOS imager/ compressor based on real-time In-pixel compressive sensingabstractA CMOS imager is presented which has the ability to perform localized compressive sensing on-chip. In-pixel convolutions of the sensed image with measurement matrices are computed in real time, and a proposed programmable two-dimensional scrambling technique guarantees the randomness of the coefficients used in successive observation. A power and area-efficient implementation architecture is presented making use of a single ADC. A 256×256 imager has been developed as a test vehicle in a 0.18μm CIS technology. Using an 11-bit ADC, a SNR of 18.6dB with a compression factor of 3.3 is achieved after reconstruction. The total power consumption of the imager is simulated at 76.7mW from a 1.8V supply voltage. Vahid Majidzadeh, Laurent Jacques, Alexandre Schmid, Pierre Vandergheynst, Yusuf Leblebici |
ISCAS | 5 |
| 2010 | Memristive devices fabricated with silicon nanowire schottky barrier transistorsabstractThis paper reports on the memory and memristive effects of Schottky barrier field effect transistors (SBFET) with gate-all-around (GAA) configuration and Si nanowire (SiNW) channel. Similar behavior has also been investigated for SBFETs with poly-Si nanowire (poly-SiNW) channel in back-gate configuration. The memristive devices presented here have the potential of a very high integration density, and they are suitable for hybrid CMOS co-fabrication with a CMOS-compatible process. We show that 2 different regimes are possible, making these devices suitable either for volatile ambipolar memory or resistive random access memory (RRAM) applications. In addition, frequency- and amplitude- dependence of the memristive behavior are reported. Davide Sacchetto, M. Haykel Ben Jamaa, Sandro Carrara, Giovanni De Micheli, Yusuf Leblebici |
ISCAS | 5 |
| 2010 | Design aspects of carry lookahead adders with vertically-stacked nanowire transistorsabstractThis paper discusses the newly introduced vertically-stacked silicon nanowire gate-all-around field-effect-transistor technology and its advantages for higher density layout design. The vertical nanowire stacking technology allows very-high density arrangement of nanowire transistors with near-ideal characteristics, and opens the possibility for design optimization by adjusting the number of nanowire stacks without affecting the footprint area of the device. Several libraries for combinational logic synthesis have been designed and implemented for the synthesis of carry-lookahead adders, using the vertically-stacked nanowire technology. The reduction in silicon active area occupancy of vertically-stacked gates are envisaged of great significance for regular cell mapping, in disruptive future applications based on nanowire transistor arrays. Davide Sacchetto, M. Haykel Ben Jamaa, Giovanni De Micheli, Yusuf Leblebici |
ISCAS | 4 |
| 2010 | Selective redundancy-based design techniques for the minimization of local delay variationsabstractIn this paper a novel approach to optimize digital integrated circuits yield with regards to speed and area/power for aggressive scaling technologies is presented. The technique is intended to reduce the effects of intra-die variations using redundancy applied only on critical parts of the circuit. The inherent property of the technique is that the improvement in the maximum frequency the circuit can run is higher for the larger variations. The work shows that the technique can be already applied for 65nm CMOS technology process where a beneficial delay vs. area/power tradeoff can be made. However, a significant benefit is expected for future nanoscale CMOS technologies such as 45nm and 32nm nodes and in low-voltage applications. Milos Stanisavljevic, Alexandre Schmid, Yusuf Leblebici |
ISCAS | 3 |
| 2010 | Output probability density functions of logic circuits: Modeling and fault-tolerance evaluationabstractThe precise evaluation of the reliability of logic circuits has a significant importance in highly-defective and future nanotechnologies. It allows efficient comparison of fault-tolerance techniques, and enables designs improvement with respect to their reliability figure. This paper presents a novel, accurate and scalable method for modeling the output probability density functions (PDFs) of logic circuits. Our method combines probability theory with concepts from logic synthesis and testing. The PDFs are modeled using the acquired circuit output probability of failure and PDFs of gates in the last two layers of the output cone. Unlike the existing output PDF modeling techniques, the proposed method is directly applicable to standard CMOS design. Simulation results of benchmark circuits demonstrate the accuracy of the method. Several potential applications of the proposed technique include the analysis of averaging (analog) fault-tolerant techniques, fine-grained redundancy insertion, and reliability-driven design optimization. Milos Stanisavljevic, Alexandre Schmid, Yusuf Leblebici |
VLSI-SoC | 3 |
| 2010 | Design and feasibility of multi-Gb/s quasi-serial vertical interconnects based on TSVs for 3D ICsabstractThis paper proposes a novel technique to exploit the high bandwidth offered by through silicon vias (TSVs). In the proposed approach, synchronous parallel 3D links are replaced by serialized links to save silicon area and increase yield. Detailed analysis conducted in 90 nm CMOS technology shows that the proposed 2-Gb/s/pin quasi-serial link requires approximately five times less area than its parallel bus equivalent at same data rate for a TSV diameter of 20 µm. Fengda Sun, Alessandro Cevrero, Panagiotis Athanasopoulos, Yusuf Leblebici |
VLSI-SoC | 4 |
| 2010 | Efficient and side-channel-aware implementations of elliptic curve cryptosystems over prime fieldsabstractElliptic curve cryptosystems (ECCs) are utilised as an alternative to traditional public-key cryptosystems, and are more suitable for resource-limited environments because of smaller parameter size. In this study, the authors carry out a thorough investigation of side-channel attack aware ECC implementations over finite fields of prime characteristic including the recently introduced Edwards formulation of elliptic curves. The Edwards formulation of elliptic curves is promising in performance with built-in resiliency against simple side-channel attacks. To our knowledge the authors present the first hardware implementation for the Edwards formulation of elliptic curves. The authors also propose a technique to apply non-adjacent form (NAF) scalar multiplication algorithm with side-channel security using the Edwards formulation. In addition, the authors implement Joye's highly regular add-always scalar multiplication algorithm both with the Weierstrass and Edwards formulation of elliptic curves. Our results show that the Edwards formulation allows increased area–time performance with projective coordinates. However, the Weierstrass formulation with affine coordinates results in the simplest architecture, and therefore has the best area–time performance as long as an efficient modular divider is available. Deniz Karakoyunlu, Frank K. Gürkaynak, Berk Sunar, Yusuf Leblebici |
IET Inf. Secur. | 4 |
| 2009 | A stochastic perturbative approach to design a defect-aware thresholder in the sense amplifier of crossbar memoriesabstractThe use of nanowire crossbars to build devices with large storage capabilities is a very promising architectural paradigm for forthcoming nanoscale memory devices. However, this new type of memory devices raises questions regarding how to test their correct operation. In particular, the variability affecting the decoder is expected to make very complex the test of these new devices. In this paper we present a method to simplify the test of these new devices by using a current thresholder to detect badly addressed nanowires. In the proposed method, the thresholder design is based on a stochastic and perturbative model of the current through the nanowires. Thus, the calculated thresholder parameters are robust against technology variation. As our experimental results indicate, the thresholder error probability is initially only ~ 10-4, which can be also reduced further (up to ~ 60×) by trading-off only ~ 35% area overhead in the memory. M. Haykel Ben Jamaa, David Atienza 0001, Yusuf Leblebici, Giovanni De Micheli |
ASP-DAC | 3 |
| 2009 | Complete nanowire crossbar framework optimized for the multi-spacer patterning techniqueabstractNanowire crossbar circuits are an emerging architectural paradigm that promises a higher integration density and an improved fault-tolerance due to its reconfigurability. In this paper, we propose for the first time the utilization of the multi-spacer patterning technique to fabricate nanowire crossbars with a high cross-point density up to 1010 cm−2. We propose a novel decoder fabrication method that can be included in a process dedicated to the multi-spacer patterning technique. We address the technology problems consisting in the variability and fabrication complexity at the design level by optimizing the encoding scheme. We show an overall reduction of the variability by 18% and a cancelation of the fabrication complexity overhead. M. Haykel Ben Jamaa, Gianfranco Cerofolini, Yusuf Leblebici, Giovanni De Micheli |
CASES | 3 |
| 2009 | A Design Flow and Evaluation Framework for DPA-Resistant Instruction Set Extensions
Francesco Regazzoni 0001, Alessandro Cevrero, François-Xavier Standaert, Stéphane Badel, Theo Kluter, Philip Brisk, Yusuf Leblebici, Paolo Ienne |
CHES | 7 |
| 2009 | Decoding nanowire arrays fabricated with the multi-spacer patterning techniqueabstractSilicon nanowires are a promising solution to address the increasing challenges of fabrication and design at the future nodes of the Complementary Metal-Oxide-Semiconductor (CMOS) Technology roadmap. Despite the attractive opportunity that offers their organization onto regular crossbars, the problem of designing the nano-wire decoder is still challenging and highly dependent on the nanowire fabrication technology. In this paper, we introduce a novel design style and encoding scheme for decoding nanowires fabricated with the Multi-Spacer-Patterning Technique (MSPT); and we present a method based on Gray codes that reduces the fabrication cost and improves the decoder reliability. We show that by arranging the code in a Gray code fashion, we decrease the fabrication complexity by 17% and the variability by 18% on average. By optimizing the decoder parameters, the simulations showed an improvement of the crossbar yield by 40% and a reduction of the effective bit area by 51% to 169 nm2. M. Haykel Ben Jamaa, Yusuf Leblebici, Giovanni De Micheli |
DAC | 2 |
| 2009 | Dynamic thermal management in 3D multicore architecturesabstractTechnology scaling has caused the feature sizes to shrink continuously, whereas interconnects, unlike transistors, have not followed the same trend. Designing 3D stack architectures is a recently proposed approach to overcome the power consumption and delay problems associated with the interconnects by reducing the length of the wires going across the chip. However, 3D integration introduces serious thermal challenges due to the high power density resulting from placing computational units on top of each other. In this work, we first investigate how the existing thermal management, power management and job scheduling policies affect the thermal behavior in 3D chips. We then propose a dynamic thermally-aware job scheduling technique for 3D systems to reduce the thermal problems at very low performance cost. Our approach can also be integrated with power management policies to reduce energy consumption while avoiding the thermal hot spots and large temperature variations. Ayse K. Coskun, José Luis Ayala, David Atienza 0001, Tajana Rosing, Yusuf Leblebici |
DATE | 5 |
| 2009 | 3D configuration caching for 2D FPGAsabstractThis poster proposes the use of 3D integration technology to enable low-overhead reconfigurable computing. In our scheme, a 64 Megabyte DRAM array is stacked on top of an FPGA using face-to-face bonding, and caches up to 289 future configurations which can be quickly loaded onto the FPGA. Past DRAMs have been designed for off-chip communication, a bottleneck that 3D stacking eliminates; hence, the DRAM array is redesigned. To reconfigure the FPGA, a configuration is read from the DRAM into a latch array while the FPGA executes; then, the configuration is loaded from the latch array into the FPGA in 5 cycles (60ns). The minimum latency between reconfigurations, 8.42s, is dominated by the time to load data from the DRAM into the latch array. The benefits, area cost, and performance of the proposed system are evaluated on three previously published FPGA implementations of multimedia applications: MP3 and MPEG-4 decoders, and JPEG compression, and are evaluated under three scenarios: No Dynamic ReConfiguration (NDRC), Off-chip Dynamic ReConfiguration (ORDC), and 3D Configuration Caching (3DCC). Our experiments demonstrate that 3D configuration caching works best when used in conjunction with FPGA-based accelerators, rather than pure FPGA-based systems; in these systems, the reconfiguration latency can easily be hidden behind software execution on the processor controlling the accelerator. This significantly reduces the amount of silicon area that must be dedicated to the accelerator, while imposing virtually no performance penalty compared to significantly larger accelerators that do not require reconfiguration. Alessandro Cevrero, Panagiotis Athanasopoulos, Hadi Parandeh-Afshar, Philip Brisk, Yusuf Leblebici, Paolo Ienne, Maurizio Skerlj |
FPGA | 5 |
| 2009 | Using 3D integration technology to realize multi-context FPGAsabstractThis paper advocates the use of 3D integration technology to stack a DRAM on top of an FPGA. The DRAM will store future FPGA contexts. A configuration is read from the DRAM into a latch array on the DRAM layer while the FPGA executes; the new configuration is loaded from the latch array into the FPGA in 60 ns (5 cycles). The latency between reconfigurations, 8.42 mus, is dominated by the time to read data from the DRAM into the latch array. We estimate that the DRAM can cache 289 FPGA contexts. Alessandro Cevrero, Panagiotis Athanasopoulos, Hadi Parandeh-Afshar, Maurizio Skerlj, Philip Brisk, Yusuf Leblebici, Paolo Ienne |
FPL | 6 |
| 2009 | A flexible DSP block to enhance FPGA arithmetic performanceabstractWe propose a new DSP block for use in modern high-performance FPGAs. Current DSP blocks contain fixed-bitwidth multipliers that can be combined efficiently to form larger multipliers. Our approach is similar, but includes a bypass layer following the partial product generator that exposes the compressor tree used for partial product reduction directly to the user. As a consequence, the proposed DSP block can accelerate multi-input addition operations in addition to multiplication. To increase the flexibility of the device, the partial product reduction tree used within our DSP block uses a fixed-function compression logic along with a field programmable compressor tree (FPCT), the latter of which is user-configurable to meet the needs of the application at hand. Multi-input addition operations can be mapped directly onto the FPCT without compromising any of the other functionality of the DSP block. Hadi Parandeh-Afshar, Alessandro Cevrero, Panagiotis Athanasopoulos, Philip Brisk, Yusuf Leblebici, Paolo Ienne |
FPT | 5 |
| 2009 | CMOS compressed imaging by Random ConvolutionabstractWe present a CMOS imager with built-in capability to perform Compressed Sensing coding by Random Convolution. It is achieved by a shift register set in a pseudo-random configuration. It acts as a convolutive filter on the imager focal plane, the current issued from each CMOS pixel undergoing a pseudo-random redirection controlled by each component of the filter sequence. A pseudo-random triggering of the ADC reading is finally applied to complete the acquisition model. The feasibility of the imager and its robustness under noise and non-linearities have been confirmed by computer simulations, as well as the reconstruction tools supporting the Compressed Sensing theory. Laurent Jacques, Pierre Vandergheynst, Alexandre Bibet, Vahid Majidzadeh, Alexandre Schmid, Yusuf Leblebici |
ICASSP | 6 |
| 2009 | Memory organization and data layout for instruction set extensions with architecturally visible storageabstractPresent application specific embedded systems tend to choose instruction set extensions (ISEs) based on limitations imposed by the available data bandwidth to custom functional units (CFUs). Adoption of the optimal ISE for an application would, in many cases, impose formidable cost increase in order to achieve the required data bandwidth. In this paper we propose a novel methodology for laying out data in memories, generating high-bandwidth memory systems by making use of existing low-bandwidth low-cost ones and designing custom functional units all with the desirable data bandwidth for only a fraction of the additional cost required by traditional techniques. Panagiotis Athanasopoulos, Philip Brisk, Yusuf Leblebici, Paolo Ienne |
ICCAD | 3 |
| 2009 | A pulse-density modulation circuit exhibiting noise shaping with single-electron neuronsabstractWe propose a bio-inspired circuit performing pulse-density modulation with single-electron devices. The proposed circuit consists of three single-electron neuronal units, receiving the same input and are connected to a common output. The output is inhibitorily fedback to the three neuronal circuits through a capacitive coupling, tuned to obtain a winners-share-all network operation. The circuit performance was evaluated through Monte-Carlo based computer simulations. We demonstrated that the proposed circuit possesses noise-shaping characteristics, where signal and noises are separated into low and high frequency bands respectively. This significantly improved the signal-to-noise ratio (SNR) by 4.34 dB in the coupled network, as compared to the uncoupled one. The noise-shaping properties are as a result of i) the inhibitory feedback between the output and the neuronal circuits, and ii) static noises (originating from device fabrication mismatches) and dynamic noises (as a result of thermally induced random tunneling events) introduced into the network. Andrew Kilinga Kikombo, Tetsuya Asai, Takahide Oya, Alexandre Schmid, Yusuf Leblebici, Yoshihito Amemiya |
IJCNN | 5 |
| 2009 | Optimization of Wire Grid Size for Differential Routing and Impact on the Power-delay-area TradeoffabstractIn this paper, the impact of the wire grid size on the power-delay-area tradeoff of VLSI digital circuits with differential routing is analyzed. To this aim, the differential MOS current-mode logic (MCML) is adopted as reference logic style, and a complete differential design flow is used. Analysis shows that the choice of the grid size in differential routing has a much stronger impact on the power-delay-area tradeoff, compared to the usual single-ended case, hence the grid size must be carefully selected. The dependence of power, delay and area on the grid size is discussed in detail through simple models and metrics. To validate the approach and show basic dependencies in practical circuits, 30 benchmark circuits with an in-house designed MCML cell library were synthesized and routed in a 0.18-mum CMOS technology. Results show that non-optimal choice of the grid size can determine a dramatic increase in power (1.7X) and area (1.3X). Interestingly, the grid size that optimizes the power-delay-area tradeoff depends very weakly on the specific circuit under design, hence a generally optimum grid size exists that optimizes a very wide range of different circuits. Massimo Alioto, Stéphane Badel, Yusuf Leblebici |
ISCAS | 3 |
| 2009 | Analysis and Design of Ultra-low Power Subthreshold MCML GatesabstractIn this paper, ultra-low power current-mode subthreshold MOS current-mode logic (MCML) gates are discussed from a modeling and design perspective. A detailed analysis of the DC characteristics is presented, and the effect of process variations is analyzed in depth. Analysis allows for understanding the main limits of sub-threshold MCML gates in terms of delay/power variability. In particular, it is shown that process variations strongly affect the DC characteristics, and moderately impact delay and power consumption. Interestingly, delay and power variations are shown to be significantly reduced compared to typical values encountered in standard subthreshold CMOS logic. Criteria to size transistors to keep variations within assigned bounds are also derived. Results of Monte Carlo simulations with a 65-nm CMOS technology are reported to validate theoretical results. Massimo Alioto, Yusuf Leblebici |
ISCAS | 2 |
| 2009 | Load Optimization of an Inductive Power Link for Remote Powering of Biomedical ImplantsabstractThis article presents the analysis of the power efficiency of the inductive links used for remote powering of the biomedical implants by considering the effect of the load resistance on the efficiency. The optimum load condition for the inductive links is calculated from the analysis and the coils are optimized accordingly. A remote powering link topology with a matching network between the inductive link and the rectifier has been proposed to operate the inductive link near its optimum load condition to improve overall efficiency. Simulation and measurement results are presented and compared for different configurations. It is shown that, the overall efficiency of the remote powering link can be increased from 9.84% to 20.85% for 6 mW and from 13.16% to 18.85% for 10 mW power delivered to the regulator, respectively. Kanber Mithat Silay, Denis Dondi, Luca Larcher, Michel J. Declercq, Luca Benini, Yusuf Leblebici, Catherine Dehollain |
ISCAS | 6 |
| 2009 | Subthreshold Leakage Reduction: A Comparative Study of SCL and CMOS DesignabstractThe large subthreshold leakage current of static CMOS logic circuits designed in modern nanometer-scale technologies is one of the main barriers for implementing ultra-low power digital systems. Subthreshold source-coupled logic (STSCL) circuits are based on an NMOS differential pair that is switching a constant tail bias current between the two output branches while biased at very low current levels. The power consumption of each STSCL gate depends on the tail bias current that can be controlled very well even for current levels in the range of few tens of pico-Amperes. The precise control on the power consumption of each gate, makes this topology very attractive for ultra-low power applications, where the power consumption of conventional static CMOS system is practically limited by the subthreshold leakage current. In this work, an analytical approach supported by simulation and measurement results will be presented to study the main issues in design of ultra-low power static CMOS and STSCL systems. Armin Tajalli, Yusuf Leblebici |
ISCAS | 2 |
| 2009 | Electrical modeling of the cell-electrode interface for recording neural activity from high-density microelectrode arrays
Neil Joye, Alexandre Schmid, Yusuf Leblebici |
Neurocomputing | 3 |
| 2009 | Field Programmable Compressor Trees: Acceleration of Multi-Input Addition on FPGAsabstractMulti-input addition occurs in a variety of arithmetically intensive signal processing applications. The DSP blocks embedded in high-performance FPGAs perform fixed bitwidth parallel multiplication and Multiply-ACcumulate (MAC) operations. In theory, the compressor trees contained within the multipliers could implement multi-input addition; however, they are not exposed to the programmer. To improve FPGA performance for these applications, this article introduces the Field Programmable Compressor Tree (FPCT) as an alternative to the DSP blocks. By providing just a compressor tree, the FPCT can perform multi-input addition along with parallel multiplication and MAC in conjunction with a small amount of FPGA general logic. Furthermore, the user can configure the FPCT to precisely match the bitwidths of the operands being summed. Although an FPCT cannot beat the performance of a well-designed ASIC compressor tree of fixed bitwidth, for example, 9×9 and 18×18-bit multipliers/MACs in DSP blocks, its configurable bitwidth and ability to perform multi-input addition is ideal for reconfigurable devices that are used across a variety of applications. Alessandro Cevrero, Panagiotis Athanasopoulos, Hadi Parandeh-Afshar, Ajay Kumar Verma, Seyed-Hosein Attarzadeh-Niaki, Chrysostomos Nicopoulos, Frank K. Gürkaynak, Philip Brisk, Yusuf Leblebici, Paolo Ienne |
ACM Trans. Reconfigurable Technol. Syst. | 9 |
| 2008 | Design space exploration for field programmable compressor treesabstractThe Field Programmable Compressor Tree (FPCT) is a programmable compressor tree (e.g., a Wallace or Dadda Tree) intended for integration in an FPGA or other reconfigurable device. This paper presents a design space exploration (DSE) method that can be used to identify the best FPCT architecture for a given set of arithmetic benchmark circuits; in practice, an FPGA vendor can use the design space exploration to tailor the FPCT to meet the needs of the most important benchmark circuits of the vendor's largest-volume clients. One novel feature of the DSE is the introduction of a metric called I/O utilization; we found that I/O utilization has a strong correlation with both the critical path delay and area of the benchmark circuits under study. Pruning the search space using I/O utilization allowed us to reduce significantly the number of FPCTs that must be synthesized and evaluated during the DSE, while giving high confidence that the best architectures are still explored. The DSE was applied to seven small-to-medium range benchmark circuits; one FPCT architecture was found that was 30% faster than the second best in terms of critical path delay, and only 3.34% larger than the smallest. Seyed-Hosein Attarzadeh-Niaki, Alessandro Cevrero, Philip Brisk, Chrysostomos Nicopoulos, Frank K. Gürkaynak, Yusuf Leblebici, Paolo Ienne |
CASES | 6 |
| 2008 | Programmable logic circuits based on ambipolar CNFETabstractRecently, it was demonstrated that the polarity of carbon nanotube field effect transistors can be electrically controlled. In this paper we show how Programmable Logic Arrays (PLA) can be built out of these devices, and we illustrate how they outperform usual PLA by internal signal inversion. The simulations show an area saving up to approximately 21% and decrease of the delay in PLA-based FPGA by 50%. We also show that this architecture is suitable for high-performance design tools and defect-tolerance approaches. M. Haykel Ben Jamaa, David Atienza 0001, Yusuf Leblebici, Giovanni De Micheli |
DAC | 3 |
| 2008 | A Generic Standard Cell Design Methodology for Differential Circuit StylesabstractIn this paper we present a generic methodology for the rapid generation and implementation of standard cell libraries for differential circuit design styles. We demonstrate a systematic approach for the classification of circuit topologies (footprints) and for generating the templates that correspond to a large number of functions. The generation of an extensive cell library with more than 4500 standard cells based on 19 footprints is demonstrated using a 180 nm CMOS technology. Stéphane Badel, Erdem Guleyupoglu, Ozgur Inac, Anna Pena Martinez, Paolo Vietti, Frank K. Gürkaynak, Yusuf Leblebici |
DATE | 7 |
| 2008 | Novel Front-End Circuit Architectures for Integrated Bio-Electronic InterfacesabstractThe prospective use of upcoming nanometer CMOS technology nodes (65 nm, 45 nm, and beyond) in bio-electronic interfaces is raising a number of important issues concerning circuit architectures and design. In particular, the advantages of scaling and higher density integration must be balanced against the requirements of low noise design, uniform power density and surface temperature distribution, better component matching, and immunity to parameter variations. Dealing with these constraints also requires more innovative approaches towards hybrid integration technologies. In this paper, we discuss the key design issues with specific examples from DNA detection, protein detection, and neuro-electronic interfaces. Carlotta Guiducci, Alexandre Schmid, Frank K. Gürkaynak, Yusuf Leblebici |
DATE | 4 |
| 2008 | Architectural improvements for field programmable counter arrays: enabling efficient synthesis of fast compressor trees on FPGAsabstractThe Field Programmable Counter Array (FPCA) was introduced to improve FPGA performance for arithmetic circuits. An FPCA is a reconfigurable IP core that can be integrated into an FPGA. To exploit the FPCA, a circuit is transformed by merging disparate addition and multiplication operations into large multi-input addition operations, which are synthesized as compressor trees on the FPCA; the remaining portion of the circuit is synthesized on the FPGA. This paper presents a series of architectural improvements to the FPCA that reduce routing delay, increase flexibility and component utilization, and simplify the integration process. Using an FPGA containing six FPCAs, we observed average and maximum speedups of 1.60x and 2.40x on a set of arithmetic benchmarks Alessandro Cevrero, Panagiotis Athanasopoulos, Hadi Parandeh-Afshar, Ajay Kumar Verma, Philip Brisk, Frank K. Gürkaynak, Yusuf Leblebici, Paolo Ienne |
FPGA | 7 |
| 2008 | Improving the power-delay product in SCL circuits using source follower output stageabstractThis article explores the effect of using source follower buffers (SFB) at the output of source coupled logic (SCL) circuits. This technique can help to improve the power-delay product (PDP) of an SCL gate approximately by a factor of two. The proposed approach has been applied to improve the PDP in sub-threshold SCL circuits that have been developed for ultra- low power applications. Designed in conventional digital 0.18mum CMOS technology, the proposed SCL gate utilizing SFB at the output achieves a PDP of 0.5fJ/fF/gate while the gate draws 10nA from a 0.6V supply voltage. Armin Tajalli, Frank K. Gürkaynak, Yusuf Leblebici, Massimo Alioto, Elizabeth J. Brauer |
ISCAS | 3 |
| 2008 | Variability-Aware Design of Multilevel Logic Decoders for Nanoscale Crossbar MemoriesabstractThe fabrication of crossbar memories with sublithographic features is expected to be feasible within several emerging technologies; in all of them, the nanowire (NW) decoder is a critical part since it bridges the sublithographic wires to the outer circuitry that is defined on the lithography scale. In this paper, we evaluate the addressing scheme of the decoder circuit for NW crossbar arrays, based on the existing technological solutions for threshold voltage differentiation of NW devices. This is equivalent to using a multivalued logic addressing scheme. With this approach, it is possible to reduce the decoder size and keep it defect tolerant. We formally define two types of multivalued codes (i.e., hot and reflexive codes), and we estimate their yield under high variability conditions. Multivalued hot decoders yield better area saving thann-ary reflexive codes, and under severe conditions, reflexive codes enable a nonvanishing part of the code space to randomly recover. The choice of the optimal combination of decoder type and logic level saves area up to 24%. We also show that the precision of the addressing voltages when a high variability affects the threshold voltages is a crucial parameter for the decoder design and permits large savings in memory area. Moreover, a precise knowledge about the variability level improves the design of memory decoders by giving the right optimal code. M. Haykel Ben Jamaa, Kirsten E. Moselund, David Atienza 0001, Didier Bouvet, Adrian M. Ionescu, Yusuf Leblebici, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2007 | Design and realization of a fault-tolerant 90nm CMOS cryptographic engine capable of performing under massive defect densityabstractThis paper presents a new approach for assessing the reliability of nanometer-scale devices prior to fabrication and a practical reliability architecture realization. A four-layer architecture exhibiting a large immunity to permanent as well as random failures is used. Characteristics of the averaging/thresholding layer are emphasized. A complete tool based on Monte Carlo simulation for a-priori functional fault tolerance analysis was used for analysis of distinctive cases and topologies. A full chip CMOS integrated design of the 128-bit AES cryptography algorithm with multiple cores that incorporate reliability architectures is shown. Milos Stanisavljevic, Frank K. Gürkaynak, Alexandre Schmid, Yusuf Leblebici, Maria Gabrani |
ACM Great Lakes Symposium on VLSI | 4 |
| 2007 | Fault-tolerant multi-level logic decoder for nanoscale crossbar memory arraysabstractSeveral technologies with sub-lithographic features are targeting the fabrication of crossbar memories in which the nanowire decoder is playing a major role. In this paper, we suggest a way to reduce the decoder size and keep it defect tolerant by using multiple threshold voltages (VT), which is enabled by our underlying technology. We define two types of multi-valued decoders and model the defects they undergo due to the VT variation. Multi-valued hot decoders yield better area saving than n-ary reflexive codes (NRC), and under severe conditions, NRC enables a non-vanishing part of the code space to recover. There are many combinations of decoder type and number of VT’s yielding equal effective memory capacities. The optimal choice saves area up to 24%. We also show that the precision of the addressing voltages for decoders with unreliable VT’s is a crucial parameter for the decoder design and permits large savings in memory area. M. Haykel Ben Jamaa, Kirsten E. Moselund, David Atienza 0001, Didier Bouvet, Adrian M. Ionescu, Yusuf Leblebici, Giovanni De Micheli |
ICCAD | 6 |
| 2007 | Breaking the Power-Delay Tradeoff: Design of Low-Power High-Speed MOS Current-Mode Logic Circuits Operating with Reduced Supply VoltageabstractIn this paper, the authors study the operation of MOS current-mode logic (MCML) gates at lower-than-nominal supply voltages. The authors show that power can be traded for speed by reducing the supply voltage below the nominal value, while the power-delay product stays nearly constant. The authors propose a negative bias strategy that enables the gates to operate at maximum speed with a reduced supply voltage, thus achieving a power saving of up to 35% at no cost for speed. Comparison with CMOS logic style are presented for three different technology nodes (0.25μm, 0.18μm and 0.13μm CMOS). Stéphane Badel, Yusuf Leblebici |
ISCAS | 2 |
| 2006 | A simulation methodology for reliability analysis in multi-core SoCsabstractReliability has become a significant challenge for system design in new process technologies. Higher integration levels dramatically increase power densities, which leads to higher temperature and adverse effects on reliability. In this paper, we introduce a simulation methodology to analyze reliability of multi-core SoCs. The proposed simulator is the first to provide system-on-chip level fine-grained reliability analysis. We use our simulation methodology to study the reliability effects of design choices such as thermal packaging and placement, as well as runtime events such as power management policies and workload distributions. Ayse K. Coskun, Tajana Rosing, Yusuf Leblebici, Giovanni De Micheli |
ACM Great Lakes Symposium on VLSI | 3 |
| 2006 | Fault-Tolerance of Robust Feed-Forward Architecture Using Single-Ended and Differential Deep-Submicron Circuits Under Massive Defect DensityabstractAn assessment of the fault-tolerance properties of single-ended and differential signaling is shown in the context of a high defect density environment, using a robust error-absorbing circuit architecture. A software tool based on Monte-Carlo simulations is used for the reliability analysis of the examined logic families. A benefit of the differential circuit over standard single-ended is shown in case of complex systems. Moreover, analysis of reliability of different circuits and discussion on the optimal granularity of redundant blocks was made. Milos Stanisavljevic, Alexandre Schmid, Yusuf Leblebici |
IJCNN | 3 |
| 2006 | Weak inversion performance of CMOS and DCVSPG logic families in sub-300 mV rangeabstractIn this paper the advantages of using differential cascode voltage switch pass gate (DCVSPG) logic with regard to standard CMOS for subthreshold operation are presented. The two families are compared in terms of their performance and energy-delay-product (EDP) figures. Multiple gates were simulated using 0.18 mum standard CMOS technology. Simulation results show that DCVSPG NAND2 gate has 71%, DCVSPG NOR2 gate has 82% and DCVSPG full adder has 66% EDP savings over the CMOS counterparts Omer Can Akgun, Yusuf Leblebici |
ISCAS | 2 |
| 2006 | Via-programmable expanded universal logic gate in MCML for structured ASIC applications: circuit designabstractThis paper presents a via-programmable expanded universal logic gate in MOS current-mode logic which can implement any 3-input Boolean function, and a significant subset of 4-input and 5-input functions. The universal logic gate is programmed with the first via mask, while metal3 and higher levels are used for cell-to-cell interconnections. Thus the cell is suitable for a structured ASIC design methodology. The circuit was used to create a functional cell library which can implement a wide range of functions. The cells are simulated to characterize delays, and a design strategy is proposed for large scale integration Elizabeth J. Brauer, Ilhan Hatirnaz, Stéphane Badel, Yusuf Leblebici |
ISCAS | 4 |
| 2006 | Analysis and modeling of jitter and frequency tolerance in gated oscillator based CDRsabstractThis paper presents an approach to analyzing and modeling of gated-oscillator (GO) -based CDRs and predicting their performance aspects such as jitter tolerance (JTOL) and frequency tolerance (FTOL). It is shown that high JTOL of this topology in addition to their acceptable FTOL and flexible topology, have made them very suitable for short-haul multi-rate applications Armin Tajalli, Paul Muller, Seyed Mojtaba Atarodi, Yusuf Leblebici |
ISCAS | 4 |
| 2006 | Implementation of Structured ASIC Fabric Using Via-Programmable Differential MCML CellsabstractThis paper presents a regular layout fabric made of via-programmable MCML universal logic cells for structured ASIC applications and the associated design flow. The proposed structured ASIC fabric offers very high noise immunity due to the differential operation, as well as low production cost due to the via-programmable properties of the universal logic cell. Implementations of a number of circuits are presented and the area/speed performances are compared with classical CMOS implementation using a commercial standard cell library in 0.18 mum CMOS technology Stéphane Badel, Ilhan Hatirnaz, Yusuf Leblebici, Elizabeth J. Brauer |
VLSI-SoC | 3 |
| 2006 | Configurable On-Line Global Energy Optimization in Multi-Core Embedded Systems Using Principles of Analog ComputationabstractThis work presents the design of an on-line energy optimizer unit, which is capable of dynamically adjusting power supply voltages and operating frequencies of multiple processing elements (PE), tailored to the instantaneous workload information and is fully adaptive to variations in process and temperature. The circuit design borrows some of the basic principles of analog computation to continuously optimize the system-wide energy dissipation of multiple cores. The analogy between the energy minimization problem under timing constraints in a general task graph and the power minimization problem under Kirchhoffs current law (KCL) constraints in an equivalent resistive network is exploited Zeynep Toprak Deniz, Yusuf Leblebici, Eric A. Vittoz |
VLSI-SoC | 2 |
| 2006 | A Predictable Communication Scheme for Embedded Multiprocessor SystemsabstractNetworks-on-chip (NoC) are emerging as a widely accepted alternative for the traditional bus architectures. However, their applicability by the system designers is far away from being intuitive due to their lack of predictability. This communication predictability can be obtained statically or dynamically. A dynamic allocation is more suitable for flexible multiprocessor systems and requires the implementation of a quality-of-service (QoS) mechanism. This paper explores the main QoS schemes suitable for such systems: connection-oriented and connectionless. The simulation results show that the connectionless scheme provides a better predictability in terms of message latency with an acceptable buffer requirement. This work provides the designer with valuable guidelines to choose a priori the QoS parameters such that they can be confident on the predicted results Mehmet Derin Harmanci, Nuria Pazos, Paolo Ienne, Yusuf Leblebici |
VLSI-SoC | 4 |
| 2005 | CONAN - A Design Exploration Framework for Reliable Nano-ElectronicsabstractIn this paper we introduce a design methodology that allows the system/circuit designer to build reliable systems out of unreliable nano-scale components. The central point of our approach is a generic (parametrical) architectural template. Configurable nanostructures for reliable nano electronics (CONAN), which embeds support for reliability at various levels of abstractions. Some of the main reliability sources are regular and decentralized structures based on simple basic computation cells designed to be robust against disturbances and noise, fault tolerance based on hardware, time and information redundancy applied at the basic cell level as well as at higher levels, self diagnosis assisted by the dynamic reconfiguration of basic computation cells and interconnect rerouting. Within the CONAN template, both technology dependent and independent models co-exists such that the more abstract layers are technology independent while the lower levels can be retargeted to various fabrication technologies. Our proposal is application-oriented and allows the designers to deal with unpredictability, and low reliability, which are unavoidable characteristics of future emerging nano-devices. When combined with the underlying software, the tools supporting the CONAN approach allow the designer to check whether the design constraints are fulfilled before performing a detailed implementation and provides means to trade area, delay, and power consumptions for reliability. As such, this proposal is a call-to-arms to mobilize the efforts of systems designers in order to achieve a systematic design methodology for reliable systems. Sorin Cotofana, Alexandre Schmid, Yusuf Leblebici, Adrian M. Ionescu, Oliver Soffke, Peter Zipf, Manfred Glesner, Antonio Rubio 0001 |
ASAP | 3 |
| 2005 | Top-Down Design of a Low-Power Multi-Channel 2.5-Gbit/s/Channel Gated Oscillator Clock-Recovery CircuitabstractWe present a complete top-down design of a low-power multi-channel clock recovery circuit based on gated current-controlled oscillators. The flow includes several tools and methods used to specify block constraints, to design and verify the topology down to the transistor level, as well as to achieve a power consumption as low as 5 mW/Gbit/s. Statistical simulation is used to estimate the achievable bit error rate in the presence of phase and frequency errors and to prove the feasibility of the concept. VHDL modeling provides extensive verification of the topology. Thermal noise modeling based on well-known concepts delivers design parameters for the device sizing and biasing. We present two practical examples of possible design improvements analyzed and implemented with this methodology. Paul Muller, Armin Tajalli, Seyed Mojtaba Atarodi, Yusuf Leblebici |
DATE | 4 |
| 2005 | Jitter Tolerance Analysis of Clock and Data Recovery Circuits
Paul Muller, Yusuf Leblebici |
FDL | 2 |
| 2005 | A low-power, multichannel gated oscillator-based CDR for short-haul applicationsabstractA gated current-controlled oscillator (GCCO) based topology is used to implement a low-power multi-channel clock and data recovery (CDR) system in a 0.18um digital CMOS technology. A systematic approach is presented to design a reliable and low-power system based on the required specifications. Behavioral simulations are also used to estimate the achievable bit error rate (BER), jitter tolerance (JTOL), and frequency offset tolerance (FTOL) of the proposed CDR. Using a single 1.8V supply voltage, the proposed 20Gbps 8-channel CDR consumes only 70.2mW or 3.51mW/Channel/Gbps while occupies 0.045mm2 silicon area Armin Tajalli, Paul Muller, Seyed Mojtaba Atarodi, Yusuf Leblebici |
ISLPED | 4 |
| 2005 | A Methodology for Reliability Enhancement of Nanometer-Scale Digital Systems Based on a-priori Functional Fault- Tolerance Analysis
Milos Stanisavljevic, Alexandre Schmid, Yusuf Leblebici |
VLSI-SoC | 3 |
| 2004 | Fault-tolerant PLA-style circuit design for failure-prone nanometer CMOS and quantum device technologiesabstractAbs. Alexandre Schmid, Yusuf Leblebici |
IJCNN | 2 |
| 2004 | Robust circuit and system design methodologies for nanometer-scale devices and single-electron transistorsabstractIn this paper, various circuit and system level design challenges for nanometer-scale devices and single-electron transistors are discussed, with an emphasis to the functional robustness and fault tolerance point of view. A set of general guidelines is identified for the design of very high-density digital systems using inherently unreliable and error-prone devices. The fundamental principles of a highly regular, redundant, and scalable design approach based on fixed-weight neural networks and multiple-valued logic are presented. It is demonstrated that the proposed design technique offers significantly improved immunity to permanent and transient faults occurring at the transistor level, and that it results in graceful degradation of circuit performance in response to device failures. Alexandre Schmid, Yusuf Leblebici |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | VLSI Realization of a Two-Dimensional Hamming Distance Comparator ANN for Image Processing Applications
Stéphane Badel, Alexandre Schmid, Yusuf Leblebici |
ESANN | 3 |
| 2003 | A VLSI Hamming artificial neural network with k-winner-take-all and k-loser-take-all capabilityabstractA novel circuit-level Hamming artificial neural network architecture based on the principle of analog charge-based computation of the neural function is proposed. k-winner-take-all and k-loser-take-all operations are performed in the time-domain, allowing for fast and compact realization of complex functions. The VLSI realization of a two-dimensional array arrangement of the Hamming network is presented, with the targeted processing applications. Stéphane Badel, Alexandre Schmid, Yusuf Leblebici |
IJCNN | 3 |
| 2000 | A compact modular architecture for high-speed binary sortingabstractA new algorithm and a new modular architecture are presented for the realization of high-speed binary sorting engines, based on efficient rank ordering. Capacitive threshold logic (CTL) gates are utilized for the implementation of the multi-input programmable majority (voting) functions required in the architecture. The overall complexity of the proposed bit-serial architecture increases linearly with the number of input vectors to be sorted (window size=m) and with the bit-length of the input vectors (word size=n), and the sorter architecture can be easily expanded to accommodate large vector sets. Detailed simulations indicate that the sorter structure can operate at sampling clock rates of up to 50 MHz, where the throughput is boosted by fine-grain pipelining. It is demonstrated that the proposed sorting engine is capable of producing a fully sorted output vector set in (m+n-1) clock cycles. Ilhan Hatirnaz, Frank K. Gürkaynak, Yusuf Leblebici |
ICASSP | 3 |
| 2000 | Higher radix Kogge-Stone parallel prefix adder architecturesabstractIn this paper, we describe the design of radix-3 and radix-4 parallel prefix adders, that theoretically have logical depths of log/sub 3/n and log/sub 4/n respectively, where n is the bit-width of the input signals. The main building blocks of the higher radix parallel prefix adders are identified and higher radix structures of Kogge-Stone Adders are presented. We show that with the higher radix architectures the logic depth can be reduced by 50% and the cell count can be reduced as much as 47% for 64-bit adders. Simulation results indicate that radix-4 adders can be more than 30% faster than radix-2 realizations. Frank K. Gürkaynak, Yusuf Leblebici, Laurent Chaouat, Patrik J. McGuinness |
ISCAS | 2 |
| 2000 | A compact modular architecture for the realization of high-speed binary sorting engines based on rank orderingabstractA new modular architecture is presented for the realization of high-speed binary sorting engines, based on efficient rank ordering. Capacitive Threshold Logic (CTL) gates are utilized for the implementation of the multi-input programmable majority (voting) functions required in the architecture. The overall complexity of the proposed bit-serial architecture increases linearly with the number of input vectors to be sorted (window size=m) and with the bit-length of the input vectors (word size=n), and the sorter architecture can be easily expanded to accommodate large vector sets. Detailed simulations indicate that the sorter structure can operate at sampling clock rates of up to 50 MHz, where the throughput is boosted by fine-grain pipelining. It is demonstrated that the proposed sorting engine is capable of producing a fully sorted output vector set in (m+n-1) clock cycles, i.e., in linear time. Ilhan Hatirnaz, Frank K. Gürkaynak, Yusuf Leblebici |
ISCAS | 3 |
| 1999 | A two-stage charge-based analog/digital neuron circuit with adjustable weightsabstractA circuit-level neuron architecture based on the principle of analog charge-based computation of neural functions has been developed with the goals of high-speed processing, adjustable weights, and support of perturbation-based learning algorithms. The two-stage architecture which is composed of nonlinear synapses, driving a linear capacitive soma, has been implemented using a conventional double-polysilicon CMOS technology. The feedforward architecture of the proposed neuron model is shown to synthesize a large number of nonlinear mappings of the 2D-1D space. Alexandre Schmid, Yusuf Leblebici, Daniel Mlynek |
IJCNN | 2 |
| 1993 | ILLIADS: a fast timing and reliability simulator for digital MOS circuitsabstractThe authors introduce ILLIADS as a fast MOS timing and reliability simulator for very large digital MOS circuits. The use of the proposed general circuit primitive not only provides better accuracy but also significantly reduces the simulation time. The use of an analytic solution embedded in the simulation engine improves both the simulation speed and the accuracy. Postponing the waveform approximation process provides better waveform approximation even for non-fully-switching waveforms and glitches. The channel length modulation effect is captured accurately in fast timing simulation with only 10% of speedup tradeoff. It is also shown that ILLIADS manifests the charge sharing problem. The modified PWCTC algorithm, PWCTC-W, which handles circuits with feedback, is introduced and shown to be superior to the dynamic-windowing scheme. It also does not manifest the window-growing problem and is insensitive to the level of strongly concerned component (SCC) hierarchy. The use of this algorithm keeps the speedup of ILLIADS over SPICE for circuits with feedbacks at the same level as that for combinational circuits.> Yung-Ho Shih, Yusuf Leblebici |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1992 | Modeling of nMOS transistors for simulation of hot-carrier-induced device and circuit degradationabstractThe authors present an accurate one-dimensional device model for the simulation of nMOS transistors with hot-carrier-induced oxide damage. The model uses a realistic charge density distribution profile to account for the localization of the oxide-interface charge near the drain. Model simulation results obtained for nMOS transistors with hot-carrier-induced oxide damage demonstrate good agreement with the experimental data. The amount and the location of the hot-carrier-induced oxide damage are simulated by using only a few parameters, which simplifies the implementation of the model in a reliability simulation environment. The proposed model has been implemented in the iSMILE circuit simulator, and the capabilities of the model have been explored by various circuit simulation examples. The damaged-MOSFET model presented offers a simple and accurate approach for simulating the circuit behavior after hot-carrier damage.> Yusuf Leblebici |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1991 | New Simulation Methods for MOS VLSI Timing and ReliabilityabstractA novel approach to incorporating the channel length modulation in a direct-equation solving fast timing simulator is presented along with a mixed event-driven and waveform relaxation algorithm to handle MOS VLSI circuits with feedback. Simulation speedup of 3N over SPICE-like simulators has been observed, where N is the number of transistors. The simulator is able to simulate circuits as large as 235000 transistors in 10 min real time. Also presented is a novel approach to fast hot-carrier reliability simulation. These methods make it possible to achieve accurate and fast hot-carrier reliability simulation of MOS circuits each with as many as hundreds of thousands of MOS transistors in a workstation environment.> Yung-Ho Shih, Yusuf Leblebici |
ICCAD | 2 |
| 1991 | An accurate analytical delay model for BiCMOS driver circuitsabstractAn analytical delay model for BiCMOS driver circuits is presented. The model is based on physical device parameters and can be used to estimate both the pull-up and the pull-down times for a variety of circuit configurations. The intrinsic delay associated with the bipolar transistors is taken into consideration by using a charge control model that incorporates the high-injection effects upon the current gain and the base transport factor. Separate sets of delay equations are derived for the pull-up and pull-down transient responses to account for significant differences between the two cases. The comparison with SPICE circuit simulation results shows that the new model predicts the respective delay times with less than 10% error in most cases. The influence of device dimensions upon the driver delay time is also investigated. The model has been applied to find an optimal area allocation between the CMOS and bipolar parts of the driver circuit when the total available area is limited as in the standard cell configuration.> Carlos H. Díaz, Yusuf Leblebici |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 1990 | An Integrated Hot-Carrier Degradation Simulator for VLSI Reliability AnalysisabstractA novel integrated simulation tool is presented for estimating the hot-carrier induced degradation of nMOS transistor characteristics and circuit performance. The proposed reliability simulation tool incorporates an accurate one-dimensional MOSFET model for representing the electrical behavior of locally damaged transistors. The hot-carrier induced oxide damage can be specified by only a few parameters, avoiding extensive parameter extractions for the characterization of device damage. The physical degradation model used in the simulation tool includes both of the fundamental device degradation mechanisms, i.e., charge trapping and interface trap generation. A repetitive simulation scheme has been adopted to ensure accurate prediction of the circuit-level degradation process under dynamic operating conditions. The simulation tool provides information on the evolution of device degradation during long-term operation, and on the performance characteristics of the damaged circuit.> Yusuf Leblebici |
ICCAD | 1 |
| 1989 | Simulation of MOS circuit performance degradation with emphasis on VLSI design-for-reliabilityabstractA framework for a reliability simulation tool to assess the hot-carrier-induced degradation of MOS circuits is presented, and the major components of this framework are examined. A method is introduced for dynamic simulation of hot-carrier-induced transistor degradation within the circuit environment. The approach accounts for the gradual degradation of terminal voltage waveforms of MOS transistors during long-term operation. It is demonstrated that the estimation of individual device lifetimes is not sufficient for circuit reliability assessment. The critical transistors that are most likely to cause circuit performance failures are identified by combining the long-term degradation estimates with the corresponding circuit performance sensitivities.> Yusuf Leblebici |
ICCD | 1 |
| 1988 | An efficient method for circuit sensitivity calculation using piecewise linear waveform modelsabstractAn efficient method for calculating the transient sensitivities in MOS circuits with respect to a large number of parameters is discussed. The approach uses simple circuit models built by piecewise-linearization of time-domain circuit responses. Closed-form expressions are derived for the calculation of transient sensitivities of the corresponding linear circuits. The transient sensitivity of the nonlinear circuit is then approximated as a simple function of individual linear circuit sensitivities. It is shown that the efficiency of this approach increases with the circuit size as well as with the number of sensitivity parameters considered.> Yusuf Leblebici |
ICCAD | 2 |