VLDB 2026 Research / reviewers in the wild / expert
Sorin Cotofana
dblp:c/SorinCotofana · also Sorin Dan Cotofana
· DBLP profile ↗
81ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0001-7132-2291ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 73 · 3 first-author · 11 since 2021Software engineering, systems software and programming languages · 10 · 1 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-speed and low-power Graphene ADC with scalable resolutionabstractContains fulltext : 333999.pdf (Publisher’s version ) (Closed access) Nicoleta Cucu Laurenciu, Charles Timmermans, Sorin Cotofana |
ISCAS | 3 |
| 2025 | A Convoluted Journey from CMOS to Spin WavesabstractIn recent years, Spin Waves (SWs) have emerged as a promising CMOS alternative technology, and SW interference-based majority gates have been proposed and experimentally realized. In this paper, we pursue a different computation avenue and introduce a SW device able to evaluate 2×2 2D convolution, which is a fundamental element for the implementation of Convolutional Neural Networks (CNNs). Assuming that the window pixels are P = [p1, p2; p3, p4] and the kernel is K = [k1, k2; k3, k4] we introduce a device which evaluates the convolution result $\sum\nolimits_{i = 1}^4 {{p_i}} {k_i}$ within the SW domain by leveraging SWs inherent mechanisms, i.e., information encoding in SW amplitude and phase, SW amplitude decay due to Gilbert damping, SW interference. After introducing the SW device structure we demonstrate its proper behaviour by means of micromagnetic simulations. We also present power consumption, area, and delay estimates and argue that due to the fact that our proposal does not rely on standard adders and multipliers, it can substantially outperform traditional CMOS-based convolution implementations. Pantazis Anagnostou, Arne Van Zegbroeck, Said Hamdioui, Christoph Adelmann, Florin Ciubotaru, Sorin Cotofana |
ISCAS | 6 |
| 2025 | Graphene-based, Frequency-Domain Self-Trigger for Low SNR Air Showers-induced PulsesabstractIn this paper we introduce a frequency-domain pulse detection method that is suitable for in-situ implementation at detector-level, for low-power, self-triggered air shower detectors. We propose a graphene-based architecture, and demonstrate its correct operation by means of SPICE simulations. The utilized graphene-based devices operate at low supply voltage, consume low energy per spike, and exhibit small footprints, which are essential properties for large-scale, energy-efficient implementations. The proposed method is particularly effective for very low (Signal-to-Noise Ratio) SNR scenarios, and is broadband noise resilient up to a certain extent, and (Radio Frequency) RF narrowband noise agnostic. Comparison results against time-domain signal-over-threshold trigger indicates that the proposed method can outperform its counterpart in terms of trigger efficiency by up to 26× and 47×, when using 1 and 2 frequency components, respectively, especially for very low SNR scenarios (up to −42 dB) where time-domain methods are largely impaired. Furthermore, the proposed method does not require RF filtering in advance, and can coexist with other noise pulses. Thus, high detection efficiency that goes in tandem with high purity (low number of false positives) becomes tenable with proposed approach. Nicoleta Cucu Laurenciu, Charles Timmermans, Sorin Cotofana |
ISCAS | 3 |
| 2025 | Benchmarking of Scaled Majority-Logic-Synthesized Spintronic Circuits Based on Magnetic Tunnel Junction TransducersabstractIt is envisaged that spintronic logic devices will ultimately be utilized in hybrid CMOS-spintronic systems where signal interconversion between magnetic and electrical domains via transducers takes place. This underscores the vital role of transducers in influencing the overall performance of such hybrid systems. This paper addresses the question: Can spintronic circuits based on Magnetic Tunnel Junction (MTJ) transducers outperform their state-of-the-art CMOS counterparts? To this end, we use the EPFL (École Polytechnique Fédérale de Lausanne) combinational benchmark sets, synthesize them in 7 nm CMOS and in MTJ transducer based spintronic technologies, and compare the two implementation methods in terms of Energy-Delay-Product (EDP). To fully utilize the technologies’ potential, CMOS and spintronic implementations are built upon standard Boolean and Majority Gates, respectively. For the spintronic circuits, we assumed that domain conversion (electric/magnetic to magnetic/electric) is performed by means of MTJs and the computation is accomplished by domain wall (DW)-based majority gates, and considered two EDP estimation scenarios: (i) Uniform Benchmarking, which ignores the circuit’s internal structure and only includes domain transducers’ power and delay contributions into the calculations, and (ii) Majority-Inverter-Graph Benchmarking, which also embeds the circuit structure, the associated critical path delay and energy consumption by DW propagation. Our results indicate that, for the uniform case, the spintronic route is better suited for the implementation of complex circuits with few inputs and outputs. On the other hand, when the circuit structure is also considered via majority and inverter synthesis, our analysis clearly indicates that in order to match and eventually outperform CMOS performance, MTJ transducers’ efficiency has to be improved by 3-4 orders of magnitude. While it is clear that for the time being the MTJ-based-spintronic way cannot compete with CMOS, further technological transducer developments may tip the balance, which, when combined with information non-volatility, may make spintronic implementation for certain applications that require a large number of calculations and have a rather limited amount of interaction with the environment. Fanfan Meng, Siang-Yun Lee, Odysseas Zografos, Mohit Gupta 0004, Van D. Nguyen, Giovanni De Micheli, Sorin Cotofana, Inge Asselberghs, Christoph Adelmann, Gouri Sankar Kar, Sebastien Couet, Florin Ciubotaru |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | An Energy-Efficient Graphene-based Spiking Neural Network Architecture for Pattern RecognitionabstractIn this paper we propose a generic graphene-based Spiking Neural Network (SNN) architecture for pattern recognition and the associated weight values initialization methodology. The SNN has a Winner-Takes-All 3-layer structure and exhibits tuneable recognition accuracy by exploiting interpatterns similarity/dissimilarity. To demonstrate the capabilities of our proposal we present an SNN instance tailored for low resolution MNIST handwritten digits recognition and evaluate its recognition accuracy by means of SPICE simulations. 2 voltage levels are initially utilized for synaptic weight values representation and the recognition accuracy varies from 75.8% to 99.2%, which, together with its compactness and energy efficient (pJ range/spike), suggests that our approach has great potential for edge device implementations. Nicoleta Cucu Laurenciu, Charles Timmermans, Sorin Cotofana |
ISCAS | 3 |
| 2024 | An Energy-Efficient Bayesian Neural Network Implementation Using Stochastic Computing MethodabstractThe robustness of Bayesian neural networks (BNNs) to real-world uncertainties and incompleteness has led to their application in some safety-critical fields. However, evaluating uncertainty during BNN inference requires repeated sampling and feed-forward computing, making them challenging to deploy in low-power or embedded devices. This article proposes the use of stochastic computing (SC) to optimize the hardware performance of BNN inference in terms of energy consumption and hardware utilization. The proposed approach adopts bitstream to represent Gaussian random number and applies it in the inference phase. This allows for the omission of complex transformation computations in the central limit theorem-based Gaussian random number generating (CLT-based GRNG) method and the simplification of multipliers as AND operations. Furthermore, an asynchronous parallel pipeline calculation technique is proposed in computing block to enhance operation speed. Compared with conventional binary radix-based BNN, SC-based BNN (StocBNN) realized by FPGA with 128-bit bitstream consumes much less energy consumption and hardware resources with less than 0.1% accuracy decrease when dealing with MNIST/Fashion-MNIST datasets. Xiaotao Jia, Huiyi Gu, Jianlei Yang 0001, Weitao Pan, Youguang Zhang, Sorin Cotofana, Weisheng Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2022 | Would Magnonic Circuits Outperform CMOS Counterparts?abstractIn the early stages of a novel technology development, it is difficult to provide a comprehensive assessment of its potential capabilities and impact. Nevertheless, some preliminary estimates can be drawn and are certainly of great interest and in this paper we follow this line of reasoning within the framework of the Spin Wave (SW) based computing paradigm. In particular, we are interested in assessing the technological development horizon that needs to be reached in order to unleash the full SW paradigm potential such that SW circuits can outperform CMOS counterparts in terms of energy consumption. In view of the zero power SWs propagation through ferromagnetic waveguides, the overall SW circuit power consumption is determined by the one associated to SWs generation and sensing by means of transducers. While current antenna based transducers are clearly power hungry recent developments indicate that magneto-electric (ME) cells have a great potential for ultra-low power SW generation and sensing. Given that MEs have been only proposed at the conceptual level and no actual experimental demonstration has been reported we cannot evaluate the impact of their utilization on the SW circuit energy consumption. However, we can perform a reverse engineering alike analysis to determine ME delay and power consumption upper bounds that can place SW circuits in the leading position. To this end, we utilize a 32-bit Brent-Kung Adder (BKA) as discussion vehicle and compute the maximum ME delay and power consumption that could potentially enable a SW implementation able to outperform its 7nm CMOS counterpart. We evaluate different BKA SW implementations that rely on conversion- or normalization-based gate cascading and consider continuous or pulsed SW generation scenarios. Our evaluations indicate that 31nW is the maximum transducer power consumption for which a 32-bit Brent-Kung SW implementation can outperform its 7nm CMOS counterpart in terms of energy consumption. Abdulqader Nael Mahmoud, Nicoleta Cucu Laurenciu, Frederic Vanderveken, Florin Ciubotaru, Christoph Adelmann, Sorin Cotofana, Said Hamdioui |
ACM Great Lakes Symposium on VLSI | 6 |
| 2022 | Non-Binary Spin Wave Based Circuit DesignabstractBy their very nature, Spin Waves (SWs) excited at the same frequency but different amplitudes, propagate through waveguides and interfere with each other at the expense of ultra-low energy consumption. In addition, all (part) of the SW energy can be moved from one waveguide to another by means of coupling effects. In this paper we make use of these SW features and introduce a novel non Boolean algebra based paradigm, which enables domain conversion free ultra-low energy consumption SW based computing. Subsequently, we leverage this computing paradigm by designing a non-binary spin wave adder, which we validate by means of micro-magnetic simulation. To get more inside on the proposed adder potential we assume a 2-bit adder implementation as discussion vehicle, evaluate its area, delay, and energy consumption, and compare it with conventional SW and 7 nm CMOS counterparts. The results indicate that our proposal diminishes the energy consumption by a factor of$3.14 \times $and$6 \times $, when compared with the conventional SW and 7 nm CMOS functionally equivalent designs, respectively. Furthermore, the proposed non-binary adder implementation requires the least number of devices, which indicates its potential for small chip real-estate realizations. Abdulqader Nael Mahmoud, Frederic Vanderveken, Florin Ciubotaru, Christoph Adelmann, Said Hamdioui, Sorin Cotofana |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | Fan-out of 2 Triangle Shape Spin Wave Logic GatesabstractHaving multi-output logic gates saves much energy because the same structure can be used to feed multiple inputs of next stage gates simultaneously. This paper proposes novel triangle shape fanout of 2 spin wave Majority and XOR gates; the Majority gate is achieved by phase detection, whereas the XOR gate is achieved by threshold detection. The proposed logic gates are validated by means of micromagnetic simulations. Furthermore, the energy and delay are estimated for the proposed structures and compared with the state-of-the-art spin wave, and 16 nm and 7 nm CMOS logic gates. The results demonstrate that the proposed structures provide energy reduction of 25%–50% in comparison to the other 2-output spin-wave devices while having the same delay, and energy reduction of 43x-0.8x when compared to the 16 nm and 7 nm CMOS counterparts while having delay overhead of 11x-40x. Abdulqader Nael Mahmoud, Christoph Adelmann, Frederic Vanderveken, Sorin Cotofana, Florin Ciubotaru, Said Hamdioui |
DATE | 4 |
| 2021 | Spin Wave Based Full AdderabstractSpin Waves (SWs) propagate through magnetic waveguides and interfere with each other without consuming noticeable energy, which opens the road to new ultra-low energy circuit designs. In this paper we build upon SW features and propose a novel energy efficient Full Adder (FA) design consisting of The FA 1 Majority and 2 XOR gates, which outputs Sum and Carry-out are generated by means of threshold and phase detection, respectively. We validate our proposal by means of MuMax3 micromagnetic simulations and we evaluate and compare its performance with state-of-the-art SW, 22nm CMOS, Magnetic Tunnel Junction (MTJ), Spin Hall Effect (SHE), Domain Wall Motion (DWM), and Spin-CMOS implementations. Our evaluation indicates that the proposed SW FA consumes 22.5% and 43% less energy than the direct SW gate based and 22nm CMOS counterparts, respectively. Moreover it exhibits a more than 3 orders of magnitude smaller energy consumption when compared with state-of-the-art MTJ, SHE, DWM, and Spin-CMOS based FAs, and outperforms its contenders in terms of area by requiring at least 22% less chip real-estate. Abdulqader Nael Mahmoud, Frederic Vanderveken, Florin Ciubotaru, Christoph Adelmann, Sorin Cotofana, Said Hamdioui |
ISCAS | 5 |
| 2021 | Graphene-Based Artificial Synapses with Tunable PlasticityabstractDesign and implementation of artificial neuromorphic systems able to provide brain akin computation and/or bio-compatible interfacing ability are crucial for understanding the human brain’s complex functionality and unleashing brain-inspired computation’s full potential. To this end, the realization of energy-efficient, low-area, and bio-compatible artificial synapses, which sustain the signal transmission between neurons, is of particular interest for any large-scale neuromorphic system. Graphene is a prime candidate material with excellent electronic properties, atomic dimensions, and low-energy envelope perspectives, which was already proven effective for logic gates implementations. Furthermore, distinct from any other materials used in current artificial synapse implementations, graphene is biocompatible, which offers perspectives for neural interfaces. In view of this, we investigate the feasibility of graphene-based synapses to emulate various synaptic plasticity behaviors and look into their potential area and energy consumption for large-scale implementations. In this article, we propose a generic graphene-based synapse structure, which can emulate the fundamental synaptic functionalities, i.e., Spike-Timing-Dependent Plasticity (STDP) and Long-Term Plasticity . Additionally, the graphene synapse is programable by means of back-gate bias voltage and can exhibit both excitatory or inhibitory behavior. We investigate its capability to obtain different potentiation/depression time scale for STDP with identical synaptic weight change amplitude when the input spike duration varies. Our simulation results, for various synaptic plasticities, indicate that a maximum 30% synaptic weight change and potentiation/depression time scale range from [-1.5 ms, 1.1 ms to [-32.2 ms, 24.1 ms] are achievable. We further explore the effect of our proposal at the Spiking Neural Network (SNN) level by performing NEST-based simulations of a small SNN implemented with 5 leaky-integrate-and-fire neurons connected via graphene-based synapses. Our experiments indicate that the number of SNN firing events exhibits a strong connection with the synaptic plasticity type, and monotonously varies with respect to the input spike frequency. Moreover, for graphene-based Hebbian STDP and spike duration of 20ms we obtain an SNN behavior relatively similar with the one provided by the same SNN with biological STDP. The proposed graphene-based synapse requires a small area (max. 30 nm 2 ), operates at low voltage (200 mV), and can emulate various plasticity types, which makes it an outstanding candidate for implementing large-scale brain-inspired computation systems. He Wang 0013, Nicoleta Cucu Laurenciu, Yande Jiang, Sorin Cotofana |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2021 | Spin Wave Normalization Toward All Magnonic CircuitsabstractThe key enabling factor for Spin Wave (SW) technology utilization for building ultra low power circuits is the ability to energy efficiently cascade SW basic computation blocks. SW Majority gates, which constitute a universal gate set for this paradigm, operating on phase encoded data are not input output coherent in terms of SW amplitude. Thus, their cascading requires information representation conversion from SW to voltage and back, which is by no means energy effective. In this paper, a novel conversion free SW gate cascading scheme is proposed that achieves SW amplitude normalization by means of a directional coupler. After introducing the normalization concept, we utilize it in the implementation of three simple circuits and, to demonstrate its bigger scale potential, of a 2-bit inputs SW multiplier. The proposed structures are validated by means of the Object Oriented Micromagnetic Framework (OOMMF) and GPU-accelerated Micromagnetics (MuMax3). Furthermore, we assess the normalization induced energy overhead and demonstrate that the proposed approach consumes 1.25× to 1.5× less energy when compared with the transducers based conventional counterpart. Finally, we introduce a normalization based SW 2-bit inputs multiplier design and compare it with functionally equivalent SW transducer based and 16nm CMOS designs. Our evaluation indicates that the proposed approach provided 1.34× and 6.25× energy reductions when compared with the conventional approach and 16nm CMOS counterpart, respectively, which demonstrates that our proposal is energy effective and opens the road towards the full utilization of the SW paradigm potential and the development of SW only circuits. Abdulqader Nael Mahmoud, Frederic Vanderveken, Christoph Adelmann, Florin Ciubotaru, Sorin Cotofana, Said Hamdioui |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | Efficient Computation Reduction in Bayesian Neural Networks Through Feature Decomposition and MemorizationabstractThe Bayesian method is capable of capturing real-world uncertainties/incompleteness and properly addressing the overfitting issue faced by deep neural networks. In recent years, Bayesian neural networks (BNNs) have drawn tremendous attention to artificial intelligence (AI) researchers and proved to be successful in many applications. However, the required high computation complexity makes BNNs difficult to be deployed in computing systems with a limited power budget. In this article, an efficient BNN inference flow is proposed to reduce the computation cost and then is evaluated using both software and hardware implementations. A feature decomposition and memorization (DM) strategy is utilized to reform the BNN inference flow in a reduced manner. About half of the computations could be eliminated compared with the traditional approach that has been proved by theoretical analysis and software validations. Subsequently, in order to resolve the hardware resource limitations, a memory-friendly computing framework is further deployed to reduce the memory overhead introduced by the DM strategy. Finally, we implement our approach in Verilog and synthesize it with a 45-nm FreePDK technology. Hardware simulation results on multilayer BNNs demonstrate that, when compared with the traditional BNN inference method, it provides an energy consumption reduction of 73% and a 4× speedup at the expense of 14% area overhead. Xiaotao Jia, Jianlei Yang 0001, Runze Liu 0001, Sorin Cotofana, Weisheng Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | n-bit Data Parallel Spin Wave Logic GateabstractDue to their very nature, Spin Waves (SWs) created in the same waveguide, but with different frequencies, can coexist while selectively interacting with their own species only. The absence of inter-frequency interferences isolates input data sets encoded in SWs with different frequencies and creates the premises for simultaneous data parallel SW based processing without hardware replication or delay overhead. In this paper we leverage this SW property by introducing a novel computation paradigm, which allows for the parallel processing of n-bit input data vectors on the same basic SW based logic gate. Subsequently, to demonstrate the proposed concept, we present 8-bit parallel 3-input Majority gate implementation and validate it by means of Object Oriented MicroMagnetic Framework (OOMMF) simulations. To evaluate the potential benefit of our proposal we compare the 8-bit data parallel gate with equivalent scalar SW gate based implementation. Our evaluation indicates that 8-bit data 3-input Majority gate implementation requires 4.16x less area than the scalar SW gate based equivalent counterpart while preserving the same delay and energy consumption figures. Abdulqader Nael Mahmoud, Frederic Vanderveken, Florin Ciubotaru, Christoph Adelmann, Sorin Cotofana, Said Hamdioui |
DATE | 5 |
| 2020 | Evolutionary bin packing for memory-efficient dataflow inference acceleration on FPGAabstractConvolutional Neural Network (CNN) dataflow inference accelerators implemented in Field-Programmable Gate Arrays (FPGAs) have demonstrated increased energy efficiency and lower latency compared to CNN execution on CPUs or GPUs. However, the complex shapes of CNN parameter memories do not typically map well to FPGA On-Chip Memories (OCM), which results in poor OCM utilization and ultimately limits the size and types of CNNs which can be effectively accelerated on FPGAs. In this work, we present a design methodology that improves the mapping efficiency of CNN parameters to FPGA OCM. We frame the mapping as a bin packing problem and determine that traditional bin packing algorithms are not well suited to solve the problem within FPGA- and CNN-specific constraints. We hybridize genetic algorithms and simulated annealing with traditional bin packing heuristics to create flexible mappers capable of grouping parameter memories such that each group optimally fits FPGA on-chip memories. We evaluate these algorithms on a variety of FPGA inference accelerators. Our hybrid mappers converge to optimal solutions in a matter of seconds for all CNN use-cases, achieve an increase of up to 65% in OCM utilization efficiency for deep CNNs, and are up to 200× faster than current state-of-the-art simulated annealing approaches. Mairin Kroes, Lucian Petrica, Sorin Cotofana, Michaela Blott |
GECCO | 3 |
| 2020 | 4-output Programmable Spin Wave Logic GateabstractTo bring Spin Wave (SW) based computing paradigm into practice and develop ultra low power Magnonic circuits and computation platforms, one needs basic logic gates that operate and can be cascaded within the SW domain without requiring back and forth conversion between the SW and voltage domains. To achieve this, SW gates have to possess intrinsic fanout capabilities, be input-output data representation coherent, and reconfigurable. In this paper, we address the first and the last requirements and propose a novel 4-output programmable SW logic gate. First, we introduce the gate structure and demonstrate that, by adjusting the gate output detection method, it can parallelly evaluate any 4-element subset of the 2-input Boolean function set {(N)AND, (N)OR, and X(N)OR}. Furthermore, we adjust the structure such that all its 4 outputs produce SWs with the same energy and demonstrate that it can evaluate Boolean function sets while providing fanout capabilities ranging from 1 to 4. We validate our approach by instantiating and simulating different gate configurations such as 4-output AND/OR, 4-output XOR/XNOR, output energy balanced 4-output AND/OR, and output energy balanced 4-output XOR/XNOR by means of Object Oriented Micromagnetic Framework (OOMMF) simulations. Finally, we evaluate the performance of our proposal in terms of delay and energy consumption and compare it against existing state-of-the-art SW and 16 nm CMOS counterparts. The results indicate that for the same functionality, our approach provides 3× and 16× energy reduction, when compared with conventional SW and 16 nm CMOS implementations, respectively. Abdulqader Nael Mahmoud, Frederic Vanderveken, Christoph Adelmann, Florin Ciubotaru, Said Hamdioui, Sorin Cotofana |
ICCD | 6 |
| 2020 | Memristive Oscillatory Circuits for Resolution of NP-Complete Logic Puzzles: Sudoku CaseabstractMemristor networks are capable of low-power and massive parallel processing and information storage. Moreover, they have presented the ability to apply for a vast number of intelligent data analysis applications targeting mobile edge devices and low power computing. Beyond the memory and conventional computing architectures, memristors are widely studied in circuits aiming for increased intelligence that are suitable to tackle complex problems in a power and area efficient manner, offering viable solutions oftenly arriving also from the biological principles of living organisms. In this paper, a memristive circuit exploiting the dynamics of oscillating networks is utilized for the resolution of very popular and NP-complete logic puzzles, like the well-known “Sudoku”. More specifically, the proposed circuit design methodology allows for appropriate usage of interconnections' advantages in a oscillation network and of memristor's switching dynamics resulting to logic-solvable puzzle-instances. The reduced complexity of the proposed circuit and its increased scalability constitute its main advantage against previous approaches and the broadly presented SPICE based simulations provide a clear proof of concept of the aforementioned appealing characteristics. Theodoros Panagiotis Chatzinikolaou, Iosif-Angelos Fyrigos, Rafailia-Eleni Karamani, Vasileios G. Ntinas, Giorgos Dimitrakopoulos, Sorin Cotofana, Georgios Ch. Sirakoulis |
ISCAS | 6 |
| 2020 | Ultra-Compact, Entirely Graphene-Based Nonlinear Leaky Integrate-and-Fire Spiking NeuronabstractDesigning and implementing artificial neuromorphic systems, which can provide biocompatible interfacing, or the human brain akin ability to efficiently process information, is paramount to the understanding of the human brain complex functionality. Energy-efficient, low-area, and biocompatible artificial neurons are key ubiquitous components of any large scale neural systems. Previous CMOS-based neurons implementations suffer from scalability drawbacks and cannot naturally mimic the analog behavior. Memristor and phase-changed neurons have variability-induced instability drawbacks, and usually rely on additional CMOS circuitry. However, graphene, despite its ballistic transport, inherently analog nature, and biocompatibility, which provide natural support for biologically plausible neuron implementations has only been considered for Boolean logic implementations. In this paper, we propose an ultra-compact, all graphene-based nonlinear Leaky Integrate-and-Fire spiking neuron. By means of SPICE simulations, we validate its basic functionality and investigate the output spikes response under stochastic noisy input spike trains with a variable firing rate, from 20 to 200 spikes per second. Simulation results indicate neuron robustness to noisy scenarios, and neuronal output firing regularity. The small area and the low energy consumption, due to 200mV supply voltage operation, can benefit the implementation of large scale neural networks, and the biologically plausible operating conditions (e.g., 2ms and 100mV spike duration and amplitude), can promote the interfacebility of graphene-based artificial neurons with biological counterparts. He Wang 0013, Nicoleta Cucu Laurenciu, Yande Jiang, Sorin Cotofana |
ISCAS | 4 |
| 2019 | A Pragmatic Gaze on Stochastic Resonance Based Variability Tolerant Memristance EnhancementabstractStochastic Resonance (SR) is a nonlinear system specific phenomenon, which was demonstrated to lead to system unexpected (counter-intuitive) performance improvements under certain noise conditions. Memristor, on the other hand, is a fundamentally nonlinear circuit element, thus susceptible to benefit from SR, which recently came in the spotlight of the emerging technologies potential candidates. However, at this time, the variability exhibited by manufactured memristor devices within the same array constitutes the main hurdle in the road towards the commercialisation of memristor-based memories and/or computing units. Thus, in this paper, memristor SR effects are explored, assuming various memristor models, and SR-based memristance range enhancement, tolerant to device-to-device variability, is demonstrated. Our experiments reveal that SR can induce significant RMAX/RMINratio increase under up to 60% variability, getting as high as 3.4× for 29 dBm noise power. Vasileios G. Ntinas, Antonio Rubio 0001, Georgios Ch. Sirakoulis, Sorin Cotofana |
ISCAS | 4 |
| 2019 | Atomistic-Level Hysteresis-Aware Graphene Structures Electron Transport ModelabstractHysteretic behavior has been experimentally observed in graphene-based structures and has a major influence on graphene surface potential and gate field modulation ability. Thus, a graphene electronic transport modelling methodology, which incorporates hysteresis effects is crucial in order to properly assess gated-controlled graphene structures response and performance. To this end, we propose an atomistic-level electronic transport model, which is non restricted to rectangular graphene geometries and captures hysteretic effects caused by near-interfacial traps, provided that interface traps trapping/detrapping time constant and density are known. We apply the model on a rectangular graphene shape and validate our results against experimentally measured drain current vs. top gate voltage hysteresis curves. Moreover, to demonstrate model's versatility we consider two non-rectangular Graphene NanoRibbons (GNRs) and investigate their hysteresis behaviour. Our experiments indicate good agreement between simulated and measured results, which qualifies the model appropriate for traps-aware exploration of the conduction behaviour of graphene-based devices and circuits. He Wang 0013, Nicoleta Cucu Laurenciu, Yande Jiang, Sorin Cotofana |
ISCAS | 4 |
| 2018 | On Carving Basic Boolean Functions on Graphene Nanoribbons Conduction MapsabstractAs CMOS feature size approaches atomic dimensions, unjustifiable static power, reliability, and economic implications are exacerbating, prompting for research and development on new materials, devices, and/or computation paradigms. Within this context, Graphene Nanoribbons (GNRs), owing to graphene's excellent electronic properties, may serve as basic blocks for carbon-based nanoelectronics. En route to GNR-based logic circuits, the ability to externally control GNRs' conduction to map a basic Boolean logic function onto its electrical characteristics, with a high ION/IOFFratio, and uncompromised carriers mobility, is the main desideratum. To this end, we augment a trapezoidal GNR with top gates as controlling inputs, and investigate its conductance G by means of the NEGF-Landauer formalism. Further, we demonstrate that the butterfly GNR can exhibit conduction maps (high G for logic “1”, and low G for logic “0”) capturing the functionality of 2 and 3-input Boolean gates, by properly adjusting its topology and dimensions. Our simulations prove butterfly GNR structure capability to capture basic Boolean logic transfer functions, while potentially providing 30× and 3000× smaller propagation delay and gate active area, respectively, when compared to 15 nm CMOS equivalent counterparts, establishing GNR's potential as basic building block for future graphene-based logic gates. Yande Jiang, Nicoleta Cucu Laurenciu, Sorin Cotofana |
ISCAS | 3 |
| 2017 | A Mixed-Size Monolithic 3D Placer with 2D Layout InheritanceabstractMonolithic 3D IC is a high integration density emerging technology in the age of both "More Moore" and "More-than-Moore". In this paper, we propose a novel method of generating mixed-size 3D placement based on transforming a 2D placement result. Experimental results indicate that, when compared with the input 2D placement, the 3D placer can reduce the wirelength by 57%, and provide a four-layer 3D chip footprint of about one quarter of the 2D counterpart. Moreover, our placer can preserve the layout information from the 2D placement input, which means that the 2D placement quality can be inherited in the 3D placement results. Compared with an analytical wirelength-driven placer, our placer achieves 34% benefit on 2D layout inheritance and 12% benefit on runtime with acceptable (4%) wirelength cost. Yao Wang 0002, Yang Guo 0003, Sorin Cotofana |
ACM Great Lakes Symposium on VLSI | 4 |
| 2017 | LDPC-Based Adaptive Multi-Error Correction for 3D MemoriesabstractIn this paper we introduce a novel error resilient memory architecture potentially applicable to a large range of memory technologies. In contrast with state of the art memory error correction schemes, which rely on (extended Hamming) Error Correcting Codes (ECC), we make use of Low Density Parity Check (LDPC) codes due to their close to the Shannon performance limit error correction capabilities. To allow for a cost-effective implementation we build our approach on top of a 3D memory organization which inherently fast and customizable wide-I/O vertical access allows for a smooth transfer of the required LDPC long code-words to/from an error correction dedicated die. To make the error correction process transparent to the memory users, e.g., processing cores, we propose an online memory scrubbing policy that performs the LDPC-based error detection and correction decoupled from the normal memory operation. For evaluation purposes we consider 3D memories protected by the proposed LDPC mechanism with various data width codes implementations. Simulation results indicate that our proposal clearly outperforms state of the art ECC schemes with fault tolerance improvements by a 4710× factor being obtained when compared to extended Hamming ECC. Furthermore, we evaluate instances of the proposed memory concept equipped with different LDPC codecs implemented on a commercial 40nm low-power CMOS technology and evaluate them on actual memory traces in terms of error correction capability, area, latency, and energy. Our results indicate that the LDPC protected memories offer substantially improved error correction capabilities, when compared to state of the art extended Hamming ECC, being able to assure clean runs for memory error rates α Mihai Lefter, George Razvan Voicu, Thomas Marconi, Valentin Savin, Sorin Cotofana |
ICCD | 5 |
| 2017 | Towards Maximum Utilization of Remained Bandwidth in Defected NoC LinksabstractTo maximize the utilization of the available networks-on-chip (NoCs) link bandwidth, partially faulty links with low fault level should be utilized while heavily defected (HD) links should be deactivated and dealt with by means of a fault tolerant routing algorithm. To reach this target, we make the following contributions in this paper: 1) we propose a flit serialization (FS) method to efficiently utilize partially faulty links. The FS approach divides the links into a number of equal width sections, and serializes sections of adjacent flits to transmit them on all fault-free link sections to mitigate the unbalance between the flit size and the actual link bandwidth; 2) we propose the link augmentation with one redundant section as a low cost mechanism to mitigate the FS drawback that a link's available bandwidth is reduced even if it contains only one faulty wire; and 3) we deactivate HD links when their fault level exceed a certain threshold to diminish congestion caused by HD links. The optimal threshold is derived by comparing the zero load packet transmission latency on the HD links and that on the shortest alternative path. Our proposal is evaluated with synthetic traffic and PARSEC benchmarks. Experimental results indicate that the FS method can achieve lower area*power/saturation_throughput value than all state of the art link fault tolerant strategies. With a redundant section in each link, the NoC saturation throughput can be largely improved than just utilizing FS, e.g., 18% when 10% of the NoC wires are broken. Simulation results we obtained at various wire broken rate configurations indicate that we achieve the highest saturation throughput if 4- or 8-section links with a flit transmission latency longer than four cycles are deactivated. Changlin Chen, Yaowen Fu, Sorin Cotofana |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Flexible, Cost-Efficient, High-Throughput Architecture for Layered LDPC Decoders with Fully-Parallel Processing UnitsabstractIn this paper, we propose a layered LDPC decoder architecture targeting flexibility, high-throughput, low cost, and efficient use of the hardware resources. The proposed architecture provides full design time flexibility, i.e., it can accommodate any Quasi-Cyclic (QC) LDPC code, and also allows redefining a number of parameters of the QC-LDPC code at the run time. The main novelty of the paper consists of: (1) a new low-cost processing unit that merges the logical functionalities of the Variable-Node Unit (VNU) and the A Posteriori Log-Likelihood Ratio (AP-LLR) unit in an efficient way, (2) a high speed, low-cost Check-Node Unit (CNU) architecture, which is executed twice at each iteration in order to complete the computation of the check-node messages, (3) a splitting of the iteration processing in two perfectly symmetric stages, executed in two consecutive clock cycles, each one using exactly the same processing resources, the processing load is perfectly balanced between the two clock cycles, thus yielding an optimal clock frequency. Synthesis results targeting a 65nm CMOS technology for a (3,6)-regular (648,1296) Quasi-Cyclic LDPC code and for the WiMax (1152,2304) irregular QC-LDPC code show significant improvements in terms of area and throughput compared to the baseline architecture discussed in this paper, as well as several state of the art implementations. Truong Nguyen 0001, Manuel Pezzin, Valentin Savin, David Declercq, Sorin Cotofana |
DSD | 6 |
| 2015 | Enabling vertical wormhole switching in 3D NoC-bus hybrid systems
Changlin Chen, Marius Enachescu, Sorin Cotofana |
DATE | 3 |
| 2015 | Hybrid adaptive clock management for FPGA processor acceleration
Alexandru Gheolbanoiu, Lucian Petrica, Sorin Cotofana |
DATE | 3 |
| 2015 | Dynamic Bitstream Length Scaling Energy Effective Stochastic LDPC DecodingabstractStochastic Computing (SC) is an attractive solution for implementing Low Density Parity Codes (LDPC) decoders due to its fault tolerance capability and low hardware requirements. However, in practical implementations, SC efficiency is limited by the Stochastic Bitstream (SB) length and by the computation inaccuracies due to non-unique SB representations. In this paper, rather than statically fixing the SB length at run-time, we propose a Dynamic Bitstream Length Scaling (DBLS) technique, which adjusts on-the-fly the SB length such that Quality of Service requirements for energy efficient LDPC decoding are fulfilled. In this way, depending on the communication channel condition, different SB lengths are adaptively utilized such that the best decoding performance vs energy consumption tradeoff is achieved. To evaluate the DBLS practical implications we selected an (1296,648) LDPC with dv=3 and dc=6 and implemented our approach and the best state-of-the-art stochastic LDPC decoder with 64-bit edge memory on a Virtex-7 FPGA. Experimental results indicate that our proposal requires 9% more FFs and 3% more LUTs while diminishing the energy consumption by 31-80% and providing 1.5-5.1x higher throughput. Thomas Marconi, Sorin Cotofana |
ACM Great Lakes Symposium on VLSI | 2 |
| 2015 | ROST-C: Reliability driven optimisation and synthesis techniques for combinational circuitsabstractTraditional logic synthesis methodologies are driven by timing, power, and area constraints. However, due to aggressive technology shrinking and lower power requirements, circuit reliability is fast turning out to be yet another major constraint in the VLSI design flow. Soft errors, which traditionally affected only the memories, are now also resulting in logic circuit reliability degradation. In this paper, we present a systematic and integrated methodology to address and improve the combinational circuit reliability measured in terms of Soft Error Rate (SER). The proposed SER reduction framework makes use of rewriting based logic optimisation technique which employs local transformations. The main idea behind our proposal is to replace parts of the circuit with functionally equivalent but more reliable counterparts chosen from a pre-computed subset of Negation-Permutation-Negation (NPN) classes of 4-variable functions. Cut enumeration and Boolean matching driven by reliability aware optimisation algorithm are used to identify best possible replacement candidates. Our experiments on a set of MCNC benchmark circuits indicate that the proposed framework can achieve up to 75% reduction of output error probability. On average, about 14% SER reduction is obtained at the expense of very low area overhead of 6.57% that results in 13.52% higher power consumption. Satish Grandhi, David McCarthy, Christian Spagnol, Emanuel M. Popovici, Sorin Cotofana |
ICCD | 5 |
| 2015 | Low cost and energy, thermal noise driven, probability modulated random number generatorabstractRandom Number Generators (RNGs) constitute the essential foundation of many applications, e.g., cryptographic technology, and statistic based computing. The vast majority of RNG proposals, generate binary sequences with uncorrelated and equiprobable bits. However, in certain applications, such as stochastic computing, the binary sequence is required to be generated not with 0.5 logic “1” probability, but with a custom specified one. This paper presents a probability modulated True Random Number Generator (TRNG), that produces binary sequences with a desired probability of logic “1”, according to a value resident in a register. The proposed circuit relies on the thermal noise as random signal source and it was implemented in 65nm CMOS technology. Simulation results reveal a quasi-linear dependence, between the output logic “1” probability sampled at 1GHz, and the modulating voltage (obtained from the register value by mean of a D/A converter) over a range of 20mV. For a desired probability of 0.5, the sequences randomness was validated by the NIST tests. Nicoleta Cucu Laurenciu, Sorin Cotofana |
ISCAS | 2 |
| 2015 | A shared polyhedral cache for 3D wide-I/O multi-core computing platformsabstract3D-Stacked IC (3D-SIC) based on Through-Silicon-Vias (TSV) is an emerging technology that enables heterogeneous integration and high bandwidth low latency interconnection. In this paper we propose a 3D novel cache architecture that leverages a wide TSV-based data link distributed on the entire memory array to support two orthogonal interfaces: (i) a vertical one, with a large data width, and, (ii) a side one, with a lower data width, but with more bank-type access ports, which reduces the bank conflict probability. Our simulations indicate that our proposal substantially outperforms planar counterparts in terms of access time, energy, and footprint while providing high bandwidth, low bank conflict rate, and an enriched access mechanism set. Mihai Lefter, George Razvan Voicu, Sorin Cotofana |
ISCAS | 3 |
| 2014 | Towards energy effective LDPC decoding by exploiting channel noise variabilityabstractIn communication systems, channel quality variation is a well known phenomenon, which fundamentally influences the decoding process. While most of the time, the transmission takes place in good signal to noise conditions, to satisfy QoS requirements in all cases, telecom platforms rely on largely over-designed hardware, which may result in energy waste during most of their operation. In this paper we propose to exploit the channel noise variability and adapt the platform operation conditions such that QoS requirements are satisfied with the minimum energy consumption. In particular, we propose a technique to exploit channel noise variability towards energy effective LDPC decoding amenable to low-energy operation. Endowed with the channel noise variability knowledge, our technique adaptively tunes the operating voltage at runtime, aiming to achieve the optimal tradeoff between decoder performance and power con-sumption, while fulfilling the QoS requirements. To demonstrate the capabilities of our proposal we implemented it and other state of the art energy reduction methods in conjunction with a fully parallel LDPC decoder on a Virtex-6 FPGA. Our experiments indicate that the proposed technique outperforms state of the art counterparts, in terms of energy reduction, with 71% to 76% and 15% to 28%, w.r.t. early termination without and with DVS, respectively, while maintaining the targeted decoding robustness. Moreover, the measurements suggest that in certain conditions Degradation Stochastic Resonance occurs, i.e., the energy consumption is unexpectedly diminished due to the fact that unpredictable underpowered components facilitate rather than impede the decoding process. Thomas Marconi, Christian Spagnol, Emanuel M. Popovici, Sorin Cotofana |
VLSI-SoC | 4 |
| 2014 | Critical transistors nexus based circuit-level aging assessment and prediction
Nicoleta Cucu Laurenciu, Sorin Cotofana |
J. Parallel Distributed Comput. | 2 |
| 2014 | Analysis of the impact of spatial and temporal variations on the stability of SRAM arrays and the mitigation technique using independent-gate devices
Yao Wang 0002, Sorin Cotofana |
J. Parallel Distributed Comput. | 2 |
| 2013 | An effective New CRT based reverse converter for a novel moduli set {22n+1 - 1, 22n+1, 22n - 1}abstractIn this paper, a novel 3-moduli set {22n+1− 1,22n+1,22n− 1}, which has larger dynamic range when compared to other existing 3-moduli sets is proposed. After providing a proof that this moduli set always results in legitimate RNS, we subsequently propose an associated reverse converter based on the New Chinese Remainder Theorem. The proposed reverse converter has a delay of (4n + 6)tfawith an area cost of (8n + 2)FAs and (An — 2)HAs, where FA, HA, and tFArepresent Full Adder, Half Adder, and delay of a Full Adder, respectively. We compared the proposed reverse converter with state of the art converters for similar or equal dynamic range RNS and our analysis indicate that for the same dynamic range the best converter requires (16n + l) FAs and exhibits a delay of (8n + 2)tFA. This indicates that, theoretically speaking, our proposal achieves about 37.5% area reduction and it is about 2 times faster than equivalent state of the art converters. Moreover, the product of the area with the square of the delay is improved up to about 84.4% over the state of the art. Edem Kwedzo Bankas, Kazeem Alagbe Gbolagade, Sorin Cotofana |
ASAP | 3 |
| 2013 | 3D stacked wide-operand adders: A case studyabstractIn this paper, we address the design of wide-operand addition units in the context of the emerging Through-Silicon Vias (TSV) based 3D Stacked IC (3D-SIC) technology. To this end we first identify and classify the potential of the direct folding approach on existing fast prefix adders, and then discuss the cost and performance of each strategy. Our analysis identifies as a major direct folding drawback the utilization of different structures on each tier. Thus, in order to alleviate this, we propose a novel 3D Stacked Hybrid Prefix/Carry-Select Adder with identical tier structure, which potentially makes the manufacturing of hardware wide-operand adders a reality. Such an N-bit carry select adder can be implemented with K identical tier stacked ICs, where each tier contains two N/K-bit fast prefix adders operating in parallel according to the computation anticipation principle. Their carry-out signals are cascaded through TSVs in order to perform the selection of the sums accordingly, which results in a delay with the asymptotic notation of O(log(N/K) + K). To evaluate the practical implications of direct folding and of the hybrid prefix/carry-select approaches we perform a thorough case study of 65 nm CMOS 3D adder implementations for different operand sizes and number of tiers, and analyze various possible design tradeoffs. Our simulations indicate the hybrid prefix/carry-select approach can achieve speed gains over 3D folding based designs of between 29% and 54%, for 512-bit up to 4096-bit adders, respectively. Even though 3D folding requires less real estate, when considering a more appropriate metric for 3D design, i.e., delay-footprint-cost product, the hybrid prefix/carry-select approach substantially outperforms the folding one and provides delay-footprint-cost reductions between 17.97% and 94.05%. George Razvan Voicu, Mihai Lefter, Marius Enachescu, Sorin Cotofana |
ASAP | 4 |
| 2013 | Is TSV-based 3D integration suitable for inter-die memory repair?abstractIn this paper we address lower level issues related to 3D inter-die memory repair in an attempt to evaluate the actual potential of this approach for current and foreseeable technology developments. We propose several implementation schemes both for inter-die row and column repair and evaluate their impact in terms of area and delay. Our analysis suggests that current state-of-the-art TSV dimensions allow inter-die column repair schemes at the expense of reasonable area overhead. For row repair, however, most memory configurations require TSV dimensions to scale down at least with one order of magnitude in order to make this approach a possible candidate for 3D memory repair. We also performed a theoretical analysis of the implications of the proposed 3D repair schemes on the memory access time, which indicates that no substantial delay overhead is expected and that many delay versus energy consumption tradeoffs are possible. Mihai Lefter, George Razvan Voicu, Mottaqiallah Taouil, Marius Enachescu, Said Hamdioui, Sorin Cotofana |
DATE | 6 |
| 2013 | An Effective Routing Algorithm to Avoid Unnecessary Link Abandon in 2D Mesh NoCsabstractIn NoCs where each interconnection between neighboring routers is composed of a pair of unidirectional links, a broken link usually leads to the abandon of the entire interconnection, even if the other one is still functional. In this paper, we propose a fault tolerant Routing Algorithm (RA) which can efficiently utilize these fault free links when their pair broken links have available misrouting-contour sides. Constraints on the usage of Virtual Channels are adaptively applied according to the fault distribution, to avoid deadlock and unnecessary resource reservation. When compared with solid fault region tolerant RAs, which always abandon the entire interconnection, the proposed algorithm has twice higher saturation point under synthetic uniform traffics, and can on average diminish the execution time overhead for the evaluated applications, sample and sat ell, by 62.6% and 76.6%, respectively. Our experiments indicate that the embedding of the proposed algorithm into a baseline router increases the area cost and power consumption by 7.43% and 4.43%, respectively, which is not that significant given that the platform area is usually dominated by the computing cores area. Changlin Chen, Sorin Cotofana |
DSD | 2 |
| 2013 | Lifetime reliability assessment with aging information from low-level sensorsabstractAggressive technology scaling has led Integrated Circuits (ICs) suffer from ever-increasing wearout effects. As a consequence, Dynamic Reliability Management (DRM) becomes an essential approach to assure IC's lifetime reliability. Accurate and efficient reliability modeling from low-level aging sensor measurements is critical to DRM systems. This work presents a Time-Sharing Sensing (TSS) method for $V_{th}$-sensor based DRM to assess the dynamic NBTI-induced degradation experienced by the circuit under monitoring. SPICE simulation results suggest that the proposed TSS method can accurately capture the circuit reliability status under random stress conditions. Yao Wang 0002, Sorin Cotofana |
ACM Great Lakes Symposium on VLSI | 2 |
| 2013 | Ultra low power NEMFET based logicabstractIn this paper, we introduce a Nano-Electro-Mechanical Field Effect Transistor (NEMFET) based logic family tailored to the implementation of low speed and ultra low energy functional units and processors. Basic Boolean gates implemented with NEMFETs only are analysed and compared against equivalent CMOS realisations. Our simulations suggest that the proposed short-circuit current free NEMFET gates exhibit up to 10x dynamic energy reduction and up to 2 orders of magnitude less leakage, at the expense of 10 to 20x slower operation, when compared with CMOS counterparts. We also analyse the fan-in influence on gate performance and observe that NEMFET the gate energy advantage increases with fan-in. Finally, we consider a 3D-Stacked hybrid NEMFET-CMOS computation platform running a heartbeat rate monitor application and demonstrate that NEMFET based logic is an enabling factor for the implementation of “zero-energy” operated systems. Marius Enachescu, Mihai Lefter, Antonios Bazigos, Adrian M. Ionescu, Sorin Cotofana |
ISCAS | 5 |
| 2013 | A direct measurement scheme of amalgamated aging effects with novel on-chip sensorabstractAggressive technology scaling has led to a significant reduction of device reliability. As a consequence Integrated Circuits (ICs) reliability became a major issue and Dynamic Reliability Management (DRM) schemes have been proposed to assure ICs' lifetime reliability. Though, up to date, various aging sensors have been proposed, few of them can provide real quantitative aging measurements. In view of this, we propose a direct measuring scheme by using the drain current as aging indicator. We designed a novel on-chip aging sensor able to detect the amalgamated aging effects of ICs caused by joint failure mechanisms. This is achieved by detecting the peak power supply current (Ipp) degradation from the device and/or circuit, which is a signature of the total drain current. Unlike the existing aging sensors which indirectly estimate the aging status of a device, the proposed sensor allows for direct aging assessment for single device and/or circuit blocks. Simulation results using the TSMC 65nm technology indicate that the proposed sensor can operate at 1GHz. Accelerated test simulation in Cadence for a set of ISCAS85 benchmark circuits indicates that the drain current exhibits a similar aging rate as the threshold voltage for the entire circuit lifetime, but with a better sensitivity towards the End-of-Life (EOL), which demonstrates the validity and practical relevance of the proposed aging monitoring framework. Nicoleta Cucu Laurenciu, Yao Wang 0002, Sorin Cotofana |
VLSI-SoC | 3 |
| 2012 | A 3D stacked high performance scalable architecture for 3D Fourier TransformabstractThis paper proposes and evaluates a novel high-performance systolic architecture for 3D Fourier Transform specially tailored for 3D stacking integration with Through Silicon Vias. Our cuboid-shaped systolic network of orthogonally connected processing elements makes use of the DFT algorithm to compute an N1×N2×N3-point 3D-FT with an asymptotic time complexity of O(N1+N2+N3) multiplications. When compared with state-of-the-art 3D-FFT implementation on the Anton machine, a physical synthesized implementation of our architecture on the same 90nm technology node achieves 7.73× and 5.88× speed improvement when computing 16×1 6×16 and 32×3 2×32 FT, respectively. George Razvan Voicu, Marius Enachescu, Sorin Cotofana |
ICCD | 3 |
| 2012 | Is the road towards "Zero-Energy" paved with NEMFET-based power management?abstractIn this paper we explore the potential that the use of state-of-the-art Nano-Electro-Mechanical (NEM) devices, i.e., NEMFETs and NEM Relays, in the implementation of power management circuitry, in combination with efficient energy harvesters through 3D stacking integration, have in meeting the tight energy budgets of “Zero-Energy” autonomous sensor systems. We propose various 3D hybrid embodiments of an openMPS430 embedded processor augmented with NEMFET and/or NEM Relay based power management mechanisms and investigate their energy consumption when executing a heart beat detection application. Our investigations indicate that the hybrid NEMFET-oriented approach, which relies on sleep transistors and associated management logic implemented on a dedicated NEMFET die, is the most promising in terms of energy consumption and reliability. Moreover, when combined with a thermal energy harvester, potentially implementable on the same die, it can enable the road towards autonomous computing. Marius Enachescu, George Razvan Voicu, Sorin Cotofana |
ISCAS | 3 |
| 2012 | A Novel Flit Serialization Strategy to Utilize Partially Faulty Links in Networks-on-ChipabstractAggressive MOS transistor size scaling substantially increase the probability of faults in NoC links due to manufacturing defects, process variations, and chip wire-out effects. Strategies have been proposed to tolerate faulty wires by replacing them with spare ones or by partially using the defective links. However, these strategies either suffer from high area and power overheads, or significantly increase the average network latency. In this paper, we propose a novel flit serialization method, which divides the links and flits into several sections, and serializes flit sections of adjacent flits to transmit them on all available fault-free link sections to avoid the complete waste of defective links bandwidth. Experimental results indicate that our method reduces the latency overhead significantly and enables graceful performance degradation, when compared with related partially faulty link usage proposals, and saves area and power overheads by up to 29% and 43.1%, respectively, when compared with spare wire replacement methods. Changlin Chen, Ye Lu 0003, Sorin Cotofana |
NOCS | 3 |
| 2011 | Analysis of delay mismatching of digital circuits caused by common environmental fluctuationsabstractEnvironmental conditions are changing all the time along the chip as a consequence of its own activity, provoking deviations on propagation time in digital circuits. In future technologies, the increment of devices sensitivity to environmental fluctuations yields to a wider range of possible time deviations, being for example, in an NOT gate designed in a 16 nm technology 1.6 times larger than for a 45 nm version. But this ratio is different for every circuit cause it depends on its fundamental structure and characteristics. In this paper the tendency of timing parameters deviations due to environmental factors fluctuation and how these deviations have deeper impact on more complex structures are analyzed. It is shown that the internal structure of the logic gates cause a mismatch between logic circuits and in future technologies it will be enlarged. Dennis Andrade, Antonio Rubio 0001, Antonio Calomarde, Sorin Cotofana |
ISCAS | 4 |
| 2011 | An Efficient FPGA Design of Residue-to-Binary Converter for the Moduli Set 2n+1, 2n, 2n-1abstractIn this paper, we propose a novel reverse converter for the moduli set {2n+1,2n,2n-1}. First, we simplify the Chinese Remainder Theorem in order to obtain a reverse converter that uses mod-(2n-1) operations. Next, we present a low complexity implementation that does not require the explicit use of modulo operation in the conversion process and we prove that theoretically speaking it outperforms state of the art equivalent converters. We also implemented the proposed converter and the best equivalent state of the art converters on Xilinx Spartan 3 field-programmable gate array. The results indicate that, on average, our proposal is about 14%, 21%, and 8% better in terms of conversion time, area cost, and power consumption, respectively. Kazeem Alagbe Gbolagade, George Razvan Voicu, Sorin Cotofana |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Memoryless RNS-to-binary converters for the {2n+1 - 1, 2n, 2n - 1} moduli setabstractIn this paper, we propose two novel memoryless reverse converters for the moduli set {2n+1– 1,2n, 2n– 1}. The first proposed converter does not entirely cover the dynamic range while the second proposed converter covers the entire dynamic range. First, we simplify the Chinese Remainder Theorem in order to obtain a reverse converter that utilizes mod-(2n+1– 1) operation. Second, we further reduce the resulting architecture to obtain a reverse converter that uses only carry save adders and carry propagate adders. FPGA implementation results indicate that, on average, the proposed limited dynamic range converter achieves about 42% area reduction. However, the second proposed converter provides only 29.48% area reduction when compared with the most effective equivalent state of the art converter. Both of the proposed converters also exhibit a small speed improvement over the state of the art equivalent converter. Kazeem Alagbe Gbolagade, George Razvan Voicu, Sorin Cotofana |
ASAP | 3 |
| 2010 | A unified addition structure for moduli set {2n-1, 2n, 2n+1} based on a novel RNS representationabstractGiven that modulo 2n±1 are the most popular moduli in Residue Number Systems (RNS), a large variety of modulo 2n±1 adder designs have been proposed based on different number representations. However, in most of the cases, these encodings do not allow the implementation of a unified adder for all the moduli of the form 2n-1, 2n, and 2n+1. In this paper, we address the modular addition issue by introducing a new encoding, namely, the stored-unibit RNS. Moreover, we demonstrate how the proposed representation can be utilized to derive a unified design for the moduli set {2n-1,2n,2n+1}. Our approach enables a unified design for the moduli set adders, which opens the possibility to design reliable RNS processors with low hardware redundancy. Moreover, the proposed representation can be utilized in conjunction with any fast state of the art binary adder without requiring any extra hardware for end-around-carry addition. Somayeh Timarchi, Mahmood Fazlali, Sorin Cotofana |
ICCD | 3 |
| 2010 | An improved RNS reverse converter for the {22n+1-1, 2n, 2n-1} moduli setabstractIn this paper, we propose a novel high speed memoryless reverse converter for the moduli set {22n+1-1, 2n, 2n-1}. First, we simplify the traditional Chinese Remainder Theorem in order to obtain a reverse converter that only requires arithmetic mod-(22n+l-1). Second, we further improve the resulting architecture to obtain a purely adder based reverse converter. The proposed converter has a critical path delay of (7n + 7) Full Adders (FA) while the best state of the art converter for this moduli set requires (10n + 5) FA on the critical path. To validate these results, the converters are implemented in a Standard Cell 0.18-μm CMOS technology and the results assert that, on average, the proposed converter achieves about 19% delay reduction at the expense of less than 3% area increase. Kazeem Alagbe Gbolagade, Ricardo Chaves, Leonel Sousa, Sorin Cotofana |
ISCAS | 4 |
| 2009 | An O(n) Residue Number System to Mixed Radix Conversion TechniqueabstractThis paper investigates the conversion of residue number system (RNS) operands to decimal, which is an important issue concerning the utilization of RNS numbers in digital signal processing applications. In this line of reasoning, we introduce an RNS to mixed radix conversion (MRC) technique, which addresses the computation of mixed radix (MR) digits in such a way that enables the MRC parallelization. Given an RNS with the set of relatively prime integer moduli {mi}i=1;n, the key idea behind the proposed technique is to maximize the utilization of the modulo-miadders and multipliers present in the RNS processor functional units. For an n-digit RNS number X = (x1; x2; x3;hellip; xn) the method requires n iterations. However, at iteration i, the modulo-miunits are utilized for the calculation of the MR digit ai, while the other modulo units are calculating intermediate results required in further iterations. Our approach results in an RNS to MRC with an asymptotic complexity, in terms of arithmetic operations, in the order of O(n), while state of the art MRCs exhibit an asymptotic complexity in the order of O(n2). More in particular, when compared with the best state of the art MRC, our technique reduces the number of arithmetic operations by 5:26% and 38:64% for moduli set of length four and ten, respectively. Kazeem Alagbe Gbolagade, Sorin Cotofana |
ISCAS | 2 |
| 2008 | Compositional, dynamic cache management for embedded chip multiprocessorsabstractThis paper proposes a dynamic cache repartitioning technique that enhances compositionality on platforms executing media applications with multiple utilization scenarios. The repartitioning among scenarios requires a cache flush, thus two undesired effects may occur: (1) the execution of critical tasks may be disturbed and (2) a performance penalty is involved. To cope with these effects we propose a method which: (1) determines, at design time, the cache footprint of each task, such that it creates the premises for critical tasks safety, and reduces the amount of required flush, and (2) enforces these footprints and further decreases the flush penalty, at run-time. We implement our dynamic cache management strategy on a CAKE multiprocessor with 4 Trimedia cores. The experimental workload consists of 6 multimedia applications, each of which formed by multiple tasks belonging to an extended MediaBench suite. For the repartitioned cache we found on average that: (1) the relative variations of critical tasks execution time are less than 0.1%, regardless the scenario switching frequency, (2) for realistic scenario switching frequencies the inter-task cache interference is at most 4% , and (3) the off-chip memory traffic reduces with 60%, and the performance (in cycles per instructions) enhances with 10%, when compared with the shared cache. Anca Mariana Molnos, Marc J. M. Heijligers, Sorin Cotofana |
DATE | 3 |
| 2008 | Bitstream compression techniques for Virtex 4 FPGAsabstractThis paper examines the opportunity of using compression for accelerating the (re)configuration of FPGA devices, focusing on the choice of compression algorithms, and their hardware implementation cost. As our purpose is the acceleration of the configuration process, estimating the decoder speed also plays a major role in our study. We evaluate a wide range of well-established compression algorithms and we also propose two methods specifically developed for compressing FPGA configuration bitstreams, one based on a static dictionary and the other on arithmetic coding. For the arithmetic coding we propose a statistical model that takes advantage of the particularities of the configuration bitstreams of the Virtex 4 FPGA family. We evaluate the efficiency of the proposed methods along with state of the art compression algorithms on a number of benchmark circuits, some selected from the available open source implementations and some synthetically generated. Our evaluations indicate that using modest resources we can achieve parity and even exceed comercial software in terms of compression ratio, and outperform all other traditional algorithms. All our implemented decompressors are shown to use less than 1.5% of the slices available on the FPGA device. Radu Andrei Stefan, Sorin Cotofana |
FPL | 2 |
| 2006 | Compositional, efficient caches for a chip multi-processorabstractIn current multi-media systems major parts of the functionality consist of software tasks executed on a set of concurrently operating processors. Those tasks interfere with each other when they share memory and other hardware components. For instance when the tasks share caches and no precautions are taken they potentially flush each other's data at random. In this case the control over the system performance is lost. However, in media processing the performance must be under tight control. In particular the performance of each individual task must be preserved if the tasks are executed concurrently in arbitrary combinations or if additional tasks are added. A system satisfying this property is addressed as being compositional. This paper proposes a novel cache partitioning technique that enhances compositionality. We assume a cache to be a rectangular array of memory elements arranged in "sets" (rows) and "ways" (columns). We perform two partitioning types. First, each task and each inter-task common data gets an exclusive part of the cache sets. Second, inside the cache sets of common data each task accessing it gets a number of ways. We apply the proposed method on a homogeneous multiprocessor using two applications: H.264 decoding and picture-in-picture-TV. Our experiments indicate that, for both applications, under our partitioning scheme the sum of misses of the individual tasks executed separately and the number of misses of all tasks executed concurrently differs at most by 4%. We conclude that compositionality is achieved within reasonable bounds. Additionally, our technique appears to improve the efficiency of the cache operation Anca Mariana Molnos, Marc J. M. Heijligers, Sorin Cotofana, Jos T. J. van Eijndhoven |
DATE | 3 |
| 2006 | Electron counting based high-radix multiplication in single electron tunneling technologyabstractThis paper investigates the implementation of high-radix multiplication based on the electron counting (EC) paradigm in single electron tunneling (SET) technology. First we propose a multiplication scheme which conceptually speaking follows the structure of traditional full-tree multipliers. The high-radix EC multiplication scheme comprises three steps and of each an implementation is presented. Second, an 8-bit radix 4 EC multiplier is designed and verified by means of simulation. The high-radix multiplication scheme is evaluated in terms of area and delay for different operand sizes and compared with corresponding binary multiplication schemes in SET technology. Both type of implementations prove to have similar delay times, but the EC based scheme requires up to five times less area Cor Meenderinck, Sorin Cotofana |
ISCAS | 2 |
| 2005 | CONAN - A Design Exploration Framework for Reliable Nano-ElectronicsabstractIn this paper we introduce a design methodology that allows the system/circuit designer to build reliable systems out of unreliable nano-scale components. The central point of our approach is a generic (parametrical) architectural template. Configurable nanostructures for reliable nano electronics (CONAN), which embeds support for reliability at various levels of abstractions. Some of the main reliability sources are regular and decentralized structures based on simple basic computation cells designed to be robust against disturbances and noise, fault tolerance based on hardware, time and information redundancy applied at the basic cell level as well as at higher levels, self diagnosis assisted by the dynamic reconfiguration of basic computation cells and interconnect rerouting. Within the CONAN template, both technology dependent and independent models co-exists such that the more abstract layers are technology independent while the lower levels can be retargeted to various fabrication technologies. Our proposal is application-oriented and allows the designers to deal with unpredictability, and low reliability, which are unavoidable characteristics of future emerging nano-devices. When combined with the underlying software, the tools supporting the CONAN approach allow the designer to check whether the design constraints are fulfilled before performing a detailed implementation and provides means to trade area, delay, and power consumptions for reliability. As such, this proposal is a call-to-arms to mobilize the efforts of systems designers in order to achieve a systematic design methodology for reliable systems. Sorin Cotofana, Alexandre Schmid, Yusuf Leblebici, Adrian M. Ionescu, Oliver Soffke, Peter Zipf, Manfred Glesner, Antonio Rubio 0001 |
ASAP | 1 |
| 2005 | High Radix Addition Via Conditional Charge Transport in Single Electron Tunneling TechnologyabstractThis paper investigates the implementation of high radix addition based on the electron counting logic design style in single electron tunneling (SET) technology. A previous proposal for such an adder assumed the presence of a conditional charge movement (MCke) block which was only described as a black box. First, this paper proposes two possible MCke block implementations, each of which is described in detail and validated by means of simulation. Second, one of the proposed MCke implementations is utilized in the design of a 6-bit radix-8 adder. The resulting adder circuit is verified by simulation and evaluation indicated that it requires 187 circuit elements, has a delay of 4.15 ns (assuming an error probability P/sub error/ and a consumed energy of 224 meV). Cor Meenderinck, Sorin Cotofana, Casper Lageweg |
ASAP | 2 |
| 2005 | Compositional Memory Systems for Multimedia Communicating TasksabstractConventional cache models are not suited for real-time parallel processing because tasks may flush each other's data out of the cache in an unpredictable manner In this way the system is not compositional so the overall performance is difficult to predict and the integration of new tasks expensive. This paper proposes a new method that imposes compositionality to the system's performance and makes different memory hierarchy optimizations possible for multimedia communicating tasks when running on embedded multiprocessor architectures. The method is based on a cache allocation strategy that assigns sets of the unified cache exclusively to tasks and to the communication buffers. We also analytically formulate the problem and describe a method to compute the cache partitioning ratio for optimizing the throughput and the consumed power. When applied to a multiprocessor with memory hierarchy our technique delivers also performance gain. Compared to the shared cache case, for an application consisting of two jpeg decoders and one edge detection algorithm 5 times less misses are experienced and for an mpeg2 decoder 6.5 times less misses are experienced. Anca Mariana Molnos, Marc J. M. Heijligers, Sorin Cotofana, Jos T. J. van Eijndhoven |
DATE | 3 |
| 2005 | Addition Related Arithmetic Operations via Controlled Transport of ChargeabstractThis work investigates the single electron tunneling (SET) technology-based computation of basic addition related arithmetic functions, e.g., addition and multiplication, via a novel computation paradigm, which we refer to as electron counting arithmetic, that is based on controlling the transport of discrete quantities of electrons within the SET circuit. First, assuming that the number of controllable electrons within the system is unrestricted, we prove that the addition of two n-bit operands can be computed with a depth-2 network composed out of 3n+1 circuit elements and that the multiplication of two n-bit operands can be computed with a depth-3 network composed out of 4n-1 circuit elements. Second, assuming that the number of controllable electrons cannot be higher than a given constant r determined by practical limitations, we prove that the addition of two n-bit operands can be computed with a depth-(n/r+3) network composed out of 3n+1+n/r circuit elements. Under the same restriction, we suggest methods to reduce the addition network depth in the order of logn/r and to perform n-bit multiplication in an O(logn/r) delay. Finally, we propose SET-based implementations for a set of basic electron counting building blocks and implement a number of circuits operating under the electron counting paradigm as follows: 4-bit digital to analog converter, 5-bit analog to digital converter, 4-bit adder, and 3-bit multiplier. All proposed implementations are verified by means of simulation. Sorin Cotofana, Casper Lageweg, Stamatis Vassiliadis |
IEEE Trans. Computers | 1 |
| 2004 | Binary Multiplication based on Single Electron Tunneling
Casper Lageweg, Sorin Cotofana, Stamatis Vassiliadis |
ASAP | 2 |
| 2004 | Efficient Hardware for Antialiasing Coverage Mask GenerationabstractAn efficient low-cost, low-power hardware implementation of a novel run-time pixel coverage mask generation algorithm for embedded 3D graphics antialiasing purposes is presented. The proposed algorithm can be incorporated in any antialiasing scheme with prefiltering that is based on algebraic representation of primitive's edges. When compared with the state of the art, the described algorithm reduces several times the size of the required hardware implementation due to the utilization of the quadrant symmetry property allowing the storage of only the coverage mask information for a few representative edges in one of the quadrants of the plane, the rest of the information being derived on the fly via computationally inexpensive operations Dan Crisu, Sorin Cotofana, Stamatis Vassiliadis, Petri Liuha |
Computer Graphics International | 2 |
| 2004 | GRAAL - A Development Framework for Embedded Graphics AcceleratorsabstractThis paper presents a versatile hardware/software co-simulation and co-design environment for embedded 3D graphics accelerators. The graphics accelerator design exploration framework (GRAAL) is an open system which offers a coherent development methodology based on an extensive library of systemC RTL models of graphics pipeline components. GRAAL incorporates tools to assist in the visual debugging of the graphics algorithms implemented in hardware, and to estimate the performance in terms of throughput, power consumption, and area. Dan Crisu, Sorin Cotofana, Stamatis Vassiliadis, Petri Liuha |
DATE | 2 |
| 2004 | Compositional Memory Systems for Data Intensive ApplicationsabstractTo alleviate the system performance unpredictability of multitasking applications running on multiprocessor platforms with shared memory hierarchies, we propose a task level set based cache partitioning. We evaluate our approach on a CAKE platform with three Trimedias, one MIPS, and a shared level 2 cache using a picture in picture benchmark. We compare the performance implications of two types of cache partitioning namely set based. Our experiments indicate that associativity-based cache partitioning induces at least 30% performance degradation, whereas set-based partitioning provides 27% performance improvement when compared to non-partitioned cache scenario. Anca Mariana Molnos, Marc J. M. Heijligers, Sorin Cotofana, Jos T. J. van Eijndhoven |
DATE | 3 |
| 2004 | Analog-to-digital converter based on single-electron tunneling transistorsabstractA novel single-electron tunneling transistors (SETTs) based analog-to-digital converter (ADC) is proposed in this paper. The scheme we propose fully utilizes Coulomb oscillation effect, can properly operate at T>0 K, and only a capacitive divider (built with 2n-2 capacitors) and n pairs of complementary SETTs are required for an n-bit ADC implementation. When compared with other state-of-the-art SET based ADCs our method provides the most compact solution measured in terms of circuit elements and has a potential advantage in terms of conversion speed. To illustrate the operation of the proposed scheme, a 4-bit ADC is demonstrated at 10K by means of simulation. Chaohong Hu, Sorin Cotofana, Jianfei Jiang 0002, Qiyu Cai |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | Pel reconstruction on FPGA-augmented TriMediaabstractThis paper presents a TriMedia processor extended with three reconfigurable designs for entropy decoding (ED), inverse quantization (IQ), and two-dimensional (2-D) inverse discrete cosine transform (IDCT), and assesses the performance gain that is provided by such extensions when performing MPEG2-compliant pel reconstruction. We first describe an extension of the TriMedia architecture, which consists of a multiple-context field programmable gate array (FPGA)-based reconfigurable functional unit (RFU), a configuration unit managing the reconfiguration of the RFU, and their associated instructions. Then, we address the computation of the ED, IQ, and 2-D IDCT tasks, and propose to provide reconfigurable hardware support for a variable-length decoder that can decode two symbols per call (VLD-2), an inverse quantizer that can dequantize four coefficients per call (IQ-4), and an 1-D IDCT (1-D IDCT). The most important aspects concerning the implementation of the FPGA-mapped VLD-2, IQ-4, and 1-D IDCT units, as well as the organization of the software routines calling these FPGA-mapped computing units are outlined. Experimental results indicate that by configuring each of the VLD-2, IQ-4, and 1-D IDCT units on a different FPGA context, and by activating the contexts as needed, the FPGA-augmented TriMedia can perform MPEG2-compliant pel reconstruction with an average speed-up of 1.4/spl times/ over the standard TriMedia. Mihai Sima, Sorin Cotofana, Stamatis Vassiliadis, Jos T. J. van Eijndhoven, Kees A. Vissers |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | On Computing Addition Related Arithmetic Operations via Controlled Transport of ChargeabstractWe investigate the implementation of basic arithmetic functions, such as addition and multiplication, in single electron tunneling (SET) technology. First, we describe the SET equivalents of Boolean CMOS gates and threshold logic gates. Second, we propose a set of building blocks, which can be utilized for a novel design style, namely arithmetic operations performed by direct manipulation of the location of individual electrons within the system. Using this new set of building blocks, we propose several novel approaches for computing addition related arithmetic operations via the controlled transport of charge (individual electrons). In particular, we prove the following: n-bit addition can be implemented with a depth-2 network built with O(n) circuit elements; n-input parity can be computed with a depth-2 network constructed with O(n) circuit elements and the same applies for n/logn counters; multiple operand addition of m n-bit operands can be implemented with a depth-2 network using O(mn) circuit elements; and finally n-bit multiplication can be implemented with a depth-3 network built with O(n) circuit elements. Sorin Cotofana, Casper Lageweg, Stamatis Vassiliadis |
IEEE Symposium on Computer Arithmetic | 1 |
| 2003 | Color Space Conversion for MPEG decoding on FPGA-augmented TriMedia ProcessorabstractA case study on color space conversion (CSC) for MPEG decoding, carried out on the FPGA-augmented TriMedia processor is presented. That is, a transform from Y'CbCr color space to R'G'B' color space is addressed. First, we outline the extension of TriMedia architecture consisting of FPGA-based reconfigurable functional units (RFU) and associated instructions. Then we analyse a CSC (RFU-specific) instruction which can process four pixels per call, and propose a scheme to implement the CSC operation on RFU(s). When mapped on an ACEX EP1K100 FPGA, the proposed CSC exhibits a latency of 10 and a recovery of 2 TriMedia@200 MHz cycles, and occupies 57% of the device. By configuring the CSC facility on the RFU(s) at application load-time, color space conversion can be computed on FPGA-augmented TriMedia with 40% speed-up over the standard TriMedia. Mihai Sima, Stamatis Vassiliadis, Sorin Cotofana, Jos T. J. van Eijndhoven |
ASAP | 3 |
| 2003 | Evaluation Methodology for Single Electron Encoded Threshold Logic Gates
Casper Lageweg, Sorin Cotofana, Stamatis Vassiliadis |
VLSI-SOC | 2 |
| 2002 | MPEG-Compliant Entropy Decoding on FPGA-Augmented TriMedia/CPU64abstractThe paper presents a Design Space Exploration (DSE) experiment which has been carried out in order to determine the optimum FPGA-based Variable-Length Decoder (VLD) computing resource and its associated instructions, with respect to an entropy decoding task which is to be executed on the FPGA-augmented TriMedia/CPU64 processor We first outline the extension of the TriMedia/CPU64 architecture, which consists of an FPGA-based Reconfigurable Functional Unit (RFU) and the associated generic instructions. Then we address entropy decoding and propose a strategy to partially break the data dependency related to variable-length decoding. Three VLDs (VLD-1, VLD-2, VLD-3) instructions which can return 1, 2, or 3 symbols, respectively, are subsequently analyzed. After completing the DSE, we determined that VLD-2 instruction leads to the most efficient entropy decoding in terms of instruction cycles and FPGA area. The FPGA-based implementation of the computing resource associated to VLD-2 instruction is subsequently presented. When mapped on an ACEX EP1K100 FPGA from Altera, VLD-2 exhibits a latency of 8 TriMedia cycles, and uses all the Electronic Array Blocks and 51% of the logic cells of the device. The simulation results indicate that the VLD-2-based entropy decoder is 43% faster than its pure software counterpart. Mihai Sima, Sorin Cotofana, Stamatis Vassiliadis, Jos T. J. van Eijndhoven, Kees A. Vissers |
FCCM | 2 |
| 2002 | Field-Programmable Custom Computing Machines - A Taxonomy -
Mihai Sima, Stamatis Vassiliadis, Sorin Cotofana, Jos T. J. van Eijndhoven, Kees A. Vissers |
FPL | 3 |
| 2002 | Alternatives in FPGA-based SAD implementationsabstractIn multimedia processing, it is well-known that the sum-of-absolute-differences (SAD) operation is the most time-consuming operation when implemented in software running on programmable processor cores. This is mainly due to the sequential characteristic of such an implementation. In this paper, we investigate several hardware implementations of the SAD operation and map the most promising one in FPGA. Our investigation shows that an adder tree based approach yields the best results in terms of speed and area requirements and has been implemented as such by writing high-level VHDL code. The design was functionally verified by utilizing the MAX+plus II 10.1 Baseline software package from Altera Corp. and then synthesized by utilizing the LeonardoSpectrum software package from Exemplar Logic Inc. Preliminary results show that the design can be clocked at 380 MHz. This result translates into a faster than real-time full search in motion estimation for the main profile/main level of the MPEG-2 standard. Stephan Wong, Bastiaan Stougie, Sorin Cotofana |
FPT | 3 |
| 2001 | Topic 15+20: Multimedia and Embedded Systems
Stamatis Vassiliadis, Francky Catthoor, Mateo Valero, Sorin Cotofana |
Euro-Par | 4 |
| 2001 | An 8x8 IDCT Implementation on an FPGA-Augmented TriMedia
Mihai Sima, Sorin Cotofana, Jos T. J. van Eijndhoven, Stamatis Vassiliadis, Kees A. Vissers |
FCCM | 2 |
| 2001 | The MOLEN rho-mu-Coded Processor
Stamatis Vassiliadis, Stephan Wong, Sorin Cotofana |
FPL | 3 |
| 2001 | MPEG Macroblock Parsing and Pel Reconstruction On An FPGA-Augmented TriMedia ProcessorabstractThis paper describes an experiment which aims to reveal the potential impact on performance yielded by augmenting a TriMedia-CPU64 processor with a multiple-context FPGA core. We first propose an extension of the TriMedia CPU64 architecture, which consists of a reconfigurable functional unit and its associated instructions. Then, we address the decoding of variable-length codes on such extended TriMedia and describe the architecture and FPGA-implementation of a variable-length decoder (VLD) computing facility. When mapped on an ACEX EP1K100 FPGA, the proposed VLD exhibits a latency of 7 cycles. Preliminary results indicate that by configuring each of the VLD and 1-D IDCT (which is described elsewhere) facilities on a different FPGA context, and by activating the contexts as needed, the augmented TriMedia can perform macroblock parsing followed up by pel reconstruction with an improvement of 20 - 25% over the standard TriMedia. Mihai Sima, Sorin Cotofana, Stamatis Vassiliadis, Jos T. J. van Eijndhoven, Kees A. Vissers |
ICCD | 2 |
| 2000 | Hashed Addressed Caches for Embedded Pointer Based Codes (Research Note)
Marian Stanca, Stamatis Vassiliadis, Sorin Cotofana, Henk Corporaal |
Euro-Par | 3 |
| 2000 | Signed Digit Addition and Related Operations with Threshold LogicabstractAssuming signed digit number representations, we investigate the implementation of some addition related operations assuming linear threshold networks. We measure the depth and size of the networks in terms of linear threshold gates. We show first that a depth-2 network with O(n) size, weight, and fan-in complexities can perform signed digit symmetric functions. Consequently, assuming radix-2 signed digit representation, we show that the two operand addition can be performed by a threshold network of depth-2 having O(n) size complexity and O(1) weight and fan-in complexities. Furthermore, we show that, assuming radix-(2n-1) signed digit representations, the multioperand addition can be computed by a depth-2 network with O(n/sup 3/) size with the weight and fan-in complexities being polynomially bounded. Finally, we show that multiplication can be performed by a linear threshold network of depth-3 with the size of O(n/sup 3/) requiring O(n/sup 3/) weights and O(n/sup 2/ log n) fan-in. Sorin Cotofana, Stamatis Vassiliadis |
IEEE Trans. Computers | 1 |
| 1999 | Vector ISA Extension for Sparse Matrix-Vector Multiplication
Stamatis Vassiliadis, Sorin Cotofana, Pyrrhos Stathis |
Euro-Par | 2 |
| 1999 | Serial binary multiplication with feed-forward neural networks
Sorin Cotofana, Stamatis Vassiliadis |
Neurocomputing | 1 |
| 1998 | Periodic symmetric functions, serial addition, and multiplication with neural networksabstractThis paper investigates threshold based neural networks for periodic symmetric Boolean functions and some related operations. It is shown that any n-input variable periodic symmetric Boolean function can be implemented with a feedforward linear threshold-based neural network with size of O(log n) and depth also of O(log n), both measured in terms of neurons. The maximum weight and fan-in values are in the order of O(n). Under the same assumptions on weight and fan-in values, an asymptotic bound of O(log n) for both size and depth of the network is also derived for symmetric Boolean functions that can be decomposed into a constant number of periodic symmetric Boolean subfunctions. Based on this results neural networks for serial binary addition and multiplication of n-bit operands are also proposed. It is shown that the serial addition can be computed with polynomially bounded weights and a maximum fan-in in the order of O(log n) in O(n= log n) serial cycles, where a serial cycle comprises a neural gate and a latch. The implementation cost is in the order of O(log n), in terms of neural gates, and in the order of O(log2 n), in terms of latches. Finally, it is shown that the serial multiplication can be computed in O(n) serial cycles with O(log n) size neural gate network, and with O(n log n) latches. The maximum weight value in the network is in the order of O(n2) and the maximum fan-in is in the order of O(n log n). Sorin Cotofana, Stamatis Vassiliadis |
IEEE Trans. Neural Networks | 1 |
| 1996 | Serial Binary Addition with Polynominally Bounded Weights
Sorin Cotofana, Stamatis Vassiliadis |
ICANN | 1 |
| 1996 | 2-1 Additions and Related Arithmetic Operations with Threshold LogicabstractIn this paper we investigate the reduction of the size for small depth feed-forward linear threshold networks performing binary addition and related functions. For n bit operands we propose a depth-3 O(n2/log n) asymptotic size network for the binary addition with O polynomially bounded weights. We propose also a depth-3 addition of optimal O(n) asymptotic sits network and a depth-2 comparison of O(√n) asymptotic size network, both with O(2√n) asymptotic size of weight values. For existing architectural formats we show that our schemes, with equal or smaller depth networks, substantially outperform existing schemes in terms of size and fan-in requirements and on occasions in weight requirements. Stamatis Vassiliadis, Sorin Cotofana, Koen Bertels |
IEEE Trans. Computers | 2 |