Martin Margala

dblp:m/MartinMargala · DBLP profile ↗
← Back
75ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-0034-0369ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 67 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 MalHDC: A Lightweight Hyperdimensional Computing Approach to Malware Image Classification
abstract
Malware image classification has emerged as an effective approach for identifying malware families by transforming binary files into visual representations and learning discriminative image patterns. While deep learning methods have achieved strong performance on this task, they often rely on computationally intensive training and large model sizes. In this paper, we present MalHDC, a lightweight hyperdimensional computing framework for malware image classification that encodes each malware image as a compact hypervector representation that captures both intensity and spatial information. Class prototypes are learned directly from training samples through non-backpropagation prototype learning, and inference is performed using similarity-based matching. We evaluate the proposed framework on the Malimg dataset, which contains 25 malware families, and analyze the impact of hypervector dimensionality on classification performance. Experimental results show that MalHDC achieves 98.07% classification accuracy for identifying the family of a malware sample, using only 125,000 stored hypervector elements, while also offering significant efficiency advantages over prior top-performing methods. In particular, MalHDC is 2.23 × faster than the next-fastest baseline in inference time and 1.32 × smaller than the most compact competing model, while maintaining higher or comparable accuracy to existing approaches. These results demonstrate that MalHDC offers a highly compact and efficient alternative to conventional deep learning methods for malware image classification.
Alaaddin Goktug Ayar, Martin Margala
ACM Great Lakes Symposium on VLSI2
2026 Efficient Hyperdimensional Monitoring for Out-of-Distribution Detection
abstract
Deep neural networks deployed in real-world systems may produce unreliable predictions when encountering inputs outside the training distribution. Detecting such out-of-distribution (OOD) inputs during inference is therefore critical for reliable deployment. We propose a lightweight monitoring framework that observes intermediate neural network embeddings and detects abnormal inputs using hyperdimensional computing (HDC). The monitor converts embeddings into binary hypervectors via random projection and sign encoding, and evaluates distance to class prototypes using only XOR and population-count operations. Experiments using CIFAR-10 as the in-distribution dataset and SVHN and CIFAR-100 as OOD datasets show that the proposed method achieves strong detection performance with highly efficient binary computation.
Alaaddin Goktug Ayar, Martin Margala
ACM Great Lakes Symposium on VLSI2
2026 TLG2LUT6: Threshold Logic Mapping for Sub-Nanosecond Binary Neural Network Inference
abstract
Binary Neural Network (BNN) inference on FPGAs is bottlenecked by deep XNOR/popcount adder trees, while LUT-based alternatives are restricted to N ≤ 6 inputs per neuron. We propose TLG2LUT6, a compiler that encodes binary weights and thresholds directly into the 64-bit INIT strings of Xilinx LUT6 primitives, bounding a 256-input neuron to just 4 LUT stages. Out-of-context synthesis achieves Fmax = 1, 086 MHz, a 3.56 × speedup over XNOR+popcount, with 15.1 × fewer LUTs and zero DSP usage. The compact footprint fits the entire 64-neuron layer within a single SLR, eliminating inter-SLR routing penalties. Full-system deployment on an Alveo U200 confirms 300 MHz timing closure (limited by the fixed Vitis platform shell; the TLG core itself sustains 1,086 MHz), 99.2 MCPS measured throughput (XRT/DMA-limited in single-batch mode; kernel pipeline: 300 MCPS), and zero SLL crossings, with zero hardware/software mismatches confirmed across 64,000 MNIST classification bits. To the best of our knowledge, this is the first LUT-based approach supporting dense N = 256 binary neurons above 1 GHz. A pipelined three-layer cascade sustains 754 MHz with \(86.91\%\) MNIST top-1 and zero layer-boundary mismatches.
Abdullah Sahruri, Martin Margala
ACM Great Lakes Symposium on VLSI2
2026 Late Breaking Results - High-Throughput Metastability Characterization of Arbiters using HDC-based BIST
Abdullah Sahruri, Martin Margala
VTS2
2026 A novel weight-optimized LSTM for dynamic pricing solutions in e-commerce platforms based on customer buying behaviour
Martin Margala, S. Siva Shankar, Prasun Chakrabarti
Soft Comput.2
2026 Guest Editorial: Special Issue on Intelligence of Social Things-Enabled Cooperative Learning for Behavioral-Cultural Modeling
abstract
editorial reviewed
Chinmay Chakraborty, Bhuvan Unhelkar, Saïd Mahmoudi, Martin Margala, Sayonara Barbosa
IEEE Trans. Comput. Soc. Syst.4
2026 Kalman-Based Adaptive Moment Estimation Optimisation Algorithm to Enhance GPT in LLMs for Medical Sentiment Analysis of Patient Health-Related Feedback
abstract
The progress in Natural Language Processing (NLP) using Large Language Models (LLMs) has greatly improved medical sentiment analysis of patient feedback extraction from health-related question and answer. However, using LLMs to analyze such data often requires significant training data and computational resources, resulting in considerable increases in training costs and durations, which is one of the primary issues in applying LLMs to real-world healthcare scenarios. To tackle these challenges, a novel optimization algorithm named KAdam-EnGPT4LLM, based on Kalman filters and Adaptive Moment Estimation, is proposed to enhance training efficiency and reduce training costs of LLMs for analyzing patient feedback sentiment. Furthermore, the optimization algorithm KAdam-EnGPT4LLM is employed in training the LLM model GPT4ALL for medical sentiment analysis, resulting in the development of GPT4ALL-MediSentAly-KAdam, which leds to faster convergence and more stable training specifically for medical questions and answer in the context of healthcare. The results show that our GPT4ALL-MediSentAly-KAdam with the optimization algorithm KAdam-EnGPT4LLM achieved better performance that include the best Accuracy, Recall, F1-score, and Runtime for both datasets, outperforming traditional fine-tuned LLMs such as the classic GPT4ALL, Ada, Babbage, Curie, and Duvinci.
Xingchi Chen, Dazhou Li, Fa Zhu, Sidheswar Routray, Manisha Guduri, Martin Margala
IEEE J. Biomed. Health Informatics7
2025 Single-Pass Symbolic Learning for Real-Time Embedded Security
Alaaddin Goktug Ayar, Sercan Aygün, Martin Margala
ACM Great Lakes Symposium on VLSI3
2025 HyperEncoding: Spiking Neural Networks with Hyperdimensional Encoding for Robust Edge Intelligence
Alaaddin Goktug Ayar, Anthony S. Maida, Martin Margala
ACM Great Lakes Symposium on VLSI3
2025 TLGLock: A New Approach in Logic Locking Using Key-Driven Charge Recycling in Threshold Logic Gates
abstract
Logic locking remains one of the most promising defenses against hardware piracy, yet current approaches often face challenges in scalability and design overhead. In this paper, we present TLGLock, a new design paradigm that leverages the structural expressiveness of Threshold Logic Gates (TLGs) and the energy efficiency of charge recycling to enforce keydependent functionality at the gate level. By embedding the key into the gate’s weighted logic and utilizing dynamic charge sharing, TLGLock provides a stateless and compact alternative to conventional locking techniques. We implement a complete synthesis-to-locking flow and evaluate it using ISCAS, ITC, and MCNC benchmarks. Results show that TLGLock achieves up to 30% area, 50% delay, and 20% power savings compared to latchbased locking schemes. In comparison with XOR and SFLL-HD methods, TLGLock offers up to $3 \times$ higher SAT attack resistance with significantly lower overhead. Furthermore, randomized keyweight experiments demonstrate that TLGLock can reach up to 100% output corruption under incorrect keys, enabling tunable security at minimal cost. These results position TLGLock as a scalable and resilient solution for secure hardware design.
Abdullah Sahruri, Martin Margala
VLSI-SoC2
2025 ML4FPGA: An LLM Framework for Electronic Design Automation and Verification on FPGA
abstract
The application of Large Language Models (LLMs) to various machine learning tasks, including natural language processing (NLP), text, and code generation, has improved efficiency in these tasks. Despite their increasing prevalence, a comprehensive framework capable of generating and verifying Verilog code via a combination of LLMs and timing diagrams is lacking. This study introduces an LLM-based transformer framework that accelerates the process of reverse engineering electronic designs by analyzing the timing diagram via image recognition and generating and verifying Verilog design via pre-trained LLMs. We also compare the performance of our trained custom LLM model with much larger base models like Llama 4, GPT-4, and DeepSeek V3, Claude 3. The training of our model is done with datasets from HDLBits and publicly available Verilog code scrubbed from the web. Our custom model is trained/finetuned on Google A100 GPU accelerator and profiled. Our results indicate improvements compared to existing research in this domain and prospects of broader applications.
Uchechukwu Leo Udeji, Martin Margala
VLSI-SoC2
2024 Word2HyperVec: From Word Embeddings to Hypervectors for Hyperdimensional Computing
abstract
Word-aware sentiment analysis has posed a significant challenge over the past decade. Despite the considerable efforts of recent language models, achieving a lightweight representation suitable for deployment on resource-constrained edge devices remains a crucial concern. This study proposes a novel solution by merging two emerging paradigms, the Word2Vec language model and Hyperdimensional Computing, and introduces an innovative framework named Word2HyperVec. Our framework prioritizes model size and facilitates low-power processing during inference by incorporating embeddings into a binary space. Our solution demonstrates significant advantages, consuming only 2.2 W, up to 1.81 × more efficient than alternative learning models such as support vector machines, random forest, and multi-layer perceptron.
Alaaddin Goktug Ayar, Sercan Aygün, M. Hassan Najafi, Martin Margala
ACM Great Lakes Symposium on VLSI4
2020 Sparse Persistent GEMM Accelerator using OpenCL for Intel FPGAs
abstract
Optimizations of primitive routines such as General Matrix Multiplication (GEMM) continue to advance the state of the art performance for their applications. Various libraries such as Basic Linear Algebra Subprograms (BLAS) exist that provide API interfaces to highly tuned, hardware-specific implementations. Applications such as deep learning push the limits of what is possible from these subroutines by necessitating unique optimizations like lowering numeric precision and data type bit-width, exploiting resiliency, and removing redundancy. Hardware plays a considerable role in the performance of a subroutine because of the inherent layout of memory and arithmetic structures such as the SIMD processors found in general CPU and GPU architectures. FPGAs play a unique role in this space because the reconfigurable circuits and routing provide a pipelined architecture capable of both SIMD and MIMD-like architectures. Within a pipeline, the capability of accelerating operations on FPGA through low-bit and fine-grained designs is typically only seen in ASICs. In this paper, we provide an overview of an OpenCL based GEMM accelerator design that exploits sparsity and compression to persist data in FPGA on-chip, fine-grained SRAM. The design includes support for successive GEMMs with added activation functions, making it suitable for some machine learning applications. Results are measured running on Intel's Arria 10 GX 1150 FPGA. Compared to non-sparse and non-persistent designs, we achieve a speedup towards the theoretical limit of 6.5× for our fine-grained implementation and 8× for our structured implementation. We measure over 1 TOP/s utilizing only 17% of Arria 10 DSP blocks for a 90% sparse design.
Philip Colangelo, Shayan Sengupta, Martin Margala
ISCAS3
2018 Exploration of Low Numeric Precision Deep Learning Inference Using Intel® FPGAs
abstract
Convolutional neural networks (CNN) have been shown to maintain reasonable classification accuracy when quantized to lower precisions, however, quantizing to sub 8-bit activations and weights can result in classification accuracy falling below an acceptable threshold. Techniques exist for closing the accuracy gap of limited numeric precision networks typically by means of increasing computation. This results in a trade-off between throughput and accuracy and can be tailored for different networks through various combinations of activation and weight data widths. Customizable hardware architectures like FPGAs provide the opportunity for data width specific computation through unique logic configurations leading to highly optimized processing that is unattainable by full precision networks. Specifically, ternary and binary weighted networks offer an efficient method of inference for 2-bit and 1-bit data respectively. Most hardware architectures can take advantage of the memory storage and bandwidth savings that come along with a smaller datapath, but very few architectures can take full advantage of limited numeric precision at the computation level. In this paper, we present a hardware design for FPGAs that takes advantage of the bandwidth, memory, power, and computation savings of limited numerical precision data. We provide insights into the trade-offs between throughput and accuracy for various networks and how they map to our framework. Further, we show how limited numeric precision computation can be efficiently mapped onto FPGAs for both ternary and binary cases. Starting with Arria 10, we show a 2-bit activation and ternary weighted AlexNet running in hardware that achieves 3,700 images per second on the ImageNet dataset with a top-1 accuracy of 0.49. Using a hardware modeler designed for our low numeric precision framework we project performance most notably for a 55.5 TOPS Stratix 10 device running a modified ResNet-34 with only 3.7% accuracy degradation compared with single precision.
Philip Colangelo, Nasibeh Nasiri, Eriko Nurvitadhi, Asit K. Mishra, Martin Margala, Kevin Nealis
FCCM5
2018 Exploration of Low Numeric Precision Deep Learning Inference Using Intel® FPGAs: (Abstract Only)
abstract
Convolutional neural networks have been shown to maintain reasonable classification accuracy when quantized down to 8-bits, however, quantizing to sub 8-bit activations and weights can result in classification accuracy falling below an acceptable threshold. Techniques exist for increasing accuracy of sub 8-bit networks typically by means of increasing computation resulting in a trade-off between throughput and accuracy and can be tailored for different networks through combinations of activation and weight precisions. Customizable hardware architectures like FPGAs provide opportunity for data width specific computation through unique logic configurations leading to highly optimized processing that is unattainable by full precision networks. Specifically, ternary and binary weighted networks offer an efficient method of inference for 2-bit and 1-bit data respectively. In this paper, we present a hardware design for FPGAs that takes advantage of the bandwidth, memory, and computation savings of limited numerical precision data. We provide insights into the trade-offs between throughput and accuracy for various networks and how they map to our framework. Further, we show how limited numeric precision computation can be efficiently mapped onto FPGAs for both ternary and binary cases. Starting with Arria 10, we show a 2-bit activation and ternary weighted AlexNet running in hardware that achieves 3,700 images per second on the ImageNet dataset with a top-1 accuracy of 0.49. Using a hardware modeler designed for our low numeric precision framework we project performance most notably for a 55.5 TOPS Stratix 10 device running a modified ResNet-34 with only 3.7% accuracy degradation compared with single precision.
Philip Colangelo, Nasibeh Nasiri, Eriko Nurvitadhi, Asit K. Mishra, Martin Margala, Kevin Nealis
FPGA5
2018 Leakage-Aware Droop Measurement Built-in Self-Test Circuit for Digital Low-Dropout Regulators
Aydin Dirican, Cagatay Ozmen, Martin Margala
J. Electron. Test.3
2017 Fine-Grained Acceleration of Binary Neural Networks Using Intel® Xeon® Processor with Integrated FPGA
abstract
Summary form only given. Binary weighted networks (BWN) for image classification reduce computation for convolutional neural networks (CNN) from multiply-adds to accumulates with little to no accuracy loss. Hardware architectures such as FPGA can take full advantage of BWN computations because of the irability to express weights represented as 0 and 1 efficiently through customizable logic. In this paper, we present an implementation on Intel®'s Xeon®processor with integrated FPGA to accelerate binary weighted networks. We interface Intel's Accelerator Abstraction Layer (AAL) with Caffe to provide a robust framework used for accelerating CNN. Utilizing the low latency Quick Path Interconnect (QPI) between the Broadwell Xeon®processor and Arria10 FPGA, we can perform fine-grained offloads for specific portions of the network. Due to convolution layers making up most of the computation in our experiments, we offload the feature and weight data to customized binary hardware in the FPGA for faster execution. An initial proof of concept design shows that by using both the Xeon processor and FPGA together we can improve the throughput by 2× on some layers and by 1.3× overall while utilizing only a small percentage of FPGA core logic.
Philip Colangelo, Randy Huang, Enno Lübbers, Martin Margala, Kevin Nealis
FCCM4
2017 Design of a Low-Power Non-Volatile Programmable Inverter Cell for COGRE-based Circuits
abstract
This paper proposes a low-power non-volatile programmable inverter cell (NVPINV) that can be used with a COGRE (i.e. a compactly organized generic reconfigurable element) circuit to store the correct information for programming when establishing the desired logic function. The programmable data in the cell is read from a non-volatile SRAM (NVSRAM); two RMs (racetrack memories) are utilized as non-volatile elements. The RM is selected as non-volatile memory element due to its capability for independent operations (read and write), thus making possible a parallel execution of the programming process. The NVSRAM operates as a programmable circuit, i.e. a programmable circuit under control as either a buffer, or an inverter. The cell is extensively analyzed in terms of its operations with respect to different figures of merit, such as delay, power dissipation and power delay product (PDP). Simulation results show that in addition to low-power operation, the proposed NVPINV cell provides significant advantages (such as low delay and non-volatile storage) compared to an SRAM based Look-Up-Table (LUT) implementation.
Pilin Junsangsri, Fabrizio Lombardi, Salin Junsangsri, Martin Margala
ACM Great Lakes Symposium on VLSI4
2017 A high performance Full Adder based on Ballistic Deflection Transistor technology
abstract
In this paper, we propose a 1-bit Full Adder circuit built with Ballistic Deflection Transistors (BDT). BDT is a disruptive technology based on AlGaAs/InGaAs heterostructure. Different combinational circuits were successfully realized using BDT NAND gate and General Purpose Gate (GPG) structures. The developed circuit is an extension of BDT GPG and different from that of the previously implemented adder circuit. The proposed Adder consists of Sum and Carryout structures, comprising seven and five BDTs, respectively. Monte Carlo modeling of a BDT NAND Gate, which consists of associating two BDTs, has been performed and the obtained I-V characteristics were integrated with Verilog AMS to investigate the feasibility of the proposed circuit. Simulation results performed on Cadence Spectre simulator indicate the correct functionality of the proposed full adder.
Poorna Marthi, Nazir Hossain, Huan Wang 0009, Jean-François Millithaler, Martin Margala, Ignacio Iñiguez-de-la-Torre, Javier Mateos, Tomás González 0001
ISCAS5
2016 Modeling and Study of Two-BDT-Nanostructure based Sequential Logic Circuits
abstract
In this paper, study of different digital logic circuits developed using two-BDT ballistic nanostructure is presented. New D flip-flop (DFF) based on the same nanostructure is also proposed. The logic structure comprises two ballistic deflection transistors (BDTs) that are experimentally proven to operate at Terahertz frequencies. The non-linear behavior of the BDT's transfer characteristic has been perfectly reproduced by means of Monte Carlo simulations, where a specific attention has been devoted to surface charges. An analytical model built on the results of advanced MC simulations has been integrated into a behavioral Verilog AMS module to confirm the functionality of the circuit design. The module is used to analyze operating conditions of different combinational circuits and to investigate the feasibility of DFF design using BDT nanostructure. The simulation results indicate successful operation of both combinational and sequential circuits developed using two-BDT logic structure under proper biasing of gate and source terminals. The operating voltages of the proposed DFF are estimated to be + 225mV.
Poorna Marthi, Sheikh Rufsan Reza, Nazir Hossain, Jean-François Millithaler, Martin Margala, Ignacio Iñiguez-de-la-Torre, Javier Mateos, Tomás González 0001
ACM Great Lakes Symposium on VLSI5
2016 A new level sensitive D Latch using Ballistic nanodevices
abstract
In this paper, a D-Latch design using Ballistic Deflection Transistors (BDT) is presented. BDT technology was developed and experimentally proven to operate at THz frequencies. A simple, compact fit based analytical BDT model, developed previously to aid circuit design was utilized in this paper. The empirical device model is integrated into a behavioral Verilog A module to facilitate the investigation of the D-latch design. The D-latch design is based on the concept of BDT multiplexer structure and has been built using two instances of a single BDT modeled in Cadence AMS simulator. The simulation results confirm the correct operation of the D-latch.
Poorna Marthi, Nazir Hossain, Jean-François Millithaler, Martin Margala
ISCAS4
2016 A CMOS Ripple Detector for Voltage Regulator Testing
abstract
This paper presents an RMS based ripple sensor for testing of fully integrated voltage regulators. A DC signal which is proportional to the input ripple amplitude is generated. Final digital pass/fail signal is obtained with a clocked comparator. The sensor can detect a peak-to-peak ripple voltage of up to 50 millivolts on the 1.2 V supply rail and has 220 MHz bandwidth. The sensor is designed using IBM 90 nm CMOS technology and its functionality is verified in Cadence Virtuoso simulation environment.
Hieu Nguyen 0007, Cagatay Ozmen, Aydin Dirican, Nurettin Tan, Martin Margala
J. Electron. Test.5
2016 Low-Power Split-Radix FFT Processors Using Radix-2 Butterfly Units
abstract
Split-radix fast Fourier transform (SRFFT) is an ideal candidate for the implementation of a low-power FFT processor, because it has the lowest number of arithmetic operations among all the FFT algorithms. In the design of such processors, an efficient addressing scheme for FFT data as well as twiddle factors is required. The signal flow graph of SRFFT is the same as radix-2 FFT, and therefore, the conventional address generation schemes of FFT data could also be applied to SRFFT. However, SRFFT has irregular locations of twiddle factors and forbids the application of radix-2 address generation methods. This brief presents a shared-memory low-power SRFFT processor architecture. We show that SRFFT can be computed by using a modified radix-2 butterfly unit. The butterfly unit exploits the multiplier-gating technique to save dynamic power at the expense of using more hardware resources. In addition, two novel address generation algorithms for both the trivial and nontrivial twiddle factors are developed. Simulation results show that compared with the conventional radix-2 shared-memory implementations, the proposed design achieves over 20% lower power consumption when computing a 1024-point complex-valued transform.
Zhuo Qian, Martin Margala
IEEE Trans. Very Large Scale Integr. Syst.2
2015 High Level Programming of Document Classification Systems for Heterogeneous Environments using OpenCL (Abstract Only)
abstract
Document classification is at the heart of several of the applications that have been driving the proliferation of the internet in our daily lives. The ever growing amounts of data and the need for higher throughput, more energy efficient document classification solutions motivated us to investigate alternatives to the traditional homogenous CPU based implementations. We investigate a heterogeneous system where CPUs are combined with FPGAs as system accelerators. Incorporating FPGAs as accelerators in a heterogeneous computing environment allows for the creation of flexible custom hardware solutions that can potentially offer increased power efficiency and performance gains. One of the main issues delaying wide spread adoption of FPGAs as standard heterogeneous system accelerators is the difficulty in programming them. The OpenCL standard offers a unified C programming model for any device that adheres to its standards. An Altera OpenCL FPGA based implementation of a document classification system is investigated in which a stream of HTML documents is scored according to a profile on a document-by-document basis. The results show that the throughput of the document classification application with and without Bloom Filters is 312MB/s and 343MB/s respectively, when running on CPU, and 354MB/s and 452MB/s respectively, when running on an FPGA. Our results also show up to 32% power efficiency improvement for the FPGA implementation over the CPU implementation. We would like to thank Davor Capalija from Altera for his invaluable advice during our work on the FPGA version of the algorithm.
Nasibeh Nasiri, Oren Segal, Martin Margala, Wim Vanderbauwhede, Sai Rahul Chalamalasetti
FPGA3
2015 A Novel Coefficient Address Generation Algorithm for Split-Radix FFT (Abstract Only)
abstract
Split-Radix Fast Fourier Transform (SRFFT) has the lowest number of arithmetic operations among all the FFT algorithms. Since arithmetic operations dramatically contribute to the dynamic power consumption, SRFFT is an ideal candidate for the implementation of a low power FFT processor. In the design of such processors, an efficient addressing scheme for FFT data as well as coefficients is required. The signal flow graph of split-radix algorithm is the same as radix-2 FFT except for the location and value of coefficients, therefore conventional radix-2 FFT data address generation scheme could also be applied to SRFFT. However, the mixed radix property of SRFFT algorithm leads to irregular locations of coefficients and forbids any conventional address generation algorithm. This paper presents a novel coefficient address generation algorithm for shared-memory based SRFFT processor. The core part of the proposed algorithm is to use two control variables to track trivial and non-trivial multiplications. We found the relationship between the value of the control variables and the butterfly and pass counter. The corresponding hardware implementation is simple consisting of a shift register and a dual port RAM bank. Compared to look-up table approach, which pre-computes the addresses of all coefficients and stores the addresses in memory units, the proposed algorithm is scalable and only requires small amount of memory to find the correct addresses of coefficients.
Zhuo Qian, Martin Margala
FPGA2
2014 FPGA implementation of low-power split-radix FFT processors
abstract
Fast Fourier Transform (FFT) is one of the fundamental operations in digital signal processing area. Split-radix Fast Fourier Transform (SRFFT) approximates the minimum number of multiplications by theory among all the FFT algorithms, therefore SRFFT is a good candidate for the implementation of a low power FFT processor. In this PhD work, we aim to implement a novel low power Split-Radix FFT processor using shared-memory architecture and extend this work to a parallel structure based on FPGA. We started by designing a new radix-2 butterfly unit using clock gating approach to block unnecessary switching activity in the multiplier. Compared to existing SRFFT processors which are based on the “L” shaped butterfly, our implementation simplifies the address generation process for FFT data. Furthermore, because the number of multiplications required by SRFFT algorithm significantly decreases as the FFT size increases, it is reasonable to assume the proposed architecture will save more power when it comes to larger points of FFT.
Zhuo Qian, Nasibeh Nasiri, Oren Segal, Martin Margala
FPL4
2014 High level programming framework for FPGAs in the data center
abstract
Heterogeneous computing offers a promising solution for energy efficient computing in the data center. FPGA based heterogeneous computing is an especially promising direction since it allows for the creation of custom hardware solutions for data centric parallel applications. One of the main issues delaying wide spread adoption of FPGAs as main stream high performance computing devices is the difficulty in programming them. OpenCL was meant to address the difficulties and the non-uniformity related to programming heterogeneous devices, unfortunately because of its complexity it sets the bar high for many software programmers, preventing them from directly benefiting from the computing power and energy efficiency that OpenCL and heterogeneous computing have to offer. This work presents an effort to bridge the gap by extending an existing Java programming framework (APARAPI), based on OpenCL, so that it can be used to program FPGAs at a high level of abstraction and increased ease of programmability. We run several real world algorithms to assess the performance of the APARAPI framework on both a low end and a high end system. On the low end and high and systems respectively we find up to 78-80 percent power reduction and 4.8X-5.3X speed increase running NBody simulation, as well as up to 65-80 percent power reduction and 6.2X-7X speed increase for a K-Means MapReduce algorithm running on top of the Hadoop framework and APARAPI.
Oren Segal, Martin Margala, Sai Rahul Chalamalasetti, Mitch Wright
FPL2
2014 A novel low-power and in-place split-radix FFT processor
abstract
Split-radix Fast Fourier Transform (SRFFT) approximates the minimum number of multiplications by theory among all the FFT algorithms. Since multiplications significantly contribute to the overall system power consumption, SRFFT is a good candidate for implementation of a low power FFT processor. In this paper we present a novel low power SRFFT processor using a modified radix-2 butterfly structure. With the proposed butterfly unit, the address generation scheme for conventional radix-2 FFT could be applied to SRFFT and therefore it can avoid the complexity of address generation and interim data registers. Simulation results show that compared with a conventional radix-2 implementation, power consumption of the new processor is reduced by an amount of 11.7% and 18.3% for 16-point and 32-point FFT respectively.
Zhuo Qian, Martin Margala
ACM Great Lakes Symposium on VLSI2
2013 High throughput filtering using FPGA-acceleration
abstract
With the rise in the amount information of being streamed across networks, there is a growing demand to vet the quality, type and content itself for various purposes such as spam, security and search. In this paper, we develop an energy-efficient high performance information filtering system that is capable of classifying a stream of incoming document at high speed. The prototype parses a stream of documents using a multicore CPU and then performs classification using Field-Programmable Gate Arrays (FPGAs). On a large TREC data collection, we implemented a Naive Bayes classifier on our prototype and compared it to an optimized CPU based-baseline. Our empirical findings show that we can classify documents at 10Gb/s which is up to 94 times faster than the CPU baseline (and up to 5 times faster than previous FPGA based implementations). In future work, we aim to increase the throughput by another order of magnitude by implementing both the parser and filter on the FPGA.
Wim Vanderbauwhede, Anton Frolov 0002, Leif Azzopardi, Sai Rahul Chalamalasetti, Martin Margala
CIKM5
2013 An FPGA memcached appliance
abstract
Providing low-latency access to large amounts of data is one of the foremost requirements for many web services. To address these needs, systems such as Memcached have been created which provide a distributed, all in-memory key-value store. These systems are critical and often deployed across hundreds or thousands of servers. However, these systems are not well matched for commodity servers, as they require significant CPU resources to achieve reasonable network bandwidth, yet the core Memcached functions do not benefit from the high performance of standard server CPUs. In this paper, we demonstrate the design of an FPGA-based Memcached appliance. We take Memcached, a complex software system, and implement its core functionality on an FPGA. By leveraging the FPGA's design and utilizing its customizable logic to create a specialized appliance we are able to tightly integrate networking, compute, and memory. This integration allows us to overcome many of the bottlenecks found in standard servers. Our design provides performance on-par with baseline servers, but consumes only 9% of the power of the baseline. Scaled out, we see benefits at the data center level, substantially improving the performance-per-dollar while improving energy efficiency by 3.2X to 10.9X.
Sai Rahul Chalamalasetti, Kevin T. Lim, Mitch Wright, Alvin AuYoung, Parthasarathy Ranganathan, Martin Margala
FPGA6
2013 An IDDQ BIST approach to characterize phase-locked loop parameters
abstract
In this work, a new IDDQ built-in self-test (BIST) solution is proposed to provide accurate on-chip current measurements for phase-locked loops (PLLs) found in deep-submicron system-on-chip (SoC) products. The proposed method characterizes PLL loop parameters to increase test quality with minimum additional test time and 4.5% accuracy in IBM 65 nm technology. A self-correction mechanism accompanies the proposed BIST to recover performance variations resulting from excessive process variation found in high-volume manufacturing (HVM). The proposed IDDQ BIST circuit's performance is evaluated in silicon using 0.18μm technology and achieves 2% accuracy with only 1.7% additional PLL area overhead. Extensions to other analog mixed signal circuit blocks should be possible.
Samed Maltabas, Osman Kubilay Ekekon, Kemal Kulovic, Anne Meixner, Martin Margala
VTS5
2013 Throughput/Resource-Efficient Reconfigurable Processor for Multimedia Applications
abstract
This brief presents the implementation and evaluation of an 8-bit adaptable processor core to be part of the power-throughput-area efficient multimedia oriented reconfigurable architecture reconfigurable array. The design of the processor core was custom implemented in IBM's 90 nm CMOS technology and occupies 0.115 mm2silicon area with approximately 70% area utilized by core circuits. The processor shows a peak throughput performance of 75 MOPS/mW. Benchmarking results show estimated throughputs of 9.5, 21.36, 39.78, 170.88, and 4.54 MSamples/s for variants of 2-D discrete cosine transform (DCT), 4 × 4 H.264 integer transform, and 2-D discrete wavelet transform, respectively. Our analysis shows that the proposed design provides approximately 4-8 times higher throughput for 2-D DCT when compared against popular architectures.
Sohan Purohit, Sai Rahul Chalamalasetti, Martin Margala, Wim Vanderbauwhede
IEEE Trans. Very Large Scale Integr. Syst.3
2013 Design and Evaluation of High-Performance Processing Elements for Reconfigurable Systems
abstract
In this paper, we present the design and evaluation of two new processing elements for reconfigurable computing. We also present a circuit-level implementation of the data paths in static and dynamic design styles to explore the various performance-power tradeoffs involved. When implemented in IBM 90-nm CMOS process, the 8-b data paths achieve operating frequencies ranging over 1 GHz both for static and dynamic implementations, with each data path supporting single-cycle computational capability. A novel single-precision floating point processing element (FPPE) using a 24-b variant of the proposed data paths is also presented. The full dynamic implementation of the FPPE shows that it operates at a frequency of 1 GHz with 6.5-mW average power consumption. Comparison with competing architectures shows that the FPPE provides two orders of magnitude higher throughput. Furthermore, to evaluate its feasibility as a soft-processing solution, we also map the floating point unit onto the Virtex 4 and 5 devices, and observe that the unit requires less than 1% of the total logic slices, while utilizing only around 4% of the DSP blocks available. When compared against popular field-programmable-gate-array-based floating point units, our design on Virtex 5 showed significantly lower resource utilization, while achieving comparable peak operating frequency.
Sohan Purohit, Sai Rahul Chalamalasetti, Martin Margala, Wim Vanderbauwhede
IEEE Trans. Very Large Scale Integr. Syst.3
2012 Evaluating FPGA-acceleration for real-time unstructured search
abstract
Emerging data-centric workloads that operate on and harvest useful insights from large amounts of unstructured data require corresponding new data-centric system architecture optimizations. In particular, with the growing importance of power and cooling costs, a key challenge for such future designs is to achieve increased performance at high energy efficiency. At the same time, recent trends towards better support for reconfigurable logic enable the use of energy-efficient accelerators. Combining these trends, in this paper, we examine the applicability of acceleration in future data-centric system architectures. We focus on an important class of data-centric workloads, real-time unstructured search, or information filtering, where large collections of documents are scored against specific topic profiles, and present an FPGA-based implementation to accelerate such workloads. Our implementation, based on the GiDEL PROCStar IV board using Altera Stratix IV FPGAs, demonstrates excellent performance and energy efficiency, 20 to 40 times better than baseline server systems for typical usage scenarios. Our results also highlight interesting insights for the design of accelerators in future data-centric systems.
Sai Rahul Chalamalasetti, Martin Margala, Wim Vanderbauwhede, Mitch Wright, Parthasarathy Ranganathan
ISPASS2
2012 Time-Based Embedded Test Instrument with Concurrent Voltage Measurement Capability
Kemal Kulovic, Martin Margala
J. Electron. Test.2
2012 Novel Practical Built-in Current Sensors
Samed Maltabas, Kemal Kulovic, Martin Margala
J. Electron. Test.3
2012 Investigating the Impact of Logic and Circuit Implementation on Full Adder Performance
abstract
This paper presents the design and characterization of 12 full-adder circuits in the IBM 90-nm process. These include three new full-adder circuits using the recently proposed split-path data driven dynamic logic. Based on the logic function realized, the adders were characterized for performance and power consumption when operated under various supply voltages and fan-out loads. The adders were then further deployed in a 32 bit ripple carry adder and 8×4 multiplier to evaluate the impact of sum and carry propagation delays on the performance, power of these systems. Performance characterization of the adder circuits in the presence of process and voltage variations was also performed through Monte Carlo simulations. Besides analyzing and comparing circuit performance, the possible impact of the choice of logic function has also been underlined in this study.
Sohan Purohit, Martin Margala
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Resource-efficient implementation of Blue Midnight Wish-256 hash function on Xilinx FPGA platform
abstract
This paper presents the design and analysis of an area efficient Blue Midnight Wish compression function with digest size of 256 bits (BMW-256) on FPGA platforms. The proposed architecture achieves significant improvements in system throughput with reduced area. We demonstrate the performance of the proposed BMW hash function core using VIRTEX 5 FPGA implementation. The new BMW hash function design allows for 16X speed up in performance while consuming significantly lower area than previously reported (i.e. just 445 slices).
Mohamed El-Hadedy 0001, Martin Margala, Danilo Gligoroski, Svein J. Knapskog
IAS2
2010 A C++-embedded Domain-Specific Language for programming the MORA soft processor array
abstract
MORA is a novel platform for high-level FPGA programming of streaming vector and matrix operations, aimed at multimedia applications. It consists of soft array of pipelined low-complexity SIMD processors-in-memory (PIM). We present a Domain-Specific Language (DSL) for high-level programming of the MORA soft processor array. The DSL is embedded in C++, providing designers with a familiar language framework and the ability to compile designs using a standard compiler for functional testing before generating the FPGA bitstream using the MORA toolchain. The paper discusses the MORA-C++ DSL and the compilation route into the assembly for the MORA machine and provides examples to illustrate the programming model and performance.
Wim Vanderbauwhede, Martin Margala, Sai Rahul Chalamalasetti, Sohan Purohit
ASAP2
2010 A new built-in IDDQ testing method using programmable BICS
abstract
The paper presents a novel programmable Built In Current Sensor (BICS) topology in IBM 65 nm CMOS technology. Proposed topology has 2.086 GHz bandwidth and 38.9ps detection time. Moreover, a new built-in IDDQ test flow is proposed. Proposed test flow is applied to a charge pump. The results show 100% fault coverage for the defects that affects the output of the charge pump (CP). 97.87% overall fault coverage is achieved for the same test.
Samed Maltabas, Osman Kubilay Ekekon, Martin Margala
ETS3
2010 Topology impact on the room temperature performance of THz-range ballistic deflection transistors
abstract
In this paper, the Ballistic Deflection Transistor (BDT) is reviewed for variations in performance of the device with respect to geometry modifications like change in gate length, side branch angle and permittivity. Experimental study is done for different gate lengths and same study is then done using Silvaco modeling tool to corroborate the behavior of model used. It is confirmed that the performance of BDT can be enhanced by using a complete gate configuration where gate runs along the channel. Silvaco simulations are also used to study the effect of different side channel angles. It is found that for a specific angle of 60 degrees, the output current is maximum. Any angle other than this gives smaller current values. Simulations are also used to study the effect of variation in the dielectric constant filled in the trenches between channel and gate. It is observed that high permittivity dielectrics enhance the output.
Vikas Kaushal, Ignacio Iñiguez-de-la-Torre, Martin Margala
ACM Great Lakes Symposium on VLSI3
2010 Design of self correcting radiation hardened digital circuits using decoupled ground bus
abstract
With shrinking device sizes, modern digital circuits are becoming increasingly susceptible to transient errors due to charged particle strikes on the sensitive nodes of the circuit. In this paper we present a simple technique to implement self correcting, radiation hardened digital circuits through the use of a decoupled ground bus. The technique relies on using the error signal as a design variable in the logic being realized and employs a state machine based design approach for combinational logic design. Simulation results show that the proposed technique is reliable over all corners and robust against random process and mismatch variations. Our technique results in average 29.43% delay, 13.8% power and negligible area overhead than non hardened static and domino designs.
Sohan Purohit, Sai Rahul Chalamalasetti, Martin Margala
ACM Great Lakes Symposium on VLSI3
2010 Novel programmable built-in current-sensor for analog, digital and mixed-signal circuits
abstract
The paper presents a novel programmable BICS topology in IBM 65 nm CMOS technology. Programmability, which is the main motivation of this approach, is shown on three different CUTs, simultaneously. It is shown that proposed topology has 473 MHz bandwidth, 200.6ps detection time with 1.5μA sensitivity. Moreover, these specifications are achieved with only one control pin. The advantage on area overhead becomes more significant, as the number of CUTs increases in the system. These results show that proposed programmable BICS is superior when it is compared with recent works.
Osman Kubilay Ekekon, Samed Maltabas, Martin Margala
ISCAS3
2010 An area efficient design methodology for SEU tolerant digital circuits
abstract
This paper presents a new circuit design style for radiation hardened digital circuits. The proposed design methodology is based on the well known DCVSL circuit style. The original DCVSL has been modified to build in robustness against single event upsets due to particle strikes. The paper presents simulation results for logic gates and arithmetic circuits built using the proposed design scheme. The circuits were found to have area savings of about 26% over their static CMOS counterparts at the cost of 7% extra delay. The design style described in this paper was found to be extremely reliable with 100% SEU mitigation for a wide spectrum of charge and current profiles at relatively lower cost than previously published circuit schemes for SEU mitigation.
Sohan Purohit, David Harrington, Martin Margala
ISCAS3
2009 A low cost reconfigurable soft processor for multimedia applications: Design synthesis and programming model
abstract
This paper presents an FPGA implementation of a low cost 8 bit reconfigurable processor core for media processing applications. The core is optimized to provide all basic arithmetic and logic functions required by the media processing and other domains, as well as to make it easily integrable into a 2D array. This paper presents an investigation of the feasibility of the core as a potential soft processing architecture for FPGA platforms. The core was synthesized on the entire Virtex FPGA family to evaluate its overall performance, scalability and portability. A special feature of the proposed architecture is its simple programming model which allows low level programming. Throughput results for popular benchmarks coded using the programming model and cycle accurate simulator are presented.
Sai Rahul Chalamalasetti, Wim Vanderbauwhede, Sohan Purohit, Martin Margala
FPL4
2009 Study of leakage current mechanisms in ballistic deflection transistors
abstract
In this paper, the Ballistic Deflection Transistor (BDT) is reviewed for variations in performance of the device including leakage with respect to geometry modifications. Monte Carlo and Silvaco modeling tools are used to study current leakage mechanism in BDT. Low power selection criteria and theory behind position of deflector in the device are examined. Since ballistic conduction is not dissipative, power loss should be low. Leakage can be reduced by placing deflector at about 25% of its own length lower than the exact centre of the device. Current leakages that occurred during device operation are compared with each other and with the output current. It is observed that magnitude of leakage current is distinct at different ports of the device. For a specific set of parameters, leakage is comparable to the output which essentially motivates to choose optimum device architecture.
Vikas Kaushal, Quentin Diduck, Martin Margala
ACM Great Lakes Symposium on VLSI3
2009 Varicap threshold logic
abstract
In this paper, a highly compact novel Threshold Logic (TL) Gate approach called "Varicap TL (VcTL)" is proposed and described. The novel feature of the design is in using variable MOSFET capacitances which reduces the area. The electrical analysis of this variable MOSFET capacitance is presented and its variability is explained. Varicap TL (VcTL) gate is created by using a latch type decision circuit topology. Parallel counter implementations of (7,3) in 0.13µm and 0.18µm technology are realized by using proposed Varicap TL based on Minnick TL Network. Comparison of these implementations with Boolean Logic (BL) based dynamic (7,3) counter is shown. VcTL approach in 0.13µm offers 41% smaller area which is a significant result of the approach and 27% higher speed and only 24% higher power consumption compared to BL realization in 0.13µm technology. The results also show VcTL's scalability. As VcTL is scaled from 0.18µm to 0.13µm, the speed is increased by 21%, the area is decreased by 33% and power by 37%.
Samed Maltabas, Martin Margala, Ugur Çilingiroglu
ACM Great Lakes Symposium on VLSI2
2009 A 1.2v, 1.02 ghz 8 bit SIMD compatible highly parallel arithmetic data path for multi-precision arithmetic
abstract
Coarse grained arithmetic and logic units have long been the primary computational units for media processing. This paper presents the organization and VLSI implementation of a new 8bit, Single Instruction Multiple Data (SIMD) compatible ALU for fast, area and power efficient arithmetic and logic operations. An array of 8 such units, along with the interconnect network to perform 16 bit multiplication is shown. The array was custom implemented in IBM 0.13 CMOS process. Post layout simulation results show cell operation at 1.02GHz with a power consumption of 1.34mW. The proposed cell consumes 22-52% less power than competing architectures, while providing GHz range operating speeds. The array is found to provide almost 6 times the performance of dedicated 16 bit multiplication units, while still providing 60% power improvement. A generalized mapping scheme for implementing higher precision arithmetic operations using the proposed ALU as the basic building block is shown.
Sohan Purohit, Sai Rahul Chalamalasetti, Martin Margala
ACM Great Lakes Symposium on VLSI3
2009 New performance/power/area efficient, reliable full adder design
abstract
Arithmetic circuits have always played one of the most important roles in the designs of processors, FPGAs, and the rapidly evolving domain of media processing architectures. The full adder cell forms the basic building block of majority of these arithmetic circuits. In this paper we describe a hybrid pseudo static full adder cell designed using Data Driven Dynamic Logic. Simulation results show the adder to out perform its competitors, both static as well as dynamic topologies in terms of performance, while maintaining relatively similar area and power characteristics. This paper presents a complete characterization of the popular adder cells in terms of delay, area, power, noise margin and reliability analysis for both super threshold and sub threshold operating regimes.
Sohan Purohit, Martin Margala, Marco Lanuzza, Pasquale Corsonello
ACM Great Lakes Symposium on VLSI2
2009 Ballistic Deflection Transistors and the Emerging Nanoscale Era
abstract
This paper presents a brief survey of the state of the art in nanoscale electronics, with special emphasis on room-temperature nanoscale ballistic deflection transistors (BDTs) and T-branch junctions (TBJs). Both devices are planar structures etched into a two-dimensional electron gas (2DEG). Extremely low capacitances (~0.2 fF) in the 2DEG system and low switching voltages (~0.15 V) predict THz performance and ultra-low power consumption, making BDTs and TBJs among the most promising and versatile of ballistic nanoelectronic devices. Obstacles in circuit and logic design using the BDT are presented along with potential solutions. I-V characteristics from a fabricated BDT and simulation results from a two-input BDT NAND gate are provided. Future plans to facilitate large-scale integration are discussed.
David Wolpert 0001, Hiroshi Irie, Roman Sobolewski, Paul Ampadu, Quentin Diduck, Martin Margala
ISCAS6
2008 A portless SRAM Cell using stunted wordline drivers
abstract
A minimum area portless SRAM cell is presented along with a stunted wordline driver. When compared to an iso-area 6T cell using logic design rules, the portless cell exhibits 22% higher static noise margin and 14% lower leakage with a 51% penalty in the on/off cell current ratio. Measurements from a fabricated test chip in 0.5 mum CMOS demonstrate functionality of the proposed cell and driver in an array, and for the first time, validate the portless concept in an isolated test cell.
Michael Wieckowski, Martin Margala
ISCAS2
2007 Novel Process and Temperature-Stable BICS for Embedded Analog and Mixed-Signal Test
abstract
This paper proposes a new current sensor design for the arduous task of testing embedded analog and mixed-signal circuits. This proposed, wide-band, minimally-intrusive IDD sensor operates up to 230MHz, which is 2.3X faster than previously proposed designs, and occupies 78.3% less area than another competing design, while achieving an inherent tolerance to process and temperature variations without sacrificing significant area. A BiST utilizing this novel IDD sensor is created and tested on a CMOS op-amp and mixer showing high fault detection sensitivity, while maintaining the performance of the DUT (device-under-test). The experiments are implemented in 0.18mum TSMC CMOS mixed- signal technology.
John C. Liobe, Martin Margala
IOLTS2
2007 Process and Temperature Calibration of PLLs with BiST Capabilities
abstract
This paper presents two self-calibrating charge pump phase locked loop (CP-PLL) architectures, the first utilizing a ring VCO (Voltage Controlled Oscillator) and the second implemented using a LC-tank VCO. Each design utilizes frequency-modulated analog-to-digital converter(s) (FM ADCs) as part of a calibration circuit to compensate for process and temperature (PT) variations. For the ring VCO-based PLL, compensation is achieved over all four process corners and for temperatures of 0°C, 27°C, and 60°C by dynamically modifying the charge-pump current. In the LC-tank VCO-based PLL design, the tuning range of the VCO is tuned according to the detected process shift. Each FM ADC consumes only 729μW of power for the worst-case scenario and 0.0016mm2of area. For the two 0.18μm CMOS PLL case studies investigated here, calibration is achieved for non-ideal operating environments.
Richard Geisler, John C. Liobe, Martin Margala
ISCAS3
2007 A Self-Biased Charge-Transfer Sense Amplifier
abstract
A Self-Biased Charge-Transfer Sense Amplifier (SB-CTSA) is proposed for applications in high performance static memory. The new design incorporates an internal biasing mechanism along with a static output latch that can store the result from a read cycle indefinitely at no additional cost in power. It exhibits 23% faster sensing delay and a 37% reduction in read energy when compared to a recently proposed CTSA structure. In addition, the SB-CTSA design exhibits low sensitivity to both input capacitance and input capacitance mismatch.
Sandeep Patil, Michael Wieckowski, Martin Margala
ISCAS3
2007 A Processor-In-Memory Architecture for Multimedia Compression
abstract
This paper presents the design and development of a novel, low-complexity processor-in-memory (PIM) architecture for image and video compression. By integrating a novel-processing element with SRAM, bandwidth is improved and latency is greatly reduced. This paper also presents PIM design techniques for reduced power, area, and complexity for rapid deployment and reduced cost. A design methodology is presented and followed by an analysis of the processing element performance and capabilities. The proposed datapath solution delivers between 2 to 40 times higher performance compared to other presented solutions. The architecture executes a discrete cosine and wavelet transforms achieving up to 40% higher throughput per watt and occupying as little as 0.9% area compared to a commercial digital signal processing and other application-specified integrated circuit implementations while maintaining precision. A comprehensive comparative analysis is also provided. The proposed processor-in-memory is implemented in 1.8-V 0.18-mum CMOS technology and operates with a 300-MHz clock
Brandon J. Jasionowski, Michelle K. Lay, Martin Margala
IEEE Trans. Very Large Scale Integr. Syst.3
2006 A 2.4-GHz auto-calibration frequency synthesizer with on-chip built-in-self-test solution
abstract
A novel single chip 2.4-GHz phase-lock loop (PLL) with auto-calibration mechanism, which is adaptive to process variation, is presented. The proposed synthesizer with self-calibration block is connected to the on-chip jitter measurement circuit. The PLL adjusts the voltage controlled oscillator (VCO) input control voltage in response to the measured jitter by the jitter measurement block. The response of the VCO is adjusted inherently towards the desired center frequency by the self-calibration process to reduce the jitter of its 2.4-GHz output for ZigBee application. This method could also be used to implement built-in-self-test for PLL. The synthesizer, designed in TSMC 0.18-mum technology, has an edge jitter standard deviation of 0.86-ps having a 20% improvement over PLL with no self-calibration. The digitally controlled VCO has a wide tuning range of 0.6-GHz. The synthesizer has a programmable frequency divider that operates with a division range of 456-496
Sadeka Ali, Martin Margala
ISCAS2
2006 An integrated countermeasure against differential power analysis for secure smart-cards
abstract
This paper presents a new hardware technique for the realization of secure smart-cards. The proposed strategy represents a valid countermeasure against non-invasive attacks, such as power analysis. It is based on a simple sub-circuit (Kocher et al., 1999) that can be easily integrated into the smart-card chip. It has been proven that the new technique decorrelates the power consumed by any digital circuit from the internally elaborated data, thus avoiding extraction of secret information from smart cards during the execution of their internal computations
Pasquale Corsonello, Stefania Perri, Martin Margala
ISCAS3
2006 Process tolerant calibration circuit for PLL applications with BIST
abstract
A process-invariant calibration circuit, capable of correcting performance errors in charge-pump based PLLs is described. Process variations detrimentally affect all building blocks of standard PLL architectures. Utilizing a novel ADC, these variations are sensed and corrected. The self-calibration circuitry is non-intrusive and requires minimal area and power overhead. A case study of a 2.4GHz ring VCO-based PLL designed in a TSMC's 0.18mum CMOS mixed-signal technology is given. The calibration circuitry is able to sense and calibrate under all four process corners as well as detect under high temperature conditions
Quentin Diduck, John C. Liobe, Sadeka Ali, Martin Margala
ISCAS4
2006 A versatile computation module for adaptable multimedia processors
abstract
This paper describes a low cost, low power, versatile computation module that can be used as a coarse-grain building block in multimedia processors. The module, which has a datapath and a controller integrated with its local data memory, performs various arithmetic operations on different data types, i.e., 8-bit integer, 16-bit integer, 32-bit integer and single precision floating point numbers. Running in parallel, the module provides high data throughput at low hardware cost. Multiple modules can be connected in a multimedia processor operated in mixed SIMD and MIMD modes, providing great flexibility for data parallel, computation intensive multimedia applications. Compared to recently proposed datapaths, the proposed module has more computational abilities, more flexibility and it is between 2/spl times/ and 10/spl times/ more energy efficient.
Yunan Xiang, R. Pettibon, Martin Margala
ISCAS3
2006 Design of a wireless test control network with radio-on-chip technology for nanometer system-on-a-chip
abstract
The continued push to smaller geometries, higher frequencies, and larger chip sizes rapidly resulted in an incompatibility between interconnect needs and projected interconnect performance. As stated in the 2003 International Technology Roadmap for Semiconductors (ITRS'03) report, revolutionary interconnect methodologies such as radio frequency (RF)/wireless will deliver the foreseen progress in semiconductor technology. Recent advances in silicon integrated circuit technique are making possible tiny low-cost transceivers to be integrated on chip, namely "radio-on-chip" (ROC) technology. This paper proposes the idea of using wireless radios to transmit test data and control signals to resolve the acerbated core accessibility problem. Three types of wireless test micronetworks are first presented, i.e., miniature wireless local area network (LAN), multihop wireless test control network (MTCNet), and distributed multihop MTCNet. Then, the test control overhead and system resource partitioning in on-chip wireless micronetworks are analyzed. Several challenging system design problems such as RF node placement, core clustering, and control routing are studied, and the test control resources (i.e., the on-chip RF nodes for intrachip communication) are properly distributed and system optimization is performed in terms of test control cost. A simulation study shows the feasibility and applicability of intrachip MTCNet.
Danella Zhao, Shambhu J. Upadhyaya, Martin Margala
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2005 A new SoC test architecture with RF/wireless connectivity
abstract
When moving into the billion-transistor era, the direct or bus interconnects in conventional SoC test control models are rather restricted in not only system performance, but also signal integrity and transmission with continued scaling of the feature size. Recent advances in silicon integrated circuit technology are making possible tiny low-cost transceivers to be integrated on chip. In this paper, we propose a new distributed multihop wireless test control network based on the recent development in "radio-on-chip" technology. Under the multilevel tree structure, the system optimization is performed on control constrained resource partitioning and distribution. Several challenging system design issues, such as RF nodes placement, clustering, and routing are studied, with the integrated resource distribution and system optimization on TAM design and test scheduling. Experimental results show that the proposed algorithm can efficiently minimize the overall testing cost.
Danella Zhao, Shambhu J. Upadhyaya, Martin Margala
ETS3
2005 Low-Cost Fully Reconfigurable Data-Path for FPGA-Based Multimedia Processor
abstract
This paper describes novel data-path architecture for FPGA-based multimedia processors. The proposed circuit can adapt itself at run-time to different operations and data wordlengths avoiding time and power consuming reconfiguration. The new data-path can operate in SIMD fashion and guarantees high parallelism levels when operations on lower precisions are executed. It also supports IEEE-754 compliant single precision floating-point addition and multiplication. The proposed circuit has been characterized using VIRTEXII XILINX devices, but it can be efficiently used also in other FPGA families.
Marco Lanuzza, Stefania Perri, Martin Margala, Pasquale Corsonello
FPL3
2005 Cost-effective low-power processor-in-memory-based reconfigurable datapath for multimedia applications
abstract
Multimedia applications have become a dominant computing workload for computer systems as well as for wireless-based devices. Due to their repetitive computing and memory intensive nature, they can take effective advantage from Processor-In-Memory (PIM) technology. In this paper, a new low-power PIM-based 32-bit reconfigurable datapath optimized for multimedia applications is presented. The new circuit efficiently performs parallel arithmetic operations on either 8-, 16-, or 32-bit integer data or on 32-bit single precision floating-point data. As a result, high flexibility is provided at a very low hardware cost. When implemented using the UMC 0.18 μm 1.8 V CMOS technology, the proposed datapath exhibits a 285 MHz running frequency, dissipates just 0.12 mW/MHz and occupies a silicon area of only 107,323 μm2. When performing 2D-DCT, proposed architecture consumes 74% less power and is 28% more power efficient compared to top-of-the-line commercial TI DSP
Marco Lanuzza, Martin Margala, Pasquale Corsonello
ISLPED2
2005 Defect detection in analog and mixed circuits by neural networks using wavelet analysis
abstract
An efficient defect-oriented parametric test method for analog & mixed-signal integrated circuits based on neural network classification of a selected circuit's parameter using wavelet decomposition preprocessing is proposed in this paper. The neural network has been used for detecting catastrophic defects in two experimental analog & mixed-signal CMOS circuits by sensing the abnormalities in selected parameters, observed under defective conditions and by their consequent classification into a proper category. To reduce complexity of the neural network, wavelet decomposition is used to perform preprocessing of the analyzed parameter. Moreover, we show that wavelet analysis brings significant enhancement in the correct classification, and makes the neural network-based test method extremely efficient & versatile for detecting hard-detectable catastrophic defects in analog & mixed-signal circuits.
Viera Stopjaková, Pavol Malosek, Marek Matej, Vladislav Nagy, Martin Margala
IEEE Trans. Reliab.5
2005 Design of wireless on-wafer submicron characterization system
abstract
A wireless technique for the testing of very large scale ICs and wafers is presented. This test technique uses standard CMOS to achieve wireless parametric testing. This technique has virtually no area overhead, minimal power requirements, and no process or design changes are required. Most compelling is that wafer contact is not required, thereby enabling the in-line process control/monitoring of the manufacture of VLSI wafers or chips. Simulations of representative VLSI antenna designs are presented along with experimental results from the implementation of the antenna coupling and communications link. Also presented are specific circuit simulations showing the characteristics of operation under a range of conditions. The technique is demonstrated experimentally in discrete form with operation at voltages as low as 1 V with submilliwatt power levels. This technique can be implemented with a requirement of 1/10 000th the area of a Pentium-class VLSI circuit, allowing contactless testing of wafers before packaging
Brian Moore 0001, Martin Margala, Christopher J. Backhouse
IEEE Trans. Very Large Scale Integr. Syst.2
2004 Design of Wireless Sub-Micron Characterization System
abstract
A wireless technique for characterization of very large scale ICs and wafers is presented. Presented is a test technique that uses standard CMOS without the use of inductors to achieve wireless parametric testing. In terms of existing technologies, this system has virtually no area overhead, minimal power requirements, and no process or design changes are required. A major feature is that wafer/chip contact is not required. Presented are specific circuit simulations showing characteristics of operation under varying conditions. The circuit operation is shown to work down to 1 volt and sub-milliwatt power level at the same time as being 1/10,000/sup th/ the area of a Pentium class VLSI circuit.
Brian Moore 0001, Christopher J. Backhouse, Martin Margala
VTS3
2004 Classification of Defective Analog Integrated Circuits Using Artificial Neural Networks
Viera Stopjaková, Pavol Malosek, Daniel Micusik, Marek Matej, Martin Margala
J. Electron. Test.5
2004 New approach to design for reusability of arithmetic cores in systems-on-chip
Martin Margala, Hongfan Wang
Integr.1
2003 1.8V 0.18µm CMOS Novel Successive Approximation ADC
Martin Margala, Quentin Diduck, Eric Moule
VLSI-SOC1
2003 1-V ADPCM Processor for Low-Power Wireless Applications
Martin Margala, Magdy A. El-Moursy, Ali El-Moursy, Junmou Zhang, Wendi B. Heinzelman
VLSI-SOC1
2003 Deep-Submicron CMOS Design Methodology for High-Performance Low-Power Analog-to-Digital Converters
Martin Margala, John C. Liobe, Quentin Diduck
VLSI-SOC1
2002 Novel design and verification of a 16 x 16-b self-repairable reconfigurable inner product processor
abstract
A novel self-repairable and reconfigurable inner-product processor with low-power, fast CMOS circuits and DFT techniques is presented. It takes the advantage of recently proposed decomposition based arithmetic circuit design approach for simple implementation of the reconfigurations, component replacements, and high-quality tests.The processor can be dynamically reconfigured for two types operations: 4 x 8 x 8-b inner product computation and 16 x 16-b multiplication. The self-repair is provided by choosing a fault-free one from 17 possible architectures during the test, which covers more than 52% transistors for the specified faults. Only one extra bit is needed for all reconfigurations, repairs, and tests. The proposed exhaustive DFT technique greatly reduces the test vector length, from 17*232 to 1.5*213, which is as short as that required by the pseudo-exhaustive DFT method recently reported in literature.
Rong Lin, Martin Margala
ACM Great Lakes Symposium on VLSI2
2002 Minimizing concurrent test time in SoC's by balancing resource usage
abstract
We present a novel test scheduling algorithm for embedded core-based SoC's. Given a system integrated with a set of cores and a set of test resources, we select a test for each core from a set of alternative test sets, and schedule it in a way that evenly balances the resource usage, and ultimately reduce the test application time. Furthermore, we propose a novel approach that groups the cores and assigns higher priority to those with smaller number of alternate test sets. In addition, we also extend the algorithm to allow multiple test sets selection from a set of alternatives to facilitate testing for various fault models.
Danella Zhao, Shambhu J. Upadhyaya, Martin Margala
ACM Great Lakes Symposium on VLSI3
2000 Optimization techniques for maximum power-efficiency of deep sub-micron CMOS digital circuits
abstract
This paper presents results of a study to locate the optimal operating power supply voltage for a maximum power-efficiency operation of CMOS digital circuits in a deep sub-micron environment. The results show that the optimal V/sub DD/ is a strong function of NMOS and PMOS device threshold voltages and their sizes. Depending on the capacitive loading, the optimal V/sub DD/=/spl xi/(|V/sub tp/|+V/sub tn/), where /spl xi/ is (0.9-1.1). The study targets low-power low-voltage applications of digital pass-transistor circuits with multiple operating voltages. The analysis has been performed using BSIM Level 28 models and verified with HSPICE. The experimental circuits were designed in 0.35 /spl mu/m CMOS technology.
Daniel S. C. Kwok, Martin Margala
ISCAS2
1995 A 33 MHz 16-bit gradient calculator for real-time volume imaging
Martin Margala, Nelson G. Durdle, Scott Juskiw, V. James Raso, Doug L. Hill
Comput. Graph.1