Roman Genov

dblp:03/3535 · DBLP profile ↗
← Back
46ranked-venue papers
4as first author
18since 2021 · last 2025
0000-0001-7506-1746ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 16 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3
YearPublicationVenuePosition
2025 MATSORT: Fast and Efficient Approximate Matrix Sorting Algorithm for DNN Applications
abstract
This paper proposes Matrix Sorting (MATSORT), a novel algorithm for compressing sparse matrices to enable efficient execution of sparse Multiply-Accumulate (MAC) operations in parallel computing architectures for Deep Neural Networks (DNNs). Pruning in DNNs often results in weight matrices with a large proportion of zero weights, leading to significant resource and time wastage when processed by parallel computing architecture. MATSORT outperforms state-of-the-art (SoTA) techniques in both compression speed and compression rate—two key factors for efficient sparse MAC operations. The algorithm employs an approximate sorting method and a novel merging strategy, reducing the sorting time complexity while resolving long-standing element conflict issues. These optimizations achieve nearly 100% compression rates across all tested matrices, substantially enhancing power and resource efficiency. Evaluation results show that MATSORT delivers a maximum speedup of 441× for a 1280 × 1280 matrix and a minimum speedup of 1.63× for a 110 × 128 matrix, outperforming existing methods.
Xuanhao Lu, Alireza Ahrar, Mostafa Rahimi Azghadi, Roman Genov, Amirali Amirsoleimani
ISCAS4
2024 In-Memory Transformer Self-Attention Mechanism Using Passive Memristor Crossbar
abstract
Transformers have emerged as the state-of-the-art architecture for natural language processing (NLP) and computer vision. However, they are inefficient in both conventional and in-memory computing architectures as doubling their sequence length quadruples their time and memory complexity due to their self-attention mechanism. Traditional methods optimize self-attention using memory-efficient algorithms or approximate methods, such as locality-sensitive hashing (LSH) attention that reduces time and memory complexity from O(L2) to O(L log L). In this work, we propose a hardware-level solution that further improves the computational efficiency of LSH attention by utilizing in-memory computing with semi-passive memristor arrays. We demonstrate that LSH can be performed with low-resolution, energy-efficient 0T1R arrays performing stochastic memristive vector-matrix multiplication (VMM). Using circuit-level simulation, we show our proposed method is feasible as a drop-in approximation in Large Language Models (LLMs) with no degradation in evaluation metrics. Our results set the foundation for future works on computing the entire transformer architecture in-memory.
Jack Cai, Muhammad Ahsan Kaleem, Roman Genov, Mostafa Rahimi Azghadi, Amirali Amirsoleimani
ISCAS3
2024 Advancing Image Classification with Phase-coded Ultra-Efficient Spiking Neural Networks
abstract
Conventional surrogate Back-Propagation-Through-Time learning in Spiking Neural Networks (SNN) demands excessive energy consumption when simulating over extended time intervals. Moreover, the spike encoding process necessitates intricate hardware support, thus undermining overall efficiency. Additionally, their classification accuracies fall short in comparison to artificial neural networks due to the inherent information loss in spike translation. Therefore, there is a critical need for efficient techniques that can enhance performance without compromising accuracy. In this study, we introduce a novel learning scheme that harnesses lossless phase coding. This approach allows us to achieve minimal inference latency, requiring a maximum of at most 8 simulation steps. Furthermore, our training times exhibit significant reductions when compared to previous single-spike networks. Our experimental results demonstrate that Phase-SNN attains state-of-the-art accuracy levels, achieving 98.6% and 89.6% accuracies on the MNIST and Fashion-MNIST datasets, respectively.
Zhengyu Cai, Hamid Rahimian Kalatehbali, Ben Walters, Mostafa Rahimi Azghadi, Roman Genov, Amirali Amirsoleimani
ISCAS5
2024 NURODE: In-Memory Crossbar Core for Hodgkin-Huxley Model ODE-Based Computations
abstract
In this work, we present a memristor crossbar array-based hardware architecture designed to solve a system of differential equations for the Hodgkin-Huxley neuron model. The system extends from previous works to perform more computations in-memory, reducing the load of software processing and distributing them onto memristive hardware along with cheap shift-and-add operations. We demonstrate the properties of the solutions produced by the system follow the expected behaviours of the Hodgkin-Huxley model under various external currents.
Andy Gong, Mostafa Rahimi Azghadi, Roman Genov, Amirali Amirsoleimani
ISCAS3
2024 BITLITE: Light Bit-wise Operative Vector Matrix Multiplication for Low-Resolution Platforms
abstract
As machine learning (ML) algorithms, particularly neural networks (NN), expand in popularity and capacity, the quest for more efficient computation methods gains momentum. Memristor crossbar technology emerges as a promising alternative to traditional computing units, aiming to address traditional computing challenges. However, conventional matrix-vector multiplication (MVM) methods on these platforms are often plagued by device imperfections and drift. In this work, we introduce an innovative lightweight calculation approach leveraging bit-transformation for MVM, significantly enhancing operation precision and, consequently, the performance of ML algorithms on memristor crossbar platforms. We provide details of the core algorithm and its extensions, furnish digital validation, and simulate its efficacy using an autoencoder (AE) neural network with an extended VTEAM model. Our tests demonstrate an average reconstruction precision improvement of approximately 53.5%. This work’s applicability extends beyond NNs, offering a foundational method for conducting more precise analog MVM operations.
Vince Tran, Demeng Chen, Roman Genov, Mostafa Rahimi Azghadi, Amirali Amirsoleimani
ISCAS3
2024 Spiking Auto-Encoder Using Error Modulated Spike Timing Dependant Plasticity
abstract
Auto-encoders are capable of performing input re-construction through an encoder-decoder structure. These net-works can serve many purposes such as noise removal and anomaly detection, whilst being trained without the need for labelled data. Spiking auto-encoders can utilise asynchronous spikes to potentially improve power and simplify the required hardware. In this work, we propose an efficient spiking auto-encoder with novel error-modulated STDP learning. Our auto-encoder uses the Time To First Spike (TTFS) encoding scheme and needs to update all synaptic weights only once per input. Also, it needs only an average of 8 spikes in its hidden layer for reconstruction, leading to a very sparse and hence potentially power-efficient implementation. We demonstrate decent reconstruction ability for MNIST and the challenging Caltech Face/Motorbike datasets and achieve excellent noise removal from MNIST images.
Ben Walters, Zhengyu Cai, Hamid Rahimian Kalatehbali, Amirali Amirsoleimani, Roman Genov, Jason Kamran Eshraghian, Mostafa Rahimi Azghadi
ISCAS5
2024 SAR-MemPipe: A Hybrid Pipeline-SAR Memristive ADC for Analog Resistive Arrays
abstract
This paper presents a hybrid pipeline SAR ADC with a loop-unrolled structure to reduce the crossbar’s ADC power while maintaining high speed. The 1ststage memristive SAR ADC can fully utilize the TIA originally in the crossbar and avoid the extra MDAC in pipeline ADC. Further, memristive weight calibration and a new resistive alternated binary search are implemented on the 1ststage to maintain the TIA gain and ADC’s accuracy. Both stage’s ADC are loop-unroll to eliminate the delay brought by SAR logic for high speed. Through multiple simulations, the design is demonstrated to be robust to the frequency and mismatch variations with the highest sampling frequency reaching 300MHz and SNDR up to 65.1dB in 9MHz input. The power consumption is designed to be as low as 6.7mW, which helps the ADC to achieve a 15.2fJ/conv FoM. Not limited to the crossbar, the presented ADC also shows promising potential for applications in various fields (biomedical, IoT etc.) for general purposes.
Hao You, Jianxiong Xu, Amirali Amirsoleimani, Mostafa Rahimi Azghadi, Roman Genov
ISCAS5
2024 WALLAX: A memristor-based Gaussian random number generator
Xuening Dong, Amirali Amirsoleimani, Mostafa Rahimi Azghadi, Roman Genov
Neurocomputing4
2024 SITU: Stochastic input encoding and weight update thresholding for efficient memristive neural network in-situ training
Xuening Dong, Roman Genov, Mostafa Rahimi Azghadi, Amirali Amirsoleimani
Neurocomputing3
2023 SEVDA: Singular Value Decomposition Based Parallel Write Scheme for Memristive CNN Accelerators
abstract
Von Neumann architecture-based deep neural network architectures are fundamentally bottlenecked by the need to transfer data from memory to compute units. Memristor crossbar-based accelerators overcome this by leveraging Kirchoff's law to perform matrix-vector multiplication (MVM) in-memory. They still, however, are relatively inefficient in their device programming schemes, requiring individual devices to be written sequentially or row-by-row. Parallel writing schemes have recently emerged, which program entire crossbars simultaneously through the outer product of bit-line and word-line voltages and pulse widths respectively. We propose a scheme that leverages singular value decomposition and low-rank approximation to generate all word-line and bit-line vectors needed to program a convolutional neural network (CNN) onto a memristive crossbar-based accelerator. Our scheme reduces programming latency by 90% from row-by-row programming schemes, while maintaining high test accuracy on state of the art image classification models.
Ali Alshaarawy, Roman Genov, Amirali Amirsoleimani
ISCAS2
2023 HESSPROP: Mitigating Memristive DNN Weight Mapping Errors with Hessian Backpropagation
abstract
A universal objective function to minimize mem-ristive crossbar deep neural network weight mapping errors through Hessian backpropagation (HessProp) is presented. Hes-sProp minimizes the$L_{2}$norm of the neural network gradient to achieve a flat minima in a neural network's weight space. We hypothesize that this leads to robustness against small perturbations of weights. The stochastic weight mapping phe-nomenon on memristor crossbars is simulated, and the proposed method was evaluated on image classification tasks using the MNIST dataset. The result demonstrates on average 40.81% and 41.45% groundbreaking accuracy increase for distilled and large memristive convolutional neural networks in worst-case scenarios.
Jack Cai, Muhammad Ahsan Kaleem, Amirali Amirsoleimani, Roman Genov
ISCAS4
2023 A Survey of Ensemble Methods for Mitigating Memristive Neural Network Non-idealities
abstract
In this work, ensemble methods are presented and tested as universal ways to improve the performance of Memristive Deep Neural Networks (MDNNs) with non-idealities. The Generalized Ensemble Method and Weighted Voting ensemble methods improve the accuracy of classification on the MNIST dataset by 6.5% and 6.6% respectively, thus showing that they are more effective than basic Ensemble Averaging which has been investigated before, as well as other methods such as Voting. Different weighting schemes for Weighted Voting were tested, and we present Algorithm 1 and 2, which are the theoretically and experimentally optimal weighting schemes respectively. Our work serves as a guideline for choosing ensemble methods for MDNNs.
Muhammad Ahsan Kaleem, Jack Cai, Amirali Amirsoleimani, Roman Genov
ISCAS4
2023 HUXIN: In-Memory Crossbar Core for Integration of Biologically Inspired Stochastic Neuron Models
abstract
In this work, we solve nonlinear systems of ordinary differential equations coupled to noisy forcing, commonly used for models of neurons such as the Hodgkin-Huxley equation, over a memristor crossbar based computing system. We demonstrate stability and faithfulness of the distributions even under the effects of nonidealities of the memristors and the system itself. We investigate the properties of the dynamical systems under quantization faithfulness, varying the level of precision of the fixed point integer representation and concluding that 24 bits is enough for solution of the Hodgkin-Huxley equations, demonstrating that our solver can operate with both high precision and achieve speedups with low precision approximate computation.
Louis Primeau, Xuening Dong, Amirali Amirsoleimani, Roman Genov
ISCAS4
2023 SSCAE: A Neuromorphic SNN Autoencoder for sc-RNA-seq Dimensionality Reduction
abstract
Single-cell RNA sequencing is an emerging technique in the field of biology that departs radically from the previous assumption of gene-expression homogeneity within a tissue. The large quantity of data generated by this technology enables discoveries of cellular biology and disease mechanics that were previously not possible, and calls for accurate, scalable, and efficient processing pipelines. In this work, we propose SSCAE (spiking single-cell autoencoder), a novel SNN-based autoencoder for sc-RNA-seq dimensionality reduction. We apply this architecture to a variety of datasets, and the results show that it can match and surpass the performance of current state-of-the-art techniques. Moreover, the potential of this technique lies in its ability to be scaled up and to take advantage of neuromorphic hardware, circumventing the memory bottleneck that currently limits the size of sequencing datasets that can be processed.
Tim Zhang, Amirali Amirsoleimani, Jason Kamran Eshraghian, Mostafa Rahimi Azghadi, Roman Genov, Yu Xia 0002
ISCAS5
2022 PRUNIX: Non-Ideality Aware Convolutional Neural Network Pruning for Memristive Accelerators
abstract
In this work, PRUNIX, a framework for training and pruning convolutional neural networks is proposed for deployment on memristor crossbar based accelerators. PRUNIX takes into account the numerous non-ideal effects of memristor crossbars including weight quantization, state-drift, aging and stuck-at-faults. PRUNIX utilises a novel Group Sawtooth Regularization intended to improve non-ideality tolerance as well as sparsity, and a novel Adaptive Pruning Algorithm (APA) intended to minimise accuracy loss by considering the sensitivity of different layers of a CNN to pruning. We compare our regularization and pruning methods with other standards on multiple CNN architectures, and observe an improvement of 13% test accuracy when quantization and other non-ideal effects are accounted for with an overall sparsity of 85%, which is similar to other methods.
Ali Alshaarawy, Amirali Amirsoleimani, Roman Genov
ISCAS3
2022 HYPERLOCK: In-Memory Hyperdimensional Encryption in Memristor Crossbar Array
abstract
We present a novel cryptography architecture based on memristor crossbar array, binary hypervectors, and neural network. Utilizing the stochastic and unclonable nature of memristor crossbar and error tolerance of binary hypervectors and neural network, implementation of the algorithm on memristor crossbar simulation is made possible. We demonstrate that with an increasing dimension of the binary hypervectors, the nonidealities in the memristor circuit can be effectively controlled. At the fine level of controlled crossbar non-ideality, noise from memristor circuit can be used to encrypt data while being sufficiently interpretable by neural network for decryption. We applied our algorithm on image cryptography for proof of concept, and to text en/decryption with 100% decryption accuracy despite crossbar noises. Our work shows the potential and feasibility of using memristor crossbars as an unclonable stochastic encoder unit of cryptography on top of their existing functionality as a vectormatrix multiplication acceleration device.
Jack Cai, Amirali Amirsoleimani, Roman Genov
ISCAS3
2022 Design Space Exploration of Dense and Sparse Mapping Schemes for RRAM Architectures
abstract
The impact of device and circuit-level effects in mixed-signal Resistive Random Access Memory (RRAM) accelerators typically manifest as performance degradation of Deep Learning (DL) algorithms, but the degree of impact varies based on algorithmic features. These include network architecture, capacity, weight distribution, and the type of inter-layer connections. Techniques are continuously emerging to efficiently train sparse neural networks, which may have activation sparsity, quantization, and memristive noise. In this paper, we present an extended Design Space Exploration (DSE) methodology to quantify the benefits and limitations of dense and sparse mapping schemes for a variety of network architectures. While sparsity of connectivity promotes less power consumption and is often optimized for extracting localized features, its performance on tiled RRAM arrays may be more susceptible to noise due to under-parameterization, when compared to dense mapping schemes. Moreover, we present a case study quantifying and formalizing the trade-offs of typical non-idealities introduced into l-Transistor-l-Resistor (ITIR) tiled memristive architectures and the size of modular crossbar tiles using the CIFAR-10 dataset.
Corey Lammie, Jason Kamran Eshraghian, Chenqi Li, Amirali Amirsoleimani, Roman Genov, Wei Lu 0003, Mostafa Rahimi Azghadi
ISCAS5
2022 SDEX: Monte Carlo Simulation of Stochastic Differential Equations on Memristor Crossbars
abstract
Here we present stochastic differential equations (SDEs) on a memristor crossbar, where the source of gaussian noise is derived from the random conductance due to ion drift in the devices during programming. We examine the effects of line resistance on the generation of normal random vectors, showing the skew and kurtosis are within acceptable bounds. We then show the implementation of a stochastic differential equation solver for the Black-Scholes SDE, and compare the distribution with the analytic solution. We determine that the random number generation works as intended, and calculate the energy cost of the simulation.
Louis Primeau, Amirali Amirsoleimani, Roman Genov
ISCAS3
2020 End-to-End Video Compressive Sensing Using Anderson-Accelerated Unrolled Networks
abstract
Compressive imaging systems with spatial-temporal encoding can be used to capture and reconstruct fast-moving objects. The imaging quality highly depends on the choice of encoding masks and reconstruction methods. In this paper, we present a new network architecture to jointly design the encoding masks and the reconstruction method for compressive high-frame-rate imaging. Unlike previous works, the proposed method takes full advantage of denoising prior to provide a promising frame reconstruction. The network is also flexible enough to optimize full-resolution masks and efficient at reconstructing frames. To this end, we develop a new dense network architecture that embeds Anderson acceleration, known from numerical optimization, directly into the neural network architecture. Our experiments show the optimized masks and the dense accelerated network respectively achieve 1.5 dB and 1 dB improvements in PSNR without adding training parameters. The proposed method outperforms other state-of-the-art methods both in simulations and on real hardware. In addition, we set up a coded two-bucket camera for compressive high-frame-rate imaging, which is robust to imaging noise and provides promising results when recovering nearly 1,000 frames per second.
Miao Qi, Rahul Gulve, Mian Wei, Roman Genov, Kiriakos N. Kutulakos, Wolfgang Heidrich
ICCP5
2018 Coded Two-Bucket Cameras for Computer Vision
Mian Wei, Navid Sarhangnejad, Zhengfan Xia, Nikita Gusev, Nikola Katic, Roman Genov, Kiriakos N. Kutulakos
ECCV (3)6
2017 Two-electrode impedance-sensing cardiac rhythm monitor for charge-aware shock delivery in cardiac arrest
abstract
This paper presents a two-electrode cardiac rhythm-monitoring amplifier that also senses the impedance between the same electrodes. Online monitoring of the body and electrodes impedances during high-voltage (HV) energy delivery or defibrillating, is needed to precisely estimate the amount of charge being delivered to the cardiac muscle. By adjusting the HV stimulus waveform, the amount of charge and thus energy delivered to the cardiac muscle can be precisely controlled for each patient. Additionally, online monitoring of the cardiac impedance ensures proper electrodes connectivity during both electrocardiogram (ECG) signal monitoring and HV energy delivery. The proposed amplifier employs two current sources of the same value and opposite direction (one at each electrode) to adaptively source through the chest an out-of-band (1 kHz) sinusoidal current and measure the resulting voltage to infer the cardiac impedance. This is done in superposition with ECG monitoring due to the same modality of the signal (i.e. voltage). A functional prototype with low-cost off-the-shelf components was implemented and tested demonstrating the feasibility of the proposed technique.
Reza Pazhouhandeh, Omid Shoaei, Roman Genov
ISCAS3
2017 Bit-pragmatic deep neural network computing
abstract
Deep Neural Networks expose a high degree of parallelism, making them amenable to highly data parallel architectures. However, data-parallel architectures often accept inefficiency in individual computations for the sake of overall efficiency. We show that on average, activation values of convolutional layers during inference in modern Deep Convolutional Neural Networks (CNNs) contain 92% zero bits. Processing these zero bits entails ineffectual computations that could be skipped. We propose Pragmatic (PRA), a massively data-parallel architecture that eliminates most of the ineffectual computations on-the-fly, improving performance and energy efficiency compared to state-of-the-art high-performance accelerators [5]. The idea behind PRA is deceptively simple: use serial-parallel shift-and-add multiplication while skipping the zero bits of the serial input. However, a straightforward implementation based on shift-and-add multiplication yields unacceptable area, power and memory access overheads compared to a conventional bit-parallel design. PRA incorporates a set of design decisions to yield a practical, area and energy efficient design.
Jorge Albericio, Alberto Delmas Lascorz, Patrick Judd, Sayeh Sharify, Gerard O'Leary, Roman Genov, Andreas Moshovos
MICRO6
2016 Battery-less modular responsive neurostimulator for prediction and abortion of epileptic seizures
abstract
An inductively-powered implantable microsystem for monitoring and treatment of intractable epilepsy is presented. The miniaturized system is comprised of two mini-boards and a power receiver coil. The first board hosts a 24-channel neurostimulator SoC developed in a 0.13μm CMOS technology and performs neural recording, electrical stimulation and onchip digital signal processing. The second board communicates recorded brain signals as well as signal processing results wirelessly, and generates different supply and bias voltages for the neurostimulator SoC and other external components. The multi-layer flexible coil receives inductively-transmitted power and sends it to the second board for power management. The system is sized at 2 × 2 × 0.7 cm3, weighs 6 grams, and is validated in control of chronic seizures in vivo in freely-moving rats.
Hossein Kassiri, Nima Soltani, Muhammad Tariqus Salam, José Luis Pérez Velazquez, Roman Genov
ISCAS5
2016 A compact low-power VLSI architecture for real-time sleep stage classification
abstract
A wearable-optimized implementation of a sleep stage classification algorithm that has low detection latency, high detection accuracy and low resource consumption is developed and successfully implemented on a low-power FPGA microsystem for closed-loop electrical brain stimulation. This implementation uses EEG and EMG signals as inputs and classifies stages of sleep. By structurally merging multichannel FIR and window averaging filters into one reconfigurable, multipurpose filter, the new implementation maintains a sleep detection accuracy of 79.7%, a REM detection sensitivity of 98.2%, a REM detection specificity of 89.2% and a detection latency of 0.982 ms, while consuming 6.8 times fewer logic elements and 96.28% less power compared with the current state of the art implementation. With its high performance and low resource usage, this implementation enables a low-power wearable microsystem to perform neural recording, real-time REM sleep stage detection, and closed-loop responsive brain stimulation as a tool to study the mechanisms of neurodegenerative diseases.
Peter Zhi Xuan Li, Hossein Kassiri, Roman Genov
ISCAS3
2016 Tradeoffs between wireless communication and computation in closed-loop implantable devices
abstract
This paper discusses general tradeoffs between wireless communication and computation in closed-loop implantable medical devices for neurological applications. Closed-loop devices enable neural monitoring, automated diagnostics and treatment of neurological disorders. Several topologies for the loop a re discussed, including within the implant, as well as implemented with a wearable, handheld or stationary processor. Common wireless communication data rate and range requirements and algorithmic computational requirements are summarized. As a case study, a 0.13 μm CMOS neurostimulator SoC for closed-loop treatment of intractable epilepsy is presented. Its triple-band radio with a 1m 230Mbps pulse-radio, a 2m 46Mbps pulse-radio 2, and a 10m 1.2Mbps FSK radio provides a versatile transcutaneous interface. The in-implant processor has constrained computational resources which results in a limited detection performance - seizure detection sensitivity of 87%. A higher-performance signal processing algorithm implemented on a stationary device within a loop enhances the seizure detection performance which was improved to a sensitivity of 98% with three times fewer false alarms. This comes at the cost of an increased wireless transmitter power budget, if communicated directly. These results illustrate a fundamental tradeoff between the communication and computation in closed-loop electronic therapies for neurological disorders.
Muhammad Tariqus Salam, Hossein Kassiri, Nima Soltani, José Luis Pérez Velazquez, Roman Genov
ISCAS6
2013 Similarity-index early seizure detector VLSI architecture
abstract
A low power VLSI architecture implementing an algorithm for early seizure detection in epileptic patients using intracranial or scalp EEG data is proposed. The algorithm tested over more than 40 hours of recording from standard databases achieves a best-case result of 100% sensitivity at a false positive rate of 0.2 per hour. The algorithm is programmed on an FPGA and was experimentally validated along with a neural recording SoC chip to demonstrate a real-time seizure detection microsystem.
Amogh Vidwans, Karim Abdelhalim, Roman Genov
ISCAS3
2012 Compact chopper-stabilized neural amplifier with low-distortion high-pass filter in 0.13µm CMOS
abstract
A compact and low-distortion neural recording amplifier is presented. The amplifier consists of two stages of amplification using capacitive feedback to set a gain of 54dB. To minimize flicker noise in the 1st stage, internal chopping is utilized at the folded node of the OTA, resulting in flicker noise contribution from the input differential pair only. A low-distortion constant-VGSfeedback circuit to set a low frequency high-pass pole is introduced. It is less sensitive to the output swing than the conventional sub-threshold MOS circuit. The amplifier fabricated in a standard 1.2V 0.13µm CMOS technology occupies 125×175µm2and achieves an NEF of 4.4, an input-referred noise of 4.7µV over a 5kHz bandwidth, a CMRR of 75dB and a THD of −50dB for a 0.6V output swing.
Karim Abdelhalim, Roman Genov
ISCAS2
2012 CMOS 3-T digital pixel sensor with in-pixel shared comparator
abstract
A CMOS digital pixel sensor (DPS) VLSI architecture with in-pixel one-bit quantization is presented. A single column-parallel comparator is shared by all pixels in the column. This results in a compact 3-T pixel implementation. By eliminating the in-pixel source follower the pixel effective power dissipation is reduced by over two orders of magnitude compared to a conventional 3-T pixel. A 64×64 DPS test prototype with 10μm pixel pitch has been fabricated in 0.35μm standard CMOS and experimentally characterized.
Derek Ho, P. Glenn Gulak, Roman Genov
ISCAS3
2012 Single-filter multi-color CMOS fluorescent contact sensing microsystem
abstract
A multi-color fluorescent contact sensing microsystem is presented. The microsystem employs a CMOS field-modulated color sensor (FCS) to spectrally detect and differentiate among multiple emission bands, requiring only one on-CMOS longpass filter. A FCS prototype has been fabricated in a standard 0.35μm CMOS technology. The multi-color imaging capability of the FCS microsystem has been validated in the detection of green-emitting and red-emitting quantum dots (QDs) with QD concentration detection limits of 313nM and 78nM, respectively.
Derek Ho, M. Omair Noor, Ulrich Krüll, P. Glenn Gulak, Roman Genov
ISCAS5
2012 Bidirectional current conveyer with chopper stabilization and dynamic element matching
abstract
A compact and accurate current conveyer for interfacing with a three-electrode electrochemical sensor is presented. It employs chopper stabilization to reduce transistor sizes for a given flicker noise. A current-copying circuit generates a mirror image of the sensor current for bidirectional conveying. It employs dynamic element matching to remove the mismatch in the current mirrors. The current conveyer prototyped in 0.13µm CMOS consumes 4µW from a 1.2V supply. It achieves an input-referred noise of 0.13pArms over a 1kHz bandwidth with a dynamic range of 8.6pA to 350nA. The prototype has been validated in DNA catalytic reporter detection.
Hamed Mazhab-Jafari, Roman Genov
ISCAS2
2011 CMOS DAC-sharing stimulator for neural recording and stimulation arrays
abstract
A compact biphasic neural stimulator for use in multi-channel integrated neural recording and stimulation interfaces is presented. The stimulator is a part of an envisioned closed-loop implantable microsystem for adaptive neural stimulation. The stimulator reuses the capacitive DAC inside the recording SAR ADC to provide 8-bit current amplitude resolution and reconfigures the digital SAR controller logic to provide 4-bit tunability of the duty cycle of the current pulse. A voltage-to-current converter and a current driver are the only additional circuits required. An OTA is reused to provide accurate current matching and to remove excess charge on the electrode. The stimulator is implemented in a standard 0.13μm CMOS technology. Post-layout and Monte Carlo simulation results show 8-bit functionality from 5μA to 1.2mA, a 2.5V swing and 1 percent current matching between positive and negative pulses.
Karim Abdelhalim, Roman Genov
ISCAS2
2010 CMOS current-copying neural stimulator with OTA-sharing
abstract
We present a compact current mode stimulator for utilization in neural stimulation arrays. The stimulator is implemented in a 0.35μm CMOS technology and occupies an area of 50μm×400μm. The memory in every current driver allows for simultaneous stimulation on multiple active channels when used as part of an array of stimulators. Circuit reuse in the stimulator and utilization of a single DAC yield a compact and low-power implementation. The current driver dissipates a quiescent power of 2.76μW. The stimulator can output current in the range of 10μA-250μA.
Ruslana Shulyzki, Karim Abdelhalim, Roman Genov
ISCAS3
2009 A Fully Differential CMOS Potentiostat
abstract
A CMOS potentiostat for chemical sensing in a noisy environment is presented. The potentiostat measures bidirectional electrochemical redox currents proportional to the concentration of a chemical down to pico-ampere range. The fully differential architecture with differential recording electrodes suppresses the common mode interference. A 200 mumtimes200 mum prototype was fabricated in a standard 0.35 mum standard CMOS technology and yields a 70 dB dynamic range. The in-channel analog-to-digital converter (ADC) performs 16-bit current-to-frequency quantization. The integrated potentiostat functionality is validated in electrical and electrochemical experiments.
Meisam Honarvar Nazari, Roman Genov
ISCAS2
2009 CMOS Image Compression Sensor with Algorithmically-multiplying ADCs
abstract
A 128times128 CMOS image compression sensor fabricated in a 0.35 mum CMOS process is reported. It computes block-matrix and convolutional image transforms with digital kernels of up to 8times8 pixels directly on the focal plane. A pixel output is sampled only when the corresponding bit of the kernel coefficient is one. Bit-wise accumulation of adjacent pixel outputs in a column is performed by the switched-capacitor accumulator circuit. A column-parallel algorithmic multiplying ADC performs binary-weighted summation by adding the accumulator circuit outputs with cyclic residues of the same binary weight. The signal range is maintained by generating two bits per cycle. The imager performs three computations per pixel readout. Image compression experimental results at 30 fps and 8-bit output resolution are presented.
Alireza Nilchi, Joseph N. Y. Aziz, Roman Genov
ISCAS3
2009 128-channel Fully Differential Digital Neural Recording and Stimulation Interface
abstract
We present a fully differential 128-channel integrated neural interface. It consists of an array of 8times16 low-power low-noise signal recording and generation channels for electrical neural activity monitoring and stimulation, respectively. The recording channel has two stages of signal amplification and conditioning with a programmable gain of 54 dB to 73 dB, and a fully differential 8-bit column-parallel successive approximation (SAR) analog-to-digital converter (ADC). The design is implemented in a 0.35 mum CMOS technology with the channel pitch of 200 mum. The total measured power consumption of each recording channel including the SAR ADC is 15.5 muW. The measured input referred noise is 6.08 muVrmsover a 5 kHz bandwidth.
Farzaneh Shahrokhi, Karim Abdelhalim, Roman Genov
ISCAS3
2009 A Hybrid Thin-film/CMOS Fluorescence Contact Imager
abstract
A hybrid thin film/CMOS microsystem for fluorescence contact imaging is presented. The microsystem integrates a high-performance optical filter and a 128times128-pixel imager fabricated in a 0.35 mum technology. The thin-film filter is fabricated and characterized prior to assembly. Its optical density (OD) is over 6.0 at the wavelength of interest. The performance of the microsystem is experimentally validated by imaging conventional Cy3 fluorophore spots using a low-cost pen-sized laser. The emission intensity as a function of fluorophore concentration is measured with the estimated sensitivity of 20 fluorophore/ mum2.
Ritu Raj Singh, Derek Ho, Alireza Nilchi, Roman Genov, P. Glenn Gulak
ISCAS4
2007 In Vitro Epileptic Seizure Prediction Microsystem
abstract
The architecture and VLSI implementation of an epileptic seizure prediction microsystem are presented. The microsystem comprises a neural recording interface and a seizure prediction processor. The two functional blocks have been prototyped in a 0.35μCMOS technology and experimentally characterized. The integrated microsystem is validated in predicting the onsets of seizures off line in an in vitro epilepsy model of recurrent spontaneous seizures in the hippocampus of mice.
Joseph N. Y. Aziz, Rafal Karakiewicz, Roman Genov, Alan W. L. Chiu, Berj L. Bardakjian, Miron Derchansky, Peter L. Carlen
ISCAS3
2006 Electro-chemical multi-channel integrated neural interface technologies
abstract
We present a comparative review of two multichannel integrated neural interface technologies. The first integrated neural interface prototype performs simultaneous current-mode acquisition of 16 independent channels of redox currents ranging five orders of magnitude in dynamic range over four scales down to hundreds of picoamperes. The second neural interface acquires neural field potentials in microvolts to millivolts range on a 16times16-electrode microarray in voltage mode. Each microsystem features programmable gain amplifiers, tunable band filters, configurable sample-and-hold circuits, and is ready for external analog-to-digital conversion. The current-mode and voltage-mode neural interface prototypes have been experimentally validated in chemical and electrical neural activity monitoring respectively. Side-by-side quantitative comparison of the two neural interface technologies is given
Joseph N. Y. Aziz, Roman Genov
ISCAS2
2006 256-channel integrated neural interface and spatio-temporal signal processor
abstract
We present an architecture and VLSI implementation of a distributed neural interface and spatio-temporal signal processor. The integrated neural interface records neural activity simultaneously on 256 voltage-mode channels. Each channel implements differential signal acquisition, amplification and band-pass filtering. An array of in-channel double-memory sample-and-hold cells stores two 16 /spl times/ 16 electronic images of distributed neural activity consecutively in time. A column-parallel double sampling circuit performs frame differencing in order to identify spatio-temporal neural activity patterns. A 3 mm /spl times/ 4.5 mm integrated prototype was fabricated in a 0.35 /spl mu/m CMOS technology. The functionality of the neural interface was experimentally demonstrated in extracellular in vitro recordings from the hippocampus of mice. The utility of the on-sensory-plane signal processor was validated in simulated wavefront detection performed on experimentally measured distributed neural activity recording.
Joseph N. Y. Aziz, Roman Genov, Berj L. Bardakjian, Miron Derchansky, Peter L. Carlen
ISCAS2
2006 Real-time seizure monitoring and spectral analysis microsystem
abstract
We present a neural recording and spectral analysis integrated microsystem. It is the instrumentational and computational core of an envisioned miniature implantable brain implant for automated epileptic seizure therapy. The microsystem combines two functional blocks: the neural recording interface and the spectral analysis processor. The neural interface contains 256 signal acquisition channels recording neural field potentials from an array of 16 times 16 electrodes simultaneously, in a distributed fashion. The spectral analysis processor computes a wavelet-based time-frequency map (spectrogram) of the neural recording. We demonstrate the functionality of the integrated microsystem in real-time epileptic seizure monitoring and spectral analysis, as necessary for subsequent automated seizure prediction and prevention
Joseph N. Y. Aziz, Rafal Karakiewicz, Roman Genov, Berj L. Bardakjian, Miron Derchansky, Peter L. Carlen
ISCAS3
2006 Algorithmic Delta-Sigma-modulated FIR filter
abstract
We present an algorithmic DeltaSigma-modulated FIR filter which computes digital convolution of a continuous-time analog input signal with a programmable digital impulse response. Selective sampling of the input signal controlled by unary-encoded FIR coefficients yields bit-serial analog-digital multiplication. A DeltaSigma-modulated analog-to-digital converter samples a time-varying input at multiple instances in time generating a quantized version of the average of all weighted samples. Computational throughput of an arbitrary FIR filter is maximized by algorithmic resampling of the modulation residue to obtain higher resolution bits. This yields a bit resolution linear in the number of conversion cycles. A 1.9mm times 1.3mm 128-channel FIR filter integrated prototype was fabricated in a 0.35 mum CMOS technology. It yields a computational throughput of up to 3.8 GMACS, with computational quantization time, power dissipation, and integration area comparable to those in a conventional oversampling analog-to-digital converter
Ashkan Olyaei, Roman Genov
ISCAS2
2003 Silicon Support Vector Machine with On-Line Learning
abstract
Training of support vector machines (SVMs) amounts to solving a quadratic programming problem over the training data. We present a simple on-line SVM training algorithm of complexity approximately linear in the number of training vectors, and linear in the number of support vectors. The algorithm implements an on-line variant of sequential minimum optimization (SMO) that avoids the need for adjusting select pairs of training coefficients by adjusting the bias term along with the coefficient of the currently presented training vector. The coefficient assignment is a function of the margin returned by the SVM classifier prior to assignment, subject to inequality constraints. The training scheme lends efficiently to dedicated SVM hardware for real-time pattern recognition, implemented using resources already provided for run-time operation. Performance gains are illustrated using the Kerneltron, a massively parallel mixed-signal VLSI processor for kernel-based real-time video recognition.
Roman Genov, Shantanu Chakrabartty, Gert Cauwenberghs
Int. J. Pattern Recognit. Artif. Intell.1
2003 Kerneltron: support vector "machine" in silicon
abstract
Detection of complex objects in streaming video poses two fundamental challenges: training from sparse data with proper generalization across variations in the object class and the environment; and the computational power required of the trained classifier running real-time. The Kerneltron supports the generalization performance of a support vector machine (SVM) and offers the bandwidth and efficiency of a massively parallel architecture. The mixed-signal very large-scale integration (VLSI) processor is dedicated to the most intensive of SVM operations: evaluating a kernel over large numbers of vectors in high dimensions. At the core of the Kerneltron is an internally analog, fine-grain computational array performing externally digital inner-products between an incoming vector and each of the stored support vectors. The three-transistor unit cell in the array combines single-bit dynamic storage, binary multiplication, and zero-latency analog accumulation. Precise digital outputs are obtained through oversampled quantization of the analog array outputs combined with bit-serial unary encoding of the digital inputs. The 256 input, 128 vector Kerneltron measures 3 mm/spl times/3mm in 0.5 /spl mu/m CMOS, delivers 6.5 GMACS throughput at 5.9 mW power, and attains 8-bit output resolution.
Roman Genov, Gert Cauwenberghs
IEEE Trans. Neural Networks1
2002 Neuromorphic processor for real-time biosonar object detection
abstract
Real-time classification of objects from active sonar echo-location requires a tremendous amount of computation, yet bats and dolphins perform this task effortlessly. To bridge the gap between human-engineered and biosonar system performance, we developed special-purpose hardware tailored to the parallel distributed nature of the computation performed in biology. The implemented architecture contains a cochlear filterbank front-end performing time-frequency feature extraction, and a kernel-based neural classifier for object detection. Based on analog programmable components, the front-end can be configured as a parallel or cascaded bandpass filterbank of up to 34 channels spanning the 10 to 150 kHz range. The classifier is implemented with the Kerneltron, a massively parallel mixed-signal Support Vector “Machine” in silicon delivering a throughput in excess of a trillion (1012) multiply-accumulates per second for every Watt of power dissipation. The system has been evaluated on detection of mine-like objects using linear frequency modulation active sonar data (LFM2, CSS Panamy City), achieving an out-or-sample performance of 93% correct single-ping detection at 5% false positives, and a real-time throughput of 250 pings per second.
Gert Cauwenberghs, R. Timothy Edwards, Yunbin Deng, Roman Genov, David Lemonds
ICASSP4
2001 Stochastic Mixed-Signal VLSI Architecture for High-Dimensional Kernel Machines
abstract
A mixed-signal paradigm is presented for high-resolution parallel inner- product computation in very high dimensions, suitable for efficient im- plementation of kernels in image processing. At the core of the externally digital architecture is a high-density, low-power analog array performing binary-binary partial matrix-vector multiplication. Full digital resolution is maintained even with low-resolution analog-to-digital conversion, ow- ing to random statistics in the analog summation of binary products. A random modulation scheme produces near-Bernoulli statistics even for highly correlated inputs. The approach is validated with real image data, and with experimental results from a CID/DRAM analog array prototype in 0.5
Roman Genov, Gert Cauwenberghs
NIPS1
1999 Learning to navigate from limited sensory input: experiments with the Khepera microrobot
abstract
The goal of this work is to augment reinforcement learning techniques for autonomous robot navigation with a state space encoding more representative of the actual state of the robot in its environment, than available from direct sensor readings. A second goal is to demonstrate the approach in a real-world setting, using the microrobot Khepera (K-Team, Lausanne, Switzerland). The choice of state representation is one of the most critical factors in the performance of reinforcement learning algorithms. The technique of inferring relative positional information indirectly from sensor readings, through unsupervised learning, is an important novel contribution of this work. As demonstrated in the robot experiments, the technique allows to optimally perform sensor fusion and avoids the need of more elaborate sensors conveying explicit information on position coordinates.
Roman Genov, Srinadh Madhavapeddi, Gert Cauwenberghs
IJCNN1