EDBT 2026 Demo / reviewers in the wild / expert
Miguel E. Figueroa
dblp:50/3882
· DBLP profile ↗
40ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-5033-432XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 4 since 2021Artificial intelligence and machine learning · 11 · 1 first-authorComputer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Streaming algorithm and hardware accelerator for high-throughput entropy estimation of network flows in sliding windows
Yaime Fernández, Javier E. Soto, Carolina Gallardo-Pavesi, Yasmany Prieto, Cecilia Hernández, Miguel E. Figueroa |
Comput. Commun. | 6 |
| 2025 | A streaming algorithm and hardware accelerator for top-K flow detection in network trafficabstractIdentifying the largest K flows in network traffic is an important task for applications such as flow scheduling and anomaly detection, which aim to improve network efficiency and security. However, accurately estimating flow frequencies is challenging due to the large number of flows and increasing network speeds. Hardware accelerators are often used in this endeavor due to their high computational power, but their limited amount of on-chip memory constrains their performance. Various sketch-based algorithms have been proposed to estimate properties of traffic such as frequency, with lower memory usage and theoretical bounds, but they often under perform with the skewed distribution of network traffic. In this work, we propose an algorithm for top- K identification using a modified TowerSketch and a priority queue array. Tested on real traffic traces, we identify the top- K flows, with K up to 32,768, with a precision of more than 0.94, and estimate their frequency with an average relative error under $1.96 \%$. We designed and implemented an accelerator for this algorithm on an AMD Virtex U280 UltraScale+ FPGA, which processes one packet per cycle at 392 MHz, reaching a minimum line rate of more than 200 Gbps. Carolina Gallardo-Pavesi, Yaime Fernández, Javier E. Soto, Cecilia Hernández, Miguel E. Figueroa |
DSD | 5 |
| 2025 | Using virtual function replacement to mitigate 0-day attacks in a multi-vendor NFV-based networkabstractNetwork Function Virtualization (NFV) is an enabling technology to handle today’s wide variety of services and to match the traffic demands to the available network resources dynamically. However, like all software solutions, they are prone to 0-day vulnerabilities that can be shared among several implementations. Consequently, a large portion of the network may be compromised during a correlated attack. In this work, we propose a Virtual Network Function (VNF) replacement strategy to minimize the impact of successive attacks after a 0-day attack has been launched to an NFV-based network. Our method consists of two steps. First, given a specific set of NFV platforms impaired by the attack, we estimate, for each VNF, the conditional probability that the VNF was targeted. Second, we solve a bi-objective optimization problem which delivers a set of possible VNF replacements, that balance financial cost and connectivity metrics. The network administrator can use this information to select a tailored solution to face successive attacks that exploit the same vulnerability. Our results show that, although a solution that guarantees full network connectivity might not exist for all attack patterns, our method is always able to provide a set of optimal solutions that mitigate the effect of correlated attacks on the network. • NFV-based networks are prone to correlated attacks through shared vulnerabilities. • Bayes theorem allows detecting targeted VNFs based on the impaired nodes. • Compromised VNFs can be replaced to avoid successive attacks. • Connectivity metrics allow measuring post-replacement network reliability. • A suitable replacement balances connectivity metrics and financial cost. Yasmany Prieto, Christian Vega Caicedo, Miguel E. Figueroa |
Comput. Networks | 3 |
| 2024 | A Hardware Accelerator for Quantile Estimation of Network Packet AttributesabstractMeasuring statistical properties of network traffic can improve our understanding of traffic distribution and help us detect short and long-term anomalies. However, computing the exact value of these properties requires significant storage and computation, which limits their application in high-speed networks. Hardware accelerators provide the computational power to process a large sequence of network packets with high throughput and low latency, but their performance is ultimately limited by the amount of on-chip memory available on the device. Consequently, researchers have proposed sketch-based algorithms to estimate properties of a data stream with sub linear memory and theoretical estimation error bounds. In this paper, we present a streaming algorithm and hardware accelerator for quantile estimation, which is based on the architecture of the KLL sketch. Implemented on an AMD Virtex XCU55 UltraScale+ FPGA, the accelerator operates at a clock frequency of 356 MHz, thereby achieving a minimum line rate of 182 Gbps and a maximum estimation latency of 4.33 µs. When processing a set of 10 real traffic traces of up to 123 million packets, the accelerator estimates 1000 packet-size quantiles per trace with a median error of 0.39% or less, and a maximum error of 1.3% or less across all traces. Carolina Gallardo-Pavesi, Yaime Fernández, Javier E. Soto, Cecilia Hernández, Miguel E. Figueroa |
DSD | 5 |
| 2024 | Machine learning controller for data rate management in science DMZ networks
Christian Vega Caicedo, Elie F. Kfoury, Jorge E. Pezoa, Miguel E. Figueroa, Jorge Crichigno |
Comput. Networks | 5 |
| 2023 | A Sketch-Based Algorithm for Network-Flow Entropy Estimation on Programmable Switches Using P4abstractThe empirical Shannon entropy is a popular metric for anomaly detection in network traffic. However, computing its exact value in real time requires fast access to a large number of counters, which is unfeasible in high-speed networks. Approximate approaches using sketches can estimate the entropy with low memory usage. However, achieving good estimation accuracy still requires large data structures, making their implementation difficult in dedicated hardware and programable switches. In this paper, we present an entropy-estimation algorithm and its implementation in a programmable switch, which achieves good accuracy for large traffic traces with low memory usage. The algorithm uses sketches to track the packet count of only the most-frequent flows and models the rest of the traffic with a uniform distribution. The implementation operates within the restrictions imposed by the P4 switch programming language, achieving a 1.72% average estimation error on 12 real-world large traces from public repositories. Javier E. Soto, Sofía Vera, Yaime Fernández, Daniel Yunge, Cecilia Hernández, Miguel E. Figueroa |
DSD | 6 |
| 2023 | A streaming algorithm and hardware accelerator to estimate the empirical entropy of network flows
Yaime Fernández, Javier E. Soto, Sofía Vera, Yasmany Prieto, Cecilia Hernández, Miguel E. Figueroa |
Comput. Networks | 6 |
| 2023 | JACC-FPGA: A hardware accelerator for Jaccard similarity estimation using FPGAs in the cloud
Javier E. Soto, Cecilia Hernández, Miguel E. Figueroa |
Future Gener. Comput. Syst. | 3 |
| 2020 | A hardware architecture for Multiscale Retinex with Chromacity Preservation on an FPGAabstractImage-processing algorithms based on Retinex theory aim to model human color perception to enhance images with low contrast or poor illumination. In particular, the Multiscale Retinex with Chromacity Preservation (MSRCP) algorithm improves on the original Retinex by processing the image at multiple scales and adding a color balance step in postprocessing. Despite their advantages, multiscale Retinex algorithms are computationally intensive, and real-time video processing is not generally possible with general-purpose processor architectures. In this paper, we present a special-purpose hardware accelerator for the MSRCP algorithm. The accelerator introduces tradeoffs to the original formulation of MSRCP by reducing the magnitude of the scales and using a cumulative histogram in the color-balance stage. Despite these modifications, we show that the accelerator produces images that are visually almost identical to a software implementation of the original MSRCP algorithm. We implement our design on a Xilinx XC7A200T-1SBG484C FPGA, which is capable of processing 1280×720-pixel video at up to 94 frames per second, a speedup of 123x compared to a desktop computer running a software version of the algorithm. Jorge Andrés Palacios, Vincenzo Caro, Miguel Durán, Miguel E. Figueroa |
DSD | 4 |
| 2020 | A hardware accelerator for entropy estimation using the top-k most frequent elementsabstractEstimating the empirical entropy of the elements in a dataset is an important task in data analysis. In particular, empirical entropy can be effectively used to detect anomalies in network traffic. However, computing the empirical entropy of a large dataset is computationally expensive and requires a large amount of memory. This is particularly important in high-speed network traffic analysis, where computing the entropy of a data flow in real time requires using hardware accelerators with restricted on-chip memory and arithmetic resources. In this work, we propose a method to estimate the entropy using a streaming algorithm with sublinear space requirements. Our approach uses a sketch to estimate the frequency of the elements in the stream, and a priority queue to store the top-k most frequent elements. We show that our method can provide a good approximation of the entropy of the dataset, and present the design of a hardware accelerator that can compute the entropy of the stream with a throughput of one packet per clock cycle. Implemented on a Xilinx Zynq UltraScale + MPSoC ZCU102 FPGA, our accelerator can operate at line rates above 181 Gbps, consuming 511 mW and using less than 24% of the resources available on the device. Javier E. Soto, Paulo Ubisse, Cecilia Hernández, Miguel E. Figueroa |
DSD | 4 |
| 2019 | A Hardware Accelerator for Edge Detection in High-Definition Video using Cellular Neural NetworksabstractThis paper presents the architecture of a hardware accelerator for a cellular neural network (CeNN) with an application to real-time edge detection on visible-range and infrared video. The accelerator features fully-pipelined processing elements (PEs) that exploit the data parallelism in the algorithm to perform an iteration of the CeNN on a stream of video data with high throughput. The memory architecture exploits the locality of reference in the CeNN, so that each PE uses only 5 line buffers to store pixel, state, and output data, thus achieving low on-chip memory utilization. Implemented on a Xilinx XC7A200T FPGA running at 245MHz, the accelerator performs edge detection on 1080p video using a single CeNN iteration with a throughput of 118 frames per second (fps), a total latency of 15.7us, and 618mW of power consumption. The architecture features static reconfiguration to store built-in kernels and to add more PEs to support multiple iterations of the CeNN algorithm. More kernels can be added dynamically through a serial interface. Ignacio Pérez, Wladimir E. Valenzuela, Miguel E. Figueroa |
DSD | 3 |
| 2019 | Hardware Acceleration of k-Mer Clustering using Locality-Sensitive HashingabstractClustering is an essential operation in many data analysis applications. In particular, bioinformatics and genome analysis use clustering to group similar components in sequence data, in order to find important patterns such as DNA motifs. In this paper, we present an algorithm that clusters DNA data using locality-sensitive hashing with MinHash to group similar subsequences in large Chip-seq datasets. Tested on a standard mESC dataset, the algorithm builds clusters that contain subsequences with high-score matches to known DNA motifs. We also describe the architecture and implementation of a hardware accelerator on a Xilinx Kintex-7 XC7K325T FPGA, that exploits the parallelism of the algorithm to cluster data with a throughput of one k-mer per clock cycle at 350MHz. The accelerator achieves a speedup of 91 compared to a parallel software implementation of the algorithm on a 24-core server. Javier E. Soto, Thomas Krohmer, Cecilia Hernández, Miguel E. Figueroa |
DSD | 4 |
| 2018 | Multimodal Image Registration between SWIR and LWIR Images in an Embedded SystemabstractWe present an algorithm for multimodal registration between short-wave and long-wave infrared images. We use a histogram of oriented gradients to extract features in each image, the Chi-square distance to match the features, a projective transformation to map the objective image onto the reference system, and bilinear interpolation to obtain the pixel values. We designed a heterogeneous embedded system that combines a custom hardware accelerator to perform coordinate transformation and pixel value interpolation, and a programmable processor core to perform feature extraction, feature association, and to compute the transformation parameters. We implemented our design on a Xilinx Zynq XC7Z020 system-on-a-chip, which uses 2.525W of power, 30% of the logic resources of the chip, and 60% of the available on-chip memory. The system runs at 66.6MHz, which allows us to process 640x512-pixel images at more than 60 frames per second after the initial calibration to obtain the transformation parameters. Javier Cardenas, Miguel E. Figueroa |
DSD | 2 |
| 2018 | Heavy-Hitter Detection Using a Hardware Sketch with the Countmin-CU AlgorithmabstractWe present a custom hardware architecture for fast heavy hitter detection in large data streams. The architecture probabilistically estimates the frequency of each element in the data stream using the Countmin-CU sketch with the H3 family of hash functions. The sketch is stored in on-chip memory, and the architecture exploits the parallelism available in the data by simultaneously processing each row of the sketch. The hash functions map each element to a set of counters on the sketch, and the sketch increments the counters that hold the minimum value, which corresponds to the estimated frequency of the element. The hash functions and sorting network are implemented in hardware as fully pipelined circuits, in order to maximize their operating clock frequency. We show a prototype of the architecture running on a Xilinx Kintex-7 XC7K325T FPGA operating with a 300MHz clock, which can process a stream of 3,982,496 32-bit elements and detect the heavy hitters with an 4x16,384-element sketch in 13.27ms, achieving a speedup of 768 compared to a modern desktop computer. Antonio Saavedra, Cecilia Hernández, Miguel E. Figueroa |
DSD | 3 |
| 2016 | Embedded Multimodal Registration of Visible Images on Long-Wave Infrared Video in Real TimeabstractWe describe an embedded system architecture that implements a real-time multimodal registration method which enables multicamera spatio-temporal feature extraction from the combination of visible and long-wave infrared image sequence. Image registration is performed by matching common features between each frame of a visible image to each frame of an infrared image sequence, in order to estimate an affine transformation between each pair of images. The parameters of this affine transformation are estimated recursively on line with the video, thus enabling image registration in real time. The registration algorithm is implemented using a combination of embedded software and dedicated hardware units on a heterogeneous reconfigurable system-on-a-chip. The hardware performs feature detection and extraction, whereas the software estimates the transformation parameters and maps each infrared video frame onto the visible image coordinates. Our prototype implementation runs with a 135MHz clock, consumes 1.8W and utilizes 29% and 54% of the configurable logic and hardware multiplier units available on the chip, respectively. Fabian Inostroza, Javier Cardenas, Sebastián E. Godoy, Miguel E. Figueroa |
DSD | 4 |
| 2016 | An Embedded Hardware Architecture for Real-Time Super-Resolution in Infrared CamerasabstractThis paper presents a custom hardware architecture for fast, low-power super-resolution in infrared images. The architecture performs multiple-frame registration using the Lukas-Kanade optical flow algorithm with Gaussian pyramids, and uses a custom-designed reconstruction algorithm that generates a high-resolution image from multiple consecutive video frames with comparatively low computational cost. We show a prototype of the architecture running on a Xilinx Spartan-6 LX45 FPGA, which can scale a 160 × 120-pixel region of interest to 640 × 480 pixels using information from eight consecutive video frames. The circuit interfaces directly to the digital output of a FLIR Tau 2 infrared camera core and can process video at up to 150 frames per second while consuming 776mW of power. Rodolfo Redlich, Luis Araneda, Antonio Saavedra, Miguel E. Figueroa |
DSD | 4 |
| 2016 | Integrating Dynamic-TDMA Communication Channels into COTS Ethernet NetworksabstractReal-time Ethernet (RTE) is widely recognized for its potential to provide a unified communication backbone for next-generation heterogeneous distributed systems. However, most of the existing research in RTE technologies has traditionally focused on formal models and theoretical analyzes of timing properties, usually omitting the associated implementation challenges for testing them in practice. This gap between theory and practice prevents experimental validation of the claimed properties, which in turn hinders the pace of innovation and adoption of the technology in industrial settings. This paper aims at narrowing the theory-practice gap by characterizing a comprehensive open-source RTE framework that explores emerging challenges in real-time networking, including the provision of ultra-low latency and jitter, dynamic bandwidth management, and segmentation within large networks. This work integrates research on formal abstractions for dynamic time-division multiple access arbitration and technological insights from modern hardware infrastructure, and uses a representative distributed video processing application to provide reproducible evidence of the achieved properties in multihop Ethernet settings. By leveraging readily available technology and an open-source design, the proposed framework facilitates further exploration and experimental validation of properties that are beyond the scope of current commercial technologies, encouraging evidence-based discussions to accelerate development and adoption of new standards for next-generation industrial networks. Gonzalo Carvajal, Luis Araneda, Alejandro Wolf, Miguel E. Figueroa, Sebastian Fischmeister |
IEEE Trans. Ind. Informatics | 4 |
| 2014 | Real-Time Digital Video Stabilization on an FPGAabstractWe present a hardware architecture for real-time digital video stabilization in high-performance embedded systems. The stabilization algorithm analyzes the current and past video frames and obtains a motion estimation vector, which is then filtered to isolate unwanted camera movements from intentional panning. The vector is then used to correct the output video frame. We designed a hardware architecture for motion estimation, filtering and correction and implemented it on a Xilinx Spartan-6 LX45 Field Programmable Gate Array (FPGA). Running on the 640x480-pixel video output of an infrared camera, the circuit successfully compensates involuntary camera motion at a maximum throughput of 104.15 frames per second and dissipates 24.16mW of power with a 100MHz clock. Luis Araneda, Miguel E. Figueroa |
DSD | 2 |
| 2014 | Model, analysis, and evaluation of the effects of analog VLSI arithmetic on linear subspace-based image recognition
Gonzalo Carvajal, Miguel E. Figueroa |
Neural Networks | 2 |
| 2013 | Atacama: An Open FPGA-Based Platform for Mixed-Criticality Communication in Multi-segmented Ethernet NetworksabstractEthernet is widely recognized as an attractive networking technology for modern distributed real-time systems. However, standard Ethernet components require specific modifications and hardware support to provide strict latency guarantees necessary for safety-critical applications. Although this is a well-stated fact, the design of hardware components for real-time communication remains mostly unexplored. This becomes evident from the few solutions reporting prototypes and experimental validation, which hinders the consolidation of Ethernet in real-world distributed applications. This paper presents Atacama, the first open-source framework based on reconfigurable hardware for mixed-criticality communication in multi-segmented Ethernet networks. Atacama uses specialized modules for time-triggered communication of real-time data, which seamlessly integrate with a standard infrastructure using regular best-effort traffic. Atacama enables low and highly predictable communication latency on multi-segmented 1Gbps networks, easy optimization of devices for specific application scenarios, and rapid prototyping of new protocol characteristics. Researchers can use the open-source design to verify our results and build upon the framework, which aims to accelerate the development, validation, and adoption of Ethernet-based solutions in real-time applications. Gonzalo Carvajal, Miguel E. Figueroa, Robert Trausmuth, Sebastian Fischmeister |
FCCM | 2 |
| 2013 | A digital architecture for real-time nonuniformity correction of infrared focal-plane arraysabstractWe present a custom digital architecture that implements the Constant Range algorithm for nonuniformity correction of infrared focal plane arrays. This scene-based technique uses the statistics of the acquired video stream to compensate the gain and offset nonuniformity of the infrared imager online. We evaluate the performance of our architecture while processing raw infrared video at 60 frames per second, with a resolution of 640×480 14-bit pixels. Our implementation of the Constant Range algorithm on a Xilinx Spartan-6 LX45 FPGA uses 1.8% of logic and 26% of arithmetic resources of the chip, consuming only 29.5mW of power. Rodolfo Redlich, Miguel E. Figueroa |
FPL | 2 |
| 2013 | FPGA v/s DSP Performance Comparison for a VSC-Based STATCOM Control ApplicationabstractDigital signal processors (DSPs) and field-programmable gate arrays (FPGAs) are predominant in the implementation of digital controllers and/or modulators for power converter applications. This paper presents a systematic comparison between these two technologies, depicting the main advantages and drawbacks of each one. Key programming and implementation aspects are addressed in order to give an overall idea of their most important features and allow the comparison between DSP and FPGA devices. A classical linear control strategy for a well-known voltage-source-converter (VSC)-based topology used as Static Compensator (STATCOM) is considered as a driving example to evaluate the performance of both approaches. A proof-of-concept laboratory prototype is separately controlled with the TMS320F2812 DSP and the Spartan-3 XCS1000 FPGA to illustrate the characteristics of both technologies. In the case of the DSP, a virtual floating-point library is used to accelerate the control routines compared to double precision arithmetic. On the other hand, two approaches are developed for the FPGA implementation, the first one reduces the hardware utilization and the second one reduces the computation time. Even though both boards can successfully control the STATCOM, results show that the FPGA achieves the best computation time thanks to the high degree of parallelism available on the device. Christian Alberto Garcia-Sepulveda, Javier Muñoz 0001, José R. Espinoza, Miguel E. Figueroa, Carlos R. Baier |
IEEE Trans. Ind. Informatics | 4 |
| 2012 | FPGA-based Neural Network for Nonuniformity Correction on Infrared Focal Plane ArraysabstractDespite recent technological advances which improve their performance and reduce their cost, Focal Plane Arrays for infrared imagers suffer from spatial nonuniformity that renders their output unusable unless a suitable correction method is applied. This paper describes an embedded hardware implementation of Scribner's algorithm for online nonuniformity correction. Our implementation on a Xilinx Spartan XC3S1200E FPGA achieves a throughput of more than 130 frames per second on 320x240-pixel IR video, which greatly exceeds real-time requirements. The power consumption of our system is 329mW, which is two orders of magnitude smaller than a software implementation of the algorithm on a traditional processor, and can be greatly reduced with a custom-VLSI implementation of the architecture. Nicolas Celedon, Rodolfo Redlich, Miguel E. Figueroa |
DSD | 3 |
| 2011 | An FPGA-based real-time nonuniformity correction system for Infrared Focal Plane ArraysabstractSpatial and temporal nonuniformity in Infrared Focal Plane Arrays (IRFPA) severely degrades the quality of images obtained from modern infrared cameras. An efficient implementation of a nonuniformity correction algorithm is therefore necessary in real-time thermal-image visualization systems. This paper presents an FPGA-based implementation of the scene-based Constant Range algorithm for adaptive nonuniformity correction. The system processes an NTSC infrarred video signal at 30fps in real time and consumes only 157 mW of power. The performance of our system is currently limited by the input video frame rate and the external memory bandwidth, but can be readily scaled to a frame rate of more than 250fps. Rodolfo Redlich, Gonzalo Carvajal, Miguel E. Figueroa |
ASAP | 3 |
| 2011 | Analysis and Compensation of the Effects of Analog VLSI Arithmetic on the LMS AlgorithmabstractAnalog very large scale integration implementations of neural networks can compute using a fraction of the size and power required by their digital counterparts. However, intrinsic limitations of analog hardware, such as device mismatch, charge leakage, and noise, reduce the accuracy of analog arithmetic circuits, degrading the performance of large-scale adaptive systems. In this paper, we present a detailed mathematical analysis that relates different parameters of the hardware limitations to specific effects on the convergence properties of linear perceptrons trained with the least-mean-square (LMS) algorithm. Using this analysis, we derive design guidelines and introduce simple on-chip calibration techniques to improve the accuracy of analog neural networks with a small cost in die area and power dissipation. We validate our analysis by evaluating the performance of a mixed-signal complementary metal-oxide-semiconductor implementation of a 32-input perceptron trained with LMS. Gonzalo Carvajal, Miguel E. Figueroa, Daniel G. Sbarbaro-Hofer, Waldo Valenzuela |
IEEE Trans. Neural Networks | 2 |
| 2010 | An FPGA-Based Accelerator for Analog VLSI Artificial Neural Network EmulationabstractAnalog VLSI circuits are being used successfully to implement Artificial Neural Networks (ANNs). These analog circuits exhibit nonlinear transfer function characteristics and suffer from device mismatches, degrading network performance. Because of the high cost involved with analog VLSI production, it is beneficial to predict implementation performance during design. We present an FPGA-based accelerator for the emulation of large (500+ synapses, 10k+ test samples) single-neuron ANNs implemented in analog VLSI. We used hardware time-multiplexing to scale network size and maximize hardware usage. An on-chip CPU controls the data flow through various memory systems to allow for large test sequences. We show that Block-RAM availability is the main implementation bottleneck and that a trade-off arises between emulation speed and hardware resources. However, we can emulate large amounts of synapses on an FPGA with limited resources. We have obtained a speedup of 30.5 times with respect to an optimized software implementation on a desktop computer. Barend van Liempd, Daniel Herrera, Miguel E. Figueroa |
DSD | 3 |
| 2010 | Modern development methods and tools for embedded reconfigurable systems: A survey
Lech Józwiak, Nadia Nedjah, Miguel E. Figueroa |
Integr. | 3 |
| 2009 | Image Recognition in Analog VLSI with On-Chip Learning
Gonzalo Carvajal, Waldo Valenzuela, Miguel E. Figueroa |
ICANN (1) | 3 |
| 2008 | Blind Source-Separation in Mixed-Signal VLSI Using the InfoMax Algorithm
Waldo Valenzuela, Gonzalo Carvajal, Miguel E. Figueroa |
ICANN (2) | 3 |
| 2007 | Subspace-Based Face Recognition in Analog VLSIabstractWe describe an analog-VLSI neural network for face recognition based on subspace methods. The system uses a dimensionality-reduction network whose coefficients can be either programmed or learned on-chip to per- form PCA, or programmed to perform LDA. A second network with user- programmed coefficients performs classification with Manhattan distances. The system uses on-chip compensation techniques to reduce the effects of device mismatch. Using the ORL database with 12x12-pixel images, our circuit achieves up to 85% classification performance (98% of an equivalent software implementation). Gonzalo Carvajal, Waldo Valenzuela, Miguel E. Figueroa |
NIPS | 3 |
| 2006 | Effects of Analog-VLSI Hardware on the Performance of the LMS Algorithm
Gonzalo Carvajal, Miguel E. Figueroa, Seth Bridges |
ICANN (1) | 2 |
| 2004 | On-Chip Compensation of Device-Mismatch Effects in Analog VLSI Neural NetworksabstractDevice mismatch in VLSI degrades the accuracy of analog arithmetic circuits and lowers the learning performance of large-scale neural net- works implemented in this technology. We show compact, low-power on-chip calibration techniques that compensate for device mismatch. Our techniques enable large-scale analog VLSI neural networks with learn- ing performance on the order of 10 bits. We demonstrate our techniques on a 64-synapse linear perceptron learning with the Least-Mean-Squares (LMS) algorithm, and fabricated in a 0.35m CMOS process. Miguel E. Figueroa, Seth Bridges, Chris Diorio |
NIPS | 1 |
| 2002 | Field-Programmable Learning ArraysabstractThis paper introduces the Field-Programmable Learning Array, a new paradigm for rapid prototyping of learning primitives and machine- learning algorithms in silicon. The FPLA is a mixed-signal counterpart to the all-digital Field-Programmable Gate Array in that it enables rapid prototyping of algorithms in hardware. Unlike the FPGA, the FPLA is targeted directly for machine learning by providing local, parallel, on- line analog learning using floating-gate MOS synapse transistors. We present a prototype FPLA chip comprising an array of reconfigurable computational blocks and local interconnect. We demonstrate the via- bility of this architecture by mapping several learning circuits onto the prototype chip. Seth Bridges, Miguel E. Figueroa, David Hsu, Chris Diorio |
NIPS | 2 |
| 2002 | Adaptive Quantization and Density Estimation in SiliconabstractWe present the bump mixture model, a statistical model for analog data where the probabilistic semantics, inference, and learning rules derive from low-level transistor behavior. The bump mixture model relies on translinear circuits to perform probabilistic infer- ence, and floating-gate devices to perform adaptation. This system is low power, asynchronous, and fully parallel, and supports vari- ous on-chip learning algorithms. In addition, the mixture model can perform several tasks such as probability estimation, vector quanti- zation, classification, and clustering. We tested a fabricated system on clustering, quantization, and classification of handwritten digits and show performance comparable to the E-M algorithm on mix- tures of Gaussians. David Hsu, Seth Bridges, Miguel E. Figueroa, Chris Diorio |
NIPS | 3 |
| 2002 | Adaptive CMOS: from biological inspiration to systems-on-a-chipabstractLocal long-term adaptation is a well-known feature of the synaptic junctions in nerve tissue. Neuroscientists have demonstrated that biology uses local adaptation both to tune the performance of neural circuits and for long-term learning. Many researchers believe it is key to the intelligent behavior and the efficiency of biological organizms. Although engineers use adaptation in feedback circuits and in software neural networks, they do not use local adaptation in integrated circuits to the same extent that biology does in nerve tissue. A primary reason is that locally adaptive circuits have proved difficult to implement in silicon. We describe complementary metal-oxide-semiconductor (CMOS) devices called synapse transistors that facilitate local long-term adaptation in silicon. We show that synapse transistors enable self-tuning analog circuits in digital CMOS, facilitating mixed-signal systems-on-a-chip. We also show that synapse transistors enable silicon circuits that learn autonomously, promising sophisticated learning algorithms in CMOS. Chris Diorio, David Hsu, Miguel E. Figueroa |
Proc. IEEE | 3 |
| 2002 | Prolog to adaptive CMOS: from biological inspiration to systems-on-a-chip
Chris Diorio, David Hsu, Miguel E. Figueroa, Richard O'Donnell |
Proc. IEEE | 3 |
| 2002 | Competitive learning with floating-gate circuitsabstractCompetitive learning is a general technique for training clustering and classification networks. We have developed an 11-transistor silicon circuit, that we term an automaximizing bump circuit, that uses silicon physics to naturally implement a similarity computation, local adaptation, simultaneous adaptation and computation and nonvolatile storage. This circuit is an ideal building block for constructing competitive-learning networks. We illustrate the adaptive nature of the automaximizing bump in two ways. First, we demonstrate a silicon competitive-learning circuit that clusters one-dimensional (1-D) data. We then illustrate a general architecture based on the automaximizing bump circuit; we show the effectiveness of this architecture, via software simulation, on a general clustering task. We corroborate our analysis with experimental data from circuits fabricated in a 0.35-mum CMOS process. David Hsu, Miguel E. Figueroa, Chris Diorio |
IEEE Trans. Neural Networks | 2 |
| 2000 | An FPGA-Based Array Processor for an Ionospheric-Imaging RadarabstractAtmospheric scientists need to observe fluctuations in the ionosphere, both to probe the underlying atmospheric physics and to remove the effects of these fluctuations from other measurements. We have built an FPGA-based, pipelined array processor that allows us to make these observations in real-time, using passive radar techniques. Our array processor time-multiplexes 16 multiply-accumulators across 1536 radar ranges, performing a pipelined correlation and integration of the radar signal for each range. A DSP-based postprocessor generates real-time range-Doppler profiles of the ionospheric targets. Tim Tuan, Miguel E. Figueroa, Frank D. Lind, Chucai Zhou, Chris Diorio, John D. Sahr |
FCCM | 2 |
| 2000 | A Silicon Primitive for Competitive LearningabstractCompetitive learning is a technique for training classification and clustering networks. We have designed and fabricated an 11- transistor primitive, that we term an automaximizing bump circuit, that implements competitive learning dynamics. The circuit per(cid:173) forms a similarity computation, affords nonvolatile storage, and implements simultaneous local adaptation and computation. We show that our primitive is suitable for implementing competitive learning in VLSI, and demonstrate its effectiveness in a standard clustering task. David Hsu, Miguel E. Figueroa, Chris Diorio |
NIPS | 2 |
| 1991 | A decoupled access/execute processor for matrix algorithms: architecture and programmingabstractThe authors describe a processor for the execution of a class of matrix algorithms according to the multimesh graph (MMG) mapping method, which is suitable as the processing cell in an application-specific array. The processor uses the decoupled access-execute model of computation, so that it consists of two programmable units: a processing unit (PU) and an access unit (AU). The two programs synchronize their execution through queues. The instruction set includes single-instruction loops with no overhead, and block- loops with just one extra instruction. All storage modules are accessed as FIFO queues, without the need for addressing mechanisms. The efficiency of the resulting code is high: for a class of matrix algorithms frequently used in signal processing applications, about 90% of the instructions executed correspond to arithmetic operations.> Jaime H. Moreno, Miguel E. Figueroa |
ASAP | 2 |