EDBT 2026 Demo / reviewers in the wild / expert
Andre Guntoro
dblp:95/6393 · also Andre Teguh Guntoro
· DBLP profile ↗
28ranked-venue papers
5as first author
10since 2021 · last 2024
0000-0003-4144-0283ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 2 since 2021Security and privacy · 2 · 1 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting Networks?abstractIn the design of traffic monitoring solutions for optimizing the urban mobility infrastructure, acoustic vehicle counting models have received attention due to their cost effectiveness and energy efficiency. Although deep learning has proven effective for visual traffic monitoring, its use has not been thoroughly investigated in the audio domain, likely due to real-world data scarcity. In this work, we propose a novel approach to acoustic vehicle counting by developing: i) a traffic noise simulation framework to synthesize realistic vehicle pass-by events; ii) a strategy to mix synthetic and real data to train a deep-learning model for traffic counting. The proposed system is capable of simultaneously counting cars and commercial vehicles driving on a two-lane road, and identifying their direction of travel under moderate traffic density conditions. With only 24 hours of labeled real-world traffic noise, we are able to improve counting accuracy on real-world data from 63% to 88% for cars and from 86% to 94% for commercial vehicles. Stefano Damiano, Luca Bondi, Shabnam Ghaffarzadegan, Andre Guntoro, Toon van Waterschoot |
ICASSP | 4 |
| 2023 | Exploiting Subword Permutations to Maximize CNN Compute Performance and EfficiencyabstractNeural networks (NNs) are quantized to decrease their computational demands and reduce their memory foot-print. However, specialized hardware is required that supports computations with low bit widths to take advantage of such optimizations. In this work, we propose permutations on subword level that build on top of multi-bit-width multiply-accumulate operations to effectively support low bit width computations of quantized NNs. By applying this technique, we extend the data reuse and further improve compute performance for convolution operations compared to simple vectorization using SIMD (single-instruction-multiple-data). We perform a design space exploration using a cycle accurate simulation with MobileNet and VGG16 on a vector-based processor. The results show a speedup of up to$3.7\times$and a reduction of up to$1.9\times$for required data transfers. Additionally, the control overhead for orchestrating the computation is decreased by up to$3.9\times$. Michael Beyer, Sven Gesper, Andre Guntoro, Guillermo Payá-Vayá, Holger Blume |
ASAP | 3 |
| 2023 | Real-Time Acoustic Perception for Automotive ApplicationsabstractIn recent years the automotive industry has been strongly promoting the development of smart cars, equipped with multi-modal sensors to gather information about the surroundings, in order to aid human drivers or make autonomous decisions. While the focus has mostly been on visual sensors, also acoustic events are crucial to detect situations that require a change in the driving behavior, such as a car honking, or the sirens of approaching emergency vehicles. In this paper, we summarize the results achieved so far in the Marie Sklodowska-Curie Actions (MSCA) Eruopean Industrial Doctorates (EID) project “Intelligent Ultra Low-Power Signal Processing for Automotive (I-SPOT)”. On the algorithmic side, the I-SPOT Project aims to enable detecting, localizing and tracking environmental audio signals by jointly developing microphone array processing and deep learning techniques that specifically target automotive applications. Data generation software has been developed to cover the I-SPOT target scenarios and research challenges. This tool is currently being used to develop low-complexity deep learning techniques for emergency sound detection. On the hardware side, the goal impels workflows for hardware-algorithm co-design to ease the generation of architectures that are sufficiently flexible towards algorithmic evolutions without giving up on efficiency, as well as enable rapid feedback of hardware implications of algorithmic decision. This is pursued though a hierarchical workflow that breaks the hardware-algorithm design space into reasonable subsets, which has been tested for operator-level optimizations on state-of-the-art robust sound source localization for edge devices. Further, several open challenges towards an end-to-end system are clarified for the next stage of I-SPOT. Jun Yin 0001, Stefano Damiano, Marian Verhelst, Toon van Waterschoot, Andre Guntoro |
DATE | 5 |
| 2023 | ACCO: Automated Causal CNN Scheduling Optimizer for Real-Time Edge AcceleratorsabstractSpatio-Temporal Convolutional Neural Networks (ST-CNN) allow extending CNN capabilities from image processing to consecutive temporal-pattern recognition. Generally, state-of-the-art (SotA) ST-CNNs inflate the feature maps and weights from well-known CNN backbones to represent the additional time dimension. However, edge computing applications would suffer tremendously from such large computation/memory overhead. Fortunately, the overlapping nature of ST-CNN enables various optimizations, such as the dilated causal convolution structure and Depth-First (DF) layer fusion to reuse the computation between time steps and CNN sliding windows, respectively. Yet, no hardware-aware approach has been proposed that jointly explores the optimal strategy from a scheduling as well as a hardware point of view.To this end, we present ACCO, an automated optimizer that explores efficient Causal CNN transformation and DF scheduling for ST-CNNs on edge hardware accelerators. By cost-modeling the computation and data movement on the accelerator architecture, ACCO automatically selects the best scheduling strategy for the given hardware-algorithm target. Compared to the fixed dilated causal structure, ST-CNNs with ACCO reach an ~8.4× better Energy-Delay-Product. Meanwhile, ACCO improves ~20% in layer-fusion optimals compared to the SotA DF exploration toolchain. When jointly optimizing ST-CNN on the temporal and spatial dimension, ACCO’s scheduling outcomes are on average 19× faster and 37× more energy-efficient than spatial DF schemes. Jun Yin 0001, Linyan Mei, Andre Guntoro, Marian Verhelst |
ICCD | 3 |
| 2023 | Online Quantization Adaptation for Fault-Tolerant Neural Network Inference
Michael Beyer, Jan Micha Borrmann, Andre Guntoro, Holger Blume |
SAFECOMP | 3 |
| 2022 | FELIX: A Ferroelectric FET Based Low Power Mixed-Signal In-Memory Architecture for DNN AccelerationabstractToday, a large number of applications depend on deep neural networks (DNN) to process data and perform complicated tasks at restricted power and latency specifications. Therefore, processing-in-memory (PIM) platforms are actively explored as a promising approach to improve the throughput and the energy efficiency of DNN computing systems. Several PIM architectures adopt resistive non-volatile memories as their main unit to build crossbar-based accelerators for DNN inference. However, these structures suffer from several drawbacks such as reliability, low accuracy, large ADCs/DACs power consumption and area, high write energy, and so on. In this article, we present a new mixed-signal in-memory architecture based on the bit-decomposition of the multiply and accumulate (MAC) operations. Our in-memory inference architecture uses a single FeFET as a non-volatile memory cell. Compared to the prior work, this system architecture provides a high level of parallelism while using only 3-bit ADCs. Also, it eliminates the need for any DAC. In addition, we provide flexibility and a very high utilization efficiency even for varying tasks and loads. Simulations demonstrate that we outperform state-of-the-art efficiencies with 36.5 TOPS/W and can pack 2.05 TOPS with 8-bit activation and 4-bit weight precision in an area of 4.9 mm 2 using 22 nm FDSOI technology. Employing binary operation, we obtain 1169 TOPS/W and over 261 TOPS/W/mm 2 on system level. Taha Soliman, Nellie Laleni, Tobias Kirchner, Franz Müller 0001, Thomas Kämpfe, Andre Guntoro, Norbert Wehn |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2021 | Efficient Accuracy Recovery in Approximate Neural Networks by Systematic Error ModellingabstractApproximate Computing is a promising paradigm for mitigating the computational demands of Deep Neural Networks (DNNs), by leveraging DNN performance and area, throughput or power. The DNN accuracy, affected by such approximations, can be then effectively improved through retraining. In this paper, we present a novel methodology for modelling the approximation error introduced by approximate hardware in DNNs, which accelerates retraining and achieves negligible accuracy loss. To this end, we implement the behavioral simulation of several approximate multipliers and model the error generated by such approximations on pre-trained DNNs for image classification on CIFAR10 and ImageNet. Finally, we optimize the DNN parameters by applying our error model during DNN retraining, to recover the accuracy lost due to approximations. Experimental results demonstrate the efficiency of our proposed method for accelerated retraining (11 x faster for CIFAR10 and 8x faster for ImageNet) for full DNN approximation, which allows us to deploy approximate multipliers with energy savings of up to 36% for 8-bit precision DNNs with an accuracy loss lower than 1%. Cecilia De la Parra, Andre Guntoro, Akash Kumar 0001 |
ASP-DAC | 2 |
| 2021 | A Novel DRAM-Based Process-in-Memory Architecture and its Implementation for CNNsabstractProcessing-in-Memory (PIM) is an emerging approach to bridge the memory-computation gap. One of the key challenges of PIM architectures in the scope of neural network inference is the deployment of traditional area-intensive arithmetic multipliers in memory technology, especially for DRAM-based PIM architectures. Hence, existing DRAM PIM architectures are either confined to binary networks or exploit the analog property of the sub-array bitlines to perform bulk bit-wise logic operations. The former reduces the accuracy of predictions, i.e. Quality-of-results, while the latter increases overall latency and power consumption. Chirag Sudarshan, Taha Soliman, Cecilia De la Parra, Christian Weis, Leonardo Ecco, Matthias Jung 0001, Norbert Wehn, Andre Guntoro |
ASP-DAC | 8 |
| 2021 | Knowledge Distillation and Gradient Estimation for Active Error Compensation in Approximate Neural NetworksabstractApproximate computing is a promising approach for optimizing computational resources of error-resilient applications such as Convolutional Neural Networks (CNNs). However, such approximations introduce an error that needs to be compensated by optimization methods, which typically include a retraining or fine-tuning stage. To efficiently recover from the introduced error, this fine-tuning process needs to be adapted to take CNN approximations into consideration. In this work, we present a novel methodology for fine-tuning approximate CNNs with ultralow bit-width quantization and large approximation error, which combines knowledge distillation and gradient estimation to recover the lost accuracy due to approximations. With our proposed methodology, we demonstrate energy savings of up to 38% in complex approximate CNNs with weights quantized to 4 bits and 8-bit activations, with less than 3% accuracy loss w.r.t. the full precision model. Cecilia De la Parra, Xuyi Wu, Andre Guntoro, Akash Kumar 0001 |
DATE | 3 |
| 2021 | Exploiting Resiliency for Kernel-Wise CNN Approximation Enabled by Adaptive Hardware DesignabstractEfficient low-power accelerators for Convolutional Neural Networks (CNNs) largely benefit from quantization and approximation, which are typically applied layer-wise for efficient hardware implementation. In this work, we present a novel strategy for efficient combination of these concepts at a deeper level, which is at each channel or kernel. We first apply layer-wise, low bit-width, linear quantization and truncation-based approximate multipliers to the CNN computation. Then, based on a state-of-the-art resiliency analysis, we are able to apply a kernel-wise approximation and quantization scheme with negligible accuracy losses, without further retraining. Our proposed strategy is implemented in a specialized framework for fast design space exploration. This optimization leads to a boost in estimated power savings of up to 34% in residual CNN architectures for image classification, compared to the base quantized architecture. Cecilia De la Parra, Ahmed El-Yamany, Taha Soliman, Akash Kumar 0001, Norbert Wehn, Andre Guntoro |
ISCAS | 6 |
| 2020 | Efficient FeFET Crossbar Accelerator for Binary Neural NetworksabstractThis paper presents a novel ferroelectric field-effect transistor (FeFET) in-memory computing architecture dedicated to accelerate Binary Neural Networks (BNNs). We present in-memory convolution, batch normalization and dense layer processing through a grid of small crossbars with reduced unit size, which enables multiple bit operation and value accumulation. Additionally, we explore the possible operations parallelization for maximized computational performance. Simulation results show that our new architecture achieves a computing performance up to 2.46 TOPS while achieving a high power efficiency reaching 111.8 TOPS/Watt and an area of 0.026 mm2in 22nm FDSOI technology. Taha Soliman, Ricardo Olivo, Tobias Kirchner, Cecilia De la Parra, Maximilian Lederer, Thomas Kämpfe, Andre Guntoro, Norbert Wehn |
ASAP | 7 |
| 2020 | Next Generation Arithmetic for Edge ComputingabstractArithmetic is a key component and is ubiquitous in today’s digital world, ranging from embedded to high-performance computing systems. With machine learning at the fore in a wide range of application domains from wearables to automotive to avionics to weather prediction, sufficiently accurate yet low-cost arithmetic is the need for the day. Recently, there have been several advances in the domain of computer arithmetic, which includes high-precision anchored numbers from ARM, posit arithmetic, bfloat16, etc. as an alternative to IEEE 754-2008 compliant arithmetic. Optimizations on fixed-point and integer arithmetic are also being pursued actively for low-power computing architectures. Furthermore, approximate computing and transprecision/mixed-precision computing have been exciting areas of research forever. While academic research in the domain of computer arithmetic has a long history, industrial adoption of some of these new data types and techniques is in its early stages and expected to increase in the future. bfloat16 is an excellent example for this. In this paper, we bring academia and industry together to discuss the latest results and future directions for research in the domain of next-generation computer arithmetic, especially for edge computing. Andre Guntoro, Cecilia De la Parra, Farhad Merchant, Florent de Dinechin, John L. Gustafson, Martin Langhammer, Rainer Leupers, Sangeeth Nambiar |
DATE | 1 |
| 2020 | ProxSim: GPU-based Simulation Framework for Cross-Layer Approximate DNN OptimizationabstractThrough cross-layer approximation of Deep Neural Networks (DNN) significant improvements in hardware resources utilization for DNN applications can be achieved. This comes at the cost of accuracy degradation, which can be compensated through different optimization methods. However, DNN optimization is highly time-consuming in existing simulation frameworks for cross-layer DNN approximation, as they are usually implemented for CPU usage only. Specially for large-scale image processing tasks, the need of a more efficient simulation framework is evident. In this paper we present ProxSim, a specialized, GPU-accelerated simulation framework for approximate hardware, based on Tensorflow, which supports approximate DNN inference and retraining. Additionally, we propose a novel hardware-aware regularization technique for approximate DNN optimization. By using ProxSim, we report up to 11× savings in execution time, compared to a multi-thread CPU-based framework, and an accuracy recovery of up to 30% for three case studies of image classification with MNIST, CIFAR-10 and ImageNet. Cecilia De la Parra, Andre Guntoro, Akash Kumar 0001 |
DATE | 2 |
| 2020 | Full Approximation of Deep Neural Networks through Efficient OptimizationabstractApproximate Computing is a promising paradigm for mitigating computational requirements of Deep Neural Networks (DNN), by taking advantage of their inherent error resilience. Specifically, the use of approximate multipliers in DNN inference can lead to significant improvements in power consumption of embedded DNN applications. This paper presents a methodology for efficient approximate multiplier selection and for full and uniform approximation of large DNNs, through retraining and minimization of the approximation error. We evaluate our methodology using 422 approximate multipliers from the EvoApprox library, with three different Residual architectures trained with Cifar10, and achieve energy savings of up to 18% surpassing the original floating-point accuracy, and of up to 58% with an accuracy loss of 0.73%. Cecilia De la Parra, Andre Guntoro, Akash Kumar 0001 |
ISCAS | 2 |
| 2020 | Improving approximate neural networks for perception tasks through specialized optimization
Cecilia De la Parra, Andre Guntoro, Akash Kumar 0001 |
Future Gener. Comput. Syst. | 2 |
| 2020 | Automated design of error-resilient and hardware-efficient deep neural networks
Christoph Schorn, Thomas Elsken, Sebastian Vogel, Armin Runge, Andre Guntoro, Gerd Ascheid |
Neural Comput. Appl. | 5 |
| 2019 | An Efficient Bit-Flip Resilience Optimization Method for Deep Neural NetworksabstractDeep neural networks usually possess a high overall resilience against errors in their intermediate computations. However, it has been shown that error resilience is generally not homogeneous within a neural network and some neurons might be very sensitive to faults. Even a single bit-flip fault in one of these critical neuron outputs can result in a large degradation of the final network output accuracy, which cannot be tolerated in some safety-critical applications. While critical neuron computations can be protected using error correction techniques, a resilience optimization of the neural network itself is more desirable, since it can reduce the required effort for error correction and fault protection in hardware. In this paper, we develop a novel resilience optimization method for deep neural networks, which builds upon a previously proposed resilience estimation technique. The optimization involves only few steps and can be applied to pre-trained networks. In our experiments, we significantly reduce the worst-case failure rates after a bit-flip fault for deep neural networks trained on the MNIST, CIFAR-10 and ILSVRC classification benchmarks. Christoph Schorn, Andre Guntoro, Gerd Ascheid |
DATE | 2 |
| 2019 | Guaranteed Compression Rate for Activations in CNNs using a Frequency Pruning ApproachabstractConvolutional Neural Networks have become state of the art for many computer vision tasks. However, the size of Neural Networks prevents their application in resource constrained systems. In this work, we present a lossy compression technique for intermediate results of Convolutional Neural Networks. The proposed method offers guaranteed compression rates and additionally adapts to performance requirements. Our experiments with networks for classification and semantic segmentation show, that our method outperforms state-of-the-art compression techniques used in CNN accelerators. Sebastian Vogel, Christoph Schorn, Andre Guntoro, Gerd Ascheid |
DATE | 3 |
| 2019 | Self-Supervised Quantization of Pre-Trained Neural Networks for Multiplierless AccelerationabstractTo host intelligent algorithms such as Deep Neural Networks on embedded devices, it is beneficial to transform the data representation of neural networks into a fixed-point format with reduced bit-width. In this paper we present a novel quantization procedure for parameters and activations of pre-trained neural networks. For 8 bit linear quantization, our procedure achieves close to original network performance without retraining and consequently does not require labeled training data. Additionally, we evaluate our method for power-of-two quantization as well as for a two-hot quantization scheme, enabling shift-based inference. To underline the hardware benefits of a multiplierless accelerator, we propose the design of a shift-based processing element. Sebastian Vogel, Jannik Springer, Andre Guntoro, Gerd Ascheid |
DATE | 3 |
| 2019 | Bit-Shift-Based Accelerator for CNNs with Selectable Accuracy and ThroughputabstractHardware accelerators for compute intensive algorithms such as convolutional neural networks benefit from number representations with reduced precision. In this paper, we evaluate and extend a number representation based on power-of-two quantization enabling bit-shift-based processing of multiplications. We found that weights of a neural network can either be represented by a single 4 bit power-of-two value or with two 4 bit values depending on accuracy requirements. We evaluate the classification accuracy of VGG-16 and ResNet50 on the ImageNet dataset with weights represented in our novel number format. To include a more complex task, we additionally evaluate the format on two networks for semantic segmentation. In addition, we design a novel processing element based on bit-shifts which is configurable in terms of throughput (4 bit mode) and accuracy (8 bit mode). We evaluate this processing element in an FPGA implementation of a dedicated accelerator for neural networks incorporating a 32-by-64 processing array running at 250 MHz with 1 TOp/s peak throughput in 8 bit mode. The accelerator is capable of processing regular convolutional layers and dilated convolutions in combination with pooling and upsampling. For a semantic segmentation network with 108.5 GOp/frame, our FPGA implementation achieves a throughput of 7.0 FPS in the 8 bit accurate mode and upto 11.2 FPS in the 4 bit mode corresponding to 760.1 GOp/s and 1,218 GOp/s effective throughput, respectively. Finally, we compare the novel design to classical multiplier-based approaches in terms of FPGA utilization and power consumption. Our novel multiply-accumulate engines designed for the optimized number representation uses 9 % less logical elements while allowing double throughput compared to a classical implementation. Moreover, a measurement shows 25 % reduction of power consumption at same throughput. Therefore, our flexible design offers a solution to the trade-off between energy efficiency, accuracy, and high throughput. Sebastian Vogel, Rajatha B. Raghunath, Andre Guntoro, Kristof Van Laerhoven, Gerd Ascheid |
DSD | 3 |
| 2019 | Efficient Acceleration of CNNs for Semantic Segmentation on FPGAsabstractWe present a Vector Processing Engine (VPE) designed for the acceleration of Convolutional Neural Networks (CNNs) for semantic segmentation. Most CNN accelerators focus on classification. However, CNNs for semantic segmentation incorporate special layer types. Our accelerator supports not only regular convolutional layers, but also dilated convolutions and convolutions in combination with down- or up-sampling. These features are implemented in dedicated address generators which load the corresponding input vector from an input line buffer. The VPE is designed as a 64x64-array where up to 64 output features and 64 input features of a convolutional layer can be unrolled in parallel. The array has a peak performance of 4.12 TOp/s and achieves 3.85 TOp/s on a CNN for semantic segmentation - resulting in an average utilization of 93 %. The design is prototypically implemented on a Virtex UltraScale+ device with a clock rate of 250 MHz. In addition to the overall architecture, we present the two-hot quantization scheme. A value in two-hot quantization can be regarded as a combination of two power-of-two values. Hence, instead of bulky multipliers, two small bit-shifts are implemented. We design and implement dedicated arithmetic engines for this quantization scheme. Additionally, we evaluate this quantization scheme on the rather complex task of semantic segmentation. We show that the performance of an 8 bit two-hot quantization scheme is marginally lower in comparison to a regular 8 bit fixed-point variant. Sebastian Vogel, Jannik Springer, Andre Guntoro, Gerd Ascheid |
FPGA | 3 |
| 2018 | Accurate neuron resilience prediction for a flexible reliability management in neural network acceleratorsabstractDeep neural networks have become a ubiquitous tool for mastering complex classification tasks. Current research focuses on the development of power-efficient and fast neural network hardware accelerators for mobile and embedded devices. However, when used in safety-critical applications, for example autonomously operating vehicles, the reliability of such accelerators becomes a further optimization criterion which can stand in contrast to power-efficiency and latency. Furthermore, ensuring hardware reliability becomes increasingly challenging for shrinking structure widths and rising power densities in the nanometer semiconductor technology era. One solution to this challenge is the exploitation of fault tolerant parts in deep neural networks. In this paper we propose a new method for predicting the error resilience of neurons in deep neural networks and show that this method significantly improves upon existing methods in terms of accuracy as well as interpretability. We evaluate prediction accuracy by simulating hardware faults in networks trained on the CIFAR-10 and ILSVRC image classification benchmarks and protecting neurons according to the resilience estimations. In addition, we demonstrate how our resilience prediction can be used for a flexible trade-off between reliability and efficiency in neural network hardware accelerators. Christoph Schorn, Andre Guntoro, Gerd Ascheid |
DATE | 2 |
| 2018 | Efficient hardware acceleration of CNNs using logarithmic data representation with arbitrary log-baseabstractEfficient acceleration of Deep Neural Networks is a manifold task. In order to save memory requirements and reduce energy consumption we propose the use of dedicated accelerators with novel arithmetic processing elements which use bit shifts instead of multipliers. While a regular power-of-2 quantization scheme allows for multiplierless computation of multiply-accumulate-operations, it suffers from high accuracy losses in neural networks. Therefore, we evaluate the use of powers-of-arbitrary-log-bases and confirmed their suitability for quantization of pre-trained neural networks. The presented method works without retraining of the neural network and therefore is suitable for applications in which no labeled training data is available. In order to verify our proposed method, we implement the log-based processing elements into a neural network accelerator on an FPGA. The hardware efficiency is evaluated in terms of FPGA utilization and energy requirements in comparison to regular 8-bit-fixed-point multiplier based acceleration. Using this approach hardware resources are minimized and power consumption is reduced by 22.3%. Sebastian Vogel, Mengyu Liang, Andre Guntoro, Walter Stechele, Gerd Ascheid |
ICCAD | 3 |
| 2018 | Efficient On-Line Error Detection and Mitigation for Deep Neural Network Accelerators
Christoph Schorn, Andre Guntoro, Gerd Ascheid |
SAFECOMP | 2 |
| 2009 | A flexible floating-point wavelet transform and wavelet packet processorabstractThe richness of wavelet transformation is known in many fields. There exist different classes of wavelet filters that can be used depending on the application. In this paper, we propose an IEEE 754 floating-point lifting-based wavelet processor that can perform various forward and inverse Discrete Wavelet Transforms (DWTs) and Discrete Wavelet Packets (DWPs). Our architecture is based on processing elements that can perform either prediction or update on a continuous data stream in every two clock cycles. We also consider the normalization step that takes place at the end of the forward DWT/DWP or at the beginning of the inverse DWT/DWP. To cope with different wavelet filters, we feature a multi-context configuration to select among various DWTs/DWPs. Different memory sizes and multi-level transformations are supported. For the 32-bit implementation, the estimated area of the proposed processor with 2times512 words memory and 8 PEs in a 0.18-mum process is 3.7 mm square and the estimated operating speed is 353 MHz. Andre Guntoro, Manfred Glesner |
DATE | 1 |
| 2008 | Novel approach on lifting-based DWT and IDWT processor with multi-context configuration to support different wavelet filtersabstractIn this paper, we propose a lifting-based DWT processor that can perform various forward and inverse transforms. Contrary to other lifting-based processors which focus on JPEG2000, our design is based on the fact that the wavelet transformations are not used only in the area of image processing and wavelet filters may not be represented as integer numbers. The proposed architecture is based on NxM processing elements which require only one multiplier and one adder to perform prediction/update on a continuous data stream. The multi-context feature allows the processor to be configured for different types of transformations in a simple manner. Andre Guntoro, Manfred Glesner |
ASAP | 1 |
| 2008 | A lifting-based DWT and IDWT processor with multi-context configuration and normalization factorabstractIn this paper, we propose a generalized lifting-based wavelet processor that can be configured to compute various forward and inverse DWTs. Our architecture is based on NxM PEs which can perform either prediction or update on a continuous data stream in every two clock cycles. We also consider the normalization step which takes place at the end of the forward DWT or at the beginning of the inverse DWT. To cope with different wavelet filters, we feature a multi-context configuration to select among various DWTs. The proposed processor can be operated at around 186 MHz depending on the data width selection. Andre Guntoro, Manfred Glesner |
FPL | 1 |
| 2008 | High-performance fpga-based floating-point adder with three inputsabstractIn this paper, we present the design and the implementation of an FPGA-based floating-point adder with three inputs. The design is based on a 5-level pipeline stage in order to distribute the critical paths and to maximize the performance. We examine the data dependencies to minimize the number of the pipeline stages and to reduce the resource allocation. Our design is parameterisable in order to cope with different floating-point formats, including the standard IEEE 754 formats and the custom configurations. The proposed design with the single precision, 32-bit floating-point format, can be operated at 143 MHz on Xilinx Virtex2Pro XC2VP30-7. Andre Guntoro, Manfred Glesner |
FPL | 1 |