EDBT 2026 Demo / reviewers in the wild / expert
René van Leuken 0001
dblp:62/5653 · also Rene van Leuken 0001, T. G. R. M. van Leuken
· DBLP profile ↗
27ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0003-0638-7595ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Jumping Shift: A Logarithmic Quantization Method for Low-Power CNN AccelerationabstractLogarithmic quantization for Convolutional Neural Networks (CNN): a) fits well typical weights and activation distributions, and b) allows the replacement of the multiplication operation by a shift operation that can be implemented with fewer hardware resources. We propose a new quantization method named Jumping Log Quantization (JLQ). The key idea of JLQ is to extend the quantization range, by adding a coefficient parameter “s” in the power of two exponents$(2^{sx+i})$. This quantization strategy skips some values from the standard logarithmic quantization. In addition, we also develop a small hardware-friendly optimization called weight de-zero. Zero-valued weights that cannot be performed by a single shift operation are all replaced with logarithmic weights to reduce hardware resources with almost no accuracy loss. To implement the Multiply-And-Accumulate (MAC) operation (needed to compute convolutions) when the weights are JLQ-ed and de-zeroed, a new Processing Element (PE) have been developed. This new PE uses a modified barrel shifter that can efficiently avoid the skipped values. Resource utilization, area, and power consumption of the new PE standing alone are reported. We have found that JLQ performs better than other state-of-the-art logarithmic quantization methods when the bit width of the operands becomes very small. Longxing Jiang, David Aledo, René van Leuken 0001 |
DATE | 3 |
| 2022 | Near-Precise Parameter Approximation for Multiple Multiplications on a Single DSP BlockabstractDSP blocks are one of the efficient solutions to implement multiply-accumulate (MAC) operations on FPGA’s. However, since the DSP blocks have wide multiplier and adder blocks, MAC operations using low bit-length parameters lead to an underutilization. Hence, an efficient approximation technique is introduced. The technique includes manipulation and approximation of the low bit-length parameters based upon a Single DSP - Multiple Multiplication (SDMM) execution. The accuracy of the developed optimization technique was evaluated for different CNN weight bit precisions using the Alexnet and VGG-16 networks and the ImageNet ILSVRC-2012 dataset. The optimization can be implemented without loss of accuracy in almost all cases, while it causes slight accuracy losses in a few cases. Through these optimizations, multiple parameter multiplications are performed in a single DSP block at the cost of a small hardware overhead. As a result of our optimizations, the parameters are represented in a different format on off-chip memory, providing up to 33% compression without any hardware cost. A prototype systolic array architecture was implemented employing our optimizations on a Xilinx Zynq FPGA. It reduced the number of DSP blocks by 66.6%, 75%, and 83.3% for 8, 6, and 4-bit input variables, respectively. Ercan Kalali, René van Leuken 0001 |
IEEE Trans. Computers | 2 |
| 2021 | A Power-Efficient Parameter Quantization Technique for CNN AcceleratorsabstractQuantization techniques are widely used in CNN inference to reduce the cost of hardware at the expense of small accuracy losses. However, after the quantization, there is still a multiplication cost for the fixed-point quantized CNN weights. Therefore, a novel CNN quantization technique is introduced, which can be implemented without using any multiplier. We evaluated our quantization technique using VGG-16 and Alexnet networks, and the Tiny ImageNet dataset. The quantization technique causes 0.39% and 0.98% accuracy losses for the 8-bit CNN weights compared to floating-point implementations of VGG-16 and Alexnet, respectively. After, a fine-tuning method for our quantization is introduced, which further reduces the accuracy loss. The fine-tuning reduced the accuracy losses on 8-bit quantized VGG-16 and Alexnet to 0.24% and 0.39%, respectively. Two different processing element architectures, which do not include any multiplier hardware, are designed to perform multiply-accumulate (MAC) operations of CNN models quantized by our technique. Two different systolic array prototypes are designed employing the two PE architectures to compare with the traditional fixed-point MAC implementation. The systolic array architectures containing our processing element designs reduced the power consumption of the systolic array up to 14.2% and 21.6%. Ercan Kalali, René van Leuken 0001 |
DSD | 2 |
| 2018 | Uncertainty in Noise-Driven Steady-State Neuromorphic Network for ECG Data ClassificationabstractThe pathophysiological processes underlying the ECG tracing demonstrate significant heart rate and the morphological pattern variations, for different or in the same patient at diverse physical/temporal conditions. Within this framework, spiking neural networks (SNN) may be a compelling approach to ECG pattern classification based on the individual characteristics of each patient. In this paper, we study electrophysiological dynamics in the self-organizing map SNN when the coefficients of the neuronal connectivity matrix are random variables. We examine synchronicity and noise-induced information processing, influence of the uncertainty on the system signal-to-noise ratio, and impact on the clustering accuracy of cardiac arrhythmia. Amir Zjajo, Johan Mes, Eralp Kolagasioglu, Sumeet S. Kumar, René van Leuken 0001 |
CBMS | 5 |
| 2017 | Fighting Dark Silicon: Toward Realizing Efficient Thermal-Aware 3-D Stacked MultiprocessorsabstractThis paper investigates the challenges of dark silicon that impede the performance and reliability of 3-D stacked multiprocessors. It presents a multipronged approach toward addressing the thermal issues arising from high-density integration in die stacks, spanning architectural techniques, design methodologies, and runtime temperature management. Importantly, this paper provides novel insights into the causes of hotspot formation in 3-D ICs and details a practical approach toward exploring and mitigating performance-limiting thermal behavior early in the system design flow. Sumeet S. Kumar, Amir Zjajo, René van Leuken 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | A 2.7μW 10b 640kS/s time-based A/D converter for implantable neural recording interfaceabstractIn this paper, we propose a time-based, programmable-gain A/D converter allowing for an easily-scalable, and power-efficient, implantable, biomedical recording system. The converter circuit is realized in a 90 nm CMOS technology, operates at 640 kS/s, occupy an area of 0.022 mm2, and consumes less than 2.7 μW corresponding to a figure of merit of 6.2 fJ/conversion-step. Amir Zjajo, Santosh Astigimath, René van Leuken 0001 |
ISCAS | 3 |
| 2016 | Determining Performance Boundaries on High-Level System SpecificationsabstractWe can significantly reduce the time required to realize designs if it is possible to find limits to the performance of an embedded system, solely based on high-level system specifications. For that purpose, we present in this paper the cprof profiler, which determines the number of clock cycles needed to execute a C-program in hardware. The cprof tool is based on the Clang compiler front-end to parse C-programs and to produce instrumented source code for the profiling. Using cprof, we determine a lower and upper bound limit for all 29 cases of the PolyBench/C benchmark suite. The lower and upper bound are determined using the absolute performance estimations assuming all statement are mapped onto the same processing resource and unbounded performance estimations assuming unlimited resources. We also compared the clock cycles found by cprof with RTL implementations for all 29 Polybench/C cases and found that cprof determines with 1.2% accuracy the correct number of clock cycles. It does this in a fraction of the time compared to the time needed to do a full RTL simulation. Wouter van Teijlingen, René van Leuken 0001, Carlo Galuzzi, Bart Kienhuis |
SCOPES | 2 |
| 2015 | Iterative learning cascaded multiclass kernel based support vector machine for neural spike data classificationabstractIn this paper, we develop an iterative learning framework based on multiclass kernel support vector machine (SVM) for adaptive classification of neural spikes. For efficient algorithm execution, we transform a multiclass problem with the Kesler's construction and extend iterative greedy optimization reduced set vectors approach with a cascaded method. Since obtained classification function is highly parallelizable, the problem is sub-divided and parallel units are instantiated for the processing of each sub-problem via energy-scalable kernels. After partition of the data into disjoint subsets, we optimize the data separately with multiple SVMs. We construct cascades of such (partial) approximations and use them to obtain the modified objective function, which offers high accuracy, has small kernel matrices and low computational complexity. Amir Zjajo, René van Leuken 0001 |
CIBCB | 2 |
| 2015 | Physical characterization of steady-state temperature profiles in three-dimensional integrated circuitsabstractThe thermal performance of three-dimensional integrated circuits is influenced by a number of design and technology parameters. However, the relationship between these parameters and thermal behaviour of die stacks is complex and not well understood. In this paper, we perform a detailed evaluation of the influence of stack composition and depth, thickness of dies, physical location of power dissipating elements and stack power density on steady-state temperature profiles. We examine how each of these parameters affects heat spread within 3D ICs, and highlight the causes for hotspot formation. The results of our analysis illustrate the implications of effective thermal conductivity on temperature sensing zones on dies, and the significant impact of stack power density on overall operating temperature. Sumeet S. Kumar, Amir Zjajo, René van Leuken 0001 |
ISCAS | 3 |
| 2015 | Stochastic noise analysis of neural interface front endabstractA time-domain methodology for noise analysis of neural interface front-end with arbitrary deterministic neuron model excitations is presented. Rather than estimating noise behavior by a population of realizations, the neural interface front-end is described as a set of stochastic differential equations and closure approximations are introduced to obtain the noise variances, covariances and cross-correlations between any electrical quantity and any stochastic source as a function of time. Statistical simulation shows that the proposed method offer an accurate and an efficient solution closely approximating those from a time-domain Monte Carlo analysis. Amir Zjajo, Carlo Galuzzi, René van Leuken 0001 |
ISCAS | 3 |
| 2015 | Ctherm: An Integrated Framework for Thermal-Functional Co-simulation of Systems-on-ChipabstractThis paper presents therm, an integrated framework for cycle-accurate thermal and functional evaluation of systems-on-chip. The presented framework enables accurate characterization of thermal behaviour by generating detailed physical models for components based on input specifications, and simulating them within a tightly integrated co-simulation platform with an embedded thermal simulator. Therm's fine-grained modelling approach yields 70% higher accuracy in hotspot resolution as compared to conventional approaches that abstract component internals. Simulation runtime time is reduced by up to 36% over conventional continuous approaches through the use of thermal check pointing, enabling the fast-forwarding of thermal simulations without loss of thermal continuity. Sumeet S. Kumar, Amir Zjajo, René van Leuken 0001 |
PDP | 3 |
| 2014 | Improving data cache performance using Persistence Selective CachingabstractThis paper presents Persistence Selective Caching (PSC), a selective caching scheme that tracks the reusability of L1 data cache (L1D) lines at runtime, and moves lines with sufficient potential for reuse to a low-latency, low-energy assist cache from where subsequent references to them are serviced. The selectivity of PSC is configurable, and can be adjusted to suit the varying memory access characteristics of different applications, unlike existing schemes. By effectively identifying reusable cache lines and storing them in the assist, PSC reduces average memory access time by upto 59% as compared to competing schemes and conventional data caches. Furthermore, by ensuring that only reusable lines are cached by the assist, PSC reduces cache line movements, and thus decreases average energy per access by upto 75% over other assists. Sumeet S. Kumar, René van Leuken 0001 |
ISCAS | 2 |
| 2014 | System Level Methodology for Interconnect Aware and Temperature Constrained Power Management of 3-D MP-SOCsabstractModern 3-D multiprocessor systems-on-chip (MP-SoC) incorporate processing elements (PEs) and memories within die-stacks interconnected using through-silicon vias (TSVs). The resulting power density of these systems necessitates the inclusion of thermal effects in the architecture space exploration stage of the design process. The number and placement of TSVs influences the thermal conductivity in the vertical direction in die-stacks, and consequently these must be considered during thermal analysis. However, the special requirement of keep out zones (KOZs) for TSVs due to mechanical stress considerations complicates the design of the vertical interconnect, potentially impacting its electrical performance as well. This paper presents an integrated methodology that allows for TSV topology exploration to evaluate the best vertical interconnect structure while considering crosstalk, area overheads, and KOZ requirements using an initial system floorplan. After incorporating feedback from the exploration, the resulting vertical interconnect is included within a temperature-power simulation that estimates the thermal profile of the 3-D stack. Within this methodology, a novel power management scheme for 3-D MP-SoCs that considers both temperature as well as positional information and thermal relationships between PEs, while performing dynamic voltage-frequency scaling (DVFS), is introduced. The scheme effectively maintains smooth temperature profiles, decreases fluctuations in voltage-frequency levels, and increases the aggregate frequency of operation at a lower total power dissipation. Further, the scheme is applied to a stack partitioned into voltage islands, where it is shown to match the conventional per-core DVFS schemes in its performance. Sumeet S. Kumar, Arnica Aggarwal, Radhika Sanjeev Jagtap, Amir Zjajo, René van Leuken 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2014 | Dynamic Thermal Estimation Methodology for High-Performance 3-D MPSoCabstractIn 3-D integrated circuits, accurate runtime sensing of on-chip temperature is required to establish dynamic thermal management instruction sets. Placement restrictions and excessive runtime thermal variations, however, compromise the performance and reliability of the sensor readings. Within this framework, a novel methodology for thermal estimation based on unscented Kalman filter, augmented only with a limited number of temperature sensors at a few selected locations, is proposed. In addition, we extend discontinuous Galerkin finite-element method to include coupling mechanism between neighboring grid cells for accurate thermal profile estimation and introduce a balanced stochastic truncation to find a low-dimensional but accurate approximation of the thermal network over the whole frequency domain. As the experimental results show, the runtime thermal estimation method reduces temperature estimation errors by an order of magnitude. Amir Zjajo, N. P. van der Meijs, René van Leuken 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2012 | A Methodology for Early Exploration of TSV Placement Topologies in 3D Stacked ICsabstractAs planar scaling to achieve higher chip integration seems to be on the brink of saturation, three-dimensional (3D) integration has emerged as a promising technology. It is critical to have efficient early stage estimation methodologies to build high performance digital systems as well as to shorten design time. In this paper, a novel methodology is proposed which takes into account key physical effects and explores Through-Silicon-Via (TSV) placement topologies for a 2-tier 3D stack. It estimates the interconnect electrical performance and TSV area penalty across two TSV performance corners. The methodology offers flexibility in selection of the CMOS technology node and the 3D stacking level. A SystemC implementation provides for parameterizability and modeling with ease and also enables integration into a high-level system simulation framework. Using our methodology, TSV placement topologies were explored for a 7-port 3D router. Our results present optimal topologies for the router for typical 45 nm and 32 nm technology nodes. They also point out unreliable topologies and give important feedback for 3D system design. Radhika Sanjeev Jagtap, Sumeet S. Kumar, René van Leuken 0001 |
DSD | 3 |
| 2012 | Multi-user leo-satellite receiver for robust space detection of AIS messagesabstractThe coverage of the terrestrial automatic identification system (AIS) is limited to close areas off the coast. Low earth orbit (LEO) satellites can expand the service of AIS to a global range but it brings large variance in Doppler shift, path loss and propagation delay. The communication between ships and LEO satellites becomes asynchronous. The collision of AIS messages from thousands of ground cells results in loss of all collided messages. Previous papers discussed the use of a single user receiver to detect AIS messages under heavy co-channel interference but the problem was never well solved. In this paper, we present a multi-user receiver equipped with an antenna array on LEO satellites, which explores the spatial multiplexing in space detection of AIS messages and significantly improves the detection performance. The proposed receiver performs rank tracking and subspace intersection based on the signed URV decomposition ahead of blind source separation to provide robust separation of user data for single user receivers. The proposed receiver is tested in an exact dynamic AIS model. Mu Zhou, Alle-Jan van der Veen, René van Leuken 0001 |
ICASSP | 3 |
| 2012 | Memory and computation reduction for least-square channel estimation of mobile OFDM systemsabstractMobile OFDM refers to OFDM systems with fast moving transceivers, contrastive to traditional OFDM systems whose transceivers are stationary or have a low velocity. In this paper, we use Basis Expansion Models (BEM) to model the time-variation of channels, based on which two least-squares (LS) channel estimators are presented. The first channel estimator allows for a general BEM assumption, and thus is called as the general implementation, whereas the second one is particularly tailored for a specific Critically-sampled Complex Exponential (CCE) BEM assumption, leading to a simplified architecture. The experimental results show that the simplified estimator is an appealing alternative, which achieves roughly a reduction of 59% for the ASIC core area, 89% for the ROM size and 53% of the processing latency at the cost of a slight estimation accuracy penalty compared to the general implementation method. Tao Xu 0001, Zijian Tang, René van Leuken 0001 |
ISCAS | 4 |
| 2012 | A 11 µW 0°C-160°C temperature sensor in 90 nm CMOS for adaptive thermal monitoring of VLSI circuitsabstractThis paper reports design, efficiency and measurement results of the temperature sensor based on substrate bipolar transistors and a PTAT multiplier for adaptive thermal monitoring of deep-submicron VLSI circuits. The prototype temperature sensor with un-calibrated 3σ accuracy of 0.9°C within a 0°C-160°C temperature range has been fabricated in standard single poly, six metal 90nm CMOS, consumes only 11μW at 1V power supply and measures 0.05mm2. Amir Zjajo, N. P. van der Meijs, René van Leuken 0001 |
ISCAS | 3 |
| 2011 | Extracting behavior and dynamically generated hierarchy from SystemC modelsabstractWe present a novel approach to extract the dynamically generated module hierarchy and its behavior from a SystemC model. SystemC is a popular modeling language which can be used to specify systems at a high abstraction level. The module hierarchy of a SystemC model is dynamically constructed during the execution of the elaboration phase of the model. This means that a system designer can build regular structures using loops and conditional statements. Currently, most SystemC tools can not cope with SystemC models for which the module hierarchy depends on dynamic parameters. In our approach this hierarchical information is retrieved by controlling and monitoring the executing of the elaboration phase of the model using a GDB debugger. Thereafter, the behavioral information is retrieved by using a GCC plug-in. This plug-in produces abstract syntax trees in static single assignment form. This behavioral information is linked with the hierarchical information. Our approach is completely non-intrusive. The SystemC model and the SystemC reference implementation can be used without any modification. We have implemented our approach in a SystemC front-end called SHaBE (SystemC Hierarchy and Behavior Extractor). This front-end facilitates the development of future SystemC visualization, debugging, static verification, and synthesis tools. Harry Broeders, René van Leuken 0001 |
DAC | 2 |
| 2011 | A Scalable Distributed Asynchronous Control Network for High Level Synthesis of Digital CircuitsabstractThis paper presents a scalable asynchronous distributed control network. The control circuit allows for true asynchronous operation of all digital resources and as a result of its scalable distributed topology allows unlimited resource sharing. We start with the description of a data flow graph, and using traditional scheduling algorithms, generate an asynchronous distributed control network and the asynchronous data path. The distributed controllers are implemented such that they can be created by connecting a small number of pre-designed sub-controllers which are presented in this paper. Prototype IP-blocks of these sub-controller circuits have been designed in a 90nm ASIC design process. To prove the effectiveness of our method, we present some key performance parameters: area and power under timing constraints. Tom Van Leeuwen 0002, René van Leuken 0001 |
DSD | 2 |
| 2011 | Systemc-AMS model of a dynamic large-scale satellite-based AIS-like network
Mu Zhou, René van Leuken 0001 |
FDL | 2 |
| 2010 | MB-LITE: A robust, light-weight soft-core implementation of the MicroBlaze architectureabstractDue to the ever increasing number of microprocessors which can be integrated in very large systems on chip the need for robust, easily modifiable microprocessors has emerged. Within this paper a light-weight cycle compatible implementation of the MicroBlaze architecture called MB-LITE is presented in an attempt to fill the gap in quality between commercial and open source processors. Experimental results showed that MB-LITE obtains very high performance compared with other open source processors while using very few hardware resources. The microprocessor can be easily extended with existing IP thanks to an easily configurable data memory bus and a wishbone bus adapter. All components are modular to optimize design reuse and are developed using a two-process design methodology for improved performance, simulation and synthesis speeds. All components have been thoroughly tested and verified on a FPGA. Currently an architecture with four MB-LITE cores in a NoC architecture is in development which will be implemented in 90 nm process technology. Tamar Kranenburg, René van Leuken 0001 |
DATE | 2 |
| 2006 | A multistandard FFT processor for wireless system-on-chip implementationsabstractThis paper presents a high performance FFT ASIP. The resulting programmable solution is scalable for the order of the FFT and capable of satisfying performance requirements of various OFDM wireless standards. The IEEE 802.15.3a ultra wideband OFDM - being the most time critical of these standards because of the computation of a 128-point FFT within 312.5 ns - has been the primary performance target of the scalable ASIP. The resulting ASIP adopts a vectorial ultra-long instruction word (ULIW) approach. The design decisions are evaluated with regards to processing speed, area and power dissipation. Ramesh Chidambaram, René van Leuken 0001, Marc Quax, Ingolf Held, Jos Huisken |
ISCAS | 2 |
| 2006 | Design of a practical scheme for ultra wideband communicationabstractIn the design of a packet-oriented impulse-radio UWB communication system, the main challenge at the receiver is to have a fast synchronization to the coded pulses, along with a detection of the message. We consider schemes that are straightforward to implement in practical systems and propose two methods to realize the synchronization algorithms: a serial and a parallel method. The algorithms for synchronization and demodulation are implemented in a receiver prototype based on an FPGA. Yiyin Wang, René van Leuken 0001, Alle-Jan van der Veen |
ISCAS | 2 |
| 2002 | General Purpose Prototyping Platform for Data-Processor Research and Development
Filip Miletic 0001, René van Leuken 0001, Alexander de Graaf |
FPL | 2 |
| 1988 | Concurrency Control in a VLSI Design Database
Ing Widya, René van Leuken 0001, Pieter van der Wolf |
DAC | 2 |
| 1988 | Object Type Oriented Data Modeling for VLSI Data Management
Pieter van der Wolf, René van Leuken 0001 |
DAC | 2 |