Peeter Ellervee

dblp:04/4257 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-0745-6743ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Architectural Exploration and Implementation of CERN LHC Trigger Algorithm With FPGA
abstract
The upcoming high-luminosity upgrade of the Large Hadron Collider (LHC) at CERN will increase data rates to values far exceeding the capabilities of software-based processing systems. As a result, new methods are required to efficiently extract scientifically valuable information from the massive data streams produced by the LHC’s particle detectors. This article discusses the design of the tau lepton trigger algorithm using the field-programmable gate array (FPGA) technology. Given its complexity and demanding technical requirements, realizing the algorithm in FPGA is a challenging task. This article details the algorithm development using high-level synthesis (HLS), a technique to generate hardware descriptions from the C++ code. We discuss architectural solutions and optimizations explored during the design process, including algorithm partitioning and pipelining, optimization of pipeline stages, floorplanning, and probing of implementation strategies. The performed design space exploration helped to improve latency, solve area-related issues, and reduce routing congestion, enabling implementation of tau lepton trigger on FPGA.
Sergei Devadze, Christine Elizabeth Nielsen, Natalia Cherezova, Dmitri Mihhailov, Peeter Ellervee
IEEE Trans. Very Large Scale Integr. Syst.5
2023 ML-Based Online Design Error Localization for RISC-V Implementations
abstract
The accelerated growth of computing systems' complexity makes comprehensive design verification challenging and time-consuming. In practice, hard-to-model complex environments are unfeasible to be simulated exhaustively within a reasonable time frame. Therefore, some corner-case conditions can be overlooked and design errors might escape to the final product. This means that it is imperative for the system to be able to detect and locate bugs to enable self-repair. This is particularly crucial during long-term remote missions in order to apply graceful degradation. This paper proposes a novel online design error localization methodology for microprocessors by immediate analysis of traced and buffered signals upon a failure detection event, using a pre-trained Neural Network (NN) and existing processor components, i.e. trace buffers and AI accelerators. An in-house Neural Architecture Search (NAS) framework is used to train a tailored Multi-Layer Perceptron (MLP) NN for error localization at the microprocessor module-level resolution. The proposed approach is validated by simulating a RISC-V implementation with different workload programs. It is demonstrated to be capable of localizing the microprocessor module of bug origin with 92.81% accuracy, on average.
Hardi Selg, Maksim Jenihhin, Peeter Ellervee, Jaan Raik
IOLTS3
2022 Surviving the Unforeseen - Teaching IT and Engineering Students During COVID-19 Outbreak
abstract
In this Full Paper on Innovative Practise, we sum up our teaching experience during the COVID-19 pandemic. We focus on several courses taught at the Department of Computer Systems. The main technique to cope with COVID-19 restrictions in higher education is to use online teaching solutions. The first step was to make the traditional lectures available online through recordings enabling students to attend lectures as needed. Yet, for several courses, a quick reconfiguration was required for laboratory exercises as COVID restrictions were established in the middle of the study semester. Thereby, an irreplaceable solution was to use remote laboratories. One of the main concerns was how to set up the online environment within days and weeks and not in months. In addition to remote laboratories, we propose an online teaching and custom assessment setup that to the best of our knowledge has not been used before. For example, instead of using Proctorio for online assessment, we propose a two-stage, licence-free, online examination setup. The proof of concept was carried out with more than 200 students.
Priit Ruberg, Peeter Ellervee, Kalle Tammemäe, Uljana Reinsalu, Andres Rähni, Tarmo Robal
FIE2
2021 CLD: An Accurate, Cost-Effective and Scalable Run-Time Cache Leakage Detector
abstract
Cache logical side channel attacks pose a significant threat to the security of modern computer systems. This is a result of exploitation of cache information leakages arising from cache contention. Detection of such leakages can be inferred from cache behavior and processes' access patterns during run time. To achieve this, a detection template that uses available information on cache outputs and process accesses at run-time is required. In this work, such template is proposed and implemented as a hardware monitor called Cache Leakage Detector (CLD). CLD is a high-accuracy, cost-effective and scalable run-time cache information leakage detector. CLD uses cache signals and process IDs to detect exploitable cache access patterns. It does so by identifying potential information leakage patterns. Accuracy of CLD is evaluated by using several benchmarks and injecting attacks into a 128-bit key AES algorithm. The experiments demonstrate that CLD has far higher detection accuracy (0.7964 vs 0.3195) and lower percentage of false positive detections (1.2% vs 30.6%) compared to a state-of-the-art hardware detector. Moreover, CLD introduces a very low area overhead of 0.002% to the total area of the cache. Experimental result section reports the above claims in detail.
Ameer Shalabi, Tara Ghasempouri, Peeter Ellervee, Jaan Raik
DDECS3
2020 SCAAT: Secure Cache Alternative Address Table for mitigating cache logical side-channel attacks
abstract
Interest in memory systems' security has increased during the last decade due to their vulnerabilities to be exploited by logical side channels attacks. A promising approach for attack detection at run-time is to monitor the cache memory's behavior. However, designing an environment capable of detecting and mitigating these attacks is very challenging. In current monitoring systems, attack mitigation has been largely neglected. To overcome these shortcomings, in this work, we present a secure cache called SCAAT. SCAAT is equipped with an attack mitigation system to handle attacks by remapping where data is stored in the cache to random locations. In addition, SCAAT uses an attack monitor that identifies suspicious behavior that indicates cache logical side-channel attacks. The effectiveness of SCAAT is analyzed and evaluated for several cache configurations in terms of area overhead and performance.
Ameer Shalabi, Tara Ghasempouri, Peeter Ellervee, Jaan Raik
DSD3
2017 Standards-based tools and services for building lifelong learning pathways
abstract
The decade that we have embarked upon presents enormous challenges for Europe. The 2020 strategy for smart, sustainable and inclusive growth recognises the key role higher education must play if the ambitions for Europe in a fast-changing global reality are to be realised. This implies widening access to lifelong learning to as many European citizens as possible and it is vital that measures are implemented to transform our reality towards this direction. Notably, labour markets increasingly require more graduates with specialized knowledge and competences and substantial investment has to be made in education systems to ensure that this demand is met. The COMPASS project aims to the Composition of Lifelong Learning Opportunity Pathways through Standards-based Services investing on the establishment of a cohesive, strategic partnership for the longterm promotion of LLL-related European strategies and the development of instruments that will raise the awareness of both learning opportunity providers and learning opportunity seekers and stimulate the design of policies for enhancing ET with alternative, flexible pathways on the basis of easy, technology-enhanced access to learning opportunities.
Cleo Sgouropoulou, Ioannis Voyiatzis, Anastasios Koutoumanos, Said Hamdioui, Peyman Pouyan, M. Comte, Paolo Prinetto, Giuseppe Airo Farulla, Peeter Ellervee, Carlos Delgado Kloos, Raquel M. Crespo García
EDUCON9
2016 Polymorphic Configuration Architecture for CGRAs
abstract
In the era of platforms hosting multiple applications with arbitrary reconfiguration requirements, static configuration architectures are neither optimal nor desirable. The static reconfiguration architectures either incur excessive overheads or cannot support advanced features (like time-sharing and runtime parallelism). As a solution to this problem, we present a polymorphic configuration architecture (PCA) that provides each application with a configuration infrastructure tailored to its needs.
Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Ahmed Hemani, Kolin Paul, Juha Plosila, Peeter Ellervee, Hannu Tenhunen
IEEE Trans. Very Large Scale Integr. Syst.6
2014 Morphable Compression Architecture for Efficient Configuration in CGRAs
abstract
Today, Coarse Grained Reconfigurable Architectures (CGRAs) host multiple applications. Novel CGRAs allow each application to exploit runtime parallelism and time sharing. Although these features enhance the power and silicon efficiency, they significantly increase the configuration memory overheads (up to 50% area of the overall platform). As a solution to this problem researchers have employed statistical compression, intermediate compact representation, and multicasting. Each of these techniques has different properties (i.e. compression ratio and decoding time), and is therefore best suited for a particular class of applications (and situation). However, existing research only deals with these methods separately. In this paper we propose a morphable compression architecture that interleaves these techniques in a unique platform. The proposed architecture allows each application to enjoy a separate compression/decompression hierarchy (consisting of various types and implementations of hardware/software decoders) tailored to its needs. Thereby, our solution offers minimal memory while meeting the required configuration deadlines. Simulation results, using different applications (FFT, Matrix multiplication, and WLAN), reveal that the choice of compression hierarchy has a significant impact on compression ratio (from configware replication to 52%) and configuration cycles (from 33 nsec to 1.5 secs) for the tested applications. Synthesis results reveal that introducing adaptivity incurs negligible additional overheads (1%) compared to the overall platform area.
Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Masoud Daneshtalab, Ahmed Hemani, Kolin Paul, Peeter Ellervee, Juha Plosila, Hannu Tenhunen
DSD6
2014 Customizable Compression Architecture for Efficient Configuration in CGRAs
abstract
Today, Coarse Grained Reconfigurable Architectures (CGRAs) host multiple applications. Novel CGRAs allow each application to exploit runtime parallelism and time sharing. Although these features enhance the power and silicon efficiency, they significantly increase the configuration memory overheads. As a solution to this problem researchers have employed statistical compression, intermediate compact representation, and multicasting. Each of these techniques has different properties, and is therefore best suited for a particular class of applications. However, existing research only deals with these methods separately. In this paper we propose a morphable compression architecture that interleaves these techniques in a unique platform.
Syed M. A. H. Jafri, Muhammad Adeel Tajammul, Masoud Daneshtalab, Ahmed Hemani, Kolin Paul, Peeter Ellervee, Juha Plosila, Hannu Tenhunen
FCCM6
2012 Multisine signal generation method for a bioimpedance measurement device
abstract
Implementing a high-speed multisine-wave synthesiser in hardware is, although common, but hardly a trivial task. In this paper we review and analyse several prominent approaches for generating a multisine signal. The most appropriate method for a given bioimpedance measurement device is considered and improved.
Maksim Gorev, Vadim Pesonen, Peeter Ellervee
DDECS3
2011 Communication modelling and synthesis for NoC-based systems with real-time constraints
abstract
This paper addresses the communication modelling and synthesis problem for applications implemented on networks-on-chip. Due to the communication complexity of such systems it is difficult to estimate the communication delay. On the other hand, guaranteeing the timing constraints without detailed know-how about the communication is impossible. In this work we propose a communication modelling and synthesis approach for networks-on-chip where communication infrastructure is not able to provide communication interleaving (such as TDMA, virtual channels) or to guarantee communication delays. The idea is, to design a communication synthesis method, which would not be run off-chip as a CAD tool on a workstation, but on-chip and being activated whenever the system-on-chip (SoC) is re-configured.
Mihkel Tagel, Peeter Ellervee, Thorsten Hollstein, Gert Jervan
DDECS2
2005 Improved Fault Emulation for Synchronous Sequential Circuits
abstract
Current paper presents new alternatives for accelerating the task of fault simulation for sequential circuits by hardware emulation on FPGA. Fault simulation is an important subtask in test pattern generation and it is frequently used throughout the test generation process. The problems associated to fault emulation for sequential circuits are explained and alternative implementations are discussed. An environment for hardware emulation of fault simulation is presented. It incorporates hardware support for fault dropping. The proposed approach allows simulation speed-up of 40 to 500 times as compared to the state-of-the-art in fault simulation. Average speedup provided by the method is 250 that is about an order of magnitude higher than previously cited in the literature. Based on the experiments, we can conclude that it is beneficial to use emulation when large numbers of test vectors is required.
Jaan Raik, Peeter Ellervee, Valentin Tihhomirov, Raimund Ubar
DSD2
2004 Evaluating Fault Emulation on FPGA
Peeter Ellervee, Jaan Raik, Valentin Tihhomirov, Kalle Tammemäe
FPL1
2001 System-level data-format exploration for dynamically allocated datastructures
abstract
System-level exploration of memory organizations is a key issue in successful implementation of data dominated applications based on dynamically allocated data structures involving records and access keys. This paper presents a formalized technique for exploring different memory data-format alternatives when only the system level functional behavior of the application has been defined. Our data-format exploration approach allows to substantially minimize the number of accessed bits by rearranging the format of the data records. The technique exploits parallelism in the data transfer by analyzing the dependencies between data-record accesses. As a result, significant reduction in memory size, bandwidth, and power are obtained. We have validated our techniques using several real-life asynchronous transfer mode cell processing applications, where we have obtained reductions in memory size (up to 20%), power (up to a 60%), and bandwidth.
Peeter Ellervee, Miguel Corbalan, Francky Catthoor, Ahmed Hemani
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2000 System-level data format exploration for dynamically allocated data structures
abstract
Memory bandwidth and pow er consumption are important design bottlenecks for data dominated applications. We propose a systematic system level exploration approach and formalised techniques to alleviate these bottlenec ks based on rearranging the format of the data records that are later stored in memory. The technique exploits parallelism in the data transfer and reduction in bit waste. Using our approach on sev eral real-life ATM processing applications, significant reduction in size, bandwidth and hence pow er consumption are obtained.
Peeter Ellervee, Miguel Corbalan, Francky Catthoor, Ahmed Hemani
DAC1
2000 TOP: An Algorithm for Three-Level Optimization of PLDs
abstract
Summary form only given. Presents an heuristic algorithm TOP (Three-level Optimization of PLDs), targeting a three-level logic expression of type g/sub 1/ o g/sub 2/, where g/sub 1/ and g/sub 2/ are sum-of-products and "o" is a binary operation. Such an expression can be implemented by a three-level Programmable Logic Device (PLD) consisting of PLA 1 and PLA2, implementing the first two levels of logic, and a set of two-input logic expanders, implementing the third level. Each logic expander can be programmed to realize any function of two variables. PLDs of this type seem to give a good trade-off between the speed of a flat PLA and density of a multi-level network of PLAs. TOP chooses the functionality of the logic expanders so that the area of the PLAs is minimized.
Elena Dubrova, Peeter Ellervee, D. Michael Miller, Jon C. Muzio
DATE2
1999 Lowering Power Consumption in Clock by Using Globally Asynchronous Locally Synchronous Design Style
abstract
Power consumption in clock of large high performance VLSIs can be reduced by adopting Globally Asynchronous, Locally Synchronous design style (GALS).GALS has small overheads for the global asynchronous communication and local clock generation.We propose methods to a) evaluate the benefits of GALS and account for its overheads, which can be used as the basis for partitioning the system into optimal number/size of synchronous blocks, and b) automate the synthesis of the global asynchronous communication.Three realistic ASICs, ranging in complexity from 1 to 3 million gates, were used to evaluate GALS benefits and overheads.The results show an average power saving of about 70% in clock with negligible overheads. Lowering power consumption in clock by using Globally AsynchronousLocally Synchronous design style.
Ahmed Hemani, Thomas Meincke, Shashi Kumar, Adam Postula, Thomas Olsson 0001, Peter Nilsson 0001, Johnny Öberg, Peeter Ellervee, Dan Lundqvist
DAC8