EDBT 2026 Demo / reviewers in the wild / expert
Endri Kaja
dblp:304/3632
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NURSE: A Distributed Architecture for Runtime Fault Management in Processor Designs
Ashwin Santhosh, Artur Jutman, Endri Kaja, Wolfgang Ecker, Maksim Jenihhin |
ETS | 3 |
| 2024 | PaGoRi:A Scalable Parallel Golomb-Rice DecoderabstractDeep Neural Networks (DNNs) have created opportunities to address real-world issues and expand the application of Artificial Intelligence (AI). Despite significant accuracy enhancements, DNNs pose a challenge when deployed on resource-limited edge devices commonly used in Internet of Things (IoT) applications. Inference execution of the DNNs requires accessing millions of parameters responsible for most energy consumption. Compression of weights is one possible solution, but most of the existing hardware decompression units could be more efficient in terms of power, area, and energy. This paper presents a scalable version of a hardware-efficient Parallel Golomb-Rice decoder (PaGoRi). The decoder has been integrated with an industry-strength Neural Network (NN) accelerator and evaluated with three TinyML benchmarks. The PaGoRi decoder achieves optimal trade-offs between power consumption and throughput, supporting decoding capacities of four and eight weights, consuming 0.43 mW and 0.79 mW of power, respectively, while achieving a throughput of 888 MBps and 1.3 GBps, respectively. Mounika Vaddeboina, Endri Kaja, Alper Yilmazer, Uttal Ghosh, Wolfgang Ecker |
DDECS | 2 |
| 2023 | Bits, Flips and RISCsabstractElectronic systems can be submitted to hostile environments leading to bit-flips or stuck-at faults and, ultimately, a system malfunction or failure. In safety-critical applications, the risks of such events should be managed to prevent injuries or material damage. This paper provides a comprehensive overview of the challenges associated with designing and verifying safe and reliable systems, as well as the potential of the RISC-V architecture in addressing these challenges.We present several state-of-the-art safety and reliability verification techniques in the design phase. These include a highly-automated verification flow, an automated fault injection and analysis tool, and an AI-based fault verification flow. Furthermore, we discuss core hardening and fault mitigation strategies at the design level. We focus on automated SoC hardening using model-driven development and resilient processing based on sensing and prediction for space and avionic applications.By combining these techniques with the inherent flexibility of the RISC-V architecture, designers can develop tailored solutions that balance cost, performance, and fault tolerance to meet the requirements of various safety-critical applications in different safety domains, such as avionics, automotive, and space. The insights and methodologies presented in this paper contribute to the ongoing efforts to improve the dependability of computing systems in safety-critical environments. Nicolas Gerlin, Endri Kaja, Fabian Vargas 0001, Anselm Breitenreiter, Junchao Chen 0001, Markus Ulbricht 0002, Maribel Gomez, Ares Tahiraga, Sebastian Siegfried Prebeck, Eyck Jentzsch, Milos Krstic, Wolfgang Ecker |
DDECS | 2 |
| 2023 | Parallel Golomb-Rice Decoder with 8-bit Unary Decoding for Weight Compression in TinyML ApplicationsabstractDue to the recent advances in AI, the requirement for Artificial Intelligence (AI) has increased exponentially in the domain of Internet of Things (IoT). Running Deep Neural Networks (DNNs) on edge devices gives the advantage of privacy, security, and lower latency. It is challenging to deploy them on embedded devices with constrained hardware resources since a lot of compute and memory resources are required. Memory access contributes to the majority of the energy requirements on edge devices. Although data compression plays a critical role in reducing storage and memory bandwidth requirements, most of the hardware decoders are inefficient in terms of power, area, and throughput. In this work, a hardware Parallel Golomb-Rice decoder is presented that can decode 8-bits of unary encoded data every cycle. The design has been integrated with a Neural Network (NN) accelerator and experimented with state-of-the-art benchmark models. Lossless compression is performed with an offline Golomb-Rice encoder. It encodes the weights of each layer with an optimum Golomb-Rice parameter. Applied to the benchmarks Anomaly Detection, Image Classification and Visual Wake Words the memory access during inference is reduced by 26.8%, 6.62% and 5.54% respectively. The decoder dissipates 0.4216 mW of power and delivers an average throughput of 860 MBps. The design has been synthesised with 40 nm technology and compared with state-of-the-art works. Mounika Vaddeboina, Endri Kaja, Alper Yilmayer, Sebastian Siegfried Prebeck, Wolfgang Ecker |
DSD | 2 |
| 2022 | Design of a Tightly-Coupled RISC-V Physical Memory Protection Unit for Online Error DetectionabstractWhile semiconductors are becoming more efficient generation after generation, the continuous technology scaling leads to numerous reliability issues due, amongst others, to variations in transistors characteristics, manufacturing defects, component wear-out, or interference from external and internal sources. Induced bit flips and stuck-at-faults can lead to a system failure. Security-critical systems often use Physical Memory Protection (PMP) modules to enforce memory isolation. The standard loosely-coupled approach eases the implementation but creates overhead in area and performance, limiting the number of protected areas and their size. While delivering great support against malicious software and induced faults, better performance would benefit safety tasks by preventing the program from jumping into an undesired region and giving wrong outputs.We propose a novel model-driven approach to resolve these limitations by generating a tightly-coupled RISC-V PMP, which reduces the impact of run-time reconfiguration. We also discuss guidelines on configuring a PMP to minimize the overhead on performance and memory, and provide an area estimation for each possible PMP design instance. We formally verified a RISC-V Core with a PMP and evaluated its performance with the Dhrystone Benchmark. The presented architecture shows a performance gain of about 3 times against the standard implementation. Furthermore, we observed that adding the PMP feature to a RISC-V SoC led to a negligible performance loss of less than 0.1% per thousand PMP reconfigurations. Nicolas Gerlin, Endri Kaja, Monideep Bora, Keerthikumara Devarajegowda, Dominik Stoffel, Wolfgang Kunz, Wolfgang Ecker |
VLSI-SoC | 2 |
| 2022 | Fast and Accurate Model-Driven FPGA-based System-Level Fault EmulationabstractSafety-critical designs need to ensure reliable operations even under a hostile working environment with a certain degree of confidence. Continuous technology scaling has resulted in designs being more susceptible to the risk of failure. As a result, the safety requirements are constantly evolving and becoming more stringent. For validating and measuring the robustness of safety-critical designs, fault injection methods are employed within the design flows. To ensure safety requirements’ compliance, and at the same time to cope with the ever-increasing complexity of modern SoCs, the existing design flows become inadequate as the process is repetitive, time-tedious, and requires high manual efforts. In this paper, a fully automated, fast and accurate, fault emulation framework based on the FPGA platform is proposed that enables a high level of controllability and observability for fault injection. The approach uses model-driven engineering concepts and automates various fault injection campaigns, namely, statistical fault injection (SFI), direct fault injection (DFI), and exhaustive fault injection (EFI). A novel design architecture tailored for the FPGA platform is also proposed to improve the overall productivity of performing fault emulation. The proposed approach scales to a wide variety of RISC-V based CPU subsystems with varying complexity in size and features. The experimental results demonstrate a significant gain in the fault emulation performance by a factor of 2.75x to 47.57x when compared to the standard simulation-based fault injection methods. Endri Kaja, Nicolas Gerlin, Monideep Bora, Gabriel Rutsch, Keerthikumara Devarajegowda, Dominik Stoffel, Wolfgang Kunz, Wolfgang Ecker |
VLSI-SoC | 1 |
| 2021 | ISA Modeling with Trace Notation for Context Free Property GenerationabstractThe scalable and extendable RISC-V ISA introduced a new level of flexibility in designing highly customizable processors. This flexibility in processor designs adds to the complexity of already complex functional verification process. Although formal methods are increasingly used to exhaustively verify the processors, the required manual effort and verification expertise become the major hurdles in industrial flows. Furthermore, efficient ISA modeling techniques are required that are scalable to multiple ISA extensions and to different architectural variants of a processor. This paper proposes a trace notation for ISA definition to capture the implicit execution behavior of a processor and specific characteristics. The proposed trace notation can be annotated with timing and hierarchy information to adapt the trace to any kind of processor architecture. From this trace notation, a complete set of properties are generated to detect all functional bugs in a processor implementation. The approach requires significantly less manual effort compared to the contemporary techniques. Its industry strength has been demonstrated by formally verifying a wide variety of RISC-V processor implementations with one or more ISA extensions (RV32I, C, Zicsr, M and custom extensions supporting AI acceleration and safety features). Keerthikumara Devarajegowda, Endri Kaja, Sebastian Siegfried Prebeck, Wolfgang Ecker |
DAC | 2 |