Johnny Öberg

dblp:98/2596 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-8072-1742ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 9 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Exploring the Potential of LSTM On Emulating Multiple-bit Fault Injection in SRAM-FPGA
Trishna Rajkumar, Johnny Öberg
SAFECOMP2
2024 Guided Fault Injection Strategy for Rapid Critical Bit Detection in Radiation-Prone SRAM-FPGA
abstract
Fault injection test is vital for assessing the reliability of SRAM-FPGAs used in radiative environments. Considering the scale and complexity of modern FPGAs, exhaustive fault injection is tedious and computationally expensive. A common approach to optimising the injection campaign involves targeting a subset of the configuration memory containing essential and critical bits crucial for the system's functionality. Identifying Essential bits in an FPGA design is often feasible through manufacturer documen-tation. However, detecting Critical bits requires complex reverse engineering to map the correspondence between the configuration bits and the FPGA modules. This task requires substantial amount of details about the logic layout and the bitstream, which is not easily available due to their proprietary nature. In some cases, manual floorplanning becomes necessary, which could impact the performance of the application. Given these limitations, we examine the potential of Monte Carlo Tree Search in guiding the fault injection process to identify critical bits with minimal injections. The key benefit of this approach is its ability to harness the spatial relations among the configuration bits without relying on reverse engineering or offline campaign planning. Evaluation results demonstrate that the proposed approach achieves a 99 % coverage using 18 % fewer injections than traditional methods. Notably, 95% of the critical bits were detected in under 50% injections, achieving at least 2X higher sensitivity to critical bits with a minimal overhead of 0.04%.
Trishna Rajkumar, Johnny Öberg
DATE2
2022 A Markovian Approach for Detecting Failures in the Xilinx SEM core
abstract
The soft error mitigation (SEM) core is an internal scrubber used to detect and correct single event upsets in the configuration memory. Although the core can mitigate errors with a high accuracy, recent studies have found it to be vulnerable to radiation errors owing to its implementation in the FPGA fabric. As the reliability of the system depends on the correctness of the scrubber, undetected SEM failure is hazardous in critical applications. In this study, we investigate the effectiveness of Markov chains in detecting such failures. In order to minimise the effects of single event upsets, the detection scheme is implemented external to the FPGA and leverages log analysis to monitor the SEM health. We evaluated our approach on the Xilinx ZCU104 Ultrascale+ board using fault injection. The results show that the SEM failures caused by single and double bit errors could be detected with an$F_{1}$score of 0.90 and 0.99 respectively. To the best of our knowledge, this is the first custom approach for failure detection in the SEM core.
Trishna Rajkumar, Johnny Öberg
FPT2
2022 AnoDe: A Log-based Self-Supervised Framework to Detect Scrubber Failures in SRAM-FPGA
abstract
SRAM-FPGAs used in radiative environment are integrated with a scrubber to protect the configuration memory from radiation effects. Any malfunctions in the scrubber degrades the reliability of the system and can have catastrophic consequences in critical applications. Existing solutions for a reliable scrubber focus on masking or detecting the scrubber errors through redundant modules implemented in the FPGA. While these approaches improve the overall reliability, they are not completely radiation-proof owing to their implementation on FPGA. In order to improve the scrubber reliability, a complementary scheme external to the FPGA is required. Based on this consideration, we propose AnoDe, a failure detection framework running on a supervisory processor external to the FPGA board. AnoDe leverages the logs generated by the scrubber to detect failures in real-time using an Autoencoder network. AnoDe provides a self-supervised solution right from generating the labelled training data to dynamically adapting to the prevailing radiation conditions. We evaluated the effectiveness of our approach on a Xilinx Ultrascale+ MPSoC ZCU104 board using fault injection. The results demonstrated a detection performance comparable to that of a custom approach with an F1 score of 0.85 for single bit upsets and 0.93 for multi bit upsets. Overall, the proposed approach could reduce the scrubber SEU sensitivity from 6 % to 1 %.
Trishna Rajkumar, Johnny Öberg
PRDC2
2018 Exploring Power and Throughput for Dataflow Applications on Predictable NoC Multiprocessors
abstract
System level optimization for multiple mixed-criticality applications on shared networked multiprocessor platforms is extremely challenging. Substantial complexity arises from the interdependence between the multiple subproblems of mapping, scheduling and platform configuration under the consideration of several, potentially orthogonal, performance metrics and constraints. Instead of using heuristic algorithms and problem decomposition, novel unified design space exploration (DSE) approaches based on Constraint Programming (CP) have in the recent years shown promising results. The work in this paper takes advantage of the modularity of CP models, in order to support heterogeneous multiprocessor Network-on-Chip (NoC) with Temporally Disjoint Networks (TDNs) aware message injection. The DSE supports a range of design criteria, in particular the optimization and satisfaction of power and throughput. In addition, the DSE now provides a valid configuration for the TDNs that guarantees the performance required to fulfil the design goals. The experiments show the capability of the approach to find low-power and high-throughput designs, and validate a resulting design on a physical TDN-based NoC implementation.
Kathrin Rosvall, Tage Mohammadat, George Ungureanu, Johnny Öberg, Ingo Sander
DSD4
2016 SAFEPOWER Project: Architecture for Safe and Power-Efficient Mixed-Criticality Systems
abstract
With the ever increasing industrial demand for bigger, faster and more efficient systems, a growing number of cores is integrated on a single chip. Additionally, their performance is further maximized by simultaneously executing as many processes as possible not regarding their criticality. Even safety critical domains like railway and avionics apply these paradigms under strict certification regulations. As the number of cores is continuously expanding, the importance of cost-effectiveness grows. One way to increase the cost-efficiency of such System on Chip (SoC) is to enhance the way the SoC handles its power resources. By increasing the power efficiency, the reliability of the SoC is raised, because the lifetime of the battery lengthens. Secondly, by having less energy consumed, the emitted heat is reduced in the SoC which translates into fewer cooling devices. Though energy efficiency has been thoroughly researched, there is no application of those power saving methods in safety critical domains yet. The EU project SAFEPOWER1 targets this research gap and aims to introduce certifiable methods to improve the power efficiency of mixed-criticality real-time systems (MCRTES). This paper will introduce the requirements that a power efficient SoC has to meet and the challenges such a SoC has to overcome.
Alina Lenz, Mikel Azkarate-askatsua, Javier Coronel, Alfons Crespo, Simon Davidmann, Juan Carlos Diaz Garcia, Nera González Romero, Kim Grüttner, Roman Obermaisser, Johnny Öberg, Jon Pérez 0001, Ingo Sander, Ingemar Söderquist
DSD10
2014 From Simulink to NoC-based MPSoC on FPGA
abstract
Network-on-chip (NoC) based multi-processor systems are promising candidates for future embedded system platforms. However, because of their complexity, new high level modeling techniques are needed to design, simulate and synthesize embedded systems targeting NoC-based MPSoC. Simulink is a popular modeling environment suitable to model at system level. However, there is no clear standard to synthesize Simulink models into SW and HW towards a NoC-based MPSoC implementation. In addition, many of the proposed solutions require large overhead in terms of SW components and memory requirements, resulting in complex and customized multi-processor platforms. In this paper we present a novel design flow to synthesize Simulink models onto a NoC-based MPSoC running on low-cost FPGAs. Our design flow constrains the MPSoC and the Simulink model to share a common semantics domain. This permits to reduce the need of resource consuming SW components, reducing the memory requirements on the platform. At the same time, performances (throughput) of dataflow applications can increase when the number of processors of the target platform is increased. This is shown through a case study on FPGA.
Francesco Robino, Johnny Öberg
DATE2
2013 The RecoBlock SoC platform: a flexible array of reusable run-time-reconfigurable IP-blocks
abstract
Run-time reconfigurable (RTR) FPGAs combine the flexibility of software with the high efficiency of hardware. Still, their potential cannot be fully exploited due to increased complexity of the design process. Consequently, to enable an efficient design flow, we devise a set of prerequisites to increase the flexibility and reusability of current FPGA-based RTR architectures. We apply these principles to design and implement the RecoBlock SoC platform, which main characterization is (1) a RTR plug-and-play IP-Core whose functionality is configured at run-time; (2) flexible inter-block communication configured via software, and (3) built-in buffers to support data-driven streams and inter-process communications. We illustrate the potential of our platform by a tutorial case study using an adaptive streaming application to investigate different combinations of reconfigurable arrays and schedules. The experiments underline the benefits of the platform and shows resource utilization.
Byron Navas, Ingo Sander, Johnny Öberg
DATE3
2007 Toward a scalable test methodology for 2D-mesh Network-on-Chips
abstract
This paper presents a BIST strategy for testing the NoC interconnect network, and investigates if the strategy is a suitable approach for the task. All switches and links in the NoC are tested with BIST, running at full clock-speed, and in a functional-like mode. The BIST is carried out as a go/no-go BIST operation at start up, or on command. It is shown that the proposed methodology can be applied for different implementations of deflecting switches, and that the test time is limited to a few thousand-clock cycles with fault coverage close to 100%
Kim Petersén, Johnny Öberg
DATE2
2004 System design for DSP applications in transaction level modeling paradigm
abstract
In this paper, we systematically define three transaction level models (TLMs), which reside at different levels of abstraction between the functional and the implementation model of a DSP system. We also show a unique language support to build the TLMs. Our results show that the abstract TLMs can be built and simulated much faster than the implementation model at the expense of a reasonable amount of simulation accuracy.
Abhijit K. Deb, Axel Jantsch, Johnny Öberg
DAC3
2004 System Design for DSP Applications Using the MASIC Methodology
abstract
Expensive top-down iterations are often required in the design cycle of complex DSP systems. In this paper, we introduce two levels of abstraction in the design flow by systematically categorizing the architectural decisions. As a result, the top-down iteration loop is broken. We also present a technique to capture and inject the architectural decisions such that the system models can be created and simulated efficiently. The concepts are illustrated by a realistic speech processing example, which is implemented using the AMBA on-chip architecture. Our methodology offers a smooth path from the functional modeling phase to the implementation level, facilitates the reuse of HW and SW components, and enjoys existing tool support at the backend.
Abhijit K. Deb, Axel Jantsch, Johnny Öberg
DATE3
2004 A study on the implementation of 2-D mesh-based networks-on-chip in the nanometre regime
Dinesh Pamunuwa, Johnny Öberg, Lirong Zheng 0001, Mikael Millberg, Axel Jantsch, Hannu Tenhunen
Integr.2
2004 Special issue on networks on chip
Axel Jantsch, Johnny Öberg, Hannu Tenhunen
J. Syst. Archit.2
2003 Simulation and Analysis of Embedded DSP Systems Using MASIC Methodology
Abhijit K. Deb, Johnny Öberg, Axel Jantsch
DATE2
2003 Load Distribution with the Proximity Congestion Awareness in a Network on Chip
Erland Nilsson, Mikael Millberg, Johnny Öberg, Axel Jantsch
DATE3
2003 Layout, Performance and Power Trade-Offs in Mesh-Based Network-on-Chip Architectures
Dinesh Pamunuwa, Johnny Öberg, Lirong Zheng 0001, Mikael Millberg, Axel Jantsch
VLSI-SOC2
2001 Grammar-based design of embedded systems
Johnny Öberg, Mattias O'Nils, Axel Jantsch, Adam Postula, Ahmed Hemani
J. Syst. Archit.1
2000 Grammar-based hardware synthesis from port-size independent specifications
abstract
A protocol defines how systems communicate. There are two ways of specifying the protocol, the language of communication. One way is to specify the automaton that recognizes the language, and this is the approach taken by SDL, etc. The other more abstract way ss to specify the grammar of the language and let a tool synthesize the automaton. Directly specifying the automaton makes the specification implementation dependent in two ways: the time behavior is specified in terms of states, and the width of the inputs and outputs is fixed. By specifying the grammar, the specification is potentially independent of both these implementation details and allows design space exploration in these dimensions. This paper presents a grammar-based language, called Program, that supports a port-size independent specifications methodology and its application to parts of the Operation and Maintenance protocol, a typical application from the ATM world. The methodology has also been applied to another test set of example designs and compared to standard RTL synthesis and HLS in order to evaluate the quality of the produced designs.
Johnny Öberg, Ahmed Hemani
IEEE Trans. Very Large Scale Integr. Syst.1
1999 Lowering Power Consumption in Clock by Using Globally Asynchronous Locally Synchronous Design Style
abstract
Power consumption in clock of large high performance VLSIs can be reduced by adopting Globally Asynchronous, Locally Synchronous design style (GALS).GALS has small overheads for the global asynchronous communication and local clock generation.We propose methods to a) evaluate the benefits of GALS and account for its overheads, which can be used as the basis for partitioning the system into optimal number/size of synchronous blocks, and b) automate the synthesis of the global asynchronous communication.Three realistic ASICs, ranging in complexity from 1 to 3 million gates, were used to evaluate GALS benefits and overheads.The results show an average power saving of about 70% in clock with negligible overheads. Lowering power consumption in clock by using Globally AsynchronousLocally Synchronous design style.
Ahmed Hemani, Thomas Meincke, Shashi Kumar, Adam Postula, Thomas Olsson 0001, Peter Nilsson 0001, Johnny Öberg, Peeter Ellervee, Dan Lundqvist
DAC7
1998 Scheduling of Outputs in Grammar-based Hardware Synthesis of Data Communication Protocols
abstract
We present a grammar based specification method for hardware synthesis of data communication protocols in which the specification is independent of the port size. Instead, it is used during the synthesis process as a constraint. When the width of the output assignments exceed the chosen output port width, the assignments are split and scheduled over the available states. We present a solution to this problem and results of applying it to some relevant problems.
Johnny Öberg, Ahmed Hemani
DATE1