VLDB 2026 Research / reviewers in the wild / expert
Joonas Multanen
dblp:173/8991
· DBLP profile ↗
12ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-4438-2031ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beaivi: A 22-nm 1-GHz+ Exposed Datapath RISC-V DSP for Low-Power ApplicationsabstractLow-power digital signal processing is required for edge devices operating in energy-constrained environments. Static multi-issue machines excel in such use cases but lack the required flexibility for maintaining high code density while exploiting instruction-level parallelism. This paper introduces a novel RISC-V-based DSP architecture, "Beaivi", that extends the processor with an exposed datapath multi-issue mode for exploiting instruction-level parallelism efficiently in performance-critical code regions while preserving high code density in noncritical phases with a RISC-V mode. The dynamic code density is further improved by leveraging a dictionary compression method that programs the dictionaries on a loop basis via compiler-driven static analysis. We demonstrate the real-world applicability of the architecture by taping out the processor using a commercial 22-nm technology. The design meets timing at 1.0 GHz and draws 50 mW under a neural network inference workload. Kari Hepola, Joonas Multanen, Väinö-Waltteri Granat, Jakub Zádník, Roope Keskinen, Karri Palovuori, Pekka Jääskeläinen |
DATE | 2 |
| 2026 | Composable Open Source Toolchain for Synthesizing Hardware Accelerators from OpenCL Command BuffersabstractHigh-level synthesis tools enable developers to use high-level programming languages such as C, C++, or OpenCL to design FPGA accelerators, reducing the complexity of hardware design. While OpenCL provides a portable programming model for heterogeneous systems, the existing FPGA toolchains require non-portable program modifications, pragmas, and pre-generation of bitstreams. To this end, we present a composable open source toolchain for automated synthesis of hardware accelerators from OpenCL, bridging the gap between high-level parallel programming and reconfigurable hardware design. Our toolchain supports the standard OpenCL input, and the runtime handles bitstream generation, reconfiguration, and kernel execution. In addition to supporting unmodified OpenCL applications, this work is the first to even partially implement the OpenCL command buffer extension to automatically generate specialized FPGA bitstreams. The toolchain is modular and extensible, supporting multiple HLS tools and a vendor-agnostic FPGA integration flow, demonstrated on FPGAs from two different vendors. Compared to AMD’s OpenCL FPGA implementation, the proposed toolchain achieves equal or better performance on 70% of PolybenchGPU benchmarks using standard OpenCL and on 80% when using the command buffer extension. Topi Leppänen, Leevi Leppänen, Zainab Jamil, Jan Solanti, Joonas Multanen, Pekka Jääskeläinen |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2026 | Headsail: One-Year Tape-Out of a 25-mm2 Linux-Capable RISC-V MPSoCabstractThe Internet-of-Things (IoT) devices feature a broad range of power, memory, and performance requirements. Ultralow-power, low-performance controllers are at one end of the spectrum, while high-performance, power-intensive systems-on-chip (SoCs) are at the other. Heterogeneous and specialized multiprocessor SoC (MPSoC) architectures have emerged as the most effective paradigm for delivering high performance and energy efficiency across a wide range of application workloads. This work introducesHeadsail, an MPSoC application-specific integrated circuit (ASIC) designed by SoC Hub at Tampere University, Finland.Headsailfeatures a 512-KiB primary data buffer, 128 KiB of shared on-chip SRAM, seven CPU cores (including four CVA6 64-bit RISC-V processors), a low power-DDR2 (LP-DDR2) memory controller, two unique chip-to-chip (C2C) interfaces, and shared peripherals.Headsailhas been successfully implemented using a TSMC 22-nm low-power CMOS technology. Testing results show that the samples can support a maximum operating frequency of 1 GHz and achieve a peak performance of 1100 giga operations per second (GOPS), with an implementation area of 25 mm2and a power-consumption range of 64 mW–1.5 W. Matti Käyrä, Thomas Szymkowiak, Antti Rautakoura, Antti Nurmi, Kari Hepola, Henri Lunnikivi, Toni Jääskeläinen, Abdesattar Kalache, Petteri Toivanen, Roope Keskinen, Andreas Stergiopoulos, Väinö-Waltteri Granat, Arto Oinonen, Joonas Multanen, Pekka Jääskeläinen, Karri Palovuori, Timo Hämäläinen 0001, Syed Mohsin Abbas |
IEEE Trans. Very Large Scale Integr. Syst. | 15 |
| 2025 | Automatically Retargeting Hardware and Code Generation for RISC-V Custom InstructionsabstractCustom instruction (CI) set extensions are beneficial for increasing performance and energy efficiency in a set of target applications. For rapid prototyping of these types of application-specific processors, designers leverage hardware (HW)/software (SW) co-design to create hardware implementations and retarget the compiler using a high-level description of the instruction set extension. Ideally, the architecture description should be flexible enough to support both hardware generation and compiler retargeting from the same description format. The challenge with these methods lies in coupling hardware extensions with the processor core, because using microarchitecture-specific interfaces leads to low design reuse and increased verification effort. To mitigate these challenges, we introduce a HW/SW co-design toolset capable of adapting to a user-defined architecture description that captures the instruction set extension semantics. Based on the architecture description, the toolset can both retarget the compiler and generate co-processors interfacing with the Core-V eXtension interface (CV-X-IF) and Rocket custom co-processor interface (RoCC) protocols that are widely used standard interfaces for RISC-V processors. To demonstrate our methods, we integrate the co-processors with two different variations of CVA6 and Rocket core. The resulting execution time reduction is up to 40% on average, with an area overhead of 8% for the CVA6. For the Rocket core, the execution time reduction is 27% with a 6% area overhead. Kari Hepola, Tharaka Ranasinghe Arachchige, Joonas Multanen, Pekka Jääskeläinen |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | Energy-Efficient Exposed Datapath Architecture With a RISC-V Instruction Set ModeabstractTransport triggered architectures (TTAs) follow the static programming model of very long instruction word (VLIW) processors but expose additional information of the processor datapath in the programming interface, which enables low-level code optimizations but results in lower code density. Multi-instruction-set architectures add flexiblity via their ability to switch instruction sets during execution. The added flexibility is interesting for VLIW-style processors because it enables reducing the large instruction stream energy footprint by using an instruction set with enhanced code density in regions with limited opportunities for exploitation of instruction level parallelism. In this article, we introduce a dual instruction-set architecture, ”Dual-IS”, that implements both RISC-V and TTA instruction sets with shared datapath resources by means of a lightweight microcode unit. In order to utilize the flexible architecture automatically, we introduce a compilation method that is able to independently target code for both instruction sets based on static code analysis and a microarchitectural model of the processor. Compared to a single-ISA TTA processor, we were able to lower the instruction stream energy consumption 45% on average in the best design point, which resulted in a total energy consumption reduction of 26% and a 0.4% lower run time. Kari Hepola, Joonas Multanen, Pekka Jääskeläinen |
IEEE Trans. Computers | 2 |
| 2024 | Bitstream Database-Driven FPGA Programming Flow Based on Standard OpenCLabstractField-programmable gate array (FPGA) vendors provide high-level synthesis (HLS) compilers with accompanying OpenCL runtimes to enable easier use of their devices by non-hardware experts. However, the current runtimes provided by the vendors are not OpenCL-compliant, limiting the application portability and making it difficult to integrate FPGA devices in heterogeneous computing platforms. We propose an automated FPGA management tool AFOCL, with a guiding principle that the software programmer should only need to use the standard OpenCL API to manage FPGA acceleration tasks. This improves portability since the same OpenCL program will work on any OpenCL-compliant computation device able to execute the same kernels, including CPUs, GPUs, and FPGAs. The proposed approach is based on pre-optimized FPGA bitstreams implementing well-defined OpenCL built-in kernels. This enables a clean separation of responsibilities between a hardware developer preparing the FPGA bitstreams containing the kernel implementations, a software developer launching computation tasks as OpenCL built-in kernels, and a bitstream distributor providing preoptimized FPGA IPs to end-users. The automated FPGA programming tool fetches bitstream files as needed from the distributor, reconfigures the FPGA, and manages the communication with the accelerator. We demonstrate that it is possible to achieve similar performance as the current FPGA vendor OpenCL implementations, while abstracting all FPGA-specific details from the software programmer. The cross-vendor potential of AFOCL is shown by porting the implementation to FPGAs from two different vendors (AMD and Altera), and to two different FPGA types [PCIe and system-on-chip (SoC)], and controlling all these systems with the same OpenCL host program. Topi Leppänen, Leevi Leppänen, Joonas Multanen, Pekka Jääskeläinen |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | OpenASIP 2.0: Co-Design Toolset for RISC-V Application-Specific Instruction-Set ProcessorsabstractApplication-specific instruction-set processors (ASIPs) are interesting for improving performance or energy-efficiency for a set of applications of interest while supporting flexibility via compiler-supported programmability. In the past years, the open source hardware community has become extremely active, mainly fueled by the massive popularity of the open-standard RISC-V instruction set architecture. However, the community still lacks an open source ASIP co-design tool that supports rapid customization of RISC-V-based processors with an automatically retargetable programming toolchain. To this end, we introduce OpenASIP 2.0: A co-design toolset that is built on top of our earlier ASIP customization toolset work by extending it to support customization of RISC-V-based processors. It enables RTL generation as well as high-level language programming of RISC-V processors with custom instructions. In this paper, in addition to describing the toolset's key technical internals, we demonstrate it with customization cases for AES, CRC and SHA applications. With the example custom instructions easily integrated using the toolset, the run time was reduced by 44% on average compared to the standard RISC-V ISA. The speedups were achieved with a negligible datapath area overhead of 1.5%, and a 1.4% reduction in the maximum clock frequency. Kari Hepola, Joonas Multanen, Pekka Jääskeläinen |
ASAP | 2 |
| 2022 | Energy-Efficient Instruction Delivery in Embedded Systems With Domain Wall MemoryabstractAs performance and energy-efficiency improvements from technology scaling are slowing down, new technologies are being researched in hopes of disrupting results. Domain wall memory (DWM) is an emerging non-volatile technology that promises extreme data density, fast access times and low power consumption. However, DWM access time depends on the memory location distance from access ports, requiring expensive shifting. This causes overheads on performance and energy consumption. In this article, we implement our previously proposed shift-reducing instruction memory placement (SHRIMP) on a RISC-V core in RTL, provide the first thorough evaluation of the control logic required for DWM and SHRIMP and evaluate the effects on system energy and energy-efficiency. SHRIMP reduces the number of shifts by 36% on average compared to a linear placement in CHStone and Coremark benchmark suites when evaluated on the RISC-V processor system. The reduced shift amount leads to an average reduction of 14% in cycle counts compared to the linear placement. When compared to an SRAM-based system, although increasing memory usage by 26%, DWM with SHRIMP allows a 73% reduction in memory energy and 42% relative energy delay product. We estimate overall energy reductions of 14%, 15% and 19% in three example embedded systems. Joonas Multanen, Kari Hepola, Asif Ali Khan, Jerónimo Castrillón, Pekka Jääskeläinen |
IEEE Trans. Computers | 1 |
| 2020 | Programmable Dictionary Code Compression for Instruction Stream Energy EfficiencyabstractWe propose a novel instruction compression scheme based on fine-grained programmable dictionaries. In its core is a compile-time region-based control flow analysis to selectively update the dictionary contents at runtime, minimizing the update overheads, while maximizing the beneficial use of the dictionary slots. Unlike in the previous work, our approach selects regions of instructions to compress at compile time and changes dictionary contents in a fine-grained manner at runtime with the primary goal of reducing the energy footprint of the processor instruction stream. The proposed instruction compression scheme is evaluated using RISC-V as an example instruction set architecture. The energy savings are compared to an instruction scratch pad and a filter cache as the next level storage. The method reduces instruction stream energy consumption up to 21 % and 5.5 % on average when compared to the RISC-V C extension with a 1% runtime overhead and a negligible hardware overhead. The previous state-of-the-art programmable dictionary compression method provides a slightly better compression ratio, but induces about 30 % runtime overhead. Joonas Multanen, Kari Hepola, Pekka Jääskeläinen |
ICCD | 1 |
| 2019 | SHRIMP: Efficient Instruction Delivery with Domain Wall MemoryabstractDomain Wall Memory (DWM) is a promising emerging memory technology but suffers from the expensive shifts needed to align memory locations with access ports. Previous work on DWM concentrates on data, while, to the best of our knowledge, techniques to specifically target instruction streams have not yet been studied. In this paper, we propose Shift-Reducing Instruction Memory Placement (SHRIMP), the first instruction placement strategy suited for DWM which is accompanied with a supporting instruction fetch and memory architecture. The proposed approach reduces the number of shifts by 40% in the best case with a small memory overhead. In addition, SHRIMP achieves a best case of 23% reduction in total cycle counts. Joonas Multanen, Pekka Jääskeläinen, Asif Ali Khan, Fazal Hameed, Jerónimo Castrillón |
ISLPED | 1 |
| 2019 | LordCore: Energy-Efficient OpenCL-Programmable Software-Defined Radio CoprocessorabstractThis paper proposes a single instruction multiple data (SIMD) processor, which is programmed with high-level OpenCL language. The low-power processor is customized for executing multiple-input-multiple-output (MIMO) detection algorithms at a high performance while consuming very little power making it suitable for software-defined radio (SDR) applications. The novel combination of SIMD operations on a transport programmed multicore datapath allows saving power on both the execution front end and the back end, leading to very good energy efficiency with a compiler programmable design. We demonstrate the feasibility of the architecture with the layered orthogonal lattice detector and minimum mean-square-error MIMO algorithms, which can be used as a software-defined radio implementation of the 3GPP local thermal equilibrium r11 standard. Compared to other state-of-the-art SDR architectures, the proposed design adds features that improve programmer productivity with an insignificant power and area impact. Heikki Kultala, Timo Viitanen, Heikki Berg, Pekka Jääskeläinen, Joonas Multanen, Mikko Kokkonen, Kalle Raiskila, Tommi Zetterman, Jarmo Takala |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2018 | Instantaneous foveated preview for progressive Monte Carlo renderingabstractProgressive rendering, for example Monte Carlo rendering of 360° content for virtual reality headsets, is a time-consuming task. If the 3D artist notices an error while previewing the rendering, they must return to editing mode, make the required changes, and restart rendering. We propose the use of eye-tracking-based optimization to significantly speed up previewing of the artist’s points of interest. The speed of the preview is further improved by sampling with a distribution that closely follows the experimentally measured visual acuity of the human eye, unlike the piecewise linear models used in previous work. In a comprehensive user study, the perceived convergence of our proposed method was 10 times faster than that of a conventional preview, and often appeared to be instantaneous. In addition, the participants rated the method to have only marginally more artifacts in areas where it had to start rendering from scratch, compared to conventional rendering methods that had already generated image content in those areas. Matias Koskela, Kalle Immonen, Timo Viitanen, Pekka Jääskeläinen, Joonas Multanen, Jarmo Takala |
Comput. Vis. Media | 5 |