EDBT 2026 Demo / reviewers in the wild / expert
Pekka Jääskeläinen
dblp:72/2938
· DBLP profile ↗
44ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0001-5707-8544ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 since 2021Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beaivi: A 22-nm 1-GHz+ Exposed Datapath RISC-V DSP for Low-Power ApplicationsabstractLow-power digital signal processing is required for edge devices operating in energy-constrained environments. Static multi-issue machines excel in such use cases but lack the required flexibility for maintaining high code density while exploiting instruction-level parallelism. This paper introduces a novel RISC-V-based DSP architecture, "Beaivi", that extends the processor with an exposed datapath multi-issue mode for exploiting instruction-level parallelism efficiently in performance-critical code regions while preserving high code density in noncritical phases with a RISC-V mode. The dynamic code density is further improved by leveraging a dictionary compression method that programs the dictionaries on a loop basis via compiler-driven static analysis. We demonstrate the real-world applicability of the architecture by taping out the processor using a commercial 22-nm technology. The design meets timing at 1.0 GHz and draws 50 mW under a neural network inference workload. Kari Hepola, Joonas Multanen, Väinö-Waltteri Granat, Jakub Zádník, Roope Keskinen, Karri Palovuori, Pekka Jääskeläinen |
DATE | 7 |
| 2026 | Composable Open Source Toolchain for Synthesizing Hardware Accelerators from OpenCL Command BuffersabstractHigh-level synthesis tools enable developers to use high-level programming languages such as C, C++, or OpenCL to design FPGA accelerators, reducing the complexity of hardware design. While OpenCL provides a portable programming model for heterogeneous systems, the existing FPGA toolchains require non-portable program modifications, pragmas, and pre-generation of bitstreams. To this end, we present a composable open source toolchain for automated synthesis of hardware accelerators from OpenCL, bridging the gap between high-level parallel programming and reconfigurable hardware design. Our toolchain supports the standard OpenCL input, and the runtime handles bitstream generation, reconfiguration, and kernel execution. In addition to supporting unmodified OpenCL applications, this work is the first to even partially implement the OpenCL command buffer extension to automatically generate specialized FPGA bitstreams. The toolchain is modular and extensible, supporting multiple HLS tools and a vendor-agnostic FPGA integration flow, demonstrated on FPGAs from two different vendors. Compared to AMD’s OpenCL FPGA implementation, the proposed toolchain achieves equal or better performance on 70% of PolybenchGPU benchmarks using standard OpenCL and on 80% when using the command buffer extension. Topi Leppänen, Leevi Leppänen, Zainab Jamil, Jan Solanti, Joonas Multanen, Pekka Jääskeläinen |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2026 | Scalable Hole-Filling for Real-Time Multi-GPU Light Field Path TracingabstractLight field (LF) displays address the mismatch in focus cues present in traditional displays by triggering natural defocus blur and enabling motion parallax. They rely on geometrical optics, displaying rays from multiple angles of view. LF path tracing is computationally expensive for real-time applications, since it requires rendering multiple views. To reduce this computational complexity, spatially reprojecting pixels between views is commonly performed. Reusing pixels that are already rendered is cheaper than path tracing additional ones. However, when occluded areas are uncovered in some views, reprojection is not possible, creating holes in these views. Filling-in the holes requires extra path tracing computation. This paper investigates scalable hole-filling strategies for LF path tracing, using multiple GPUs to reach real-time performance. We propose an algorithm search optimization procedure to determine whether a specific assignment algorithm can be generalized across scenes, using the hole-filling time as a minimization function. In addition, we introduce DaSH (Discarded and Subsampled path traced Hole-filling), a novel method that reduces computation and divergence overhead in fixed-size hardware thread-blocks. Based on local pixel sparsity within pixel patches, DaSH adaptively subsamples and discards hole-filling rays. Our evaluation demonstrates that DaSH achieves significant performance gains while preserving the visual and structural quality of refocused light field images at the retina plane. The experiments demonstrate an average speedup factor of $1.8\times$1.8× for DaSH, compared to prior work, in a multi-GPU rendering system. Erwan Leria, Markku Mäkitalo, Pekka Jääskeläinen, Mårten Sjöström |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | Headsail: One-Year Tape-Out of a 25-mm2 Linux-Capable RISC-V MPSoCabstractThe Internet-of-Things (IoT) devices feature a broad range of power, memory, and performance requirements. Ultralow-power, low-performance controllers are at one end of the spectrum, while high-performance, power-intensive systems-on-chip (SoCs) are at the other. Heterogeneous and specialized multiprocessor SoC (MPSoC) architectures have emerged as the most effective paradigm for delivering high performance and energy efficiency across a wide range of application workloads. This work introducesHeadsail, an MPSoC application-specific integrated circuit (ASIC) designed by SoC Hub at Tampere University, Finland.Headsailfeatures a 512-KiB primary data buffer, 128 KiB of shared on-chip SRAM, seven CPU cores (including four CVA6 64-bit RISC-V processors), a low power-DDR2 (LP-DDR2) memory controller, two unique chip-to-chip (C2C) interfaces, and shared peripherals.Headsailhas been successfully implemented using a TSMC 22-nm low-power CMOS technology. Testing results show that the samples can support a maximum operating frequency of 1 GHz and achieve a peak performance of 1100 giga operations per second (GOPS), with an implementation area of 25 mm2and a power-consumption range of 64 mW–1.5 W. Matti Käyrä, Thomas Szymkowiak, Antti Rautakoura, Antti Nurmi, Kari Hepola, Henri Lunnikivi, Toni Jääskeläinen, Abdesattar Kalache, Petteri Toivanen, Roope Keskinen, Andreas Stergiopoulos, Väinö-Waltteri Granat, Arto Oinonen, Joonas Multanen, Pekka Jääskeläinen, Karri Palovuori, Timo Hämäläinen 0001, Syed Mohsin Abbas |
IEEE Trans. Very Large Scale Integr. Syst. | 16 |
| 2025 | Open Software Stack for Compression-Aware Adaptive Edge OffloadingabstractOffloading a computationally complex task from an edge device can improve latency and its battery life. The additional network transfers increase power consumption and latency, but can be mitigated with compression at the cost of additional computation and distortion. Thus, a balance between compression efficiency and complexity must be found and maintained as the network conditions change. In this paper, we propose an open -source edge offloading software stack that decides between local and remote task execution and chooses the optimal compression method based on continuously monitored system metrics under user-defined constraints. We evaluate the system offloading a semantic segmentation task from a smartphone over WiFi-6 and 5G networks, using latency, intersection over union (IoU), and power consumption metrics. Portability, multi-tenancy, and granular profiling are achieved by leveraging the PoCL-R OpenCL implementation. In simulated network impairments, dynamically selecting the compression strategy achieves 2.1-10.7% average latency improvement but maintains the highest possible quality when network conditions allow meeting the latency budget. If the computational overhead of compression surpasses the network transfer overhead, the system can transmit images uncompressed. Field measurements under network impairments confirm the usability of the system and its ability to fall back to local execution. Jakub Zádník, Robin Bijl, Jan Solanti, Erno Joensuu, Markku Mäkitalo, Pekka Jääskeläinen |
WCNC | 6 |
| 2025 | CV-Cast: Computer Vision-Oriented Linear Coding and TransmissionabstractRemote inference allows lightweight edge devices, such as autonomous drones, to perform vision tasks exceeding their computational, energy, or processing delay budget. In such applications, reliable transmission of information is challenging due to high variations of channel quality. Traditional approaches involving spatio-temporal transforms, quantization, and entropy coding followed by digital transmission may be affected by a sudden decrease in quality (thedigital cliff) when the channel quality is less than expected during design. This problem can be addressed by using Linear Coding and Transmission (LCT), a joint source and channel coding scheme relying on linear operators only, allowing to achieve reconstructed per-pixel error commensurate with the wireless channel quality. In this paper, we propose CV-Cast: the first LCT scheme optimized for computer vision task accuracy instead of per-pixel distortion. Using this approach, for instance at 10 dB channel signal-to-noise ratio, CV-Cast requires transmitting 28% less symbols than a baseline LCT scheme in semantic segmentation and 15% in object detection tasks. Simulations involving a realistic 5G channel model confirm the smooth decrease in accuracy achieved with CV-Cast, while images encoded by JPEG or learned image coding (LIC) and transmitted using classical schemes at low Eb/N0 are subject to digital cliff. Jakub Zádník, Michel Kieffer, Anthony Trioux, Markku Mäkitalo, Pekka Jääskeläinen |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Correction to "CV-Cast: Computer Vision-Oriented Linear Coding and Transmission"abstractIn the above article [1], on page 1151, eq. (6), there is an error in the equation. The correct equation is: \begin{equation*} \min.\,\,D,\,\,\text{s.t.} \sum\limits_{k = 1}^K {{{\lambda }_k}\beta _k^2 \leqslant P.} \tag{6} \end{equation*} min.D,s.t.∑k=1Kλkβk2⩽P.(6) Jakub Zádník, Michel Kieffer, Anthony Trioux, Markku Mäkitalo, Pekka Jääskeläinen |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Automatically Retargeting Hardware and Code Generation for RISC-V Custom InstructionsabstractCustom instruction (CI) set extensions are beneficial for increasing performance and energy efficiency in a set of target applications. For rapid prototyping of these types of application-specific processors, designers leverage hardware (HW)/software (SW) co-design to create hardware implementations and retarget the compiler using a high-level description of the instruction set extension. Ideally, the architecture description should be flexible enough to support both hardware generation and compiler retargeting from the same description format. The challenge with these methods lies in coupling hardware extensions with the processor core, because using microarchitecture-specific interfaces leads to low design reuse and increased verification effort. To mitigate these challenges, we introduce a HW/SW co-design toolset capable of adapting to a user-defined architecture description that captures the instruction set extension semantics. Based on the architecture description, the toolset can both retarget the compiler and generate co-processors interfacing with the Core-V eXtension interface (CV-X-IF) and Rocket custom co-processor interface (RoCC) protocols that are widely used standard interfaces for RISC-V processors. To demonstrate our methods, we integrate the co-processors with two different variations of CVA6 and Rocket core. The resulting execution time reduction is up to 40% on average, with an area overhead of 8% for the CVA6. For the Rocket core, the execution time reduction is 27% with a 6% area overhead. Kari Hepola, Tharaka Ranasinghe Arachchige, Joonas Multanen, Pekka Jääskeläinen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Dynamic load balancing for real-time multiview path tracing on multi-GPU architecturesabstractStereoscopic and multiview rendering are used for virtual reality and the synthetic generation of light fields from three-dimensional scenes. Because rendering multiple views using ray tracing techniques is computationally expensive, the utilization of multiprocessor machines is necessary to achieve real-time frame rates. In this study, we propose a dynamic load-balancing algorithm for real-time multiview path tracing on multi-compute device platforms. The proposed algorithm was adapted to heterogeneous hardware combinations and dynamic scenes in real time. We show that on a heterogeneous dual-GPU platform, our implementation reduces the rendering time by an average of approximately 30%–50% compared with that of a uniform workload distribution, depending on the scene and number of views. Erwan Leria, Markku Mäkitalo, Julius Ikkala, Pekka Jääskeläinen |
Virtual Real. Intell. Hardw. | 4 |
| 2024 | Interactive Multi-GPU Light Field Path Tracing Using Multi-Source Spatial ReprojectionabstractPath tracing combined with multiview displays enables progress towards achieving ultrarealistic virtual reality. However, multiview displays based on light field technology impose a heavy workload for real-time graphics due to the large number of views to be rendered. In order to achieve low latency performance, computational effort can be reduced by path tracing only some views (source views), and synthesizing the remaining views (target views) through spatial reprojection, which reuses path traced pixels from source views to target views. Deciding the number of source views with respect to the computational resources is not trivial, since spatial reprojection introduces dependencies in the otherwise trivially parallel rendering pipeline and path tracing multiple source views increases the computation time. Erwan Leria, Markku Mäkitalo, Pekka Jääskeläinen, Mårten Sjöström |
VRST | 3 |
| 2024 | Performance of Linear Coding and Transmission in Low-Latency Computer Vision OffloadingabstractImage communication increasingly involves machine-to-machine delivery. For example, images acquired by an autonomous drone can be compressed and sent to an edge server over a wireless network for resource-intensive processing. Traditional compression techniques involving transform, quantization, and entropy coding reach high compression efficiency, but channel conditions worse than expected may lead to a sharp decrease in the decoded image quality. As an alternative, Linear Coding and Transmission (LCT) systems have been proposed to avoid this digital cliff problem: The reconstructed image quality decreases gradually as channel conditions degrade. This paper presents a comprehensive evaluation of computer vision tasks with input images processed and transmitted using LCT. It also analyses the benefits of network retraining, accounting for impairments due to LCT and noisy channel. Considering object detection and semantic segmentation over images transmitted and received by LCT systems, we show that the task accuracy degrades smoothly when the channel quality decreases, avoiding the cliff effect. Retraining with noisy images processed by LCT restores detection mAP degradation from 23.8% to 4.4% and segmentation mIoU degradation from 43.2% to 8.1 % when the channel signal-to-noise ratio is 10 dB. Jakub Zádník, Anthony Trioux, Michel Kieffer, Markku Mäkitalo, François-Xavier Coudoux, Patrick Corlay, Pekka Jääskeläinen |
WCNC | 7 |
| 2024 | Energy-Efficient Exposed Datapath Architecture With a RISC-V Instruction Set ModeabstractTransport triggered architectures (TTAs) follow the static programming model of very long instruction word (VLIW) processors but expose additional information of the processor datapath in the programming interface, which enables low-level code optimizations but results in lower code density. Multi-instruction-set architectures add flexiblity via their ability to switch instruction sets during execution. The added flexibility is interesting for VLIW-style processors because it enables reducing the large instruction stream energy footprint by using an instruction set with enhanced code density in regions with limited opportunities for exploitation of instruction level parallelism. In this article, we introduce a dual instruction-set architecture, ”Dual-IS”, that implements both RISC-V and TTA instruction sets with shared datapath resources by means of a lightweight microcode unit. In order to utilize the flexible architecture automatically, we introduce a compilation method that is able to independently target code for both instruction sets based on static code analysis and a microarchitectural model of the processor. Compared to a single-ISA TTA processor, we were able to lower the instruction stream energy consumption 45% on average in the best design point, which resulted in a total energy consumption reduction of 26% and a 0.4% lower run time. Kari Hepola, Joonas Multanen, Pekka Jääskeläinen |
IEEE Trans. Computers | 3 |
| 2024 | R-Blocks: an Energy-Efficient, Flexible, and Programmable CGRAabstractEmerging data-driven applications in the embedded, e-Health, and internet of things (IoT) domain require complex on-device signal analysis and data reduction to maximize energy efficiency on these energy-constrained devices. Coarse-grained reconfigurable architectures (CGRAs) have been proposed as a good compromise between flexibility and energy efficiency for ultra-low power (ULP) signal processing. Existing CGRAs are often specialized and domain-specific or can only accelerate simple kernels, which makes accelerating complete applications on a CGRA while maintaining high energy efficiency an open issue. Moreover, the lack of instruction set architecture (ISA) standardization across CGRAs makes code generation using current compiler technology a major challenge. This work introduces R-Blocks; a ULP CGRA with HW/SW co-design tool-flow based on the OpenASIP toolset. This CGRA is extremely flexible due to its well-established VLIW-SIMD execution model and support for flexible SIMD-processing, while maintaining an extremely high energy efficiency using software bypassing, optimized instruction delivery, and local scratchpad memories. R-Blocks is synthesized in a commercial 22-nm FD-SOI technology and achieves a full-system energy efficiency of 115 MOPS/mW on a common FFT benchmark, 1.45× higher than a highly tuned embedded RISC-V processor. Comparable energy efficiency is obtained on multiple complex workloads, making R-Blocks a promising acceleration target for general-purpose computing. Barry de Bruin, Kanishkan Vadivel, Mark Wijtvliet, Pekka Jääskeläinen, Henk Corporaal |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2024 | Bitstream Database-Driven FPGA Programming Flow Based on Standard OpenCLabstractField-programmable gate array (FPGA) vendors provide high-level synthesis (HLS) compilers with accompanying OpenCL runtimes to enable easier use of their devices by non-hardware experts. However, the current runtimes provided by the vendors are not OpenCL-compliant, limiting the application portability and making it difficult to integrate FPGA devices in heterogeneous computing platforms. We propose an automated FPGA management tool AFOCL, with a guiding principle that the software programmer should only need to use the standard OpenCL API to manage FPGA acceleration tasks. This improves portability since the same OpenCL program will work on any OpenCL-compliant computation device able to execute the same kernels, including CPUs, GPUs, and FPGAs. The proposed approach is based on pre-optimized FPGA bitstreams implementing well-defined OpenCL built-in kernels. This enables a clean separation of responsibilities between a hardware developer preparing the FPGA bitstreams containing the kernel implementations, a software developer launching computation tasks as OpenCL built-in kernels, and a bitstream distributor providing preoptimized FPGA IPs to end-users. The automated FPGA programming tool fetches bitstream files as needed from the distributor, reconfigures the FPGA, and manages the communication with the accelerator. We demonstrate that it is possible to achieve similar performance as the current FPGA vendor OpenCL implementations, while abstracting all FPGA-specific details from the software programmer. The cross-vendor potential of AFOCL is shown by porting the implementation to FPGAs from two different vendors (AMD and Altera), and to two different FPGA types [PCIe and system-on-chip (SoC)], and controlling all these systems with the same OpenCL host program. Topi Leppänen, Leevi Leppänen, Joonas Multanen, Pekka Jääskeläinen |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | BrainTTA: A 28.6 TOPS/W Compiler Programmable Transport-Triggered NN SoCabstractAccelerators designed for deep neural network (DNN) inference with extremely low operand widths, down to 1-bit, have become popular due to their ability to significantly reduce energy consumption during inference. This paper introduces a compiler-programmable flexible System-on-Chip (SoC) with mixed-precision support. This SoC is based on a Transport-Triggered Architecture (TTA) that facilitates efficient implementation of DNN workloads. By shifting the complexity of data movement from the hardware scheduler to the exposed-datapath compiler, DNN workloads can be implemented in an energy efficient yet flexible way. The architecture is fully supported by a compiler and can be programmed using C/C++/OpenCL. The SoC is implemented using 22nm FDX technology and achieves a peak energy efficiency of 28.6/14.9/2.47 TOPS/W for binary, ternary, and 8-bit precision, respectively, while delivering a throughput of 614/307/77 GOPS. Compared to state-of-the-art (SotA), this work achieves up to 3.3x better energy efficiency compared to other programmable solutions. Maarten Molendijk, Floran de Putter, Manil Dev Gomony, Pekka Jääskeläinen, Henk Corporaal |
ICCD | 4 |
| 2022 | OpenASIP 2.0: Co-Design Toolset for RISC-V Application-Specific Instruction-Set ProcessorsabstractApplication-specific instruction-set processors (ASIPs) are interesting for improving performance or energy-efficiency for a set of applications of interest while supporting flexibility via compiler-supported programmability. In the past years, the open source hardware community has become extremely active, mainly fueled by the massive popularity of the open-standard RISC-V instruction set architecture. However, the community still lacks an open source ASIP co-design tool that supports rapid customization of RISC-V-based processors with an automatically retargetable programming toolchain. To this end, we introduce OpenASIP 2.0: A co-design toolset that is built on top of our earlier ASIP customization toolset work by extending it to support customization of RISC-V-based processors. It enables RTL generation as well as high-level language programming of RISC-V processors with custom instructions. In this paper, in addition to describing the toolset's key technical internals, we demonstrate it with customization cases for AES, CRC and SHA applications. With the example custom instructions easily integrated using the toolset, the run time was reduced by 44% on average compared to the standard RISC-V ISA. The speedups were achieved with a negligible datapath area overhead of 1.5%, and a 1.4% reduction in the maximum clock frequency. Kari Hepola, Joonas Multanen, Pekka Jääskeläinen |
ASAP | 3 |
| 2022 | Real-Time Light Field Path Tracing
Markku Mäkitalo, Erwan Leria, Julius Ikkala, Pekka Jääskeläinen |
CGI | 4 |
| 2022 | Prebypass: Software Register File Bypassing for Reduced Interconnection ArchitecturesabstractExposed Datapath Architectures (EDPAs) with aggressively pruned data-path connectivity, where not all function units in the design have connections to a centralized register file, are promising solutions for energy-efficient computation. A direct bypassing of data between function units without temporary copies to the register file is a prime optimization for programming such architectures. However, traditional compiler frameworks, such as LLVM, assume function-units connect to register-files and allocate all live variables in register-files. This leads to schedule inefficiencies in terms of instruction-level parallelism and reg-ister accesses in the EDPAs. To address these inefficiencies, we propose Prebypass; a new optimization pass for EDPA compiler backends. Experimental results on an EDPA class of architecture, Transport- Triggered Architecture, show that Prebypass improves the runtime, register reads, and register writes up to 16%, 26 %, and 37 % respectively, when the datapath is extremely pruned. Evaluation in a 28-nm FDSOI technology reveals that Prebypass improves the core-level Energy by 17.5 % over the current heuristic scheduler. Kanishkan Vadivel, Barry de Bruin, Roel Jordans, Henk Corporaal, Pekka Jääskeläinen |
DSD | 5 |
| 2022 | Pruned Lightweight Encoders for Computer VisionabstractLatency-critical computer vision systems, such as autonomous driving or drone control, require fast image or video compression when offloading neural network inference to a remote computer. To ensure low latency on a near-sensor edge device, we propose the use of lightweight encoders with constant bitrate and pruned encoding configurations, namely, ASTC and JPEG XS. Pruning introduces significant distortion which we show can be recovered by retraining the neural network with compressed data after decompression. Such an approach does not modify the network architecture or require coding format modifications. By retraining with compressed datasets, we reduced the classification accuracy and segmentation mean intersection over union (mIoU) degradation due to ASTC compression to 4.9-5.0 percentage points (pp) and 4.4-4.0 pp, respectively. With the same method, the mIoU lost due to JPEG XS compression at the main profile was restored to 2.7-2.3 pp. In terms of encoding speed, our ASTC encoder implementation is 2.3x faster than JPEG. Even though the JPEG XS reference encoder requires optimizations to reach low latency, we showed that disabling significance flag coding saves 22–23% of encoding time at the cost of 0.4-0.3 mIoU after retraining. Jakub Zádník, Markku Mäkitalo, Pekka Jääskeläinen |
MMSP | 3 |
| 2022 | Energy-Efficient Instruction Delivery in Embedded Systems With Domain Wall MemoryabstractAs performance and energy-efficiency improvements from technology scaling are slowing down, new technologies are being researched in hopes of disrupting results. Domain wall memory (DWM) is an emerging non-volatile technology that promises extreme data density, fast access times and low power consumption. However, DWM access time depends on the memory location distance from access ports, requiring expensive shifting. This causes overheads on performance and energy consumption. In this article, we implement our previously proposed shift-reducing instruction memory placement (SHRIMP) on a RISC-V core in RTL, provide the first thorough evaluation of the control logic required for DWM and SHRIMP and evaluate the effects on system energy and energy-efficiency. SHRIMP reduces the number of shifts by 36% on average compared to a linear placement in CHStone and Coremark benchmark suites when evaluated on the RISC-V processor system. The reduced shift amount leads to an average reduction of 14% in cycle counts compared to the linear placement. When compared to an SRAM-based system, although increasing memory usage by 26%, DWM with SHRIMP allows a 73% reduction in memory energy and 42% relative energy delay product. We estimate overall energy reductions of 14%, 15% and 19% in three example embedded systems. Joonas Multanen, Kari Hepola, Asif Ali Khan, Jerónimo Castrillón, Pekka Jääskeläinen |
IEEE Trans. Computers | 5 |
| 2021 | DDISH-GI: Dynamic Distributed Spherical Harmonics Global Illumination
Julius Ikkala, Petrus E. J. Kivi, Joel Alanko, Markku Mäkitalo, Pekka Jääskeläinen |
CGI | 5 |
| 2020 | Design and management of image processing pipelines within CPS: 2 years of experience from the FitOptiVis ECSEL ProjectabstractCyber-Physical Systems (CPS) are dynamic and reactive systems interacting with processes, environment and, sometimes, humans. They are often distributed with sensors and actuators, smart, adaptive, predictive and react in real-time. Indeed, as sight for human beings, image- and video-processing pipelines are a prime source for environmental information for systems allowing them to take better decisions according to what they see. Therefore, in FitOptiVis we are developing novel methods and tools to integrate complex image and video processing pipelines. FitOptiVis aims to deliver a reference architecture for describing and optimizing quality and resource management for imaging and video pipelines in CPS both at design- and run-time. The architecture is concretized in low-power, high-performance, smart components, and in methods and tools for combined design-time and run-time multi-objective optimization and adaptation within system and environment constraints. Luigi Pomante, Francesca Palumbo, Claudia Rinaldi, Giacomo Valente, Carlo Sau, Tiziana Fanni, Frank van der Linden 0001, Twan Basten, Marc Geilen, Geran Peeren, Jirí Kadlec, Pekka Jääskeläinen, Marcos Martinez de Alejandro, Jukka Saarinen, Tero Säntti, Maria Katiuscia Zedda, Victor Sanchez, Dip Goswami, Zaid Al-Ars, Ad de Beer |
DSD | 12 |
| 2020 | TTA-SIMD Soft Core ProcessorsabstractSoft processors are an important tool in the Field Programmable Gate Array (FPGA) designer's toolkit, and their Single Instruction Multiple Data (SIMD) organizations are an efficient means to utilize the parallelism of FPGAs. However, the state-of-the-art SIMD processors are hindered by the additional logic complexity resulting from dynamic features. By minimizing such constructs, it is possible to design soft processors that are efficient but still flexible enough to operate within an application domain. To this end, we propose a family of instruction set programmable multi-issue wide SIMD soft cores. The template is based on a highly static Transport Triggered Architecture (TTA) and a design time customizable shuffle unit to minimize inefficient dynamic features while remaining compiler programmable. The cores are evaluated on the PYNQ-Z1 board against the ARM A9 hard processor system with NEON vector extensions. The proposed cores reach up to 2.4x performance improvement over the ARM, can fit up to 1024 bit wide SIMD units onto the relatively small FPGA, while still operating at above 100 MHz. The scalability of TTA enables state of the art vector widths. The multicore scalability of the template is preliminarily tested with a 14-core design on a XCZU9EG FPGA customized for real-time convolutional neural net inference. Kati Tervo, Samawat Malik, Topi Leppänen, Pekka Jääskeläinen |
FPL | 4 |
| 2020 | Programmable Dictionary Code Compression for Instruction Stream Energy EfficiencyabstractWe propose a novel instruction compression scheme based on fine-grained programmable dictionaries. In its core is a compile-time region-based control flow analysis to selectively update the dictionary contents at runtime, minimizing the update overheads, while maximizing the beneficial use of the dictionary slots. Unlike in the previous work, our approach selects regions of instructions to compress at compile time and changes dictionary contents in a fine-grained manner at runtime with the primary goal of reducing the energy footprint of the processor instruction stream. The proposed instruction compression scheme is evaluated using RISC-V as an example instruction set architecture. The energy savings are compared to an instruction scratch pad and a filter cache as the next level storage. The method reduces instruction stream energy consumption up to 21 % and 5.5 % on average when compared to the RISC-V C extension with a 1% runtime overhead and a negligible hardware overhead. The previous state-of-the-art programmable dictionary compression method provides a slightly better compression ratio, but induces about 30 % runtime overhead. Joonas Multanen, Kari Hepola, Pekka Jääskeläinen |
ICCD | 3 |
| 2019 | The FitOptiVis ECSEL project: highly efficient distributed embedded image/video processing in cyber-physical systemsabstractCyber-Physical Systems (CPS) are systems that are in feedback with their environment, possibly with humans in the loop. They are often distributed with sensors and actuators, smart, adaptive and predictive and react in real-time. Image- and video-processing pipelines are a prime source for environmental information improving the possibilities of active, relevant feedback. In such a context, FitOptiVis aims to provide end-to-end multi-objective optimization for imaging and video pipelines of CPS, with emphasis on energy and performance, leveraging on a reference architecture, supported by low-power, high-performance, smart devices, and by methods and tools for combined design-time and run-time multi-objective optimization within system and environment constraints. Zaid Al-Ars, Twan Basten, Ad de Beer, Marc Geilen, Dip Goswami, Pekka Jääskeläinen, Jirí Kadlec, Marcos Martinez de Alejandro, Francesca Palumbo, Geran Peeren, Luigi Pomante, Frank van der Linden 0001, Jukka Saarinen, Tero Säntti, Carlo Sau, Maria Katiuscia Zedda |
CF | 6 |
| 2019 | AEx: Automated Customization of Exposed Datapath Soft-CoresabstractHigh-level synthesis tools aim to produce hardware designs out of software descriptions with a goal to lower the bar in FPGA usage for software engineers. Despite their recent progress, however, HLS tools still require FPGA target specific pragmas and other modifications to the originally processor-targeting source code descriptions. Customized soft core based overlay architectures provide a software programmable layer on top of the FPGA fabric. The benefit of this approach is that a platform independent compiler target is presented to the programs, which lowers the porting burden, and online repurposing the same configuration is natural by just switching the executed program. The main drawback, like with any overlay architecture, are the additional implementation overheads the overlay imposes to the resource consumption and the maximum operating frequency. In this paper we show how by utilizing the efficient structure of Transport-Triggered Architectures (TTA), soft-cores can be customized automatically to benefit from the flexible FPGA fabric while still presenting a comfortable software layer to the users. The results compared to previously published non-specialized TTA soft cores indicate equal or better execution times, while the program image size is reduced by up to 49%, and overall resource utilization improved from 10% to 60%. Alex Hirvonen, Kati Tervo, Heikki Kultala, Pekka Jääskeläinen |
DSD | 4 |
| 2019 | SHRIMP: Efficient Instruction Delivery with Domain Wall MemoryabstractDomain Wall Memory (DWM) is a promising emerging memory technology but suffers from the expensive shifts needed to align memory locations with access ports. Previous work on DWM concentrates on data, while, to the best of our knowledge, techniques to specifically target instruction streams have not yet been studied. In this paper, we propose Shift-Reducing Instruction Memory Placement (SHRIMP), the first instruction placement strategy suited for DWM which is accompanied with a supporting instruction fetch and memory architecture. The proposed approach reduces the number of shifts by 40% in the best case with a small memory overhead. In addition, SHRIMP achieves a best case of 23% reduction in total cycle counts. Joonas Multanen, Pekka Jääskeläinen, Asif Ali Khan, Fazal Hameed, Jerónimo Castrillón |
ISLPED | 2 |
| 2019 | Towards Efficient Code Generation for Exposed Datapath ArchitecturesabstractCoarse-grained reconfigurable architectures and other exposed datapath architectures such as transport-triggered architectures come with a high energy efficiency promise for accelerating data oriented workloads. Their main drawback results from the push of complexity from the architecture to the programmer; compiler techniques that allow starting from a higher-level programming language and generate code efficiently to such architectures robustly is still an open research area. In this article we survey the known main sources of challenges and outline a generic processor architecture template that covers the most common architecture variations along with a proposal for a common code generation framework for such challenging architectures. Kanishkan Vadivel, Roel Jordans, Sander Stuijk, Henk Corporaal, Pekka Jääskeläinen, Heikki Kultala |
SCOPES | 5 |
| 2019 | Blockwise Multi-Order Feature Regression for Real-Time Path-Tracing ReconstructionabstractPath tracing produces realistic results including global illumination using a unified simple rendering pipeline. Reducing the amount of noise to imperceptible levels without post-processing requires thousands of samples per pixel (spp), while currently it is only possible to render extremely noisy 1 spp frames in real time with desktop GPUs. However, post-processing can utilize feature buffers, which contain noise-free auxiliary data available in the rendering pipeline. Previously, regression-based noise filtering methods have only been used in offline rendering due to their high computational cost. In this article we propose a novel regression-based reconstruction pipeline, called Blockwise Multi-Order Feature Regression (BMFR), tailored for path-traced 1 spp inputs that runs in real time. The high speed is achieved with a fast implementation of augmented QR factorization and by using stochastic regularization to address rank-deficient feature data. The proposed algorithm is 1.8× faster than the previous state-of-the-art real-time path-tracing reconstruction method while producing better quality frame sequences. Matias Koskela, Kalle Immonen, Markku Mäkitalo, Alessandro Foi, Timo Viitanen, Pekka Jääskeläinen, Heikki Kultala, Jarmo Takala |
ACM Trans. Graph. | 6 |
| 2019 | LordCore: Energy-Efficient OpenCL-Programmable Software-Defined Radio CoprocessorabstractThis paper proposes a single instruction multiple data (SIMD) processor, which is programmed with high-level OpenCL language. The low-power processor is customized for executing multiple-input-multiple-output (MIMO) detection algorithms at a high performance while consuming very little power making it suitable for software-defined radio (SDR) applications. The novel combination of SIMD operations on a transport programmed multicore datapath allows saving power on both the execution front end and the back end, leading to very good energy efficiency with a compiler programmable design. We demonstrate the feasibility of the architecture with the layered orthogonal lattice detector and minimum mean-square-error MIMO algorithms, which can be used as a software-defined radio implementation of the 3GPP local thermal equilibrium r11 standard. Compared to other state-of-the-art SDR architectures, the proposed design adds features that improve programmer productivity with an insignificant power and area impact. Heikki Kultala, Timo Viitanen, Heikki Berg, Pekka Jääskeläinen, Joonas Multanen, Mikko Kokkonen, Kalle Raiskila, Tommi Zetterman, Jarmo Takala |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Instantaneous foveated preview for progressive Monte Carlo renderingabstractProgressive rendering, for example Monte Carlo rendering of 360° content for virtual reality headsets, is a time-consuming task. If the 3D artist notices an error while previewing the rendering, they must return to editing mode, make the required changes, and restart rendering. We propose the use of eye-tracking-based optimization to significantly speed up previewing of the artist’s points of interest. The speed of the preview is further improved by sampling with a distribution that closely follows the experimentally measured visual acuity of the human eye, unlike the piecewise linear models used in previous work. In a comprehensive user study, the perceived convergence of our proposed method was 10 times faster than that of a conventional preview, and often appeared to be instantaneous. In addition, the participants rated the method to have only marginally more artifacts in areas where it had to start rendering from scratch, compared to conventional rendering methods that had already generated image content in those areas. Matias Koskela, Kalle Immonen, Timo Viitanen, Pekka Jääskeläinen, Joonas Multanen, Jarmo Takala |
Comput. Vis. Media | 4 |
| 2017 | Fast Hardware Construction and Refitting of Quantized Bounding Volume HierarchiesabstractAbstract There is recent interest in GPU architectures designed to accelerate ray tracing, especially on mobile systems with limited memory bandwidth. A promising recent approach is to store and traverse Bounding Volume Hierarchies (BVHs), used to accelerate ray tracing, in low arithmetic precision. However, so far there is no research on refitting or construction of such compressed BVHs, which is necessary for any scenes with dynamic content. We find that in a hardware‐accelerated tree update, significant memory traffic and runtime savings are available from streaming, bottom‐up compression. Novel algorithmic techniques of modulo encoding and treelet‐based compression are proposed to reduce backtracking inherent in bottom‐up compression. Together, these techniques reduce backtracking to a small fraction. Compared to a separate top‐down compression pass, streaming bottom‐up compression with the proposed optimizations saves on average 42% of memory accesses for LBVH construction and 56% for refitting of compressed BVHs, over 16 test scenes. In architectural simulation, the proposed streaming compression reduces LBVH runtime by 20% compared to a single‐precision build, and 41% compared to a single‐precision build followed by top‐down compression. Since memory traffic dominates the energy cost of refitting and LBVH construction, energy consumption is expected to fall by a similar fraction. Timo Viitanen, Matias Koskela, Pekka Jääskeläinen, Kalle Immonen, Jarmo Takala |
Comput. Graph. Forum | 3 |
| 2017 | MergeTree: A Fast Hardware HLBVH Constructor for Animated Ray TracingabstractRay tracing is a computationally intensive rendering technique traditionally used in offline high-quality rendering. Powerful hardware accelerators have been recently developed that put real-time ray tracing even in the reach of mobile devices. However, rendering animated scenes remains difficult, as updating the acceleration trees for each frame is a memory-intensive process. This article proposes MergeTree, the first hardware architecture for Hierarchical Linear Bounding Volume Hierarchy (HLBVH) construction, designed to minimize memory traffic. For evaluation, the hardware constructor is synthesized on a 28nm process technology. Compared to a state-of-the-art binned surface area heuristic sweep (SAH) builder, the present work speeds up construction by a factor of 5, reduces build energy by a factor of 3.2, and memory traffic by a factor of 3. A software HLBVH builder on a graphics processing unit (GPU) requires 3.3 times more memory traffic. To take tree quality into account, a rendering accelerator is modeled alongside the builder. Given the use of a toplevel build to improve tree quality, the proposed builder reduces system energy per frame by an average 41% with primary rays and 13% with diffuse rays. In large ( > 500K triangles) scenes, the difference is more pronounced, 62% and 35%, respectively. Timo Viitanen, Matias Koskela, Pekka Jääskeläinen, Heikki Kultala, Jarmo Takala |
ACM Trans. Graph. | 3 |
| 2016 | Integer Linear Programming-Based Scheduling for Transport Triggered ArchitecturesabstractStatic multi-issue machines, such as traditional Very Long Instructional Word (VLIW) architectures, move complexity from the hardware to the compiler. This is motivated by the ability to support high degrees of instruction-level parallelism without requiring complicated scheduling logic in the processor hardware. The simpler-control hardware results in reduced area and power consumption, but leads to a challenge of engineering a compiler with good code-generation quality. Transport triggered architectures (TTA), and other so-called exposed datapath architectures, take the compiler-oriented philosophy even further by pushing more details of the datapath under software control. The main benefit of this is the reduced register file pressure, with a drawback of adding even more complexity to the compiler side. In this article, we propose an Integer Linear Programming (ILP) -based instruction scheduling model for TTAs. The model describes the architecture characteristics, the particular processor resource constraints, and the operation dependencies of the scheduled program. The model is validated and measured by compiling application kernels to various TTAs with a different number of datapath components and connectivity. In the best case, the cycle count is reduced to 52% when compared to a heuristic scheduler. In addition to producing shorter schedules, the number of register accesses in the compiled programs is generally notably less than those with the heuristic scheduler; in the best case, the ILP scheduler reduced the number of register file reads to 33% of the heuristic results and register file writes to 18%. On the other hand, as expected, the ILP-based scheduler uses distinctly more time to produce a schedule than the heuristic scheduler, but the compilation time is within tolerable limits for production-code generation. Tomi Äijö, Pekka Jääskeläinen, Tapio Elomaa, Heikki Kultala, Jarmo Takala |
ACM Trans. Archit. Code Optim. | 2 |
| 2014 | Heuristics for greedy transport triggered architecture interconnect explorationabstractMost power dissipation in Very Large Instruction Word (VLIW) processors occurs in their large, multi-port register files. Transport Triggered Architecture (TTA) is a VLIW variant whose exposed datapath reduces the need for RF accesses and ports. However, the comparative advantage of TTAs suffers in practice from a wide instruction word and complex interconnection network (IC). We argue that these issues are at least partly due to suboptimal design choices. The design space of possible TTA architectures is very large, and previous automated and ad-hoc design methods often produce inefficient architectures. We propose a reduced design space where efficient TTAs can be generated in a short time using excecution trace-driven greedy exploration. The proposed approach is evaluated by optimizing the equivalent of a 4-issue VLIW architecture. The algorithm finishes quickly and produces a processor with 10% reduced core energy product compared to a fully-connected TTA. Since the generated processor has low IC power and a shorter instruction word than a typical 4-issue VLIW, the results support the hypothesis that these drawbacks of TTA can be worked around with efficient IC design. Timo Viitanen, Heikki Kultala, Pekka Jääskeläinen, Jarmo Takala |
CASES | 3 |
| 2014 | Parallel programming of a symmetric transport-triggered architecture with applications in flexible LDPC encodingabstractExposed-datapath architectures yield small, low-power processors that trade instruction word length for aggressive compile-time scheduling and a high degree of instruction-level parallelism. In this paper, we present a general-purpose parallel accelerator consisting of a main processor and eight symmetric clusters, all in a single core. Use of a lightweight and memory-efficient application programming interface allows for the first high-performance program executing both sequential and data-parallel code on the same TTA processor. We use the processor for LDPC encoding, a popular method of forward error correction. Demonstrating the flexibility of software-defined radio, we benchmark the processor with two programs, one which can handle almost any sort of LDPC code, and another which is optimized for a specific standard. We achieve a throughput of 5 Mb/s with the flexible program and 92 Mb/s with the standard-specific one, while consuming only 95 mW at a clock frequency of 1175 MHz. Blaine Rister, Pekka Jääskeläinen, Olli Silvén, Jari Hannuksela, Joseph R. Cavallaro |
ICASSP | 2 |
| 2014 | A high throughput LDPC decoder using a mid-range GPUabstractA standard-throughput-approaching LDPC decoder has been implemented on a mid-range GPU in this paper. Turbo-Decoding Message-Passing algorithm is applied to achieve high throughput. Different from traditional host managed multi-streams to hide host-device transfer delay, we use kernel maintained data transfer scheme to achieve implicit data transfer between host memory and device shared memory, which eliminates an intermediate stage of global memory. Data type optimization, memory accessing optimization, and low complexity Soft-In Soft-Out algorithm are also used to improve efficiency. Through these optimization methods, the 802.11n LDPC decoder on NVIDIA GTX480 GPU, which is released in 2010 with Fermi architecture, has achieved a high throughput of 295Mb/s when decoding 512 codewords simultaneously, which is close to highest bit rate 300Mb/s with 20MHz bandwidth in 802.11n standard. Decoding 1024 and 4096 codewords achieve 330 and 365Mb/s. A 802.16e LDPC decoder is also implemented, 374Mb/s (512 codewords), 435Mb/s (1024 codewords) and 507Mb/s (4096 codewords) throughputs have been achieved. Xie Wen, Xianjun Jiao, Pekka Jääskeläinen, Heikki Kultala, Canfeng Chen, Heikki Berg, Zhisong Bie |
ICASSP | 3 |
| 2014 | Efficient software synthesis of dynamic dataflow programsabstractThis paper introduces advanced software synthesis techniques that enhance the implementation of dynamic dataflow programs. These techniques have been implemented into open-source tools and demonstrated on well-known video decoders including one based on the new High Efficiency Video Coding (HEVC) standard. The results show an improvement of more than 100% of the frame-rate over previously proposed implementations, and achieve real-time decoding of high definition video sequences. Hervé Yviquel, Alexandre Sanchez, Pekka Jääskeläinen, Jarmo Takala, Mickaël Raulet, Emmanuel Casseau |
ICASSP | 3 |
| 2014 | Grover: Looking for Performance Improvement by Disabling Local Memory Usage in OpenCL KernelsabstractDue to the diversity of processor architectures and application memory access patterns, the performance impact of using local memory in OpenCL kernels has become unpredictable. For example, enabling the use of local memory for an OpenCL kernel can be beneficial for the execution on a GPU, but can lead to performance losses when running on a CPU. To address this unpredictability, we propose an empirical approach: by disabling the use of local memory in OpenCL kernels, we enable users to compare the kernel versions with and without local memory, and further choose the best performing version for a given platform. To this end, we have designed Grover, a method to automatically remove local memory usage from OpenCL kernels. In particular, we create a correspondence between the global and local memory spaces, which is used to replace local memory accesses by global memory accesses. We have implemented this scheme in the LLVM framework as a compiling pass, which automatically transforms an OpenCL kernel with local memory to a version without it. We have validated Grover with 11 applications, and found that it can successfully disable local memory usage for all of them. We have compared the kernels with and without local memory on three different processors, and found performance improvements for more than a third of the test cases after Grover disabled local memory usage. We conclude that such a compiler pass can be beneficial for performance, and, because it is fully automated, it can be used as an auto-tuning step for OpenCL kernels. Jianbin Fang, Henk J. Sips, Pekka Jääskeläinen, Ana Lucia Varbanescu |
ICPP | 3 |
| 2013 | Simplified floating-point division and square rootabstractDigital Signal Processing (DSP) algorithms on low-power embedded platforms are often implemented using fixed-point arithmetic due to expected power and area savings over floating-point computation. However, recent research shows that floating-point arithmetic can be made competitive by using a reduced-precision format instead of, e.g., IEEE standard single precision, thereby avoiding the algorithm design and implementation difficulties associated with fixed-point arithmetic. This paper investigates the effects of simplified floating-point arithmetic applied to an FMA-based floating-point unit and the associated software division and square root operations. Software operations are proposed which attain near-exact precision with twice the performance of exact algorithms and resolve overflow-related errors with inexpensive exponent-manipulation special instructions. Timo Viitanen, Pekka Jääskeläinen, Otto Esko, Jarmo Takala |
ICASSP | 2 |
| 2013 | A 122Mb/s Turbo decoder using a mid-range GPUabstractParallel implementations of Turbo decoding has been studied extensively. Traditionally, the number of parallel sub-decoders is limited to maintain acceptable code block error rate performance loss caused by the edge effect of code block division. In addition, the sub-decoders require synchronization to exchange information in the iterative process. In this paper, we propose loosening the synchronization between the sub-decoders to achieve higher utilization of parallel processor resources. Our method allows high degree of parallel processor utilization in decoding of a single code block providing a scalable software-based implementation. The proposed implementation is demonstrated using a graphics processing unit. We achieve 122.8Mbps decoding throughput using a medium range GPU, the Nvidia GTX480. This is, to the best of our knowledge, the fastest Turbo decoding throughput achieved with a GPU-based implementation. Xianjun Jiao, Canfeng Chen, Pekka Jääskeläinen, Vladimír Guzma, Heikki Berg |
IWCMC | 3 |
| 2013 | Turbo decoding on tailored OpenCL processorabstractTurbo coding is commonly used in the current wireless standards such as 3G and 4G. However, due to the high computational requirements, its software-defined implementation is challenging. This paper proposes a static multi-issue exposed datapath processor design tailored for turbo decoding. In order to utilize the parallel processor datapath efficiently without resorting to low level assembly programming, the turbo decoder is implemented using OpenCL, a parallel programming standard for heterogeneous devices. The proposed implementation includes only a small set of Turbo-specific custom operations to accelerate the most critical parts of the algorithm. Most of the computation is performed using general-purpose integer operations. Thus, the processor design can be used as a general-purpose OpenCL accelerator for arbitrary integer workloads as well. The proposed processor design was evaluated both by implementing it using a Xilinx Virtex 6 FPGA and by ASIC synthesis using 130 nm and 40 nm technology libraries. The implementation achieves over 63 Mbps Turbo decoding throughput on a single low-power core. According to the ASIC synthesis, the maximum operating clock frequency is 344 MHz/1 050 MHz (130 nm/40 nm). Heikki Kultala, Otto Esko, Pekka Jääskeläinen, Vladimír Guzma, Jarmo Takala, Xianjun Jiao, Tommi Zetterman, Heikki Berg |
IWCMC | 3 |
| 2010 | Customized Exposed Datapath Soft-Core Design Flow with Compiler SupportabstractA popular way to exploit high level programming languages in FPGA designs is to use a soft-core with accompanying software development tools. However, a common shortcoming with the current soft-core offerings is their limited software execution capability: the required performance for the implementation can be often reached only with instruction set extensions. In this paper, we propose and evaluate an application-specific processor design toolset that uses a multi-issue exposed data path processor architecture template. The main benefit of the architecture is scalability with respect to instruction-level parallelism (ILP). The design flow allows the designer to freely customize the data path resources in the core to exploit the ILP available in computation intensive kernels. The design toolset includes a retargetable C compiler and an architecture simulator, making design space exploration feasible. The experiments show that a relatively small soft-core tailored with the toolset provides significant speedups on software execution without using any instruction set extensions. The best measured speedup in comparison to the major commercial soft-cores was fourfold in applications from the CHStone benchmark suite, while the amount of consumed FPGA resources remained moderate. Otto Esko, Pekka Jääskeläinen, Pablo Huerta, Carlos S. de La Lama, Jarmo Takala, José Ignacio Martínez |
FPL | 2 |
| 2008 | Resource conflict detection in simulation of function unit pipelines
Pekka Jääskeläinen, Vladimír Guzma, Viljami Korhonen |
J. Syst. Archit. | 1 |