VLDB 2026 Research / reviewers in the wild / expert
Timo Hämäläinen 0001
dblp:h/TimoHamalainen · also Timo D. Hämäläinen
· DBLP profile ↗
123ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-7867-0800ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 66 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27Software engineering, systems software and programming languages · 8Artificial intelligence and machine learning · 4 · 1 first-authorComputer networks · 4Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Work in Progress: Hardware Support for EDF Scheduling on Bare-Metal Systems
Antti Nurmi, Justin Beaurivage, Pawel Dzialo, Per Lindgren, Timo Hämäläinen 0001 |
RTAS | 5 |
| 2026 | Headsail: One-Year Tape-Out of a 25-mm2 Linux-Capable RISC-V MPSoCabstractThe Internet-of-Things (IoT) devices feature a broad range of power, memory, and performance requirements. Ultralow-power, low-performance controllers are at one end of the spectrum, while high-performance, power-intensive systems-on-chip (SoCs) are at the other. Heterogeneous and specialized multiprocessor SoC (MPSoC) architectures have emerged as the most effective paradigm for delivering high performance and energy efficiency across a wide range of application workloads. This work introducesHeadsail, an MPSoC application-specific integrated circuit (ASIC) designed by SoC Hub at Tampere University, Finland.Headsailfeatures a 512-KiB primary data buffer, 128 KiB of shared on-chip SRAM, seven CPU cores (including four CVA6 64-bit RISC-V processors), a low power-DDR2 (LP-DDR2) memory controller, two unique chip-to-chip (C2C) interfaces, and shared peripherals.Headsailhas been successfully implemented using a TSMC 22-nm low-power CMOS technology. Testing results show that the samples can support a maximum operating frequency of 1 GHz and achieve a peak performance of 1100 giga operations per second (GOPS), with an implementation area of 25 mm2and a power-consumption range of 64 mW–1.5 W. Matti Käyrä, Thomas Szymkowiak, Antti Rautakoura, Antti Nurmi, Kari Hepola, Henri Lunnikivi, Toni Jääskeläinen, Abdesattar Kalache, Petteri Toivanen, Roope Keskinen, Andreas Stergiopoulos, Väinö-Waltteri Granat, Arto Oinonen, Joonas Multanen, Pekka Jääskeläinen, Karri Palovuori, Timo Hämäläinen 0001, Syed Mohsin Abbas |
IEEE Trans. Very Large Scale Integr. Syst. | 18 |
| 2025 | Efficient and Predictable Context Switching for Mixed-Criticality and Real-Time SystemsabstractContext switching is both a highly utilized and highly repetitive routine in interrupt-driven systems, such as safety-critical control systems. Conventional context switching routines are sequential and dependent on data memory access, which may be detrimental to time-predictability. This publication explores the use of stacked register files for efficient and predictable context switching. Two complementary microarchitectures are characterized: combinationally addressed register windowing, and a novel parallel context stack (PCS). Both implementations enable minimal latency and inherent predictability in context switching. To efficiently utilize the benefit of stacked register files while limiting hardware costs, the heterogeneous interrupt (HETI) architecture is proposed. HETI integrates a small stacked register file for accelerating a dynamically selected subset of high-priority interrupts. Automatic firmware generation is contributed to enable seamless utilization of the HETI architecture. A total of four HETI configurations on an open-source RISC-V microcontroller are evaluated against the baseline platform and an implementation of Cortex-M style hardware-assisted stacking. Implementations on a TSMC 22nm technology demonstrate low area overhead for small HETI configurations and favorable frequency characteristics against the hardware-assisted stacking implementation. A representative layout of the full system with a HETI-4 instance is presented with a gate count overhead of 1.2% and no frequency detriment in relation to the baseline design. The functional performance evaluated in a synthetic case study demonstrates how the HETI design can reduce retired instruction count by up to 26% and allow for 21% more sleep in comparison to the software baseline and Cortex-M style solution, promising significant improvements to real-time response and energy efficiency. Antti Nurmi, Abdesattar Kalache, Henri Lunnikivi, Per Lindgren, Timo Hämäläinen 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | An Efficient High-level Synthesis Implementation of the MUSIC DoA Algorithm for FPGAabstractHigh-level synthesis (HLS) promises to increase the design and verification productivity for digital hardware systems. However, the industry still predominantly uses more time-consuming manual register-transfer level techniques instead of HLS. To accelerate the adoption of HLS, it is vital to explore if it is possible to achieve competitive results with this method. To that end, this paper demonstrates an HLS implementation of the well-known MUSIC algorithm for estimating the direction of arrival of a radio signal. We use as a receiver a four-antenna uniform linear array with one signal source and a resolution of one degree. For the computationally heavy eigenvalue decomposition within MUSIC, we employ the iterative Jacobi algorithm. We target two different Virtex FPGAs for synthesis and obtain results faring well in comparison to the previous literature, with$5.0\ \mu \mathrm{s}$microseconds latency, high accuracy, and low resource consumption. The results show that HLS is suitable for implementing these kinds of algorithms on FPGA. Sakari Lahti, Tuomas Aaltonen, Elizaveta Rastorgueva-Foi, Jukka Talvitie, Bo Tan 0003, Timo Hämäläinen 0001 |
DDECS | 6 |
| 2024 | Keelhaul: Processor-Driven Chip Connectivity and Memory Map Metadata Validator for Large Systems-on-ChipabstractThe integration of large-scale systems-on-chip warrants thorough verification both at the level of the individual component and at the system level. In this article, we address the automated testing of system-level memory maps. The golden reference is the IEEE 1685/IP-XACT hardware description, which includes implementation agnostic definitions for the global memory map. The IP-XACT description is used as a specification for implementing the registers and memory regions in a register transfer-level (RTL) language, and for implementing the corresponding hardware-dependent software. The challenge is that hardware design changes might not always propagate to firmware and applications developers, which causes errors and faults. We present a method and a tool called Keelhaul which takes as input the CMSIS-SVD format commonly used for firmware development and generates automated software tests that attempt to access all available memory mapped input/output registers. During development of a large-scale research-focused multiprocessor system-on-chip, we ran a total of 32 automatically generated test suites per pipeline comprising 882 test cases for each of its two CPU subsystems. A total of 15 distinct issues were found by the tool in the lead-up to tapeout. Another research-focused SoC was validated posttapeout with 984 test cases generated for each core, resulting in the discovery of four distinct issues. Keelhaul can be used with any IP-XACT or CMSIS-SVD-based systems-on-chip that include processors for accessing implemented registers and memory regions. Henri Lunnikivi, Roni Hämäläinen, Timo Hämäläinen 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | AnTiQ: A Hardware-Accelerated Priority Queue Design with Constant Time Arbitrary Element RemovalabstractThe recent trend towards open architectures and open source hardware enables agile development of domain specific architectures. In this paper, we explore and propose the design and implementation of AnTiQ, a hardware-accelerated priority queue tailored to the needs of timer queue implementations and other cases where removal of arbitrary elements, a feature we found lacking in related work, is desirable. The key novelty of AnTiQ is the arbitrary remove or DROP operation, which along with PUSH, POP and PEEK is performed in constant time. The design was conceived by reviewing existing implementations of hardware priority queues and analysing the fundamental characteristics of three queue architectures: the shift register queue, the systolic array queue and the binary heap queue. The systolic array architecture was chosen for our design due to its pipelined structure and ability to provide constant time responses to queue operations regardless of the queue depth. The architecture and implementation of our RTL prototype is presented and relevant behavioural corner cases in the design are identified. The implementation of our design achieved a clock frequency of 350 MHz on a Xilinx VCU 118 FPGA. The scalability of the design is evaluated through FPGA synthesis for 32, 64, 128 and 256 depth configurations and compared to the related work. We elaborate on directions for further development and future work towards monotonic timer implementations for the RTIC framework, a domain specific Rust language extension for concurrent programming targeting bare metal microcontrollers. Antti Nurmi, Per Lindgren, Thomas Szymkowiak, Timo Hämäläinen 0001 |
DSD | 4 |
| 2023 | Leveraging Modern C++ in High-Level SynthesisabstractHigh-level synthesis (HLS) enables the automated conversion of high-level language algorithms into synthesizable register-transfer level code, allowing computation-intensive algorithms to be accelerated on FPGAs. Most HLS tools have C++ as their input language, as it is widely known in both software and hardware industry. However, even though C++ receives a new standard every three years, the HLS tool vendors have mostly provided support and examples using C++98/03. Limiting to early C++ standards imposes a productivity penalty, since the newer standards provide both compilation time reductions and more concise, expressive, and maintainable way of writing code. In this study, we make the case for adopting modern C++ in HLS. We inspect the language features of C++11 and forward, and consider their benefits for HLS. We also test the present support for the modern language features with two state-of-the-art commercial HLS tools. Finally, we provide an extended example, demonstrating the increased clarity of code achieved using the newer standards. We note that the investigated HLS tools already have good support for modern C++ features, and urge their adoption to increase designer productivity. Sakari Lahti, Matti Rintala, Timo Hämäläinen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Does SoC Hardware Development Become Agile by Saying So: A Literature Review and Mapping StudyabstractThe success of agile development methods in software development has raised interest in System-on-Chip (SoC) design, which involves high architectural and development process complexity under time and project management pressure. This article discovers the current state of agile hardware development with the questions (1) how well literature covers the SoC development process, (2) what agile methods and practices are applied or (3) what proposals are made to increase the agility, and (4) what is the impact for the SoC community. To answer the questions, a mapping study and literature review were performed. Seven hundred thirty papers were first studied, and eventually, after a rigorous filtering process, 25 papers were thoroughly analyzed. The results show that the popular agile SW development methods are applied in 5 cases, ideas adapted from the agile Hardware manifesto in 9 cases, and 11 cases do not define the Agile HW development method. Most of the papers address shorter development time by better methodologies and tools that indirectly shape the SoC development toward agility. The focus of agile hardware development is mostly on the SoC artifacts and methodological improvements have not been quantified. However, the literature indicates a significant impact on many academic chip prototypes. The challenges are better understood and the interest in agile methods is clearly increasing. The methodological gaps in the prevalent situation encourage further research and more accurate reporting of the development in addition to the SoC artifacts. Antti Rautakoura, Timo Hämäläinen 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | A Resilient System Design to Boot a RISC-V MPSoCabstractThis paper presents a highly resilient boot process design for Ballast, a new RISC- V based multiprocessor system-on-chip (SoC). An open source RISC- V SoC was adapted as a bootstrap processor and customized to meet our requirement for guaranteed chip wake-up. We outline the characteristic challenges of implementing a large program into a read-only memory (ROM) used for booting and propose generally applica-ble workflows to verify the boot process for application specific integrated circuit (ASIC) synthesis. We implemented four distinct boot modes. Two modes that load a software bootloader autonomously from an SD card are implemented for a secure digital input output (SDIO) interface and for a serial peripheral interface (SPI), respectively. Another SDIO based mode allows for direct program execution from external memory, while the last mode is based on usage of a RISC- V debug module. The boot process was verified with instruction set simulation, register transfer level simulation, gate-level simulation and field-programmable gate array prototyping. We received the fabricated ASIC samples and were able to successfully boot the chip via all boot modes on our custom circuit board. Antti Nurmi, Antti Rautakoura, Henri Lunnikivi, Timo Hämäläinen 0001 |
DSD | 4 |
| 2022 | Ballast: Implementation of a Large MP-SoC on 22nm ASIC TechnologyabstractChips have become the critical asset of the technology, and increasing effort is put to design System-on-Chips (SoC) faster and more affordable. Typically the focus of the research has been on the Power, Performance and Area optimization of the specific component or sub-system. To improve the situation we report design effort for complex SoC counted from specification to ASIC tape-out to lay out a solid reference for the community. Ballast is the first SoC-Hub chip taped out on 22nm technology. It includes six sub-systems on 15 mm2area and reaches 1.2GHz top speed. The design team included 24 persons and spent 21 200 person hours to tape-out in one calendar year from scratch. This is an outstanding achievement and sets the baseline to SoC design productivity development. Antti Rautakoura, Timo Hämäläinen 0001, Ari Kulmala, Tero Lehtinen, Mehdi Duman |
DSD | 2 |
| 2022 | High-Level Synthesis Implementation of an Embedded Real-Time HEVC Intra Encoder on FPGA for Media ApplicationsabstractHigh Efficiency Video Coding (HEVC) is the key enabling technology for numerous modern media applications. Overcoming its computational complexity and customizing its rich features for real-time HEVC encoder implementations, calls for automated design methodologies. This article introduces the first complete High-Level Synthesis (HLS) implementation for HEVC intra encoder on FPGA. The C source code of our open-source Kvazaar HEVC encoder is used as a design entry point for HLS that is applied throughout the whole encoder design process, from data-intensive coding tools like intra prediction and discrete transforms to more control-oriented tools such as context-adaptive binary arithmetic coding (CABAC). Our prototype is run on Nokia AirFrame Cloud Server equipped with 2.4 GHz dual 14-core Intel Xeon processors and two Intel Arria 10 PCIe FPGA accelerator cards with 40 Gigabit Ethernet. This proof-of-concept system is designed for hardware-accelerated HEVC encoding and it achieves real-time 4K coding speed up to 120 fps. The coding performance can be easily scaled up by adding practically any number of network-connected FPGA cards to the system. These results indicate that our HLS proposal not only boosts development time, but also provides previously unseen design scalability with competitive performance over the existing FPGA and ASIC encoder implementations. Panu Sjövall, Ari Lemmetti, Jarno Vanne, Sakari Lahti, Timo Hämäläinen 0001 |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2021 | A Survey on System-on-a-Chip Design Using Chisel HW Construction LanguageabstractThis paper presents a survey of functional programming languages in System-on-a-Chip (SoC) design. The motivation is improving the design productivity by better source code expressiveness, increased abstraction level in design entry, or improved automation. The survey focuses on Chisel that is one of the most potential High Level Language (HLL) based design frameworks. We include 26 papers that report implementations ranging from IP blocks to complete chips. The result is that functional programming languages are viable for SoC design and can also be deployed in production use. However, Chisel does not increase the abstraction level in a similar way as High Level Synthesis (HLS), since it is used to create circuit generators instead of direct descriptions. Additional benefit is that Chisel offloads user effort from control and connectivity structures, and makes reusability and configurability improved over traditional Hardware Description Language (HDL) designs. Matti Käyrä, Timo Hämäläinen 0001 |
IECON | 2 |
| 2020 | Kamel: IP-XACT compatible intermediate meta-model for IP generationabstractAutomatic code generation is used to implement Intellectual Property (IP) blocks for System-on-Chip (SoC), but the challenge is how to describe the IP as a model and what is a feasible meta-model. IEEE 1685 IP-XACT standard and many domain-specific meta-models are not compatible and tool flows are too specific for general use. We present Kamel that is a new intermediate IP meta-model. It is used to generate behavioral code to complete IP-XACT structural models. The key idea is light modeling overhead while automating the majority of the RTL IP development tasks. Kamel uses Model Driven Architecture (MDA) to integrate IP-XACT and Kamel modeling together. Python Mako template-based code generation framework is used to generated different views from the models. The compatibility with IP-XACT is demonstrated with the Kactus2 tool. Our case study is modeling and code generation for Kvazaar HEVC video intra encoder IP block on FPGA. The results confirm that the Kamel and introduced tool flow can provide 5×-10× productivity gain when measured on time spent on model entry and Lines of Code used for model entry. Antti Rautakoura, Matti Käyrä, Timo Hämäläinen 0001, Wolfgang Ecker, Esko Pekkarinen, Mikko Teuho |
DSD | 3 |
| 2020 | Implementation of a Nonlinear Self-Interference Canceller using High-Level SynthesisabstractHigh-level synthesis (HLS) aims to improve the productivity of digital logic design over traditional register-transfer level (RTL) methods. This paper shows that HLS can replace RTL when implementing a complex data path oriented signal processing algorithm under strict throughput constraints. Our system is a nonlinear spline-based Hammerstein self-interference (SI) canceller for full-duplex transceiver capable of achieving high SI suppression, while maintaining low computational complexity. The achieved suppression of the SI is superb 45 dB, while consuming 29 026 of the available LUTs, 17992 of registers, and 655 of the DSP slices on Kintex-7 XC7K410T FPGA. Our paper also compares the usability of two commercial HLS tools that were used in this work. Sakari Lahti, Pablo Pascual Campo, Vesa Lampu, Lauri Anttila, Mikko Valkama, Timo Hämäläinen 0001 |
ISCAS | 6 |
| 2019 | Visualization of Dynamic Resource Allocation for HEVC Encoding in FPGA-Accelerated SDN CloudabstractThis paper describes a demonstration setup to visualize dynamic resource allocation for real-time HEVC encoding services in FPGA-accelerated cloud. The demonstrated application is Kvazaar HEVC intra encoder, whose functionality is partitioned between FPGAs and processors. During the demonstration, several encoding services can be invoked with requests to the resource manager, which is responsible for allocation, deallocation, and load balancing of resources in the network. The manager provides JSON data to the visualizer, which uses D3 JavaScript library to visualize 1) the physical network structure; 2) running services; and 3) performance of the network elements. This interactive demonstration allows users to request new video streams, view the encoded streams, observe the visualization of the network and services, and manually turn on/off resources to test the robustness of the system. Panu Sjövall, Mikko Teuho, Arto Oinonen, Jarno Vanne, Timo Hämäläinen 0001 |
VCIP | 5 |
| 2019 | Are We There Yet? A Study on the State of High-Level SynthesisabstractTo increase productivity in designing digital hardware components, high-level synthesis (HLS) is seen as the next step in raising the design abstraction level. However, the quality of results (QoRs) of HLS tools has tended to be behind those of manual register-transfer level (RTL) flows. In this paper, we survey the scientific literature published since 2010 about the QoR and productivity differences between the HLS and RTL design flows. Altogether, our survey spans 46 papers and 118 associated applications. Our results show that on average, the QoR of RTL flow is still better than that of the state-of-the-art HLS tools. However, the average development time with HLS tools is only a third of that of the RTL flow, and a designer obtains over four times as high productivity with HLS. Based on our findings, we also present a model case study to sum up the best practices in comparative studies between HLS and RTL. The outcome of our case study is also in line with the survey results, as using an HLS tool is seen to increase the productivity by a factor of six. In addition, to help close the QoR gap, we present a survey of literature focused on improving HLS. Our results let us conclude that HLS is currently a viable option for fast prototyping and for designs with short time to market. Sakari Lahti, Panu Sjövall, Jarno Vanne, Timo Hämäläinen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Rate-Distortion-Complexity Optimized Coding Scheme for Kvazaar HEVC Intra EncoderabstractThis paper summarizes a low-complexity rate-distortion optimization (RDO) scheme for Kvazaar HEVC intra encoder (github.com/ultravideo/kvazaar). Our work particularly addresses RDO quantization (RDOQ) since it is the most complex intra coding tool taking almost 60% of the Kvazaar complexity. Ari Lemmetti, Eemeli Kallio, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001 |
DCC | 5 |
| 2018 | Modeling RISC-V Processor in IP-XACTabstractIP-XACT is the most used standard in IP (Intellectual Property) integration. It is intended as a language neutral golden reference, from which RTL and HW dependent SW is automatically generated. Despite its wide popularity in the industry, there are practically no public and open design examples for any part of the design flow from IP-XACT to synthesis. One reason is the difficulty of creating IP-XACT models for existing RTL projects. In this paper, we address the issues by modeling the PULPino RISC-V microprocessor that is written in SystemVerilog (SV) and the project distributed over several repositories. We propose how to solve the mismatching concepts between SV project and IP-XACT, and based on the findings propose improvements for the Kactus2 IP-XACT tool. In addition, the final PULPino model contributes to the rare public non-trivial examples for better adoption of the IP-XACT methodology. Esko Pekkarinen, Timo Hämäläinen 0001 |
DSD | 2 |
| 2018 | Visualization of Memory Map Information in Embedded System DesignabstractData compression is a common requirement for displaying large amounts of information. The goal is to reduce visual clutter. The approach given in this paper uses an analysis of a data set to construct a visual representation. The visualization is compressed using the address ranges of the memory structure. This method produces a compressed version of the initial visualization, retaining the same information as the original. The presented method has been implemented as a Memory Designer tool for ASIC, FPGA and embedded systems using IP-XACT. The Memory Designer is a user-friendly tool for model based embedded system design, providing access and adjustment of the memory layout from a single view, complementing the "programmer's view" to the system. Mikko Teuho, Esko Pekkarinen, Timo Hämäläinen 0001 |
DSD | 3 |
| 2018 | Feasibility of FPGA Accelerated IPsec on CloudabstractLine-rate speed requirements for performance hungry network applications like IPsec are getting problematic due to the virtualization trend. A single virtual network application hardly can provide 40 Gbps operation. This research considers the IPsec packet processing without IKE to be offloaded on an FPGA in a network. We propose an IPsec accelerator in an FPGA and explain the details that need to be considered for a production ready design. Based on our evaluation, Intel Arria 10 FPGA can provide 10 Gbps line-rate operation for the IPsec accelerator and to be responsible for 1000 IPsec tunnels. The research points out that for future data centers it is beneficial to rely on HW acceleration in terms of speed and energy efficiency for applications like IPsec. Markku Vajaranta, Vili Viitamäki, Arto Oinonen, Timo Hämäläinen 0001, Ari Kulmala, Jouni Markunmäki |
DSD | 4 |
| 2018 | Low Latency Edge Rendering Scheme for Interactive 360 Degree Virtual Reality GamingabstractThis paper describes the core functionality and a proof-of-concept demonstration setup for remote 360 degree stereo virtual reality (VR) gaming. In this end-to-end scheme, the execution of a VR game is off-loaded from an end user device to a cloud edge server in which the executed game is rendered based on user's field of view (FoV) and control actions. Headset and controller feedback is transmitted over the network to the server from which the rendered views of the game are streamed to a user in real-time as encoded HEVC video frames. This approach saves energy and computation load of the end terminals by making use of the latest advancements in network connection speed and quality. In the showcased demonstration, a VR game is run in Unity on a laptop powered by i7 7820HK processor and GTX 1070 GPU. The 360 degree spherical view of the game is rendered and converted to a rectangular frame using equirectangular projection (ERP). The ERP video is sliced vertically and only the FoV is encoded with Kvazaar HEVC encoder in real time and sent over the network in UDP packets. Another laptop is used for playback with a HTC Vive VR headset. Our system can reach an end-to-end latency of 30 ms and bit rate of 20 Mbps for stereo 1080p30 format. Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001, Ari Kulmala |
ICDCS | 3 |
| 2018 | FPGA-Powered 4K120p HEVC Intra EncoderabstractThis paper presents a hardware-accelerated Kvazaar HEVC intra encoder for 4K real-time video coding at up to 120 fps. The encoder is implemented on a Nokia AirFrame Cloud Server featuring a 2.4 GHz dual 14-core Intel Xeon processor and two Arria 10 PCI Express FPGA accelerator cards. The presented encoder is a speed-optimized version of our 1st generation 4K40p HEVC intra encoder. The proposed speedup techniques include 1) Increasing the number of FPGA cards to two; 2) Remapping the simplest multiplications from DSP blocks to logic for better FPGA utilization; 3) Making task scheduling more flexible to improve utilization rate of hardware accelerators; and 4) Increasing the pipeline depth and duplicating time-sensitive resources in the hardware accelerator. As a result, up to three hardware accelerator instances can be accommodated in a single Arria 10 so the encoder is able to make use of six accelerators. According to our experiments, the proposed encoder obtains threefold speedup over our 1st generation encoder. Our proposal is also shown to outperform all other encountered FPGA and ASIC implementations. Panu Sjövall, Vili Viitamäki, Jarno Vanne, Timo Hämäläinen 0001, Ari Kulmala |
ISCAS | 4 |
| 2018 | Live Demonstration: 4K100p HEVC Intra EncoderabstractThis paper describes a demonstration setup for real-time 4K HEVC intra coding. The system is built on Kvazaar open-source HEVC encoder partitioned between 22-core Xeon processor and two Arria 10 FPGAs. The demonstrator supports 1) live streaming of up to three 4K30p videos; or 2) offline video streaming up to 4K100p format. Live feeds are shot by three cameras whereas offline video is accessed from a local hard drive. In both cases, encoded bit stream is sent over a wired connection and played back by laptop(s). The demonstrated HEVC coding speed is over three times as fast as that of a pure software solution. Vili Viitamäki, Panu Sjövall, Jarno Vanne, Timo Hämäläinen 0001, Ari Kulmala |
ISCAS | 4 |
| 2018 | Live Demonstration: Kvazzup 4K HEVC Video CallabstractThis paper describes a demonstration setup for an end-to-end 4K video call with Kvazzup open-source HEVC video call application. The Kvazzup clients are installed on a desktop and a laptop computer powered by Intel 22-core Xeon and Intel 4-core i7 processors, respectively. The proposed two-way peer-to-peer video call setup is shown to support 2160p30 video stream from the desktop to the laptop and 720p30 stream in the reverse direction. Joni Räsänen, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001 |
ISM | 4 |
| 2018 | Eye-Controlled Region of Interest HEVC EncodingabstractThis paper presents a demonstrator setup for real-time HEVC encoding with gaze-based region of interest (ROI) detection. This proof-of-concept system is built on Kvazaar open-source HEVC encoder and Pupil eye tracking glasses. The gaze data is used to extract the ROI from live video and the ROI is encoded with higher quality than non-ROI regions. This demonstration illustrates that performing HEVC encoding with non-uniform quality reduces bit rate by 40-90% and complexity by 10-35% over that of the conventional approaches with negligible to minor deterioration in subjective quality. Joose Sainio, Arttu Ylä-Outinen, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001 |
ISM | 5 |
| 2018 | Open framework for error-compensated gaze data collection with eye tracking glassesabstractEye tracking is nowadays the primary method for collecting training data for neural networks in the Human Visual System modelling. Our recommendation is to collect eye tracking data from videos with eye tracking glasses that are more affordable and applicable to diverse test conditions than conventionally used screen based eye trackers. Eye tracking glasses are prone to moving during the gaze data collection but our experiments show that the observed displacement error accumulates fairly linearly and can be compensated automatically by the proposed framework. This paper describes how our framework can be used in practice with videos up to 4K resolution. The proposed framework and the data collected during our sample experiment are made publicly available. Kari Siivonen, Joose Sainio, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001 |
ISM | 5 |
| 2017 | Analysis and Visualization of Product Memory Layout in IP-XACTabstractModern ASIC and FPGA based embedded products use model based design, in which both hardware and software are developed in parallel. Previously HW was completed first and the information handed over to SW team, typically in the form of register tables. The information was even manually copied to SW code, making any changes error-prone and laborious. IP-XACT is the most feasible standard to model HW also for the SW needs. The HW design connectivity and overall memory layout may change due to component instantiations, configurations and conditional operation states, which makes it difficult to create register tables even for documentation. Current register design tools fall short in serving the needs of both the HW and SW designer. All of them rely on low-level IP-XACT XML description that leaves much of the interpretation responsibility to the user, and most do not have any user-friendly visualization of the whole. Our approach is to perform a design connectivity analysis through hierarchies and configurations, analyze the memory structure based on it, and finally provide a novel, interactive visualization of system-wide memory maps. Our new Memory Designer for Kactus2 serves the hardware design by accessing and adjusting the memory layout from a single view disclosing end-to-end paths between masters and slaves. For SW development the Memory Designer shows the "programmer's view" to the system by hiding intermediate address translations and other connectivity related issues. As far as we know, this is the first visual, user-friendly memory tool for model based embedded system design. Esko Pekkarinen, Mikko Teuho, Timo Hämäläinen 0001 |
DSD | 3 |
| 2017 | High-level synthesis implementation of HEVC 2-D DCT/DST on FPGAabstractThis paper presents the first known high-level synthesis (HLS) implementation of integer discrete cosine transform (DCT) and discrete sine transform (DST) for High Efficiency Video Coding (HEVC). The proposed approach implements these 2-D transforms by two successive 1-D transforms using a well-known row-column and Even-Odd decomposition techniques. Altogether, the proposed architecture is composed of a 4-point DCT/DST unit for the smallest transform blocks (TBs), an 8/16/32-point DCT unit for the other TBs, and a transpose memory for intermediate results. On Arria II FPGA, the low-cost variant of the proposed architecture is able to support encoding of 1080p format at 60 fps and at the cost of 10.0 kALUTs and 216 DSP blocks. The respective figures for the proposed high-speed variant are 2160p at 30 fps with 13.9 kALUTs and 344 DSP blocks. These cost-performance characteristics outperform respective non-HLS approaches on FPGA. Panu Sjövall, Vili Viitamäki, Jarno Vanne, Timo Hämäläinen 0001 |
ICASSP | 4 |
| 2017 | Complexity based test cases for log file analyzersabstractLog files are used in many big data applications. If the log is meant for a different purpose, the analysis and finding the best log analyzer can be very complex. Our solution is to create a generic test case framework to model and create representative log data. Related work model the behavior as state machines, but our model uses a composition of elementary acyclic graphs, thus addressing the log file size, variation, branching and confidence. We have created test cases originally based on real Intelligent Transportation Systems (ITS) data, and evaluated our LOGDIG analyzer against it. We can easily generate hundreds of test cases with our model, and modify the cases as needed. Esa Heikkinen, Timo Hämäläinen 0001 |
INDIN | 2 |
| 2017 | High-level synthesized 2-D IDCT/IDST implementation for HEVC codecs on FPGAabstractThis paper presents efficient inverse discrete cosine transform (IDCT) and inverse discrete sine transform (IDST) implementations for High Efficiency Video Coding (HEVC). The proposal makes use of high-level synthesis (HLS) to implement a complete HEVC 2-D IDCT/IDST architecture directly from the C code of a well-known Even-Odd decomposition algorithm. The final architecture includes a 4-point IDCT/IDST unit for the smallest transform blocks (TB), an 8/16/32-point IDCT unit for the other TBs, and a transpose memory for intermediate results. On Arria II FPGA, it supports real-time (60 fps) HEVC decoding of up to 2160p format with 12.4 kALUTs and 344 DSP blocks. Compared with the other existing HLS approach, the proposed solution is almost 5 times faster and is able to utilize available FPGA resources better. Vili Viitamäki, Panu Sjövall, Jarno Vanne, Timo Hämäläinen 0001 |
ISCAS | 4 |
| 2017 | Kvazzup: Open Software for HEVC Video CallsabstractThis paper introduces an open-source HEVC video call application called Kvazzup. This academic proposal is the first HEVC-based end-to-end video call system with a user-friendly Graphical User Interface for call management. Kvazzup is built on the Qt framework and it makes use of four open-source tools: Kvazaar for HEVC encoding, OpenHEVC for HEVC decoding, Opus codec for audio coding, and Live555 for managing RTP/RTCP traffic. In our experiments, Kvazzup is prototyped with low-complexity VGA and high-quality 720p video calls between two desktops. On an Intel 4-core i5 processor, the VGA call accounts for 17% of the total CPU time. Averagely, it requires a bit rate of 0.31 Mbit/s out of which 0.26 Mbit/s is taken by video and 0.05 Mbit/s by audio. In the 720p call, the respective figures are 46%, 1.13 Mbit/s, 1.08 Mbit/s, and 0.05 Mbit/s. These test cases also validate the feasibility of HEVC in different types of video calls. HEVC coding is shown to account for around 34% of the Kvazzup processing time in the VGA call and 45% in the 720p call. Joni Räsänen, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001 |
ISM | 4 |
| 2017 | Kvazaar: HEVC/H.265 4K30p Intra EncoderabstractThis paper demonstrates the usage of Kvazaar open-source HEVC intra encoder in 4K real-time video encoding. In this setup, a raw 4K video is shot by an action camera, captured by an HDMI capture card, encoded in real-time by Kvazaar ultrafast preset on a 22-core Intel Xeon processor, sent to a laptop, and decoded by OpenHEVC decoder for playback. The encoding process is visualized on the fly by Kvazaar run-time visualizer. Arttu Ylä-Outinen, Ari Lemmetti, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001 |
ISM | 5 |
| 2016 | AVX2-optimized Kvazaar HEVC intra encoderabstractThis paper presents efficient SIMD optimizations for the open-source Kvazaar HEVC intra encoder. The C implementation of Kvazaar is accelerated by Intel AVX2 instructions whose effect on Kvazaar ultrafast preset is profiled. According to our profiling results, C functions of SATD, DCT, quantization, and intra prediction account for over 60% of the total intra coding time of Kvazaar ultrafast preset. This work shows that optimizing primarily these functions doubles the coding speed of a single-threaded Kvazaar intra encoder for the same rate-distortion performance. The highest performance boost is obtained by deploying the proposed optimizations jointly with multithreading. On the Intel 8-core i7 processor, the AVX2-optimized 16-threaded Kvazaar ultrafast preset achieves real-time (30 fps) intra coding speed up to 1080p resolution. Compared to AVX2-optimized ultrafast preset of x265, Kvazaar is 20% times faster and still obtains 9.1% bit rate gain for the same quality. These results justify that Kvazaar is currently the leading open-source HEVC intra encoder in terms of real-time coding speed and efficiency. Ari Lemmetti, Ari Koivula, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001 |
ICIP | 5 |
| 2016 | Behavior Mining Language for mining expected behavior from log filesabstractLog files are often the only way to identify and locate errors in a deployed system. This paper proposes a novel Behavior Mining Language (BML) for a log file analyzing framework called LOGDIG. It is proposed for logs that include temporal data (timestamps) and extra-log system-specific data (e.g. spatial data with coordinates of moving objects), which are present e.g. in Real Time Passenger Information Systems (RTPIS). BML is state-machine-based, and specifies searches for desired events from the log files by adjustable accuracy. The analysis output is static behavioral knowledge and human friendly composite log files for reporting the results in legacy tools. Field data from a commercial RTPIS is used as a proof-of-concept case study. BML is Python-based for excellent development support, as well as self-explanatory and self-descriptive for correct-by-construct usage. Compared to a general language approach, BML code is much shorter and easier to maintain. In the RTPIS case study we compare BML to the closest related log file analysis language LFAL2. BML can be applied to complicated cases that are not possible to capture using LFAL2, but at the penalty of more setting-up effort. Thus, BML is positioned between the simple specific and the fully general programming languages used in log file analysis. BML fulfills specifically the needs for log analysis on RTPIS kind of systems. Esa Heikkinen, Timo Hämäläinen 0001 |
IECON | 2 |
| 2016 | Designing a clock cycle accurate application with high-level synthesisabstractDuring the recent years, high-level synthesis (HLS) has gained traction as a viable alternative to traditional handwritten register transfer level code in describing digital systems. This has been attributed to the maturing of the HLS tools and improving quality of their results. However, most published applications are data path intensive as HLS offers good tools for loop optimization, such as pipelining and loop unrolling. HLS is seldom applied to control-oriented applications since clock is not explicitly present in HLS source code. In this paper, we show how a clock cycle accurate application can be described with HLS. We give as a proof of concept an implementation of an FPGA-based I2C bus controller for an audio codec using Catapult C, and present a generalized work flow. Compared with a corresponding handwritten VHDL implementation, the HLS version consumes 84% more area at the same performance but productivity is increased by 100% at the first design time and even more with further design iterations. Sakari Lahti, Jarno Vanne, Timo Hämäläinen 0001 |
IECON | 3 |
| 2016 | Live demonstration: Run-time visualization of Kvazaar HEVC intra encoderabstractThis demonstrator presents a run-time visualization tool for Kvazaar HEVC intra encoder. The implemented open-source tool is seamlessly integrated into Kvazaar to provide instant visual feedback of the encoding process. The visualization overlays the reconstructed HEVC video with the boundaries of the used block partitioning structure and associated intra prediction modes. The tool is also able to illustrate all Kvazaar parallelization schemes at run time: Wavefront Parallel Processing, tiles, and picture-level parallel processing. The displayed visualization information can be gradually adjusted to user needs. The tool is primarily designed for Kvazaar debugging but it also suits educational purposes. Marko Viitanen, Ari Koivula, Jarno Vanne, Timo Hämäläinen 0001 |
ISCAS | 4 |
| 2016 | RTP/RTCP Reception Hint Tracks for Video Call Recording and PlaybackabstractThis paper proposes to use RTP (Real-time Transport Protocol) Reception Hint Tracks for convenient recording and playback of a video call in MP4 format. The feasibility of RTP Reception Hint Tracks is validated as a part of an implemented end-to-end video call system. The proposed approach records a bidirectional Linphone video call and multiplexes it as MP4 RTP Reception Hint Tracks with the L-SMASH software library. It also stores RTCP (RTP Control Protocol) Reception Hint Tracks for additional timing information. Playback of the MP4 file is performed with VLC Media Player that is made compatible with RTP Reception Hint Tracks. The proposed proof-of-concept setup meets particularly well the needs of multi-codec solutions where different audio and video codecs can be used for a video call recording and playback. According to our analysis, recording RTP reception Hint Tracks increases the Linphone CPU time by under 1% and the bitrate by under 2% over the bare bitrate of the recorded RTP packets. Joni Räsänen, Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001, Miska M. Hannuksela, Vinod Kumar Malamal Vadakital |
ISM | 4 |
| 2016 | Kvazaar: Open-Source HEVC/H.265 EncoderabstractKvazaar is an academic software video encoder for the emerging High Efficiency Video Coding (HEVC/H.265) standard. It provides students, academic professionals, and industry experts a free, cross-platform HEVC encoder for x86, x64, PowerPC, and ARM processors on Windows, Linux, and Mac. Kvazaar is being developed from scratch in C and optimized in Assembly under the LGPLv2.1 license. The development is being coordinated by Ultra Video Group at Tampere University of Technology (TUT) and the implementation work is carried out by an active community on GitHub. Developer friendly source code of Kvazaar makes joining easy for new developers. Currently, Kvazaar includes all essential coding tools of HEVC and its modular source code facilitates parallelization on multi and manycore processors as well as algorithm acceleration on hardware. Kvazaar is able to attain real-time HEVC coding speed up to 4K video on an Intel 14-core Xeon processor. Kvazaar is also supported by FFmpeg and Libav. These de-facto standard multimedia frameworks boost Kvazaar popularity and enable its joint usage with other well-known multimedia processing tools. Nowadays, Kvazaar is an integral part of teaching at TUT and it has got a key role in three Eureka Celtic-Plus projects in the fields of 4K TV broadcasting, virtual advertising, Video on Demand, and video surveillance. Marko Viitanen, Ari Koivula, Ari Lemmetti, Arttu Ylä-Outinen, Jarno Vanne, Timo Hämäläinen 0001 |
ACM Multimedia | 6 |
| 2016 | Log File Analyzing in Intelligent Transportation Systems Development
Esa Heikkinen, Timo Hämäläinen 0001 |
PROFES | 2 |
| 2016 | PROMOTE: A Process Mining Tool for Embedded System Development
Arttu Leppäkoski, Timo Hämäläinen 0001 |
PROFES | 2 |
| 2015 | High-Level Synthesis Design Flow for HEVC Intra Encoder on SoC-FPGAabstractThis paper presents a High-Level Synthesis (HLS) flow for mapping a software HEVC encoder into Altera CycloneV SoC-FPGA. The starting point is a C implementation of an open-source Kvazaar HEVC intra encoder, which is minimally refined for SystemC design space exploration and automatic Catapult-C RTL generation. The final implementation involves Kvazaar encoder executed in Linux on dual-core ARM, and HW accelerated intra prediction on FPGA. Changing the SW/HW partitioning or modifying the implementation takes hours instead of weeks with Catapult-C HLS. In addition, the design is portable to other platforms without major manual re-writing. We obtained 9 fps full-HD intra prediction speed with a single accelerator on Altera Cyclone V SX on Terasic VEEK-MT-C5SoC board including video capture and HEVC video streaming via Ethernet. To the best of our knowledge, this is the first reported HLS assisted implementation of HEVC encoder on SoC-FPGA. Panu Sjövall, Janne Virtanen, Jarno Vanne, Timo Hämäläinen 0001 |
DSD | 4 |
| 2015 | Resolving parameter reference management in IP-XACT using Kactus2abstractModern VLSI and FPGA chip designs utilize automated generation of the structure and component configuration for different product variations. This is based on re-usable, parametrized library components, and tools for definition, assembly, configuration and generation of final HW and SW code. A product version includes several structural hierarchies, in which each component is independently reusable and must be configured for the specific product context. IEEE 1685 "IP-XACT" standardizes the component and design descriptions and the overall process. Still the challenges are very large parameter space, name-based referencing and propagation of parameter values. Practical user problems are careless parameter renaming, duplicate names, and removing a parameter definition without first removing all references to it. In this paper, we present solutions to these problems implemented in Kactus2 v2.7 that is an open-source IP-XACT tool. Our basis is automatic identifier generation and referencing. This required major changes for Kactus2 import wizards, generators as well as expression editors and evaluators. The implementation was carried out in C++/Qt5 and we modified and added 5k LOC compared to Kactus2 v2.6. According to several use cases analysis the new solution practically eliminates the user errors in the parameter referencing, which significantly improves productivity. Esko Pekkarinen, Mikko Teuho, Erno Salminen, Timo Hämäläinen 0001 |
IECON | 4 |
| 2015 | Kvazaar HEVC encoder for efficient intra codingabstractThis paper presents an open-source Kvazaar encoder for HEVC intra coding. This academic software encoder has been developed from the scratch using C as an implementation language by prioritizing modularity, portability, and readability of the source code. Kvazaar implements almost the same intra coding functionality as HEVC reference encoder (HM) but its rewritten source code makes it significantly faster. In all-intra (AI) coding, a single-threaded C implementation of Kvazaar is 2.3 times faster than HM at a cost of 1.7% bit rate increase. The respective values with a high speed preset of Kvazaar are 10.6 and 8.8%. Compared to a single-threaded C++ implementation of x265, Kvazaar improves rate-distortion performance and increases encoding speed in both high-quality and high-speed test cases. Kvazaar has a particular edge in the high-speed test case where it almost halves the BD-rate loss and more than doubles the performance. Marko Viitanen, Ari Koivula, Ari Lemmetti, Jarno Vanne, Timo Hämäläinen 0001 |
ISCAS | 5 |
| 2015 | Kvazaar HEVC still image coding on Raspberry Pi 2 for low-cost remote surveillanceabstractThis demonstrator serves as a proof-of-concept of our multi-camera remote surveillance system that supports 1080p still image capture with a 10-second refresh rate. Image capture, compression, and broadcast are implemented in each Raspberry Pi 2 camera node. Image compression is conducted with an open-source Kvazaar HEVC encoder that outputs HEVC images in BPG format. The BPG images are broadcast from camera nodes to terminals over the Internet through WebSocket protocol. The images can be played back with most Web browsers in remote locations with Internet access. Marko Viitanen, Ari Koivula, Jarno Vanne, Timo Hämäläinen 0001 |
VCIP | 4 |
| 2014 | Comparative study of 8 and 10-bit HEVC encodersabstractThis paper compares the rate-distortion-complexity (RDC) characteristics of the HEVC Main 10 Profile (M10P) and Main Profile (MP) encoders. The evaluations are performed with HEVC reference encoder (HM) whose M10P and MP are benchmarked with different resolutions, frame rates, and bit depths. The reported RD results are based on bit rate differences for equal PSNR whereas complexities have been profiled with Intel VTune on Intel Core 2 processor. With our 10-bit 4K 120 fps test set, the average bit rate decrements of M10P over MP are 5.8%, 11.6%, and 12.3% in the all-intra (AI), random access (RA), and low-delay B (LB) configurations, respectively. Decreasing the bit depth of this test set to 8 lowers the RD gain of Ml OP only slightly to 5.4% (AI), 11.4% (RA), and 12.1% (LB). The similar trend continues in all our tests even though the RD gain of M10P is decreased over MP with lower resolutions and frame rates. M10P introduces no computational overhead in HM, but it is anticipated to increase complexity and double the memory usage in practical encoders. Hence, the 10-bit HEVC encoding with 8-bit input video is the most recommended option if computation and memory resources are adequate for it. Jarno Vanne, Marko Viitanen, Ari Koivula, Timo Hämäläinen 0001 |
VCIP | 4 |
| 2014 | Efficient Mode Decision Schemes for HEVC Inter PredictionabstractThe emerging High Efficiency Video Coding (HEVC) standard reduces the bit rate by almost 40% over the preceding state-of-the-art Advanced Video Coding (AVC) standard with the same objective quality but at about 40% encoding complexity overhead. The main reason for HEVC complexity is inter prediction that accounts for 60%-70% of the whole encoding time. This paper analyzes the rate-distortion-complexity characteristics of the HEVC inter prediction as a function of different block partition structures and puts the analysis results into practice by developing optimized mode decision schemes for the HEVC encoder. The HEVC inter prediction involves three different partition modes: square motion partition, symmetric motion partition (SMP), and asymmetric motion partition (AMP) out of which the decision of SMPs and AMPs are optimized in this paper. The key optimization techniques behind the proposed schemes are: 1) a conditional evaluation of the SMP modes; 2) range limitations primarily in the SMP sizes and secondarily in the AMP sizes; and 3) a selection of the SMP and AMP ranges as a function of the quantization parameter. These three techniques can be seamlessly incorporated in the existing control structures of the HEVC reference encoder without limiting its potential parallelization, hardware acceleration, or speed-up with other existing encoder optimizations. Our experiments show that the proposed schemes are able to cut the average complexity of the HEVC reference encoder by 31%-51% at a cost of 0.2%-1.3% bit rate increase under the random access coding configuration. The respective values under the low-delay B coding configuration are 32%-50% and 0.3%-1.3%. Jarno Vanne, Marko Viitanen, Timo Hämäläinen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2012 | Complexity analysis of next-generation HEVC decoderabstractThis paper analyzes the complexity of the HEVC video decoder being developed by the JCT-VC community. The HEVC reference decoder HM 3.1 is profiled with Intel VTune on Intel Core 2 Duo processor. The analysis covers both Low Complexity (LC) and High Efficiency (HE) settings for resolutions varying from WQVGA (416 × 240 pixels) up to 1600p (2560 × 1600 pixels). The yielded cycle-accurate results are compared with the respective results of H.264/AVC Baseline Profile (BP) and High Profile (HiP) reference decoders. HEVC offers significant improvement in compression efficiency over H.264/AVC: the average BD-rate saving of LC is around 51% over BP whereas the BD-rate gain of HE is around 45% over HiP. However, the average decoding complexities of LC and HE are increased by 61% and 87% over BP and HiP, respectively. In LC, the most complex functions are motion compensation (MC) and loop filtering (LF) that account on average for 50% and 14% of the decoder complexity. The decoding complexity of HE configuration is on average 42% higher than that of the LC configuration. Majority of the difference is caused by extra LF stages. In HE, the complexities of MC and LF are 37% and 32%, respectively. In practice, a standard 3 GHz dual core processor is expected to be able to decode 1080p HEVC content in real-time. Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001, Moncef Gabbouj, Jani Lainema |
ISCAS | 3 |
| 2012 | MARTE profile extension for modeling dynamic power management of embedded systems
Tero Arpinen, Erno Salminen, Timo Hämäläinen 0001, Marko Hännikäinen |
J. Syst. Archit. | 3 |
| 2012 | Comparative Rate-Distortion-Complexity Analysis of HEVC and AVC Video CodecsabstractThis paper analyzes the rate-distortion-complexity of High Efficiency Video Coding (HEVC) reference video codec (HM) and compares the results with AVC reference codec (JM). The examined software codecs are HM 6.0 using Main Profile (MP) and JM 18.0 using High Profile (HiP). These codes are benchmarked under the all-intra (AI), random access (RA), low-delayB(LB), and low-delayP(LP) coding configurations. In order to obtain a fair comparison, JM HiP anchor codec has been configured to conform to HM MP settings and coding configurations. The rate-distortion comparisons rely on objective quality assessments, i.e., bit rate differences for equal PSNR. The complexities of HM and JM have been profiled at the cycle level with Intel VTune on Intel Core 2 Duo processor. The coding efficiency of HEVC is drastically better than that of AVC. According to our experiments, the average bit rate decrements of HM MP over JM HiP are 23%, 35%, 40%, and 35% under the AI, RA, LB, and LP configurations, respectively. However, HM achieves its coding gain with a realistic overhead in complexity. Our profiling results show that the average software complexity ratios of HM MP and JM HiP encoders are 3.2× in the AI case, 1.2× in the RA case, 1.5× in the LB case, and 1.3× in the LP case. The respective ratios with HM MP and JM HiP decoders are 2.0×, 1.6×, 1.5×, and 1.4×. This paper also reveals the bottlenecks of HM codec and provides implementation guidelines for future real-time HEVC codecs. Jarno Vanne, Marko Viitanen, Timo Hämäläinen 0001, Antti Hallapuro |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Kactus2: Environment for Embedded Product Development Using IP-XACT and MCAPIabstractKey challenge for embedded system companies is management of product configurations over lifecycle. Either new functionality is implemented on an old platform, legacy code on a new one, or both at the same time. We propose Kactus2, a product integration environment suitable for small and mid-size enterprises (SME) utilizing FPGAs. We combine IP-XACT for HW integration and Multicore Association Communications API (MCAPI) for SW integration. The programmer has a uniform view of the system as MCAPI nodes regardless of their implementation in SW or HW. Kactus2 is a work-in-progress open source project implemented in C++ and QT 4.7 currently containing over 40k lines of code and gSOAP as TGI IP-XACT interface. Antti Kamppi, Lauri Matilainen, Joni-Matti Määttä, Erno Salminen, Timo Hämäläinen 0001, Marko Hännikäinen |
DSD | 5 |
| 2011 | Link Quality-Based Channel Selection for Resource Constrained WSNs
Markku Hänninen, Jukka Suhonen, Timo Hämäläinen 0001, Marko Hännikäinen |
GPC | 3 |
| 2009 | Evaluating UML2 modeling of IP-XACT objects for automatic MP-SoC integration onto FPGAabstractIP-XACT is a standard for describing intellectual property metadata for System-on-Chip (SoC) integration. Recently researchers have proposed visualizing and abstracting IP-XACT objects using structural UML2 model elements and diagrams. Despite the number of proposals at conceptual level, experiences on utilizing this representation in practical SoC development environments are very limited. This paper presents how UML2 models of IP-XACT features can be utilized to efficiently design and implement a multiprocessor SoC prototype on FPGA. The main contribution of this paper is the experimental development of a multiprocessor platform on FPGA using UML2 design capture, IP-XACT compatible components, and design automation tools. In addition, modeling concepts are improved from earlier work for the utilized integration methodology. Tero Arpinen, Tapio Koskinen, Erno Salminen, Timo Hämäläinen 0001, Marko Hännikäinen |
DATE | 4 |
| 2009 | Energy-efficient neighbor discovery protocol for mobile wireless sensor networks
Mikko Kohvakka, Jukka Suhonen, Mauri Kuorilehto, Ville Kaseva, Marko Hännikäinen, Timo Hämäläinen 0001 |
Ad Hoc Networks | 6 |
| 2009 | A Configurable Motion Estimation Architecture for Block-Matching AlgorithmsabstractThis paper introduces a configurable motion estimation architecture for a wide range of fast block-matching algorithms (BMAs). Contemporary motion estimation architectures are either too rigid for multiple BMAs or the flexibility in them is implemented at the cost of reduced performance. The proposed architecture overcomes both of these limitations. The configurability of the proposed architecture is based on a new BMA framework that can be adjusted to support the desired set of BMAs. The chosen framework configuration is implemented by an intelligent control logic which is integrated to an efficient parallel memory system and distortion computation unit. The flexibility of the framework is demonstrated by mapping five different BMAs (BBGDS, DS, CDS, HEXBS, and TSS) to the architecture. The total execution time of the mapped BMAs is shown to be almost directly proportional to the number of tested checking points in the search area, so the architecture is very tolerant of different BMA-specific search strategies and search patterns. In addition, a run-time switching between supported BMAs can be done without performance compromises. With a 0.13-mum CMOS technology, the proposed architecture configured for HEXBS, BBGDS, and TSS requires only 14.2 kgates and 2.5 KB of memory at 200 MHz operating frequency. A performance comparison to the reference programmable architectures reveals that only the proposed implementation is able to process real-time (30 fps) fixed block-size motion estimation (1 reference frame) at full HDTV resolution (1920 times1080). Jarno Vanne, Eero Aho, Kimmo Kuusilinna, Timo Hämäläinen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | Rapid design and evaluation framework for wireless sensor networks
Mauri Kuorilehto, Marko Hännikäinen, Timo Hämäläinen 0001 |
Ad Hoc Networks | 3 |
| 2008 | Performance model for IEEE 802.11s wireless mesh network deployment design
Timo Vanhatupa, Marko Hännikäinen, Timo Hämäläinen 0001 |
J. Parallel Distributed Comput. | 3 |
| 2008 | Editorial
Mladen Berekovic, Andy D. Pimentel, Timo Hämäläinen 0001 |
J. Syst. Archit. | 3 |
| 2008 | A Parallel Memory System for Variable Block-Size Motion Estimation AlgorithmsabstractThis paper proposes an efficient parallel memory system for algorithms applied in fixed and variable block-size motion estimation (VBSME). The proposed system is implemented by a novel combination of two parallel memory architectures. The distribution of data among the memory modules is modified over contemporary approaches and the optimized address computation unit enables a rapid address generation for accessed memory locations. Furthermore, the introduced data permutation scheme organizes data efficiently for storage and retrieval. The proposed system enables up to 4 X speedup in data storage and retrieves data up to 55% faster for VBSME compared with the reference implementations. With a 0.18- mum CMOS technology, the proposed memory addressing and data permutation scheme can be clocked at 980 MHz operating frequency with a cost of less than 6 kgates. On FPGA, the system can operate at 200 MHz with less than 700 logic elements. The results show that the proposed system is applicable to real-time VBSME at HDTV resolution. Jarno Vanne, Eero Aho, Timo Hämäläinen 0001, Kimmo Kuusilinna |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | Compact hardware design of Whirlpool hashing core
Timo Alho, Panu Hämäläinen, Marko Hännikäinen, Timo Hämäläinen 0001 |
DATE | 4 |
| 2007 | Cost-aware capacity optimization in dynamic multi-hop WSNsabstractLow energy consumption and load balancing are required for enhancing lifetime at wireless sensor networks (WSN). In addition, network dynamics and different delay, throughput, and reliability requirements demand cost-aware traffic adaptation. This paper presents a novel capacity optimization algorithm targeted at locally synchronized, low-duty cycle WSN MACs. The algorithm balances the traffic load between contention and contention free channel access. The energy-inefficient contention access is avoided, whereas the more reliable contention free access is preferred. The algorithm allows making cost-aware trade-off between delay, energy-efficiency, and throughput guided by routing layer. Analysis results show that the algorithm has 10% to 100% better energy-efficiency than IEEE 802.15.4 LR-WPAN in a typical sensing application, while providing comparable good put and delay Jukka Suhonen, Mikko Kohvakka, Mauri Kuorilehto, Marko Hännikäinen, Timo Hämäläinen 0001 |
DATE | 5 |
| 2007 | Evaluating the Model Accuracy in Automated Design Space ExplorationabstractDesign space exploration is used to shorten the design time of System-on-Chips (SoCs). The models used in the exploration need to be both accurate and fast to simulate. This paper introduces a multi-level communication cost to improve the accuracy of the abstracted system models. During the simulation, one of three different communication costs is applied for each inter-task communication event based on the mapping of the communicating tasks. The accuracy of three system abstraction models including the presented communication cost is evaluated using a Motion-JPEG (M-JPEG) application described in Unified Modeling Language (UML). According to the results, the average error in frames per second (FPS) is 3.8% for the trace model, 4.3% for the modulo model, and 12.8% for the probabilistic model compared to FPGA execution. The results show that with the multi-level communication cost the accuracy is increased significantly, and accurate results can be achieved with arbitrary mappings. Kalle Holma, Mikko Setälä, Erno Salminen, Timo Hämäläinen 0001 |
DSD | 4 |
| 2007 | On network-on-chip comparisonabstractThis paper presents the state-of-the-art in the field of network-on-chip (NoC) benchmarking and comparison. The study identifies the mainstream approaches, how NoCs are currently evaluated, and shows which aspects have been covered and those needing more research effort. No single article can cover all the aspects, and therefore, possibility to compare results from various sources must be ensured by proper scientific reporting. Basic guidelines for achieving that are given. Erno Salminen, Ari Kulmala, Timo Hämäläinen 0001 |
DSD | 3 |
| 2007 | Modeling Embedded Software Platforms with a UML Profile
Tero Arpinen, Mikko Setälä, Petri Kukkala, Erno Salminen, Marko Hännikäinen, Timo Hämäläinen 0001 |
FDL | 6 |
| 2007 | Evaluation of throughput estimation models and algorithms for WLAN frequency planning
Timo Vanhatupa, Marko Hännikäinen, Timo Hämäläinen 0001 |
Comput. Networks | 3 |
| 2007 | Editorial
Timo Hämäläinen 0001, Stephan Wong, C. John Glossner, Stamatis Vassiliadis |
J. Syst. Archit. | 1 |
| 2007 | Automated memory-aware application distribution for Multi-processor System-on-Chips
Heikki Orsila, Tero Kangas, Erno Salminen, Timo Hämäläinen 0001, Marko Hännikäinen |
J. Syst. Archit. | 4 |
| 2007 | Benchmarking mesh and hierarchical bus networks in system-on-chip context
Erno Salminen, Tero Kangas, Vesa Lahtinen, Jouni Riihimäki, Kimmo Kuusilinna, Timo Hämäläinen 0001 |
J. Syst. Archit. | 6 |
| 2007 | Editorial
Jarmo Takala, Timo Hämäläinen 0001, Andy D. Pimentel, Stamatis Vassiliadis |
J. Syst. Archit. | 2 |
| 2006 | Configurable multiprocessor platform with RTOS for distributed execution of UML 2.0 designed applicationsabstractThis paper presents the design and full prototype implementation of a configurable multiprocessor platform that supports distributed execution of applications described in UML 2.0. The platform is comprised of multiple Altera Nios II softcore processors and custom hardware accelerators connected by the heterogeneous IP block interconnection (HIBI) communication architecture. Each processor has a local copy of eCos real-time operating system for the scheduling of multiple application threads. The mapping of a UML application into the proposed platform is presented by distributing a WLAN medium access control protocol onto multiple CPUs. The experiments performed on FPGA show that our approach raises system design to a new level. To our knowledge, this is the first real implementation combining a high-level design flow with a synthesizable platform Tero Arpinen, Petri Kukkala, Erno Salminen, Marko Hännikäinen, Timo Hämäläinen 0001 |
DATE | 5 |
| 2006 | Design and Implementation of Low-Area and Low-Power AES Encryption Hardware CoreabstractThe Advanced Encryption Standard (AES) algorithm has become the default choice for various security services in numerous applications. In this paper we present an AES encryption hardware core suited for devices in which low cost and low power consumption are desired. The core constitutes of a novel 8-bit architecture and supports encryption with 128-bit keys. In a 0.13 mum CMOS technology our area optimized implementation consumes 3.1 kgates. The throughput at the maximum clock frequency of 153 MHz is 121 Mbps, also in feedback encryption modes. Compared to previous 8-bit implementations, we achieve significantly higher throughput with corresponding area. The energy consumption per processed block is also lower Panu Hämäläinen, Timo Alho, Marko Hännikäinen, Timo Hämäläinen 0001 |
DSD | 4 |
| 2006 | Comparison of GALS and Synchronous Architectures with MPEG-4 Video Encoder on Multiprocessor System-on-Chip FPGAabstractIn large system-on-chip (SoC) architectures, balancing the clock network is increasingly difficult. Globally asynchronous locally synchronous (GALS) removes the need for global clock net, and also provides efficient means for managing the complexity and re-use in large architectures. However, quantitative comparisons of GALS against similar synchronous structures are rare for full SoC architectures. In this paper, we compare our SoC GALS architectures to a synchronous architecture with a fully functional MPEG-4 video encoder on FPGA. The results show that the area and performance overhead of GALS is only 1%. That is negligible compared to the benefits of the GALS architecture such as multiple clock frequencies for intellectual property (IP) blocks and dynamic frequency/voltage scaling, clock tree removal, and re-usability. Our architecture does not require modifications to the IP blocks already used with synchronous architectures, providing an ideal solution for rapid switch to GALS architecture Ari Kulmala, Timo Hämäläinen 0001, Marko Hännikäinen |
DSD | 2 |
| 2006 | Reliable GALS Implementation of MPEG-4 Encoder with Mixed Clock FIFO on Standard FPGAabstractGlobally asynchronous locally synchronous (GALS) is a paradigm for complexity management and re-use of large system-on-chip (SoC) architectures. GALS is most often based on specific ASIC design components or special FPGA platforms with custom development tools. In this paper we present a multiprocessor GALS implementation on a standard commercial FPGA with standard development tools. The key building block is a novel, reliable RTL mixed clock FIFO. A complete MPEG-4 video encoder with four processors is implemented for proofing the concept. The area overhead compared to a fully synchronous design is shown to be only 2% and the performance overhead is 3%. This is negligible compared to the benefits that are much better flexibility, ASIC or FPGA vendor independency, and reduced design time. Furthermore, the mixed-clock interfaces allow easy re-usability, since the RTL-level blocks do not need to be re-verified in design iterations Ari Kulmala, Timo Hämäläinen 0001, Marko Hännikäinen |
FPL | 2 |
| 2006 | WSN API: Application Programming Interface for Wireless Sensor NetworksabstractIn this paper, an application programming interface for wireless sensor networks (WSN API) is presented. The WSN API consists of a client-side API (gateway API) and a sensor-side API (node API). The WSN API conceals the complexities of WSN communication protocols and architectures, and provides a well-defined and easy-to-use way to collect data from sensors. Also, easy expandability for new sensor components and applications is provided. The WSN API is implemented practically in TUTWSN prototype platforms Jari K. Juntunen, Mauri Kuorilehto, Mikko Kohvakka, Ville Kaseva, Marko Hännikäinen, Timo Hämäläinen 0001 |
PIMRC | 6 |
| 2006 | Transmission Power Based Path Loss Metering for Wireless Sensor NetworksabstractThe metering of path loss to neighbouring nodes is vital in multi-hop wireless sensor networks (WSN) for managing network self-configuration, robust data routing, and node localization. Conventionally, path loss is measured by a received signal strength indicator (RSSI) mechanism of a transceiver. Yet, the lowest hardware complexity, energy consumption and cost are reached by transceivers without RSSI. We present a new simple transmission power based path loss metering method for WSNs without RSSI mechanism. The method determines path loss from frames transmitted at different power levels. Performance measurements with physical WSN prototypes indicate sufficient accuracy for network management. According to performance analysis, the power consumption of the proposed method is below 10 mu. An energy analysis in a large IEEE 802.15.4 network indicates 51% to 69% energy saving compared to RSSI-equipped transceivers by using the path loss metering method together with a simple and low power transceiver Mikko Kohvakka, Jukka Suhonen, Marko Hännikäinen, Timo Hämäläinen 0001 |
PIMRC | 4 |
| 2006 | Configurable Protocol Engine for Runtime-Configurable Communication Subsystems on Multiprocessor SoCabstractThis paper presents a configurable protocol engine (CPE) to implement runtime-configurable communication subsystems, which are able to adapt their protocol stacks to varying service requirements. The communication subsystems with CPE are designed and implemented using a UML-based design methodology and automated design flow. CPE has been applied to implementing wireless protocol stacks on multiprocessor system-on-chip (SoC) platforms. As a design case study, we present the implementation of a WSN-to-WLAN bridge on a multiprocessor SoC on FPGA. Experiences with CPE proved its feasibility in rapid implementation of communication subsystems with very decent performance Petri Kukkala, Marko Hännikäinen, Timo Hämäläinen 0001 |
PIMRC | 3 |
| 2006 | Experimenting TCP/IP for Low-Power Wireless Sensor NetworksabstractThis paper presents the analysis and real experiments on TCP/IP communication in a low-power monitoring wireless sensor network (WSN). TCP/IP flow control, addressing, and packet fragmentation are adapted in a gateway that relays TCP/IP communication to WSN. The performance of TCP/IP is evaluated between endpoint PCs communicating over TUTWSN (Tampere University of Technology WSN), which is used as a transparent communication medium for TCP/IP data. The evaluation results show that the window-based flow control algorithms of TCP perform too aggressively in WSNs, where random bit errors and topology changes are main reasons for errors instead of congestion. Further, a frequent duty cycle is needed in WSN to compensate the TCP/IP overhead. In TUTWSN, compared to an environmental monitoring application with two second activity cycle and native WSN transport, the enabling of TCP/IP consumes five times more power Mauri Kuorilehto, Jukka Suhonen, Mikko Kohvakka, Marko Hännikäinen, Timo Hämäläinen 0001 |
PIMRC | 5 |
| 2006 | Cost-Aware Dynamic Routing Protocol for Wireless Sensor Networks - Design and Prototype ExperimentsabstractThis paper presents an energy-efficient multi-hop routing protocol for wireless sensor networks. The protocol uses cost metrics to create gradients from a source to a destination node. The cost metrics consist of energy, node load, delay, and link reliability information that provide a trade-off between performance and energy usage. A node can query routes from its neighbors, which allows efficient recovery from route losses. The protocol is the first cost-field based WSN routing protocol suitable for low processing and memory capacity nodes that is tested in a practical real-world environment. The protocol performance is evaluated on full scale prototype implementation consisting of 38 ultra-low power nodes in indoor environment. Compared to traditional flooding, the protocol requires only 25% of the bandwidth, while having smaller end-to-end delays Jukka Suhonen, Mauri Kuorilehto, Marko Hännikäinen, Timo Hämäläinen 0001 |
PIMRC | 4 |
| 2006 | Evaluation of throughput estimation models and algorithms for WLAN frequency planningabstractFrequency optimization is required to maximize the WLAN throughput in environments where several networks coexist. This paper evaluates four throughput estimation models and two optimization algorithms. Throughput is selected as the optimization criteria for channel assignment. Thus, the result of a throughput estimation model is used as an input for an optimization algorithm. The target is a frequency plan that maximizes multi-cell WLAN throughput. The throughput estimation models are based on radio spectrum usage, practical throughput measurements, WLAN protocol behavior, and theoretical coverage estimations. The models use separate functions for defining the minimum channel distance. In the evaluation, Genetic Algorithm (GA) and a distributed optimization algorithm produce the final frequency plan. A dedicated simulator has been implemented for the evaluation. As overall results, the evaluated throughput estimation models did not produce significant WLAN throughput improvements compared to each other. Still, the selection of the throughput estimation model and optimization algorithm pair is significant for results, since certain combinations cause poor results. Timo Vanhatupa, Marko Hännikäinen, Timo Hämäläinen 0001 |
QSHINE | 3 |
| 2006 | A High-Performance Sum of Absolute Difference Implementation for Motion EstimationabstractThis paper presents a high-performance sum of absolute difference (SAD) architecture for motion estimation, which is the most time-consuming and compute-intensive part of video coding. The proposed architecture contains novel and efficient optimizations to overcome bottlenecks discovered in existing approaches. In addition, designed sophisticated control logic with multiple early termination mechanisms further enhance execution speed and make the architecture suitable for general-purpose usage. Hence, the proposed architecture is not restricted to a single block-matching algorithm in motion estimation, but a wide range of algorithms is supported. The proposed SAD architecture outperforms contemporary architectures in terms of execution speed and area efficiency. The proposed architecture with three pipeline stages, synthesized to a 0.18-mum CMOS technology, can attain 770-MHz operating frequency at a cost of less than 5600 gates. Correspondingly, performance metrics for the proposed low-latency 2-stage architecture are 730 MHz and 7500 gates Jarno Vanne, Eero Aho, Timo Hämäläinen 0001, Kimmo Kuusilinna |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | UML-based multiprocessor SoC design frameworkabstractThis paper describes a complete design flow for multiprocessor systems-on-chips (SoCs) covering the design phases from system-level modeling to FPGA prototyping. The design of complex heterogeneous systems is enabled by raising the abstraction level and providing several system-level design automation tools. The system is modeled in a UML design environment following a new UML profile that specifies the practices for orthogonal application and architecture modeling. The design flow tools are governed in a single framework that combines the subtools into a seamless flow and visualizes the design process. Novel features also include an automated architecture exploration based on the system models in UML, as well as the automatic back and forward annotation of information in the design flow. The architecture exploration is based on the global optimization of systems that are composed of subsystems, which are then locally optimized for their particular purposes. As a result, the design flow produces an optimized component allocation, task mapping, and scheduling for the described application. In addition, it implements the entire system for FPGA prototyping board. As a case study, the design flow is utilized in the integration of state-of-the-art technology approaches, including a wireless terminal architecture, a network-on-chip, and multiprocessing utilizing RTOS in a SoC. In this study, a central part of a WLAN terminal is modeled, verified, optimized, and prototyped with the presented framework. Tero Kangas, Petri Kukkala, Heikki Orsila, Erno Salminen, Marko Hännikäinen, Timo Hämäläinen 0001, Jouni Riihimäki, Kimmo Kuusilinna |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2005 | UML 2.0 Profile for Embedded System DesignabstractThe unified modeling language (UML) 2.0 is emerging in the area of embedded system design. This paper presents a new UML 2.0 profile - called TUT-profile - that introduces a set of stereotypes and design rules for an application, platform, and mapping. The profile classifies different application and platform components, and enables their parameterization. The TUT-profile concentrates on the structure of an application and platform, and utilizes standard UML 2.0 for the behavioral modeling. The application is seen as a set of active classes with an internal behavior. Correspondingly, the platform is seen as a component library with a parameterized presentation in UML 2.0 for each library component. Petri Kukkala, Jouni Riihimäki, Marko Hännikäinen, Timo Hämäläinen 0001, Klaus Kronlöf |
DATE | 4 |
| 2005 | Design of Transport Triggered Architecture Processors for Wireless EncryptionabstractTransport triggered architecture (TTA) offers a cost-effective trade-off between the size and performance of ASICs and the programmability of general-purpose processors. In this paper TTA processors for the RC4 and AES encryption algorithms of the new IEEE 802.11i WLAN security standard are designed. Special operations efficiently supporting the ciphers are developed. The TTA design flow is utilized for finding configurations with the best performance-size ratios. The size of the configuration supporting both the algorithms is 69.4 kgates and the throughput 100 Mb/s for RC4 and 68.5 Mb/s for AES at 100 MHz in the 0.13 /spl mu/m CMOS technology. Compared to commercial processors of the same wireless application domain, higher throughputs are achieved at significantly smaller area and lower clock speed, which also results in decreased energy consumption. Panu Hämäläinen, Jari Heikkinen, Marko Hännikäinen, Timo Hämäläinen 0001 |
DSD | 4 |
| 2005 | Wireless Sensor Network Implementation for Industrial Linear Position MeteringabstractThis paper presents the design and performance measurements of a prototype wireless sensor network (WSN) for industrial linear position metering. Design includes two different prototype platforms and a user application. Prototypes combine energy efficient commercial off-the-shelf components including a 2.4 GHz radio, and the custom TUTWSN communication protocols resulting high robustness, autonomous operation and very low power consumption. The user application displays sensor data graphically and enables further data analysis. Measurements contain component power analysis and prototype performance measurements. The measurements indicate 200 /spl mu/W to 400 /spl mu/W average node power consumption, as 16-bit sample is measured with 1 Hz sample rate, and routed to a WSN gateway with 1 s latency per hop and 512 bps throughput between nodes. Predicted lifetime of implemented WSN is 2 months with a small rechargeable battery or over 2 years with two AA batteries. Mikko Kohvakka, Marko Hännikäinen, Timo Hämäläinen 0001 |
DSD | 3 |
| 2005 | Co-simulation of Wireless Local Area Network Terminals with Protocol Software Implemented in SDLabstractThis paper presents the verification of our WLAN terminal (TUTWLAN), with its medium access control protocol and test applications, using cycle-accurate hardware/ software co-simulation. The protocol software has been implemented using SDL and automatic C code generation. The hardware implementation of the terminal contains hardware accelerators for time-critical protocol functions. Full system co-simulations were used for both the functional verification and performance evaluation of a single TUTWLAN terminal as well as a network of terminals. With simulations, the performance bottlenecks were identified, and the results enable the implementing of the next generation TUTWLAN terminal as a single-chip. Petri Kukkala, Marko Hännikäinen, Timo Hämäläinen 0001 |
DSD | 3 |
| 2005 | A Parallel MPEG-4 Encoder for FPGA Based Multiprocessor SoCabstractA parallel MPEG-4 simple profile encoder for FPGA based multiprocessor system-on-chip (SoC) is presented. The goal is a computationally scalable framework independent of platform. The scalability is achieved by spatial parallelization where images are divided to horizontal slices. Slice coding tasks are mapped to the multiprocessor consisting of four soft-cores arranged into master-slave configuration. Also, the shared memory model is adopted where large images are stored in shared external memory while small on-chip buffers are used for processing. The interconnections between memories and processors are realized with our HIBI network. Our main contributions are the scalable encoder framework as well as methods for coping with limited memory of FPGA. The current software only implementation processes 6 QCIF frames/s with three encoding slaves. In practice, speed-ups of 1.7 and 2.3 have been measured with two and three slaves, respectively. FPGA utilization of current implementation is 59% requiring 24 207 logic elements on Altera Stratix EP1S40. Olli Lehtoranta, Erno Salminen, Ari Kulmala, Marko Hännikäinen, Timo Hämäläinen 0001 |
FPL | 5 |
| 2005 | Software and hardware prototypes of the IEEE 1588 precision time protocol on wireless LANabstractIEEE 1588 is a standard for precise clock synchronization for networked measurement and control systems in LAN environment. This paper presents the design and implementation of two IEEE 1588 prototypes for wireless LAN (WLAN). The first one is implemented using a Linux PC platform and a standard IEEE 802.11 WLAN with modifications to the network device driver. The second prototype is implemented using an embedded WLAN development board that implements the synchronization functionality using an embedded processor with programmable logic device (PLD) circuits. The measured results show that 1.1 ns average clock offset can be reached on HW based implementation, while Linux PC network driver enables 660 ns with a standard WLAN. Although WLAN is an extremely difficult environment for the synchronization, the results achieved with the prototype are fully comparable to those achieved with wired LAN implementations. Juha Kannisto, Timo Vanhatupa, Marko Hännikäinen, Timo Hämäläinen 0001 |
LANMAN | 4 |
| 2005 | Frequency management tool for multi-cell WLAN performance optimizationabstractCareful configuration of frequencies for WLAN access points (AP) is crucial for ensuring the optimal performance of the network. Thus, advanced tools for WLAN frequency management tasks are needed. This paper presents a practical tool for the frequency management in IEEE 802.11 WLANs. The tool minimizes the interference between APs and consequently maximizes the effective capacity of the network. It also contains a graphical user interface (GUI) showing an illustrative view of the network state in frequency domain. The tool has been designed for administrators of medium and large WLAN networks, such as used in companies, airports, and campus areas. In the presented case, the throughput of the optimized network was 60 % higher compared to the original network with random channels Timo Vanhatupa, Marko Hännikäinen, Timo Hämäläinen 0001 |
LANMAN | 3 |
| 2005 | Ultra low energy wireless temperature sensor network implementationabstractCondition monitoring in buildings is one of the most potential and foreseen applications for wireless sensor networks (WSN). This paper presents the design and full scale prototype implementation of WSN for temperature monitoring. The prototypes are implemented using low power commercial of-the-shelf components including a 2.4 GHz radio, microcontrollers, and a custom TUTWSN communication protocol. A user application provides a graphical data analysis. Measurements indicate 183 muW to 390 muW average node power consumptions, as temperature is measured at 5 s intervals and data is multi-hop routed to a gateway. Predicted lifetime with two AA batteries is up to 4.9 years. In addition, experiments indicate that time accuracy is extremely important in hardware prototypes Mikko Kohvakka, Marko Hännikäinen, Timo Hämäläinen 0001 |
PIMRC | 3 |
| 2005 | Energy optimized beacon transmission rate in a wireless sensor networkabstractResource constrained wireless sensor networks (WSN) require an energy efficient medium access control (MAC) protocol that minimizes the radio active time (duty cycle). Time slotted MAC schemes provide lowest duty cycles by dividing time into consecutive data exchange and sleep periods. Synchronization for data exchange and network maintenance is achieved by exchanging beacons. For detecting changes in network topology, nodes periodically perform scanning during which beacons are received from neighbors. This is energy consuming, and the energy required equals to the transmission of thousands of packets. This paper shows that the energy consumption is mainly depending on the beacon transmission rate, and that an optimal rate is a function of three parameters: a network scanning interval, beacon transmission energy, and radio reception power. The optimal beacon transmission rate is derived for a TUTWSN prototype by power analysis and energy models. According to the analysis, the optimization decreases the average network energy consumption up to an order of magnitude. For the prototype, the optimal beacon transmission rate is 3.7 Hz, when network scanning is performed with 2 minutes intervals Mikko Kohvakka, Marko Hännikäinen, Timo Hämäläinen 0001 |
PIMRC | 3 |
| 2005 | A middleware for task allocation in wireless sensor networksabstractResource constrained platforms, dynamic nature, and complex applications set challenges to the development of wireless sensor networks (WSN). Sophisticated tasking and networking control is required in WSNs to reach lifetimes in order of years. This paper presents a WSN node middleware, which controls task allocation and WSN topology according to the current requirements of the application. The middleware uses a lightweight algorithm that balances communication and computation load between nodes. The discovering of resources and application tasks are comprised by a tuple space that selectively disperses information to nodes. The middleware has been implemented and evaluated in a wireless sensor network simulator (WISENES) that models resource usage and network operation accurately. The results show that in a static network configuration the obtained lifetime with our middleware is 6.8 times longer compared to an uncontrolled network while the increase in processing is negligible and the peek data memory usage increases by 11.6%. In a dynamically changing network the lifetime increases by a factor 3.9. Our middleware does not limit the applications and networks and improves the performance and predictability of WSNs significantly Mauri Kuorilehto, Marko Hännikäinen, Timo Hämäläinen 0001 |
PIMRC | 3 |
| 2005 | Multihop IEEE 802.11b WLAN Performance for VoIPabstractThis paper evaluates the performance of IEEE 802.11b WLAN for supporting multihop voice over IP (VoIP) service. Evaluation is carried out using the NS-2 network simulator and the mean opinion score (MOS) as a criteria for measuring the quality of a VoIP connection. The results show that the mean number of hops between a VoIP transmitter and receiver has the main effect on the number of calls with acceptable quality. On a small network where connections cause interference to each other already three hops cause problems. The mean number of hops can be decreased with a supporting access point (AP) infrastructure. Also the type of interfering traffic affects. The voice quality in VoIP is sensitive to transmission losses, and VoIP cannot compete equally with high data rate applications Timo Vanhatupa, Marko Hännikäinen, Timo Hämäläinen 0001 |
PIMRC | 3 |
| 2005 | Comments on "Winscale: an image-scaling algorithm using an area pixel Model"abstractIn the paper by Kim et al. (2003), the authors propose a new image scaling method called winscale. The presented method can be used for scaling up and down. However, scaling down utilizing the winscale concept gives exactly the same results as the well-known bilinear interpolation. Furthermore, compared to bilinear, scaling up with the proposed winscale "overlap stamping" method has very similar calculations. The basic winscale upscaling differs from the bilinear method. Eero Aho, Jarno Vanne, Kimmo Kuusilinna, Timo Hämäläinen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2004 | UML 2.0 implementation of an embedded WLAN protocolabstractThis paper evaluates the suitability of the new Unified Modelling Language (UML) 2.0 for the design and implementation of embedded wireless local area network (WLAN) protocols. UML 2.0 introduces several new features and improvements to semantics, diagram types, and notations of the language. Behaviour in UML 2.0 can be described accurately and precisely, which enables automatic source code (C/C++ for example) generation for creating executable models. A medium access control (MAC) protocol called TUTMAC was designed and implemented with UML 2.0, and integrated into a hardware platform. The implementation with UML 2.0 reached adequate performance for the protocol with time-critical functionality on the target platform. Petri Kukkala, Väinö Helminen, Marko Hännikäinen, Timo Hämäläinen 0001 |
PIMRC | 4 |
| 2003 | Distributing SoC Simulations over a Network of ComputersabstractThis paper presents parallelization of System-on-Chip (SoC) simulation over a network of computers. A custom C-language-based SoC exploration and simulation tool, Discrete Time Network Simulator (DTNS), is used to examine the problem. Parallelization is implemented with either CORBA or TCP/IP sockets. The distributed DTNS architecture facilitates the analysis of computation requirements for reasonable distribution. The minimum execution time per distributed process for the parallelization to be profitable is 1.2ms with CORBA and 0.4ms with TCP/IP implementation. These results are based on our networked PCs running the Linux operating system. The same network is used to evaluate this distribution method in case of video encoder SoC simulation. Jouni Riihimäki, Väinö Helminen, Kimmo Kuusilinna, Timo Hämäläinen 0001 |
DSD | 4 |
| 2003 | Design of a Management System for Wireless Home Area Networking
Tapio Rantanen, Janne Sikiö, Marko Hännikäinen, Timo Vanhatupa, Olavi Karasti, Timo Hämäläinen 0001 |
Euro-Par | 6 |
| 2003 | Offline architecture for real-time bettingabstractTraditional betting systems do not enable betting in real-time during an event and require decisions and preparations beforehand. In this paper a novel architecture for real-time betting is presented. It disposes of the up-front effort and enables frequent bet announcements and placements during an ongoing event. The novel situation is achieved by broadcasting announcements, time-stamping and storing the placements locally, and collecting them after the event has been finished. While solving the processing problems, the architecture requires reliable cryptographic and physical protection. Currently, e.g. DVB and LAN technologies offer potential platforms for providing the service. The implemented LAN demonstrator has shown that user interfaces have to be simple and bets should not be announced too often. It has also shown that real-time operation makes betting more inspiring. Panu Hämäläinen, Marko Hännikäinen, Timo Hämäläinen 0001, Riku Soininen |
ICME | 3 |
| 2003 | Positioning with IEEE 802.11b wireless LANabstractThis paper presents the design and implementation of a local positioning prototype. The prototype implements received signal power level based positioning on IEEE 802.11b wireless LAN (WLAN) platform. In addition to a WLAN adapter, the prototype includes a host PC, which runs a positioning application. The application computes and displays position estimates on the basis of measurements performed by the WLAN adapter. A signal propagation model and an extended Kalman filter are used for enhancing the position estimates. With the used WLAN hardware, the mean absolute error of positioning was measured to be 2.6 m. Antti Kotanen, Marko Hännikäinen, Helena Leppäkoski, Timo Hämäläinen 0001 |
PIMRC | 4 |
| 2003 | Class of service control protocol for wireless LANabstractThis paper presents a class of service (CoS) protocol to be used on wireless LANs (WLAN). The protocol provides two modes for bandwidth control and uses centrally controlled topology. The first mode reserves a fixed bandwidth for a station, while in the second mode the station must request a permission to send data. A traffic scheduler for the network is located in the central controller station. A prototype of the protocol operating on IEEE802.11 WLAN has been implemented for testing. Results show that the protocol provides CoS support that differentiates bandwidth and transfer delay for each class. Jukka Suhonen, Marko Hännikäinen, Timo Hämäläinen 0001 |
PIMRC | 3 |
| 2003 | Comparison of video protection methods for wireless networks
Olli Lehtoranta, Jukka Suhonen, Marko Hännikäinen, Ville Lappalainen, Timo Hämäläinen 0001 |
Signal Process. Image Commun. | 5 |
| 2003 | Complexity of optimized H.26L video decoder implementationabstractAn analysis of computational complexity is presented for an H.26L video decoder, based on extensive experiments on a general-purpose processor. In addition, platform-independent techniques to optimize an H.26L decoder implementation are given. Comparisons are carried out between our highly optimized version of H.26L, the public reference implementation of H.26L, and a highly optimized H.263+ implementation. Both QCIF and CIF-sized image sequences are used. The results show that with equal visual quality, the bit-rate savings range from 28% to 58%, while the frame decoding speed of H.26L is about 11% better than that of a highly optimized H.263+. Ville Lappalainen, Antti Hallapuro, Timo Hämäläinen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2002 | Enhanced Configurable Parallel Memory ArchitectureabstractContemporary multimedia processors and applications are increasingly limited by their data accessing capabilities. However, the designed Configurable Parallel Memory Architecture (CPMA) alleviates these multimedia data accessing requirements; achieving significant performance improvements over traditional memory architectures. CPMA decreases considerably the processor-memory bottleneck by widening the memory bandwidth, decreasing the number of memory accesses, and diminishing the significance of memory latency. To further enhance the performance of CPMA, this paper introduces a novel architectural extension called CPMA access instruction correlation recognition. The presented method is intended for accelerating the execution rate of consecutive, temporally conflict-free, CPMA memory accesses. As demonstrated in this paper, the superior CPMA performance can also be maintained in the case of limited access widths. In addition, the presented results confirm that CPMA can have an acceptable silicon area. Jarno Vanne, Eero Aho, Kimmo Kuusilinna, Timo Hämäläinen 0001 |
DSD | 4 |
| 2002 | Trends in personal wireless data communications
Marko Hännikäinen, Timo Hämäläinen 0001, Markku Niemi, Jukka Saarinen |
Comput. Commun. | 2 |
| 2002 | Configurable parallel memory architecture for multimedia computers
Kimmo Kuusilinna, Jarno K. Tanskanen, Timo Hämäläinen 0001, Jarkko Niittylahti |
J. Syst. Archit. | 3 |
| 2002 | Overview of research efforts on media ISA extensions and their usage in video codingabstractThis paper summarizes the results of over 25 research groups or individual researchers that have presented video coding implementations on general-purpose processors with the new single instruction multiple data media instruction set architecture extensions. The extensions are introduced and the fundamentals for extensions, as well as some inherent problems, are explained. The reported attempts to utilize the extensions are divided into kernel- and application-level, as well as platform dependent and independent optimizations. Optimized applications include, in addition to some proprietary methods, all of the major video coding standards such as H.261, H.263, MPEG-4, MPEG-1, and MPEG-2. These optimized implementations include a complete video codec, several decoders, and several encoders. Additionally, a performance comparison is given for four representative encoder implementations based on the reported results. Also included is an overview of future trends for new instructions and architectural speed-up techniques. Ville Lappalainen, Timo Hämäläinen 0001, Petri Liuha |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | Performance analysis of low bit rate H.26L video encoderabstractA new video encoder proposal, H.26L, is compared against H.263 and H.263+. In the comparison, both computational complexity and compression performance are analyzed. Moreover, the trade-off possibilities between the complexity and compression performance within H.26L are presented. Experimental comparisons with H.263 and H.263+ show that H.26L reduces the output bit rate about 30% with the same quality. The computation time increases about three times compared to H.263 and leads into the encoding speed of 3-6 fps for QCIF sequences on a 400 MHz Pentium III processor. Realtime operation can be achieved by applying additional, algorithmic and platform-specific optimizations. Antti Hallapuro, Ville Lappalainen, Timo Hämäläinen 0001 |
ICASSP | 3 |
| 2001 | Configurable hardware implementation of triple-DES encryption algorithm for wireless local area networkabstractThis paper presents three implementations of triple data encryption standard (3DES) algorithm on a configurable platform. Implementations are aimed at the medium access control (MAC) protocol of a multimedia-capable wireless local area network (WLAN). For this reason, very strict timing constraints as well as demands for area-efficiency are present. The MAC processing is handled by a digital signal processor (DSP) and a Xilinx Virtex field programmable gate array (FPGA) chip. The latter one is also used for the presented encryption implementations. As a result of the study, 3DES implementations with small area and reasonable throughput and, on the contrary, with large area and very high throughput are realized. Even though 3DES turns out to be quite large and resource-demanding, the implementations still leave enough chip area for the other MAC functions. Consequently, the set requirements are met and the cipher can be integrated into the system. Panu Hämäläinen, Marko Hännikäinen, Timo Hämäläinen 0001, Jukka Saarinen |
ICASSP | 3 |
| 2001 | Architecture of a passenger information system for public transport servicesabstractA passenger information system (PIS) called TUTPIS has been developed for networking passengers with companies that provide public transport services. TUTPIS supports a passenger with personalised, real time information services in all phases of a journey. Services include timetables, travel route searching, route reservations, and electronic payment. TUTPIS is targeted to operate on the developing wireless network infrastructure of mobile, local, and personal networking technologies. A TUTPIS model concentrating on bus services has been implemented using the specification and description language (SDL). The model contains functional implementations of a TUTPIS server and mobile terminals. In addition, the behaviour of external entities, such as buses, trains, and third party systems is implemented for simulations. By simulations, the operability of TUTPIS can be verified and performance estimations for system servers, wireless network throughputs, and mobile terminals derived. Marko Hännikäinen, Antti Laitinen, Timo Hämäläinen 0001, Ilkka Kaisto, Kimmo Leskinen |
VTC Fall | 3 |
| 2000 | Embedding SDL implemented protocols into DSPabstractSpecification and Description Language (SDL) is an efficient tool for specifying and implementing complex real time systems, such as communications protocols.A Medium Access Control (MAC) protocol for a Wireless Local Area Network (WLAN) has been implemented using the language.The source C code for the protocol application is automatically generated from SDL.In this paper, the embedding of the generated protocol into a Digital Signal Processor (DSP) is described.As the high abstraction SDL model does not contain any dependencies to the target operational environment the application is adapted to DSP using tailored environment functions.These functions convert the asynchronous signalling used inside the SDL system into physical events of the DSP.In addition, as no operating system is used, the environment functions must implement other services required by the SDL generated application, such as platform clock and interrupt handling. Antti Takko, Marko Hännikäinen, Jarno Knuutila, Timo Hämäläinen 0001, Jukka Saarinen |
CASES | 4 |
| 2000 | Parallel DSP implementation of wavelet transform in image compressionabstractA parallel implementation of the 2D discrete wavelet transform on a distributed memory multiprocessor system called PARNEU is presented. The mapping has been chosen with consideration to load balancing and communication methods in order to achieve the best possible scalability and performance in transforming one single image. Detailed performance figures are included. Experimental results show that significant parallel speedup is reached with this mapping. Kaisa Haapala, Pasi Kolinummi, Timo Hämäläinen 0001, Jukka Saarinen |
ISCAS | 3 |
| 2000 | Scalable implementation of H.263 video encoder on a parallel DSP systemabstractA mapping and realization of H.263 video encoder for video conferencing applications are given using our DSP based multiprocessor system, called PAKNEU. PARNEU system is built for computationally intensive applications like image and video processing as well as soft computing applications. PARNEU has flexible communication architecture and thus it allows different mapping possibilities. The presented data parallel mapping has low communication and memory requirements, which allows encoding of any of the five standard H.263 picture formats. With a prototype system using four ADSP-21062 DSPs, a real-time encoding is achieved with QCIF sized picture. Phase times for each step of H.263 encoder are presented. Pasi Kolinummi, Juha Sarkijarvi, Timo Hämäläinen 0001, Jukka Saarinen |
ISCAS | 3 |
| 2000 | Real-time H.263 encoding of QCIF-images on TMS320C6201 fixed point DSPabstractIn this paper a real-time video encoder following ITU-T H.263 recommendation is described and relative computational loads of various encoding phases are evaluated. The current implementation shows that QCIF size (176/spl times/144) images can be coded in real-time by TMS320C6201 when computationally light motion estimation routines and basic H.263 coding mode is used. The real-time performance can be achieved using C-code and optimizing some key functions like SAD (sum of absolute difference) with linear assembler. In addition, careful memory management design for program code and application data is required. With the presented coding scheme, up to 29 fps can be achieved. Olli Lehtoranta, Timo Hämäläinen 0001, Jukka Saarinen |
ISCAS | 2 |
| 2000 | Chained backplane communication architecture for scalable multiprocessor systems
Pasi Kolinummi, Timo Hämäläinen 0001, Jukka Saarinen |
J. Syst. Archit. | 2 |
| 2000 | Parallel Implementation of Self-Organizing Map on the Partial Tree Shape Neurocomputer
Pasi Kolinummi, Pasi Pulkkinen, Timo Hämäläinen 0001, Jukka Saarinen |
Neural Process. Lett. | 3 |
| 1999 | Architecture for a Windows NT wireless LAN multimedia terminalabstractThe support for multimedia services in wireless access networks is challenging to implement, due to an unreliable and bandwidth-limited wireless medium. In addition, the widely used applications and protocols generally lack the support for quality-of-service (QoS) parameters. This paper presents the architecture of a multimedia wireless LAN terminal. The terminal is implemented using a custom network demonstrator platform connected to a Windows NT workstation. The system provides a separate management plane for configuring the service parameters of a proprietary wireless medium access control (MAC) protocol. Also, native applications that are capable of accessing the MAC QoS parameters directly are supported. Marko Hännikäinen, Timo Vanhatupa, Jussi Lemiläinen, Timo Hämäläinen 0001, Jukka Saarinen |
MMSP | 4 |
| 1999 | Software support for low latency communication in scalable multimedia DSP systemabstractThe software development and optimization of the communication is presented for an experimental scalable multiprocessor system purposed for multimedia applications. For convenient application mappings, presented C-code primitives hide the communication topology and allow the system to be expanded without needs to rewrite programs. For real-time applications, fast hardware communication need a strong support from software to prevent large latency times. Measured latency times and obtained throughput in communication are compared between user level application and maximum available hardware performance. Pasi Kolinummi, Pasi Pulkkinen, Timo Hämäläinen 0001, Jukka Saarinen |
MMSP | 3 |
| 1998 | TUTMAC: a medium access control protocol for a new multimedia wireless local area networkabstractThis paper presents a medium access control (MAC) protocol called TUTMAC for a new wireless local area network (TUTWLAN). The design objective has been to develop a simple, multimedia service capable protocol that provides sufficient medium utilisation efficiency and guarantees QoS (quality of service) parameters. The developed system utilises a centralised (base station controlled) network architecture. A limited number of portable stations can be associated with the same base station, i.e. in the same TUTWLAN cell. TUTMAC is connection oriented: the bandwidth is allocated deploying constant bit-rate TDMA based data channels that are reserved by exchanging short control messages. The connection parameters can be dynamically altered during the data exchange session. Currently, a TUTWLAN prototype is being developed comprising both TUTMAC software and platform hardware modules. The prototype will support up to eight simultaneous data-transfer connections each having 64 to 512 kbit/s data transmission bandwidth. Marko Hännikäinen, Jarno Knuutila, Ari Letonsaari, Timo Hämäläinen 0001, Jari Jokela, Juha Ala-Laurila, Jukka Saarinen |
PIMRC | 4 |
| 1998 | Security design for a new wireless local area network TUTWLANabstractThis paper presents a security scheme for a medium access control protocol in a new wireless local area network TUTWLAN (Tampere University of Technology WLAN). The design objective has been to develop a security scheme that will be scalable for various needs and offer high security for demanding applications. The designed security scheme provides both privacy of wireless data communications and the authenticity of communicating parties. Our authentication scheme allows also the communicating entities to establish a shared secret key for secure communication session. Data security schemes have also been introduced. There are three optional data security modes that offer flexible ciphering and data security level. Karri-Tuomas Salli, Timo Hämäläinen 0001, Jarno Knuutila, Jukka Saarinen |
PIMRC | 2 |
| 1997 | Mapping of Radial Basis Function Networks to Partial Tree Shape Parallel Neurocomputer
Pasi Kolinummi, Timo Hämäläinen 0001, Jukka Saarinen |
ICANN | 2 |
| 1997 | Parallel realizations of Kanerva's sparse distributed memory on a tree-shaped computerabstractThis paper presents two parallel realizations of sparse distributed memory (SDM) on a tree-shaped computer. The original model of SDM is introduced in terms of generalized computer memory and artificial neural networks (ANNs). For parallellization purposes, addressing, storage and retrieval operations are explained in detail. Some existing implementations in various computing platforms are considered before introducing the tree-shaped parallel computer, TUTNC (Tampere University of Technology Neural Computer). Two mappings are given, each utilizing parallelism with different granularities, and compared in terms of measured execution time, task partitioning and load balancing. Performance estimates are given for a larger system. The results show that SDM can be well parallelized in TUTNC. © 1997 John Wiley & Sons, Ltd. Timo Hämäläinen 0001, Harri Klapuri, Jukka Saarinen, Kimmo Kaski |
Concurr. Pract. Exp. | 1 |
| 1997 | Mapping of SOM and LVQ Algorithms on a Tree Shape Parallel Computer System
Timo Hämäläinen 0001, Harri Klapuri, Jukka Saarinen, Kimmo Kaski |
Parallel Comput. | 1 |
| 1996 | Linearly Expandable Partial Tree Shape Architecture for Parallel Neurocomputer
Timo Hämäläinen 0001, Pasi Kolinummi, Kimmo Kaski |
ICANN | 1 |
| 1996 | Mapping of Multilayer Perceptron Networks to Partial Tree Shaped Parallel Neurocomputer
Pasi Kolinummi, Timo Hämäläinen 0001, Harri Klapuri, Kimmo Kaski |
ICANN | 2 |
| 1996 | Accelerating genetic algorithm computation in tree shaped parallel computer
Timo Hämäläinen 0001, Harri Klapuri, Jukka Saarinen, Pekka Ojala, Kimmo Kaski |
J. Syst. Archit. | 1 |