Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Perttu Salmela

dblp:10/5181 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorComputer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Integrated circuit design · 77% Memory systems · 23%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Integrated circuit design
digital circuit design
0.112005
Systematic approach for path metric access in Viterbi decoders · IEEE Trans. Commun. 2005
Integrated circuit design › digital circuit design › VLSI architecture
viterbi decoder
0.112005
Systematic approach for path metric access in Viterbi decoders · IEEE Trans. Commun. 2005
Memory systems
memory access
0.012005
Systematic approach for path metric access in Viterbi decoders · IEEE Trans. Commun. 2005
Memory systems › memory access
parallel memory access
0.012005
Systematic approach for path metric access in Viterbi decoders · IEEE Trans. Commun. 2005

Methods — techniques the papers use, named apart from their topics

permutation-based memory access scheduling · 0.1
YearPublicationVenuePosition
2011 Fixed- versus floating-point implementation of MIMO-OFDM detector
abstract
In this paper, we investigate the opportunities offered by floating-point arithmetics in enabling an assembly and intrinsics free high-level language based development. We compare the characteristics of floating- and fixed-point arithmetics by simulating a MIMO-OFDM soft output detector in a 3G LTE link level simulator. The hardware complexity and energy dissipation are analyzed by implementing three programmable processors supporting 32- and 12-bit floating-point and 16-bit fixed-point arithmetics. The processors are based on the transport triggered architecture (TTA) that has a very low programmability overhead. The analysis shows that at the same goodput rate a floating-point implementation can achieve a lower gate count and better power efficiency than a fixed-point design.
Janne Janhunen, Perttu Salmela, Olli Silvén, Markku Juntti
ICASSP2
2008 Complex-valued QR decomposition implementation for MIMO receivers
abstract
Multiple input multiple output (MIMO) transmission is an emerging technique targeted at 3G long term evolution (LTE) systems. One vital baseband function in MIMO receivers is QR decomposition of the channel matrix. In this paper, a processor based complex-valued QR decomposition is presented. The processor is enhanced with complex arithmetic and inverse square root function units. The proposed processor fits well with the real-time requirements of the MIMO receiver. The computing power is tailored for typical MIMO systems. Due to the generality of the applied computing resources it can also be used for other tasks. Also, the presented principles can be applied on any customizable processor architectures to accelerate QR decomposition.
Perttu Salmela, Adrian Burian, Harri Sorokin, Jarmo Takala
ICASSP1
2006 Evaluation of stride permutation networks
abstract
By exploiting the inherent parallelism in digital signal processing algorithms, significant savings in area and power consumption may be achieved. Completely parallel computation can lead to excessive area, thus mapping the algorithm onto reduced computational resources becomes beneficial. As a drawback, data interconnections become more complex and require storage in order to maintain computationally correct processing. We have proposed a systematic design methodology for managing data interconnections called stride permutations. These stride permutations are found in several algorithms, including fast Fourier transforms and Viterbi decoding. The proposed methodology leads to regular and scalable permutation networks which support power-of-two strides. In addition, the networks reach the lower bound in the number of registers indicating area-efficiency. In this paper, the proposed networks are evaluated in terms of control, area, power consumption, and timing
Tuomas Järvinen, Perttu Salmela, Konsta Punkka, Jarmo Takala
ISCAS2
2005 Complex Fixed-Point Matrix Inversion Using Transport Triggered Architecture
abstract
Fixed-point simulations for inverting matrices using transport triggered architectures are performed. Several methods are implemented in fixed-point: the Cholesky decomposition as a direct method, Newton iterations as an iterative method, and Strassen Newton algorithm as a combined recursive method. Fixed-point implementations of these matrix inversion algorithms are tested and analyzed. A division-free implementation is targeted.
Adrian Burian, Perttu Salmela, Jarmo Takala
ASAP2
2005 256-State Rate 1/2 Viterbi Decoder on TTA Processor
abstract
Efficient and flexible Viterbi decoding is an important problem in implementation of modern telecommunications systems. In this paper, a 256-state, rate 1/2 Viterbi decoder is implemented on a transport triggered architecture processor. Due to the processor-based platform, the implementation is flexible and it can achieve relatively high decoding speed. The decoder is implemented by tailoring the processor to meet the requirements of Viterbi decoding. The processor is enhanced with a number of special function units, which accelerate the decoding. As a result, the decoding can be carried out efficiently on a processor-based platform.
Perttu Salmela, Tuomas Järvinen, Teemu Sipilä, Jarmo Takala
ASAP1
2005 Systematic approach for path metric access in Viterbi decoders
abstract
A systematic approach for the path metric memory management in Viterbi decoders is presented. Between the parallel computation units and memory modules, a permutation of path metrics is required in order to access the path metrics in correct order. We propose a parallel memory access scheme, which reduces the interconnection complexity between parallel computation units and memory modules by rescheduling the path metric computations.
Tuomas Järvinen, Perttu Salmela, Teemu Sipilä, Jarmo Takala
IEEE Trans. Commun.2
2004 Stride Permutation Networks for Array Processors
Tuomas Järvinen, Perttu Salmela, Harri Sorokin, Jarmo Takala
ASAP2
2003 On allocation of turbo decoder iterations
abstract
The computational requirements of turbo decoders are heavy and depend on the number of decoding iterations. However, only a few iterations are often enough for obtaining an error free outcome. Instead of idling during the unused iterations, the released time slots can be used for enabling extra iterations required by the decoding, when the errors are more severe. Since the number of required iterations for each code block is not known beforehand, the allocation of iterations must adapt according to the current demand and the past iterations. In this paper, the dynamic adaptation is achieved by an additional buffering of incoming code blocks and the effect of utilizing unused iterations are studied.
Perttu Salmela, Tuomas Järvinen, Teemu Sipilä, Jarmo Takala
PIMRC1
2003 Parallel memory access in turbo decoders
abstract
The memory requirements of turbo decoders are high since long code block lengths are preferred. Especially, the extrinsic information memory is accessed frequently with both linear and interleaved access patterns. In this paper, a parallel access scheme into extrinsic information memory is developed for a 3GPP turbo decoder. A single port memory is divided into parallel accessible modules and the memory throughput requirements and both the linear and interleaved access patterns are considered as module and word address generating functions are developed. As a result, the throughput of the parallel access scheme allows high-speed decoding and the usage of the dual port memory can be avoided and savings in chip area are achieved.
Perttu Salmela, Tuomas Järvinen, Teemu Sipilä, Jarmo Takala
PIMRC1
2003 In-Place Storage of Path Metrics in Viterbi Decoders
Tuomas Järvinen, Perttu Salmela, Teemu Sipilä, Jarmo Takala
VLSI-SOC2
2001 Multi-port interconnection networks for radix-R algorithms
abstract
In array processors, complex data reordering is often needed to realize the interconnection topologies between the computational nodes in algorithms. Several important algorithms, e.g., discrete trigonometric transforms and Viterbi decoding, can be represented in a radix-R form where the principal topology is stride by R permutation. A general factorialization of stride permutations is derived, which can be mapped onto register-based structures for constructing area-efficient multi-port interconnection networks. The networks can be modified to support several stride permutations and sequence sizes.
Jarmo Takala, Tuomas Järvinen, Perttu Salmela, David Akopian
ICASSP3