EDBT 2026 Demo / reviewers in the wild / expert
Jari Nurmi
dblp:36/6529
· DBLP profile ↗
75ranked-venue papers
6as first author
23since 2021 · last 2025
0000-0003-2169-4606ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 52 · 5 first-author · 12 since 2021Computer networks · 7 · 5 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | REAM: A Reinforcement Learning-based Energy-Efficient and Adaptive Multi-Modal Routing Protocol for Underwater Acoustic NetworksabstractEnhancing reliability in dense underwater networks with hefty data traffic often raises energy consumption due to constant packet listening and retransmissions caused by packet loss. To address the challenging energy demand in Underwater Acoustic Sensor Networks (UASNs), we propose a Reinforce-ment learning (RL)-based Energy-efficient and Adaptive Multi-modal routing protocol abbreviated as REAM, that integrates Q-learning with multi-modal communication to enhance energy efficiency in underwater networks by adaptively selecting the optimal mode for packet transmission. We compare the performance of two variants of the proposed REAM protocol-REAM-MM, which uses two modems, and REAM-SM, which uses a single modem, against two other state-of-the-art protocols namely QELAR and MARLIN-Q. Our results demonstrate the effective-ness of using multiple modems instead of a single modem in reducing energy consumption and improving reliability. REAM-MM reduces energy consumption per bit by up to 81.2%, 79.6%, and 72.9% under low, medium, and high traffic scenarios, respec-tively, compared to the best-performing alternative, MARLIN-Q. REAM-MM achieves a consistently comparable Packet Delivery Ratio (PDR) to MARLIN-Q and a higher PDR than QELAR and REAM-SM. Additionally, it maintains the lowest energy consumption under all traffic conditions and dense networks. Rabia Qadar, Waleed Bin Qaim, Bo Tan 0003, Jari Nurmi |
WCNC | 4 |
| 2025 | Novel Direction-of-Arrival-Based Localization in Massive DECT-2020 5G NR NetworksabstractThis research investigates an affordable, energy-efficient direction-of-arrival (DOA)-based localization solution for digital enhanced cordless telecommunications (DECTs) 2020 new radio (NR), a new standard lacking a native positioning feature. This standard enables massive Internet of Things (IoT) networks, a vast 5G network interconnecting an unparalleled number of low-cost and battery-operated smart sensors. However, integrating DOA localization into such networks is challenging due to cost constraints and power limitations. We propose a potentially cost-effective solution using a single radio-frequency (RF) chain for uniform L-shaped antenna arrays. Each antenna takes turns sampling the orthogonal frequency division multiplexing (OFDM) signal via an RF switch, enabled by time-dividing the OFDM signal into sample and switch slots. Further, we introduce a novel DOA method optimized for single Line-of-Sight (LOS) OFDM signals and array sequential sampling. This method leverages the dual shift-invariant properties of L-shaped antenna arrays and the array frequency response to estimate the azimuth and elevation angles. Experiments in an indoor environment reveal that at a signal-to-noise ratio (SNR) of 15 dB, over 50% of data achieve subdegree angular accuracy, increasing to 75% at 20 dB. Thus, over 50% of position estimations fall below the submeter error level at 15 dB SNR, rising to nearly 75% at 25 dB SNR. Our findings also indicate that halving the slot rate by proportionately reducing active subcarriers does not compromise accuracy. Experiments on the nRF52480 system-on-chip show the new DOA method is both fast and energy-efficient, taking only 0.76–2.26 ms and consuming 5.08–15.1 nWh. Tiago Troccoli, Hans Jakob Damsgaard, Juho Pirskanen, Elena Simona Lohan, Aleksandr Ometov, Jorge Morte Palacios, Jari Nurmi, Ville Kaseva |
IEEE Internet Things J. | 7 |
| 2025 | Parallel Accurate Minifloat MACCs for Neural Network Inference on Versal FPGAsabstractMachine learning (ML) is ubiquitous in contemporary applications. Its need for efficient acceleration has driven vast research efforts into the quantization of neural networks with low-precision numerical formats. Models quantized with minifloat formats of eight or fewer bits have proven capable of outperforming models quantized into same-size integers. However, unlike integers, minifloats require accurate accumulation to prevent the introduction of rounding errors. We explore the design space of parallel accurate minifloat multiply-accumulators (MACCs) targeting the AMD VersalTM FPGA fabric. We experiment with three variations of the multiply-and-shift and adder tree components of a minifloat MACC. For comparison, we apply similar alterations to a parallel integer MACC. Our results show that custom compressor trees with external sign-inversion gates reduce the mean area of the minifloat MACCs by 17.7% and increase their clock frequency by 16.2%. In comparison, custom compressor trees with absorbed partial product generation gates reduce the mean area of integer MACCs by 28.1% and increase their clock frequency by 3.60%. Comparing the best-performing designs, we observe that minifloat MACCs consume 20% to 180% more resources than integer ones with same-size operands without accounting for a conversion back into a floating-point format, and 60% to 300% more resources when including it. Our data enable engineers to make informed decisions in their designs of deeply integrated embedded ML solutions when trading off training and fine-tuning effort versus resource cost. Hans Jakob Damsgaard, Konstantin Hoßfeld, Jari Nurmi, Thomas B. Preußer |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Guest Editorial: Selected Papers From IEEE Nordic Circuits and Systems Conference (NorCAS) 2024abstractNon peer reviewed Jari Nurmi, Snorre Aunet, Alireza Saberkari |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | Adaptive approximate computing in edge AI and IoT applications: A reviewabstractRecent advancements in hardware and software systems have been driven by the deployment of emerging smart health and mobility applications. These developments have modernized the traditional approaches by replacing conventional computing systems with cyber-physical and intelligent systems combining the Internet of Things (IoT) with Edge Artificial Intelligence. Despite the many advantages and opportunities of these systems within various application domains, the scarcity of energy, extensive computing needs, and limited communication must be considered when orchestrating their deployment. Inducing savings in these directions is central to the Approximate Computing (AxC) paradigm, in which the accuracy of some operations is traded off with energy, latency, and/or communication reductions. Unfortunately, the dynamics of the environments in which AxC-equipped IoT systems operate have been paid little attention. We bridge this gap by surveying adaptive AxC techniques applied to three emerging application domains, namely autonomous driving, smart sensing and wearables, and positioning, paying special attention to hardware acceleration. We discuss the challenges of such applications, how adaptive AxC can aid their deployment, and which savings it can bring based on traits of the data and devices involved. Insights arising thereof may serve as inspiration to researchers, engineers, and students active within the considered domains. Hans Jakob Damsgaard, Antoine Grenier, Dewant Katare, Zain Taufique, Salar Shakibhamedan, Tiago Troccoli, Georgios Chatzitsompanis, Anil Kanduri, Aleksandr Ometov, Aaron Yi Ding, Nima Taherinejad, Georgios Karakonstantis, Roger F. Woods, Jari Nurmi |
J. Syst. Archit. | 14 |
| 2024 | Coarse-grained reconfigurable architectures for radio baseband processing: A surveyabstractEmerging communication technologies, such as 5G and beyond, have introduced diverse requirements that demand high performance and energy efficiency at all levels. Furthermore, the real-time requirements of different services vary significantly — increasing the baseband processor design complexity and demand for flexible hardware platforms. This paper identifies the key characteristics of hardware platforms for baseband processing and describes the existing processing limitations in traditional architectures. In this paper, Coarse-Grained Reconfigurable Architecture (CGRA) is examined as a prospective hardware platform and its characteristic features are highlighted as compared to traditionally employed architectures that make it a suitable candidate for incorporation as a domain-specific accelerator in baseband processing applications. We survey various CGRAs from the last two decades (2004-2023) and analyze their distinct architectural features which can serve as a reference while designing CGRAs for baseband processing applications. Moreover, we investigate the existing challenges toward developing CGRAs for baseband processing and explore their potential solutions. We also provide an overview of the emerging research directions for CGRA and how they can contribute toward the development of advanced baseband processors. Lastly, we highlight a conceptual RISC-V+CGRA framework that can serve as a potential direction toward integrating CGRA in future baseband processing systems. Aleksandr Ometov, Elena Simona Lohan, Jari Nurmi |
J. Syst. Archit. | 4 |
| 2024 | EWOk: Towards Efficient Multidimensional Compression of Indoor Positioning DatasetsabstractIndoor positioning performed directly at the end-user device ensures reliability in case the network connection fails but is limited by the size of the RSS radio map necessary to match the measured array to the device’s location. Reducing the size of the RSS database enables faster processing, and saves storage space and radio resources necessary for the database transfer, thus cutting implementation and operation costs, and increasing the quality of service. In this work, we propose EWOk, an Element-Wise cOmpression using k-means, which reduces the size of the individual radio measurements within the fingerprinting radio map while sustaining or boosting the dataset’s positioning capabilities. We show that the 7-bit representation of measurements is sufficient in positioning scenarios, and reducing the data size further using EWOk results in higher compression and faster data transfer and processing. To eliminate the inherent uncertainty of k-means we propose a data-dependent, non-random initiation scheme to ensure stability and limit variance. We further combine EWOk with principal component analysis to show its applicability in combination with other methods, and to demonstrate the efficiency of the resulting multidimensional compression. We evaluate EWOk on 25 RSS fingerprinting datasets and show that it positively impacts compression efficiency, and positioning performance. Lucie Klus, Roman Klus, Joaquín Torres-Sospedra, Elena Simona Lohan, Carlos Granell, Jari Nurmi |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | High-efficiency Compressor Trees for Latest AMD FPGAsabstractHigh-fan-in dot product computations are ubiquitous in highly relevant application domains, such as signal processing and machine learning. Particularly, the diverse set of data formats used in machine learning poses a challenge for flexible efficient design solutions. Ideally, a dot product summation is composed from a carry-free compressor tree followed by a terminal carry-propagate addition. On FPGA, these compressor trees are constructed from generalized parallel counters whose architecture is closely tied to the underlying reconfigurable fabric. This work reviews known counter designs and proposes new ones in the context of the new AMD Versal™ fabric. On this basis, we develop a compressor generator featuring variable-sized counters, novel counter composition heuristics, explicit clustering strategies, and case-specific optimizations like logic gate absorption. In comparison to the Vivado™ default implementation, the combination of such a compressor with a novel, highly efficient quaternary adder reduces the LUT footprint across different bit matrix input shapes by 45% for a plain summation and by 46% for a terminal accumulation at a slight cost in critical path delay still allowing an operation well above 500 MHz. We demonstrate the aptness of our solution at examples of low-precision integer dot product accumulation units. Konstantin Hoßfeld, Hans Jakob Damsgaard, Jari Nurmi, Michaela Blott, Thomas B. Preußer |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2024 | Guest Editorial Selected Papers From IEEE Nordic Circuits and Systems Conference (NorCAS) 2022abstractThis special section is composed of substantially extended handpicked papers from the IEEE Nordic Circuits and Systems Conference (NorCAS) 2022 that took place in October 2022 in Oslo, Norway. Jari Nurmi, Snorre Aunet, Alireza Saberkari |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | Guest Editorial Selected Papers From IEEE Nordic Circuits and Systems Conference (NorCAS) 2023
Jari Nurmi, Snorre Aunet, Alireza Saberkari |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | SPHERE-DNA: Privacy-Preserving Federated Learning for eHealthabstractThe rapid growth of chronic diseases and medical conditions (e.g. obesity, depression, diabetes, respiratory and musculoskeletal diseases) in many OECD countries has become one of the most significant wellbeing problems, which also poses pressure to the sustainability of healthcare and economies. Thus, it is important to promote early diagnosis, intervention, and healthier lifestyles. One partial solution to the problem is extending long-term health monitoring from hospitals to natural living environments. It has been shown in laboratory settings and practical trials that sensor data, such as camera images, radio samples, acoustics signals, infrared etc., can be used for accurately modelling activity patterns that are related to different medical conditions. However, due to the rising concern related to private data leaks and, consequently, stricter personal data regulations, the growth of pervasive residential sensing for healthcare applications has been slow. To mitigate public concern and meet the regulatory requirements, our national multi-partner SPHERE-DNA project aims to combine pervasive sensing tech-nology with secured and privacy-preserving distributed privacy frameworks for healthcare applications. The project leverages local differential privacy federated learning (LDP-FL) to achieve resilience against active and passive attacks, as well as edge computing to avoid transmitting sensitive data over networks. Combinations of sensor data modalities and security architectures are explored by a machine learning architecture for finding the most viable technology combinations, relying on metrics that allow balancing between computational cost and accuracy for a desired level of privacy. We also consider realistic edge computing platforms and develop hardware acceleration and approximate computing techniques to facilitate the adoption of LDP-FL and privacy preserving signal processing to lightweight edge processors. A proof-of-concept (PoC) multimodal sensing system will be developed and a novel multimodal dataset will be collected during the project to verify the concept. Jari Nurmi, Yinda Xu, Jani Boutellier, Bo Tan 0003 |
DATE | 1 |
| 2023 | Generating CGRA Processing Element Hardware with CGRAgenabstractThe popularity of the Internet of Things and next-generation wireless networks calls for a greater distribution of small but high-performance and energy-efficient compute devices at the networks' Edge. These devices must integrate hardware acceleration to meet the latency requirements of relevant use cases. Existing work has highlighted Coarse-Grained Reconfigurable Arrays (CGRAs) as suitable compute architectures for this purpose. However, like other modern hardware design, research and design space exploration into CGRAs is hindered by long development times needed for Register Transfer Level implementation. In this paper, we propose mitigating these by extending the open-source CGRAgen tool with a Chisel-based hardware backend capable of transforming abstract Processing Element (PE) descriptions into synthesizable Verilog code. We present how CGRAgen's internal module representation is transformed to Chisel modules and demonstrate this on a selection of PE architectures from the literature. Finally, we outline future work on extending this flow to generate entire CGRAs. Hans Jakob Damsgaard, Aleksandr Ometov, Jari Nurmi |
DSD | 3 |
| 2023 | Towards Coarse-Grained Reconfigurable Approximate Computing with CGRAgenabstractModern Edge Computing devices execute applications that must meet strict latency requirements as per traditional standardization activities. Achieving the needed performance implies a need for efficiency in all aspects, thus, flexible solutions are needed. In this Ph.D. project, we address this issue for error-tolerant applications by using Coarse-Grained Reconfigurable Arrays (CGRAs) enriched with Approximate Computing (AxC) features. To do so, we aim to develop a CGRA architecture modeling, mapping, and hardware generation flow complete with AxC hardware primitives and significance analysis. Hans Jakob Damsgaard, Aleksandr Ometov, Jari Nurmi |
FPL | 3 |
| 2023 | Approximate computing in B5G and 6G wireless systems: A survey and future outlookabstractAs modern 5G systems are being deployed, researchers question whether they are sufficient for the oncoming decades of technological evolution.Growing numbers of interconnected intelligent devices put these networks under tremendous pressure, demanding their development.Paving the way for beyond 5G and 6G systems, commonly denoted by B5G herein, therefore means seeking enablers to increase efficiency from different perspectives.One novel look on this is the application of inexact computations where nine 9s reliability is not needed, for example, in non-critical mobile broadband traffic.The paradigm of Approximate Computing (AxC) focuses on such areas where constrained quality degradation results in savings that benefit the users and operators.This paper surveys the state-of-the-art publications on the intersection of AxC and B5G systems, identifying and emphasizing trends and tendencies in existing work and directions for future research.The work highlights resource allocation algorithms as particularly mesmerizing in the former, while research related to Intelligent Reflective Surfaces appears the most prominent in the latter.In both, problems are often NP-hard and, thus, only solvable using heuristics or approximations, Successive Convex Approximation and Reinforcement Learning are most frequently applied. Hans Jakob Damsgaard, Aleksandr Ometov, Md. Munjure Mowla, Adam Flizikowski, Jari Nurmi |
Comput. Networks | 5 |
| 2023 | Scalable and Efficient Clustering for Fingerprint-Based PositioningabstractIndoor positioning based on IEEE 802.11 wireless LAN (Wi-Fi) fingerprinting needs a reference data set, also known as a radio map, in order to match the incoming fingerprint in the operational phase with the most similar fingerprint in the data set and then estimate the device position indoors. Scalability problems may arise when the radio map is large, e.g., providing positioning in large geographical areas or involving crowdsourced data collection. Some researchers divide the radio map into smaller independent clusters, such that the search area is reduced to less dense groups than the initial database with similar features. Thus, the computational load in the operational stage is reduced both at the user devices and on servers. Nevertheless, the clustering models are machine-learning algorithms without specific domain knowledge on indoor positioning or signal propagation. This work proposes several clustering variants to optimize the coarse and fine-grained search and evaluates them over different clustering models and data sets. Moreover, we provide guidelines to obtain efficient and accurate positioning depending on the data set features. Finally, we show that the proposed new clustering variants reduce the execution time by half and the positioning error by$\approx 7$% with respect to fingerprinting with the traditional clustering models. Joaquín Torres-Sospedra, Darwin Quezada-Gaibor, Jari Nurmi, Yevgeni Koucheryavy, Elena Simona Lohan, Joaquín Huerta |
IEEE Internet Things J. | 3 |
| 2022 | Towards Approximate Computing for Achieving Energy vs. Accuracy Trade-offsabstractDespite the recent advances in semiconductor technology and energy-aware system design, the overall energy consumption of computing and communication systems is rapidly growing. On the one hand, the pervasiveness of these technologies everywhere in the form of mobile devices, cyber-physical embedded systems, sensor networks, wearables, social media and context-awareness, intelligent machines, broadband cellular networks, Cloud computing, and Internet of Things (IoT) has drastically increased the demand for computing and communications. On the other hand, the user expectations on features and battery life of online devices are increasing all the time, and it creates another incentive for finding good trade-offs between performance and energy consumption. One of the opportunities to address this growing demand is to utilize an Approximate Computing approach through software and hardware design. The APROPOS project aims at finding the balance between accuracy and energy consumption, and this short paper provides an initial overview of the corresponding roadmap, as the project is still in the initial stage. Aleksandr Ometov, Jari Nurmi |
DATE | 2 |
| 2022 | SURIMI: Supervised Radio Map Augmentation with Deep Learning and a Generative Adversarial Network for Fingerprint-based Indoor PositioningabstractIndoor Positioning based on Machine Learning has drawn increasing attention both in the academy and the industry as meaningful information from the reference data can be extracted. Many researchers are using supervised, semi-supervised, and unsupervised Machine Learning models to reduce the positioning error and offer reliable solutions to the end-users. In this article, we propose a new architecture by combining Convolutional Neural Network (CNN), Long short-term memory (LSTM) and Generative Adversarial Network (GAN) in order to increase the training data and thus improve the position accuracy. The proposed combination of supervised and unsupervised models was tested in 17 public datasets, providing an extensive analysis of its performance. As a result, the positioning error has been reduced in more than 70% of them. Darwin Quezada-Gaibor, Joaquín Torres-Sospedra, Jari Nurmi, Yevgeni Koucheryavy, Joaquín Huerta |
IPIN | 3 |
| 2022 | Data Cleansing for Indoor Positioning Wi-Fi Fingerprinting DatasetsabstractWearable and IoT devices requiring positioning and localisation services grow in number exponentially every year. This rapid growth also produces millions of data entries that need to be pre-processed prior to being used in any indoor positioning system to ensure the data quality and provide a high Quality of Service (QoS) to the end-user. In this paper, we offer a novel and straightforward data cleansing algorithm for WLAN fingerprinting radio maps. This algorithm is based on the correlation among fingerprints using the Received Signal Strength (RSS) values and the Access Points (APs)'s identifier. We use those to compute the correlation among all samples in the dataset and remove fingerprints with low level of correlation from the dataset. We evaluated the proposed method on 14 independent publicly-available datasets. As a result, an average of 14% of fingerprints were removed from the datasets. The 2D positioning error was reduced by 2.7% and 3D positioning error by 5.3% with a slight increase in the floor hit rate by 1.2% on average. Consequently, the average speed of position prediction was also increased by 14%. Darwin Quezada-Gaibor, Lucie Klus, Joaquín Torres-Sospedra, Elena Simona Lohan, Jari Nurmi, Carlos Granell, Joaquín Huerta |
MDM | 5 |
| 2022 | Underwater Optical Communication Module: An Extension to the ns-3 Network SimulatorabstractIn the last decade, the field of wireless optical communication has gathered immense interest due to its adoption in growing bandwidth-hungry underwater applications. The expensive and non-standardized on-field research measurements call for a reliable simulation tool that allows researchers to realistically design and assess the performance of Underwater Optical Communication (UOC) systems before conducting actual underwater experiments. In this paper, we present a UOC module as an extension to the network simulator ns-3. The module can study the impact of different water conditions on underwater optical networks from the physical layer to the network layer. The proposed UOC module realizes physical layer models of the UOC channels where the added noise and interference effects are modeled as Additive White Gaussian Noise (AWGN). Results show the capability of our module to facilitate large underwater optical network design and optimization. Since ns-3 is an open-source software, the module has the flexibility and reusability to be further developed by the worldwide research community. Rabia Qadar, Waleed Bin Qaim, Bo Tan 0003, Jari Nurmi |
VTC Fall | 4 |
| 2021 | When wearable technology meets computing in future networks: a road aheadabstractRapid technology advancement, economic growth, and industrialization have paved the way for developing a new niche of small body-worn personal devices, gathered together under a wearable-technology title. The triggers stimulated by end-users interest have introduced the first generation of mass-consumer wearables in just the past decade. Evidently, the trailblazing ones were not designed with strict energy-consumption restrictions in mind. Thus, wearable-computing-related research remained fragmented. Advanced and sophisticated batteries and communication technologies could be already procurable on devices. Additional solutions for efficient utilization of processing power are still a white spot on the wearable technology roadmap. A-WEAR EU project aims to enhance the understanding of how the superimposition of those technologies would improve wearable devices' energy efficiency, with the research area being far from saturation. We foresee enormous room for research as the Edge computing paradigm is emerging towards hand-held devices. Aleksandr Ometov, Olga Chukhno, Nadezhda Chukhno, Jari Nurmi, Elena Simona Lohan |
CF | 4 |
| 2021 | Run-to-Completion versus Pipelined: The Case of 100 Gbps Packet ParsingabstractPacket parsing is the initial step in processing of network packets. It is encountered in any environment in which packets must be processed. Examples include switches, routers, firewalls, and kernel of operating system. In recent years, there has been focus on programmable and protocol-independent packet processing hardware. The two main hardware architectures for packet processing are run-to-completion and pipelined organization of functional units. This applies to packet parsing as well. Both run-to-completion and pipelined organization have pros and cons and the debate as to which provides greater overall benefit is endless. In this paper, we consider this problem from the perspective of programmable 100 Gbps packet parsing. We will see that the pipelined parser provides 40x throughput compared to the run-to-completion architecture despite running at the same operating frequency and using the same functional units in each pipeline stage. Hesam Zolfaghari, Haseeb Mustafa, Jari Nurmi |
HPSR | 3 |
| 2021 | Lightweight Wi-Fi Fingerprinting with a Novel RSS Clustering AlgorithmabstractNowadays, several indoor positioning solutions sup-port Wi-Fi and use this technology to estimate the user position. It is characterized by its low cost, availability in indoor and outdoor environments, and a wide variety of devices support Wi-Fi technology. However, this technique suffers from scalability problems when the radio map has a large number of reference fingerprints because this might increase the time response in the operational phase. In order to minimize the time response, many solutions have been proposed along the time. The most common solution is to divide the data set into clusters. Thus, the incoming fingerprint will be compared with a specific number of samples grouped by, for instance similarity (clusters). Many of the current studies have proposed a variety of solutions based on the modification of traditional clustering algorithms in order to provide a better distribution of samples and reduce the computational load. This work proposes a new clustering method based on the maximum Received Signal Strength (RSS) values to join similar fingerprints. As a result, the proposed fingerprinting clustering method outperforms three of the most well-known clustering algorithms in terms of processing time at the operational phase of fingerprinting. Darwin Quezada-Gaibor, Joaquín Torres-Sospedra, Jari Nurmi, Yevgeni Koucheryavy, Joaquín Huerta |
IPIN | 3 |
| 2021 | Towards Ubiquitous Indoor Positioning: Comparing Systems across Heterogeneous DatasetsabstractThe evaluation of Indoor Positioning Systems (IPSs) mostly relies on local deployments in the researchers' or partners' facilities. The complexity of preparing comprehensive experiments, collecting data, and considering multiple scenarios usually limits the evaluation area and, therefore, the assessment of the proposed systems. The requirements and features of controlled experiments cannot be generalized since the use of the same sensors or anchors density cannot be guaranteed. The dawn of datasets is pushing IPS evaluation to a similar level as machine-learning models, where new proposals are evaluated over many heterogeneous datasets. This paper proposes a way to evaluate IPSs in multiple scenarios, that is validated with three use cases. The results prove that the proposed aggregation of the evaluation metric values is a useful tool for high-level comparison of IPSs. Joaquín Torres-Sospedra, Ivo Silva, Lucie Klus, Darwin Quezada-Gaibor, Antonino Crivello, Paolo Barsocchi, Cristiano G. Pendão, Elena Simona Lohan, Jari Nurmi, Adriano J. C. Moreira |
IPIN | 9 |
| 2019 | Reducing Crossbar Costs in the Match-Action PipelineabstractSoftware Defined Networking (SDN) is a new networking paradigm in which the control plane and data plane are decoupled. Throughout the recent years, a number of architectures have emerged for protocol-independent packet processing. One such architecture is the Protocol Independent Switch Architecture (PISA). It is a programmable and protocol-independent architecture composed of a number of Match and Action stages. Inside each of these stages is a crossbar to generate the search key and another crossbar to provide the input to the Action Units. In this paper, we design and explore alternative interconnection schemes with the aim of finding the most area- and power-efficient interconnection structure. Moreover, we propose further modifications to the interconnection structure, as a result of which the on-chip area of both match and action crossbars will be reduced by more than 70 % and power dissipation will be reduced by 25.8 % and 23.1 % for match and action crossbars respectively. Hesam Zolfaghari, Davide Rossi 0001, Jari Nurmi |
HPSR | 3 |
| 2018 | An Explicitly Parallel Architecture for Packet Parsing in Software Defined NetworksabstractPacket parsing is the first step in processing of packets in devices such as network switches and routers. The process of packet parsing has become more challenging due to the increase in line rates and emergence of Software Defined Networking which leads to new protocols being adopted. In this paper, we present a novel architecture for parsing of packets. The architecture is fully programmable and is not tied to any specific protocol. It can be programmed to parse any protocol making it suitable for Software Defined Networks. Compared with the parser used in the Reconfigurable Match Tables, our parser improves supported throughput by a factor of 3.2. Moreover, to achieve the target throughput of 640 Gbps, our parser needs only 2 percent of the number of gates used in the parsers of Reconfigurable Match Tables. Hesam Zolfaghari, Davide Rossi 0001, Jari Nurmi |
ASAP | 3 |
| 2018 | Delay-Accuracy Tradeoff in Opportunistic Time-of-Arrival LocalizationabstractWhile designing a positioning network, the localization performance is traditionally the main concern. However, collection of measurements together with channel access methods require a nonzero time, causing a delay experienced by network nodes. This fact is usually neglected in the positioning-related literature. In terms of the delay-accuracy tradeoff, broadcast schemes have an advantage over unicast, provided nodes can be properly synchronized. In this letter, we analyze the delay-accuracy tradeoff for localization schemes in which the position estimates are obtained based on broadcasted ranging signals. We find that for dense networks, the tradeoff is the same for cooperative and noncooperative networks, and cannot exceed a certain threshold value. Ondrej Daniel, Henk Wymeersch, Jari Nurmi |
IEEE Signal Process. Lett. | 3 |
| 2018 | Errata to "Evaluation of a Heterogeneous Multicore Architecture by Design and Test of an OFDM Receiver"abstractPresents corrections to the paper, “Evaluation of a heterogeneous multicore architecture by design and test of an OFDM receiver,” (Nouri, S. et al,), IEEE Trans. Parallel Dist. Syst., vol. 28, no. 11, pp. 3171–3187, Nov. 2017. Sajjad Nouri, Waqar Hussain 0001, Jari Nurmi |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | A comparison of Bayesian localization methods in the presence of outliersabstractLocalization of a user in a wireless network is challenging in the presence of malfunctioning or malicious reference nodes, since if they are not accounted for, large localization errors can ensue. We evaluate three Bayesian methods to statistically identify outliers during localization: an exact method, an expectation maximization (EM) method proposed earlier, and a new method based on Variational Bayesian EM (VBEM). Simulation results indicate similar performance for the latter two schemes, with the VBEM algorithm able to provide a statistical description of the user location, rather than an estimate as in the simpler EM case. In contrast to previous studies, we find that there is a significant gap between the approximate methods and the exact method, the cause of which is discussed. Giorgia Nunzia Ferrara, Henk Wymeersch, Elena Simona Lohan, Jari Nurmi |
IWCMC | 4 |
| 2017 | FPGA Implementation Issues of a Flexible Synchronizer Suitable for NC-OFDM-Based Cognitive Radios
Farid Shamani, Roberto Airoldi, Vida Fakour Sevom, Tapani Ahonen, Jari Nurmi |
J. Syst. Archit. | 5 |
| 2017 | Evaluation of a Heterogeneous Multicore Architecture by Design and Test of an OFDM ReceiverabstractThis paper presents an evaluation of a Heterogeneous Multicore Architecture (HMA) by implementing Orthogonal Frequency-Division Multiplexing (OFDM) receiver blocks as designs for the test of functionality. OFDM receiver consists of computationally intensive and general-purpose processing tasks that can provide maximum coverage to test and evaluate a massively-parallel as well as a general-purpose platform like the HMA. The blocks of the receiver are primarily designed by crafting template-based Coarse-Grained Reconfigurable Array (CGRA) devices and then arranging them in a sequence over a Network-on-Chip (NoC) structure along with a few RISC cores for complete OFDM processing. The OFDM blocks such as Fast Fourier Transform (FFT) and Time Synchronization are computationally intensive and require parallel processing. The OFDM receiver also contains tasks such as frequency offset estimation which require the processing of Taylor series and CORDIC algorithms that are serial in nature. Such a combination of serial and parallel algorithms can perform a thorough exploration and evaluation of almost all the design features of an HMA. The OFDM implementation has led to scale CGRAs to different dimensions, instantiate Processing Elements (PEs) as multiple arithmetic resources and to establish almost all possible ways of PE interconnections. It further explores time-multiplexed patterns for data placement in the CGRA memories. Nevertheless, the data can also be exchanged among different nodes over NoC structure simultaneously and independently by using direct memory access devices. In this experimental work, the performance of each CGRA, the collective performance of the whole platform and the NoC traffic are recorded in terms of the number of clock cycles and several high-level performance metrics. Today's HMAs are generally over or under resourced for the applications that they are designed for and thus not an optimal choice for the end user. Apart from the interesting comparisons to the other state-of-the-art, our experimental setup has provided important insight and guidelines that the designers can use to implement near-optimal solutions for their target applications. Sajjad Nouri, Waqar Hussain 0001, Jari Nurmi |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | MULTI-POS: Marie Curie network in multi-technology positioning
Jari Nurmi, Elena Simona Lohan |
DATE | 1 |
| 2016 | Blind sub-Nyquist GNSS signal detectionabstractA satellite navigation receiver traditionally searches for positioning signals using an acquisition procedure. In situations, in which the required information is only a binary decision whether at least one positioning signal is present or absent, the procedure represents an unnecessarily complex solution. This paper presents a different approach for the binary detection problem with significantly reduced computational complexity. The approach is based on a novel decision metric which is utilized to design two binary detectors. The first detector operates under the theoretical assumption of additive white Gaussian noise and is evaluated by means of Receiver Operating Characteristics. The second one considers also additional interferences and is suitable to operate in a real environment. Its performance is verified using a signal captured by a receiver front-end. Ondrej Daniel, Jussi Raasakka, Pekka Peltola, Markus Fröhle, Alejandro Rivero Rodriguez, Henk Wymeersch, Jari Nurmi |
ICASSP | 7 |
| 2014 | Design of an accelerator-rich architecture by integrating multiple heterogeneous coarse grain reconfigurable arrays over a network-on-chipabstractThis paper presents an accelerator-rich system-on-chip (SoC) architecture integrating many heterogeneous Coarse Grain Reconfigurable Arrays (CGRA) connected through a Network-on-Chip (NoC). The architecture is designed to maximize the reconfigurable processing capacity for the execution of massively parallel algorithms. The central node of the NoC contains a Reduced Instruction Set Computer (RISC) core that manages distribution of computing functions and data within the SoC while the other nodes contain CGRAs of application-specific sizes. Prior approaches coupled only a few accelerators with a RISC core using special instructions and/or a direct memory access device. In contrast, our design couples a RISC core to many CGRAs through the NoC. This approach provides for independent and simultaneous execution of multiple computing kernels. Furthermore, the proposed architecture mitigates power dissipation as CGRA sizes are tailored for the individual application kernels. We present a proof-of-concept design with a total of 408 reconfigurable processing elements. This instance and its sub-systems are customized and tested for different computationally-intensive signal processing algorithms. The overall single-chip computing system is synthesized for a Field Programmable Gate Array device. We present comparison to and evaluation against some of the existing multicore systems in terms of multiple performance metrics. Waqar Hussain 0001, Roberto Airoldi, Henry Hoffmann, Tapani Ahonen, Jari Nurmi |
ASAP | 5 |
| 2014 | High-level parameterizable area estimation modeling for ASIC designs
Ville Eerola, Jari Nurmi |
Integr. | 2 |
| 2014 | MPSoC based on Transport Triggered Architecture for baseband processing of an LTE receiver
Omer Anjum, Tapani Ahonen, Jari Nurmi |
J. Syst. Archit. | 3 |
| 2014 | Transport triggered architecture to perform carrier synchronization for LTEabstractIn this article implementation of carrier frequency offset estimate for 20MHz LTE baseband processing is discussed. LTE (Long Term Evolution) is a wireless communication standard that makes use of some innovative techniques to gain very high data rates (>100Mbps). This goal for such a high throughput also imposes design challenges for the industry and academia such as in the case of handheld mobile devices where the power budget is very limited. Implicitly high throughput means we need more computation power and more energy. On the other hand industry is also struggling for a flexible hardware solution, or software defined a radio (SDR), to amortize the huge cost of required hardware changes as the wireless standards have kept evolving. Design innovations are now needed to confront those challenges of low power and flexible design without changing the hardware. The implementation is made on Transport Triggered Architecture (TTA), which is a unique concept in computer architecture design, based on the single instruction, “MOVE”. The power consumption of the architecture when synthesized on 180nm technology at 180MHz and 1.8V is 18.39mW. The total area occupied excluding memory is 0.6mm 2 . The proposed TTA solution has been compared with, a more ASIC (application specific integrated circuits), like ASIP (application specific instruction processor) solution and a coprocessor accelerator-based solution. The proposed solution is more flexible: easily programmable due to high level language support, easily scalable, and still efficient in energy consumption needed to complete the CFO (carrier frequency offset) estimation task. Because of these attractive characteristics, TTA is also a potential candidate for SDR platforms. Omer Anjum, Mubashir Ali, Teemu Pitkänen, Jari Nurmi |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2013 | A Reconfigurable Application-specific Instruction-set Processor for Fast Fourier Transform processingabstractIn this paper, we have presented a Reconfigurable Application-specific Instruction-set Processor (rASIP) that processes mixed-radix(2, 4) 64 and 128-point Fast Fourier Transform (FFT) algorithms while satisfying the partial execution-time requirements of IEEE-802.11n standard. The rASIP was designed by integrating a template-based Coarse-Grain Reconfigurable Array (CGRA) in the datapath of a simple Reduced Instruction-Set Computing (RISC) Processor. The instruction set of the RISC processor was extended to add special instructions to enable cycle-accurate processing by the CGRA. The rASIP is synthesized for Field Programmable Gate Arrays for the measurement of resource utilization and execution time. The postfit gate-level netlist of rASIP was simulated to estimate the power and energy consumption. Based on our measurements and estimates, we have studied the advantages of using rASIP in comparison with other systems. Waqar Hussain 0001, Gerd Ascheid, Jari Nurmi |
ASAP | 4 |
| 2013 | Correctly rounded architectures for Floating-Point multi-operand addition and dot-product computationabstractThis study presents hardware architectures performing correctly rounded Floating-Point (FP) multioperand addition and dot-product computation, both of which are widely used in various fields, such as scientific computing, digital signal processing, and 3D graphic applications. A novel realignment method is proposed to solve the catastrophic cancellation and multi-sticky bits. Only one rounding operation is performed in both of the proposed FP multi-operand adder and dot-product computation unit. Implementation results show that our architectures not only can produce correctly rounded results, whose errors are less than 0.5 ULP (Unit in the Last Place), but also have reduced delay comparing with the traditional network architecture, which uses 2-operand FP adders and multipliers to perform multi-operand addition and dot-product computation. Deyuan Gao, Xiaoya Fan, Jari Nurmi |
ASAP | 4 |
| 2013 | Exploiting RSS measurements among neighbouring devices: A matter of trustabstractIn this paper we present experimental evaluations of cooperative ranging-based approaches in mobile positioning. Our main contribution is represented by experimental investigations, analysis of the errors introduced in the distance estimation and exploitation of the knowledge of environmental perturbations with a NLLS algorithm, highlighting the limits of such cooperative schemes, proving that the effect of cooperation might be very limited in real-life scenarios. Francescantonio Della Rosa, Tommi Paakki, Jari Nurmi, Mauro Pelosi |
IPIN | 3 |
| 2012 | Reconfigurable multi-processor architecture for streaming applicationsabstractThis paper presents a work-in-progress design of a reconfigurable multi-processor architecture. The architecture is composed of nine nodes arranged in a 3x3mesh topology. The central node of the architecture hosts a RISC processor, which acts as master of the platform, taking care of data and task scheduling. The surrounding nodes hosts a reconfigurable engine and the actual processing. The system was prototyped on an Altera FPGA device and RTL simulations of the architecture were carried out to ensure the correct functionality of the systems. Future works will focus on the implementation of significant kernels of streaming applications on the platform as well as the implementation of power saving techniques in order to achieve a high power efficiency of the system. Leyla S. Ghazanfari, Roberto Airoldi, Jari Nurmi, Tapani Ahonen |
FPL | 3 |
| 2011 | System Level Performance Simulation of Distributed GENESYS Applications on Multi-core PlatformsabstractModern high end mobile devices employ multi-core platforms and support diverse distributed applications due to increased computational power. A brisk performance evaluation phase is required after the application modelling to evaluate feasibility of new distributed applications on the multi-core mobile platforms. GENESYS modelling methodology which employs service-oriented and component based distributed application design has been extended for this purpose such that application level services are refined to platform-level services allowing mapping of GENESYS application architecture to workload models used in performance evaluation. This results in easy extraction of application workload models, reducing the time and effort in the performance evaluation phase needed for architectural exploration. This article presents the way brisk performance evaluation of distributed GENESYS applications is achieved by employing extended GENESYS distributed application architecture. The approach is experimented with a case study. UML2.0 MARTE profile, Papyrus UML2.0 modelling tool and SystemC were used for modelling and simulation. Subayal Khan, Jukka Saastamoinen, Kari Tiensyrjä, Jari Nurmi |
DASC | 4 |
| 2011 | Relative positioning of mass market devices in ad-hoc networksabstractIn this paper we present a practical approach to relative positioning of heterogeneous mass market devices. We propose a user-centric solution exploiting measurements from on-the-fly fixed reference points detected during scanning procedures. The technique implemented on mobile devices is able to locate neighboring nodes, demonstrating to be a feasible and practical solution for sensing spatial social contexts in ad-hoc environments. Francescantonio Della Rosa, Helena Leppäkoski, Jari Nurmi |
IPIN | 4 |
| 2010 | Evaluation of Radix-2 and Radix-4 FFT processing on a reconfigurable platformabstractIn this paper, we present the mapping of Radix-2 and Radix-4 FFT algorithms using CREMA, a Coarse-Grain Reconfigurable Array (CGRA) with mapping adaptiveness. CREMA is employed to generate special purpose accelerators tailored on the specified mapping. We analyze the results for 64-point FFT targeted for OFDM processing. Both of the implementations are (4-5X) smaller and showed a speed-up of 4X when compared with a general-purpose CGRA. Waqar Hussain 0001, Fabio Garzia, Jari Nurmi |
DDECS | 3 |
| 2010 | Instantiating GENESYS Application Architecture Modeling via UML 2.0 Constructs and MARTE ProfileabstractModeling of complex and computationally intense applications supported by modern mobile devices via standard modeling languages is a challenging task. Within the GENESYS process model the application modeling phase is thus of key importance. GENESYS manages complexity by employing cross domain and platform-based application design. The main contribution of this article is to describe the instantiation of GENESYS application architecture modeling via MARTE profile and describe a methodology for validation of nonfunctional properties annotated in the application model. Subayal Khan, Kari Tiensyrjä, Jari Nurmi |
DSD | 3 |
| 2010 | Control Techniques for Coupling a Coarse-Grain Reconfigurable Array with a Generic RISC CoreabstractThis paper presents three different control techniques to couple a Coarse-Grain Reconfigurable Architecture (CGRA) with a generic RISC processor. In the architecture under study the CGRA, i.e., a coarse-grain array, works as co-processor and is used to accelerate a kernel selected by the application developer. The array is meant to perform the data processing operations of the kernel, while the RISC processor takes care of the control of the kernel global execution. The control techniques proposed here do not add any dedicated instructions to the ISA of the RISC processor, but are mostly based on load/store operations. The first method is the slowest one, in which the control operations are executed sequentially with the array processing. The second approach enables some parallelism between control operations and array processing. The parallelism is guaranteed by the replacement of a single control register with a two-register delay chain. The last approach is the fastest, which replaces the delay chain with a FIFO. In particular, a 64-point FFT test case shows that the last solution is 2.5× faster than the second one. Fabio Garzia, Waqar Hussain 0001, Jari Nurmi |
FPL | 3 |
| 2010 | Ad-hoc networks aiding indoor calibrations of heterogeneous devices for Fingerprinting applicationsabstractIn this paper we propose a Collaborative Mapping (CM) method based on the exploitation of the WLAN Received Signal Strength (RSS) measured from short-range ad-hoc links between neighboring devices. The estimated spatial proximity allows on-the-fly calibrations of heterogeneous clients for Location Fingerprinting (LF) applications. This method can avoid long time-consuming and battery-draining calibrations when implementing Fingerprinting Location applications running on heterogeneous mass market devices. Francescantonio Della Rosa, Helena Leppäkoski, Stefano Biancullo, Jari Nurmi |
IPIN | 4 |
| 2010 | A coarse-grain reconfigurable architecture for multimedia applications supporting subword and floating-point calculations
Claudio Brunelli, Fabio Garzia, Davide Rossi 0001, Jari Nurmi |
J. Syst. Archit. | 4 |
| 2009 | CREMA: A coarse-grain reconfigurable array with mapping adaptivenessabstractThis paper presents CREMA, a coarse-grain reconfigurable array with mapping adaptiveness. Mapping adaptiveness consists of tailoring the array to a specific application requirements. Run-time reconfigurability allows the re-usage of same PE with different functionality and interconnections among the ones supported. We proved this approach very efficient if compared with a standard CGRA. In our test cases CREMA gets a performance speed-up from 1.5X to 4X, reducing in the same time the area occupation by 80%-90% in comparison with butter CGRA. Fabio Garzia, Waqar Hussain 0001, Jari Nurmi |
FPL | 3 |
| 2008 | Improving the Efficiency of Run Time Reconfigurable Devices by Configuration LockingabstractRun-time reconfigurable logic is a very attractive alterative in the design of SoC. However, configuration overhead can largely decrease the system performance. In this work, we present a novel configuration locking technique to reduce the effect of the overhead. The idea is to at run-time lock a number of the most frequently used tasks on the configuration memory so that they cannot be evicted by other tasks. With real applications in validation, the results show that using proper amount of resources to lock tasks can significantly outperform simply using more resources. In addition, an algorithm has been developed for estimating the lock ratio. Experimental results show that the estimates are close to optimal results and the measured computer runtime is less than 4 us in a commercial embedded processor. Juha-Pekka Soininen, Jari Nurmi |
DATE | 3 |
| 2008 | A dedicated DMA logic addressing a time multiplexed memory to reduce the effects of the system bus bottleneckabstractA very common problem which affects the performance of bus-based computing systems arises from the fact that the bus is a common resource which needs to be shared between a number of master devices. The common resource contention forces to stall temporarily the execution of one or more of the bus masters, slowing down the execution. Moreover, the width of the bus is usually relatively small, forcing the bus master to perform several bus cycles in order to transfer a data block from the main memory to a peripheral (or to a processing element), and the other way around. The combination of these factors leads to problems and inefficiencies which designers need to solve. In this paper we present a dedicated hardware used to allow an external accelerator to access the system memory independently from the main microprocessor. The proposed device is able to exchange data with the memory in a DMA-like fashion, to generate properly memory addresses in order to access it in an efficient way. Results show that using such a solution it is possible to reach a considerable speed-up in the execution of a given algorithm. Claudio Brunelli, Fabio Garzia, Carmelo Giliberto, Jari Nurmi |
FPL | 4 |
| 2008 | Reconfigurable hardware: The holy grail of matching performance with programming productivityabstractMany reconfigurable hardware architectures have been proposed so far, ranging from FPGAs to coarse grained architectures. Reconfigurability can be intended in several ways, and a number of diverse solutions have been proposed. One of the most relevant issues that have emerged is that the performance gain offered by reconfigurable hardware is balanced by relevant difficulties in their programming, which often inhibit its utilization in many appealing fields and ultimately jeopardize its diffusion. The solution for enabling higher productivity in application mapping likely do not reside only in the development of better tools, but also of more usable designs. This paper gives an overview of different reconfigurable architectures and related design flows proposed over the last years, including commercial offers and efforts coming from academia, analyzes the challenges they pose to the application developer and focuses on the latest alternatives to mitigate the design productivity issue. Claudio Brunelli, Fabio Garzia, Jari Nurmi, Fabio Campi, Damien Picard |
FPL | 3 |
| 2008 | Implementation of a floating-point matrix-vector multiplication on a reconfigurable architectureabstractThis paper describes the implementation of a floating-point 4times4 matrix-vector multiplication on a reconfigurable system. The 4times4 matrix-vector multiplication is meant to be used to perform two steps (transformation and perspective projection) of a 3D graphics application. The target system is based on a bus architecture with a general purpose core as master and the reconfigurable array as main accelerator. The system has been prototyped on a FPGA device. The matrix-vector multiplication has been successfully implemented on the reconfigurable block. Compared to the general purpose implementation it is convenient if the number of vectors to process is higher than seven. If hundreds of vectors are processed, the speed-up achievable reaches 3times. Fabio Garzia, Claudio Brunelli, Davide Rossi 0001, Jari Nurmi |
IPDPS | 4 |
| 2008 | Design space exploration of an open-source, IP-reusable, scalable floating-point engine for embedded applications
Claudio Brunelli, Fabio Campi, Claudio Mucci, Davide Rossi 0001, Tapani Ahonen, Juha Kylliäinen, Fabio Garzia, Jari Nurmi |
J. Syst. Archit. | 8 |
| 2007 | Interactive presentation: Using dynamic voltage scaling to reduce the configuration energy of run time reconfigurable devicesabstractIn this paper, an approach that uses dynamic voltage scaling (DVS) to reduce the configuration energy of runtime reconfigurable devices is proposed. The basic idea is to use configuration prefetching and parallelism to create excessive system idle time and apply DVS on the configuration process when such idle time can be utilized. A genetic algorithm is developed to solve the task scheduling and voltage assignment problem. With real applications, the results show that up to 19.3% of configuration energy can be reduced. When considering the reduction of the configuration energy, the results show that using more computation resources is more favorable when the configuration latency is relatively small, and using more configuration controllers is more favorable for relatively large latency Juha-Pekka Soininen, Jari Nurmi |
DATE | 3 |
| 2007 | A Genetic Algorithm for Scheduling Tasks onto Dynamically Reconfigurable HardwareabstractIn this paper, a genetic algorithm (GA) for scheduling tasks onto dynamically reconfigurable devices is presented. The scheduling problem is NP-hard and more complicated than multiprocessor scheduling, because both the task allocation and the configurations need to be carefully managed. The approach has been validated with a number of random task graphs. The results show that the GA approach has good convergence and it is in average 8.6% better than a list-based scheduler for large task graphs of various sizes. Juha-Pekka Soininen, Jari Nurmi |
ISCAS | 3 |
| 2007 | System-Level Design for Partially Reconfigurable HardwareabstractThis paper presents a SystemC-based approach for system-level design of partially reconfigurable hardware. The main focuses are resource estimation to support system analysis, reconfiguration modeling for fast performance simulation, automatic generation of reconfigurable components and a static prefetch scheduler. The approach was applied in a real design case of a part of a WCDMA decoding algorithm on a commercial reconfigurable platform. Kari Tiensyrjä, Juha-Pekka Soininen, Jari Nurmi |
ISCAS | 4 |
| 2007 | Static scheduling techniques for dependent tasks on dynamically reconfigurable devices
Juha-Pekka Soininen, Jari Nurmi |
J. Syst. Archit. | 3 |
| 2007 | Applying CDMA Technique to Network-on-ChipabstractThe issues of applying the code-division multiple access (CDMA) technique to an on-chip packet switched communication network are discussed in this paper. A packet switched network-on-chip (NoC) that applies the CDMA technique is realized in register-transfer level (RTL) using VHDL. The realized CDMA NoC supports the globally-asynchronous locally-synchronous (GALS) communication scheme by applying both synchronous and asynchronous designs. In a packet switched NoC, which applies a point-to-point connection scheme, e.g., a ring topology NoC, data transfer latency varies largely if the packets are transferred to different destinations or to the same destination through different routes in the network. The CDMA NoC can eliminate the data transfer latency variations by sharing the data communication media among multiple users concurrently. A six-node GALS CDMA on-chip network is modeled and simulated. The characteristics of the CDMA NoC are examined by comparing them with the characteristics of an on-chip bidirectional ring topology network. The simulation results reveal that the data transfer latency in the CDMA NoC is a constant value for a certain length of packet and is equivalent to the best case data transfer latency in the bidirectional ring network when data path width is set to 32 bits. Tapani Ahonen, Jari Nurmi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2006 | A parallel configuration model for reducing the run-time reconfiguration overheadabstractMultitasking on reconfigurable logic can achieve very high silicon reusability. However, configuration latency is a major limitation and it can largely degrade the system performance. One reason is that tasks can run in parallel but configurations of the tasks can be done only in sequence. This work presents a novel configuration model to enable configuration parallelism. It consists of multiple homogeneous tiles and each tile has its own configuration SRAM that can be individually accessed. Thus multiple configuration controllers can load tasks in parallel and more speedups can be achieved. We used a prefetch scheduling technique to evaluate the model with randomly generated tasks. The experiment results reveal that in average using multiple controllers can reduce the configuration overheads by 21%. Compared to best cases of using multiple tiles with a single controller, additional 40% speedup can be achieved using multiple controllers. Juha-Pekka Soininen, Jari Nurmi |
DATE | 3 |
| 2006 | On-Line Reconfigurable XGFT Network-on-Chip Designed for Improving the Fault-Tolerance and Manufacturability of the MPSoC ChipsabstractLarge System-on-Chip (SoC) circuits will contain an increasing number of processors which will communicate with each other across Networks-on-Chip (NOC). The faulty processors could be replaced with faultless ones, whereas only a single defect in the NOC can make the whole chip unusable. Therefore, the fault-tolerance of the NOC is a crucial component of the fault-tolerance and manufacturability of the SoCs. This paper presents a fault-tolerant extended generalized fat tree (XGFT) NOC developed for future multi-processor SoCs (MPSoC). Its fault-tolerance is improved with a new version of fault-diagnosis-and-repair (FDAR) system, which makes it possible to diagnose and repair the NOC on-line. It detects such static, dynamic and transient faults which block packets or produce bit errors, and reconfigures the faulty switches to operate correctly. Processors can also use it for reconfiguring the faulty switch nodes after the faults are located with other test methods. Simulation and synthesis results show that slightly defected XGFTs are able to achieve good performance after they are repaired with the FDAR while the costs of the FDAR remain tolerable Heikki Kariniemi, Jari Nurmi |
FPL | 2 |
| 2006 | Prototyping a Globally Asynchronous Locally Synchronous Network-On-Chip on a Conventional FPGA Device Using Synchronous Design ToolsabstractAn FPGA prototype of a four-node globally-asynchronous locally-synchronous network-on-chip is described. The network for global communication operates asynchronously at the link level and synchronously within a node. Two C-element control pipelines constitute the control logic for the asynchronous part. C-element and asynchronous arbiter realizations on FPGA using standard synchronous design tools are presented Tapani Ahonen, Jari Nurmi |
FPL | 3 |
| 2006 | A Wireless MIMO STC OFDM System ImplementationabstractWe present the study and implementation of a 2 times 2 MIMO STC wireless OFDM system. Synchronization and channel estimation algorithms have been inserted in this implementation. The latter is the main focus of the work presented in this paper. The VHDL FPGA implementations have been applied in Elektrobit's OFDM test environment for 4G MIMO systems (EB4G) which consists of high-speed, FPGA-based programmable units. The verification of the design and intensive measurements under various ETSI BRAN channels have been performed using the Elektrobit Propsim C8 hardware multi-channel emulator. The performance follow those of floating-point simulations. The implementation shows the performance attainable by such MIMO techniques in practical conditions and it also runs much faster than software simulations. The design allows further developments and studies towards 4G types of system like OFDMA, for example, with minor changes, as well as different types of impairments, in a quick and practical manner Sandrine Boumard, Matti Weissenfelt, Huageng Chi, Jari Nurmi |
PIMRC | 4 |
| 2005 | Integration of a NoC-Based Multimedia Processing PlatformabstractAt Tampere University of Technology we are developing a multimedia processing platform using previously designed IP components. The utilized components include the Proteo network-on-chip, the coffee processor, the milk floating point coprocessor, and the transport triggered TACO for protocol processing. Unlike shared buses, networks-on-chip support varying levels of communication parallelism depending on the topology. This design case illustrates the need to match the network topology and the interfaces to the computation models. Characteristics of the platform prototype on FPGA are described together with our approach to enable efficient utilization of the communication resources through the bus-oriented standard interfaces used. Tapani Ahonen, Jari Nurmi |
FPL | 2 |
| 2005 | Fault-Tolerant XGFT Network-On-Chip for Multi-Processor System-on-Chip CircuitsabstractThis paper presents a fault-tolerant eXtended Generalized Fat Tree (XGFT) Network-On-Chip (NOC) implemented with a new fault-diagnosis-and-repair (FDAR) system. The FDAR system is able to locate faults and reconfigure switch nodes in such a way that the network can route packets correctly despite the faults. This paper presents how the FDAR finds the faults and reconfigures the switches. Simulation results are used for showing that faulty XGFTs could also achieve good performance, if the FDAR is used. This is possible if deterministic routing is used in faulty parts of the XGFTs and adaptive Turn-Back (TB) routing is used in faultless parts of the network for ensuring good performance and Quality-of-Service (QoS). The XGFT is also equipped with parity bit checks for detecting bit errors from the packets. Heikki Kariniemi, Jari Nurmi |
FPL | 2 |
| 2005 | An Efficient Approach to Hide the Run-Time Reconfiguration from SW ApplicationsabstractDynamically reconfigurable logic is becoming an important design unit in SoC system. A method to make the reconfiguration management transparent to software applications is required in order to make easier the design with such devices. In this paper, we present an efficient approach similar to the cache miss and the data replacement in modern computer system for the task. The main advantage is that the reconfiguration can be correctly issued without extra instructions inserted either manually by SW application programmers or automatically by compilers. The approach was validated in a real case design. In the Virtex2P20 implementation platform, the resource overhead was 2.45% in terms of the number of LUTs. Performance is measured in cycle-accurate simulation environment. The overhead is about equal when compared with an OS-based equivalent design that uses system calls and critical section code to manage the reconfiguration. Juha-Pekka Soininen, Jari Nurmi |
FPL | 3 |
| 2005 | A programmable baseband receiver platform for WCDMA/OFDM mobile terminalsabstractA programmable system platform that enables software defined implementations of WCDMA and OFDM baseband receivers is presented. The design complexity of future wireless terminals and the shrinking time-to-market constraints are the motivators for adopting the platform-based design methodology. The presented hardware platform comprises a RISC core and three tightly coupled coprocessors that are used for the most intensive computation kernels. The application program interface of the platform is provided in the form of a library of special C-functions that are used to pass the input parameters to the coprocessors and initiate the execution. The SystemC-based simulation environment of the platform is described and the simulation results are given. Lasse Harju, Jari Nurmi |
WCNC | 2 |
| 2004 | Virtualizing the Dimensions of a Coarse-Grained Reconfigurable Array
Tapio Ristimäki, Jari Nurmi |
FPL | 2 |
| 2004 | A baseband receiver architecture for UMTS-WLAN interworking applicationsabstractThis paper presents a programmable hardware platform for dual-mode WCDMA/OFDM receiver implementations. The platform is targeted for mobile terminals capable of operating in tight coupling UMTS-WLAN interworking systems. The proposed platform comprises a RISC core and three coprocessors that are used for the most intensive computation kernels. The receiver algorithms needed in WCDMA and OFDM receivers are overviewed and the needed computation resources are specified based on the analysis. The high-level architecture of the dual-mode receiver is also presented. A software development model is specified for the platform. Lasse Harju, Jari Nurmi |
ISCC | 2 |
| 2004 | Issues in the development of a practical NoC: the Proteo concept
David A. Sigüenza-Tortosa, Tapani Ahonen, Jari Nurmi |
Integr. | 3 |
| 2003 | Variable-Length Instruction Compression for Area MinimizationabstractMemories comprise a significant part of chips in embedded applications, thus also contributing considerably to the costs. We present a variable-length compression scheme for reducing program memory footprint in a 32-bit DSP processor. The compression method is based on a static program code analysis. Short operand fields do not provide sufficient repetition individually, so the compression is realized by handling all operands of an instruction as one field. The more a certain operand combination is used, the shorter it is coded. The original combination is placed on a look-up table and the coded bit pattern forms an index to that table. The compression results that are achieved in two test applications are 46% (audio decoder) and 51% (video decoder). Piia Simonen, Ilkka Saastamoinen, Jari Nurmi |
ASAP | 3 |
| 2003 | Reprogrammable Algorithm Accelerator IP Block
Tapio Ristimäki, Jari Nurmi |
VLSI-SOC | 2 |
| 1995 | A Processor Core for 32 kbit/s G.726 ADPCM CodecsabstractThis paper describes an application specific DSP core designed to be used in a CCITT 32 kbit/s G.726 Adaptive Differential Pulse Code Modulation (ADPCM) codec. The instruction set architecture and the programming model of the DSP core were derived from an algorithm profile and complexity analysis and the core was implemented using VHDL and logic synthesis. Architecture design efforts were concentrated on finding the minimum amount of hardware resources which could implement the required functionality within the clock cycle count limit. The result is a Harvard architecture processor core which can be used to implement the 32 kbit/s G.726 ADPCM encoding/decoding functions with very modest external instruction and data memory requirements. In a typical configuration the processor can perform a full encode/decode operation for one sample in less than 1100 clock cycles. A gate-level implementation of less than 4000 gates of silicon area was created using logic synthesis for a standard cell technology. Juhani Vehvilainen, Jari Nurmi |
ISCAS | 2 |
| 1994 | A DSP core for speech coding applicationsabstractAn application specific processor core for mobile speech coding applications has been designed and implemented. Since the architecture is tailored to the application, it has a very low power consumption, making it attractive for handheld devices. The low power consumption, flexible design and high performance have been achieved by a standby mode, optimized full custom design, and a low clock frequency, together with a highly parallel architecture. All the parallelism is accessible to the user on the assembler level. Engineering samples of the processor have been fabricated and tested. The silicon area required for the core is approximately 25 mm/sup 2/ using 1.0 /spl mu/m CMOS. A typical average power consumption for a GSM full rate speech codec implementation using this core is less than 50 mW at 5 V operating voltage, and the complex algorithm is executed in less than 5 ms for each 20 ms speech frame (including encode, decode, VAD and DTX operations).> Jari Nurmi, Ville Eerola, Erwin Ofner, Andreas Gierlinger, Jürgen Jernej, Teppo Karema, Tommi Raita-aho |
ICASSP (2) | 1 |
| 1994 | An Overall FIR Filter Optimization Tool for High Granularity Implementation TechnologiesabstractA toolbox for the optimization of FIR filters for quantized high granularity technologies has been implemented. The main principle is to make the best possible filter that fits into the specified number of logic elements on the target technology. The design environment consists of the toolbox and commercial tools Matlab, Synopsys and XACT. Xilinx XC4000 series FPGAs (field programmable gate arrays) have been used to demonstrate the performance of the optimization tool.> Jouni Isoaho, Jari Nurmi |
ISCAS | 2 |
| 1994 | Multipurpose Chip for Physiological MeasurementsabstractThe idea of a multipurpose chip for the measurement of physiological signals is discussed here. First, the general idea of the chip consisting of an amplifier, a sigma-delta analog-to-digital converter (ADC) and a digital filter is presented. Then the possible implementation alternatives, specifications and, finally, the linear phase filters to maintain some important features in the signal are discussed.> Maini Williams, Jari Nurmi |
ISCAS | 2 |