EDBT 2026 Demo / reviewers in the wild / expert
Steffen Malkowsky
dblp:137/9263
· DBLP profile ↗
11ranked-venue papers
1as first author
4since 2021 · last 2024
0000-0003-2902-6071ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Computer networks · 3Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | The LuViRA Dataset: Synchronized Vision, Radio, and Audio Sensors for Indoor LocalizationabstractWe present a synchronized multisensory dataset for accurate and robust indoor localization: the Lund University Vision, Radio, and Audio (LuViRA) Dataset. The dataset includes color images, corresponding depth maps, inertial measurement unit (IMU) readings, channel response between a 5G massive multiple-input and multiple-output (MIMO) testbed and user equipment, audio recorded by 12 microphones, and accurate six degrees of freedom (6DOF) pose ground truth of 0.5 mm. We synchronize these sensors to ensure that all data is recorded simultaneously. A camera, speaker, and transmit antenna are placed on top of a slowly moving service robot, and 89 trajectories are recorded. Each trajectory includes 20 to 50 seconds of recorded sensor data and ground truth labels. Data from different sensors can be used separately or jointly to perform localization tasks, and data from the motion capture (mocap) system is used to verify the results obtained by the localization algorithms. The main aim of this dataset is to enable research on sensor fusion with the most commonly used sensors for localization tasks. Moreover, the full dataset or some parts of it can also be used for other research areas such as channel estimation, image classification, etc. Our dataset is available at: https://github.com/ilaydayaman/LuViRA_Dataset Ilayda Yaman, Guoda Tian, Martin Larsson, Patrik Persson, Michiel Sandra, Alexander Dürr, Erik Tegler, Nikhil Challa, Henrik Garde, Fredrik Tufvesson, Kalle Åström, Ove Edfors, Steffen Malkowsky, Liang Liu 0002 |
ICRA | 13 |
| 2022 | An Application Specific Vector Processor for Efficient Massive MIMO ProcessingabstractThis paper presents an implementation for a baseband massive multiple-input multiple-output (MIMO) application-specific instruction set processor (ASIP). The ASIP is geared with vector processing capabilities in the form of single instruction multiple data (SIMD), and furthermore exploits instruction level parallelism by employing a very large instruction word (VLIW) architecture. Additionally, a systolic array is built into the pipeline which is tuned to speed up matrix calculations. A parallel memory subsystem and stand-alone accelerators are integrated into the ASIP architecture in order to meet the processing requirement. The processor is synthesized in 22 nm FD-SOI technology running at a clock frequency of 800 MHz. The system achieves a maximum detection throughput of 0.75 Gb/s/mm2for a$128\times 8$massive MIMO system. Mohammad Attari, Lucas Ferreira, Liang Liu 0002, Steffen Malkowsky |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | An Application Specific Vector Processor for CNN-Based Massive MIMO PositioningabstractThis paper sets out to create an implementation for fingerprint-based positioning using massive multiple-input multiple-output (MIMO) technology, by means of deep convolutional neural networks (CNN), and utilizing the wireless channel state information (CSI). Due to the sheer volume of computational requirements imposed by CNN processing, an accelerator- assisted design is well-suited to the task at hand. Consequently, an application specific instruction set processor (ASIP) is designed to combine flexibility with implementation efficiency. This ASIP is equipped with vector processing capabilities employing a single instruction multiple data (SIMD) scheme, and additionally has a very large instruction word (VLIW) architecture to further exploit instruction-level parallelism. A configurable 2D array of processing engines (PE) is integrated into the processor, in a tightly coupled manner, to accelerate the CNN operation. Synthesis results will be demonstrated using the GF-22 nm FD- SOI technology with a clock frequency of 555 MHz. The system can achieve a throughput of 271 positionings/s, with an average positioning error of 3.5 λ (40 cm) at a carrier frequency of 2.6 GHz. Mohammad Attari, Jesus Rodriguez Sanchez, Liang Liu 0002, Steffen Malkowsky |
ISCAS | 4 |
| 2021 | Reconfigurable Multi-Access Pattern Vector Memory for Real-Time ORB Feature ExtractionabstractThis work presents an on-chip memory subsystem envisioned for real-time applications performing Oriented FAST and Rotated Brief (ORB) feature extraction for Simultaneous Localization and Mapping (SLAM) systems. For autonomous navigation of battery-powered devices, feature-based SLAM is a computationally frugal alternative to direct methods. This paper thoroughly analyses ORB multiple memory access patterns, exploring possible systematic parallelism and hardware-biased algorithmic enhancements, alleviating requirements on bandwidth and reducing redundant accesses. Enabling those, a suitable multi-bank parallel memory featuring run-time reconfigurable address generation, image allotment, and close-to-memory data-shuffling is proposed. As case study, a 30 Frames-Per-Second (FPS) VGA-resolution ORB-capable 8-bank memory is evaluated using 22 FDX technology, running at 909 MHz, with a negligible area overhead of 0.3%, reducing operand accesses between 54 - 160× relative to Sudoku-like and scalar memories. Lucas Ferreira, Steffen Malkowsky, Patrik Persson, Kalle Åström, Liang Liu 0002 |
ISCAS | 2 |
| 2019 | A Programmable 16-Lane SIMD ASIP for Massive MIMOabstractThis paper presents a 16-lane, 16-bit complex application-specific instruction processor (ASIP) for baseband processing in massive multiple-input multiple-output (MIMO). The architecture utilizes a 3/4-way very large instruction word (VLIW) with highly efficient pre- and post-processing units specifically trimmed for massive MIMO requirements. Architecture optimizations include features like single cycle vector-dot-product, vector indexing and broadcasting, hardware loops and full complex accumulator to provide high performance for various massive MIMO algorithms. Moreover, the ASIP is fully C-programmable, which is crucial for adapting to the evolving 5G standard. In our evaluation, a full massive MIMO up-link detection is executed in ≈11k clock cycles while synthesis results in ST 28 nm FD-SOI suggest a clock frequency of 900 MHz equating in a detection throughput of 330 Mb/s for a 128×16 massive MIMO system. Steffen Malkowsky, Hemanth Prabhu, Liang Liu 0002, Ove Edfors, Viktor Öwall |
ISCAS | 1 |
| 2017 | Temporal Analysis of Measured LOS Massive MIMO Channels with MobilityabstractThe first measured results for massive multiple-input, multiple-output (MIMO) performance in a line-of-sight (LOS) scenario with moderate mobility are presented, with 8 users served by a 100 antenna base Station (BS) at 3.7 GHz. When such a large number of channels dynamically change, the inherent propagation and processing delay has a critical relationship with the rate of change, as the use of outdated channel information can result in severe detection and precoding inaccuracies. For the downlink (DL) in particular, a time division duplex (TDD) configuration synonymous with massive MIMO deployments could mean only the uplink (UL) is usable in extreme cases. Therefore, it is of great interest to investigate the impact of mobility on massive MIMO performance and consider ways to combat the potential limitations. In a mobile scenario with moving cars and pedestrians, the correlation of the MIMO channel vector over time is inspected for vehicles moving up to 29km/h. For a 100 antenna system, it is found that the channel state information (CSI) update rate requirement may increase by 7 times when compared to an 8 antenna system, whilst the power control update rate could be decreased by at least 5 times relative to a single antenna system. Paul Harris 0001, Steffen Malkowsky, Joao Vieira, Fredrik Tufvesson, Wael Boukley Hasan, Liang Liu 0002, Mark A. Beach, Simon Armour, Ove Edfors |
VTC Spring | 2 |
| 2017 | Performance Characterization of a Real-Time Massive MIMO System With LOS Mobile ChannelsabstractThe first measured results for massive multiple-input, multiple-output (MIMO) performance in a line-of-sight scenario with moderate mobility are presented, with eight users served in real time using a 100-antenna base station at 3.7 GHz. When such a large number of channels dynamically change, the inherent propagation and processing delay has a critical relationship with the rate of change, as the use of outdated channel information can result in severe detection and precoding inaccuracies. For the downlink (DL) in particular, a time-division duplex configuration synonymous with massive MIMO deployments could mean only the uplink (UL) is usable in extreme cases. Therefore, it is of great interest to investigate the impact of mobility on massive MIMO performance and consider ways to combat the potential limitations. In a mobile scenario with moving cars and pedestrians, the massive MIMO channel is sampled across many points in space to build a picture of the overall user orthogonality, and the impact of both azimuth and elevation array configurations are considered. Temporal analysis is also conducted for vehicles moving up to 29 km/h and real-time bit-error rates for both the UL and DL without power control are presented. For a 100-antenna system, it is found that the channel state information update rate requirement may increase by seven times when compared with an eight-antenna system, whilst the power control update rate could be decreased by at least five times relative to a single antenna system. Paul Harris 0001, Steffen Malkowsky, Joao Vieira, Erik L. Bengtsson, Fredrik Tufvesson, Wael Boukley Hasan, Liang Liu 0002, Mark A. Beach, Simon Armour, Ove Edfors |
IEEE J. Sel. Areas Commun. | 2 |
| 2017 | Reciprocity Calibration for Massive MIMO: Proposal, Modeling, and ValidationabstractThis paper presents a mutual coupling-based calibration method for time-division-duplex massive MIMO systems, which enables downlink precoding based on uplink channel estimates. The entire calibration procedure is carried out solely at the base station (BS) side by sounding all BS antenna pairs. An expectation-maximization (EM) algorithm is derived, which processes the measured channels in order to estimate calibration coefficients. The EM algorithm outperforms the current state-of-the-art narrow-band calibration schemes in a mean squared error and sum-rate capacity sense. Like its predecessors, the EM algorithm is general in the sense that it is not only suitable to calibrate a co-located massive MIMO BS, but also very suitable for calibrating multiple BSs in distributed MIMO systems. The proposed method is validated with experimental evidence obtained from a massive MIMO testbed. In addition, we address the estimated narrow-band calibration coefficients as a stochastic process across frequency, and study the subspace of this process based on measurement data. With the insights of this study, we propose an estimator which exploits the structure of the process in order to reduce the calibration error across frequency. A model for the calibration error is also proposed based on the asymptotic properties of the estimator, and is validated with measurement results. Joao Vieira, Fredrik Rusek, Ove Edfors, Steffen Malkowsky, Liang Liu 0002, Fredrik Tufvesson |
IEEE Trans. Wirel. Commun. | 4 |
| 2016 | Transmission Schemes for Multiple Antenna Terminals in Real Massive MIMO SystemsabstractIn massive MIMO performance evaluations it is often assumed that the terminal has a single antenna. The combination of multiple antennas in a terminal and massive MIMO precoding at the base station side can further improve overall system performance. We present measurement results for multi antenna terminals operating in different transmission schemes and how they perform under varying loading conditions. Gain expressions are derived that enable easy comparison between the transmission schemes. The evaluation is performed on realistic antennas integrated into Sony Xperia handsets tuned to 3.7 GHz and operated together with the Lund University massive MIMO (LuMaMi) test bed. It is concluded that the approach used in today's mobile systems, where up link and down link are addressed independently, will not provide the best performance. The performance can be improved by the selection of transmission schemes optimized for massive MIMO. Erik L. Bengtsson, Peter C. Karlsson, Fredrik Tufvesson, Joao Vieira, Steffen Malkowsky, Liang Liu 0002, Fredrik Rusek, Ove Edfors |
GLOBECOM | 5 |
| 2015 | Stress Test of Vehicular Communication Transceivers Using Software Defined RadioabstractWireless vehicular communication is, in contrast to other terrestrial types of wireless communications, more dynamic in nature. Both the transmitter and the receiver are moving at high speeds relative to each other, which generates highly dynamic wireless channels. Such channels are characterized by short stationarity regions and large Doppler spreads. Modem manufacturers face a challenge when designing and implementing equipment for such environments. Similarly, for testing and evaluation real-life measurements with vehicles are required, which often is an expensive and slow process. This paper tackles this problem by proposing a method for stress testing transceivers based on the design and implementation of a real-time wireless channel emulator for wireless vehicular communications using a software defined radio (SDR). The emulator together with the proposed test methodology enable quick on-bench evaluation of wireless modems. In the paper we also apply the test on two different IEEE 802.11p modem implementations and characterize the packet error rate performance for different Doppler-delay combinations. Dimitrios Vlastaras, Steffen Malkowsky, Fredrik Tufvesson |
VTC Spring | 2 |
| 2013 | A 65-nm CMOS area optimized de-synchronization flow for sub-VT designsabstractThis paper proposes a process independent post layout de-synchronization flow implemented in tool command language working on designs operating in the sub-VTregime. The overhead due to the self-timed operation is combated by introducing full-custom delay elements and latches for a standard 65-nm CMOS process. The flow offers the possibility to adjust granularity based on user requirements. Case studies with different reference designs manifested an average reduction of area and power overhead from 105% to 9% and 174% to 58% in comparison to a full standard cell de-synchronization approach. Christoph Thomas Muller, Steffen Malkowsky, Oskar Andersson, Jens Sparsø, Joachim Neves Rodrigues |
VLSI-SoC | 2 |