VLDB 2026 Research / reviewers in the wild / expert
Liang Liu 0002
dblp:10/6178-2
· DBLP profile ↗
44ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0001-9491-8821ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 5 first-author · 7 since 2021Computer networks · 9 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Software acceleration of multi-user MIMO uplink detection on GPUabstractThis paper presents the exploration of GPU-accelerated block-wise decompositions for zero-forcing (ZF) based QR and Cholesky methods applied to massive multiple-input multiple-output (MIMO) uplink detection algorithms. Three algorithms are evaluated: ZF with block Cholesky decomposition, ZF with block QR decomposition (QRD), and minimum mean square error (MMSE) with block Cholesky decomposition. The latter was the only one previously explored, but it used standard Cholesky decomposition. Our approach achieves an 11% improvement over the previous GPU-accelerated MMSE study. Through performance analysis, we observe a trade-off between precision and execution time. Reducing precision from FP64 to FP32 improves execution time but increases bit error rate (BER), with ZF-based QRD reducing execution time from 2 . 04 μ s to 1 . 24 μ s for a 128 × 8 MIMO size. The study also highlights that larger MIMO sizes, particularly 2048 × 32, require GPUs to fully utilize their computational and memory capabilities, especially under FP64 precision. In contrast, smaller matrices are compute-bound. Our results recommend GPUs for larger MIMO sizes, as they offer the parallelism and memory resources necessary to efficiently handle the computational demands of next-generation networks. This work paves the way for scalable, GPU-based massive MIMO uplink detection systems. Ali Nada, Hazem Ismail Abdel Aziz Ali, Liang Liu 0002, Yousra Al-Kabani |
Parallel Comput. | 3 |
| 2024 | The LuViRA Dataset: Synchronized Vision, Radio, and Audio Sensors for Indoor LocalizationabstractWe present a synchronized multisensory dataset for accurate and robust indoor localization: the Lund University Vision, Radio, and Audio (LuViRA) Dataset. The dataset includes color images, corresponding depth maps, inertial measurement unit (IMU) readings, channel response between a 5G massive multiple-input and multiple-output (MIMO) testbed and user equipment, audio recorded by 12 microphones, and accurate six degrees of freedom (6DOF) pose ground truth of 0.5 mm. We synchronize these sensors to ensure that all data is recorded simultaneously. A camera, speaker, and transmit antenna are placed on top of a slowly moving service robot, and 89 trajectories are recorded. Each trajectory includes 20 to 50 seconds of recorded sensor data and ground truth labels. Data from different sensors can be used separately or jointly to perform localization tasks, and data from the motion capture (mocap) system is used to verify the results obtained by the localization algorithms. The main aim of this dataset is to enable research on sensor fusion with the most commonly used sensors for localization tasks. Moreover, the full dataset or some parts of it can also be used for other research areas such as channel estimation, image classification, etc. Our dataset is available at: https://github.com/ilaydayaman/LuViRA_Dataset Ilayda Yaman, Guoda Tian, Martin Larsson, Patrik Persson, Michiel Sandra, Alexander Dürr, Erik Tegler, Nikhil Challa, Henrik Garde, Fredrik Tufvesson, Kalle Åström, Ove Edfors, Steffen Malkowsky, Liang Liu 0002 |
ICRA | 14 |
| 2023 | High-Precision Machine-Learning Based Indoor Localization with Massive MIMO SystemabstractHigh-precision cellular-based localization is one of the key technologies for next-generation communication systems. In this paper, we investigate the potential of applying machine learning (ML) to a massive multiple-input multiple-output (MIMO) system to enhance localization accuracy. We analyze a new ML-based localization pipeline that has two parallel fully connected neural networks (FCNN). The first FCNN takes the instantaneous spatial covariance matrix to capture angular information, while the second FCNN takes the channel impulse responses to capture delay information. We fuse the estimated coordinates of these two FCNNs for further accuracy improvement. To test the localization algorithm, we performed an indoor measurement campaign with a massive MIMO testbed at 3.7 GHz. In the measured scenario, the proposed pipeline can achieve centimeter-level accuracy by combining delay and angular information. Guoda Tian, Ilayda Yaman, Michiel Sandra, Xuesong Cai, Liang Liu 0002, Fredrik Tufvesson |
ICC | 5 |
| 2023 | LuMaMi28: Real-Time Millimeter-Wave Multi-User MIMO Systems With Antenna SelectionabstractThis paper presents LuMaMi28, a real-time 28 GHz multi-user (MU) multiple-input multiple-output (MIMO) testbed. In this testbed, the base station has 16 transceiver chains with a fully-digital beamforming architecture (with different pre-coding algorithms) and simultaneously supports multiple user equipments (UEs) with spatial multiplexing. The UEs are equipped with a beam-switchable antenna array for real-time antenna selection where the one with the highest channel magnitude, out of four pre-defined beams, is selected. For the beam-switchable antenna array, we consider two kinds of UE antennas, with different beam-width and different peak-gain. Based on this testbed, we provide measurement results for millimeter-wave (mmWave) MU-MIMO performance in different real-life scenarios with static and mobile UEs. We explore the potential benefit of the mmWave MU-MIMO systems with antenna selection based on measured channel data, and discuss the performance results through real-time measurements. MinKeun Chung, Liang Liu 0002, Andreas Johansson, Sara Willhammar, Zhinong Ying, Olof Zander, Kamal Samanta, Chris Clifton, Toshiyuki Koimori, Shinya Morita, Satoshi Taniguchi, Fredrik Tufvesson, Ove Edfors |
IEEE Trans. Wirel. Commun. | 2 |
| 2022 | System Design and Performance for Antenna Reservation in Massive MIMOabstractPeak to average power (PAPR) reduction of OFDM signals is critical in order to improve power amplifier (PA) efficiency in base stations. For massive MIMO, the complexity of these methods can become a real bottleneck in implementing low power digital signal processing chains. In this work, we consider an antenna reservation technique, which uses a low complexity clipping method to reduce signal peaks and leverages the benefit of massive antennas, by reserving a subset of antennas in order to compensate for the clipping distortion. Reserving antennas on the other hand reduces the potential array gain in the massive MIMO system, complicating the application of antenna reservation. This work explores various design space parameters in antenna reservation such as number of reserved antennas, amount of peak reduction and clipping methods. We investigate the impact of these parameters on the error vector magnitude at the user and on the adjacent channel power ratio at both transmitter and user positions. Our results enable a deeper understanding of antenna reservation as a low complexity PAPR reduction method in massive MIMO systems. Sidra Muneer, Jesus Rodriguez Sanchez, Liesbet Van der Perre, Ove Edfors, Henrik Sjöland, Liang Liu 0002 |
VTC Fall | 6 |
| 2022 | An Application Specific Vector Processor for Efficient Massive MIMO ProcessingabstractThis paper presents an implementation for a baseband massive multiple-input multiple-output (MIMO) application-specific instruction set processor (ASIP). The ASIP is geared with vector processing capabilities in the form of single instruction multiple data (SIMD), and furthermore exploits instruction level parallelism by employing a very large instruction word (VLIW) architecture. Additionally, a systolic array is built into the pipeline which is tuned to speed up matrix calculations. A parallel memory subsystem and stand-alone accelerators are integrated into the ASIP architecture in order to meet the processing requirement. The processor is synthesized in 22 nm FD-SOI technology running at a clock frequency of 800 MHz. The system achieves a maximum detection throughput of 0.75 Gb/s/mm2for a$128\times 8$massive MIMO system. Mohammad Attari, Lucas Ferreira, Liang Liu 0002, Steffen Malkowsky |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | Spatially Coupled Serially Concatenated Codes: Performance Evaluation and VLSI Design TradeoffsabstractSpatially coupled serially concatenated codes (SC-SCCs) are constructed by coupling several classical turbo-like component codes. The resulting spatially coupled codes provide a close-to-capacity performance and low error floor, which have attracted a lot of interest in the past few years. The aim of this paper is to perform a comprehensive design space exploration to reveal different aspects of SC-SCCs, which is missing in the literature. More specifically, we investigate the effect of block length, coupling memory, decoding window size, and number of iterations on the decoding performance, complexity, latency, and throughput of SC-SCCs. To this end, we propose two decoding algorithms for the SC-SCCs:block-wiseandwindow-wisedecoders. For these, we present VLSI architectural templates and explore them based on building blocks implemented in 12nm FinFET technology. Linking architectural templates with the new algorithms, we demonstrate various tradeoffs between throughput, silicon area, latency, and decoding performance. Mojtaba Mahdavi 0001, Stefan Weithoffer, Matthias Herrmann, Liang Liu 0002, Ove Edfors, Norbert Wehn, Michael Lentmaier |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | Phase-Noise Compensation for OFDM Systems Exploiting Coherence Bandwidth: Modeling, Algorithms, and AnalysisabstractPhase-noise (PN) estimation and compensation are crucial in millimeter-wave (mmWave) communication systems to achieve high reliability. The PN estimation, however, suffers from high computational complexity due to its fundamental characteristics, such as spectral spreading and fast-varying fluctuations. In this paper, we propose a new framework for low-complexity PN compensation in orthogonal frequency-division multiplexing systems. The proposed framework also includes a pilot allocation strategy to minimize its overhead. The key ideas are to exploit the coherence bandwidth of mmWave systems and to approximate the actual PN spectrum with its dominant components, resulting in a non-iterative solution by using linear minimum mean squared-error estimation. The proposed method obtains a reduction of more than$2.5 \times $in total complexity, as compared to the existing methods. Furthermore, we derive closed-form expressions for normalized mean squared-errors (NMSEs) as a function of critical system parameters, which help in understanding the NMSE behavior in low and high signal-to-noise ratio regimes. Lastly, we study a trade-off between performance and pilot-overhead to provide insight into an appropriate approximation of the PN spectrum. MinKeun Chung, Liang Liu 0002, Ove Edfors |
IEEE Trans. Wirel. Commun. | 2 |
| 2021 | An Application Specific Vector Processor for CNN-Based Massive MIMO PositioningabstractThis paper sets out to create an implementation for fingerprint-based positioning using massive multiple-input multiple-output (MIMO) technology, by means of deep convolutional neural networks (CNN), and utilizing the wireless channel state information (CSI). Due to the sheer volume of computational requirements imposed by CNN processing, an accelerator- assisted design is well-suited to the task at hand. Consequently, an application specific instruction set processor (ASIP) is designed to combine flexibility with implementation efficiency. This ASIP is equipped with vector processing capabilities employing a single instruction multiple data (SIMD) scheme, and additionally has a very large instruction word (VLIW) architecture to further exploit instruction-level parallelism. A configurable 2D array of processing engines (PE) is integrated into the processor, in a tightly coupled manner, to accelerate the CNN operation. Synthesis results will be demonstrated using the GF-22 nm FD- SOI technology with a clock frequency of 555 MHz. The system can achieve a throughput of 271 positionings/s, with an average positioning error of 3.5 λ (40 cm) at a carrier frequency of 2.6 GHz. Mohammad Attari, Jesus Rodriguez Sanchez, Liang Liu 0002, Steffen Malkowsky |
ISCAS | 3 |
| 2021 | Reconfigurable Multi-Access Pattern Vector Memory for Real-Time ORB Feature ExtractionabstractThis work presents an on-chip memory subsystem envisioned for real-time applications performing Oriented FAST and Rotated Brief (ORB) feature extraction for Simultaneous Localization and Mapping (SLAM) systems. For autonomous navigation of battery-powered devices, feature-based SLAM is a computationally frugal alternative to direct methods. This paper thoroughly analyses ORB multiple memory access patterns, exploring possible systematic parallelism and hardware-biased algorithmic enhancements, alleviating requirements on bandwidth and reducing redundant accesses. Enabling those, a suitable multi-bank parallel memory featuring run-time reconfigurable address generation, image allotment, and close-to-memory data-shuffling is proposed. As case study, a 30 Frames-Per-Second (FPS) VGA-resolution ORB-capable 8-bank memory is evaluated using 22 FDX technology, running at 909 MHz, with a negligible area overhead of 0.3%, reducing operand accesses between 54 - 160× relative to Sudoku-like and scalar memories. Lucas Ferreira, Steffen Malkowsky, Patrik Persson, Kalle Åström, Liang Liu 0002 |
ISCAS | 5 |
| 2021 | The Effect of Coupling Memory and Block Length on Spatially Coupled Serially Concatenated CodesabstractSpatially coupled serially concatenated codes (SC-SCCs) are a class of spatially coupled turbo-like codes, which have a close-to-capacity performance and low error floor. In this paper, we perform a comprehensive design space exploration, revealing different aspects of SC-SCCs and discussing various design trade-offs. In particular, we investigate the impact of coupling memory, block length, decoding window size, and number of iterations on the performance, complexity, and latency of SC-SCCs. As a result, we propose design guidelines to make the code design independent of the block length. By introducing a modified window decoding schedule, we are able to demonstrate that the block length and coupling memory can be exchanged flexibly without changing the latency and complexity of decoding and without performance loss. Thus, thanks to spatial coupling, a certain code strength and performance can be achieved by either a very small block length or a large one, while the complexity and latency are fixed. Moreover, our results show that using higher coupling memory with smaller blocks can even improve the performance without increasing the latency and complexity. For all considered cases we observe that the performance of SC-SCCs is improved with respect to the uncoupled ensembles for a fixed latency and complexity. Mojtaba Mahdavi 0001, Muhammad Umar Farooq 0001, Liang Liu 0002, Ove Edfors, Viktor Öwall, Michael Lentmaier |
VTC Spring | 3 |
| 2021 | Power Scaling Laws for Radio Receiver Front EndsabstractIn this paper, we combine practically verified results from circuit theory with communication-theoretic laws. As a result, we obtain closed-form theoretical expressions linking fundamental system design and environment parameters with the power consumption of analog front ends (AFEs) for communication receivers. This collection of scaling laws and bounds is meant to serve as a theoretical reference for practical low power AFE design. We show how AFE power consumption scales with bandwidth,SNDR, andSIR. We build our analysis based on two well established power consumption studies and show that although they have different design approaches, they lead to the same scaling laws. The obtained scaling laws are subsequently used to derive relations between AFE power consumption and several other important communication system parameters, namely, digital modulation constellation size, symbol error probability, error control coding gain, and coding rate. Such relations, in turn, can be used when deciding which system design strategies to adopt for low-power applications. For instance, we show how AFE power scales with environment parameters if the performance is kept constant and we use these results to illustrate that adapting to fading fluctuations can theoretically reduce AFE power consumption by at least 20x. Muris Sarajlic, Ashkan Sheikhi, Liang Liu 0002, Henrik Sjöland, Ove Edfors |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2020 | Low-Complexity Fully-Digital Phase Noise Suppression for Millimeter-Wave SystemsabstractPhase noise (PN) estimation and compensation is needed for millimeter-wave (mmWave) communication systems to achieve high data rates. Conventional approaches for PN suppression suffer from high computational complexity. To overcome this, we present a low-complexity fully-digital PN suppression for mmWave systems. The key ideas are to exploit the coherence bandwidth of a mmWave system and an approximation of the PN spectrum, based on its Npdominant components. Utilizing these features, the joint estimation problem of PN and channel can be reformulated into a system with the same number of equations and unknowns, which enables low-complexity PN suppression by using a linear minimum mean square error (LMMSE) estimator. Furthermore, we propose a method to reduce the hardware cost of the LMMSE estimator by using a B-bit signal-to-noise ratio (SNR) quantization, which has a very low complexity of O(Np2+ NNp), where N is the number of subcarriers in an OFDM symbol. As a proof-of-concept, a low-cost VLSI architecture is presented to realize the proposed method. The proposed architecture in 28 nm CMOS (post-synthesis) results in an area cost of 258 K gate count and power consumption of 19.3 mW at 250 MHz clock rate. MinKeun Chung, Hemanth Prabhu, Farhana Sheikh, Ove Edfors, Liang Liu 0002 |
ISCAS | 5 |
| 2019 | A VLSI Implementation of Angular-Domain Massive MIMO DetectionabstractThis paper presents an angular-domain massive MIMO detector which exploits sparsity in massive MIMO channel along with a reconfigurable systolic array architecture to achieve high area efficiency. The underlying idea is to perform signal detection in the angular domain, where the channel matrix can have much lower dimension due the limited number of dominant angles (of arrival and departure) of the wireless signal. Evaluated using the measured massive MIMO channel, the proposed method results in 40%-70% reduction in processing complexity and memory requirements compared to traditional antenna-domain detection. This complexity reduction enables extensive hardware reuse where all the operations of detection processing are mapped to a condensed reconfigurable systolic array. The angular-domain zero-forcing detector, which supports 128 base station antennas and 16 users is implemented in a 28 nm FD-SOI technology. Synthesis result shows that our design attains a throughput of 510 MSps with an area of 537 kG. Mojtaba Mahdavi 0001, Ove Edfors, Viktor Öwall, Liang Liu 0002 |
ISCAS | 4 |
| 2019 | A Programmable 16-Lane SIMD ASIP for Massive MIMOabstractThis paper presents a 16-lane, 16-bit complex application-specific instruction processor (ASIP) for baseband processing in massive multiple-input multiple-output (MIMO). The architecture utilizes a 3/4-way very large instruction word (VLIW) with highly efficient pre- and post-processing units specifically trimmed for massive MIMO requirements. Architecture optimizations include features like single cycle vector-dot-product, vector indexing and broadcasting, hardware loops and full complex accumulator to provide high performance for various massive MIMO algorithms. Moreover, the ASIP is fully C-programmable, which is crucial for adapting to the evolving 5G standard. In our evaluation, a full massive MIMO up-link detection is executed in ≈11k clock cycles while synthesis results in ST 28 nm FD-SOI suggest a clock frequency of 900 MHz equating in a detection throughput of 330 Mb/s for a 128×16 massive MIMO system. Steffen Malkowsky, Hemanth Prabhu, Liang Liu 0002, Ove Edfors, Viktor Öwall |
ISCAS | 3 |
| 2019 | Decentralized Equalizer Construction for Large Intelligent SurfacesabstractIn this paper we present fully decentralized methods for calculating an approximate zero-forzing (ZF) equalizer in a large intelligent surface (LIS). A LIS is intended for wireless communication and facilitates unprecedented MU- MIMO performance, far superior to that of Massive MIMO. Antenna modules in the grid connect to their neighbors to exchange messages of information needed for interference cancellation in a fully-decentralized fashion, making the system scalable. By a careful design of how the messages are routed, we show that the proposed method is able to cancel inter-user interference sufficiently well without any centralized coordination, opening the door for the realization of this type of structures. Juan Vidal Alegría, Jesus Rodriguez Sanchez, Fredrik Rusek, Liang Liu 0002, Ove Edfors |
VTC Fall | 4 |
| 2018 | Impact of Relay Cooperation on the Performance of Large-Scale Multipair Two-Way Relay NetworksabstractWe consider a multipair two-way relay communication network, where pairs of user devices exchange information via a relay system. The communication between users employs time division duplex, with all users transmitting simultaneously to relays in one time slot and relays sending the processed information to all users in the next time slot. The relay system consists of a large number of single antenna units that can form groups. Within each group, relays exchange channel state information (CSI), signals received in the uplink and signals intended for downlink transmission. On the other hand, per-group CSI and uplink/downlink signals (data) are not exchanged between groups, which perform the data processing completely independently. Assuming that the groups perform zero-forcing in both uplink and downlink, we derive a lower bound for the ergodic sumrate of the described system as a function of the relay group size. By close observation of this lower bound, it is concluded that the sumrate is essentially independent of group size when the group size is much larger than the number of user pairs. This indicates that a very large group of cooperating relays can be substituted by a number of smaller groups, without incurring any significant performance reduction. Moreover, this result implies that relay cooperation is more efficient (in terms of resources spent on cooperation) when several smaller relay groups are used in contrast to a single, large group. Muris Sarajlic, Liang Liu 0002, Fredrik Rusek, Farhana Sheikh, Ove Edfors |
GLOBECOM | 2 |
| 2017 | A Cholesky decomposition based massive MIMO uplink detector with adaptive interpolationabstractAn adaptive uplink detection scheme for a Massive MIMO (MaMi) base station serving up to 16 users is presented. Considering user distribution in a cell, selective matched filtering (MF) is proposed for non-interference limited users and a Cholesky decomposition (CD) based zero-forcing (ZF) detector is implemented for the remaining users. Channel conditions such as coherence bandwidth are exploited to lower computational complexity by interpolating CD outputs. Performance evaluations on measured MaMi channels indicate a reduction in computation count by 60 times with a less than 1 dB loss at an uncoded bit error rate of 10-3. For the CD, a reconfigurable processor optimized for 8×8 matrices with block decomposition extension to support up to 16×16 matrices is presented. Circuit level optimizations in 28 nm FD-SOI resulted in an energy of 1.4 nJ/CD at 400 MHz, and post-layout simulations indicate a 50% reduction in power dissipation when operating with the proposed interpolation based detection scheme compared to traditional ZF detection. Rakesh Gangarajaiah, Hemanth Prabhu, Ove Edfors, Liang Liu 0002 |
ISCAS | 4 |
| 2017 | A low latency and area efficient FFT processor for massive MIMO systemsabstractA low-latency and area-efficient FFT/IFFT scheme is presented. The main idea is to utilize OFDM guard bands to reduce the operation counts and processing time, which results in 42% latency reduction compared to the reported pipelined schemes. To realize this idea, a modified pipelined architecture and an efficient data scheduling scheme are proposed. Furthermore, the proposed architecture is scalable to different FFT sizes and is also reconfigurable to support a wide range of applications. A 2048-point FFT/IFFT processor based on the proposed scheme has been designed, resulting in 1200 clock cycles latency, which can address the low latency demand of massive MIMO systems. Synthesis results in a 28 nm CMOS technology show that proposed design attains a throughput of 1 GS/s when clocked at 500 MHz. Mojtaba Mahdavi 0001, Ove Edfors, Viktor Öwall, Liang Liu 0002 |
ISCAS | 4 |
| 2017 | Temporal Analysis of Measured LOS Massive MIMO Channels with MobilityabstractThe first measured results for massive multiple-input, multiple-output (MIMO) performance in a line-of-sight (LOS) scenario with moderate mobility are presented, with 8 users served by a 100 antenna base Station (BS) at 3.7 GHz. When such a large number of channels dynamically change, the inherent propagation and processing delay has a critical relationship with the rate of change, as the use of outdated channel information can result in severe detection and precoding inaccuracies. For the downlink (DL) in particular, a time division duplex (TDD) configuration synonymous with massive MIMO deployments could mean only the uplink (UL) is usable in extreme cases. Therefore, it is of great interest to investigate the impact of mobility on massive MIMO performance and consider ways to combat the potential limitations. In a mobile scenario with moving cars and pedestrians, the correlation of the MIMO channel vector over time is inspected for vehicles moving up to 29km/h. For a 100 antenna system, it is found that the channel state information (CSI) update rate requirement may increase by 7 times when compared to an 8 antenna system, whilst the power control update rate could be decreased by at least 5 times relative to a single antenna system. Paul Harris 0001, Steffen Malkowsky, Joao Vieira, Fredrik Tufvesson, Wael Boukley Hasan, Liang Liu 0002, Mark A. Beach, Simon Armour, Ove Edfors |
VTC Spring | 6 |
| 2017 | Reducing On-Chip Memory for Massive MIMO Baseband Processing Using Channel CompressionabstractEmploying a large number of antennas at the base station, massive MIMO significantly improves spectral efficiency and transmit power efficiency. On the other hand, massive MIMO also introduces unprecedented implementation challenges, especially in terms of processing and storage of large-size channel state information (CSI) matrices. Since on-chip memory is generally very expensive and has limited storage capacity, this paper uses the concept of on-chip CSI data compression and decompression to reduce memory requirements during baseband processing. To achieve this, massive MIMO channel properties are explored using a hardware-friendly DFT-based compression algorithm. The proposed method is evaluated with measured channel data at 2.6 GHz using a 128-antenna linear array. Simulation results show that aggressive CSI compression can be adopted without significant loss in communication performance, while the DFT-based compression can be conveniently integrated into the on-chip memory. This enables a large reduction of required on-chip memory, with negligible hardware overhead for compression/decompression. Yangxurui Liu, Ove Edfors, Liang Liu 0002, Viktor Öwall |
VTC Fall | 3 |
| 2017 | Performance Characterization of a Real-Time Massive MIMO System With LOS Mobile ChannelsabstractThe first measured results for massive multiple-input, multiple-output (MIMO) performance in a line-of-sight scenario with moderate mobility are presented, with eight users served in real time using a 100-antenna base station at 3.7 GHz. When such a large number of channels dynamically change, the inherent propagation and processing delay has a critical relationship with the rate of change, as the use of outdated channel information can result in severe detection and precoding inaccuracies. For the downlink (DL) in particular, a time-division duplex configuration synonymous with massive MIMO deployments could mean only the uplink (UL) is usable in extreme cases. Therefore, it is of great interest to investigate the impact of mobility on massive MIMO performance and consider ways to combat the potential limitations. In a mobile scenario with moving cars and pedestrians, the massive MIMO channel is sampled across many points in space to build a picture of the overall user orthogonality, and the impact of both azimuth and elevation array configurations are considered. Temporal analysis is also conducted for vehicles moving up to 29 km/h and real-time bit-error rates for both the UL and DL without power control are presented. For a 100-antenna system, it is found that the channel state information update rate requirement may increase by seven times when compared with an eight-antenna system, whilst the power control update rate could be decreased by at least five times relative to a single antenna system. Paul Harris 0001, Steffen Malkowsky, Joao Vieira, Erik L. Bengtsson, Fredrik Tufvesson, Wael Boukley Hasan, Liang Liu 0002, Mark A. Beach, Simon Armour, Ove Edfors |
IEEE J. Sel. Areas Commun. | 7 |
| 2017 | Architecture Design of a Memory Subsystem for Massive MIMO Baseband ProcessingabstractThis brief presents an on-chip memory subsystem for massive multiple-input-multiple-output (MIMO) baseband processing at the base station. In massive MIMO systems, the required memory bandwidth and capacity are orders of magnitude higher than those used in conventional wireless systems, due to the large number of serving antennas. These are further combined with design targets on low access latency and flexibility in data organization and access modes. This brief applies and improves the concept of parallel memories to achieve the challenging design target with low hardware overhead. As a case study, a memory subsystem for 128-antenna and 16-user massive MIMO systems is evaluated using ST 28-nm technology. According to postlayout simulation results, the proposed memory subsystem provides 512-Gb/s throughput and offers 1-Mb capacity with a cost of 0.30 mm2. Yangxurui Liu, Liang Liu 0002, Viktor Öwall |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Reciprocity Calibration for Massive MIMO: Proposal, Modeling, and ValidationabstractThis paper presents a mutual coupling-based calibration method for time-division-duplex massive MIMO systems, which enables downlink precoding based on uplink channel estimates. The entire calibration procedure is carried out solely at the base station (BS) side by sounding all BS antenna pairs. An expectation-maximization (EM) algorithm is derived, which processes the measured channels in order to estimate calibration coefficients. The EM algorithm outperforms the current state-of-the-art narrow-band calibration schemes in a mean squared error and sum-rate capacity sense. Like its predecessors, the EM algorithm is general in the sense that it is not only suitable to calibrate a co-located massive MIMO BS, but also very suitable for calibrating multiple BSs in distributed MIMO systems. The proposed method is validated with experimental evidence obtained from a massive MIMO testbed. In addition, we address the estimated narrow-band calibration coefficients as a stochastic process across frequency, and study the subspace of this process based on measurement data. With the insights of this study, we propose an estimator which exploits the structure of the process in order to reduce the calibration error across frequency. A model for the calibration error is also proposed based on the asymptotic properties of the estimator, and is validated with measurement results. Joao Vieira, Fredrik Rusek, Ove Edfors, Steffen Malkowsky, Liang Liu 0002, Fredrik Tufvesson |
IEEE Trans. Wirel. Commun. | 5 |
| 2016 | Compressor design for silicon debugabstractThe objective of this paper is to design a compressor for silicon debug that is suitable in an industrial Nexus environment. The compressor must operate in real-time and must be lossless. Important for the compressor is high compression ratio, low hardware cost and high throughput. We implemented the compression on an FPGA and we compared our implementation in terms of throughput and hardware cost against other approaches. Lars-Johan Fritz, Liang Liu 0002, Erik Larsson |
ETS | 3 |
| 2016 | Transmission Schemes for Multiple Antenna Terminals in Real Massive MIMO SystemsabstractIn massive MIMO performance evaluations it is often assumed that the terminal has a single antenna. The combination of multiple antennas in a terminal and massive MIMO precoding at the base station side can further improve overall system performance. We present measurement results for multi antenna terminals operating in different transmission schemes and how they perform under varying loading conditions. Gain expressions are derived that enable easy comparison between the transmission schemes. The evaluation is performed on realistic antennas integrated into Sony Xperia handsets tuned to 3.7 GHz and operated together with the Lund University massive MIMO (LuMaMi) test bed. It is concluded that the approach used in today's mobile systems, where up link and down link are addressed independently, will not provide the best performance. The performance can be improved by the selection of transmission schemes optimized for massive MIMO. Erik L. Bengtsson, Peter C. Karlsson, Fredrik Tufvesson, Joao Vieira, Steffen Malkowsky, Liang Liu 0002, Fredrik Rusek, Ove Edfors |
GLOBECOM | 6 |
| 2015 | Modified forced convergence decoding of LDPC codes with optimized decoder parametersabstractReducing the complexity of decoding algorithms for LDPC codes is an important prerequisite for their practical implementation. In this work we propose a reduction of computational complexity targeting the highly reliable codeword bits and show that this approach can be seamlessly merged with the forced convergence scheme. We also show how the minimum achievable complexity of the resulting scheme for given performance constraints can be found by solving a constrained optimization problem, and successfully apply a gradient-descent based stochastic approximation (SA) method for solving this problem. The proposed methods are tested on LDPC codes from the IEEE 802.11n standard. Computational complexity reduction of 55% and a 75% reduction of memory access have been observed. Muris Sarajlic, Liang Liu 0002, Ove Edfors |
PIMRC | 2 |
| 2014 | Hardware efficient approximative matrix inversion for linear pre-coding in massive MIMOabstractThis paper describes a hardware efficient linear precoder for Massive MIMO Base Stations (BSs) comprising a very large number of antennas, say, in the order of 100s, serving multiple users simultaneously. To avoid hardware demanding direct matrix inversions required for the Zero-Forcing (ZF) precoder, we use low complexity Neumann series based approximations. Furthermore, we propose a method to speed-up the convergence of the Neumann series by using tri-diagonal precondition matrices, which lowers the complexity even further. As a proof of concept a flexible VLSI architecture is presented with an implementation supporting matrix inversion of sizes up-to 16×16. In 65 nm CMOS, a throughput of 0.5M matrix inversions per sec is achieved at clock frequency of 420MHz with a 104K gate count. Hemanth Prabhu, Ove Edfors, Joachim Neves Rodrigues, Liang Liu 0002, Fredrik Rusek |
ISCAS | 4 |
| 2014 | Energy efficient SQRD processor for LTE-A using a group-sort update schemeabstractThis paper presents an energy-efficient sorted QR-decomposition (SQRD) processor for 3GPP LTE-Advanced (LTE-A) systems. The processor adopts a hybrid decomposition scheme to reduce computational complexity and provides a wide-range of performance-complexity trade-offs. Based on the energy distribution of spatial channels, it switches between the brute-force SQRD and a low-complexity group-sort QR-update strategy, which is proposed in this work to effectively utilize the LTE-A pilot pattern. As a proof of concept, a run-time reconfigurable vector processor is developed to efficiently implement this adaptive-switching QR decomposition algorithm. In a 65 nm CMOS technology, the proposed SQRD processor occupies 0.71mm2core area and has a throughput of up to 100MQRD/s. Compared to the brute-force approach, an energy reduction of 5 ~ 33% is achieved. Chenxin Zhang, Hemanth Prabhu, Liang Liu 0002, Ove Edfors, Viktor Öwall |
ISCAS | 3 |
| 2014 | Digitally assisted adaptive non-linearity suppression scheme for RF front endsabstractThis paper presents a robust and low-complexity non-linearity suppression scheme for radio frequency (RF) transceiver building blocks to efficiently mitigate intermodulation distortion. The scheme consists of tunable RF components assisted by an auxiliary path equipped with an adaptive digital signal processing algorithm to provide the tuning control. This proposed concept of digitally-assisted tuning is capable of handling a large range of non-linear behaviours without any complexity increase in the expensive RF circuitry and is robust to process, voltage and temperature variations. A case study on the third order intermodulation of the channel select filter for a full 10MHz Long Term Evolution (LTE) reception bandwidth is used to demonstrate the feasibility and effectiveness of the technique. Rakesh Gangarajaiah, Mohammed Abdulaziz, Liang Liu 0002, Henrik Sjöland |
PIMRC | 3 |
| 2014 | Low complexity adaptive channel estimation and QR decomposition for an LTE-A downlinkabstractThis paper presents a link adaptive processor to perform low-complexity channel estimation and QR decomposition (QRD) in Long Term Evolution-Advanced (LTE-A) receivers. The processor utilizes frequency domain correlation of the propagation channel to adaptively avoid unnecessary computations in the received signal processing, achieving significant complexity reduction with negligible performance loss. More specifically, a windowed Discrete Fourier transform (DFT) algorithm is used to detect channel conditions and to compute a minimum number of sparse subcarrier channel estimates required for low complexity linear QRD interpolation. Furthermore, the sparsity of subcarrier channel estimates can be adaptively changed to handle different channel conditions. Simulation results demonstrate a reduction of 40%-80% in computational complexity for different channel models specified in the LTE-A standard. Rakesh Gangarajaiah, Peter Nilsson 0001, Ove Edfors, Liang Liu 0002 |
PIMRC | 4 |
| 2014 | Reducing the complexity of LDPC decoding algorithms: An optimization-oriented approachabstractThis paper presents a structured optimization framework for reducing the computational complexity of LDPC decoders. Subject to specified performance constraints and adaptive to environment conditions, the proposed framework leverages the adjustable performance-complexity tradeoffs of the decoder to deliver satisfying performance with minimum computational complexity. More specifically, two constraint scenarios are studied: the “good-enough” performance and “as-good-as-possible performance”. Moreover, we also investigate the effects of different degrees of freedom in performance-complexity tradeoff adjustments. The effectiveness of the proposed method has been verified by simulating a set of LDPC codes used in IEEE 802.11 and IEEE 802.16 standards. Computational complexity reductions of up to 35% have been observed. Muris Sarajlic, Liang Liu 0002, Ove Edfors |
PIMRC | 2 |
| 2013 | High-throughput hardware-efficient soft-input soft-output MIMO detector for iterative receiversabstractThis paper presents a high-throughput, hardware-efficient iterative soft-input soft-output signal detector for spatial-multiplexing system. The detector provides near-optimal performance with much reduced complexity by adopting imbalanced tree-travel strategy, LLR correction techniques, and iteration-adaptive node-selection method. A multi-stage highly-parallel VLSI architecture is employed to implement the detection algorithm. In a 65-nm CMOS technology, the detector occupies 0.64 mm2core area and shows a peak throughput of 1.2 Gb/s. The energy consumed in the proposed detector is 116.5 pJ/b. Liang Liu 0002 |
ISCAS | 1 |
| 2013 | Energy-Efficient MIMO Detection Using Link-Adaptive Parameter AdjustmentabstractThis paper presents a link-adaptive parameter adjustment scheme for energy efficient signal detection in multiple-input multiple-output (MIMO) systems. Equipped with multiple detection schemes of different performance-energy levels, the proposed detector adapts to instantaneous channel conditions and is adjusted at run-time with the detection scheme satisfying performance requirements with minimum power. To enable an efficient link adaptation, we develop a performance prediction method based on the equivalent SNR mapping technique. It provides accurate-enough BER estimation for non-linear signal detection under selectively fading channels. To demonstrate the effectiveness of the link adaptation technique on power saving, we carried out case studies by simulating a simplified LTE downlink system, where the parameter K of K-Best detection is adaptively adjusted. Simulation results demonstrate that the proposed link-adaptive detection offers significant power savings (53% on average and up to 86% for a 4×4 64-QAM system) compared to the static detection scheme targeting the worst-case environment. Liang Liu 0002 |
VTC Spring | 1 |
| 2013 | A highly parallelized MIMO detector for vector-based reconfigurable architecturesabstractThis paper presents a highly parallelized MIMO signal detection algorithm targeting vector-based reconfigurable architectures. The detector achieves high data-level parallelism and near-ML performance by adopting a vector-architecture-friendly technique - parallel node perturbation. To further reduce the computational complexity, imbalanced node and successive partial node expansion schemes in conjunction with sorted QR decomposition are applied. The effectiveness of the proposed algorithm is evaluated by simulations performed on a simplified 4×4 MIMO LTE-A testbed and operation analysis. Compared to the K-Best detector and fixed-complexity sphere decoder (FSD), the number of visited nodes in the proposed algorithm is reduced by 15 and 1.9 times respectively, with less than 1 dB performance degradation. Benefiting from the fully deterministic non-iterative dataflow structure, reconfiguration rate is 95% less than that of the K-Best detector and 17% less than the case of FSD. Chenxin Zhang, Liang Liu 0002, Meifang Zhu, Ove Edfors, Viktor Öwall |
WCNC | 2 |
| 2013 | VLSI Implementation of a Soft-Output Signal Detector for Multimode Adaptive Multiple-Input Multiple-Output SystemsabstractThis paper presents a multimode soft-output multiple-input multiple-output (MIMO) signal detector that is efficient in hardware cost and energy consumption. The detector is capable of dealing with spatial-multiplexing (SM), space-division-multiple-access (SDMA), and spatial-diversity (SD) signals of 4 × 4 antenna and 64-QAM modulation. Implementation-friendly algorithms, which reuse most of the mathematical operations in these three MIMO modes, are proposed to provide accurate soft detection information, i.e., log-likelihood ratio, with much reduced complexity. A unified reconfigurable VLSI architecture has been developed to eliminate the implementation of multiple detector modules. In addition, several block level technologies, such as parallel metric update and fast bit-flipping, are adopted to enable a more efficient design. To evaluate the proposed techniques, we implemented the triple-mode MIMO detector in a 65-nm CMOS technology. The core area is 0.25 mm2with 83.7 K gates. The maximum detecting throughput is 1 Gb/s at 167-MHz clock frequency and 1.2-V supply, which archives the data rate envisioned by the emerging long-term evolution advanced standard. Under frequency-selective channels, the detector consumes 59.3-, 10.5-, and 169.6-pJ energy per bit detection in SM, SD, and SDMA modes, respectively. Liang Liu 0002, Johan Löfgren, Peter Nilsson 0001, Viktor Öwall |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | A unified multi-mode MIMO detector with soft-outputabstractThis paper presents an area/energy efficient soft-output MIMO detector that supports the detection of spatial-multiplexing (SM), spatial-diversity (SD), and space-division-multiple-access (SDMA) signals. The developed near-optimal detection algorithms for these three modes share most of the mathematical operations to enable extensive hardware reuse. A unified VLSI architecture is accordingly designed to be reconfigured to different modes. The detector is implemented using a 65-nm CMOS technology with 0.25 mm2core area, representing a 70% hardware-resource saving to state-of-the-art detectors. Operating at 167-MHz clock frequency with 1.2-V supply, the detector achieves 1 Gb/s peak throughput and only consumes 59.3 pJ/b energy. Liang Liu 0002, Johan Löfgren, Peter Nilsson 0001 |
ISCAS | 1 |
| 2012 | Mapping channel estimation and MIMO detection in LTE-advanced on a reconfigurable cell arrayabstractThis paper presents a flexible architecture suitable for performing both channel estimation and signal detection in a MIMO-OFDM downlink. Extensive hardware sharing between two tasks is achieved by algorithm and architecture co-design, where robust MMSE sliding window channel estimation and MMSE-based signal detection with symbol perturbation scheme are adopted. The proposed architecture is based on a coarse-grained reconfigurable cell array with fast context switching capabilities. High flexibility is provided by the architecture, which allows task-level resource sharing and dynamic adoption of different algorithms onto the same platform. Simulation and analysis results have confirmed the efficiency of the proposed design solution, where more than 75% hardware resources are reused between the adopted algorithms. Chenxin Zhang, Liang Liu 0002, Viktor Öwall |
ISCAS | 2 |
| 2011 | Detecting multi-mode MIMO signals: Algorithm and architecture designabstractThis paper presents an efficient and reconfigurable MIMO detector design solution targeting on the emerging LTE-A downlink. The detector supports signal detection of multiple MIMO modes which are spatial-multiplexing (SM), spatial-diversity (SD), and space-division-multiple-access (SDMA). Cost-reduction is achieved by algorithm and architecture co-design where low-complexity, near-maximum-likelihood (ML) detection algorithms are proposed for these three MIMO modes respectively while keeping in mind that operations could be reused among different MIMO modes. A multi-stage VLSI architecture is accordingly developed that achieves run-time reconfigurability without extra hardware overhead. Simulation and analysis results have confirmed the efficiency of the proposed solution. Liang Liu 0002, Peter Nilsson 0001 |
ISCAS | 1 |
| 2011 | Low complexity soft-output signal detector for spatial-multiplexing MIMO systemabstractThis paper presents a cost-efficient soft-output signal detector design solution targeting on the spatial-multiplexing MIMO system. The detector achieves low hardware cost and near-optimal detection performance based on the modification to the fixed-complexity sphere decoder (FSD) using several implementation-oriented algorithm-level improvements, which are early-pruning with polygon-shaped constraint, symbol-level bit-flipping, and ℓ1-norm approximation. To evaluate the proposed method, we implement the MIMO detector in a 65-nm standard VTCMOS technology. The core area is 0.14 mm2with 69 K equivalent gates, representing a 60% hardware-resource saving to the state-of-the-art in the open literature. The detecting throughput is up to 1.5Gb/s at 250-MHz clock frequency and 1.2-V supply. The normalized energy consumption of 36.4 pJ/b is shown to be the most energy-efficient design compared with other soft-output detectors. Liang Liu 0002, Johan Löfgren, Peter Nilsson 0001 |
PIMRC | 1 |
| 2010 | Joint estimation and compensation for front-end imperfection in MB-OFDM UWB systemsabstractIn this paper, we investigate the analog front-end imperfection in multi-band orthogonal frequency division multiplexing (MB-OFDM) ultra-wideband (UWB) systems and propose a joint estimation and compensation scheme. The estimation of CFO and SFO is presented on frequency domain, which is robust to frequency dependent I/Q imbalance. After that, I/Q imbalance and channel response is jointly estimated with partial compensation to CFO and SFO. Finally, pilot sub-carriers within OFDM symbols are used to track the residual phase distortion. Simulation results show that the proposed joint estimation and compensation scheme achieves 0.3dB SNR advantage in practical MB-OFDM UWB systems comparing to traditional joint estimation algorithm. Jun Zhou 0002, Liang Liu 0002, Fan Ye 0001, Junyan Ren |
ISCAS | 2 |
| 2008 | Design of Highly-Parallel, 2.2Gbps Throughput Signal Detector for MIMO SystemsabstractThis paper presents a field-programmable gate array (FPGA) implementation of a new multiple-input multiple-output (MIMO) signal detection algorithm applicable to ultra-high throughput MIMO communication systems. The algorithm simplifies the computation significantly compared to traditional K-Best algorithm, and with negligible bit error ratio (BER) degradation. A highly-parallel structure is implemented on the Xilinx Virtex-4 (XC4VLX200) platform, which achieves 2.2 Gbps detection throughput and is about four times over previous implementation. Moreover, a pre-processing method is realized to reduce the number of multipliers inside the detector and shrinks the critical path delay down to 6.79 ns. Together with candidate-sharing-architecture to further save the hardware cost, a high speed, compact signal detector for MIMO systems is demonstrated. Liang Liu 0002, Xiaojing Ma 0001, Fan Ye 0001, Junyan Ren |
ICC | 1 |
| 2007 | Design of Low-Power, 1GS/s Throughput FFT Processor for MIMO-OFDM UWB Communication SystemabstractA new 8PBF structure for 64/128 flexible point FFT processor is proposed. The processor, which is based on 8*8*2 mixed radix algorithm, can deal with multiple inputs more efficiently for MIMO applications. The 8PFB structure efficiently brings the throughput of the processor up to 1GS/s and the chances of register reverse down, reducing the power dissipation remarkably. Meanwhile the modified shift-add algorithm can remove complex multipliers in the FFT processor. Liang Liu 0002, Junyan Ren, Xuejing Wang, Fan Ye 0001 |
ISCAS | 1 |
| 2007 | A Novel Synchronizer for OFDM-based UWB System on New Preamble DesignabstractIn OFDM-based UWB baseband, synchronizer plays a key role to the performance of the whole system. Therefore achieving an area and power efficient architecture at least SNR loss becomes the main challenge of synchronizer design. To meet 528Msamples/s throughput, the parallel architecture is used in current approaches. But it leads to higher area and power consumption. In this paper, we propose an efficient synchronizer for OFDM-based UWB system, consisting of the filter-free symbol synchronizer, frequency synchronizer with improved Moose algorithm and sign-based auto-correlator. According to the results of synthesis using SMIC, 0.13μm CMOS process, the proposed design can achieve the throughput requirement with only 24% gate count and 25% power consumption (in idle state) of the conventional four-parallelism approach. Xuejing Wang, Liang Liu 0002, Fan Ye 0001, Junyan Ren |
PIMRC | 2 |