Chathura Jayawardena

dblp:194/6943 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-7846-548XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 ViPer NL-COMM: Making Vector Perturbation Precoding Practical
abstract
Large multiple-input multiple-output (MIMO) systems rely on efficient downlink precoding to enhance data rates and improve connectivity through spatial multiplexing. However, currently employed linear precoding techniques, such as minimum mean square error (MMSE) precoding, significantly limit the achievable spectral efficiency. To meet practical error-rate targets, existing linear methods require an excessively high number of access point (AP) antennas relative to the number of supported users, leading to disproportionate increases in power consumption. Efficient non-linear processing frameworks for uplink MIMO transmissions, such as NL-COMM, have been proposed. However, downlink non-linear precoding methods, such as Vector Perturbation (VP), remain impractical for real-world deployment due to their exponentially increasing computational complexity with the number of supported MIMO streams. This work presents ViPer NL-COMM, the first practical algorithmic and implementation framework for VP-based downlink precoding. ViPer NL-COMM extends the core principles of NL-COMM to the precoding problem, enabling scalable parallelization and real-time computational performance while maintaining the substantial spectral-efficiency benefits of VP precoding. ViPer NL-COMM consists of a novel mathematical framework and an FPGA prototype capable of supporting large MIMO configurations (up to 16×16), high-order modulation (256-QAM), and wide bandwidths (100 MHz) within practical power and resource budgets. System-level evaluations demonstrate that ViPer NL-COMM achieves target error rates using only half the number of transmit antennas required by linear precoding, yielding net power savings on the order of hundreds of Watts at the RF front end. Moreover, ViPer NL-COMM enables supporting more information streams than available AP antennas when the streams are of low-rate, paving the way for enhanced massive-connectivity scenarios in next-generation wireless networks.
Thomas James Thomas, Georgios Ntavazlis Katsaros, Chathura Jayawardena, Konstantinos Nikitopoulos
IEEE Trans. Mob. Comput.3
2025 NL-COMM: Enhanced Video Streaming via Advanced Non-Linear Processing
abstract
With video streaming now accounting for the majority of internet traffic, wireless networks face increasing demands, especially in densely populated areas where limited spectral resources are shared among many devices. While multi-user (MU)-MIMO technology aims to improve spectral efficiency by enabling concurrent transmissions over the same frequency and time resources, traditional linear processing methods fall short of fully utilizing available channel capacity. These methods require a substantial number of antennas and RF chains, to support a much smaller number of MIMO streams, leading to increased power consumption and operational costs, even when the supported streams are of low rate. In this demo, we present NL-COMM, an advanced non-linear MIMO processing framework, demonstrated for the first time with commercial off-the-shelf (COTS) user equipment (UEs) in a fully 3GPP-compliant environment. In addition, also for the first time, the audience will compare and assess the quality of live, over-the-air video transmission from four concurrently transmitting UE devices, alternating between current state-of-the-art MIMO detection algorithms and NL-COMM. Key gains of NL-COMM include improved stream quality, halving the number of required base station antennas without compromising stream quality compared to linear approaches, as well as achieving antenna overloading factors of 400%.
Marcin Filo, Georgios Ntavazlis Katsaros, Chathura Jayawardena, Konstantinos Nikitopoulos
WCNC3
2024 Ultra-Low-Complexity, Non-Linear Processing for MU-MIMO Systems
abstract
Non-linear detection schemes can substantially improve the achievable throughput and connectivity capabilities of uplink MU-MIMO systems that employ linear detection. However, the complexity requirements of existing non-linear soft detectors that provide substantial gains compared to linear ones are at least an order of magnitude more complex, making their adoption challenging. In particular, joint soft information computation involves solving multiple vector minimization problems, each with a complexity that scales exponentially with the number of users. This work introduces a novel ultra-low-complexity, non-linear detection scheme that performs joint Detection and Approximate Reliability Estimation (DARE). For the first time, DARE can substantially improve the achievable throughput (e.g., $40 \%$) with less than $2 \times$ the complexity of linear MMSE, making non-linear processing extremely practical. To enable this, DARE includes a novel procedure to approximate the reliability of the received bits based on the region of the received observable that can efficiently approach the accurately calculated soft detection performance. In addition, we show that DARE can achieve a better throughput than linear detection when using just half the base station antennas, resulting in substantial power savings (e.g., 500 W). Consequently, DARE is a very strong candidate for future power-efficient MU-MIMO developments, even in the case of software-based implementations, as in the case of emerging Open-RAN systems. Furthermore, DARE can achieve the throughput of the state-of-the-art non-linear detectors with complexity requirements that are orders of magnitude lower.
Chathura Jayawardena, Konstantinos Nikitopoulos
PIMRC1
2023 Joint Frequency Offset Compensation and Detection for Multi-User MIMO-OFDM Systems with Frequency Asynchronous User Access
abstract
Multi-user (MU) MIMO-OFDM systems with aggressive spatial multiplexing are promising to enhance throughput and enable massive connectivity. In such systems, residual carrier frequency offsets (CFOs), due to the instability of oscillators and doppler shifts, can substantially degrade the achievable uplink throughput, especially when the number of connected devices becomes large. Existing approaches to mitigate CFOs in MU scenarios, typically involve closed-loop feedback that can result in high signaling overhead and/or significant residual CFO. Being able to compensate for the CFO of the multiple users at the receiver side, can enable the joint transmission of frequency asynchronous users, can obviate the need for high overhead synchronization procedures, can enable the use of cheaper oscillators, and can potentially unlock new user access schemes. However, as we discuss here in detail, compensating for the multiple user CFOs at the receiver is currently impractical due to the corresponding exponential complexity requirements. At the same time, methods that are typically used in single-user MIMO-OFDM systems are inappropriate for MU-MIMO scenarios and, as we show, can result in substantial (e.g., 80%) throughput degradation. To fill this gap, for the first time, we propose a joint CFO compensation and MU detection scheme that can support a large number of spatially transmitted information streams with practical processing complexity and latency requirements. We show that the proposed scheme enables frequency asynchronous user transmission and approaches the performance of perfectly synchronized systems with complexity requirements that are comparable to current MU-MIMO detection schemes that assume perfect synchronization.
Chathura Jayawardena, Konstantinos Nikitopoulos
ICC1
2023 MU-MIMO, Open-RAN PHY with Linear and Massively Parallelizable Non-Linear Processing
abstract
Multi-user multiple-input, multiple-output (MU-MIMO) designs can substantially increase the achievable throughput and connectivity capabilities of wireless systems. However, existing MU-MIMO deployments typically employ linear processing that, despite its practical benefits, can leave capacity and connectivity gains unexploited. On the other hand, traditional non-linear processing solutions (e.g., sphere decoders) promise improved throughput and connectivity capabilities, but can be impractical in terms of processing complexity and latency, and with questionable practical benefits that have not been validated in actual system realizations. At the same time, emerging new Open Radio Access Network (Open-RAN) designs call for physical layer (PHY) processing solutions that are also practical in terms of realization, even when implemented purely on software. This work demonstrates the gains that our highly efficient, massively parallelizable, non-linear processing (MPNL) framework can provide, both in the uplink and downlink, when running in real-time and over-the-air, using our new 5G-New Radio (5G-NR) and Open-RAN compliant, software-based PHY. We showcase that our MPNL framework can provide substantial throughput and connectivity gains, compared to traditional, linear approaches, including increased throughput, the ability to halve the number of base-station antennas without any performance loss compared to linear approaches, as well as the ability to support a much larger number of users than base-station antennas, without the need for any traditional Non-Orthogonal Multiple Access (NOMA) techniques, and with overloading factors that can be up to 300%.
Konstantinos Nikitopoulos, Marcin Filo, Georgios Ntavazlis Katsaros, Chathura Jayawardena, Rahim Tafazolli
MobiCom4
2022 Reduced Complexity Matrix Inversions in Slow Time-Varying MIMO Channels
abstract
The intensifying demand for data rate and connectivity has resulted in multi-user multiple-input multiple-output (MU-MIMO) deployments. MU-MIMO allows multiple data streams to transmit concurrently in the same spectrum band. These mutually interfering streams need to be processed at the base station (BS), leading to substantial computational complexity requirements. Linear MIMO detectors/precoders are popular due to their relatively low complexity. However, matrix inversion is a challenging task in linear detectors/precoders. Especially in experimental platforms, software-based inversions are infeasible for a large number of users. This work presents Matrix Inversion on Channel Approximation (MICA), a novel method that aims to reduce the complexity of matrix inversion by exploiting the characteristics of channel correlation in the time domain. In low-mobility scenarios (user speeds less than 20km/h), MICA can reduce the average complexity and processing latency required for computing the inverse of 64 × 12 channel matrices by about 90% compared to a conventional scheme, while maintaining almost the same error rate performance.
Chathura Jayawardena, Konstantinos Nikitopoulos
ICC2
2020 Evaluating Non-Linear Beamforming in a 3GPP-Compliant Framework Using the SWORD Platform
abstract
It is well documented that the achievable throughput of MIMO systems that employ linear beamforming can significantly degrade when the number of concurrently transmitted information streams approaches the number of base-station antennas. To increase the number of the supported streams, and therefore, to increase the achievable net throughput, non-linear beamforming techniques have been proposed. These beamforming approaches are typically evaluated via simulations or via simplified over-the-air experiments that are sufficient for validating their basic principles, but they neither provide insights about potential practical challenges when trying to adopt such approaches in a standards-compliant framework, nor they provide any indication about the achievable performance when they are part of a standards-compliant protocol stack. In this work, for first time, we evaluate non-linear beamforming in a 3GPP standards-compliant framework, using our recently-proposed SWORD research platform. SWORD is a flexible, open for research, software-driven platform that enables the rapid evaluation of advanced algorithms without extensive hardware optimizations that can prevent promising algorithms from being evaluated in a standards-compliant stack. We show that in an indoor environment, vector perturbation-based non-linear beamforming can provide up to 46% throughput gains compared to linear approaches for 4×4 MIMO systems, while it can still provide gains of nearly 10% even if the number of base-station antennas is doubled.
Marcin Filo, Juan Carlos De Luna Ducoing, Chathura Jayawardena, Christopher Husmann, Rahim Tafazolli, Konstantinos Nikitopoulos
PIMRC3
2020 G-MultiSphere: Generalizing Massively Parallel Detection for Non-Orthogonal Signal Transmissions
abstract
The increasing demand for connectivity and throughput, despite the spectrum limitations, has triggered a paradigm shift towards non-orthogonal signal transmissions. However, the complexity requirements of near-optimal detection methods for such systems becomes impractical, due to the large number of mutually interfering streams and to the rank-deficient or ill-determined nature of the corresponding interference matrix. This work introduces g-MultiSphere; a generic massively parallel and near-optimal sphere-decoding-based approach that, in contrast to prior work, applies to both well- and ill-determined non-orthogonal systems. We show that g-MultiSphere is the first approach that can support large uplink multi-user MIMO systems with numbers of concurrently transmitting users that exceed the number of receive antennas by a factor of two or more, while attaining throughput gains of up to 60% and with reduced complexity requirements in comparison to known approaches. By eliminating the need for sparse signal transmissions for non-orthogonal multiple access (NOMA) schemes, g-MultiSphere can support more users than existing systems with better detection performance and practical complexity requirements. In comparison to state-of-the-art detectors for NOMA schemes and non-orthogonal signal waveforms (e.g., SEFDM) g-MultiSphere can be up to an order of magnitude less complex, and can provide throughput gains of up to 60%.
Chathura Jayawardena, Konstantinos Nikitopoulos
IEEE Trans. Commun.1
2019 Massively Parallel Tree Search for High-Dimensional Sphere Decoders
abstract
The recent paradigm shift towards the transmission of large numbers of mutually interfering information streams, as in the case of aggressive spatial multiplexing, combined with requirements towards very low processing latency despite the frequency plateauing of traditional processors, initiates a need to revisit the fundamental maximum-likelihood (ML) and, consequently, the sphere-decoding (SD) detection problem. This work presents the design and VLSI architecture of MultiSphere; the first method to massively parallelize the tree search of large sphere decoders in a nearly-concurrent manner, without compromising their maximum-likelihood performance, and by keeping the overall processing complexity comparable to that of highly-optimized sequential sphere decoders. For a 10 × 10 MIMO spatially multiplexed system with 16-QAM modulation and 32 processing elements, our MultiSphere architecture can reduce latency by 29× against well-known sequential SDs, approaching the processing latency of linear detection methods, without compromising ML optimality. In MIMO multicarrier systems targeting exact ML decoding, MultiSphere achieves processing latency and hardware efficiency that are orders of magnitude improved compared to approaches employing one SD per subcarrier. In addition, for 16×16 both “hard”and “soft”-output MIMO systems, approximate MultiSphere versions are shown to achieve similar error rate performance with state-of-the art approximate SDs having akin parallelization properties, by using only one tenth of the processing elements, and to achieve up to approximately 9× increased energy efficiency.
Konstantinos Nikitopoulos, Georgios Georgis, Chathura Jayawardena, Daniil Chatzipanagiotis, Rahim Tafazolli
IEEE Trans. Parallel Distributed Syst.3
2016 MultiSphere: Massively Parallel Tree Search for Large Sphere Decoders
abstract
This work introduces MultiSphere, a method to massively parallelize the tree search of large sphere decoders in a nearly-independent manner, without compromising their maximum-likelihood performance, and by keeping the overall processing complexity at the levels of highly-optimized sequential sphere decoders. MultiSphere employs a novel sphere decoder tree partitioning which can adjust to the transmission channel with a small latency overhead. It also utilizes a new method to distribute nodes to parallel sphere decoders and a new tree traversal and enumeration strategy which minimize redundant computations despite the nearly-independent parallel processing of the subtrees. For an 8 × 8 MIMO spatially multiplexed system with 16-QAM modulation and 32 processing elements MultiSphere can achieve a latency reduction of more than an order of magnitude, approaching the processing latency of linear detection methods, while its overall complexity can be even smaller than the complexity of well-known sequential sphere decoders. For 8 × 8 MIMO systems, MultiSphere's sphere decoder tree partitioning method can achieve the processing latency of other partitioning schemes by using half of the processing elements. In addition, it is shown that for a multi-carrier system with 64 subcarriers, when performing sequential detection across subcarriers and using MultiSphere with 8 processing elements to parallelize detection, a smaller processing latency is achieved than when parallelizing the detection process by using a single processing element per subcarrier (64 in total).
Konstantinos Nikitopoulos, Daniil Chatzipanagiotis, Chathura Jayawardena, Rahim Tafazolli
GLOBECOM3