Min Li 0001

dblp:82/0-1 · also Min (Leon) Li · DBLP profile ↗
← Back
26ranked-venue papers
12as first author
0since 2021 · last 2016
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-authorComputer networks · 8 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 5 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Electronic design automation · 25% Hardware accelerators and domain-specific architectures · 25% Embedded and real-time systems · 19%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
dataflow mapping
0.112010
Exploiting finite precision information to guide data-flow mapping · DAC 2010
Electronic design automation
high-level synthesis
0.112010
Exploiting finite precision information to guide data-flow mapping · DAC 2010
Embedded and real-time systems
software-defined radio
0.112008
How to let instruction set processor beat ASIC for low power wireless baseband implementation: a system level approach · DAC 2008
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.012010
Exploiting finite precision information to guide data-flow mapping · DAC 2010

Methods — techniques the papers use, named apart from their topics

fixed-point refinement · 0.1finite precision propagation · 0.1system-level design · 0.1parallel ISP architecture · 0.1
YearPublicationVenuePosition
2016 Data Flow Transformation for Energy-Efficient Implementation of Givens Rotation-Based QRD
abstract
QR decomposition (QRD), a matrix decomposition algorithm widely used in embedded application domain, can be realized in a large number of valid processing sequences that differ significantly in the number of memory accesses and computations, and hence the overall implementation energy. With modern low-power embedded processors evolving toward register files with wide memory interfaces and vector functional units (FUs), data flow in these algorithms needs to be carefully devised to efficiently utilize the costly wide memory accesses and the vector FUs. In this article, we present an energy-efficient data flow transformation strategy for the Givens rotation--based QRD.
Namita Sharma 0001, Preeti Ranjan Panda, Francky Catthoor, Min Li 0001, Prashant Agrawal
ACM Trans. Embed. Comput. Syst.4
2015 <30 mW rectangular-to-polar conversion processor in 802.11ad polar transmitter
abstract
This paper presents an energy-efficient digital signal processor (DSP) for rectangular-to-polar conversion in 802.11ad polar transmitter working on 60 GHz band. Firstly, system simulations with a complete transmission chain are conducted with regard to error vector magnitude and output spectrum, which allows to systematically optimize the design requirements on the DSP block. Secondly, algorithm and architecture co-optimization on the DSP block is explored to minimize the power consumption. Finally, the proposed DSP is synthesized using 28 nm CMOS technology, which provides a throughput of 7.04 Giga samples per second with a power consumption of 28 mW, and area of 0.01 mm2.
Chunshu Li, André Bourdoux, Marian Verhelst, Yanxiang Huang, Min Li 0001, Liesbet Van der Perre, Sofie Pollin
ICASSP5
2014 Energy efficient data flow transformation for Givens Rotation based QR Decomposition
abstract
QR Decomposition (QRD) is a typical matrix decomposition algorithm that shares many common features with other algorithms such as LU and Cholesky decomposition. The principle can be realized in a large number of valid processing sequences that differ significantly in the number of memory accesses and computations, and hence, the overall implementation energy. With modern low power embedded processors evolving towards register files with wide memory interfaces and vector functional units (FUs), the data flow in matrix decomposition algorithms needs to be carefully devised to achieve energy efficient implementation. In this paper, we present an efficient data flow transformation strategy for the Givens Rotation based QRD that optimizes data memory accesses. We also explore different possible implementations for QRD of multiple matrices using the SIMD feature of the processor. With the proposed data flow transformation, a reduction of up to 36% is achieved in the overall energy over conventional QRD sequences.
Namita Sharma 0001, Preeti Ranjan Panda, Min Li 0001, Prashant Agrawal, Francky Catthoor
DATE3
2014 Towards approaching near-optimal MIMO detection performance ONAC-programmable baseband processor
abstract
Lattice Reduction aided softoutput MIMO detectors have been demonstrated to offer a promising gain. However, computing Log-Likelihood ratios (LLR) for near-optimal MIMO detection, still poses a significant challenge for practical implementations. In this work, we present counter-ML bit-flipping algorithm for LLR generation. The proposed LLR generation algorithm has been designed to take advantage of the previously reported list generation algorithm, Multi-Tree Selective Spanning (MTSS), by maximizing the reuse of computations. Afterwards, a C-programmable MIMO detector architecture providing both data level parallelism (DLP) and instruction level parallelism (ILP), is designed for implementation. The proposed solution supports multiple MIMO detection modes, with both hard and softoutput. Performance of the proposed solution can be tuned ranging from SIC to near-ML to near-MAP, by adjusting a single parameter. In case of 4 × 4 QAM-64, it achieves peak-throughputs of 2.43Gbps and 629Mbps in case of hard and softoutput MIMO detection, with only 66.37mW and 76.14mW respective power consumption.
Ubaid Ahmad, Min Li 0001, Amir Amin, Meng Li 0012, Liesbet Van der Perre, Rudy Lauwereins, Sofie Pollin
ICASSP2
2014 Efficient duty-cycle mismatch compensation in digital transmitter
abstract
This paper presents an efficient mitigation approach for duty cycle mismatch of in-phase and quadrature upconversion signals in digital transmitters. This approach is supported by a mathematical analysis of the baseband equivalent impact of duty cycle mismatch. An efficient digital pre-distortion method is proposed to eliminate the distortion impact. Simulation results show that, for both 64-QAM and 256-QAM modulation schemes, the error-vector-magnitude can be improved from -25.1dB to less than -55dB, which leaves substantial design margin for other non-idealities distorting the transmitted signal.
Chunshu Li, Min Li 0001, Mark Ingels, Marian Verhelst, Xiaoqiang Zhang 0008, Joris Van Driessche, André Bourdoux, Liesbet Van der Perre, Sofie Pollin
ICASSP2
2013 Processor based 20Mhz 4×4 Cat-5 LTE MIMO receiver with advanced detectors
abstract
The Category-5 (Cat-5) UE defined by LTE, as the most demanding category, requires processing 20Mhz bandwidth and 4×4 MIMO transmissions. Very little progress has been reported for its feasibility on programmable processors. In fact, most related work focus on lower categories with much less throughput. Since MIMO signal processing complexity increases non-linearly even with the simplest linear MIMO detectors, 4×4 MIMO transmissions combined with 20Mhz bandwidth is much more challenging when compared to lower UE categories. Our work explores the feasibility of software defined baseband for the most demanding UE category. On a customized SDR baseband processor, we have recently accomplished a software defined downlink inner receiver for Cat-5 LTE UE. The implemented inner receiver includes fully fledged synchronization and data detection functionalities, including coarse CFO estimation/compensation, I/Q imbalance estimation/compensation, OFDMA demodulation, channel estimation, fine SCO/CFO estimation/compensation, MIMO channel processing, MIMO data detection and LLR generation. Both linear MIMO detectors and more advanced MIMO detectors have been studied. To the best of our knowledge, this is the first work experimenting practical Cat-5 LTE receivers on baseband processors.
Min Li 0001, Amir Amin, Rodolfo Torrea Duran, Ubaid Ahmad, Raf Appeltans, Antoine Dejonghe 0001, Liesbet Van der Perre
ICASSP1
2013 Adaptive filter based low complexity digital intensive harmonic rejection for SDR receiver
abstract
Harmonic rejection mixing is indispensable in software defined radio receivers employing switched mixers. Current analog multi-path mixing solution suffers from phase and gain mismatches along the paths and as a consequence cannot provide sufficient harmonic rejection. In this paper, we present a low complexity flexible digital intensive harmonic rejection architecture and show how it can be used to enhance the rejection of any single harmonic interference by adaptively combining the different mixing paths. Simulation results show that the proposed method can reject any single interferer adaptively by over 80 dB, which is sufficient for practical applications.
Chunshu Li, Min Li 0001, Marian Verhelst, Sofie Pollin, André Bourdoux, Liesbet Van der Perre
ICASSP2
2013 A unified receiver signal processing architecture for all modes of the DTMB broadcasting system
abstract
The Chinese Digital Television Terrestrial Broadcasting System has a complex PHY layer definition with many different modes including two different block transmission schemes (OFDM and SC) and three different known symbol padding cyclic extensions, some of which with phase rotation between blocks that break the cyclicity. The block sizes with or without cyclic extension are “non power of two” numbers. This plurality of modes and the unusual block sizes make the design of a signal processing architecture very difficult. In addition, the known symbol padding extensions are intended for channel estimation but have poor auto-correlation properties; hence the channel estimation in long multipath channels is degraded and not suitable for high order constellations. We have designed a novel unified receiver architecture supporting all modes of this broadcasting system, capable to start from a poor initial channel estimation. We describe in detail this architecture and provide simulation results supporting our system choices.
André Bourdoux, Min Li 0001, Hans Cappelle, Amir Amin, Raf Appeltans, Andy Folens, Antoine Dejonghe 0001
PIMRC2
2013 A computationally efficient soft-output Lattice Reduction-aided Selective Spanning Sphere Decoder for wireless MIMO systems
abstract
In recent years, the algorithmic optimizations and implementations of near-optimal Multiple-Input Multiple-Output (MIMO) detectors have been an area of active research. Lattice Reduction (LR) has shown to be a promising technique to improve the performance of linear MIMO detectors. However, LR-aided linear hard-output MIMO detection is still far from optimal. Practical systems use soft-output information to exploit gains from coded systems in order to yield near-optimal performance. In this paper, the LR-aided Selective Spanning Sphere Detection algorithm is proposed as a reduced-complexity candidate list generation method for soft-output MIMO detection, specifically optimized for practical MIMO-OFDM systems. This algorithm uses efficient and scalable heuristics based on simple processor-friendly operations that significantly contribute to lowering the computational complexity of the MIMO detection problem. Results from Monte Carlo simulations reveal that LR-aided SSSD is a promising algorithm that is capable of providing near-optimal performance whilst being especially computationally efficient, in comparison to other algorithms.
Hoang Duy Nguyen, Ubaid Ahmad, Min Li 0001, Liesbet Van der Perre, Rudy Lauwereins, Sofie Pollin
PIMRC3
2012 Exploiting frequency correlation in LTE to reduce HARQ memory
abstract
Hybrid ARQ (HARQ) combines Automatic Repeat Request (ARQ) and Forward Error Correction (FEC) to exploit information from erroneous packets after retransmissions. Due to its superior reliability, HARQ became a crucial component of several important 3G and 4G systems. Although the performance advantage is very attractive, implementing HARQ is challenging with emerging high throughput communication systems such as LTE and LTE-A. Specifically, a large amount of data needs to be stored whenever there is a retransmission. Knowing that LTE/LTE-A systems would provide Gbps or hundreds of Mbps, a straightforward implementation would require around 14 Mbit memory as HARQ buffer. With cost and power limited wireless terminals, the available memory size to support HARQ is a strict and challenging constraint. Hence, it is essential to find efficient techniques to minimize the memory footprint for storing erroneous packets with marginal or, preferably, no degradation. In the context of LTE systems, we propose a novel method to reduce the memory footprint for HARQ systems. This method has been evaluated with a fully fledged practical LTE simulation chain with MIMO transmissions. Compared to state-of-the-art solutions, we show a promising memory compression factor that is close to 5, whereas communication performance degradation is marginal.
Rodolfo Torrea Duran, Min Li 0001, Claude Desset, Sofie Pollin, Liesbet Van der Perre
GLOBECOM2
2012 Algorithm-Architecture Co-Optimization of Area-Efficient SDR Baseband for Highly Diversified Digital TV Standards
abstract
The rapidly evolving and diversifying wireless landscape demands highly flexible wireless chipsets. Due to the ultimate programmability, SDR solutions are becoming more and more attractive. However, the programmability overhead is still a concern for the silicon area cost of SDR solutions. In this work, we prove that, with algorithm and architecture co- design, SDR solutions can be very competitive even when compared to highly optimized ASICs. Specifically, we show a baseband processor design that can support ISDB-T, DVB-T and ATSC, but the area cost is still comparable to the combination of ASICs which handle the three terrestrial digital TV standards respectively.
Kiyotaka Kobayashi, Hidekuni Yomo, Min Li 0001, Raf Appeltans, Hans Cappelle, Amir Amin, Aïssa Couvreur, Matthias Hartmann, André Bourdoux, Praveen Raghavan, Antoine Dejonghe 0001, Liesbet Van der Perre
VTC Spring3
2011 Scalable Block-Based Parallel Lattice Reduction Algorithm for an SDR Baseband Processor
abstract
Lattice Reduction (LR) is a promising technique to improve the performance of linear MIMO detectors. In this paper the Scalable Block-based Parallel LR algorithm (SBP-LR) is proposed and optimized for parallel programmable baseband architectures offering ILP and DLP features. In our algorithm, architecture-friendliness is explicitly introduced from the very beginning of the algorithm/architecture co-design flow. In this context, abundant vector-parallelism is enabled with highly-regular and deterministic data-flow. Hence, SBP-LR can be easily parallelized and efficiently mapped on Software Defined Radio (SDR) baseband architectures. The proposed algorithm has been implemented on ADRES and is evaluated in the context of 3GPP LTE. Most of the previously reported algorithms are implemented for ASIC or FPGA. However, to the best of author's knowledge, this is the first reported LR algorithm explicitly optimized for a Coarse Grain Reconfigurable Array (CGRA) processor like ADRES.
Ubaid Ahmad, Amir Amin, Min Li 0001, Sofie Pollin, Liesbet Van der Perre, Francky Catthoor
ICC3
2011 Overview of a Software Defined Downlink Inner Receiver for Category-E LTE-Advanced UE
abstract
With the soaring development cost of deep sub-micron silicon and the fast-growing diversity in wireless communications, software defined baseband becomes more and more important for handheld devices. However, most software defined receivers reported in previous literatures are still far away from fulfilling the requirement of emerging wireless standards such as the LTE-Advanced. The category-E User Equipment (UE) defined in LTE-Advanced, as the most demanding category for handheld devices, requires processing 2 concurrent data streams at around 300Mbps aggregated throughput. It is not clear whether SDR baseband processors can tackle this challenge. In our work, we explore the feasibility for software defined baseband for LTE-Advanced. With a highly customized SDR baseband processor, we have recently accomplished a software defined downlink inner receiver for Category-E LTE-Advanced UE. The implemented inner receiver includes fully fledged synchronization and data detection functionalities, including coarse CFO estimation/compensation, I/Q imbalance estimation/compensation, OFDMA demodulation, channel estimation, fine SCO/CFO estimation/compensation, MIMO channel processing, MIMO data detection and LLR generation. This paper is intended to bring an overview for the work and emphasizes key aspects that enable the feasibility of the work. In this paper, we will introduce the algorithm and processor architecture co-design flow, overall receiver functionalities, important optimizations and implementation results.
Min Li 0001, Raf Appeltans, Amir Amin, Rodolfo Torrea Duran, Hans Cappelle, Matthias Hartmann, Hidekuni Yomo, Kiyotaka Kobayashi, Antoine Dejonghe 0001, Liesbet Van der Perre
ICC1
2010 Exploiting finite precision information to guide data-flow mapping
abstract
Advanced handheld applications are demanding for implementations of higher energy efficiency and higher performance. In typical implementations, the finite precision information is only known after fixed-point refinement, once the data-flow has been frozen. Instead, in this paper we suggest the propagation of finite precision information to drive data-flow transformations in order to achieve a higher mapping efficiency. Then, provided a flexible architecture with low run-time switching overhead, the data-flow under execution can opportunistically be tuned to provide the instantaneous computational accuracy required by the application. Thereby, the average number of operations and the precision of those is minimized. This principle is demonstrated with the implementation of the 128-point FFT present in a WLAN receiver. Compared to a conventional implementation, a reduction of 49% to 65% of the number of cycles can be achieved depending on conditions external to the receiver.
David Novo, Min Li 0001, Robert Fasthuber, Praveen Raghavan, Francky Catthoor
DAC2
2009 Algorithm-architecture co-design of soft-output ML MIMO detector for parallel application specific instruction set processors
abstract
Emerging SDR baseband platforms are usually based on multiple DLP+ILP processors with massive parallelism. Although these platforms would theoretically enable advanced SDR signal processing, existing work implemented basic systems and simple algorithms. Importantly, MIMO is not fully supported in most implementations. Implemented MIMO but with a simple linear detector. Our work explores the feasibility for SDR implementations of soft-output ML MIMO detectors, which brings 6-12 dB SNR gains when compared to popular linear detectors. Although soft-output ML MIMO detectors are considered to be challenging even for ASICs, we combine architecture-friendly algorithms, application specific instructions, code transformations and ILP/DLP explorations to make SDR implementations feasible. In our work, a 2times4 ADRES based ASIP with 16-way SIMD can deliver 193 Mbps for 2times2 64 QAM, and 368 Mbps for 2times2 16 QAM transmissions. To the best of our knowledge, this is the first work exploring SDR based soft-output ML MIMO detectors.
Min Li 0001, Robert Fasthuber, David Novo, Bruno Bougard, Liesbet Van der Perre, Francky Catthoor
DATE1
2009 Finite precision processing in wireless applications
abstract
Complex signal processing algorithms are often specified in floating point precision. Thus, a type conversion is needed when the targeted platform requires fixed-point precision. In this work we proposed a new method to evaluate the final impact of finite precision processing in wireless applications. The latter combines analytical analysis with simulations. This extends previous work including the effect of the decision-making errors resulting from quantization. Thereby efficient dimensioning of the minimum bit-widths that satisfy a given accuracy constraint can be deployed. The method is validated with two representative case studies, namely an OFDM inner receiver and a Near-ML MIMO (Multiple Inputs, Multiple Outputs) detector.
David Novo, Min Li 0001, Bruno Bougard, Liesbet Van der Perre, Francky Catthoor
DATE2
2009 A System Level Algorithmic Approach toward Energy-Aware SDR Baseband Implementations
abstract
Wireless communication standards are continuously evolving and getting more diverse.This requires a wide variety of baseband implementations within a short time-to-market. Besides, deep sub-micron technology significantly increases the design complexity and associated cost. These yield a growing need for reconfigurable/programmable baseband solutions. Implementing the whole base band functionality on programmable architectures, as foreseen in the tier-2 SDR, will become a must. However, the energy efficiency of SDR baseband platforms is unavoidably worse than the ASIC counterparts. This brings a challenging gap to bridge, which is even broadening further in emerging high rate standards. With a holistic view, we advocate a system level algorithmic approach to bridge this gap. Specifically, we propose to leverage the advantages (programmability) of SDR platforms to compensate for its disadvantages (energy efficiency). Highly flexible baseband algorithms are designed to exploit the abundant dynamics in the environment and the user requirements.In this way, the baseband can utilize the dynamics and substantially reduce the average energy consumption. In this paper, we present a design methodology and principles, illustrated with 3 representative case studies in HSDPA, WiMAX, and 3GPP LTE.
Min Li 0001, David Novo, Bruno Bougard, Claude Desset, Antoine Dejonghe 0001, Liesbet Van der Perre, Francky Catthoor
ICC1
2008 How to let instruction set processor beat ASIC for low power wireless baseband implementation: a system level approach
abstract
Nowadays, mobile devices are integrating an increasing variety of wireless communication standards, and each standard demands a multitude of modes. This tremendous diversity, combined with the increasing development-cost of deep-submicron silicon, desires highly flexible baseband implementations. The tier-2 SDR (Software Defined Radio) paradigm, where the entire baseband runs on programmable architectures, is very attractive to obtain the desired flexibility. Parallel ISP (Instruction Set Processor) based SDR baseband platforms have attracted extensive interest in recent years. However, such implementations typically come with a much lower energy-efficiency than traditional implementations as ASICs (Application Specific Integrated Circuits). This energy efficiency gap remains to be bridged in order to make SDR more pervasive. Importantly, the gap is becoming even more and more challenging in emerging high rate communication standards, such as 3GPP LTE, Mobile WiMAX and 802.11n. This situation demands disruptive innovations.
Min Li 0001, Bruno Bougard, David Novo, Liesbet Van der Perre, Francky Catthoor
DAC1
2008 Optimizing Near-ML MIMO Detector for SDR Baseband on Parallel Programmable Architectures
abstract
ML and near-ML MIMO detectors have attracted a lot of interest in recent years. However, almost all the reported implementations are delivered in ASICs or FPGAs. Our contribution is optimizing the near-ML MIMO detector for parallel programmable architectures, such as those with ILP and DLP features. In the proposed SSFE (selective spanning with fast enumeration), architecture-friendliness is explicitly introduced from the very beginning of the design flow. Importantly, high level algorithmic transformations make the dataflow pattern and structure fit architecture-characteristics very well. We enable abundant vector-parallelism with highly regular and deterministic dataflow in the SSFE; memory rearrangements, shuffling and non-predictable dynamism are all elaborately excluded. Hence, the SSFE can be easily parallelized and efficiently mapped onto ILP and DLP architectures. Furthermore, to fine-tune the SSFE on parallel architectures, extensive pre-compiler transformations are applied with the help of the application-level information. These optimize not only computation-operations but also address-generations and memory-accesses. Experiments show that the SSFE brings very efficient resource-utilizations on real-life VLIW architectures. Specifically, with the SSFE the percentage of NOPs instructions on VLIW is below 1%, even better than that achieved by the software-pipelined FFT. To the best of our knowledge, this is the first reported work about comprehensive optimizations of near-ML MIMO detectors for parallel programmable architectures.
Min Li 0001, Bruno Bougard, Weiyu Xu, David Novo, Liesbet Van der Perre, Francky Catthoor
DATE1
2008 Generic Multi-Phase Software-Pipelined Partial-FFT on Instruction-Level-Parallel Architectures and SDR Baseband Applications
abstract
The PFFT (Partial FFT) is an extended FFT where only part of input or output bins are used. By pruning the useless dataflow, the PFFT can potentially achieve a significant speedup in many important applications. Although theoretical aspects of the PFFT have been thoroughly studied in past three decades, efficient implementations were rarely reported. The most important obstacle is the highly irregular dataflow and the associated control flow. In addition, a size-N PFFT has 2Ndataflow possibilities, so that delivering both flexibility and efficiency in the same implementation is very challenging. This paper presents a generic scheme to map the highly irregular dataflow of arbitrary PFFT onto ILP architectures with highly efficient SWP (Software-Pipelining). Constraints and opportunities of algorithms and architecture are carefully analyzed and exploited. We introduce a multi-phase partitioning, bringing heterogeneous control structures and heterogeneous software pipelining schemes to minimize control overheads and to maximize the efficiency of SWP. The proposal has been tested with 10 representative benchmarks extracted from baseband applications. In experiments cycle-counts, instructions, NOPs, LID/LIP access/miss/hit are thoroughly analyzed. Comparing to full FFTs with efficient SWP, our work reduces 20.5% - 87.5% cycle-counts, 11.2% - 86.5% instructions, 16.1% - 79.4% LID cache accesses and 19.5% - 87.1% LIP cache accesses. To the best of our knowledge, this is the first reported work about the generic software-pipelined PFFT on ILP architectures.
Min Li 0001, David Novo, Bruno Bougard, Liesbet Van der Perre, Francky Catthoor
DATE1
2008 Adaptive SSFE Near-ML MIMO Detector with Dynamic Search Range and 80-103Mbps Flexible Implementation
abstract
In this paper, we will present a near-ML (maximum likelihood) MIMO (multiple input multiple output) detector explicitly optimized for parallel programmable baseband architectures, such as DSPs (digital signal processors) with VLIW (very long instruction word), SIMD (single instruction multiple data) or vector processing features. First, we propose the SSFE (selective spanning with fast enumeration) algorithm as an architecture friendly near-ML MIMO detector. The SSFE has a distributed and greedy algorithmic structure that brings a completely deterministic and regular dataflow. This enables efficient parallelization on programmable architectures. More importantly, in order to exploit the abundant flexibility enabled by programmable architectures, we propose an efficient online algorithm to adaptively adjust the search range of the SSFE according to the numerical properties of MIMO channel matrixes. Such adaptiveness brings significant throughput improvements at negligible performance degradations. Specifically, on VLIW DSP TI TMS320C6416, such a dynamic adaptation brings 2.62 times to 28.6 times improvements (comparing to the static SSFE) for 1/2 turbo-coded 4 times 4 64 QAM transmissions over 3GPP suburban macro channels, delivering 80 - 103 Mbps average throughput.
Min Li 0001, Bruno Bougard, David Novo, Wim Van Thillo, Liesbet Van der Perre, Francky Catthoor
GLOBECOM1
2008 Bridging the energy gap in size, weight and power constrained software defined radio: Agile baseband processing as a key enabler
abstract
The diversity and evolution of wireless communication standards are fast-pacing. This requires a wide-variety of baseband implementations within short time-to-market. Besides, always deeper submicron technology significantly increase design cost. This yields an increasing need for using reconfigurable or programmable solutions for an always larger part of wireless modems. Mapping the whole baseband functionality on a programmable architecture, as foreseen in tier-2 SDR, will become a must in future implementation. In handhelds where the multi-mode trend adds extra needs for programmability, the energy efficiency of SDR baseband is however a major concern. New processor architectures with major improvements on energy efficiency (GOPS/mW) are emerging but are still not sufficient to catch the continuously increasing complexity of wireless physical layers within the shrinking energy budget. To enable SDR in size, weight and power constrained devices, innovation is also needed at the software side. Specifically, a thorough architecture-aware algorithm implementation methodology is needed for the baseband signal processing functions, which account for most of the SDR computational complexity. We present the premise of such a methodology and illustrate its effectiveness on the design of key kernels from present and future wireless baseband systems.
Bruno Bougard, Min Li 0001, David Novo, Liesbet Van der Perre, Francky Catthoor
ICASSP2
2008 Selective Spanning with Fast Enumeration: A Near Maximum-Likelihood MIMO Detector Designed for Parallel Programmable Baseband Architectures
abstract
ML and near-ML MIMO detectors have attracted a lot of interest in recent years. However, almost all of the reported implementations are delivered in ASIC or FPGA. Our contribution is to co-optimize the near-ML MIMO detector algorithm and implementation for parallel programmable base-band architectures, such as DSPs with VLIW, SIMD or vector processing features. Although for hardware the architecture can be tuned to fit algorithms, for programmable platforms the algorithm must be elaborately designed to fit the given architecture, so that efficient resource-utilizations can be achieved. By thoroughly analyzing and exploiting the interaction between algorithms and architectures, we propose the SSFE (selective spanning with fast enumeration) as an architecture-friendly near-ML MIMO detector. The SSFE has a distributed and greedy algorithmic structure that brings a completely deterministic and regular dataflow. The SSFE has been evaluated for coded OFDM transmissions over 802.11n channels and 3GPP channels. Under the same performance constraints, the complexity of the SSFE is significantly lower than the K-Best, the most popular detector implemented in hardware. More importantly, SSFE can be easily parallelized and efficiently mapped on programmable baseband architectures. With TI TMS320C6416, the SSFE delivers 37.4 - 125.3 Mbps throughput for 4x4 64 QAM transmissions. To the best of our knowledge, this is the first reported near-ML MIMO detector explicitly designed for parallel programmable architectures and demonstrated on a real-life platform.
Min Li 0001, Bruno Bougard, Eduardo Lopez-Estraviz, André Bourdoux, David Novo, Liesbet Van der Perre, Francky Catthoor
ICC1
2007 The Quality-Energy Scalable OFDMA Modulation for Low Power Transmitter and VLIW Processor Based Implementation
abstract
The improvement of spectral efficiency comes at the cost of exponential increment of signal processing complexity [1]. Hence, the energy-efficiency of baseband has recently turned out to be the bottleneck when deploying advanced air interfaces such as that in 4G. We advocate the scalable baseband design as a system level technique to aggressively optimize the average computation-load and associated energy-consumption. The key technique is to dynamically scale the baseband processing itself to the user requirement, the environment, the platform, etc. In this paper, we present the scalable design and VLIW processor based implementation of the OFDMA modulator, which is one of the most energy consuming parts of OFDMA and MIMO- OFDMA transmitters (in IEEE 802.16e , 3GPP LTE, etc.). Our work enables the OFDMA modulator to scale the modulation- accuracy and computation-load, so that the OFDMA modulator can dynamically reconfigure and work with minimal number of operations, whereas the required modulation-accuracy is still firmly guaranteed. Our work brings significant reductions in the average computation-load and associated energy-dissipation on real-life programmable platforms. Specifically, when the user is working with 16QAM and 1/2 coding rate (Turbo Coding) in a half-loaded 8-user system, the proposed scheme reduces 84% of the cycle-count and the associated energy-consumption on TI TMS320C6713, whereas the resulted Relative Constellation Error (RCE) is still lOdB better than the required RCE in IEEE 802.16e specifications.
Min Li 0001, Bruno Bougard, Eduardo Lopez-Estraviz, André Bourdoux, Liesbet Van der Perre, Francky Catthoor
GLOBECOM1
2007 Efficient QRD for SRI-RLS Based Equalization on Programmable Architecture
abstract
Advanced adaptive filters have been shown to be very powerful for tracking time varying channels in various wireless communications system. However, the performance comes at the expense of highly resource-demanding implementations, especially in the context of programmable architecture based SDR. We present the optimizations for programmable implementation of QRD based SRI-RLS, which represents a large family of advanced adaptive filters. The key contribution of our work is to comprehensively and systematically remove the redundant operations in the QRD for SRI-RLS. Although most signal processing and scientific libraries implement householder reflection based QRD, we explore different alternatives and then choose given rotations based QRD to enable the aforementioned systematic redundancy removals. Our work significantly reduces the resource requirements (cycle count, energy consumption, etc.) of SRI-RLS implementation. Comparing to the widely accepted QRD implementation in numerical recipes, our work reduces 96.4% cycle-count on a typical baseband DSP (TI TMS320C6713), enabling efficient implementations. The paper shows that removing redundancy is very effective for modern statistical signal processing algorithms that largely rely on cascaded matrix operations.
Min Li 0001, Bruno Bougard, Javed Absar, François Horlin, Liesbet Van der Perre, Francky Catthoor
ICASSP (2)1
2006 Quality-Energy Scalable Chip Level Equalization for HSDPA
abstract
Quality-Energy scalability has been proved to be an effective technique toward the cost reduction for signal processing. However, although it has large potentials for reducing signal processing energy in wireless transceivers, it has not yet been applied to them. In this paper, we elaborate an example and present a feedback control scheme to achieve Quality-Energy scalability for the chip level equalization in High Speed Downlink Packet Access (HSDPA) receivers. In conventional equalizers, the coefficients are updated according to the worst-case assumption that the channel is very dynamic. In our proposed approach, we take into account the speed of the channel variation to adapt the updating frequency. Specifically, a low-complexity closed-loop controller is designed to vary the update-interval while limiting the maximum equalization error. The proposed control scheme takes high order channel statistics implicitly into account without the need for an explicit estimator. Simulation results show that the Quality-Energy scalability for the equalizer results in significant energy reduction with minor quality degradation for channels with large coherence times. Specifically, for a pedestrian channel, 60% of the signal processing operations, and hence energy, can be saved with only 0.25 dBBERloss.
Min Li 0001, Bruno Bougard, François Horlin, Marc Engels, Liesbet Van der Perre, Francky Catthoor
GLOBECOM1