Leilei Song

dblp:30/5437 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
2since 2021 · last 2025
0009-0003-3593-3425ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-authorComputer networks · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
2 papers
Physical-layer communications · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Integrated circuit design · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Physical-layer communications › modulation › multicarrier modulation
FBMC-OQAM
0.912025
Phase Noise Estimation and Pilot Design Suppressing Intrinsic Interference for mmWave FBMC-OQAM Systems · IEEE Trans. Commun. 2025
Physical-layer communications › modulation
multicarrier modulation
0.912025
Phase Noise Estimation and Pilot Design Suppressing Intrinsic Interference for mmWave FBMC-OQAM Systems · IEEE Trans. Commun. 2025
Physical-layer communications › signal processing for communications › statistical signal processing › estimation theory
phase noise estimation
0.912025
Phase Noise Estimation and Pilot Design Suppressing Intrinsic Interference for mmWave FBMC-OQAM Systems · IEEE Trans. Commun. 2025
Physical-layer communications › channel estimation
pilot design
0.912025
Phase Noise Estimation and Pilot Design Suppressing Intrinsic Interference for mmWave FBMC-OQAM Systems · IEEE Trans. Commun. 2025
Physical-layer communications
channel coding
0.012004
On The Performance/Complexity Tradeoff in Block Turbo Decoder Design · IEEE Trans. Commun. 2004
Physical-layer communications › channel coding › decoding algorithms › iterative decoding
turbo decoding
0.012004
On The Performance/Complexity Tradeoff in Block Turbo Decoder Design · IEEE Trans. Commun. 2004
Integrated circuit design
VLSI design
0.012004
On The Performance/Complexity Tradeoff in Block Turbo Decoder Design · IEEE Trans. Commun. 2004

Methods — techniques the papers use, named apart from their topics

frequency-domain estimation · 0.9MSE analysis · 0.9chase search · 0.1
YearPublicationVenuePosition
2025 Phase Noise Estimation and Pilot Design Suppressing Intrinsic Interference for mmWave FBMC-OQAM Systems
abstract
In this paper, we investigate the frequency domain phase noise estimation in millimeter wave filter-bank multicarrier with offset quadrature amplitude (mmWave FBMC-OQAM) systems. In the frequency domain, the primary impairment introduced by phase noise is the common phase error (CPE), whose estimation performance is significantly affected by the inter-carrier interference (ICI) and inter-symbol interference (ISI) experienced by OQAM symbols. To address the above problem, we propose novel frequency domain phase noise estimation and pilot symbol design methods to mitigate the impact of ICI and ISI. Firstly, we quantify the interferences affecting each pilot symbol and design appropriate weights to mitigate the impact of ICI and ISI on the phase noise estimation. Then, through analysis, we observe that the ICI and ISI originate from the interaction between the imaginary intrinsic interferences and the phase noise terms. Accordingly, we propose a pilot symbol design method by eliminating the primary imaginary intrinsic interferences, thereby reducing the ICI and ISI. Furthermore, we analyze the mean squared error (MSE) lower bound and the computational complexity. Simulation results demonstrate that the proposed phase noise estimation method outperforms the traditional method and the proposed pilot structure further improves the accuracy of the phase noise estimation.
Da Chen 0001, Leilei Song, Pei Liu 0004, Wei Peng 0003, Wei Wang 0050
IEEE Trans. Commun.2
2022 A novel transfer learning for recognition of overlapping nano object
Yuexing Han, Qiaochuan Chen, Leilei Song, Chuanbin Lai, Akihiko Konagaya
Neural Comput. Appl.5
2009 An efficient symbol-level combining scheme for MIMO systems with hybrid ARQ
abstract
This paper proposes a new combining scheme for multiple-input multiple-output (MIMO) systems with hybrid automatic-repeat-request (HARQ). The proposed combining scheme is proved to have the optimal decoding performance. Furthermore, the proposed combining scheme is shown to have low memory requirement and reduced complexity compared to other optimal combining schemes. Simulation results under IEEE 802.16e setting with UMTS channel models verify that the proposed combining scheme achieves the optimal decoding performance and performs much better than other suboptimal combining schemes.
Edward W. Jang, John M. Cioffi, Leilei Song
IEEE Trans. Wirel. Commun.4
2007 Concatenation-Assisted Symbol-Level Combining Scheme for MIMO Systems with Hybrid ARQ
abstract
This paper proposes a new receiver scheme for multiple-input multiple-output (MIMO) systems with hybrid automatic-repeat-request (HARQ). The proposed scheme is proved to have the optimal decoding performance in a sense that all the relevant information is fully used for decoding. Furthermore, it is shown that the proposed scheme has reduced complexity compared to the other optimal receiver scheme. Simulation results under IEEE 802.16e setting with UMTS channel model verify that the proposed receiver scheme achieves the same decoding performance with the other optimal receiver scheme and performs better than other suboptimal receiver schemes.
Edward W. Jang, Leilei Song, John M. Cioffi
GLOBECOM3
2004 On The Performance/Complexity Tradeoff in Block Turbo Decoder Design
abstract
In this letter, tradeoffs between very large scale integration implementation complexity and performance of block turbo decoders are explored. We address low-complexity design strategies on choosing the scaling factor of the log extrinsic information and on reducing the number of hard-decision decodings during a Chase search.
Zhipei Chi, Leilei Song, Keshab K. Parhi
IEEE Trans. Commun.2
2000 VLSI design of Reed-Solomon decoder architectures
abstract
This paper presents VLSI implementations of an 8-error correcting (255, 239) Reed-Solomon (RS) decoder architecture for the optical fibre systems. We present the RS decoders using Euclidean and modified Euclidean algorithms which are regular and simple, and naturally suitable for VLSI implementation. We investigate hardware complexity, clock frequency and data processing rate for those RS decoders. The RS decoder based on the modified Euclidean algorithm operates at a clock frequency of 75 MHz and has a data processing rate of 600 Mbits/s in 0.25-/spl mu/m CMOS technology with a supply voltage of 2.5 V.
Hanho Lee, Meng-Lin Yu, Leilei Song
ISCAS3
2000 Hardware/software codesign of finite field datapath for low-energy Reed-Solomon codecs
abstract
Reed-Solomon (RS) coders are used for error-control coding in many applications such as digital audio, digital TV, software radio, CD players, and wireless and satellite communications. Traditionally, RS coders have been implemented using dedicated hardware. This paper considers software-based implementation of RS codecs. A hardware-software codesign approach is used to design the finite field datapath in a domain-specific digital signal processor (DSP) with low-energy RS codecs application in mind. These datapaths are designed to accommodate programmability with respect to the primitive polynomial as well as the field degree m. A novel heterogeneous digit-serial approach is proposed, where the heterogeneity corresponds to the use of different digit sizes in the multiply-accumulate (MAC) and degree reduction (DEGRED) subarrays. The salient feature of this digit-serial approach is that only the digit cells are implemented in hardware and the finite field multiplications are performed digit-serially in software by dynamically scheduling the internal digit-level operations. Efficient scheduling strategies for digit-serial finite field multiplications are presented and applied to the design of low-energy high-performance RS codecs in software. Significant energy and energy-latency reductions can be achieved using the digit-serial datapaths, as compared with the traditional approach where a combined MAC-DEGRED (parallel multiplier) unit is used. It is concluded that for two-error-correcting RS(n, k) codes over finite field GF(2/sup 8/), datapath containing a parallel MAC unit (of digit size eight) and a DEGRED unit with digit size two (or four) leads to RS codecs with the least energy consumption and energy-latency products; with these datapath architectures and appropriate digit-serial scheduling strategies, more than 60% energy reduction and more than one-third energy-latency reduction can be achieved compared with the parallel multiplication datapath-based approach.
Leilei Song, Keshab K. Parhi, Ichiro Kuroda, Takao Nishitani
IEEE Trans. Very Large Scale Integr. Syst.1
1998 Low-energy heterogeneous digit-serial Reed-Solomon codecs
abstract
Reed-Solomon (RS) codecs are used for error control coding in many applications such as digital audio, digital TV, software radio, CD players, and wireless and satellite communications. This paper considers software-based implementation of RS codecs where special instructions are assumed to be used to program finite field multiplication datapaths inside a domain-specific programmable digital-signal processor (DS-PDSP). A heterogeneous digit-serial approach is presented, where the heterogeneity corresponds to the use of different digit-sizes in the multiply-accumulate (MAC for polynomial multiplication) and degree reduction (DEGRED for polynomial module operation) subarrays. The salient feature of this digit-serial approach is that only the digit-cells are implemented in hardware, the finite field multiplications are performed digit-serially in software by dynamically scheduling the internal digit-level operations in RS encoders and decoders. It is concluded that, for 2-error-correcting RS(n,k) codec implementations over finite field GF(2/sup 8/), a parallel MAC unit (of digit-size 8) and a DEGRED unit with digit-size 2 is the best datapath, with respect to least energy consumption and energy-delay products. With this datapath architecture and appropriate digit-serial scheduling strategies, more than 60% energy reduction and more than 1/3 energy delay reduction can be achieved compared with the parallel multiplication datapath based approach.
Leilei Song, Keshab K. Parhi, Ichiro Kuroda, Takao Nishitani
ICASSP1
1998 Efficient semisystolic architectures for finite-field arithmetic
abstract
Finite fields have been used for numerous applications including error-control coding and cryptography. The design of efficient multipliers, dividers, and exponentiators for finite field arithmetic is of great practical concern. In this paper, we explore and classify algorithms for finite field multiplication, squaring, and exponentiation into least significant bit first (LSB-first) scheme and most significant bit first (MSB-first) scheme, and implement these algorithms using semisystolic arrays. For finite field multiplication (for programmable as well as fixed field order) and exponentiation, we conclude that LSB-first algorithms are more efficient as their basic cells have less critical path computation time. Another advantage of LSB-first scheme is its capability of achieving substructure sharing among multiple operations, which could lead to savings in hardware when these arithmetic units are used as building blocks for a large system. For finite field squaring operation, it turns out that the MSB-first algorithm is more efficient as it leads to simpler architectures. Bit-level pipelined semisystolic architectures utilize broadcast signals. As a result, these require much less number of latches and lead to much smaller latency than the corresponding systolic array, with the same cycle time (the computation time in one basic cell). Efficient VLSI implementation of semisystolic multipliers, squarers and exponentiators are designed and compared with existing architectures. A novel architecture for computing AB/sup n/+C using power representation is also presented.
Leilei Song, Keshab K. Parhi
IEEE Trans. Very Large Scale Integr. Syst.2
1997 Low-area dual basis divider over GF(2M)
abstract
This paper presents a low-area finite field divider using dual basis representation. This divider is based on the division algorithm of solving Discrete Wiener-Hopf Equation using Gauss-Jordan elimination method. The hardware complexity of the matrix generation part has been reduced dramatically form O(m/sup 2/) to O(m). When it is used as a building block for a large system, this divider can achieve more savings in hardware by utilizing sub-structure sharing techniques.
Leilei Song, Keshab K. Parhi
ICASSP1
1996 Efficient Finite Field Serial/Parallel Multiplication
abstract
Finite field has received a lot of attention due to its widespread applications in cryptography, coding theory, etc. Design of efficient finite field arithmetic architectures is very important and of great practical concern. In this paper, a new bit-serial/parallel finite field multiplier is presented with standard basis representation. This design is regular and well suited for VLSI implementation. As compared to existing serial/parallel finite field multipliers, it has smaller critical path, lower latency and can be easily pipelined. When it is used as a building block for large systems, it can achieve more savings in hardware in the broadcast structures by utilizing sub-structure sharing technique. This paper also presents two generalized algorithms for finite field serial/parallel multiplication. They can be used to derive efficient bit-parallel, digit-serial or bit-serial multiplication architectures. The optimal primitive polynomials over GF(2/sup m/) (for 2/spl les/m/spl les/9) are provided which will generate structures with minimum hardware complexity and relatively more flexibilities for feasible digit-sizes with respect to the proposed algorithms. Finally a multiplier over GF(2/sup 8/) is given as an example showing how to derive finite field multipliers using the proposed algorithms. This multiplier has less number of transistors, smaller critical path and consumes less power compared to the existing semi-systolic architecture.
Leilei Song, Keshab K. Parhi
ASAP1
1996 Systematic analysis of bounds on power consumption in pipelined and non-pipelined multipliers
abstract
The paper presents a systematic theoretical approach for the analysis of bounds on power consumption in Baugh-Wooley, binary tree and Wallace tree multipliers. This is achieved by first developing state transition diagrams (STDs) for the sub circuits making up the multipliers. The STD is comprised of states and edges, with the edges representing a transition (switching activity) from one state to another in the sub circuit. Then, maximum (minimum) energy values associated with the edges constituting the STDs are used to derive the zipper (lower) bound in both non pipelined and p-bit level pipelined multipliers. It is shown that as p is decreased, the upper bound approaches the lower bound. Moreover, based on the theoretical analysis we conclude that the upper bound in a Baugh-Wooley multiplier has a cubic dependence on the word length, while that in a binary tree multiplier has a quadratic dependence on the word length.
Janardhan H. Satyanarayana, Keshab K. Parhi, Leilei Song, Yun-Nan Chang
ICCD3