EDBT 2026 Demo / reviewers in the wild / expert
Jin Sha 0001
dblp:87/1706-1
· DBLP profile ↗
24ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-0266-3583ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 3 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Toward Robust Radio Frequency Fingerprint Identification via Adaptive Semantic AugmentationabstractRadio frequency fingerprint identification (RFFI) is regarded as one of the most promising techniques for managing and regulating Internet of Things (IoT) devices. This technology analyzes the unique electromagnetic signals emitted by wireless devices to enable precise identification and authentication. Most existing RFFI methods focus on RF signals collected in specific scenarios. However, in real-world applications, signals are often collected at different times or from varying deployment locations, leading to differences between the training and testing distributions. The study of RFFI methods under these conditions remains underexplored. To address this gap, this paper introduces a cross-domain RFFI framework centered on adaptive semantic augmentation (ASA). The framework integrates a computationally efficient multi-resolution spectrogram decomposition strategy with a feature-sensitive multi-scale network. The ASA method enhances RFFI accuracy in cross-domain settings by linearly interpolating between two distinct semantic features to create new semantics for further identification. The proposed approach leverages two-dimensional discrete wavelet transform (2D-DWT) to decompose the raw spectrogram into four sub-bands, followed by a multi-scale network to extract critical semantic features for the ASA method. Simulation results show that the proposed ASA method significantly improves Unmanned Aerial Vehicle (UAV) identification performance, achieving accuracies of 93.05% and 98.90% on two different cross-domain datasets, respectively, outperforming existing data augmentation (DA) methods. Furthermore, generalizability validation demonstrates that the proposed method performs outstandingly across other Internet of Things (IoT) applications. Zhenxin Cai, Yu Wang 0078, Guan Gui 0001, Jin Sha 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | A CPU+FPGA OpenCL Heterogeneous Computing Platform for Multi-Kernel PipelineabstractOver the past decades, Field-Programmable Gate Arrays (FPGAs) have become a choice for heterogeneous computing due to their flexibility, energy efficiency, and processing speed. OpenCL is used in FPGA heterogeneous computing for its high-level abstraction and cross-platform compatibility. Previous works have introduced optimization techniques in OpenCL for FPGAs to leverage FPGA-specific advantages. However, the multi-kernel pipeline technique, which can raise throughput and resource utilization, has not performed well. This article presents a CPU+FPGA heterogeneous platform with a novel execution model to optimize multi-kernel pipeline. Firstly, we extend OpenCL by introducing new APIs and additional functions to represent the execution model. Secondly, a hardware-software co-scheduling scheme is employed to manage execution. Thirdly, we design a holistic development flow and toolkit to facilitate the deployment of algorithms on the platform or the integration of RTL IP cores to the OpenCL environment. We validate the platform using a Range Doppler algorithm. The proposed development flow and integrated toolchain enhance the efficiency of integrating traditional RTL IP cores into the OpenCL environment. Experimental results demonstrate that, with a comparable processing speed (averaging 95%) to traditional RTL implementations, the platform successfully establishes the multi-kernel pipelines. Leveraging the multi-kernel pipeline, the platform achieves a significant improvement in multi-frame processing speed compared to traditional OpenCL. Yuefei Wang, Wendong Mao, Lang Feng 0001, Jin Sha 0001, Zhongfeng Wang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | An Energy-Efficient FPGA Accelerator for Swin TransformerabstractRecently, transformers have shown strong performance in tasks such as computer vision and natural language processing. Notably, Swin Transformer has gained significant attention for its low computational complexity and impressive performance in computer vision tasks, due to its window attention mechanism and hierarchical architecture. However, these features also make hardware deployment more complicated. In this brief, we present an energy-efficient field-programmable gate array (FPGA) accelerator for Swin Transformer to support the hierarchical architecture and execute the window attention. First, we introduce a systolic array with alterable datapath (SAAD) to conduct the window attention. Second, we split the patch merging operation and design a data rearrangement module, which reduces the computing latency induced by the data rearrangement in Swin Transformer. Third, we present a parallelized dual-array dataflow to support different computing operations in Swin Transformer. We implement the accelerator on the Xilinx XCZU19EG platform. The proposed architecture achieves a throughput per digital signal processing (DSP) of 0.630 giga operations per second (GOPS)/DSP, which is$1.94\times $higher than existing works. Yuefei Wang, Wendong Mao, Huihong Shi, Jin Sha 0001, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Efficient UAV Identification Leveraging Multi-Resolution Analysis and Multi-Scale ResNetabstractAs unmanned aerial vehicles (UAVs) become increasingly prevalent in various environments, their detection is essential for ensuring safety and effective management. For non-standard transmitter waveforms, identifying a specific model poses challenges, and most researchers focus on developing deeper and broader architectures to improve performance at the expense of computational burden and speed, which is not suitable for the deployment of UAV radio frequency (RF) fingerprinting identification. In this paper, we propose a novel multi-resolution analysis-based method for UAV identification. We develop a lightweight, multi-scale convolutional network that utilizes various receptive fields to extract unique hardware intrinsic features. We employ the two-dimensional Discrete Wavelet Transform (2D-DWT) to innovatively decompose the low-frequency and the high-frequency components of the RF signal spectrogram into four distinct sub-bands. Simulation results indicate that our proposed methods outperforms other state-of-the-art UAV identification methods in terms of both performance and computational complexity. Zhenxin Cai, Jin Sha 0001, Yu Wang 0078, Guan Gui 0001 |
VTC Spring | 2 |
| 2024 | Toward Intelligent Lightweight and Efficient UAV Identification With RF FingerprintingabstractThe inherent flexibility of small unmanned aerial vehicles (UAVs) enables their deployment across various emerging markets. Unauthenticated UAVs pose a significant threat if they intrude into aviation-sensitive areas. To address this issue, deep learning (DL)-based radio frequency fingerprint identification (RFFI) has been developed as a promising approach for identifying illegal UAVs. However, these commonly used DL-based methods demand high computation and storage requirements, which are not suitable for the deployment of RFFI. In this paper, we propose an efficient and low-complexity RFFI method for UAV identification. Specifically, we design a lightweight backbone network consisting of lightweight multi-scale convolution (LMSC) blocks that can significantly reduce the model size and enhance the feature extraction ability. The simulation results indicate that our proposed UAV RFFI method outperforms other state-of-the-art and popular DL-based RFFI methods in terms of both identification performance and complexity. The identification accuracy surpasses that of all other methods at low signal-to-noise ratios (SNRs) and achieves nearly 100% accuracy at high SNRs. To further enhance model efficiency, we employ data truncation in our experimental simulations, demonstrating that a sample length of 2000 is sufficient to retain high identification performance. Additionally, we incorporate the Mixup regularization strategy, which improves accuracy without increasing the complexity, especially as sample length decreases. Zhenxin Cai, Yu Wang 0078, Guan Gui 0001, Jin Sha 0001 |
IEEE Internet Things J. | 5 |
| 2024 | Radio Frequency Fingerprint Identification Based on Variational Autoencoder for GNSSabstractInterference against global navigation satellite system (GNSS) is threatening its reliability. Radio frequency fingerprint identification (RFFI) emerges as a physical-layer security solution that can effectively identify genuine transmitters. However, external noise in the transmission is not conducive to maintaining the robustness of the RFFI, and the deep neural networks (DNNs) or large datasets for boosting robustness will consume excessive resources. To this end, this letter proposes a lightweight RFFI scheme based on variational autoencoder (VAE) and long short-term memory (LSTM) for real-field GPS signals collected in Nuremberg. The VAE aims to denoise and reconstruct the RFFs, thereby improving the identification accuracy and reducing feature dimensionality. LSTMs can extract the RFF features without any pretransformation and avoid the problem of gradient vanishing or gradient exploding. Numerical results demonstrate that our model can yield an identification accuracy of up to 95.68% on postcorrelation GPS data at low complexity. Jin Sha 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | RF Fingerprinting Identification in Low SNR Scenarios for Automatic Identification SystemabstractThe explosive growth of maritime vessels imposes high demands on the security of automatic identification system (AIS). Radio frequency fingerprinting identification (RFFI) as a physical-layer authentication method offers a new perspective on security solutions for wireless communication systems. However, RFFI for AIS satellite component suffers from poor performance under low signal-to-noise ratio (SNR) scenarios. To this end, this paper proposes a data-driven RFFI method with high performance and resilience for AIS satellite links. The bivariate variational mode decomposition (VMD) enables adaptive signal decomposition for effective denoising. The raw I/Q samples after bivariate VMD are directly taken as input to the RF fingerprinting feature extraction network without any transformation process. The complex-valued neural network used for feature extraction provides a superior representation capability compared to traditional real-valued neural networks. Furthermore, channel attention mechanisms are embedded in the feature extraction network to shed light on the correlation between channels within the signal. The following integrated spatial attention mechanisms are employed to reveal the correlations between signals. Numerical results show that the proposed RFFI method can achieve over 90.37% accuracy on real-world datasets with SNRs greater than 6 dB, and achieve a balance between complexity and performance. Jin Sha 0001 |
IEEE Trans. Wirel. Commun. | 2 |
| 2023 | The Use of SNN for Ultralow-Power RF Fingerprinting Identification With Attention Mechanisms in VDES-SATabstractdata exchange system (VDES) is explored as a navigational safety guarantor for ships worldwide. As the number of jammers and spoofers targeting VDES is on the rise, its security authentication emerges as an urgent issue. Radio frequency fingerprinting identification (RFFI) as a noncryptographic physical-layer security solution can effectively defend against increasingly sophisticated attacks. Nevertheless, the power consumption of frequent RFFI is not negligible for VDES satellite components with energy constraints. To this end, spiking neural networks (SNNs) are first employed to build an ultralow-power RFFI system for VDES, where the optimized spiking neurons can yield high accuracy with very few spikes. Moreover, multiple RF fingerprinting features are fused to highly integrate the information of the VDES signal in different dimensions, and the attention mechanism is utilized to further improve RFFI performance. Numerical results show that our SNN-based RFFI system can yield identification accuracy up to 92.59% at a signal-to-noise ratio (SNR) of 25 dB on VDES data sets and reduces power consumption by 64% compared to artificial neural networks (ANNs) of comparable accuracy on 45-nm CMOS process. Jin Sha 0001 |
IEEE Internet Things J. | 2 |
| 2023 | 1+1 <2: Efficient Automatic Standard Cell Sharing Between Digital VLSI Designs for Area SavingabstractIn the field of digital VLSI design, multimode circuits are the designs where the modes can be switched according to different application scenarios, and are commonly used in communication systems. In a multimode circuit, different modes are usually implemented by different circuits, which can lead to large circuit area consumption. For different modes, sharing their isomorphic circuit regions in the standard cell level can save the area. This goal is similar to that in the subgraph isomorphism problem, which is to check if a given graph is a subgraph of another one. However, subgraph isomorphism needs unacceptable runtime to solve as it is NP-complete. Even worse, finding the largest isomorphic regions of different circuits is a problem harder than subgraph isomorphism. In this article, we propose a novel algorithm for efficiently finding enough isomorphic circuit regions of different digital circuits in polynomial time, and give the theoretical proof of the correctness. The experiments show that the proposed approach can save 20%–25% area on average by sharing the standard cells between 2 and 4 circuits, while keeping the functional correctness. The proposed algorithm also has reasonable runtime and mostly incurs negligible timing overhead. Lang Feng 0001, Jin Sha 0001, Zhongfeng Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Polar-coded forward error correction for MLC NAND flash memory
Haochuan Song, Jen-Chien Fu, Shih-Jia Zeng, Jin Sha 0001, Zaichen Zhang, Xiaohu You 0001, Chuan Zhang 0001 |
Sci. China Inf. Sci. | 4 |
| 2017 | High-Speed Parallel LFSR Architectures Based on Improved State-Space TransformationsabstractLinear feedback shift register (LFSR) has been widely applied in BCH and CRC encoding. In order to increase the system throughput, the parallelization of LFSR is usually needed. Previously, a technique named state-space transformation was presented to reduce the complexity of parallel LFSR architectures. Exhaustive searches are performed to find good transformation matrix candidates. This brief proposes a new technique for construction of the transformation matrix together with a more efficient searching algorithm. The realization results indicate that the proposed architecture outperforms the prior arts, improving the hardware efficiency by around 35% and the corresponding searching algorithm finds the desirable transformation matrix much faster. Jin Sha 0001, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | A high throughput belief propagation decoder architecture for polar codesabstractThe belief propagation (BP) decoding algorithm not only is an alternative to the successive cancelation (SC) decoders of polar codes, but also provides soft outputs that are necessary for joint detection and decoding. The BP decoders with the flooding schedule achieve high throughput with excessive hardware cost especially when the block length is large. The soft-cancelation (SCAN) decoders for polar codes have reduced memory complexity compared to the BP decoders based on the flooding schedule. The simplified SC aided reduced complexity soft-cancelation (S-RCSC) decoders further reduce the computational and memory complexity of the SCAN decoders at the cost of negligible error performance degradation. Both the SCAN and S-RCSC decoders have limited throughput due to their serial decoding schedules. In this paper, we first propose an improved S-RCSC (IS-RCSC) decoding algorithm and then present a high throughput decoder architecture based on our IS-RCSC algorithm. Our IS-RCSC decoding algorithm performs the message passing on a binary tree representation of a polar code. Compared to the S-RCSC decoding algorithm, our IS-RCSC decoding algorithm accelerates the computing of the returned soft messages when certain types of nodes are activated. The corresponding hardware architecture of our IS-RCSC decoder is also proposed. In terms of area efficiency, the hardware implementation results demonstrate that our IS-RCSC decoders are 19% to 43% better than decoders in the literature. Jun Lin 0001, Jin Sha 0001, Li Li 0003, Chenrong Xiong, Zhiyuan Yan 0001, Zhongfeng Wang 0001 |
ISCAS | 2 |
| 2016 | Stage-combined belief propagation decoding of polar codesabstractA novel modification is introduced in this paper for the belief propagation decoder of polar codes, wherein adjacent two processing stages are efficiently combined together to speed up decoding. Corresponding path based belief estimation method is presented in detail. The proposed decoder halves the number of stages of the conventional decoder and thus can significantly reduce the decoding latency and lower message memory requirement. Jin Sha 0001, Jun Lin 0001, Zhongfeng Wang 0001 |
ISCAS | 1 |
| 2014 | Hardware architecture for list successive cancellation polar decoderabstractThis paper aims at designing an efficient hardware architecture for list successive cancellation (SC) polar decoder. Previous literatures have shown that, compared to conventional SC decoder, list SC decoder has the ability to approach the performance of maximum likelihood (ML) decoder. However, the efficient implementation of list SC decoder has not been proposed yet. To tackle this issue, first we propose a sub-optimal version of list SC decoding. Then different selections of list size L are evaluated. By introducing the pre-computation technique, the hardware architecture for a list SC decoder with L = 2 is proposed. Comparison results have shown that, for a rate-½ (1024, 512) polar code, the proposed decoder can achieve near-optimal decoding performance with less hardware cost and latency than the decoder with conventional design approach. We believe that the design approach presented in this paper will facilitate practical applications of list SC polar decoder. Chuan Zhang 0001, Xiaohu You 0001, Jin Sha 0001 |
ISCAS | 3 |
| 2014 | Efficient symbol reliability based decoding for QCNB-LDPC codesabstractAs an extension of binary low-density parity-check (LDPC) codes, non-binary LDPC (NB-LDPC) codes show significantly better performance when the code length is moderate or small. Recently, enhanced iterative hard reliability based (EIHRB) decoding algorithm is proposed to reduce the computation complexity. However, the EIHRB algorithm suffers a lot from significant performance degradation when the column weight is small. In this paper, a symbol reliability based (SRB) decoding algorithm, which also performs well when the column weight is low, is proposed for NB-LDPC decoding to improve the decoding performance. With the same maximum iteration number, around 0.38 dB extra coding gain is achieved. Furthermore, the corresponding efficient decoder architecture is proposed. Comparison results have shown that the proposed SRB algorithm can not only achieve good coding gain, but the cost for hardware implementation is reasonable. Leixin Zhou, Jin Sha 0001, Yun Chen 0001, Chuan Zhang 0001, Zhongfeng Wang 0001 |
ISCAS | 2 |
| 2013 | Memory efficient EMS decoding for non-binary LDPC codesabstractNon-binary low-density parity-check (NB-LDPC) codes are an extension of binary LDPC codes with significantly better performance when the code length is moderate. Previously, forward-backward schemes are used to implement check node processing, which need large amount of memory. In this paper, a novel approach-TCL-EMS is proposed for NB-LDPC decoding. Compared to original EMS decoding algorithm, the memory efficiency is improved and the average number of iterations is reduced significantly. Also, the overall decoder architecture is proposed. Leixin Zhou, Jin Sha 0001, Yun Chen 0001, Zhongfeng Wang 0001 |
ISCAS | 2 |
| 2012 | Memory efficient column-layered decoder design for non-binary LDPC codesabstractLow-density parity-check (LDPC) codes constructed over the Galois field GF(q) (q>;2), which are also called non-binary LDPC codes, are an extension of binary LDPC codes with significantly better performance. In this paper, an efficient column-layered decoding algorithm, which can reduce the message memory as well as the average number of iterations dramatically, is proposed for min-max decoding. In addition, a non-uniform quantization scheme is developed for reducing the word length while achieving similar performances compared to a conventional quantization scheme. Meanwhile, the corresponding decoder architecture is also proposed. Jin Sha 0001, Zhongfeng Wang 0001 |
ISCAS | 2 |
| 2012 | Efficient network for non-binary QC-LDPC decoderabstractThis paper presents approaches to develop efficient network for non-binary quasi-cyclic LDPC (QC-LDPC) decoders. By exploiting the intrinsic shifting and symmetry properties of the check matrices, significant reduction of memory size and routing complexity can be achieved. Two different efficient network architectures for Class-I and Class-II non-binary QC-LDPC decoders have been proposed, respectively. Comparison results have shown that for the code of the 64-ary (1260, 630) rate-0.5 Class-I code, the proposed scheme can save more than 70.6% hardware required by shuffle network than the state-of-the-art designs. The proposed decoder example for the 32-ary (992, 496) rate-0.5 Class-II code can achieve a 93.8% shuffle network reduction compared with the conventional ones. Meanwhile, based on the similarity of Class-I and Class-II codes, similar shuffle network is further developed to incorporate both classes of codes at a very low cost. Chuan Zhang 0001, Jin Sha 0001 |
ISCAS | 2 |
| 2012 | Unified Architecture for Reed-Solomon Decoder Combined With Burst-Error CorrectionabstractReed-Solomon (RS) codes are widely used as forward correction codes (FEC) in digital communication and storage systems. Correcting random errors of RS codes have been extensively studied in both academia and industry. However, for burst-error correction, the research is still quite limited due to its ultra high computation complexity. In this brief, starting from a recent theoretical work, a low-complexity reformulated inversionless burst-error correcting (RiBC) algorithm is developed for practical applications. Then, based on the proposed algorithm, a unified VLSI architecture that is capable of correcting burst errors, as well as random errors and erasures, is firstly presented for multi-mode decoding requirements. This new architecture is denoted as unified hybrid decoding (UHD) architecture. It will be shown that, being the first RS decoder owning enhanced burst-error correcting capability, it can achieve significantly improved error correcting capability than traditional hard-decision decoding (HDD) design. Li Li 0003, Bo Yuan 0001, Zhongfeng Wang 0001, Jin Sha 0001, Hongbing Pan, Weishan Zheng |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2010 | Low power decoder design for QC-LDPC codesabstractThis paper presents a low-power decoder design approach for generic quasi-cyclic low-density parity-check (QC-LDPC) codes based on the layered min-sum decoding algorithm. To reduce the energy consumption, a novel message length-shortening scheme is explored. The check node processing unit (CNU) is accordingly optimized using bit-serial architecture. This low cost design scheme can greatly lower the power consumption while maintaining the necessary throughput required by mobile applications. We further demonstrate the benefits of the proposed techniques by applying the new architecture to the QC-LDPC code in CMMB standard. Jin Sha 0001, Li Li 0003, Zhongfeng Wang 0001 |
ISCAS | 2 |
| 2010 | Layered decoding for non-binary LDPC codesabstractIn this paper, we present a layered decoding algorithm for non-binary LDPC codes. Differing from the flooding message-passing schedule in conventional designs, the proposed scheme updates check node messages in serial. Furthermore, since fully serial decoding will lead to long latency, the locally-parallel globally-serial schedule is adopted. Basically the check nodes can be divided into several groups, i.e. layers. The layers are processed one by one while the check nodes in each layer are processed in parallel. Simulation results show that the proposed algorithm not only brings some improvement in error correcting performance but also gives some advantage in VLSI implementation of efficient partially parallel decoders. Jin Sha 0001, Li Li 0003, Zhongfeng Wang 0001 |
ISCAS | 2 |
| 2009 | LDPC Decoder Design for IEEE 802.15 StandardabstractThis paper presents an efficient decoder design for the LDPC codes in IEEE 802.15 standard. This decoder features by high parallel level, low message memory requirement and code rate flexibility. By processing 72 columns and 72 rows in parallel, it can reach a throughput of 2.8 Gbps to fulfill the standard requirement. Furthermore, the decoder supports three different code rates by employing flexible check node processor units. Jin Sha 0001, Jun Lin 0001, Li Li 0003, Minglun Gao, Zhongfeng Wang 0001 |
ISCAS | 1 |
| 2009 | Area-efficient Reed-Solomon Decoder Design for 10-100 Gb/s ApplicationsabstractWith the extensive applications in high-speed communication systems, the current high-throughput Reed-Solomon decoders are required to achieve the target data rates from 10 Gb/s to 100 Gb/s with low hardware complexity. In this paper, pipeline interleaving inversionless Berlekamp-Massey (PI-iBM) algorithm and pipeline interleaving reformulated inversionless Berlekamp-Massey (PI-RiBM) algorithms for decoding Reed-Solomon codes are presented. Based on these two new algorithms PI-iBM and PI-RiBM Reed-Solomon decoders targeted at 10-100 Gb/s applications are developed. Compared with previously published works, the proposed designs can achieve very high throughput with relatively low hardware complexity. Thus they are well suited for modern high data rate communication systems. Bo Yuan 0001, Li Li 0003, Jin Sha 0001, Zhongfeng Wang 0001 |
ISCAS | 3 |
| 2009 | Multi-Gb/s LDPC Code Design and ImplementationabstractLow-density parity-check (LDPC) code, a very promising near-optimal error correction code (ECC), is being widely considered in next generation industry standards. The VLSI implementation of high-speed LDPC decoder remains a big challenge. This paper presents the construction of a new class of implementation-oriented LDPC codes, namelyshift-LDPCcodes. With girth optimization, this kind of codes can perform as well as computer generated random codes. More importantly, the decoder can be efficiently implemented to obtain very high decoding speeds. In addition, more than 50% of message memory can be generally saved over conventional partially parallel decoder architectures. We demonstrate the benefits of the proposed techniques with an application-specific integrated circuit (ASIC) design (in 0.18-mum CMOS) for a 8192-bit regular LDPC code, which can achieve 5 Gb/s throughput at 15 iterations. Jin Sha 0001, Zhongfeng Wang 0001, Minglun Gao, Li Li 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |