EDBT 2026 Demo / reviewers in the wild / expert
Gangqiang Yang
dblp:163/8605
· DBLP profile ↗
15ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlexMSM: A Flexible FPGA-Based Accelerator for Multi-Scalar Multiplication with Reconfigurable Modular Arithmetic and Optimized Pippenger SchedulingabstractZero-Knowledge Proofs (ZKPs), especially zk-SNARKs, rely heavily on Multi-Scalar Multiplication (MSM), a compute-intensive elliptic curve operation. While prior work targets curves with optimized operations like BLS12-377, MSM on general-purpose curves such as BLS12-381 remains challenging due to imbalanced resource usage, performance gaps between curve operations, and low utilization of point addition. This paper proposes FlexMSM to support scalable MSM cores on a single FPGA for BLS12-381 curve, delivering significant gains over existing works for input sizes from 218 to 226 . Cheng Chen 0076, Gangqiang Yang, Hongchao Zhou, Hailiang Xiong, Zhiguo Wan |
FPGA | 2 |
| 2026 | Multi-Scale Convolutional Attention Model for GNSS Jamming RecognitionabstractGlobal Navigation Satellite System (GNSS) functionality is highly vulnerable to jamming, which disrupts signals. Existing datasets and models for jamming classification are inefficient for low-power GNSS jamming. Additionally, current machine and deep learning models struggle to balance accuracy with efficiency. To address this, we develop the Multi-Featured GNSS Jamming Dataset (MFGJD), which includes 16 distinct jamming types along with a clean GNSS signal. Building on MFGJD, we propose the Multi-Scale Convolutional Attention Model (MSCAM), designed to operate in the precorrelation phase using spectrograms as input. MSCAM integrates channel, spatial, and temporal features, enabling accurate and robust jamming classification. Experimental results show that MSCAM outperforms current models, achieving 96.95% accuracy with 0.09 million parameters and 0.30 milliseconds inference time per sample, offering high efficiency and robustness. Danyal Hussain Shah, Hailiang Xiong, Fen He, Bo He 0005, Gangqiang Yang |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | High-Performance Accelerator for Constant-Time Cross-Domain Integer and Montgomery Inversion on FPGAabstractModular Inversion (MI) is one of the fundamental arithmetic operations in the finite field, which plays an essential role in various cryptographic applications and requires high performance and security. Unfortunately, the simple MI algorithm is vulnerable to side-channel attacks, such as the timing attack, which can compromise the cryptographic system by analyzing the time taken to execute cryptographic algorithms. Attackers may recover the initial data since the time can differ based on the input. Besides, the low complexity and low resource consumption of hardware implementations in MI are also challenging. In this article, we propose two novel modular inversion algorithms, named Constant-Time Integer Modular Inversion (CT-IMI) and Constant-Time Complementary Montgomery Modular Inversion (CT-CMMI). They both consist of constant iteration rounds to resist the timing attack. CT-IMI processes the data in the integer field, which is designed for common scenarios. CT-CMMI is suitable for the cross-domain case, which can directly use data in the Montgomery domain and avoid the conversion steps for some specific applications, e.g., scalar multiplication in Elliptic Curve Cryptography (ECC). In software simulations, we measure the average clock cycles for a single inversion and illustrate the relationship between various bit lengths and the latency. The significant differences between constant and non-constant algorithms demonstrate the vulnerability of modular inversion to timing attacks. In addition, we design two efficient hardware architectures on FPGA. Experimental results show that our CT-IMI can finish a single inversion in 2.56 \(\mu\) s with 4.2k LUTs, 1.8k FFs, and our CT-CMMI requires 2.45 \(\mu\) s with 2.7k LUTs, 1.6k FFs. The product of area and latency of our CT-IMI and CT-CMMI can reach 10.50 and 6.62, respectively, which shows optimal performance compared with all the results in the existing literature. Cheng Chen 0076, Gangqiang Yang, Hongchao Zhou, Hailiang Xiong, Xianye Ben, Zhiguo Wan |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | Global-local coherency contrastive learning for context-aware time series forecasting
Fengqian Ding, Chuandong Lyu, Gangqiang Yang, Hailiang Xiong, Hongchao Zhou |
Knowl. Based Syst. | 5 |
| 2025 | Customized FPGA Implementation of Authenticated Lightweight Cipher Fountain for IoT SystemsabstractAuthenticated Encryption with Associated-Data (AEAD) can ensure both confidentiality and integrity of information in encrypted communication. Distinctive variants are customized from AEAD to satisfy various requirements. In this paper, we take a 128-bit lightweight AEAD stream cipher Fountain as an example. We provide a general cryptographic solution with three Fountain variants. These three variants are for encryption, message authentication code (MAC) generation, and authenticated encryption with associated data, respectively. Besides, we propose area-saved and throughput-improved strategies for the FPGA implementation of Fountain. The conventional paralleled hardware implementation leads to much resource-consuming with higher parallel width. We propose a hybrid architecture with parallel and serial update modes simultaneously. We also analyze the trade-off between area occupation and authentication latency for those two architectures. According to our discussion, hybrid architectures can perform efficiently with higher throughput than most ciphers, including Grain-128 x32. Our Fountain keystream generator occupies 46 slices on Spartan-3 FPGAs, smaller than most ciphers with the same security level, and even smaller than the 80-bit security level cipher Trivium. In summary, the customized Fountain with optimized implementations on FPGA is suitable for various applications in the field of IoT. Zhengyuan Shi, Cheng Chen 0076, Gangqiang Yang, Hongchao Zhou, Hailiang Xiong, Zhiguo Wan |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | A Robust and Scalable Multigrid Solver for 3-D Low-Frequency Electromagnetic Diffusion ProblemsabstractMultigrid (MG) solvers, typically increasing the computational time linearly with the grid size, are suitable for large-scale forward modeling problems. However, for electromagnetic (EM) problems as frequency decreases and grid is increasingly stretched, MG solvers for EM modeling can converge slowly or even diverge. We propose an efficient four-color Line Gauss-Seidel (GS) MG for finite difference (FD) frequency-domain EM solution. In this algorithm, the edge components attached to nodes on each line in one particular direction are updated simultaneously, leading to that the solution in each local region satisfies divergence free condition. Due to the fact that each local linear system of equations is completely uncoupled with that formed for its disjoint lines, we can group all lines of grid nodes into four colors with the requirement that all local systems formed for lines with the same color are disjoint. This can be utilized to parallelize or vectorize our algorithm. The correctness is verified by comparing with the analytical solution based on a three-layered model. The numerical performance is examined by comparing with other commonly used state-of-art solvers based on three increasingly more complex models, indicating the efficiency dominance, good parallelization and excellent ability on handling grid stretching for our algorithm. Yongfei Wang, Rongwen Guo, Kejia Pan, Gangqiang Yang, Jian Li 0046, Xiaokang Deng |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | An Efficient Multigrid Solver With Two-Color Plane Gauss-Seidel Smoother for 3D Low-Frequency Electromagnetic ModelingabstractMultigrid (MG) methods are among the best choice for stable and efficient three-dimensional (3D) forward modeling of electromagnetic (EM) fields over large area due to their linear dependence of computational time on grid size. However, as the frequency decreases to near zero and/or the grid is increasingly stretched, MG solvers for EM modeling based on the curl-curl equations converge slowly or even diverge. In this letter, we develop an efficient MG algorithm combined with a two-color plane Gauss-Seidel (GS) smoother for finite difference forward modeling of EM fields particularly at low frequencies. In this algorithm, we group different planes of the grid nodes into two colors. In each color, the components attached to different planes are totally decoupled and can be solved simultaneously, which can be distributed to different processors. The Dublin Test Model 1 is used to verify the accuracy of our algorithm and examine the numerical performance of our method against MG algorithms based on a four-color cell-block GS smoother and the Bi-Conjugate Gradient stabilized (BICGstab) smoother (as four-color cell-block GS MG and BICGstab-MG, respectively), and BICGstab and Quasi-Minimal Residual (QMR) both preconditioned with block incomplete lower-upper (blockILU) decomposition (as blockILU-BICGstab and blockILU-QMR, respectively). The numerical test based on OpenMP shows the good parallelization of our algorithm. Grids stretched to different degrees are designed to examine its ability to handle grid-stretching. The numerical performance comparison indicates its remarkable dominance in efficiency and stability. Gangqiang Yang, Rongwen Guo, Chunming Liu, Yongfei Wang, Jian Li 0046, Kejia Pan |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Design Space Exploration of Galois and Fibonacci Configuration Based on Espresso Stream CipherabstractFibonacci and Galois are two different kinds of configurations in stream ciphers. Although many transformations between two configurations have been proposed, there is no sufficient analysis of their FPGA performance. Espresso stream cipher provides an ideal sample to explore such a problem. The 128-bit secret key Espresso is designed in Galois configuration, and there is a Fibonacci-configured Espresso variant proved with the equivalent security level. To fully leverage the efficiency of two configurations, we explore the hardware optimization approaches toward area and throughput, respectively. In short, the FPGA-implemented Fibonacci cipher is more suitable for extremely resource-constrained or high-throughput applications, while the Galois cipher compromises both area and speed. To the best of our knowledge, this is the first work to systematically compare the FPGA performance of cipher configurations under relatively fair cryptographic security. We hope this work can serve as a reference for the cryptography hardware architecture research community. Zhengyuan Shi, Cheng Chen 0076, Gangqiang Yang, Hailiang Xiong, Fudong Li 0002, Honggang Hu, Zhiguo Wan |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2023 | Hardware Optimizations of Fruit-80 Stream Cipher: Smaller than GrainabstractFruit-80, which emerged as an ultra-lightweight stream cipher with 80-bit secret key, is oriented toward resource-constrained devices in the Internet of Things. In this article, we propose area and speed optimization architectures of Fruit-80 on FPGAs. Our implementations include both serial and parallel structure and optimize area, power, speed, and throughput, respectively. The area optimization architecture aims to achieve the most suitable ratio of look-up-tables and flip-flops to fully utilize the reconfigurable unit. It also reuses NFSR and LFSR feedback functions to save resources for high throughput. The speed optimization architecture adopts a hybrid approach for parallelization and reduces the latency of long data paths by pre-generating primary feedback and inserting flip-flops. Besides, we recommend using the round key function to optimize serial or parallel implementations for Fruit-80 and using indexing and shifting methods for different throughput. In conclusion, our results show that the area optimization architecture occupies up to 35 slices on Xilinx Spartan-3 FPGA and 18 slices on Xilinx 7 series FPGA, smaller than that of Grain and other common stream ciphers. The optimal throughput/area ratio of the speed optimization architecture is 7.74 Mbps/slice, better than that of Grain v1, which is 5.98 Mbps/slice. The serial implementation of Fruit-80 with round key function occupies only 75 slices on Spartan-3 FPGA. To the best of our knowledge, the result sets a new record of the minimum area in lightweight cipher implementation on FPGA. Gangqiang Yang, Zhengyuan Shi, Cheng Chen 0076, Hailiang Xiong, Fudong Li 0002, Honggang Hu, Zhiguo Wan |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2022 | Work-in-Progress: Towards a Smaller than Grain Stream Cipher: Optimized FPGA Implementations of Fruit-80abstractFruit-80, an ultra-lightweight stream cipher with 80-bit secret key, is oriented toward resource constrained devices in the Internet of Things. In this paper, we propose area and speed optimization architectures of Fruit-80 on FPGAs. The area optimization architecture reuses NFSR&LFSR feedback functions and achieves the most suitable ratio of look-up-tables and flip-flops. The speed optimization architecture adopts a hybrid approach for parallelization and reduces the latency of long data paths by pre-generating primary feedback and inserting flip-flops. In conclusion, the optimal throughput-to-area ratio of the speed optimization architecture is better than that of Grain v1. The area optimization architecture occupies only 35 slices on Xilinx Spartan-3 FPGA, smaller than that of Grain and other common stream ciphers. To the best of our knowledge, this result sets a new record of the minimum area in lightweight cipher implementations on FPGA. Gangqiang Yang, Zhengyuan Shi, Cheng Chen 0076, Hailiang Xiong, Honggang Hu, Zhiguo Wan, Keke Gai, Meikang Qiu |
CASES | 1 |
| 2021 | Mobile target localization and tracking techniques in harsh environment utilizing adaptive multi-modal data fusionabstractAbstract Multi‐source cooperative positioning systems relying on federated filtering have become attractive development directions of navigation strategy in real‐time localization and tracking under challenging urban environment. However, local Kalman filters of traditional federated filtering may result in divergence when the system model or the measurements is inaccurate, and the fixed information distribution coefficient in federated filter cannot adaptively reflect the performance of each local filter. To improve the precision and robustness of the integrated navigation system, a novel adaptive federated strong tracking Kalman filter with dynamic fading factor mechanism for multi‐sensor information fusion is proposed. Through iterative computation of the fading factor and updating the adaptive weight coefficients, strong tracking filter becomes robust to the uncertainty of system model. Meanwhile, an effective adaptive information distribution estimation algorithm based on the predicted residuals is constructed to balance the contributions of the kinematic model information and measurements on the state estimates. To ensure the stability of filtering, a simplified fusion strategy is established to solve the singular problem of the global covariance matrix of estimation error. Theoretical analysis and simulation results demonstrate the validity of the proposed approach in improving the accuracy and robustness of the integrated navigation and positioning systems. The proposed integrated multi‐modal cooperative navigation and positioning algorithm will be of great significance to the implementation of real‐time mobile target localization and tracking in harsh environment or dead zone. Zhenzhen Mai, Hailiang Xiong, Gangqiang Yang, Weihong Zhu, Fen He, Ruochen Bian |
IET Commun. | 3 |
| 2018 | Towards a Cryptographic Minimal Design: The sLiSCP Family of PermutationsabstractThe security of highly resource constrained applications is often viewed in the literature from a single aspect of a specific cryptographic primitive. More precisely, most of the proposed lightweight cryptographic primitives focus on providing a single functionality within the available hardware area dedicated for security purposes. In this paper, we argue that for such applications, a cryptographic primitive that follows the cryptographic minimal design strategy maybe the only realistically adopted security solution where there is a constrained GE budget for all security functionalities. Indeed, it is reasonable, if not desirable, for the adopted cryptographic design to have well justified building components and to provide minimal overhead for multiple cryptographic functionalities including encryption, hashing, authentication, and pseudorandom bit generation. Following such a strategy, we propose the sLiSCP family of lightweight cryptographic permutations which employs two of the most hardware efficient and extensively cryptanalyzed constructions, namely a 4-subblock Type-2 Generalized Feistel-like Structure (GFS) and round-reduced unkeyed Simeck. In addition to the hardware efficiency, we follow restrictive security design goals which enable us to provide resistance against differential and linear cryptanalysis, as well as guaranteed resistance to diffusion-based, algebraic, and self-symmetry distinguishers, and accordingly, we claim that there exist no structural distinguishers for sLiSCP-b with a complexity below 2b=2 where b is the state size. Moreover, we present the sLiSCP duplex sponge mode to illustrate how the permutations can be used in a unified design that provides (authenticated) encryption, hashing, and pseudorandom bit generation functionalities. Finally, we report two efficient parallel hardware implementations for the sLiSCP unified duplex sponge mode when using sLiSCP-192 (resp. sLiSCP-256) in CMOS 65 nm ASIC with area of 2289 (resp. 3039) GE and a throughput of 29.62 (resp. 44.44) kbps, and their areas in CMOS 130 nm are 2498 (resp. 3319) GE. Riham AlTawy, Raghvendra Rohit 0001, Morgan He, Kalikinkar Mandal, Gangqiang Yang, Guang Gong |
IEEE Trans. Computers | 5 |
| 2018 | SLISCP-light: Towards Hardware Optimized Sponge-specific Cryptographic PermutationsabstractThe emerging areas in which highly resource constrained devices are interacting wirelessly to accomplish tasks have led manufacturers to embed communication systems in them. Tiny low-end devices such as sensor networks nodes and Radio Frequency Identification (RFID) tags are of particular importance due to their vulnerability to security attacks, which makes protecting their communication privacy and authenticity an essential matter. In this work, we present a lightweight do-it-all cryptographic design that offers the basic underlying functionalities to secure embedded communication systems in tiny devices. Specifically, we revisit the design approach of the sLiSCP family of lightweight cryptographic permutations, which was proposed in SAC 2017. sLiSCP is designed to be used in a unified duplex sponge construction to provide minimal overhead for multiple cryptographic functionalities within one hardware design. The design of sLiSCP follows a 4-subblock Type-2 Generalized Feistel-like Structure (GFS) with unkeyed round-reduced Simeck as the round function, which are extremely efficient building blocks in terms of their hardware area requirements. In S L I SCP-light, we tweak the GFS design and turn it into an elegant Partial Substitution-Permutation Network construction, which further reduces the hardware areas of the S L I SCP permutations by around 16% of their original values. The new design also enhances the bit diffusion and algebraic properties of the permutations and enables us to reduce the number of steps, thus achieving a better throughput in both the hashing and authentication modes. We perform a thorough security analysis of the new design with respect to its diffusion, differential and linear, and algebraic properties. For S L I SCP-light-192, we report parallel implementation hardware areas of 1,820 (respectively, 1,892)GE in CMOS 65 nm (respectively, 130 nm ) ASIC. The areas for S L I SCP-light-256 are 2,397 and 2,500GE in CMOS 65 nm and 130 nm ASIC, respectively. Overall, the unified duplex sponge mode of S L I SCP-light-192, which provides (authenticated) encryption and hashing functionalities, satisfies the area (1,958GE), power (3.97μ W ), and throughput (44.4kbps) requirements of passive RFID tags. Riham AlTawy, Raghvendra Rohit 0001, Morgan He, Kalikinkar Mandal, Gangqiang Yang, Guang Gong |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2017 | sLiSCP: Simeck-Based Permutations for Lightweight Sponge Cryptographic Primitives
Riham AlTawy, Raghvendra Rohit 0001, Morgan He, Kalikinkar Mandal, Gangqiang Yang, Guang Gong |
SAC | 5 |
| 2015 | The Simeck Family of Lightweight Block Ciphers
Gangqiang Yang, Bo Zhu 0007, Valentin Suder, Mark D. Aagaard, Guang Gong |
CHES | 1 |