EDBT 2026 Demo / reviewers in the wild / expert
Xinkuang Geng
dblp:377/4612
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-3673-237XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-author · 11 since 2021Software engineering, systems software and programming languages · 7 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SA-ANT: Efficient Low-Bit Group-Wise Quantization for Large Language Models via Sign-Asymmetric Adaptive Numeric TypeabstractLarge language models (LLMs) have demonstrated remarkable potential across diverse domains; meanwhile, their large parameter sizes pose substantial inference costs, motivating the need for efficient low-bit quantization. Group-wise quantization, which adopts finer granularity, has been widely used to improve low-bit quantization performance. Several adaptive numeric types have been proposed to further enhance low-bit group-wise quantization; however, they construct quantization grids based on symmetric numeric types, which limits their ability to model asymmetric distributions. To address this limitation, we propose SA-ANT, a sign-asymmetric adaptive numeric type for efficient low-bit group-wise quantization. SA-ANT constructs quantization grids separately on the positive and negative sides, enabling adaptive support for asymmetric and non-uniform distributions. Furthermore, the carefully designed SA-ANT not only reduces quantization errors but also ensures a unified computing across different sub numeric types, thereby facilitating hardware efficiency. To accelerate LLM inference, we develop (1) a quantization framework that transforms LLM weights into the SA-ANT and adaptively selects the sub numeric type for each group, and (2) an accelerator that maps SA-ANT inference to low-bit INT operations. Experimental results show that SA-ANT delivers 3.92%– 5.57% higher accuracy than state-of-the-art adaptive numeric types under 3-bit weight quantization, while also enabling 7.84%– 44.65% area savings and 7.80%–43.88% power reductions. Xinkuang Geng, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 1 |
| 2026 | HAP: Accelerating DNNs with Resolution-Preserved Quantization by Harnessing Adaptive-PrecisionabstractReducing the precision in post-training quantization can cause catastrophic accuracy loss in Deep Neural Networks, especially when compressing the activations. To address this problem, we present a novel adaptive-precision quantization (APQ) and accelerator design that achieves lossless activation compression by exploiting the inherent coding redundancy. Compared to existing APQ methods, this design can be generalized to implement asymmetric quantization, making it particularly suitable for activations. The accelerator offers a practical solution to mitigate the computational workload imbalance problem incurred by variable precision. A dual-precision quantization scheme further provides the flexibility to trade off accuracy and performance. Erjing Luo, Xinkuang Geng, Honglan Jiang, Leibo Liu, Jie Han 0001 |
DATE | 2 |
| 2026 | Approximate Signed Multiplier Designs for Efficient CNN Inference
Mengshuo Zhang, Xiaolu Hu, Xinkuang Geng, Honglan Jiang |
ISCAS | 4 |
| 2026 | LUT-ALMs: Trading Off Accuracy and Power for Approximate Logarithmic Multipliers via LUT OptimizationabstractLogarithmic multiplier (LM) converts fixed-point (FxP) input operands to logarithmic numbers and performs multiplication with simple shift and addition operations, which achieves distinct power reduction, yet with significant single-sided errors. This paper proposes to fuse error compensation with logarithmic conversion by using customized look-up tables (LUTs). To avoid the use of large LUTs, partition strategies are designed for the optimization of LUTs. In addition, to effectively balance the accuracy and hardware costs, two iterative algorithms are proposed for generating precision-configurable LUTs. Based on the optimized LUTs, high-accuracy and low-power approximate LMs (LUT-ALMs) are constructed for 8-bit and 16-bit multiplications. Furthermore, to enhance the flexibility of data type and bit width, a mixed-mode LM (MM-ALM) supporting eight multiplication modes is devised. Compared with the exact 8-bit signed multiplier from Synopsys DesignWare (DW) library (DW-Exact), LUT-ALMs show up to 30.96% and 23.99% reductions in the power-delay product (PDP) and area-delay product (ADP), respectively, with a mean relative error distance (MRED) of 3.45%. Compared with state-of-the-art approximate 8- bit signed multipliers, LUT-ALMs form the Pareto front in terms of PDP and MRED. For 16-bit multiplication, LUT-ALMs obtain up to 70.90% and 62.13% savings in PDP and ADP with a MRED of 2.75%, compared with the corresponding DW-Exact. Compared with the corresponding mixed-mode exact multiplier constructed of DW multipliers, MM-ALM performing 8-bit multiplication achieves up to 57.44% and 47.99% reductions in PDP and ADP, respectively. When performing 4-bit multiplications, MM-ALM can save up to 37.16% savings in PDP. With lower hardware over-heads, LUT-ALMs and MM-ALM present comparable accuracy to the corresponding exact designs in the considered convolutional neural networks (CNNs) and image processing applications. The hardware description of the devised LMs is open-sourced athttps://anonymous.4open.science/r/LUT-ALM-5FD1. Xinkuang Geng, Xiaolu Hu, Hui Wang 0023, Jianfei Jiang 0001, Qin Wang 0009, Siting Liu 0001, Jie Han 0001, Honglan Jiang |
IEEE Trans. Computers | 2 |
| 2025 | Lookup Table Refactoring: Towards Efficient Logarithmic Number System Addition for Large Language ModelsabstractCompared to integer quantization, logarithmic quantization aligns more effectively with the long-tailed distribution of data in large language models (LLMs), resulting in lower quantization errors. Moreover, the logarithmic number system (LNS) employs a fixed-point adder to perform multiplication, indicating a potential reduction in computational complexity for LLM accelerators that require extensive multiply-accumulate (MAC) operations. However, a key bottleneck is that LNS addition requires complex nonlinear functions, which are typically approximated using lookup tables (LUTs). This study aims to reduce the hardware resources needed for LUTs in LNS addition while maintaining high precision. Specifically, we investigate the specific nature of addition operations within LLMs; the relationship between the hardware parameters of the LUT and the computing errors is then mathematically derived. Based on these insights, we propose LUT refactoring to optimize the LUT for enhanced efficiency in LNS addition. With 10.93% and 19.78% reductions in area-delay product (ADP) and power-delay product (PDP), respectively, LUT refactoring results in an accuracy improvement of up to 33.5% in LLM benchmarks compared to the naive design. When compared to integer quantization, our method achieves higher accuracy while reducing area by 18.27% and power by 42.61%. Xinkuang Geng, Siting Liu 0001, Hui Wang 0023, Jie Han 0001, Honglan Jiang |
DATE | 1 |
| 2025 | Segment-Wise Accumulation: Low-Error Logarithmic Domain Computing for Efficient Large Language Model InferenceabstractLogarithmic domain computing (LDC) has great potential for reducing quantization errors and computational complexity in Large Language Models (LLMs). While logarithmic multiplication can be efficiently implemented using fixed-point addition, the primary challenge in multiply-accumulate (MAC) operations is balancing the precision of logarithmic adders with their hardware overhead. Through a detailed analysis of the errors inherent in LDC-based LLMs, we propose segment-wise accumulation (SWA) to mitigate these errors. In addition, a processing element (PE) is introduced to enable SWA in the systolic array architecture. Compared with the accumulation scheme devised for enhancing floating-point computing, the proposed SWA facilitates the integration into existing accelerator architectures, resulting in lower hardware overhead. The experimental results show that SWA allows LDC under low-precision configurations to achieve remarkable accuracy in LLMs, demonstrating higher hardware efficiency than merely increasing the precision of individual computations. Our method, while maintaining a lower hardware overhead than traditional LDC, achieves more than 13.9% improvement in average accuracy across multiple zero-shot benchmarks in LLAMA-2-7B. Furthermore, compared to integer domain computing, a logarithmic processing element array based on the proposed SWA yields reductions of 24.6% in area and 42.3% in power, while achieving higher accuracy. Xinkuang Geng, Yunjie Lu, Hui Wang 0023, Honglan Jiang |
DATE | 1 |
| 2025 | A Low-Power Mixed-Precision Integrated Multiply-Accumulate Architecture for Quantized Deep Neural NetworksabstractAs mixed-precision quantization techniques have been widely considered for balancing computational efficiency and flexibility in quantized deep neural networks (DNNs), mixed-precision multiply-accumulate (MAC) units are increasingly important in DNN accelerators. However, conventional mixed-precision MAC architectures support either signed × signed or unsigned ×unsigned multiplications. The signed ×unsigned multiplication enhancing the computing efficiency of DNNs with ReLU activations has never been considered in the design of mixed-precision MAC. Thus, this work proposes a mixed-precision MAC architecture supporting six operation modes, int8 × int8, int8 × uint8, two int4 × int4, two int4 × uint4, four int2 × int2, and four int2 × uint2. In this design, to balance the power and delay of different modes, the multiplication is implemented based on four precision-split 4×4 multipliers (PS4Ms). The accumulation is integrated into the partial product accumulation of the multiplication to eliminate redundant switching activities in separate compression. With 10% area reduction, the proposed MAC denoted as PS4MAC, reduces the power by over 35%, 42%, and 56% for 8-bit, 4-bit, and 2-bit operations, respectively, compared with the design based on the Synopsys DesignWare (DW) multipliers. Additionally, it achieves over 23% power savings for 8-bit operations compared to state-of-the-art (SotA) mixed-precision MAC designs. To save more power, an approximate computing mode for 8-bit multiplication is further designed, resulting in a MAC unit enabling eight operation modes, referred to as PS4MAC_AP. Finally, output-stationary systolic arrays (SAs) are explored using the above-mentioned MAC designs to implement DNNs operating under a 1 GHz clock. Our designs show the highest energy efficiency and outstanding area efficiency in all 8-bit, 4-bit, and 2-bit operation modes. Compared with the traditional SA with high-precision-split multipliers, PS4MAC_AP improves the energy efficiency for 8-bit operations by 0.6 TOPS/W, and PS4MAC achieves 0.4 TOPS/W - 0.7 TOPS/W improvement for all operation modes. Xiaolu Hu, Xinkuang Geng, Zhigang Mao, Jie Han 0001, Honglan Jiang |
DATE | 2 |
| 2024 | QUQ: Quadruplet Uniform Quantization for Efficient Vision Transformer InferenceabstractWhile exhibiting superior performance in many tasks, vision transformers (ViTs) face challenges in quantization. Some existing low-bit-width quantization techniques cannot effectively cover the whole inference process of ViTs, leading to an additional memory overhead (22.3%-172.6%) compared with corresponding fully quantized models. To address this issue, we propose quadruplet uniform quantization (QUQ) to deal with data of various distributions in ViT. QUQ divides the entire data range into at most four subranges that are uniformly quantized with different scale factors. To determine the partition scheme and quantization parameters, an efficient relaxation algorithm is proposed accordingly. Moreover, dedicated encoding and decoding strategies are devised to facilitate the design of an efficient accelerator. Experimental results show that QUQ surpasses state-of-the-art quantization techniques; it is the first viable scheme that can fully quantize ViTs to 6-bit with acceptable accuracy. Compared with conventional uniform quantization, QUQ leads to not only a higher accuracy but also an accelerator with lower area and power. Xinkuang Geng, Siting Liu 0001, Leibo Liu, Jie Han 0001, Honglan Jiang |
DAC | 1 |
| 2024 | Compact Powers-of-Two: An Efficient Non-Uniform Quantization for Deep Neural NetworksabstractTo reduce the demands for computation and memory of deep neural networks (DNNs), various quantization techniques have been extensively investigated. However, conventional methods cannot effectively capture the intrinsic data characteristics in DNNs, leading to a high accuracy degradation when employing low-bit-width quantization. In order to better align with the bell-shaped distribution, we propose an efficient non-uniform quantization scheme, denoted as compact powers-of-two (CPoT). Aiming to avoid the rigid resolution inherent in powers-of-two (PoT) without introducing new issues, we add a fractional part to its encoding, followed by a biasing operation to eliminate the unrepresentable region around O. This approach effectively balances the grid resolution in both the vicinity of 0 and the edge region. To facilitate the hardware implementation, we optimize the dot product for CPoT based on the computational characteristics of the quantized DNNs, where the precomputable terms are extracted and incorporated into bias. Consequently, a multiply-accumulate (MAC) unit is designed for CPoT using shifters and look-up tables (LUTs). The experimental results show that, even with a certain level of approximation, our proposed CPoT outperforms state-of-the-art methods in data-free quantization (DFQ), a post-training quantization (PTQ) technique focusing on data privacy and computational efficiency. Furthermore, CPoT demonstrates superior efficiency in area and power compared to other methods in hardware implementation. Xinkuang Geng, Siting Liu 0001, Jianfei Jiang 0001, Honglan Jiang |
DATE | 1 |
| 2024 | A Configurable Approximate Multiplier for CNNs Using Partial Product SpeculationabstractTo improve the performance and energy efficiency of the compute-intensive convolutional neural networks (CNNs), approximate multipliers have widely been investigated, taking advantage of the inherent error tolerance in CNNs. However, as per their divergencies in the capability of error tolerance, different CNN models and datasets may require various accuracy in multiplication. Thus, in this paper, we propose an energy-efficient approximate multiplier with configurable accuracy to satisfy the continuously evolving requirements of CNNs. In this design, the approximation level is configured by changing the processing scheme for inputs due to their significance to accuracy. The correlations between partial products (PPs) are utilized to eliminate the generation and accumulation of some less significant PPs that are speculated by their adjacent more significant ones. Consequently, four approximate multiplier configurations are devised for 8x8 unsigned multiplication, denoted as AMPPS_S2, AMPPS_S3, AMPPS_S4, and AMPPS_S6. Compared with existing approximate multipliers, the proposed designs show significantly higher accuracy. Compared with an exact carry-save array multiplier, AMPPS_S2 can reduce the power dissipation, delay, and area by 21.7%, 35.1 %, and 24.4%, respectively. Moreover, to enhance the efficiency of the proposed approximate multiplier in systolic array-based hardware architectures for CNNs, a novel encoding strategy is proposed for storing the pre-trained weights. Obtaining a similar classification accuracy to the accurate implementation (tested in ResNet18 and ResNet50 on ImageNet), the 32x32 systolic array using AMPPS_S2 shows a 13.5% reduction in power consumption and a 19.2% reduction in area. Overall, the experimental results demonstrate that the proposed approximate multiplier results in higher accuracy in CNN-based image classification, with lower hardware overhead, compared with state-of-the-art approximate multipliers. Xiaolu Hu, Xinkuang Geng, Zizhong Wei, Honglan Jiang |
DATE | 3 |
| 2024 | A Low-Power and High-Accuracy Approximate Adder for Logarithmic Number SystemabstractThe Logarithmic Number System (LNS) exploits the non-uniform distribution of data in convolutional neural networks (CNNs), so it leads to a high accuracy for image classification. An LNS provides an easier way to implement complex operations such as multiplication and division. However, addition and subtraction in the LNS require huge hardware resources due to the involved nonlinear operations. To mitigate this problem, we design a low-power approximate logarithmic adder with high-accuracy. Initially, a compact piecewise linear approximation (CPLA) algorithm is proposed to approximately compute the binary exponentiation and logarithm. Implemented by using simple circuits, the CPLA algorithm results in higher accuracy than the classical Mitchell’s algorithm. Consequently, three approximate logarithmic adders are devised, denoted as LA_CPLA1, LA_CPLA2, and LA_CPLA3. Compared with the logarithmic adder design based on lookup tables, the proposed LA_CPLA3 with a configuration of (e, f, n) = (7, 6, 3) achieves 35.05% and 39.80% reductions in area and power dissipation respectively, with a 0.01% mean relative error distance (MRED). We define (e, f) as the bit width of the logarithmic adder, where e and f are the bit widths of the integer and fractional parts, respectively. n is the approximate LSBs in the proposed LA_CPLAs processed by using OR gates. Compared with the multiply and accumulate (MAC) unit in a conventional system using fixed-point numbers, the MAC in the LNS using the proposed LA_CPLAs achieve a lower power by 5.96% to 32.02%, and a smaller area by 6.48% to 32.40%. To assess the efficiency of the proposed approximate adders, they are applied to the implementations of two image processing and CNN applications. The simulation results show that LA_CPLAs result in marginal accuracy loss compared to the corresponding accurate implementations. Xinkuang Geng, Qin Wang 0009, Jie Han 0001, Honglan Jiang |
ACM Great Lakes Symposium on VLSI | 2 |