EDBT 2026 Demo / reviewers in the wild / expert
Hui Chen 0015
dblp:12/417-15
· DBLP profile ↗
20ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0001-5462-8029ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 7 first-author · 18 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PIM-NoC: A NoC Architecture with In-Router Processing-in-Memory for DNN Acceleration
Haixiang Ren, Lixia Han, Cuiyu Qi, Hui Chen 0015 |
ISCAS | 5 |
| 2026 | C2NoC: A Communication-Computation Coupled NoC-based Neural Network Accelerator
Cuiyu Qi, Hui Chen 0015, Lixia Han, Chenkai Cao, Boxiang Zhang, Weiqiang Liu 0001 |
ISCAS | 3 |
| 2026 | Perception-Core: Reconfigurable Energy-Efficient Domain-Specific Architecture with Multimodal Fusion for AIoT
Xinyu Wang 0027, Yuanhua Deng, Haixiang Ren, Hui Chen 0015, Li Li 0003 |
ISCAS | 6 |
| 2026 | Low-Latency Scaling-Free Hyperbolic CORDIC Algorithm Based on Linear Rotation Angles and Leading-One Bit Detection
Fei Lyu 0006, Zongguang Yu, Weiqiang Liu 0001, Hui Chen 0015 |
ISCAS | 6 |
| 2026 | Instant-CIM: An Instant Neural Radiance Field Computing-In-Memory Architecture for Low-Power and Real-Time AR/VR RenderingabstractNovel View Synthesis is a foundational technique for creating immersive Augmented and Virtual Reality (AR/VR) experiences, aiming to generate photorealistic images of a scene from arbitrary camera viewpoints using only a limited set of source images, with Neural Radiance Fields (NeRF) emerging as the state-of-the-art solution. However, real-time NeRF rendering on low-power devices remains challenging due to its memory-intensive hash encoding and compute-intensive Multilayer Perception (MLP). In this work, we propose Instant-CIM, the fully on-chip Computing-in-Memory (CIM) architecture for efficient NeRF rendering. At the algorithm level, Instant-CIM proposes a spatially-adaptive framework that dynamically selects the number of active hash encoding levels per spatial region based on a composite importance score derived from density and gradient. The approach replaces uniform level allocation with a threshold-based strategy that activates finer encoding levels only in regions with high representation complexity. At the hardware level, Instant-CIM proposes an in-situ hash engine that implements in-memory hash query and interpolation through 3D scene grid decomposition and Z-order based mapping schemes. Meanwhile, Instant-CIM proposes a sparse MLP engine that leverages differential-based input complemented by a precision-adjustable skipping mechanism to fully exploit spatial similarities. Comprehensive evaluation across synthetic datasets demonstrates that Instant-CIM achieves 3.0×~4.9× improvement in rendering speed and 8.6×~33× enhancement in energy efficiency compared to state-of-the-art NeRF architecture. Lixia Han, Hui Chen 0015, Xueming Fu, Ke Chen 0018, Peng Huang 0004, Yijun Cui, Weiqiang Liu 0001 |
IEEE Trans. Computers | 3 |
| 2026 | Ultra-Low Latency Generalized Architecture for Complex Nth Root and Nth Power ComputationabstractThis paper proposes a novel computing architecture for high-precision, low-latency, low-power, and cost-effective computation of complex numberNth roots andNth powers. By integrating the high-precision properties of the coordinate rotation digital computer (CORDIC) algorithm with the low-latency benefits of piecewise linear (PWL) approximation, the architecture leverages the binary logarithm-antilogarithm relationship to compute roots and powers of arbitrary complex numbers. Specifically, TheNth roots andNth powers of the modulus of the input complex number are computed using normalization preprocessing and PWL, and the conversion between the plane coordinate and polar coordinate forms of the complex number is achieved using the CORDIC algorithm. The design is implemented in Verilog HDL and synthesized using 40nm CMOS technology at a frequency of 1GHz. The synthesis results show that the area consumption for the complexNth root computation is$24193.92\mu $m2, with a power consumption of 2.1352mW. The area consumption for the complexNth power computation is$20836.83\mu $m2, with a power consumption of 1.8658mW. Compared to the latest complexNth root design, the proposed architecture reduces the area by 11.67% and power consumption by 9.33%. The average accuracy exhibits only a slight reduction compared to the state-of-the-art design while remaining at the same order of magnitude. Furthermore, the computation delay for theNth root architecture is only 58.46% of the delay of the latest complexNth root design, while the delay for theNth power architecture is 94.87% of the delay of the existing real-valuedNth power design. Liangbo Xie, Mu Zhou, Hui Chen 0015 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | PQA-FGS: Piecewise Quadratic Approximation with Fine-grained Segmentation for High-Precision Non-linear ComputationabstractNon-linear functions are fundamental mathematical computations in the realms of digital signal processing and artificial intelligence applications. Despite efforts to develop hardware accelerators for non-linear functions, current methodologies exhibit shortcomings in achieving both high-speed and high-precision computations. This paper introduces a novel scheme for high-precision non-linear approximation that leverages piecewise polynomial computation. By employing quadratic approximation, we reduce the number of segments and introduce a fine-grained computation method to meticulously control computational resources. Our example design, synthesized using 40nm CMOS technology, occupies an area of 15123µm2and consumes a power of 3.48mW when operating at 1GHz. In comparison with the state-of-the-art research, the proposed design boasts a reduction of 76.6% in area and an 86.3% decrease in power consumption, while maintaining equivalent precision levels. Lianghua Quan, Fei Lyu 0006, Hui Chen 0015 |
ISCAS | 4 |
| 2025 | High-Radix Generalized Hyperbolic CORDIC and Its Hardware ImplementationabstractIn this paper, we propose a high-radix generalized hyperbolic coordinate rotation digital computer (HGH-CORDIC). This algorithm not only computes logarithmic and exponential functions with any fixed base but also significantly reduces the number of iterations required compared to traditional CORDIC methods. Initially, we present the general iteration formulas for HGH-CORDIC. Subsequently, we discuss its pivotal convergence properties and selection criteria, exemplifying these with commonly used cases. Through extensive software simulations, we validate the theoretical foundations of our approach. Finally, we explore efficient hardware implementation strategies. Our analysis indicates that, relative to state-of-the-art radix-2 GH-CORDIC, the proposed HGH-CORDIC can decrease the number of iterations by more than$50\%$while maintaining comparable accuracy. Synthesized under the 28nm CMOS technology, the reports show that the reference circuit can save about$40\%$area and power consumption averagely for$2^{x}$and$log_{2}x$calculations compared with the latest CORDIC method. Hui Chen 0015, Lianghua Quan, Ke Chen 0018, Weiqiang Liu 0001 |
IEEE Trans. Computers | 1 |
| 2025 | High-Precision Low-Latency Method and Architecture for Computing Binary and Decimal LogarithmsabstractBinary and decimal logarithms (BDLs) are commonly used in science and engineering. This brief presents a theory of the radix-4 generalized hyperbolic coordinate rotation digital computer (GH-CORDIC) to compute them directly. Compared with traditional hyperbolic CORDIC (TH-CORDIC), the two logarithms can be calculated without extra dividers or multipliers. Compared with the GH-CORDIC, this theory has low iterations under the same high precision. Through theoretical derivation and software simulation, we can find that the calculation accuracy can reach the magnitude of$10^{-7}$, and the number of iterations can be reduced by more than 50%. Through hardware implementation, the synthesis report shows that the proposed architecture can save 53.44% area and 46.36% power consumption compared with the latest radix-2 GH-CORDIC method. Hui Chen 0015, Lianghua Quan, Weiqiang Liu 0001, Zhonghai Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2024 | HGH-CORDIC: A High-Radix Generalized Hyperbolic COordinate Rotation Digital ComputerabstractIn this paper, we propose a high-radix generalized hyperbolic coordinate rotation digital computer (HGH-CORDIC), which not only can compute the logarithmic and exponential functions with any fixed base, but also can reduce the number of iterations compared with the traditional CORDIC. First, we propose the general iteration formulas for HGH-CORDIC. Then we demonstrate its important convergence property and selection criteria, and illustrate them with the most commonly used examples. Through software simulation, we further prove the correctness of the theory. Finally, we analyze how to implement it efficiently in hardware. Compared with the state-of-the-art work, HGH-CORDIC can reduce the number of iterations by more than 50% with the same accuracy. Hui Chen 0015, Lianghua Quan, Weiqiang Liu 0001 |
ARITH | 1 |
| 2023 | Low-Cost High-Precision Architecture for Arbitrary Floating-Point Nth Root ComputationabstractIn this paper, we propose a feasible architecture with high precision and low resource consumption to compute the$N$th root of a floating-point number, which is mainly based on radix-4 SRT and 2-based Coordinate Rotation Digital Computer (CORDIC). Simulation results show that our method can achieve a relative error of the magnitude of 10−7. Under the same precision requirements, the hardware implementation results show a better performance of our design in terms of area, power, and absolute delay compared with the method based on the generalized hyperbolic CORDIC. After synthesizing it under the TSMC$40n$m CMOS technology, it can be obtained that our design can achieve an area consumption of$125465.80\ \mu m^{2}$and power consumption of 97.8062 mW at the highest frequency of 3.12 GHz. Wanyuan Hong, Hui Chen 0015, Lianghua Quan, Li Li 0003 |
ISCAS | 2 |
| 2023 | High-Precision Method and Architecture for Base-2 Softmax Function in DNN TrainingabstractSoftmax is a common and complex activation function in Deep Neural Networks (DNN). However, it is a challenge to apply it efficiently in DNN training hardware accelerator. Therefore, we propose a high precision calculation method and architecture based on base-2 softmax, which has low hardware complexity than base-$e$softmax but can still be useful in DNN training. First, we simplify the hardware implementation complexity of calculating base-2 softmax. Second, we use the base-2 hyperbolic COordinate Rotation Digital Computer (CORDIC) to implement the core computation. Finally, we show that the proposed method can be used in DNN training through experiments. Moreover, with the same order of the magnitude of high precision, our hardware cost is lower than traditional base-$e$softmax or other alternative design methods. Under TSMC 28nm CMOS technology, an example design of our architecture has the area of$98787.43\mu m^{2}$and the power consumption of 24.72mW for circuit synthesis at the frequency of 1GHz. Lele Peng, Lianghua Quan, Yonggang Zhang 0005, Shubin Zheng, Hui Chen 0015 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | Huicore: A Generalized Hardware Accelerator for Complicated FunctionsabstractEmerging advanced System-on-Chip (SoC) designs contain more and more complicated functions to be accelerated. This presents a challenge to conventional design approaches which use different hardware architectures or separate hardware accelerators to implement the various functions. To tackle this challenge, for the first time, we propose a generalized hardware accelerator called “Huicore” to speed up diverse functions on the same substrate. Through the analysis and transformation of mathematical characteristics, we reveal the commonality of many complicated functions using the CORDIC algorithm. Then we explore a reconfigurable architecture to implement them. The proposed reconfigurable accelerator can not only accelerate the implementation of many complicated functions, but also has small area, low power consumption and high precision. It is very suitable for integration in a SoC system to accelerate the implementation of various applications. Hui Chen 0015, Zongguang Yu, Zhonghai Lu, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | Low-Latency Low-Complexity Method and Architecture for Computing Arbitrary Nth Root of Complex NumbersabstractThis paper presents a new architecture, based on CORDIC and parabolic synthesis methodology, for computing Nth root of a complex number. The proposed architecture uses the pretreatment for normalization and parabolic synthesis method to calculate the Nth root of modulus of the input complex number and performs the conversion between the plane coordinate form and the polar coordinate form of the complex number by CORDIC, which not only ensures the accuracy but also has an ultra-low computation latency. MATLAB simulation result indicates that our proposed method can calculate the Nth root of the complex numbers in the form of fixed-point number with an error of$2.16 \boldsymbol {\times {10^{ - 6}}}$. Under TSMC 40nm CMOS technology, the report shows that the area consumption is$27390.72 \boldsymbol {\mu m^{2}}$at the frequency of 1GHz and the power consumption is 2.3549mW. More importantly, the computation latency of the proposed architecture is only 60.18% of the latest architecture in the same calculation accuracy. Hui Chen 0015, Guoqiang He, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Base-2 Softmax Function: Suitability for Training and Efficient Hardware ImplementationabstractThe softmax function is widely used in deep neural networks (DNNs), its hardware performance plays an important role in the training and inference of DNN accelerators. However, due to the complexity of the traditional softmax, the existing hardware architectures are resource-consuming or have low precision. In order to address the challenges, we study a base-2 softmax function in terms of its suitability for neural network training and efficient hardware implementation. Compared to the classical base-$e$softmax function, the base-2 softmax function is a new softmax function that uses 2 as the exponential base instead of$e$. From the aspects of mathematical derivation and software simulation, we first demonstrate the feasibility and good accuracy of the base-2 softmax function in the application of neural network training. Then, we use the symmetric-mapping lookup table (SM-LUT) method to design a low-complexity architecture but with high precision to implement it. Under TSMC 28nm CMOS technology, an example design of our architecture has the area of$5676 ~\mu m^{2}$and the power consumption of 13.12 mW for circuit synthesis at the frequency of 3 GHz. Compared with the latest works, our architecture achieves the best performance and efficiency. Yonggang Zhang 0005, Lele Peng, Lianghua Quan, Shubin Zheng, Zhonghai Lu, Hui Chen 0015 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2021 | A General Methodology and Architecture for Arbitrary Complex Number Nth Root ComputationabstractAs the existing complex number Nth root computation methods are relatively discrete, we propose a general method and architecture based on coordinate rotation digital computer (CORDIC) to compute arbitrary complex number Nth root for the first time. Our method performs the tasks of computing complex modulus, complex phase angle, real Nth root, sine function and cosine function, which can be implemented by circular CORDIC, linear CORDIC and hyperbolic CORDIC. Based on these CORDICs, our proposed architecture can not only improve the hardware efficiency just through shift-add operations, but also flexibly adjust the precision and the input range of complex number Nth root. To prove its feasibility, we conduct a software simulation and implement an example circuit in hardware. Under the TSMC 28nm CMOS technology, we synthesize it and get the report that it has the area of 6561μm2and the power of 3.95mW at the frequency of 1.5GHz. Hui Chen 0015, Zhonghai Lu, Li Li 0003, Zongguang Yu |
ISCAS | 1 |
| 2021 | Optimizing Vertical Link Placement and Congestion Aware Dynamic Elevator Assignment for Partially Connected 3D-NoCsabstractThe fully connected 3D-NoCs in which all routers are vertically connected with their neighbors above and below need a lot of Through-Silicon-Vias (TSVs), and they will occupy a large silicon area and reduce the fabrication yield. Thus, the idea of partially connected 3D-NoCs has emerged. The optimal number and placement of the vertical links (elevators) must be determined at the chip design stage, which is a multiobjective optimization problem of the performance and the cost. However, optimizing the static elevator placement needs a great amount of calculation and we can not examine all possible solutions at design time. Therefore, we propose a hybrid heuristic strategy for the static elevator placement and assignment, in which the genetic algorithm and the tabu search are combined. The dynamic assignment method is essential for the partially connected 3D-NoCs, and it leads to different traffic distributions and therefore has a huge impact on performance. Many previous static assignment methods can not dynamically change the elevator assignment according to the real-time states of the network, thus it may lead to network congestion. A congestion-aware dynamic assignment (CDA) scheme is proposed in this article, which considers the impact of the distance factor and the congestion factor on the network performance. Experiments show that the proposed CDA method can improve the network performance by 67%-86% compared with the random selection algorithm and can improve the reliability of the partially connected 3D-NoC as well. The key component for the CDA method, the path selection module (PSM), is implemented in FPGA, and the results show that its area cost is negligible compared with a router. Chuan Zhang 0001, Wenqing Song, Qinyu Chen, Hui Chen 0015, Li Li 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Symmetric-Mapping LUT-Based Method and Architecture for Computing XY-Like FunctionsabstractWe propose a new method and hardware architecture to compute the functions expressed as XY (X and Y are arbitrary floating-point numbers), which can support arbitrary Nth root, exponential and power operations. Because of the complexity of direct computation, we usually convert it to logarithm, multiplication, and antilogarithm operations. Traditional approaches suffer from long latency, large area and high power consumption. To solve this problem, we propose a symmetric-mapping lookup table (SM-LUT) to be capable of computing log2x (x ∈ [1, 2]) and 2x(x ∈ [0, 1]) simultaneously. It lays the foundation for computing XY. To further improve hardware performance of our architecture, we propose a multi-region address searcher to speed up the calculation of SM-LUT. In addition, we use an optimized Vedic multiplier to shorten the critical path and improve the efficiency of multiplication, which is included in computing XY. Under the TSMC 40nm CMOS technology, we design and synthesize a reference circuit to compute XY with a maximum relative error of 10-3. The report shows that the reference circuit achieves the area of 14338.50 μm2and the power consumption of 4.59 mW at the frequency of 1 GHz. In comparison with the state-of-the-art work under the same input range and similar precision, it saves 78.57% area and 80.42% power consumption for N√R computation and 82.89% area and 81.89% power consumption for RN computation averagely. On top of that, our architecture reduces the computation latency by 62.77% averagely and has one more order of magnitude of energy efficiency than others. Hui Chen 0015, Heping Yang, Wenqing Song, Zhonghai Lu, Li Li 0003, Zongguang Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Low-Complexity High-Precision Method and Architecture for Computing the Logarithm of Complex NumbersabstractThis paper proposes a low-complexity method and architecture to compute the logarithm of complex numbers based on coordinate rotation digital computer (CORDIC). Our method takes advantage of the vector mode of circular CORDIC and hyperbolic CORDIC, which only needs shift-add operations in its hardware implementation. Our architecture has lower design complexity and higher performance compared with conventional architectures. Through software simulation, we show that this method can achieve high precision for logarithm computation, reaching the relative error of 10-7. Finally, we design and implement an example circuit under TSMC 28nm CMOS technology. According to the synthesis report, our architecture has smaller area, lower power consumption, higher precision and wider operation range compared with the alternative architectures. Hui Chen 0015, Zongguang Yu, Yonggang Zhang 0005, Zhonghai Lu, Li Li 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | A CORDIC-Based Architecture with Adjustable Precision and Flexible Scalability to Implement Sigmoid and Tanh FunctionsabstractIn the artificial neural networks, tanh (hyperbolic tangent) and sigmoid functions are widely used as activation functions. Past methods to compute them may have shortcomings such as low precision or inflexible architecture that is difficult to expand, so we propose a CORDIC-based architecture to implement sigmoid and tanh functions, which has adjustable precision and flexible scalability. It just needs shift-add-or-subtract operations to compute high-accuracy results and is easy to expand the input range through scaling the negative iterations of CORDIC without changing the original architecture. We adopt the control variable method to explore the accuracy distribution through software simulation. A specific case (ARCH. (1, 15, 18), RMSE: 10−6) is designed and synthesized under the TSMC 40nm CMOS technology, the report shows that it has the area of 36512.78μm2and power of 12.35mW at the frequency of 1GHz. The maximum work frequency can reach 1.5GHz, which is better than the state-of-the-art methods. Hui Chen 0015, Yuanyong Luo, Zhonghai Lu, Li Li 0003, Zongguang Yu |
ISCAS | 1 |