EDBT 2026 Demo / reviewers in the wild / expert
Yuchun Chang 0001
dblp:79/10285-1 · also Yu-Chun Chang 0001
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0001-9270-8839ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FPGA-Friendly Architecture of Processing Elements for Efficient and Accurate Quantized CNNsabstractAn FPGA-friendly processing element based on the small logarithmic floating-point (SLFP) format is proposed. The proposed processing elements not only support inner product but also perform various nonlinear activation functions (NAF), which consume 674× LUT6s and 7× DSPs and operate at 450MHz in a pipeline manner for Zynq-7000. In addition, as the distribution of SLFP numbers is not uniform, this brief revises the weight decay scheme in the quantization aware training process to explore the optimum quantized weights. Compared with INT8 based design, the proposed method balances the resource usage between lookup tables and digital signal processing blocks. The accuracy loss of the quantized model based on the 8-bit SLFP is also small due to the high dynamic range of SLFP format. Moreover, since the proposed method can support different NAFs, this brief improves the quantized model accuracy by selecting an appropriate NAF from Swish, GELU, Mish and PReLU. Compared to the baseline (parameters are FP32, NAF is ReLU), the accuracy of quantized ResNet-50 and MobileNet is increased by 2.65% and -0.33%. Botao Xiong, Shize Zhang, Xingyu Shao, Xintong He, Yuchun Chang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | A Triaxial MEMS Capacitive Accelerometer With Efficient Temperature Hysteresis Compensation Based on a Modified Prandtl-Ishlinskii ModelabstractThis paper presents a triaxial accelerometer readout application-specific integrated circuit (ASIC) based on time-division multiplexing (TDM) and a digital on-chip compensation scheme based on a modified Prandtl–Ishlinskii (MPI) model. By adopting the TDM scheme, three sensing elements share the forward path of the analog signal chain, which significantly reduces the chip area. Furthermore, the MPI model is introduced for accelerometer temperature compensation for the first time. The proposed model can accurately compensate both the major and minor hysteresis loops under different temperature variation rates. This compensation approach overcomes the limitation of conventional polynomial methods, which are typically effective only under either static or dynamic temperature conditions. The readout ASIC is fabricated in a 1P6M 180 nm BCD process and powered by a 5 V single supply. Experimental results show that the designed accelerometer achieves a measurement range of ±30 g with a power consumption of 47 mW and a bandwidth of 57 Hz per axis. The sensitivities of the three axes are 0.1273 V/g, 0.1204 V/g, 0.1255 V/g, with corresponding nonlinearities of 0.129%, 0.072%, and 0.081%, respectively. After MPI compensation, a 36-dB bias temperature drift rejection ratio is achieved from$- 5~^{\circ }$C to$55~^{\circ }$C. Hongbin Lu, Zhaohan Li, Jicheng Yang, Zhikang Ma, Yuchun Chang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | Design of Low-Cost and High-Accurate 8-bit Logarithmic Floating-Point Arithmetic CircuitsabstractRecent studies suggest that the 8-bit floating-point (FP) format plays an important role in the deep learning, where the$E4M3$(4-bit exponent, 3-bit mantissa) is suited for the natural language processing model and the$E3M4$is better on computer vision task. In this brief, the logarithmic number system (LNS) is used to simplify the design of FP8 multipliers and dividers because the multiplication and division can be performed by the addition and subtraction in the logarithmic domain. Furthermore, this brief finds that the 3- and 4-bit logarithmic and anti-logarithmic (Antilog) converters can be effectively realized by {x,$x+1$} and {x,$x-1$}. As a result, compared to the standard$E4M3$and$E3M4$multipliers, the cell area can be reduced by 32% and 40%. Compared to the standard$E4M3$and$E3M4$divider, the cell area can be reduced by 61% and 67%. In addition, compared with the INT8-based design, the area of convolution core using proposed multiplier is reduced by 33%. The accuracy loss of the quantized ResNet-50, MobileNet, and ViT-B based on the proposed convolution core are −0.12%, +0.38%, and +0.8%, which are better than the INT8-based design. In the end, the proposed divider can be used in the image change detection. The false rate is slightly reduced from 2.97% to 2.95% compared to the standard$E3M4$divider. Botao Xiong, Xingyu Shao, Shize Zhang, Yuchun Chang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2022 | Half-Precision Logarithmic Arithmetic Unit Based on the Fused Logarithmic and Antilogarithmic ConverterabstractThis brief utilizes the logarithmic number system (LNS) to realize the half-precision division (DIV), square root (SR), and inverse SR (ISR) that are widely used in both the error resilience application and high-performance computing. With the aid of similarities of logarithmic and antilogarithmic functions, the adder tree and multiplexer in the shift-and-add architecture can be shared by the Log and Antilog converter. Moreover, a novel architecture for DIV and SR based on the fused converter is proposed. Compared to the existing works, this new architecture not only achieves a good tradeoff between precision level and hardware efficiency, but also can support more operations (e.g., exponential function, multiplication) with negligible hardware resources. In addition, the proposed architecture can be easily pipelined to further increase the throughput. Furthermore, for some formulas requiring multiple basic operations (e.g., ISR), the advantage of the LNS is more significant. Botao Xiong, Sicun Li, Sheng Fan, Yuchun Chang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |