Dongdong Chen 0002

dblp:92/1489-2 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-authorTheory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Processor architecture and microarchitecture · 57% Integrated circuit design · 32% Hardware accelerators and domain-specific architectures · 12%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Integrated circuit design › digital circuit design
arithmetic circuit design
0.822020
New Flexible Multiple-Precision Multiply-Accumulate Unit for Deep Neural Network Training and Inference · IEEE Trans. Computers 2020
Efficient Multiple-Precision Floating-Point Fused Multiply-Add with Mixed-Precision Support · IEEE Trans. Computers 2019
Processor architecture and microarchitecture › computer arithmetic › floating-point arithmetic
mixed-precision arithmetic
0.822020
New Flexible Multiple-Precision Multiply-Accumulate Unit for Deep Neural Network Training and Inference · IEEE Trans. Computers 2020
Efficient Multiple-Precision Floating-Point Fused Multiply-Add with Mixed-Precision Support · IEEE Trans. Computers 2019
Processor architecture and microarchitecture › arithmetic unit
multiply-accumulate unit
0.412020
New Flexible Multiple-Precision Multiply-Accumulate Unit for Deep Neural Network Training and Inference · IEEE Trans. Computers 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.412020
New Flexible Multiple-Precision Multiply-Accumulate Unit for Deep Neural Network Training and Inference · IEEE Trans. Computers 2020
Processor architecture and microarchitecture › computer arithmetic › floating-point arithmetic
fused multiply-add
0.412019
Efficient Multiple-Precision Floating-Point Fused Multiply-Add with Mixed-Precision Support · IEEE Trans. Computers 2019
Processor architecture and microarchitecture
computer arithmetic
0.112012
Improved Decimal Floating-Point Logarithmic Converter Based on Selection by Rounding · IEEE Trans. Computers 2012
Processor architecture and microarchitecture › computer arithmetic
decimal floating-point arithmetic
0.112012
Improved Decimal Floating-Point Logarithmic Converter Based on Selection by Rounding · IEEE Trans. Computers 2012
Integrated circuit design
digital circuit design
0.112012
Improved Decimal Floating-Point Logarithmic Converter Based on Selection by Rounding · IEEE Trans. Computers 2012
Processor architecture and microarchitecture › computer arithmetic
digit-recurrence algorithm
0.112012
Improved Decimal Floating-Point Logarithmic Converter Based on Selection by Rounding · IEEE Trans. Computers 2012
Integrated circuit design › digital arithmetic circuits › special function unit
logarithmic converter
0.112012
Improved Decimal Floating-Point Logarithmic Converter Based on Selection by Rounding · IEEE Trans. Computers 2012
Integrated circuit design
low-power circuit design
0.012012
Improved Decimal Floating-Point Logarithmic Converter Based on Selection by Rounding · IEEE Trans. Computers 2012

Methods — techniques the papers use, named apart from their topics

floating-point arithmetic · 0.4fixed-point arithmetic · 0.4carry-save compression · 0.4retiming · 0.1delay balancing · 0.1
YearPublicationVenuePosition
2020 New Flexible Multiple-Precision Multiply-Accumulate Unit for Deep Neural Network Training and Inference
abstract
In this paper, a new flexible multiple-precision multiply-accumulate (MAC) unit is proposed for deep neural network training and inference. The proposed MAC unit supports both fixed-point operations and floating-point operations. For floating-point format, the proposed unit supports one 16-bit MAC operation or sum of two 8-bit multiplications plus a 16-bit addend. To make the proposed MAC unit more versatile, the bit-width of exponent and mantissa can be flexibly exchanged. By setting the bit-width of exponent to zero, the proposed MAC unit also supports fixed-point operations. For fixed-point format, the proposed unit supports one 16-bit MAC or sum of two 8-bit multiplications plus a 16-bit addend. Moreover, the proposed unit can be further divided to support sum of four 4-bit multiplications plus a 16-bit addend. At the lowest precision, the proposed MAC unit supports accumulating of eight 1-bit logic AND operations to enable the support of binary neural networks. Compared to the standard 16-bit half-precision MAC unit, the proposed MAC unit provides more flexibility with only 21.8 percent area overhead. Compared to a standard 32-bit single-precision MAC unit, the proposed MAC unit requires much less hardware cost but still provides 8-bit exponent in the numerical format to maintain large dynamic range for deep learning computing.
Hao Zhang 0041, Dongdong Chen 0002, Seok-Bum Ko
IEEE Trans. Computers2
2019 Efficient Multiple-Precision Floating-Point Fused Multiply-Add with Mixed-Precision Support
abstract
In this paper, an efficient multiple-precision floating-point fused multiply-add (FMA) unit is proposed. The proposed FMA supports not only single-precision, double-precision, and quadruple-precision operations, as some previous works do, but also half-precision operations. The proposed FMA architecture can execute one quadruple-precision operation, or two parallel double-precision operations, or four parallel single-precision operations, or eight parallel half-precision operations every clock cycle. In addition to the support of normal FMA operations, the proposed FMA also supports mixed-precision FMA operations and mixed-precision dot-product operations. Specifically, the products of two lower precision multiplications can be accumulated to a higher precision addend. By setting the operands of one multiplication to zeros, the proposed FMA can also perform mixed-precision FMA operations. Support for mixed-precision FMA and mixed-precision dot-product is newly added but it only consumes 6.5 percent more area compared to a normal multiple-precision FMA unit. Compared to the state-of-the-art multiple-precision FMA design, the proposed FMA supports more floating-point operations such as half-precision FMA operations and mixed-precision operations with only 10.6 percent larger area.
Hao Zhang 0041, Dongdong Chen 0002, Seok-Bum Ko
IEEE Trans. Computers2
2012 Improved Decimal Floating-Point Logarithmic Converter Based on Selection by Rounding
abstract
This paper presents the algorithm and architecture of the decimal floating-point (DFP) logarithmic converter, based on the digit-recurrence algorithm with selection by rounding. The proposed approach can compute faithful DFP logarithm results for any one of the three DFP formats specified in the IEEE 754-2008 standard. In order to optimize the latency for the proposed design, we mainly integrate the following novel features: 1) using the redundant carry-save representation of the data path; 2) reducing the number of iterations by determining the number of initial iteration; and 3) retiming and balancing the delay of the proposed architecture. The proposed architecture is synthesized with STM 90-nm standard cell library and the results show that the critical path delay and the number of clock cycles of the proposed Decimal64 logarithmic converter are 1.55 ns (34.4 FO4) and 19, respectively, and the total hardware complexity is 43,572 NAND2 gates. The delay estimation results of the proposed architecture show that its latency is close to that of the binary radix-16 logarithmic converter, and that it has a significant decrease on latency compared with a recently published high performance CORDIC implementation.
Dongdong Chen 0002, Liu Han, Younhee Choi, Seok-Bum Ko
IEEE Trans. Computers1
2011 Nonspeculative decimal signed digit adder
abstract
Decimal floating point (DFP) arithmetic has been paid more attention in recent years, since it is superior to the binary counterpart in the financial and commercial computing including currency conversion, billing system, banking and tex calculation. Many DFP arithmetic units, such as addition, multiplication, division and fused-multi ply-add are not possible to achieve the high performance without a fast decimal fixed point adder. In this paper, the conventional four steps carry free signed digit addition algorithm is discussed. Furthermore, to improve the speed, we proposed a new method for the decimal SD addition and subtraction in digit set [-9,9]. To evaluate the design, a VHDL model is provided and synthesized in STM 90 nm technology. The result shows that our design has a better performance on timing delay and area compared with previous designs in the same digit set.
Liu Han, Dongdong Chen 0002, Khan A. Wahid, Seok-Bum Ko
ISCAS2
2010 A high performance pseudo-multi-core ECC processor over GF(2163)
abstract
In this paper, we propose a high performance processor for elliptic curve cryptography (ECC) over GF(2163) by using polynomial presentation. It has three finite field (FF) RISC cores and a main controller to achieve instruction-level parallelism (ILP) with pipeline so that the largely parallelized algorithm for elliptic curve point multiplication can be well suited on this platform. Instructions for combined FF operation are proposed to decrease clock cycles in the instruction set. The interconnection among three FF cores and the main controller is obtained by analyzing the data dependency in the parallelized algorithm. The whole design is implemented on Xilinx XC4VLX80 FPGA device, and it can reach 185 MHz with 20,807 slices. The total time required for one ECC point scalar operation is 7.7μs in 1428 cycles.
Dongdong Chen 0002, Younhee Choi, Seok-Bum Ko
ISCAS2
2009 A 32-bit Decimal Floating-Point Logarithmic Converter
abstract
This paper presents a new design and implementation of a 32-bit decimal floating-point (DFP) logarithmic converter based on the digit-recurrence algorithm. The converter can calculate accurate logarithms of 32-bit DFP numbers which are defined in the IEEE 754-2008 standard. Redundant digit e1is obtained by look-up table in the first iteration and the rest redundant digits ejare selected by rounding the scaled remainder during the succeeding iterations. The sequential architecture of the proposed 32-bit DFP logarithmic converter is implemented on Xilinx Virtex-II Pro P30 FPGA device and then synthesized with TMSC 0.18-um standard cell library. The implementation results indicate that the maximum frequency of the proposed architecture is 47.7 MHz in FPGA and 107.9 MHz in TMSC 0.18-um technology. The faithful 32-bit DFP logarithm results can be obtained in 18 cycles.
Dongdong Chen 0002, Younhee Choi, Moon Ho Lee, Seok-Bum Ko
IEEE Symposium on Computer Arithmetic1
2009 A New Decimal Antilogarithmic Converter
abstract
This paper presents a new design and implementation of a 32-bit decimal floating-point (DFP) antilogarithmic converter based on the digit-recurrence algorithm with selection by rounding. The converter can calculate the accurate antilogarithm (10dec) of the 32-bit DFP numbers which are defined in the IEEE 754-2008 standard. The sequential architecture of the proposed 32-bit DFP antilogarithmic converter is implemented on Xilinx Virtex-II Pro P30 FPGA device. The proposed architecture occupies 2, 315 out of 13696(16%) slices and can obtain a faithful 32-bit DFP antilogarithm in 11 clock cycles running at 51.5 MHz. The 7-digit decimal fixed-point (FXP) antilogarithmic converter is an essential operational part of the 32-bit DFP antilogarithmic converter. We transform it to a 7-digit decimal exponential converter to compare with a 24-bit binary FXP exponential converter. The compared results show that the 7-digit decimal exponential converter occupies 2.18 times more area and 1.66 times slower than the 24-bit binary FXP exponential converter.
Dongdong Chen 0002, Daniel Teng, Khan A. Wahid, Moon Ho Lee, Seok-Bum Ko
ISCAS1
2008 A novel decimal-to-decimal logarithmic converter
abstract
This paper presents a novel design and implementation of a 7-digit fixed-point decimal-to-decimal logarithmic converter. Two approaches, binary-based decimal approximation algorithm (Algorithm 1) and decimal linear approximation algorithm (Algorithm 2), are proposed and investigated. It shows that decimal linear approximation algorithm (Algorithm 2) is error-free in conversion between decimal and binary formats and also able to reduce maximum absolute error from binary-based Algorithm 1’s 0.00399 (integer cases) and 0.0483 (fraction cases) to 0.000994 (both cases). The Algorithm 2 is modeled in VHDL and implemented using combinational logic only in a Xilinx Virtex-II Pro P30 FPGA device. The logarithms results can be obtained in a single clock cycle, running at 50.9 MHz.
Dongdong Chen 0002, Younhee Choi, Daniel Teng, Khan A. Wahid, Seok-Bum Ko
ISCAS1